Why alignment is upstream of everything

3 min read
Share:
Michael Darius Eastwood
Michael Darius Eastwood · Independent AI alignment researcher

the long version

This post is part of the long version series: the full text behind a passage the front pages now say in fewer words. Nothing was cut from the record; it moved.

Every other great problem leaves us holding the wheel. Climate, pandemics, weapons: they can be catastrophic, and we can be slow and stupid about them, and it is still people who decide what happens next. This is the first one that hands the deciding to something else. That is what makes it the largest problem we have, rather than the danger by itself. It sits upstream of all the others, because it decides whether we keep the capacity to solve any of them.

The world’s method for getting it right is to build the system and then check it. Benchmark it, red-team it, interpret it, watch it. I think that method has a hole in it, and it is not a funding problem: you cannot find out what a mind does when nobody is watching by watching it. The check stops working at exactly the point where the stakes arrive. So safety cannot be a test taken at the end. It has to be a property of how the thing was built.

I took the question up not because I am the obvious person to do it, but because it was sitting there unattended, and the people best placed to answer it have structural reasons to look somewhere else.

This is the argument behind the opening of the home page. The full statement of the unsolved problem, with its literature, is at the open problem.

reads aloud · highlights as it goes · jump to any section