Many of the people building AI, and many of the people working on AI safety, share a common vision of what a good AI future looks like:
- We figure out alignment,
- build superintelligent AI,
- and it takes over the world, for our benefit.
Many people will tell you that last part outright: they think human disempowerment is a good thing because the AIs will be smarter and “more moral” than us. Others don’t outright cheer for disempowerment, but you can infer it from their influences, e.g. people who say they are inspired by Iain Banks’ Culture series of novels, where benevolent superintelligent machines run the world while humans just party and play video games. This idea of benevolent disempowerment goes back
to the origins of alignment as an idea.
In this worldview, alignment is the last and most important task for humans to work on. It is also a thought-terminating cliche, because it lets you avoid any of the hard political or economic or moral questions about the post-AI world. Any objection about the aligned AI utopia can be dismissed by saying “that’s not real alignment”.
You might ask: “won’t humans be powerless in a world with superintelligent machines?”, and the answer is “aligned AI would care about human agency, so that would be a failure of alignment, which we don’t want, so we really have to get alignment right!”. Similarly: “what happens to democracy when the state doesn’t need any human labour?” can be answered by “the AIs will be in control, and since they are aligned, nothing bad will happen”. Which is completely irrefutable.
Of course if someone said “to solve our political problems, we just need to find the right totalitarian dictator. The right dictator would select the right successor, so, by induction, this system will be perfect forever!”, you would laugh at them. But replace “dictator” with “aligned ASI”, and you have the ideology of tens of thousands of the most influential people in the world.
Rhetorically, “aligned ASI” is an opaque premise from which we can prove every desirable conclusion, and refute any undesirable conclusion. Every utopian dream is realized by definition, and any dystopian outcome is averted by definition. Any “gotchas” you try to find in the utopia can be refuted by “the AI will know you better than you know yourself, and will be smarter than you, so it will predict all the bad higher-order consequences of the utopia and fix them”. This should make us suspicious that the concept of an aligned superintelligence is incoherent and born of motivated reasoning.