It seems fairly likely that the first powerful AIs will be instruction-following rather than value-aligned, and will be controlled by a small number of people. So it makes sense to worry what individual people might do with such immense power. Here intuitions diverge and careful analysis is scarce. This post presents a debate between Seth Herd and cousin_it over how risky such a scenario would be.
The debate ran under an unusual protocol. First we wrote our initial draft statements and sent them to each other in private. Then we each revised our statements to strengthen them against the other's, and sent them to each other again. We continued this for about 10 rounds over the course of about a month, until we both agreed to stop revising and publish (while still remaining in disagreement).
Here's the final pair of statements we ended up with, so you can judge for yourself:
If there is an AI-assisted overlord (or several) and everyone else is their completely powerless subjects, that situation will be historically new, but not 100% new. Large power imbalances have existed in the past too and we can learn from them. Usually, when power was more absolute and less accountable, the subjects had it worse. We can even compare the same ruler's treatment of different subjects: like King Leopold II, who was good to Belgians, but horrible to the Congolese at the same time. (One imagines the department head being more abusive toward the junior clerk than toward the senior clerk.) It's clear that the abusive treatment depends mostly on the amount of power difference, not on other details of the ruler's situation.
Maybe Leopold II isn't a good analogy: an AI-assisted overlord wouldn't be under economic pressure to exploit us like the Congolese. More like, he wouldn't really need us or our labor for anything, like the early US didn't need the Native Americans. Whoops, this example doesn't look good for us either! Maybe we need to imagine an even larger power difference: an overlord who's so rich with territory and resources that he can easily spare some for his favorite creatures. But then who says we, with all our imperfections, will be his favorite creatures? He can create new ones instead, or select some of us and discard the rest, and we'll have no recourse.
Let's say even we get lucky, and the overlord decides to be a do-gooder toward all of us. If his views are colored by ideology or religion, then he'll be free to impose them on us. He could try instituting conversion therapy for gay people, or creating a New Soviet Man, or whatever else he decides is a good idea.
Maybe we could hope that the AI itself, by virtue of being a wise adviser, would stop the overlord from doing bad things? But the problem is that the overlord won't accept such an AI to begin with. Rulers today already don't want AI that will second-guess them: we just saw the US government demanding that Anthropic's AI not restrict them in any way. Rulers have always wanted yes-men, and now they want a yes-man AI.
Which leads to yet another problem: a sycophantic yes-man AI will make the overlord free to spiral off into their own world, as has happened with some autocratic rulers in the past. The overlord's views and sense of morality might change over time, probably toward self-aggrandizement, thinking of other people as less important or less real.
And all of these problems are just with one overlord. What if multiple overlords compete with each other, economically or militarily? Since helping regular people makes an overlord less effective at competing, most likely the winners will be those overlords who care about regular people the least.
The only argument for restricting AI to a handful of overlords is the argument from lesser evil: that spreading AI out to many people would be even worse. But I don't agree with that argument.
Historically, spreading out power to those affected by it has usually been a good thing. India under British colonial rule had regular famines killing many millions, then with independence these famines instantly stopped and never happened again. So in this case, spreading out power worked out well.
What is specific to AI power that makes spreading it out a bad idea? The usual story is that AI would allow the small guy to threaten the whole world. But "being able to threaten the whole world" is a moving target, because technology advances for the world too. A virus created in a basement can be cured by someone else's AI in their own basement; an assassination drone can be identified and shot down by a police drone; a cyberattack launched from a basement computer can be stopped by the AI of the NSA. Bigger weapons, like nukes or asteroid strikes and so on, will be even more detectable and preventable by the big guy with the big AI.
Most likely, apocalypse or even large-scale terrorism will remain out of reach for the small guy. The use case for small AI will be small-scale resistance to power and some plain old self-reliance. These are good things and we should try to keep them.
At that, I'll rest my case. I should've started by saying that I'd prefer to not build powerful AI at all, or to build AI that acts according to human morality instead of obeying a specific person. But on the terms of the argument, if the choice is between restricting AI to a few overlords vs. having many AIs owned by many people, to me the latter has a much better chance of a good future.
When we imagine one or a few people in charge of the whole future, it's intuitively very scary. We imagine a future serving the values of current and historically powerful people, which typically range between lacking and horrifying. But an ASI-empowered future will be unlike the past in important ways. And whatever humans wind up in charge will probably refine their beliefs and therefore their values over time.
I argue that most people are basically good [1] (net prosocial) I also contrast this to the scenario in which we distribute power over strong AI [3] more broadly. Broad access to AI capable of creating better AI and novel weapons and tactics is unlikely to remain stable. This is a sharp contrast to historical balances of power. These have been driven by dependence on the governed, and sharply limited information and power for would-be oppressors (
The development of AGI creates a potential for historically unmatched power concentration. This both makes questions about human nature pressingly relevant to AI safety, and limits the usefulness of classic arguments on the issue. I think this topic is relatively neglected; it's important since confusion on this topic may cause us to needlessly work at cross-purposes.
I dispute the common claim that the most powerful are the most cruel. I think the powerful probably have powerful a modestly worse than average distribution of temperament, enough to worry about but not despair over. The powerful usually care for pets and children and attempt charitable works. They rarely torment individuals or treat them as "sims"; instead, they typically focus on broader accomplishments, and particularly in competing with their perceived rivals.
But the larger disagreement isn't about the starting temperament of the powerful; it's about how power changes them over time. I think the oft-quoted aphorism "power corrupts" is rarely examined, and happens to be quite wrong despite describing a strong correlation in history to date.
Instead, I think secure power probably purifies. To the extent I'm right, the average weakly prosocial person will become better over the time they hold truly secure power. I think this is likely despite the historical evidence that competition for power tends to corrupt, which has in the past made the average weakly prosocial person worse over time. I think this purifying effect will over time usually outweigh the selection and corrupting effects of competition for power (§2).
This thesis leads to a currently-unusual conclusion: maximal concentration of AGI/ASI power may be our safest route into an AI-dominated future. [4] Competition for power among multiple AGI-empowered individuals may intensify the historical dangers of power concentration (
More broadly distributed powerful AI, among the majority of humans, is an intuitively appealing solution to risks from concentration of power, but it presents new and I think greater risks since it puts destabilizing AGI (capable of RSI, inventing new weapons, and/or takeover) into more hands, making it more likely that one of them will be vicious enough to deploy it in extremely destructive ways. Defending against every conceivable type of new attack seems unlikely in the limit. Thus, preventing destruction from broadly distributed AI would seem to require some sort of panopticon surveillance. This would create concentrated ultimate power, defeating the purpose of distributing AI in the first place.
Hoping to distribute AI powerful enough to counterbalance leading AIs but not powerful enough to take over if it's used for RSI or creating superweapons seems like a difficult target. AI is not like firearms that provide a small, fixed amount of power to each individual. It's more like a gun that can turn into a nuke (§5.1). Proposals for achieving such a balance between leading AI and distributed AI need much more detail; relying on intuition from history simply isn't adequate.
(To be clear, I agree that broadly distributed near-term AI that's not capable of full RSI or easily creating superweapons, like next-gen open-source models, might well improve our odds of a good transition to AGI and ASI; that's a separate question.)
Thus, I think the fewer individuals who initially control AGI, the better off we are. Which specific individual(s) gain power matters a lot, but I think the majority of those currently in positions of power would produce very good but not ideal outcomes (§3).
The thesis, which I think is supported by the psychological literature, albeit indirectly, is roughly this: humans have many biologically determined instincts, but neurotypical humans are primarily ethically flexible. Humans' actions in the short term and their beliefs and "character" are largely shaped by their perceived circumstances. The second premise is that secure, near-absolute power is a very safe context, in sharp contrast to the limited, contested, and temporary (aging-limited) power achieved by any human in history thus far. The effect of such a unique position must be predicted from psychology, since nothing much like it has occurred yet. The safe context of secure, unlimited power should bring out the best in human nature, as defensive and competitive instincts become largely irrelevant. A supporting premise is that humans change over time much more than folk psychology suggests, so improving circumstances will not only improve behavior but will also improve character over time.
History suggests that power corrupts. But absolute, secure power is in many ways the inverse of the psychological situation produced by holding historical and studied levels of power.
I think most (but not all) people currently in positions of sufficient power are good enough to lead to good results in the long term. This is through the dynamic of continued growth. I think precommitting to a future path or ethics is unlikely if someone already holds secure power; it is giving up freedom. And I think basically-good people are likely to allow free speech and thought; they may shape culture, but directly controlling people's thinking seems pretty obviously evil. So I'd guess the scenarios range from fairly good (e.g., a future locked into traditional values of some sort, but with everyone happy) to more likely near-optimal (collective epistemic and moral growth indirectly reaches the tyrant, primarily through his servant ASI). The range of outcomes is worth considering in more depth; see §3.
A small cooperative group in control of one ASI has most of those advantages, and is probably safer due to reduced risks of exceptionally bad people getting full control. [5] And of course it's much better if that group is in turn directed by a democratic or other public-preference gathering system. I use the singular throughout for simplicity.
I don't want to overstate the case: I say secure power purifies to suggest that mostly-good people may become better, but if a truly horrible person (far in the tails of distributions on sadism and psychopathy/lack of empathy) gains absolute power, we'd have a truly horrible outcome ("s-risk") (§3 and §4 in the longer version linked below). I currently estimate this as 1%-10% likely for the individuals most likely to achieve control over AGI, but as elsewhere, my uncertainty is large. This is, however, relatively well-calibrated uncertainty; I have been unable to find better evidence or arguments in any direction, since few have considered the contextual effects of truly unlimited power.
I think the subject deserves much more analysis*. * The above stands alone as an overview, but it is also the abstract and overview from what became a longer post. The full post is here: Extreme concentration of power over ASI has non-obvious advantages. Section headings here refer to sections from that full post.
I use "basically good" to mean someone who has more prosocial (wishing good for others) than antisocial or sadistic motivation (wishing ill). I think that the vast majority of humans are in this category, even most people categorized as sociopathic/psychopathic. Power or dominance motivations, and a variety of others, are somewhat orthogonal to this primary "goodness" axis, and have important, complex effects on outcomes.
I'd be undecided on the dangers of proliferation vs. power concentration if egregious misalignment wasn't a concern. It is by any reasonable estimate a nontrivial concern, and becomes a larger one with more parties racing from human-plus AGI to takeover-capable levels of intelligence. I currently favor accepting the risks of power concentration over allowing advanced AI to proliferate, in part because that creates more individual opportunities create egregiously misaligned ASI. However, this is a compromise to practicality. Slowdown or would be better if we can get it.
Here I'm addressing only future strong AI, not current or near-future open source models, even if they're dangerous without being existentially risky. The arguments here apply to AI capable of existentially threatening humanity, particularly by takeover, creating superweapons, or rapidly creating new AI capable of those threats. The arguments don't apply to models that are dangerous in mundane ways like cyber attacks and even uplift on engineered bioweapons. I'd prefer broad distribution of power right up to the point of existential threat if that were possible.
I do not mean that concentrating AGI/ASI power is safe. While I think power concentration is safer than proliferation, the safer path is to not build AGI until we have better plans and understanding. Unfortunately, that's looking unlikely, so we're stuck taking large risks. This argument is also dependent on the argument that wide access to transformative AI creates something like an n-way non-iterated prisoner's dilemma, in which the first person to use new weapons and tactics to seize absolute power wins. This premise is also counterintuitive. I claim the situation is distinct from historical distributions of power. I lay out a brief form of this argument in If we solve alignment, do we die anyway? and Michael Nielsen makes similar points in his excellent ASI existential risk: Reconsidering Alignment as a Goal.
A small group of reasonably cooperative people controlling an ASI has many of the same advantages and risks, but one large advantage over the single-person case I focus on. If an ASI were reliably aligned so that those individuals couldn't benefit from power struggles, roughly averaging those people's desires would eliminate most of the risk of getting truly horrible values in charge of the future.