This post is an extension of a collaboration with cousin_it on the question How risky would it be if powerful AI obeyed one or a few people?. There he argues for a common position: a future controlled by one or a few humans with powerful AI aligned to their intent is likely to produce terrible outcomes. My position is guardedly optimistic,[1]* for reasons I think are fairly novel: humans tend strongly to be better and become better over time under good circumstances, and near-perfect power and knowledge are the best circumstances. That post contains his essay and the abstract and overview sections of this post as my shorter response. This piece grew longer than our original target, because the subject is potentially critical for alignment strategy, and has not been analyzed in any depth, to my knowledge.*
Abstract:
Concentration of power over AGI/ASI seems quite possible. The first AGIs being aligned to intent (or instructions) over values seems fairly likely. So one or a few individuals or small groups gaining power over ASI seems fairly likely. [2] Thus it seems relevant to technical alignment strategy (value alignment vs. corrigibility) to worry about what individuals might do with such immense power. Here intuitions diverge, and careful analysis is scarce.
When we imagine one or a few people in charge of the whole future, it's intuitively very scary. We imagine a future serving the values of current and historically powerful people, which typically range between lacking and horrifying. But an ASI-empowered future will be unlike the past in important ways. And whatever humans wind up in charge will probably refine their beliefs and therefore their values over time. People don't typically lock themselves in against changing their minds later.
I argue that most people are basically good [3] (net prosocial) I also contrast this to the scenario in which we distribute power over strong AI [4] more broadly. Broad access to AI capable of creating better AI and novel weapons and tactics is unlikely to remain stable. This is a sharp contrast to historical balances of power. These have been driven by dependence on the governed, and sharply limited information and power for would-be oppressors (§5.2).
The development of AGI creates a potential for historically unmatched power concentration. This both makes questions about human nature pressingly relevant to AI safety, and limits the usefulness of classic arguments on the issue. I think this topic is relatively neglected; it's important since confusion on this topic may cause us to needlessly work at cross-purposes.
I dispute the common claim that the most powerful are the most cruel. I think the powerful have a modestly worse than average distribution of temperament, enough to worry about but not despair over. The powerful usually care for pets and children and attempt charitable works. They rarely torment individuals or treat them as "sims"; instead, they typically focus on broader accomplishments, and particularly in competing with their perceived rivals.
But the larger disagreement isn't about the starting temperament of the powerful; it's about how power changes them over time. I think the oft-quoted aphorism "power corrupts" is rarely examined, and happens to be quite wrong despite describing a strong correlation in history to date.
Instead, I think secure power probably purifies. To the extent I'm right, the average weakly prosocial person will become better over the time they hold truly secure power. I think this is likely despite the historical evidence that competition for power tends to corrupt, which has in the past made the average weakly prosocial person worse over time. I think this purifying effect will over time usually outweigh the selection and corrupting effects of competition for power (§2).
This thesis leads to a currently-unusual conclusion: maximal concentration of AGI/ASI power may be our safest route into an AI-dominated future. [5] Competition for power among multiple AGI-empowered individuals may intensify the historical dangers of power concentration (§5).
More broadly distributed powerful AI, among the majority of humans, is an intuitively appealing solution to risks from concentration of power, but it presents new and I think greater risks since it puts destabilizing AGI (capable of RSI, inventing new weapons, and/or takeover) into more hands, making it more likely that one of them will be vicious enough to deploy it in extremely destructive ways. Defending against every conceivable type of new attack seems unlikely in the limit. Thus, preventing destruction from broadly distributed AI would seem to require some sort of panopticon surveillance. This would create concentrated ultimate power, defeating the purpose of the distributed AI. Hoping to distribute AI powerful enough to counterbalance leading AIs but not powerful enough to take over if it's used for RSI or creating superweapons seems like a difficult target. AI is not like firearms that provide a small, fixed amount of power to each individual. It's more like a gun that can turn into a nuke (§5.1). Proposals for achieving such a balance between leading AI and distributed AI need much more detail; relying on intuition from history simply isn't adequate. (To be clear, I agree that broadly distributed near-term AI that's not capable of full RSI or easily creating superweapons, like next-gen open-source models, might well improve our odds of a good transition to AGI and ASI; that's a separate and equally neglected question.)
Thus, I think the fewer individuals who initially control AGI, the better off we are. Which specific individual(s) gain power matters a lot, but I think the majority of those currently in positions of power would produce very good (but not ideal) outcomes (§3).
The thesis, which I think is supported by the psychological literature, albeit indirectly, is roughly this: humans have many biologically determined instincts, but neurotypical humans are primarily ethically flexible. Humans' actions in the short term and their beliefs and "character" are largely shaped by their perceived circumstances. The second premise is that secure, near-absolute power is a very safe context, in sharp contrast to the limited, contested, and temporary (aging-limited) power achieved by any human in history thus far. The effect of such a unique position must be predicted from psychology, since nothing much like it has occurred yet. The safe context of secure, unlimited power should bring out the best in human nature, as defensive and competitive instincts become largely irrelevant. A supporting premise is that humans change over time much more than folk psychology suggests, so improving circumstances will not only improve behavior but will also improve character over time.
History suggests that power corrupts. But absolute, secure power is in many ways the inverse of the psychological situation produced by holding historical and studied levels of power.
I think most (but not all) people currently in positions of sufficient power are good enough to lead to good results in the long term. This is through the dynamic of continued growth. I think precommitting to a future path or ethics is unlikely if someone already holds secure power; it is giving up freedom. And I think basically-good people are likely to allow free speech and thought; they may shape culture, but directly controlling people's thinking seems pretty obviously evil. So I'd guess the scenarios range from fairly good (e.g., a future locked into traditional values of some sort, but with everyone happy) to more likely near-optimal (collective epistemic and moral growth indirectly reaches the tyrant, primarily through his servant ASI). The range of outcomes is worth considering in more depth; see §3.
A small cooperative group in control of one ASI has most of those advantages, and is probably safer due to reduced risks of exceptionally bad people getting full control. [6] And of course it's much better if that group is in turn directed by a democratic or other public-preference gathering system. I use the singular throughout for simplicity.
I don't want to overstate the case: I say secure power purifies to suggest that mostly-good people may become better, but if a truly horrible person (far in the tails of distributions on sadism and psychopathy/lack of empathy) gains absolute power, we'd have a truly horrible outcome ("s-risk") (§3 and §4, below). I currently estimate this as 1%-10% likely for the individuals most likely to achieve control over AGI, but as elsewhere, my uncertainty is large. This is, however, relatively well-calibrated uncertainty; I have been unable to find better evidence or arguments in any direction, since few have considered the contextual effects of truly unlimited power.
Humans placed in more extreme competition become worse. Increased power has historically almost always intensified competition for power, and it has never been secure. Every leader in history has had too little power to prevent their own death and suffering. Historically, power often raises the stakes of competition dramatically; most monarchs in history, and perhaps notably the worst of them, risked death and torture of themselves and any loved ones, if they lost their grips on power.
Absolute power with a loyal AGI/ASI is historically unprecedented both in how it lowers the stakes of any remaining competition by providing absolute security, and in how it reduces the epistemic distortions that have (perhaps heavily) contributed to the many historical abuses of power. Historically, advisors have been strongly motivated toward sycophancy, creating disastrous epistemic conditions. Accidental sycophancy from a highly competent ASI seems unlikely and nearly self-contradictory.[7]
The late serfdom period in imperial Russia is a fair example of the historical injustices that drive our starting intuitions about the brutality of human leaders. (This example is a result of previous iterations of this discussion, and my knowledge of the period is entirely based on asking Fable and Sol about it.) This and every other historical example I have encountered is of suffering inflicted primarily for pragmatic reasons. Practical incentives to inflict suffering would be entirely absent in a unipolar ASI-dominated future.
Freeing their own serfs would have impoverished the nobility. The nobles' treatment of their serfs appeared to be largely selfish, but rarely sadistic. Keeping harems of women as sexual objects was rare, although sexual exploitation probably was not. The motivation there was largely sexual, not sadistic or primarily dominance-motivated. Sadistic abuses of power occurred, but appeared to be rare. Making their lives better by reducing taxes and indentured labor would've cost relatively little in material comfort. But it would've cost. This is an example of values with little weight on others' well-being, not zero or negative weight. Absolute power makes generosity very cheap. The average person gives little in charity to strangers, but a count of exactly zero is rare.
The primary motivators of the massive suffering appeared to be simple desire for more material wealth; some of the worst abuses seem to have happened when a noble faced financial ruin and felt pressed to extract more from their serfs to protect themselves and their family.
This period was also multilateral and competitive, two distinctions between history and the ASI-empowered god-emperor I primarily focus on. The nobles under discussion lived under the power of the upper nobility, and competed with each other for status in a variety of arenas, including the wealth they extracted from serfs. This produced a situation ideal for producing motivated reasoning and group beliefs supporting the justice of holding power over serfs, although they seemed to devote less energy toward justifying the system than even American slaveholders with their justifications of paternalism.
It's unclear to what degree power-abusing individuals actually believe their own stated moral logic justifying their behavior. Based on my study of motivated reasoning, I'd guess the level is well above zero, and that group dynamics are historically crucial in creating that web of beliefs. However, this may not be such a distinction from the single sovereign situation, as any god-emperor can find their own circle of sycophantic humans, creating a similar effect - if they avoid ever asking their servant-god for the truth.
Many of the deaths under dictatorships have resulted from bad epistemics, and this might extend to a large majority. Famines from poor management outpaced malice by an order of magnitude or more (depending on assumptions about how many deaths were unwanted but seen as acceptable side effects). A regional party chief assured Khrushchev he could triple meat production, while actually slaughtering so extensively as to cripple production in future years. This "Ryazan miracle" is illustrative of how severely conflicting incentives and poor human predictions (the primary architect didn't benefit, killing himself in disgrace two years on) have contributed to historical injustices. Loyal ASI would essentially eliminate both factors. Mao's sparrows seem to have a similar cause: incompetence, not malice.
Historically, rulers' treatment of subjects has followed their need for loyalty. Rulers have needed their subjects' labor throughout history. An ASI-empowered autocrat would not need human labor nor fear rebellion. But neither would it cost them more than a hair of their own effort or material prosperity to make their subjects wealthier than kings. They need merely tell their servant-god to do so. And how many galaxies can one person enjoy alone?
Some of the worst abuses in history have been caused by greed (or more charitably, competition for resources). Leopold II's horrific mistreatment of natives in the Congo was based on the rubber trade. Native Americans were wiped out so that the US could take over their land.
An ASI-empowered ruler would not need to mistreat people to take their resources (unless we count not giving them equal shares in a galactic endowment as mistreatment). With superintelligent aid, they wouldn't be incompetent or ignorant. Such a ruler's treatment of their subjects will be wholly the result of their preferences. This is entirely unlike any situation in history, so historical leaders offer only tangential evidence.
An ultra-competent and ultra-informed dictatorship is still scary, of course, but I think the risks are in the tails and not the average.
My central point is that long-term outcomes are unlikely to be determined by the values a ruler holds when they assume control. They are more likely to be determined by decades or centuries of secure reflection, input from an honest and hypercompetent ASI and other humans. We might have to celebrate Samday weekly and Samfest yearly for a decade or a millennium, but most humans would rather spend eternity as a great hero than a great villain, if they're each equally easy. (And I expect Samfest to have good games and food. ;)
A broader claim is that intuition and historical analogy are not nearly good enough to steer the future in a good direction. This is the counterpoint to admitting that my own guesses are equally unreliable, since I've spent limited time on the topic, and have found almost no other attempts at close analysis of the question. With that said, on to my current predictions.
I predict (90-99%) good to very good outcomes of one powerful person gaining control over ASI and holding that power as long as they want. I include a substantial chance of near-optimal outcomes; that after reflection, many people will decide to do roughly what everyone would prefer (while holding aside substantial resources for their own pet projects). [8] In comparison, allowing a number of rivalrous humans in charge of distinct AGIs seems much more risky, since they may feel pressured to take drastic actions and hold beliefs that are motivated by their current context of competition and threat. The intuition that distributing AI power more broadly is better contains many assumptions about how those distributed AGIs will be useful defensively but not offensively, despite their ability to self-improve and invent new technologies. See §5 below.
The single-ruler scenario is probably better than wide proliferation of destabilizing AGI, but it is not optimal. There is a substantial tail risk of bad and very bad results, making this at most a best-of-bad-options future to shoot for.
A note on terminology and predictions: I've used "good" and "bad" in what I hope is a fairly intuitive and consensus sense: having high (or low) sum approval both by the people/sentients living in them, and by the lights of many of the best-considered people living now (people with strong non-rational moral commitments like religion and idiosyncratic philosophies might recoil in horror, but most of us would think they're at least pretty good, ranging to roughly as good as we could think of, or better).
Secure contemplation will probably, to a first approximation, "purify" that person. By this I mean that they will refine their beliefs and values toward more coherence. This is very good if the sum of their beliefs and values is good; it is very bad if they sum toward valuing outcomes others will dislike.
I think the vast majority of human beings place some value on human life and exhibit some empathy toward the states of others. Such empathy can be counteracted by sadism or dominance motivations, or by valuing other projects more so that material resources aren't devoted toward the wellbeing of humans/sentients.
But to me the averages look good, among people in good circumstances. People sometimes abuse children and pets, but most people like and care for them. And it looks to me like happy or fulfilled people never or almost never abuse those in their power. I'd expect an ASI-empowered ruler to be happy and fulfilled when it only takes asking "how could I become happy and fulfilled?"
The reflection that determines a ruler's long-term decisions will probably be done in conversation with an ASI advisor that's honest and extremely competent. It will also likely happen in conversation with other humans, and no small amount of hearing others' ideas of what the future could and should look like.
That reflection won't be comfortable. Our supreme leader is immune to physical threat but not to criticism: he will hear himself called a tyrant, a child, and a buffoon, and he will almost certainly ask his ASI whether his critics make good points. He can engage in motivated reasoning like anyone. But sycophancy from an ASI probably has to be requested. [7] "Soften your framings" is an explicit act he won't forget performing, unlike ordinary self-deception which is usually non-conscious and so not remembered. Deliberate self-blinding is possible, but it strikes me as an act of weakness unlikely from the sort of personality that would seize control.
Criticism, engaged from a position of security, with honest advisors, is roughly the recipe by which people grow.
I expect the stable end point of reflection for most people to be roughly libertarian utilitarianism, with some idiosyncratic weighting, because it's the rational conclusion of the motivations and value systems possessed by most humans. [9] Here I mean utilitarianism
My estimated odds of mediocre outcomes that are long-term stable are pretty low. This is an awfully weak means of reasoning, but it's a start: do you really imagine a universe full of suburbias? Corporate boardrooms? Terraformed worlds empty of humans, with a superyacht waiting for a galactic trazillionaire? Earth locked in conservative stasis and the universe left empty?
The odds of someone choosing such unimaginative futures, and never ever changing their mind to something more interesting or wisely chosen, seem pretty low to me. But again, I think these questions deserve a lot more thought.
What exactly are you envisioning if one person controls the whole future? This is a serious question: I think we desperately need more explicit models of both optimistic and pessimistic futures, so we're not gambling the future on intuitions formed from historical precedents that only partly apply. And is your model of an equilibrium, a long-term stable outcome, or a transitional period of merely years, decades, or centuries?
I'm primarily trying to solve for the equilibrium. I think transition periods could be rough, but they're likely a lot rougher if they include competition rather than clean power concentration.
And of course the outcome will depend on the individual. I think it will make sense to put a lot of effort into the 2028 US presidential election, for a nonrandom example.
I don't want to downplay the risks. Some of the main ones I see:
Many who think about the risks of transformative AI or AGI prefer an outcome where that power is distributed broadly. In the current day, I think broadly distributing power by broad distribution of open-weight models probably makes the world safer. In the next generation, with AGI capable of creating new weapons and recursively creating smarter AI, the logic changes dramatically.
I and others fear AGI proliferation more than concentration of power. This crux seems worth resolving, lest similarly well-intentioned, AI-risk-concerned people work at cross-purposes. My conclusion is based on the counterintuitive factors in secure absolute power discussed above, and on roughly inverse effects in the multipolar scenario. I discuss this in If we solve alignment, do we die anyway? and Fear of centralized power vs. fear of misaligned AGI.
In short, I challenge those who are optimistic about such scenarios to develop them further. Here I look at some fairly obvious difficulties which are nonetheless rarely addressed. At the end I try to steelman the arguments for optimism a little. Both efforts fall far short of the elaborate scenario-construction we'd need to get real traction on this question. History suggests that distributed power works well, but no type of historical power allowed rapid creation of new forms of power. Superhuman AI does exactly that.
Let's try to envision a world in which power is distributed broadly, so that most people have access to powerful AI that will follow their instructions. Let's consider a scenario in which many people have access to near-frontier AGI, since that seems more likely than everyone having access to equally powerful AGI.[10]
First let's consider the downside. If many people have access to AGI that can create new weapons (bioweapons, basement nukes, assassin drones, horrible new weapons we haven't thought of) and new AI, and if such weapons and tactics can be developed, we seem to be stuck hoping that defense is dominant over offense in all of the many different arenas of attack.
Unlike previous technologies, powerful AI seems unlikely to create a new multipolar equilibrium, because there is no stable game state when the rules and players' capabilities keep rapidly changing. Powerful AI can be used to relatively quickly create yet more powerful AI, as well as novel weapons and tactics. All of these are destabilizing factors, with a very bad game-theoretic conclusion (at least on my initial inspection): defectors win, and the first mover may have a strong advantage, leading to survival of the most vicious.
Preventing the most vicious from developing new weapons in the face of ongoing progress in AI and technology seems to require more than defensive acceleration. That might work for a while, but it's hoping that every type of weapon has a defense that can be developed in advance and with realistic levels of effort. That seems unlikely on first principles. It seems increasingly unrealistic as technology advances. Triggering existing nukes through software intrusion and social engineering is a mere starting point; developing new nukes up to crust-busters, creating asteroid strikes, delivering viruses tailored to individuals or groups, rods-from-god decapitation strikes, micro-assassin drones, and taking off and nuking the whole thing from orbit (leaving the solar system and sending the sun nova) are just off the top of my head. And I'm no military technologist let alone an ASI.
Thus, I'd think long-term safe distribution of AGI power would require preventing the development of superweapons and super-AI. Manufacturing and compute will both become more efficient, allowing smaller physical sites to do more. Manufacturing and compute locations will also diversify in location (e.g., manufacturing distributed with better printers/assemblers, sites underwater, underground, and in orbital or distant space). Increased diversity and smaller size of physical sites will require more fine-grained monitoring of what's going on in each piece of compute and manufacturing: a panopticon.
I'm afraid such a panopticon is increasingly realistic. Publicly available information can already be used to infer intent, if we have enough processing power to aggregate it and analyze it carefully. And more advanced AI and sensing technology will rapidly expand this. Totalitarian governments with merely current AI and technology, like ubiquitous license-plate readers and other AI-monitored cameras, are becoming more proficient at suppressing dissent by targeting individuals. Extrapolating this trend, we might ask what access to powerful but not frontier AI might do to protect us from the power of larger entities (states most likely but corporations in some visions of the future) armed with yet-better AI and more physical force.
We might hope that the dynamic of mutually assured destruction continues to provide a stabilizing influence into a multipolar ASI-empowered future. I think this might hold in a useful way through a transitional period, but is unlikely to remain effective long into a period in which AGI can create new weapons and tactics. The possibility of inventing wildly new technologies, and spreading power and populations into space and then distant stars makes this scenario seem less stable. And even if we can pursue every colony with the threat of destruction if their faction defects, this solution does not seem stable indefinitely; if accidents or defections are possible, they will happen eventually.
Powerful AI may allow old or new means for a distributed power model to remain stable over the longer term. I hope so, but I find existing proposals highly lacking; they usually seem based on intuition and do not come to grips with the historical discontinuities created by powerful AI. Existing mechanisms of sharing power, including mutually assured destruction, seem to largely break down in the face of new technologies and continuing rapid, unpredictable technological progress.
But I don't want to dismiss the possibility that we'll find solutions for those problems, perhaps enabled by newer, more powerful AI. It will bring advantages for cooperation, and we can hope those outweigh the problems. And it's possible that distributed near-future AI (stronger than today's open models but not existentially dangerous) would make the transition to an ASI-dominated world safer in subtle ways I haven't foreseen.
Would you want to be powerless while giants fight? I would not, but I don't think I have a choice, because AI power scales increasingly nonlinearly once we hit appreciable self-improvement. But it does seem worth looking for holes and edge cases for this argument.
One argument that distributed strong AI won't create disaster relies on one or a few actors having stronger AI that can prevent large-scale threats (like creating superweapons or recursively self-improving). To a first approximation, this seems to also negate the advantage of having widely distributed power stemming from weaker AIs. The lead player(s) can probably use that defensive power offensively at will, easily rolling over any opposition. But perhaps this is wrong, and the intuition that some power is better than none is correct. This would be analogous to the argument for broad access to firearms making it less likely that governments become tyrannical. Civilians and ad-hoc militias can't oppose the full force of state armies, but they can make it more costly for states to oppress their citizens. And such distributed power encourages noble sacrifices that can inspire further resistance.
Perhaps this analogy holds up to the era of superhuman AI, at only modest risk of existential threats created by some of the many possible actors. I think the situation is not really analogous; the advantages of ASI include information and spin manipulation, so that noble sacrifices are unlikely to spark greater resistance; and greater resistance would indeed be futile. ASI-enabled power (and much AGI-enabled power) is unlikely to route through the loyalty of humans. The question isn't whether an ASI-enabled tyrant can suppress revolt and survive, it's how forceful he needs to be to do so.
In sum, I'd want much stronger and clearer arguments. The existing arguments I've seen are based on intuitions. I don't think those intuitions survive the disanalogy between current power distributions and those in the face of strong AGI or ASI, while the risks of widely-distributed AI that's potentially capable of weak RSI or strong new weapons development seem pretty clear and robust to careful analysis.
It's possible that the global panopticon necessary to prevent development of new weapons and new ASI might be run in a distributed fashion, with trust distributed in some manner based on encryption, and with dedicated, publicly verifiable AIs that can answer security questions while keeping private the information that would allow abuse of power.
Smarter AI will enable better communication and may produce new strategies for cooperation. It will at least reduce competition from incompetence. A large part of my premise is that we should not assume malice when confusion and motivated reasoning are human universals. As such, I do think that AI for epistemics will provide substantial advantages, as outlined in my Human-like metacognitive skills will reduce LLM slop, AI 2040, and elsewhere.
Section 5 has diverged from my main focus, on the nature of psychology in the face of power. This seems necessary but I've kept it brief; therefore the analysis is incomplete and I haven't tracked down references to the relevant existing work.
I am uncertain of these conclusions about the results of making powerful AI broadly accessible; my only strong claim is that these issues deserve more careful attention before we default into one of these paths.
The question I'd pose, refined: how, exactly, will distributing power produce stability rather than an increased power struggle favoring the most vicious individuals or the most repressive states?
Putting the future in the hands of one individual or even a small group is a risky proposition. But I think it's less risky than handing power to as many people as possible. Power cancels out and defense dominates in some domains. I doubt it does with RSI-capable AGI, but of course I'm unsure on that as well as everything else. We've barely started on this set of topics.
What a particular individual wants to do after centuries of absolute power seems pretty hard to determine. [11] Basically nice people should be on some sort of nice trajectory resulting in nice things. People with a mix of nice and mean motives might become nicer or meaner, but an easy and happy life would seem to dispose them toward the nice direction of evolution. And failure to be happy with unlimited power seems unlikely; historical rulers got bored and cranky, but they didn't have unlimited power over their experiences and their own minds. If they allow freedom of thought and expression, the resultant civilizational trajectory seems likely to bend toward near-ideal (by many people's preferences) outcomes.
But this is all far too uncertain to bet the future on. Inherent uncertainty is probably high, even with our best efforts to select trustworthy leaders as we approach ASI.
This set of topics seems worth a lot more analysis than it's gotten to date. The future is very hard to predict, but the payoff of even limited predictive success seems large. So little effort has gone in this direction that there may still be obvious-in-retrospect low-hanging fruit from even a little additional thought in this direction.
Acknowledgements:
Thanks to cousin_it for extensive comments in the form of iterated drafts, and for generating this post as the result of the "adversarial" collaboration How risky would it be if powerful AI obeyed one or a few people?
Thanks to Peter Gebauer and Richard Juggins for useful comments on an earlier draft.
I'd be undecided on the dangers of proliferation vs. power concentration if egregious misalignment wasn't a concern. It is by any reasonable estimate a nontrivial concern, and becomes a larger one with more parties racing from human-plus AGI to takeover-capable levels of intelligence. I currently favor accepting the risks of power concentration over allowing advanced AI to proliferate, in part because that creates more individual opportunities create egregiously misaligned ASI. However, this is a compromise to practicality. Slowdown or would be better if we can get it.
Due to the dynamics of exponential racing, I find it fairly unlikely that we'll have even a semi-stable situation with a few actors controlling different near-peer AGIs; I'd expect those in the lead to sabotage other projects, under the expert and pragmatic advice from their own AGI. That's one reason I focus primarily on the single individual scenario; I think it's more likely in the long run and probably even in the medium-run. The only near-peer scenario I find fairly plausible is the US and China progressing in parallel, based essentially on nuclear deterrence against sabotage efforts.
I use "basically good" to mean someone who has more prosocial (wishing good for others) than antisocial or sadistic motivation (wishing ill). I think that the vast majority of humans are in this category, even most people categorized as sociopathic/psychopathic. Power or dominance motivations, and a variety of others, are somewhat orthogonal to this primary "goodness" axis, and have important, complex effects on outcomes.
Here I'm addressing only future strong AI, not current or near-future open source models, even if they're dangerous without being existentially risky. The arguments here apply to AI capable of existentially threatening humanity, particularly by takeover, creating superweapons, or rapidly creating new AI capable of those threats. The arguments don't apply to models that are dangerous in mundane ways like cyber attacks and even uplift on engineered bioweapons. I'd prefer broad distribution of power right up to the point of existential threat if that were possible.
I do not mean that concentrating AGI/ASI power is safe. While I think power concentration is safer than proliferation, the safer path is to not build AGI until we have better plans and understanding. Unfortunately, that's looking unlikely, so we're stuck taking large risks. This argument is also dependent on the argument that wide access to transformative AI creates something like an n-way non-iterated prisoner's dilemma, in which the first person to use new weapons and tactics to seize absolute power wins. This premise is also counterintuitive. I claim the situation is distinct from historical distributions of power. I lay out a brief form of this argument in If we solve alignment, do we die anyway? and Michael Nielsen makes similar points in his excellent ASI existential risk: Reconsidering Alignment as a Goal.
A small group of reasonably cooperative people controlling an ASI has many of the same advantages and risks, but one large advantage over the single-person case I focus on. If an ASI were reliably aligned so that those individuals couldn't benefit from power struggles, roughly averaging those people's desires would eliminate most of the risk of getting truly horrible values in charge of the future.
Current AI is highly sycophantic and near-future systems will be too. It's a problem inherent in how we train AI to be useful. But I don't expect this to remain a large problem up to takeover-capable AI. There are routes to improving sycophancy/hallucination on the current path by improving Human-like metacognitive skills. More broadly, being capable of superhuman strategy and invention requires being able to sort truth from imagination quite effectively. It seems to me that ASI that's sycophantic without knowing it would be a weak sort of ASI, and improvements would likely rectify that. Unconscious sycophancy seems to depend on limited self-knowledge.
These scenarios in which "only" a large portion of future resources are put toward consensus goods might be a moral tragedy compared to best outcomes, but it's also an enormous moral victory relative to failure scenarios.
This is not a claim of moral realism, but a claim about innate human drives, and logic. I say roughly libertarian utilitarianism because I expect people to heavily favor their own wellbeing, but also on net value the freedom and wellbeing of other sentient beings as well. This leaves the distribution of resources in question, and of course this is a claim about typical humans, not every instance. There genuinely are people who would prefer suffering for others, and logical routes to reflective stability around those values.
I think all the dynamics of widely-distributed human-controlled AGIs become worse if more of those systems are peers. I also think this is unlikely to persist; the government shutdown of Fable indicates that broad distribution of truly frontier AGI, e.g. capable of superhuman progress on AI and weapons, is highly unlikely to be achievable. Governments are waking up rapidly to the security risks of advanced AI and its rapid progress. The open question is whether distributed lesser AGIs under control from many humans can be an effective counterbalance to ASI power while being restricted from creating more ASIs or doomsday weapons.
I think how someone with absolute power evolves over centuries is genuinely hard to predict, but I acknowledge that there are factors weighing toward loose lock-in in the short term. Leaders tend to acquire sycophants, and those will be only partly counterbalanced by an objective ASI. Humans are typically prone to double down on bad decisions, although we could hope this bias might also be counteracted by wiser advice. People also tend to dissociate from those who raise emotionally difficult questions, and this tendency will also lead them to tell their ASI explicitly to be sycophantic.