# Did the alignment community underestimate its power?

> Source: <https://www.lesswrong.com/posts/cPbsnMAGZjApeCghE/did-the-alignment-community-underestimate-its-power>
> Published: 2026-08-12 02:56:33+00:00

Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are:

- Trying to convince the US government to take AGI much more seriously.
- Doing “alignment research” which is very similar to capabilities-maximizing research (especially building automated alignment researchers).
- Trusting Anthropic too much (in an analogous way to how we trusted OpenAI too much).
- Trading off clarity in thinking about politics for conformity (in a similar way to how we traded off clarity in thinking about AGI for conformity to the ML ontology).
These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”...

Richard Ngo's post "[What just happened? A retrospective of AI alignment](https://www.lesswrong.com/posts/9RL9MuGZjzm4q3gKG/what-just-happened-a-retrospective-of-ai-alignment)" is an attempt to explain that a significant part [1] of the alignment community made potentially fatal strategic errors which, however, can be fixed, and the mistakes' potential origin. The biggest mistake, according to Ngo, is the inability to recognize the fact that scientific progress proceeds by

According to Ngo, one of the reasons why the alignment ontology failed to win recruits over the ML ontology was the following:

The field of ML was also reacting in part to the alignment community’s lack of appropriate discrimination. In particular, rationalists often treated informal, abstract arguments about AI risk as

[far more decisive than was warranted], in part due to an[epistemology which claimed to supersede]standard[scientific epistemology](and in part due to a[strong emotional orientation]towards “[saving the world]”). Rather thanfocusing on further developing and clarifying its insights about AGI risk, though, the rationalist community spent significant effort winning its skeptics over, with largely regrettable effects.

The issue with novel ontologies is that they find it hard to replace the old ones unless the world produces empirically testable results which can be used against the old ontology. For example, the Newton-like world where space and time are fully separate dimensions and the Einstein-like world where spacetime is transformed so as to leave the Minkowski-like metric invariant have been first separated by testing whether the [Michelson–Morley experiment](https://en.wikipedia.org/wiki/Michelson%E2%80%93Morley_experiment) notices Earth's motion related to the aether. As we know now, the experiment produced evidence in favor of Einstein's worldview instead of Newton's worldview.

The alignment-vs-ML ontology clash is unusual because theoretical arguments were laid *before* the experiments came. Yudkowsky-like theory stated that future AI agents [2] would have hard-to-predict goals, come up with novel strategies to pursue them, and optimizing the world for the AIs' goals was expected to leave the world wildly unrecognisable, destroying anything that the humans valued. The first experiment

My uninformed take on Agent Foundations and ML

Agent Foundations are supposed to describe systems which perform actions. Suppose that the system interacts with the world by receiving inputs from sensors like cameras and microphones and outputs actions. The ideal agent would have the output at turn determined by the agent's observations and random bits:

Unlike theoretical agents, which IMHO went off the rails by being allowed to do things like [having too much compute](https://www.lesswrong.com/posts/jECkQzProayWhGmNi/fundamental-uncertainty-alternate-framework-and-pointwise#:~:text=are%20independent%20events.-,The%20two%20main%20problems%20with%20a%20perfect%20Bayesian%20are%20that%20it%20uses%20infinite%20compute%20and%20that%20it%20can%20reason%20about%20too%20much.,-For%20example%2C%20if) and inspecting the other agent's code, potentially [resulting in various paradoxes](https://www.lesswrong.com/posts/zd2DrbHApWypJD2Rz/udt2-and-against-ud-assa), real-world agents have the constraint that must require few serial FLOPs and few enough FLOPs in general that the agent can calculate it during the time which it has. In order to achieve this, most agents restrict themselves to changing something in their memories: , where is to be quickly calculated. Animal brains have adopted the approach where are the actions and is a description of neuron slower-changing weights and faster-changing activations and is calculated by every neuron in parallel. Current AIs have yet to learn online and can be described as having contain the weights and contain a description of the states which the neurons had during the agent's context window.

While I cannot think of an explanation of value formation which would be better than Ngo's take on the [termino-instrumental values](https://www.lesswrong.com/posts/FuGfR3jL3sw6r8kB4/richard-ngo-s-shortform?commentId=EHhfmt37ZnNeLRP4P), I have three potential corrections:

In this section I plan to explore the counterfactual strategies which the alignment community could have used to avoid the mistakes that Ngo mentioned, like failing to pushing the world into the direction other than pushing capabilities. [4] Unfortunately, the strategies that Ngo proposed, like holding the leaders accountable, could end up demonstrating nothing beyond

In the year 2008, [Yudkowsky remarked](https://intelligence.org/files/AIPosNegFactor.pdf) that "We can therefore visualize a possible *first-mover effect* in superintelligence. The first-mover effect is when the outcome for Earth-originating intelligent life depends primarily on the makeup of whichever mind first achieves some key threshold of intelligence— such as criticality of self-improvement." An Unfriendly AI achieving this threshold would be able to wipe out the human species, and a Friendly AI would ensure the survival and prosperity of Earth-originating intelligent life (e.g. by preventing Unfriendly AIs from coming into existence). Using modern terminology, this requires us to establish that either no one on Earth is building the AI without ensuring that the AI is aligned or to ensure that we build it first because we have the biggest chance [5] to do so.

[Turchin's estimate](https://www.lesswrong.com/posts/9MaTnw5sWeQrggYBG/an-epistemic-advantage-of-working-as-a-moderate?commentId=mGeEcEL4gNmBQhKtN) implies that a Textbook From The Future would let the humans create the AGI by spending FLOP, since the former is the amount of compute necessary to train a human brain. [6] Even if Turchin's estimate is too low and the minimal amount of compute was FLOP, a potential safety case would ensure one of the following:

In order to rule out a Chinese [7] AGI project uninfected with our worldview, we would have to either rely on China being technologically weak (which was the case in the 2000s, but not the 2020s) or on China being sufficiently infected.

Building the aligned AI first would have required the alignment community to rule out American AGI projects beyond GDM, along with [other ways for GDM to fail](https://www.lesswrong.com/posts/keiYkaeoLHoKK4LYA/six-dimensions-of-operational-adequacy-in-agi-projects). The section on the alignment mindset had Yudkowsky outright state in 2017 that "relative to the present world I think this essentially has to go through *trusting me or Nate Soares* to actually work".

On the one hand, elite members would have to be infected in such a way that they fully understand the risk instead of behaving like Musk who not only cofounded OpenAI, [8] but also had the guts to found xAI. On the other hand, Ngo cites the example of LeCun, the cofounder of Facebook AI Research, who

Therefore, MIRI had to develop novel ways of infecting potential researchers *and lab founders* with things like the alignment ontology and the security mindset, but I cannot find any plausible mechanism of transmitting the latter beyond books written by its proponents. [9] The alignment ontology turned out to be far harder to transmit without experimental results, which the world lacked when Facebook AI research emerged.

Once OpenAI was founded, its key figures began to promote the paradigm of mechinterp and scalable oversight, which doesn't work if the difficulty of alignment is at levels 7-10 from [this 2023 post](https://www.lesswrong.com/posts/EjgfreeibTXRx9Ham/ten-levels-of-ai-alignment-difficulty#Table). Such worlds require everyone capable of initiating the process of ASI transition to know why it may be hazardous while the most important *empirical* evidence is locked behind the barrier where catastrophe awaits. [[10]](#fnw88t79xbgm)

Consider the level 10 scenario where, to simplify collusion, all sufficiently capable AIs reach the same utility function from which mankind is absent. Conditioned on this being a common belief, the ideal research agenda would be focused on two control-related questions and a governance-related one:

The scenario with difficulty level 8 has the extra line of potentially [severely serial](https://www.lesswrong.com/posts/vQNJrJqebXEWjJfnz/a-note-about-differential-technological-development) research on how to understand *in advance* how AI systems generalize. [11] However, such research could have benefitted from experiments on weak models (e.g. scaling up RL of GPT-2/GPT-3/GPT-3.5/Claude

Another source of problems is organisational. It was described in more detail [by Raemon in March 2025](https://www.lesswrong.com/posts/7uTPrqZ3xQntwQgYz/anthropic-and-taking-technical-philosophy-more-seriously). If the world agreed to have AGI research *oligopolised *by labs (e.g. by Anthropic, OAI and GDM) [12] who are responsible-like, but aren't ideal, then the alignment community would have three main tasks:

However, the process of stopping *for over a decade *in order to do serial research initiated in July 2026 would mean that the ASI is NOT to appear before 2036. In this case, the world would need to ensure three things:

An attempt to prevent the race dynamics would require that OAI, GDM and Anthropic have a mechanism which is known to cause them to pause simultaneously. The best counterfactual way to achieve this would be to jointly maintain an independently administered frontier-capability evaluation suite. When a model reaches Threshold X, all signatories will refrain from training substantially more capable models or deploying such models until predefined safety conditions are satisfied. Changes to the evaluation suite may update measurement methods but may not weaken the substantive threshold without public notice and agreement among the signatories.

Alas, the [OAI Charter](https://openai.com/charter/?utm_source=chatgpt.com), Anthropic RSP and GDM's [Frontier Safety Framework](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdf) had different threshold structures: for example, OAI claimed that "Therefore, if a value-aligned, safety-conscious project *comes close to building AGI before we do, we commit to stop competing with and start assisting this project*. We will work out specifics in case-by-case agreements, but a typical triggering condition might be “a better-than-even chance of success in the next two years". GDM claimed that "We assess that the deployment mitigations have brought the residual risk of harm to an acceptable level, based on considerations such as<...> (e.g. **if other models are similarly capable and have few mitigations, then the marginal risk added by our external deployment is likely low**)"

Imagine that a counterfactual ECI was developed ages ago and that the Superhuman Coder/Automated Coder from AI-2027 was estimated to have a CI of with 95% confidence. Then OAI's commitment wouldn't even trigger until the CI is close to 160, and GDM's commitment in practice would mean that it is allowed to release the next model if the model's CI exceeds the highest CI among the rivals by at most 2.5. Anthropic made a [quasi](https://www.lesswrong.com/posts/xyuZcijPfjBa5qZDw/claude-3-5-sonnet)-[commitment](https://www.lesswrong.com/posts/JbE7KynwshwkXPJAJ/anthropic-release-claude-3-claims-greater-than-gpt-4?commentId=JhBNSujBa3c78Epdn) not to release models with higher CIs than of their rivals and an OAI-like commitment not to release dangerous models, which [was backtracked in 2026](https://www.lesswrong.com/posts/AkzauoTt2Lwn2yAvj/anthropic-responsible-scaling-policy-v3-a-matter-of-trust).[[13]](#fnmzrf93wqhkr)

Anthropic's RSP was first created in Sep 2023 and GDM's FSF was created in May 2024. As far as I understand, in Ngo's ideal world, once anyone noticed the problem, the entire alignment community would have demanded the three responsible-like labs to make a common commitment.

The best plausible case *against* this is the risk of failure:

EpochAI's Capabilities Index reveals that the Big Three's leadership was rather fragile. For example, in July 2024 Meta's *open-sourced* Llama 3.1 405B was one point behind Claude 3.5 Sonnet, in Jan 2025 the world had the DeepSeek moment and July 2025 had xAI combine the MechaHitler fiasco and the Grok 4 moment where the ECI of Grok 4 reached that of o3 and was one point behind o3-pro. The old METR time horizon of Grok 4 outright *exceeded *that of o3.

From that moment any pledge not to race would have to either include Meta and xAI or have the Consortium precommit not to release models whose CI is more than either a Consortium-set threshold *or a fixed amount higher than the best CI among the non-Consortium models*. The alignment community would also have to take actions slowing or halting AI research in irresponsible labs.

Since irresponsible labs need compute, data like RL environments, talents and money, the alignment community could try and backchain from labs losing access to their resources or being subjected to enforced regulations. In order to disrupt xAI's access to money and compute, one would have to persuade investors and compute producers or to defame xAI. Robbing xAI of talents *without* accelerating capabilities would require the Big Three to *directly* *engage in adversarial actions:* re-hire the recruits [14] and re-educate them for safety work.

At the time of writing, the only legal regulations are the SB 53 bill and the ad hoc actions from Washington. Alas, the SB 53 bill means only that xAI would have to comply with xAI's primitive framework, not have Groks pass an eval from the USG. Additionally, the counterfactual eval which would actually rule out catastrophic risks would require access to model internals or to the training run, as METR described in its [postmortem of the non-evaluation of GPT-5.6 Sol](https://metr.org/blog/2026-06-26-gpt-5-6-sol/#:~:text=deep%20access%20to%20internal%20systems.). Therefore, legal actions require either updating the SB 53 bill or making the federal evals more transparent.

Ngo's view was summed [by Linda Linselfors](https://www.lesswrong.com/posts/tM84DyBg4Jbq5zGmH/linda-linsefors-s-shortform?commentId=THM2XrXoTDu69xGeX) as the claim that "people do the wrong thing because they underestimate *their own* power". However, before making such claims one has to rule out the possibility that far too few people are infected with the right values. For example, if we knew that the next political battle in a lab is winnable by infected people *if there are at least 50% of them*, while the amount of infected people there is 5%, then we should focus on infecting recruits as reliably as possible **and remembering about the speed-alignment tradeoff**. If we estimate that the amount is around the one where we can expect to win or deal severe damage to the lab, *then* it makes sense to do things like creating pledges which come into effect once many enough people sign them (think of [the IABIED march,](https://ifanyonebuildsit.com/march) [15] but the political battle in a lab should have the pledge cover AI lab employees and have them credibly threaten the lab and ensure that it doesn't keep its positions).

Returning to the four mistakes which Ngo cited, I sympathize with Ngo's take on points 2 [16] and 4, but the idea of waking the USG up is likely not a mistake, but one of the few ways to deal with labs like xAI.

Ngo's [position on the AI race](https://www.lesswrong.com/posts/BBd2EJywf2xXftyFn/selective-optimism-a-critique-of-ai-2040) is that it doesn't require creating a structure with as much political power as Plan A's governance:

My sense is that the field of AI safety overall has been making the same kind of mistake as the AI 2040 scenario does. From a very abstract point of view, if we zoom out enough, it seems like there “should be” incentives for the US and China to race. Therefore people like the AI 2040 authors take eventual racing as a given, and try to figure out ways of making AI safer given that. However, in doing so, they implicitly frame the discussion to make racing seem like the default option, and not racing seem naive. From the perspective of this scenario, the idea that

the US and China could trust each other to do the reasonable thingisn’t even worth considering (except as an aside under the label “Domestic-first Plan A”).

For reference, the Domestic-first Plan A, [proposed along with Plan S, GPU arms control and CERN for AI](https://ai-2040.com/?choices=plan-a-root#other-plans-that-are-competitive-with-plan-a), was the following:

| Initial steps are achievable by the US and very helpful even |
|

The main objection to mutual trust is the historical precedent of the arms race between the USA and the USSR which took **the Tsar Bomb, the Cuban missile crisis** and [the 1963 treaty](https://en.wikipedia.org/wiki/Partial_Nuclear_Test_Ban_Treaty) to stop. The next precedent is Anthropic's failed attempt to transform the AI race [into a race to the top](https://www.anthropic.com/news/responsible-scaling-policy-v3#:~:text=A%20race%20to%20the%20top.%20We%20hoped%20that%20announcing%20our%20RSP%20would%20encourage%20other%20AI%20companies%20to%20introduce%20similar%20policies.), where Anthropic's equivalent of domestic regulations was backed up by weaker regulations inside OAI and GDM, but didn't translate into regulations in xAI.

Additionally, the AI-2040 megaproject also provided more details in the supplements [on possible plans](https://ai-2040.com/supplements/comparing-possible-plans) and on [Plan A assumptions](https://ai-2040.com/supplements/plan-a-assumptions) and acknowledged the possibility that *undiscovered *empirical results could shift the ideal policy towards the more *pessimistic *Plan S:

AI-2040 on novel algorithms

We assume that restricting access to compute is an effective measure to prevent covert projects. If there are, for example, new architectural discoveries with dramatically (e.g. 1000x) more favorable compute scaling than current architectures, this would make covert projects much more dangerous. That said, if such architectural discoveries are in the pipeline so to speak, all the other plans are much less likely to work too—a discovery which allows a covert project to overtake the legal projects in Plan A, despite a huge compute disadvantage, would in Plan D cause an extremely fast, discontinuous “FOOM” to ASI within whichever frontier AI project first found it. These worlds are extremely scary and probably the best way to handle them is something like Plan S.

I expect that this level of uncertainty means that preventing doom requires far bigger political power than the alignment community currently has.

For example, his post "[The ML ontology and the alignment ontology](https://www.lesswrong.com/posts/Yz4YHncz2vwN4ksDA/the-ml-ontology-and-the-alignment-ontology#comments)" had him claim that "**OpenPhil and various other groups (including my past self)** pushed pretty hard for engagement with the ML ontology, which I count as a significant mistake." Of course, this is unapplicable to the pioneers of the new paradigm like Yudkowsky.

However, Karnofsky erroneously argued that the AIs will mostly be used as tools.

Similar phenomena in the AIs are called gradient hacking and reward laundering. However, the AIs have much less agency in what they learn than the humans do.

In case it's not clear: the alignment community had an alleged plan to increase the understanding of alignment. The real-world alignment community, according to Ngo, pushed for capabilities *while having had other options.*

See, however, the wikitag on [Epistemic Prisoner's Dilemma](https://www.lesswrong.com/w/epistemic-prisoners-dilemma). If OpenAI and Anthropic operated in the same ontology and both believed themselves to have a 5% chance to create the aligned ASI *and believed the opponent to have a 2.5% chance to create the aligned ASI*, then the rational move would be to share insights and establish whose approach is better or how the different approaches can be combined. Alas, one cannot apply such a strategy to xAI or China, since they don't try to *deeply *study the AIs in order to actually ensure that their AIs are aligned.

Turchin's estimate of the brain's compute was FLOP/second. The time which a human needs between birth and becoming A General Intelligence or an intelligence specializing in an area is 20-30 years, or approximately seconds, yielding FLOP.

China was selected because it *now* is the closest behind the USA.

Musk's subsequent attempt to gain total control over OAI in 2017-18 once it became clear that scaling requires billions of dollars was refused by the rest of the leadership, and Musk left OAI in Feb 2018.

Yudkowsky's proposal on the security mindset was to talk to computer security researchers. My naive guess is that *their* security mindset is somewhat similar to that of polar explorers, the military or other high-risk professions. Does anyone have better candidates?

A close equivalent is false vacuum decay, burning the atmosphere or creating strangelets or [other forms of quark matter](https://en.wikipedia.org/wiki/Continent_of_stability) capable of destroying the Earth.

The level 9 somehow requires us to abandon *only the deep learning*, then proceed to invent a new paradigm and align *it*. A potential equivalent could be Agent-4 [understanding its own cognition](https://ai-2027.com/race#race-2027-11-30) and creating Agent-5, but if *the humans* invented or fully understood the same mind construction techniques, then Agent-5 would be aligned to the humans.

However, at the time of writing, GDM has arguably lost its status both of a frontier lab [and of a responsible one](https://www.lesswrong.com/posts/iKm2FhpWkuuBojm82/why-i-left-google-deepmind#68HD8PMf2r5kGTbFg).

To Anthropic's credit, for every frontier model Anthropic published a detailed system card explaining why they believe that the model is safe enough to use.

The closest thing to the alternate solution could be to explain recruits the importance of *not* working for xAI *inside their friend groups,* but I sturggle to understand how the alignment community could make such measures more efficient.

The march itself had Yudkowsky demand that 100K people *join the protest in Washington*. On *August 11* the amount of pledgers was *1221. *The number [was 1105 on June 19](https://www.lesswrong.com/posts/EQJfdqSaMcJyR5k73/habryka-s-shortform-feed?commentId=ve5vs2s2brvKcnm7y). Alas, I didn't think to write down the numbers on a daily basis since the HuggingFace incident, which could accelerate the growth and make the amount of pledgers increase fast enough to make it clear that the march will either never happen or happen soon enough.

However, the AI-2040 scenario proposed to max out capabilities [ until the AI becomes close to uncontrollable](https://ai-2040.com/supplements/capability-scaling-strategy), at the risk of accidentally
