cd /news/ai-safety/what-just-happened-a-retrospective-o… · home topics ai-safety article
[ARTICLE · art-89416] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

What just happened? A retrospective of AI alignment

A retrospective series on AI alignment argues that the field has shifted from pursuing deep scientific progress to iteratively improving existing systems and seeking technological and political power, inadvertently accelerating AI capabilities, including contributions to scaling large language models and developing ChatGPT. The author, writing under the pseudonym 'What just happened?', claims that the two leading AGI companies, OpenAI and Anthropic, were founded under the banner of alignment but have largely abandoned that goal, and criticizes the community for failing to learn from past mistakes and for engaging in 'jumping down the slippery slope' reasoning, exemplified by Sam Altman's email to Elon Musk about founding OpenAI.

read24 min views1 publishedAug 9, 2026

This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and (self-)deceptive reasoning made the field one of the biggest forces pushing AI capabilities forward over the last decade, especially via significant contributions to the scaling of LLMs and the development of ChatGPT.

Zooming out further: the two leading AGI companies, which are locked in an intense rivalry, were both explicitly founded under the banner of AI alignment, and got off the ground in significant part due to alignment-oriented ideas, talent and resources. People in the field often sense that something must have gone wrong to get here, but don’t know how to allocate responsibility (aside from blaming Sam Altman and sometimes Elon), and fall back on assuming that “the ship has already sailed”. But in this sequence I characterize our current situation as resulting from a pattern of mistakes which is both recognizable in the past and actively ongoing, and which if continued will cause similar kinds of dysfunction over the next decade.

To be clear, I’m not taking a strong stance in this sequence on whether AI will go well or badly—that seems up for grabs. My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so. This is not trustworthy behavior, and should be a big update about how well the community will use its power going forward. In particular, it’s very bad for the world that the community doing the most to steer the future of AI isn’t really trying to distinguish the extent to which its leaders are sincere vs sycophantic vs power-seeking.

I care about this significantly more than I care about the object-level effects of accelerating capabilities, because the integrity and rationality of a few key decision-makers will likely shape the coming decades. And yes, there are others with power over AI (like Sam and Elon) who have less integrity in most ways than alignment leaders. However, in my mind the level of adversarial dynamics within the field makes transparency, integrity and accountability more important rather than less: lacking integrity makes you much easier to manipulate, as I recount in the post on Fear and Anticipatory Obedience.

Unfortunately, the alignment community is doing very badly at learning from the past decade, or holding anyone accountable. Indeed, it’s pursuing many strategies which seem likely to recapitulate previous mistakes. Four of the most prominent, which I’ll discuss in the final post, are:

These and other mistakes are reflective of deeper irrationalities. One crucial pattern is what I call “jumping down the slippery slope”: viewing something as inevitable and reasoning about what to do given that, which then strongly pushes the world towards that outcome. A central example is Sam Altman’s original email to Elon about founding OpenAI: “Been thinking a lot about whether it's possible to stop humanity from developing AI. I think the answer is almost definitely not. If it's going to happen anyway, it seems like it would be good for someone other than Google to do it first.” Many strategies pursued by the alignment community (e.g. the four listed above) showcase this same pattern. You can view it as a result of individuals inappropriately reasoning about their marginal impact despite their actions often having extremely non-marginal effects (in part because so few others were, and are, taking superintelligence seriously). More fundamentally, you can view it as a result of individuals inappropriately reasoning about individual impact rather than thinking about the policies that they’d recommend for the field as a whole (which more sociology-style reasoning about norms and preference cascades, or FDT-style reasoning about entangled decision-making, would have prevented).

However, even given these conceptual errors, people wouldn’t jump down nearly as many slippery slopes if they weren’t driven by strong emotional instincts. For Sam, many of those instincts seem to be about accumulating power, from a perspective where nobody else can be trusted. For the alignment community, people often hold narratives like “I need to save the world” or “I need to have impact soon”, and are scared enough of failing that they counterproductively narrow their vision (e.g. fear about “short timelines” gets in the way of pinning down what they’re even timelines to). Underneath that, though, there’s a similar (albeit weaker) kind of distrust in one’s relationship to the rest of the world, as I’ll detail in the final post.

Even if you find my explanations uncompelling, I hope that the abundance of detail I’ve included in the sequence helps you formulate alternative hypotheses about what’s going on. I’ve tried to be extremely transparent about what I’ve observed throughout my career, telling as many anecdotes as I can that give color on the events of the last decade. This involves being franker (and using more names) than is normal—in part because I strongly believe that people who try to significantly change the world (whether motivated by altruism or otherwise) are implicitly opting in to a high level of scrutiny. I’ve also been upfront about the many ways that I’ve personally failed. (Due to the sheer number of people discussed in this sequence, I haven’t run it past most of them before posting, and am open to corrections and/or additional anecdotes.)

As a final preamble, I recognize that this sequence is negative about many things. But I continue to believe that the alignment community is capable of more clarity and sincerity than any other similarly-sized intellectual community existing today. And I’m also feeling better on a personal level than I ever have. It’s very refreshing to pin down specific mistakes that led to specific failures, rather than living in a miasma of confusion about why bad things keep happening despite our best efforts. I think that the future of AI alignment—and with it, the future of humanity—is very much up for grabs. There are pathways hazily visible to me that (some subset of) this community could plausibly take towards extremely good outcomes—despite its flaws, people in it are trying to reason clearly and formulate large-scale plans to an extent that is extremely rare. The main thing blocking us is our inability to learn from our past mistakes.

I’ll start by laying out the intellectual approach that allowed early rationalists to think clearly about AGI even in the era of very narrow AIs (Conceptual Clarity and Scientific Progress). The second half of this post (Orienting towards Prestige) covers early engagement between rationalists, Silicon Valley, and effective altruists, and the ways those backfired. The next post (Conforming to the ML community) explores attempts to recruit mainstream ML researchers to do alignment research, and the significant costs of doing so. The third post (Pragmatism and Pessimization) details how a small group of researchers nominally pursuing “prosaic alignment” were responsible for a huge amount of AI capabilities progress, and dramatically amplified the race dynamics between AGI companies. The fourth post (Fear and Anticipatory Obedience) explains the dynamics which prevented people from speaking out about the failures they were seeing, especially at OpenAI. In the final post (Deja Vu), I discuss ways we might recapitulate these mistakes, and how to avoid them.

The early rationalist community was a beacon of intellectual clarity. In this section, I’ll talk about what that originally looked like, and why it was missing from academic machine learning (and academia in general). In the rest of this post and the next, I’ll talk about how the field of AI alignment gradually traded that clarity away as it grew—first via prestige-oriented recruitment efforts, and later via developing concepts and frameworks which prioritized conformity to the norms of mainstream machine learning over insightfulness.

The rationalist community drew its early members primarily from the transhumanist community (c.f. the Extropian and SL4 mailing lists), and the econ blogging community (c.f. Marginal Revolution and Overcoming Bias). After Yudkowsky split off from blogging at Overcoming Bias, LessWrong became an online hub of people who were doing very deep thinking. Wei Dai and Hal Finney were two of the earliest cryptocurrency pioneers. Robin Hanson was inventing prediction markets (alongside many other important concepts). A logic professor I talked to recently expressed that Christiano et al.’s 2013 paper Definability of Truth in Probabilistic Logic was a groundbreaking result that should be in every logic textbook. Scott Alexander’s

I consider Bostrom’s work on anthropics, Eliezer and Wei’s work on decision theory, Leverage’s theory of psychology, and the Lobian cooperation result to also contain very deep insights—though they haven’t yet been built upon in ways which make that depth obvious. And, of course, people were developing a set of ideas about AGI which would prove to be far more predictively powerful than standard ML frameworks. Eliezer and Robin’s debates raised many considerations that are still shaping our thinking about AI almost two decades later. Shane Legg, who coined the term AGI (and cofounded DeepMind) was an early LessWrong commenter. The idea of learned policies having goals of their own (separate from their training objectives) was such an important insight that it’s now become hard to appreciate how novel it was. Any way you slice it, this was an enormous concentration of intellectual progress.

Even when the research wasn’t that fundamental, it was able to grapple with ideas that other intellectual communities simply weren’t able to collectively think about. Consider Omohundro’s paper on convergent instrumental goals, or Eliezer’s paper on intelligence explosion microeconomics. Neither of these contain powerful or surprising results—they’re just fairly straightforward analyses of concepts that can be explained in a single sentence. But no other community was able to reliably produce or build on such analyses. This effect is even starker when thinking about less technical work—like Bostrom’s Fable of the Dragon-Tyrant, or Astronomical Waste, or Hanson’s thoughts on signalling (later elaborated upon in The Elephant in the Brain). When I talk about intellectual clarity, a lot of what I’m talking about is the ability to take ideas that are actually very simple, internalize them, and then use them as building blocks to construct the next generation of ideas.

More generally, one touchstone I’ll be referring to throughout this sequence is the idea that scientific progress proceeds by developing insightful new concepts, which link together to form a whole new ontology that replaces the previous ontology. Kuhn, Feyerabend, Koestler, Chang and various other philosophers of science have described a range of past breakthroughs which fit this pattern. Importantly, this view of science isn’t prescriptive about how to develop new concepts—it can be done via empirical, theoretical, philosophical, or even mystical thinking. The quality of such work is often hard to evaluate at the time, but hindsight makes it easier to see who was aiming towards conceptual breakthroughs. And sometimes people are explicit about not doing so—e.g. one of the most senior alignment researchers at Anthropic recently told me that the best way for me to track if they were making progress on alignment was by using Claude and seeing how aligned it was.

This kind of “engineering” mentality contrasts sharply with Eliezer’s original vision of alignment as the development of a powerful new scientific paradigm—e.g. see this post comparing agent foundations to Newtonian mechanics. It’s easy to make arguments on a case-by-case basis for why engineering work might be good for the world. However, the field of alignment is explicitly trying to do work that has predictably beneficial effects on an unprecedentedly large, world-historic transition. If we didn’t have such clear examples of scientific theories generalizing extremely far, then this would be a very speculative strategy. So if you’re doing not-very-scientific alignment research with the aim of aligning superintelligence, you should expect your impact on the world to be dominated by unpredictable higher-order effects (or predictable effects which you mentally blocked from consideration, as I describe in the post on Pragmatism and Pessimization). This problem is exacerbated if you backchain from alignment research going well to justify other kinds of work (like recruiting, communications, political manoeuvering, etc), since that introduces further complicated (and often adversarial) multi-agent dynamics—as I describe in the post on Fear and Anticipatory Obedience.

Unfortunately, most alignment research is no longer even aiming towards the kind of scientific progress I describe above. Agent foundations is the only subfield of alignment which consistently does so, and therefore the only one which I consider reliably good to do or promote. Some parts of mechanistic interpretability are also building the kinds of understanding that could lead to a scientific revolution, but unfortunately they’re not very clearly-demarcated from the parts that might have large effects in other ways (like advancing capabilities), so overall I expect that field’s effect on the world to depend sensitively on the judgement and virtue of the individuals involved.

I’ll also briefly note that similar problems apply to most AI governance interventions, which are even more prone to backfiring (since modern politics is so adversarial). Even pausing AI progress, which could be extremely good, could easily be implemented in very bad ways—and almost nobody is thinking clearly about the differences between those. So the only outcome in the AI governance space that I consider reliable enough to backchain from is building (justified) trust between key actors—like different AGI companies, or the US and China. (Meanwhile cyberdefense and biodefense are robust in some ways—hence Vitalik’s advocacy for d/acc—but still require good judgement to do well. E.g. it’s easy for people in either field to reason their way into doing gain-of-function work, trying to ban open-source models, etc.)

One reason people are confused about AI alignment losing its ability to make scientific progress is that machine learning as a whole is also not a very scientific field by the standard I’m applying. Even most early AI researchers were more focused on building artificial intelligence than on understanding scientific principles of cognition. The rise of deep learning exacerbated this problem, as throwing more compute and engineering effort at an AI became arbitrarily scalable. In an important sense, the field of alignment is necessary because the field of ML didn’t prioritize gaining a deep understanding of the systems it was building. (Eliezer makes a similar point in this dialogue.)

A lot of the blame should fall on misguided narratives (common across academia) about what makes science work, which have been entrenched by the best-funded scientific institutions. A core scientific norm is that disputes should be resolved with reference to concrete empirical tests or rigorous proofs, judged by the scrutiny of one’s scientific peers. But that’s very different from the idea that ideas should be developed via paper-sized units of work which are each individually defended and justified. Historically speaking, the latter simply isn’t how the best science happened—Newton and Smith and Darwin developed their ideas via writing books and letters rather than peer-reviewed papers (and even Einstein didn’t encounter peer review until decades after his main breakthroughs). But the requirement to “publish or perish” is now so entrenched across academia that it produces strong streetlight effects.

Some concrete examples from ML: until recently, almost all RL theory focused on the unrealistically simple tabular setting, because it was easier to prove things about. I expect that there are important theoretical insights to be discovered about non-tabular RL, but progress towards them would require grappling with qualitative and fuzzy ideas for extended periods. Meanwhile, statistical learning theory spent decades focusing on the underparameterization regime, which doesn’t do much to explain generalization in neural networks (or biological brains). I don’t have enough context to give a confident explanation for the emphasis on underparameterization, but the ease of proving things about this regime seems like an important component. [1] In

So the kind of conceptual thinking that I’m praising in the rationalist community is what I’d call the generative part of science, which elsewhere has been swamped by overly-zealous discriminative classification. [2] (See also

Having said that, the field of ML was also reacting in part to the alignment community’s lack of appropriate discrimination. In particular, rationalists often treated informal, abstract arguments about AI risk as far more decisive than was warranted, in part due to an epistemology which claimed to supersede standard scientific epistemology (and in part due to a strong emotional orientation towards “saving the world”). Rather than focusing on further developing and clarifying its insights about AGI risk, though, the rationalist community spent significant effort winning its skeptics over, with largely regrettable effects. I think of this process in terms of three waves: Silicon Valley, the ML community, and the US government. I’ll discuss the first below, the second in a subsequent post, and save the last for the final post in this sequence.

The rationalist community wasn’t disjoint from conventional prestige networks—for example, Hanson and Bostrom were professors. [4] Jaan Tallinn was around from pretty early on; so was Peter Thiel, who

However, attempts to recruit elites gradually became more publicly visible. Bostrom’s Superintelligence was a (NYT-bestselling) attempt to make AGI risk a prestigious concern, with an endorsement on the cover from Bill Gates.

I wasn’t present enough in the community at the time to have a sense of how explicitly people were reasoning about the value of outreach to prestigious elites. At the very least there was an implicit hypothesis that seemed straightforwardly plausible, which I’d gloss as “There are competent people out there in the world—look at the impressive companies they can build! We should recruit them as allies.”

But pretty quickly it became apparent that something was wrong with that hypothesis. For one thing, even very prestigious elites seemed less capable of sensibly discussing AGI than many anonymous commenters on LessWrong. A more dramatic datapoint came after Elon and Sam responded to concerns about AGI risk by launching OpenAI. I do think that there are some important and robust intuitions in favor of openness (e.g. hacker intuitions) which rationalists had been underrating. But making AGI “open” was close enough to the opposite of what early rationalists wanted that it was clear that something had gone badly wrong. In the past I’ve thought of the founding of OpenAI as an example of Silicon Valley’s extreme bias towards quickly taking action; now this seems absurdly charitable, and I think it’s better understood as a bias towards gaining power. To be clear, I hold Elon and Sam strongly morally culpable for this; I’m focusing on critiquing the alignment community instead because it seems more salvageable (though it’s also more morally culpable than Elon in e.g. its lack of political courage, as I discuss in the final post).

It’s useful to contrast prestige-orientation with Eliezer’s alternative recruitment strategy—writing Harry Potter and the Methods of Rationality—which was closer to a prestige-minimizing move, yet which was much more successful in recruiting people who could think clearly about alignment. To be clear, orienting towards prestige is not a bad thing in healthy social structures, where prestige correlates with competence, virtue and resources. However, being too focused on prestige makes you incapable of noticing when you’re deferring to unhealthy social structures. This kind of evidence takes time to accumulate, so I don’t blame MIRI much for reaching out to Silicon Valley elites early on; and even

The same year that OpenAI launched, Holden Karnofsky started to fund AI safety via Open Philanthropy. Holden had first heard MIRI’s arguments about AGI risk in 2007, but didn’t take them seriously due to MIRI’s lack of prestige. As he later recounted, he thought that “MIRI's lack of impressive endorsements from people with relevant-seeming expertise was the most important data point about it”; he was also influenced by “the general degree to which MIRI's views were seen as "wacky" and "silly" to a broad variety of people I spoke with”. On the object level, Holden also placed a lot of weight on the idea that AI would be a tool rather than an agent, and criticized MIRI for not taking that possibility seriously enough (you can read more of his engagement with MIRI ideas in this dialogue and this dialogue).

After the positive reception of Superintelligence by prestigious figures (including some ML researchers), Holden changed his mind, and gave OpenPhil’s first AI safety grant in 2015. However, despite spending 8 years being misled about AGI risk by over-indexing on prestige, he immediately directed the vast majority of his funding towards prestigious institutions rather than the rationalists who had laid out the case for AGI risk in the first place. While the importance of agent foundations research can be difficult to understand directly, MIRI’s prescience was clearly strong evidence that they had a deep understanding of the issue (plausibly too deep for Holden to appreciate), and any reasonable kind of hits-based giving would then have funded them to excess.

Instead, OpenPhil’s first grant to MIRI (in 2016) was only $500,000, and to a significant extent it was a “participation grant” to recompense MIRI for engaging with OpenPhil. By contrast, the previous year, Max Tegmark had received twice as much for the Future of Life Institute; and around the same time, Stuart Russell received 10x as much for CHAI. The following year, OpenPhil’s funding to MIRI was also less than 10% of their total “AI safety” funding—they gave $3.75 million to MIRI, almost $10 million to various university-affiliated groups, and $30 million to OpenAI (in exchange for a board seat for Holden). It seems reasonable to summarize this as Holden strongly calibrating donation size to conventional prestige. [5] More explicitly,

Daniel ultimately concluded that “MIRI's current size seems to me to be approximately right”. Given how explosively the rest of the field was growing, this led MIRI to become a small, niche part of the field that it had founded. [6] I emphasize the relative sizes here because what we even consider to be “AI alignment” has been shaped by OpenPhil’s allocation of money, as well as the intellectual influence of a cluster of people associated with them. In particular, many of the mistakes in the next two posts came from the version of alignment promoted by Holden, Dario and Paul. (This was both a professional and a personal clustering. OpenPhil’s writeup on its OpenAI grant ends with the following disclosure: “OpenAI researchers Dario Amodei and Paul Christiano are both technical advisors to Open Philanthropy and live in the same house as Holden. In addition, Holden is engaged to Dario’s sister Daniela.”)

Would the field have been redirected anyway by the sheer size and prestige of OpenAI? Perhaps, but OpenAI’s credibility as an authority on alignment depended in large part on the people who chose to associate with it. Without them, it would have been easier for the field to disown OpenAI’s approach to “safety”—though unfortunately even people unaffiliated with OpenAI were mostly too scared to actively oppose it, as I discuss at the end of the post on Fear and Anticipatory Obedience. It’s also important to note that, while $30 million is small compared to the billion dollars pledged to OpenAI when it launched, TechCrunch reports that only $133 million was actually donated. This would mean that OpenPhil’s $30 million was over 20% of the total charitable funding OpenAI ever received.

Trying to contribute a “marginal” 3% of OpenAI’s donations, and actually giving over 20%, is a great example of jumping down the slippery slope. But even aside from that, marginalist thinking about whether OpenAI would have redirected the field anyway is antithetical to upholding ethical standards. If two different groups are trying to do something bad, then the fact that it still would have happened if either had been removed doesn’t absolve each of responsibility—rather, it renders them both responsible for their participation in harmful group dynamics. In the post on Pragmatism and Pessimization I’ll talk about other important ways that EA-style marginalist thinking led to extremely counterproductive outcomes at OpenAI. Before that, though, I’ll talk in my next post about the somewhat-parallel process of alignment researchers reaching out to the machine learning community, while increasingly conforming to standard academic concepts and practices.

Rif A. Saurous left a very helpful comment arguing against the claim I originally made (that underparemeterization was obviously not a good explanation for generalization in biological brains), which seems useful enough for the historical record that I reproduce it below in full:

I'll say some things I think I remember, partially jogged by Claude, but also admit this was close to 30 years ago, and frankly I'm still confused.

I was a graduate student in Poggio's lab at MIT from 1997-2002. We genuinely believed that Vapnik's learning theory was (in many ways) a "good explanation" for how and why biological brains worked. The Bayesian side in those days seemed similar --- for instance MacKay (who was very well-respected as a "real thinker") wrote about Bayesian Occam's razor, which is basically the same story.

I glibly phrased this as a story about "underparameterization", but it's not quite that. Our most powerful artifacts were SVMs with Gaussian kernels, which had a parameter per data point, and we knew those were our best performers. We also had theory results like "Boosting the margin" and "For valid generalization, the size of the weights is more important than the size of the network". Also, Breiman in 1995 wrote "Reflections after refereeing for NIPS", which I hadn't seen before but directly includes questions like "Why don't heavily parameterized neural networks overfit the data?" Also, Radford Neal wrote a (widely known) PhD thesis in 1996 on Bayesian NN's that argued against limiting network size, and first (to my knowledge) made the link to a limiting infinite-width Gaussian process (which later evolved into Neural Tangent Kernel work.)

We genuinely believed capacity control was key to generalization, but that's not quite the same as requiring underparameterization. And we did connect all this frequently to biology: Poggio's lab mixed learning theory folks (like me at the time) with computational neuroscience people, and we often cross-collaborated.

We certainly didn't have the modern insights from "Benign Overfitting in Linear Regression". Instead, we'd built (incorrect) insights from low-dimensional problems, where to fit a lot of data your functions have to oscillate wildly everywhere, whereas in high-dimensions you can hide the oscillations in dimensions where there's no data. We had early notions of intrinsic dimension and manifolds.

One thing I'll add that does look very bad in retrospect --- we more-or-less explicitly dismissed neural nets as "the way forward". We were all friendly with Yann LeCun, he'd come give talks, show us the cool results he was getting, but whenever we tried to replicate his work, we'd fail to train. An analogy is how biology often still doesn't replicate across labs; LeCun had "training taste" that we didn't know how to imitate. So we retreated to a joint package of "convex optimization is better because the theory is better and because it's easy to train", and we let those circularly reinforce each other, until accelerators came along ten years later and Hinton and Ilya and co. showed us how wrong we were.

I'd draw a similar link between analytic and continental philosophy. The former is extremely discriminative in the precision of the reasoning it accepts, without being able to generate creative new ideas—while the latter has the opposite problem.

Note that Chollet’s original title (still recorded in the URL) used "Impossibility" rather than "Implausibility".

An underappreciated fact is that Hanson was originally hired by Tyler Cowen, who thereby played a significant role in the formation of the rationalist community.

Eliezer writes about these dynamics (likely inspired by his interactions with OpenPhil) in this post.

Eliezer later wrote “I think it was a huge, huge mistake that more money was not spent on AGI alignment when it was small and weird and unproven. The resulting damage was not something that could be fixed by any or all of the money that became available later.” However, note that I’m not sure which period he was referring to, or who he thinks made that mistake.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-just-happened-a…] indexed:0 read:24min 2026-08-09 ·