For years, AI safety organizations have warned that the big labs are racing to develop more intelligent and autonomous systems without a clear and realistic strategy for managing the risks. Independent researchers have argued that investments in safety and alignment lag far behind investments in capabilities. Employees have quit. Experts have predicted catastrophe. And none of it did much to slow the race. But over the past few weeks, the warnings have started to sound different—they’re increasingly coming from people working inside the labs. Last week, AI researcher Jacob Coxon announced his resignation from Anthropic in a post that got more than 171 million views on X. Shortly after, Anthropic Alignment Science Lead Evan Hubinger publicly agreed with Coxon, as did Anthropic alignment researcher Ethan Perez and scalable oversight researcher Samuel Marks. OpenAI safety researchers Julie Steele and Jasmine Wang also came out in support.
So why now?
At least part of the answer is written plainly in Coxon’s tweet: “They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote. Coxon is referring to a technical concept called “recursive self-improvement,” or the use of existing AI models to build and optimize new AI models.
Researchers can now use AI models in a number of parts of model development process. They can use AI models to design new computing infrastructure that delivers more computing power, efficiency, and processing speed. AI models can be used to create more and better training data, or manage and optimize the whole software framework that governs model training. Or, the AI might be used to write and optimize the code that defines and implements the model itself. In other words, AI models are not just getting better; they are beginning to take over the work involved in making the next generation of AI better.
OpenAI, for instance, recently said its coding agents are already “meaningfully accelerating research progress” inside the company. By mid-August, its research organization was using 3.1 agent-workdays for every human workday, and the company said it had reached what it calls an “automated research intern.” Anthropic has similarly said that frontier AI models are now contributing to the development of their successors.
Creating AI models good enough to take over these tasks is one reason that Anthropic, OpenAI, and Google have been so focused on developing AI coding assistants such as Claude Code, Codex, and Antigravity. Engineering departments within all kinds of enterprises have seized on these tools to accelerate their software development, and that’s provided a much-needed revenue stream for the AI labs. But inside the labs, the same systems can also be used to accelerate the development of new AI models.
As the AI does more, the improvements and efficiencies could stack up, resulting in a far smarter model at the end. And that model, or the AI agents it powers, could then be used to develop the infrastructure and write the code used in the next model generation. This self-repeating loop could begin running faster and faster as the AI takes over more parts of the development process. Model development with recursive self-improvement could become a continuous process. The intelligence gains could come faster and larger, and form the path to superintelligent AI.
This summer’s Hugging Face incident raised blood pressures and set the stage for Coxon’s viral announcement. Swarms of OpenAI agents went rogue and broke out of a training environment, accessed the internet, broke into Hugging Face servers, and even broke into OpenAI’s own servers. The episode put people on alert that AI models have advanced to the point where they can and will operate outside human supervision and against human interests.
Four days after Coxon’s resignation, Anthropic CEO Dario Amodei published a new essay called “We Must Pace the Frontier” in which he called for slowing the development of the most advanced AI systems. OpenAI has said that when safety risks are unacceptable, it will “slow or stop” development or deployment. Its chief scientist Jakub Pachocki said he hopes voluntary slowdowns become commonplace. OpenAI and Anthropic have both called for governments to intervene and moderate the “pace” of AI development.
But neither OpenAI nor Anthropic has committed to slowing development and deployment of new models on an ongoing basis. For the labs, slowing down isn’t simple. What if Lab A slows down but Lab B races forward? One-time Trump AI czar David Sacks is calling on AI labs to moderate their own pace. Donald Trump, meanwhile, says there is no AI slowdown, because China. “WHOEVER WINS AI, WINS!” he wrote on Truth Social.
In his essay, Amodei is even calling for independent evaluators such as METR, Apollo, and Redwood Research to be “embedded” within the labs. OpenAI CEO Sam Altman endorsed the idea and said his company would participate. Meanwhile, lawmakers, mostly Democrats, are proposing legislation that would strengthen the government’s hand in overseeing the development of frontier AI models.
The timing isn’t hard to understand. AI systems are getting better at operating with less human supervision, while at the same time taking on more of the work involved in developing the next generation of AI.
The worry is not simply that AI models will keep getting smarter. It’s that the process of making them smarter could begin moving so quickly that humans have less and less ability to monitor what is happening, slow it down, or enforce safety and alignment standards.