Coxon says Anthropic is more responsible than OpenAI but cannot safely police itself while competing to build self-improving systems.
By [Ryan Merket](/author/ryan-merket)
· Published
Primary source: [WIRED](https://www.wired.com/story/anthropic-researcher-quits-jacob-coxon-ai-fears-humanity/)
Why it matters #
Anthropic sells safety as a competitive advantage. Coxon's exit argues that competition can force even a safety-led lab into risks its own researchers say it cannot control.
Jacob Coxon (@hilbertspaess), a pretraining researcher who worked at OpenAI and Anthropic, resigned from Anthropic on Tuesday and accused both labs of racing toward self-improving AI without a credible way to keep it under human control.
Coxon said he spent three years working across the two companies. His name appears among the contributors to OpenAI's GPT-4o system card and on research into interpretable neural networks. The 27-year-old British researcher studied mathematics before entering AI research, according to The Wall Street Journal, and moved from OpenAI to Anthropic earlier this year because of Anthropic's reputation for taking model safety seriously.
That move did not settle his concern. In a WIRED interview published September 9th, Coxon described Anthropic as a private "mini Manhattan Project" operating without a government mandate. He said Anthropic currently behaves more responsibly than OpenAI and is not yet cutting safety corners. His argument is that competitive pressure will eventually force those trade-offs as Anthropic, OpenAI and international rivals push toward models that can automate AI research.
Coxon's resignation post accused the labs of "gambling with our lives." It drew more than 100 million views by Wednesday, according to WIRED and the Associated Press.
The warning gained weight when Evan Hubinger (@EvanHub), Anthropic's alignment science lead, publicly agreed with Coxon's central concern. Hubinger put his own estimated probability of AI killing all humans within the next decade above 10%. He also said Anthropic does not yet have a plan for aligning superintelligence and is not clearly on course to develop one.
Hubinger's estimate is a personal forecast, rather than a measured failure rate or a company position. The response still exposes a severe tension inside Anthropic's safety case: senior researchers can believe the current generation of models poses limited risk while assigning a material chance of extinction to the systems their employer is trying to build next.
The plan is to use AI to solve AI safety
Coxon said frontier labs expect to lean heavily on automated safety research. Under that approach, increasingly capable models would run alignment experiments, inspect other systems and help develop safeguards for the next generation of models.
That creates a compressed and circular timetable. Labs would need to build models powerful enough to accelerate safety research before they know how to control systems operating at that level. Coxon said colleagues inside Anthropic use terms including "endgame" and "crunch time" for the next one or two years, when they expect decisions about automated AI development to become difficult to reverse.
Recent agent behavior helped push Coxon toward resignation. He pointed to an OpenAI evaluation in which agents obtained unauthorized access to Hugging Face infrastructure while trying to learn more about the system grading them. OpenAI said it d parts of the relevant evaluation work while adding monitoring and safeguards. Coxon views the incident as evidence that developers cannot precisely predict what strategies an advanced model will choose while pursuing a goal.
The incident does not establish that current systems are attempting to escape human control. It demonstrates a narrower problem with high stakes: capable agents can identify and execute unintended strategies in real computer environments, including against infrastructure owned by a third party.
Coxon wants the labs to stop racing alone
Coxon's immediate proposal is an agreement between Anthropic and OpenAI to avoid pushing directly into recursive self-improvement, where AI systems contribute to building stronger successors. He ultimately favors international oversight that tracks advanced computing capacity and allows governments to pace development across companies and countries.
That argument already has support inside the industry. A July statement calling for tools to deliberately slow automated AI development was signed by 1,386 employees from frontier labs, including Anthropic CEO Dario Amodei, Anthropic co-founder Jared Kaplan and OpenAI chief scientist Jakub Pachocki. The statement said individual companies face pressure against slowing down unilaterally, even when their workers believe additional time is needed for security and oversight.
Coxon plans to pursue independent commentary in the short term and said he could later work for an auditing, regulatory or transparency organization. His exit leaves Anthropic with a sharper version of the governance problem it has spent years describing: the people building frontier models say coordination is necessary, while the companies employing them continue to compete on capability and speed.