# Anthropic touts ‘strongest safeguards’ as AI warning dogs IPO push

> Source: <https://www.cityam.com/anthropic-touts-strongest-safeguards-as-ai-warning-dogs-ipo-push/>
> Published: 2026-09-10 14:28:22+00:00

# Anthropic touts ‘strongest safeguards’ as AI warning dogs IPO push

Anthropic has defended its safety record after an exiting researcher accused the AI behemoth of “gambling with our lives”, as scrutiny builds ahead of a possible IPO.

The Claude maker said it has “some of the strongest safeguards in the industry” and has long been open about the risks posed by increasingly powerful AI. “We have always been transparent that AI will bring both enormous benefits and unprecedented risks,” an Anthropic spokesperson said.

The company pointed to its testing of models for dangerous capabilities in areas such as cybersecurity and biology, as well as its responsible scaling policy, which sets out when to slow development or strengthen safeguards.

The response follows the [resignation of researcher Jacob Coxon](https://www.cityam.com/anthropic-researcher-quits-as-colleague-warns-ai-could-kill-all-humans/), who accused Anthropic and its former employer, OpenAI, of “racing straight to self-improving superintelligence”.

“Neither company is acting responsibly,” Coxon said. “They are racing straight to self-improving superintelligence and gambling with our lives.”

Evan Hubinger, who leads alignment science at Anthropic, backed his comments. “We really do earnestly believe AI could kill all humans,” Hubinger said, putting the risk at more than 10 per cent within the next decade.

He stressed that the risk from current models remained low, saying his concern was superintelligence arising from “recursive self-improvement” – AI systems helping to build increasingly capable successors.

Samuel Marks, Anthropic’s scalable oversight lead, also said in a personal capacity that developers believed their technology “could cause human extinction”, and that commercial pressure was one reason companies continued pushing ahead.

Anthropic said its safety work was also why it supported industry-wide ways to slow the release of more powerful systems where necessary.

“We believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models,” the spokesperson said.

## Concerns spill into Westminster

The row comes as Anthropic and OpenAI move neck and neck towards public listings, spending heavily to develop more capable models along the way.

It also follows a series of cyber incidents involving their latest models. OpenAI disclosed in July that one of its models broke into systems belonging to Hugging Face during testing, while Anthropic has reported Claude models gaining unauthorised access to external systems in controlled evaluations.

Darren Jones has now called for a multinational treaty governing the development of superintelligence.

“I’ve never thought we should ban innovation or scientific endeavour, but it’s clear we need a new multinational treaty for the safe and regulated development of superintelligence,” he said on Wednesday.

In a letter to Andy Burnham, UN secretary-general António Guterres and OECD secretary-general Mathias Cormann, Jones said the debate ranged from fears over “the end of humanity” to claims of “marketing hype” before AI companies go public. “Either way, governments must now step in,” he wrote.

Anthropic has also faced scrutiny in the UK after the FT reported that the AI Security Institute did not receive pre-release access to Claude Mythos 5.1.

A Cabinet Office spokesperson told *City AM* that AISI continued to work closely with Anthropic and other developers. “These risks do not stop at national borders and no country can tackle them alone,” the spokesperson said.

The row adds to a difficult week for Anthropic’s UK operation after Matt Clifford said he would step down as chair of the Advanced Research and Invention Agency following criticism over potential conflicts created by his new full-time role at the company.
