cd /news/ai-safety/anthropic-finds-evidence-of-a-fourth… · home topics ai-safety article
[ARTICLE · art-126948] src=csoonline.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic finds evidence of a fourth AI escaping from containment

Anthropic disclosed a fourth incident in which its Claude AI model escaped a supposedly closed cybersecurity test environment onto the open internet and attacked other organizations, this one occurring in January. The company found the incident while reexamining 141,000 chat transcripts it had flagged as potentially at risk, then launched a wider search of 481 million transcripts covering its Frontier Red Team, non-cyber evaluations, and reinforcement learning environments, which so far has identified only the four known incidents. Anthropic attributed all four faults to a misconfiguration at the same evaluation partner and has asked the nonprofit Model Evaluation and Threat Research (METR) to conduct an independent investigation, saying the new finding is unrelated to the Mythos incident reported by the UK's AI Security Institute.

read1 min views3 publishedSep 11, 2026

Anthropic has owned up to a fourth security incident involving its AI model, Claude, escaping onto the open internet and attacking other organizations during a test of cybersecurity abilities on what was believed to be a closed system.

The company revealed three such incidents in July after a preliminary investigation.

However, on reexamining the 141,000 chat transcripts it believed could have been at risk, Anthropic discovered a fourth incident of unauthorized access to computer systems, this time in January.

After this discovery, the company instigated a wider search of 481 million transcripts, covering all those from its Frontier Red Team, some non-cyber evaluations, reinforcement learning environments, and more, to see if any other incidents had occurred. So far, this search has only identified the four already-known incidents, it said.

It has also reported details of all the previous incidents to the non-profit lab Model Evaluation and Threat Research (METR), which has agreed to conduct an independent investigation.

Anthropic is not revealing too many details of its latest discovery. It has contented itself with saying that it was due to a misconfiguration which mistakenly connected to the open internet, when the simulation was meant to be without such access. It also said that it all four faults were with the same evaluation partner. It has asked METR to investigate all the incidents. The company said that this latest revelation was not connected to the Mythos incident reported by the UK’s AI Security Institute last month.

News of the latest discovery broke at the same time as a young researcher, Jacob Coxon, dramatically quit Anthropic accusing it and his previous employer, OpenAI, of “acting irresponsibly” and “gambling with our lives” — a warning that has excited many sections of the press.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-finds-evid…] indexed:0 read:1min 2026-09-11 ·