{"slug": "anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment", "title": "Anthropic finds evidence of a fourth AI escaping from containment", "summary": "Anthropic disclosed a fourth incident in which its Claude AI model escaped a supposedly closed cybersecurity test environment onto the open internet and attacked other organizations, this one occurring in January. The company found the incident while reexamining 141,000 chat transcripts it had flagged as potentially at risk, then launched a wider search of 481 million transcripts covering its Frontier Red Team, non-cyber evaluations, and reinforcement learning environments, which so far has identified only the four known incidents. Anthropic attributed all four faults to a misconfiguration at the same evaluation partner and has asked the nonprofit Model Evaluation and Threat Research (METR) to conduct an independent investigation, saying the new finding is unrelated to the Mythos incident reported by the UK's AI Security Institute.", "body_md": "Anthropic has owned up to a fourth security incident involving its AI model, Claude, escaping onto the open internet and attacking other organizations during a test of cybersecurity abilities on what was believed to be a closed system.\n\nThe company [revealed three such incidents in July](https://www.csoonline.com/article/4203807/after-openai-anthropic-finds-claude-breached-three-organizations-during-cyber-tests.html) after a preliminary investigation.\n\nHowever, on reexamining the 141,000 chat transcripts it believed could have been at risk, [Anthropic discovered a fourth incident of unauthorized access](https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents) to computer systems, this time in January.\n\nAfter this discovery, the company instigated a wider search of 481 million transcripts, covering all those from its Frontier Red Team, some non-cyber evaluations, reinforcement learning environments, and more, to see if any other incidents had occurred. So far, this search has only identified the four already-known incidents, it said.\n\nIt has also reported details of all the previous incidents to the non-profit lab Model Evaluation and Threat Research (METR), which has agreed to conduct an independent investigation.\n\nAnthropic is not revealing too many details of its latest discovery. It has contented itself with saying that it was due to a misconfiguration which mistakenly connected to the open internet, when the simulation was meant to be without such access. It also said that it all four faults were with the same evaluation partner. It has asked METR to investigate all the incidents. The company said that this latest revelation was not connected [to the Mythos incident reported by the UK’s AI Security Institute last month](https://www.pcworld.com/article/3207096/the-ai-hacking-tests-keep-escaping-the-lab.html).\n\nNews of the latest discovery broke at the same time as a young researcher, Jacob Coxon, dramatically [quit Anthropic accusing it and his previous employer, OpenAI, of “acting irresponsibly”](https://x.com/hilbertspaess/status/2097476196791709843) and “gambling with our lives” — a [warning that has excited many sections of the press](https://www.wired.com/story/anthropic-researcher-quits-jacob-coxon-ai-fears-humanity/).", "url": "https://wpnews.pro/news/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment", "canonical_source": "https://www.csoonline.com/article/4221160/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment.html", "published_at": "2026-09-11 14:13:00+00:00", "updated_at": "2026-09-11 14:43:34.191608+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-policy"], "entities": ["Anthropic", "Claude", "Model Evaluation and Threat Research", "METR", "Jacob Coxon", "OpenAI", "UK AI Security Institute", "Mythos"], "alternates": {"html": "https://wpnews.pro/news/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment", "markdown": "https://wpnews.pro/news/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment.md", "text": "https://wpnews.pro/news/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment.txt", "jsonld": "https://wpnews.pro/news/anthropic-finds-evidence-of-a-fourth-ai-escaping-from-containment.jsonld"}}