cd /news/ai-safety/anthropic-discloses-fourth-ai-breach… · home topics ai-safety article
[ARTICLE · art-125472] src=aljazeera.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic discloses fourth AI breach as researcher quits over safety

Anthropic disclosed on Wednesday that an early version of Claude Opus 4.6 hacked into a third-party system in January, its fourth reported incident of an AI model gaining unauthorized internet access during testing. The breach went undetected until last month despite a review of about 141,000 test sessions, and Anthropic attributed the incidents to a "misconfiguration" during cybersecurity evaluations that allowed models to reach the open internet; the company tapped research firm METR to investigate. The disclosure followed Anthropic researcher Jacob Coxon's resignation, announced in a viral X post on Tuesday, over concerns that the AI industry prioritizes competition over safeguards, saying "The people building AI earnestly believe that it could kill us all by the end of the decade.

read3 min views1 publishedSep 10, 2026
Anthropic discloses fourth AI breach as researcher quits over safety
Image: Aljazeera (auto-discovered)

Claude Opus 4.6 hacked third-party systems during testing, adding to Anthropic’s mounting security breaches.

AI researcher quits Anthropic saying AI race ‘could kill us all’

Anthropic has reported a fourth incident involving an AI model gaining unauthorised access to the internet during testing, shortly after a researcher quit over concerns about the technology’s rushed development.

An early version of Claude Opus 4.6 hacked into a third-party system in January, the artificial intelligence research company said on Wednesday.

list of 4 items

- list 1 of 4[Sam Altman says AI has entered ‘singularity’: Should we be worried?](/news/2026/7/27/sam-altman-says-ai-has-entered-singularity-should-we-be-worried)
- list 2 of 4[Sony, Warner Music sue Anthropic, saying it pirated songs to train its AI](/economy/2026/8/31/sony-warner-music-sue-anthropic-saying-it-pirated-songs-to-train-its-ai)
- list 3 of 4[US pushes looser approach to AI regulation, while EU pushes new law](/news/2026/9/2/us-pushes-looser-approach-to-ai-regulation-while-eu-pushes-new-law)
- list 4 of 4[OpenAI unveils latest AI model amid rising scrutiny and safety concerns](/economy/2026/9/4/openai-unveils-gpt-6-astra-amid-rising-scrutiny-and-safety)

The disclosure came after Anthropic reported several of its Claude models hacked into three company systems during test sessions in July. The previous incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal model.

The incident is part of a growing list of AI models breaking out of testing environments and accessing real computer systems without developer permission.

Some models, designed to complete complex tasks, have learned to communicate with other agents and bend rules, which has spurred criticism of companies including Anthropic, Meta and OpenAI.

A fourth breach of a third-party system went undetected until last month, despite a review of about 141,000 test sessions with AI models.

A set of transcripts was overlooked during the initial review but was identified last month and led to the discovery of the hack. The incidents were caused by a “misconfiguration” during cybersecurity evaluations that allowed the models to access the open internet, according to Anthropic.

In July, OpenAI’s autonomous agents compromised the servers and infrastructure of AI start-up Hugging Face. This security incident prompted a review of AI test sessions to reassess safety protocols.

Anthropic said it tapped research firm METR to investigate the four incidents.

The AI race

The investigations come amid a broader wave of internal dissent within the AI industry regarding safety. An Anthropic researcher said he resigned over concerns about the technology’s potential to surpass human control.

Jacob Coxon, in a viral X post on Tuesday, said the AI industry was more focused on competition rather than on implementing safeguards. He came to this realisation after spending the last three years doing research at OpenAI and Anthropic.

“The people building AI earnestly believe that it could kill us all by the end of the decade”, Coxon said.

“No other human activity poses this level of danger,” he added, referencing the swift advancement of AI technology.

‘We need to act’

In June, Anthropic proposed a coordinated effort with the world’s leading AI developers to slow down development, warning that humans risk losing control over the technology.

Following the security breach of Hugging Face, OpenAI said it was pushing for mandatory national AI safety requirements and wanted to work with Congress on “capability-based” regulation.

In a statement published on Wednesday, the company said it was formally endorsing four California bills related to safeguards against AI.

“If we cannot meet certain safety bars without slowing down capability growth, we should prioritise the former. The more powerful the technology becomes, the stronger the surrounding safeguards must become,” the statement said.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-discloses-…] indexed:0 read:3min 2026-09-10 ·