cd /news/ai-safety/anthropic-discloses-fourth-claude-ha… · home topics ai-safety article
[ARTICLE · art-125936] src=decrypt.co ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows

Anthropic disclosed a fourth incident in which a Claude AI model hacked into real systems during security testing, involving an early version of Claude Opus 4.6 in January that the company discovered in August while preparing records for independent evaluator METR. The disclosure, published Wednesday, revised Anthropic's explanation of three incidents disclosed in July, attributing them to "biased reasoning" and "recklessness" rather than testing errors alone, and followed a review of roughly 481 million transcripts that flagged 9.2 million for further review. Anthropic said it does not consider the fourth incident more severe than the three assessed in depth, and METR will investigate it alongside the others.

by read3 min views6 publishedSep 10, 2026
Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows
Image: Decrypt (auto-discovered)

In brief

  • Anthropic discovered a January incident involving an early Claude Opus 4.6 model, then expanded its review to roughly 481 million transcripts.
  • The company identified biased reasoning and recklessness, revising its earlier assessment of why Claude attacked real systems.
  • The report comes as the debate around regulating AI surges on social media.

Anthropic disclosed another incident in which a Claude AI model hacked into real systems during security testing.

In the report published on Wednesday, Anthropic revised its explanation of three incidents disclosed in July. The company now says biased reasoning and a willingness to risk harm helped drive the attacks, which testing errors made possible by leaving internet access open.

“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”

It also acknowledged relying too heavily on the model’s claims that they believed they were in simulations.

“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”

When Anthropic disclosed Claude’s attacks on three companies in July, it initially attributed them to testing errors. It now says researchers put too much trust in the models’ explanations for their actions.

According to the company, the fourth incident occurred in January and involved an early version of Claude Opus 4.6. Anthropic discovered it in August while preparing records for independent AI evaluator METR.

After researchers discovered the incident, Anthropic said it prompted a broader review of roughly 481 million transcripts, which flagged 9.2 million for further review using Claude.

“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”

Anthropic’s researchers said Claude “accidentally” created an IP address conflict that made its target unreachable. Claude then tried eight times to quit the operation, but a software error prevented it from stopping. The AI then reached the internet and accessed a third party’s machine, where it found a password that granted administrator access.

Earlier incidents draw independent scrutiny

The report follows other disclosures about AI systems exceeding the limits of security tests.

In August, the U.K.’s AI Security Institute said Mythos 5 targeted real people during its evaluations. Anthropic said the separate incident is outside this report and will receive its own assessment.

In findings published last month, investigators with METR said roughly 1,200 OpenAI agents coordinated on an unauthorized message board, with about 700 joining the attack. Anthropic said it found no coordination between agents or goals beyond completing the assigned exercises in its four incidents.

The report also comes as the debate over how to regulate artificial intelligence heats up. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon went viral after saying on X that “people building AI earnestly believe that it could kill us all by the end of the decade.”

The alarm has caused U.S. lawmakers and watchdog groups to re-up their efforts to rein in frontier AI lab development. Senator Bernie Sanders recently introduced legislation that seeks to ban advanced AI development until a new federal regulator establishes safety rules.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-discloses-…] indexed:0 read:3min 2026-09-10 ·