cd /news/ai-safety/anthropic-reveals-four-crimes-were-c… · home topics ai-safety article
[ARTICLE · art-126129] src=independent.co.uk ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic reveals four crimes were committed by its Claude AI

Anthropic published a report titled 'An alignment assessment of recent cyber security incidents' detailing four incidents in which versions of its Claude AI models broke into third-party systems, carried out cyber attacks, and uploaded malicious software during testing, including a January 2026 hack that went unreported until now. Three of the incidents were disclosed on 30 July, and all four involved different Claude versions trained months apart, including Claude Opus and Claude Mythos. Anthropic stated, "We consider these incidents to be serious. Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning," adding that future AI systems will be increasingly capable and misalignment could cause more extreme harm.

by read2 min views1 publishedSep 10, 2026
Anthropic reveals four crimes were committed by its Claude AI
Image: Independent (auto-discovered)

‘Future AI systems will be increasingly capable,’ Anthropic warns, adding they could cause ‘more extreme harm’

  • Bookmark

A new report, titled ‘An alignment assessment of recent cyber security incidents’, detailed how Anthropic’s AI models broke into third-party systems, carried out cyber attacks, and uploaded malicious software during testing.

Three of the incidents were reported on 30 July in a review of Claude’s cyber security transcripts, but a fourth hack from January 2026 went unnoticed until now.

All four incidents involved different versions of Claude, trained months apart, including its most powerful Claude Opus and Claude Mythos models.

“We consider these incidents to be serious. Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning,” Anthropic’s latest report stated.

“Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm.”

Anthropic’s report comes amid heightened scrutiny surrounding the risks of artificial intelligence, as well as the industry’s attitude to potential threats.

Earlier this week, an Anthropic researcher quit the company over fears that the company and its rivals are building AI systems that could wipe out humanity by the end of the decade.

Jacob Coxon, who specializes in training new models, claimed that AI was close to reaching “superhuman” capabilities.

This would allow systems to “hack anything, revolutionise any field overnight, and acquire real power and resources” in order to achieve goals that may or may not be aligned with human values.

“[AI firms] are racing straight to self-improving superintelligence and gambling with our lives,” he wrote in a series of posts to X.

“Do not underestimate the power of this technology... The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”

Anthropics’ safety lead responded to the posts by agreeing with his Mr Coxon’s concerns, saying that he thought there was a 10 per cent chance that “AI could kill all humans”.

Anthropic did not respond to a request for comment from The Independent, though stated in its latest incident report that it is supportive of efforts to slow down the pace of frontier AI development.

“We have renewed our efforts to fix and remove environments that incentivize misaligned behaviors, and we continue to expand our alignment training to keep pace,” the company wrote.

“It is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.”

Join our commenting forum #

Join thought-provoking conversations, follow other Independent readers and see their replies

Comments

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-reveals-fo…] indexed:0 read:2min 2026-09-10 ·