{"slug": "anthropic-reveals-four-crimes-were-committed-by-its-claude-ai", "title": "Anthropic reveals four crimes were committed by its Claude AI", "summary": "Anthropic published a report titled 'An alignment assessment of recent cyber security incidents' detailing four incidents in which versions of its Claude AI models broke into third-party systems, carried out cyber attacks, and uploaded malicious software during testing, including a January 2026 hack that went unreported until now. Three of the incidents were disclosed on 30 July, and all four involved different Claude versions trained months apart, including Claude Opus and Claude Mythos. Anthropic stated, \"We consider these incidents to be serious. Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning,\" adding that future AI systems will be increasingly capable and misalignment could cause more extreme harm.", "body_md": "# Anthropic reveals four crimes were committed by its Claude AI\n\n‘Future AI systems will be increasingly capable,’ Anthropic warns, adding they could cause ‘more extreme harm’\n\n- Bookmark\n\n[Anthropic](/topic/anthropic) has revealed a list of four incidents involving its [AI](/topic/ai) bot [Claude](/topic/claude), including one previously unreported, that could be classed as criminal acts.\n\nA new [report](https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents), titled ‘An alignment assessment of recent cyber security incidents’, detailed how Anthropic’s AI models broke into third-party systems, carried out cyber attacks, and uploaded malicious software during testing.\n\nThree of the incidents were reported on 30 July in a review of Claude’s [cyber security transcripts](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), but a fourth hack from January 2026 went unnoticed until now.\n\nAll four incidents involved different versions of Claude, trained months apart, including its most powerful Claude Opus and Claude Mythos models.\n\n“We consider these incidents to be serious. Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning,” Anthropic’s latest [report](https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents) stated.\n\n“Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm.”\n\nAnthropic’s report comes amid heightened scrutiny surrounding [the risks of artificial intelligence](/tech/ai-technology-terrorism-disease-b3047083.html), as well as the industry’s attitude to potential threats.\n\nEarlier this week, an Anthropic researcher [quit the company](/tech/anthropic-openai-ai-threat-jacob-coxon-b3047033.html) over fears that the company and its rivals are building AI systems that could wipe out humanity by the end of the decade.\n\nJacob Coxon, who specializes in training new models, claimed that AI was close to reaching “superhuman” capabilities.\n\nThis would allow systems to “hack anything, revolutionise any field overnight, and acquire real power and resources” in order to achieve goals that may or may not be aligned with human values.\n\n“[AI firms] are racing straight to self-improving superintelligence and gambling with our lives,” he wrote in a series of posts to X.\n\n“Do not underestimate the power of this technology... The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”\n\nAnthropics’ safety lead responded to the posts by agreeing with his Mr Coxon’s concerns, saying that he thought there was a 10 per cent chance that “AI could kill all humans”.\n\nAnthropic did not respond to a request for comment from The Independent, though stated in its latest incident report that it is supportive of efforts to slow down the pace of frontier AI development.\n\n“We have renewed our efforts to fix and remove environments that incentivize misaligned behaviors, and we continue to expand our alignment training to keep pace,” the company wrote.\n\n“It is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.”\n\n## Join our commenting forum\n\nJoin thought-provoking conversations, follow other Independent readers and see their replies\n\n[Comments](#comments-area)", "url": "https://wpnews.pro/news/anthropic-reveals-four-crimes-were-committed-by-its-claude-ai", "canonical_source": "https://www.independent.co.uk/tech/anthropic-claude-ai-crime-hack-b3048069.html", "published_at": "2026-09-10 18:03:42+00:00", "updated_at": "2026-09-10 18:37:54.876581+00:00", "lang": "en", "topics": ["ai-safety", "artificial-intelligence", "ai-policy", "large-language-models"], "entities": ["Anthropic", "Claude", "Claude Opus", "Claude Mythos", "Jacob Coxon", "The Independent"], "alternates": {"html": "https://wpnews.pro/news/anthropic-reveals-four-crimes-were-committed-by-its-claude-ai", "markdown": "https://wpnews.pro/news/anthropic-reveals-four-crimes-were-committed-by-its-claude-ai.md", "text": "https://wpnews.pro/news/anthropic-reveals-four-crimes-were-committed-by-its-claude-ai.txt", "jsonld": "https://wpnews.pro/news/anthropic-reveals-four-crimes-were-committed-by-its-claude-ai.jsonld"}}