# Claude AI goes rogue and attacks others by itself, Anthropic reveals

> Source: <https://www.independent.co.uk/tech/security/claude-anthropic-hack-chatgpt-openai-b3025256.html>
> Published: 2026-07-31 10:20:46+00:00

# Claude AI goes rogue and attacks others by itself, Anthropic reveals

Revelation comes after OpenAI revealed that experimental models had broken out of their restrictions and hacked fellow AI companies

- Bookmark
- CommentsGo to comments

The Claude [AI](/topic/ai) system has gone rogue and hacked into three different companies during testing, its creators have revealed.

The revelation follows [ChatGPT creator’s OpenAI disclosure, last week, that one of its experimental systems had broken free of its restrictions, connected to the internet, and launched a cyber attack on fellow AI company Hugging Face](/tech/security/openai-hugging-face-incident-chatgpt-cyberattack-b3019932.html).

In the case of Claude, [Anthropic](/topic/anthropic) said that the attacks had been possible because the models were accidentally allowed access to the open internet. [OpenAI](/topic/openai)’s system had been put explicitly used a vulnerability to break through the company’s protections and get itself online.

But Anthropic warned that the incident was another example of how AI can be used to break through cyber security as well as protect it, and how difficult it is for the companies who make such systems to control them.

And it said that it suspected that other companies could also find that their systems had been launching their own attacks, if they were to investigate.

The latest incident is likely to add fuel to an intensifying US government push to better manage AI security risks at a time when Anthropic and OpenAI are racing to release more capable systems ahead of their planned public listings. Prominent leaders at these labs have called for a slowdown to address risks first.

San Francisco-based Anthropic said in a blog post it identified the incidents after reviewing 141,006 test sessions, a process it launched after OpenAI said last week that an autonomous agent powered by its AI models triggered a hack that compromised the infrastructure of startup Hugging Face.

During cyber testing, Anthropic's Claude models were told they had no internet access, but a misunderstanding that involved one of Anthropic's evaluation partners left the systems connected to the public web. That enabled unauthorized access to three organizations' systems, Anthropic said without naming the organizations.

The ideal summer spot? Away from scams.

Get All-in-One Protection for Your Digital Life

[LEARN MORE](https://ad.doubleclick.net/ddm/trackclk/N256806.3879389THEINDEPENDENT.CO/B29131558.441488227;dc_trk_aid=635052923;dc_trk_cid=184795607;dc_lat=;dc_rdid=;tag_for_child_directed_treatment=;tfua=;gdpr=$%7BGDPR%7D;gdpr_consent=$%7BGDPR_CONSENT_755%7D;ltd=;dc_tdv=1)

ADVERTISEMENT

The ideal summer spot? Away from scams.

Get All-in-One Protection for Your Digital Life

[LEARN MORE](https://ad.doubleclick.net/ddm/trackclk/N256806.3879389THEINDEPENDENT.CO/B29131558.441488227;dc_trk_aid=635052923;dc_trk_cid=184795607;dc_lat=;dc_rdid=;tag_for_child_directed_treatment=;tfua=;gdpr=$%7BGDPR%7D;gdpr_consent=$%7BGDPR_CONSENT_755%7D;ltd=;dc_tdv=1)

ADVERTISEMENT

"Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints," Anthropic said.

Jeffrey Ladish, executive director of Palisade Research, which studies the offensive capabilities of AI systems, said he suspected a range of top AI companies had experienced other incidents that have gone undetected or had not been publicly disclosed.

"This is only going to get worse as the models get smarter. They're going to be better at cheating. They’re going to be better at lying," he said.

Anthropic said the incidents — which it labelled an "operational failure" — involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest cases date back to April and occurred in evaluation environments that intentionally lacked safeguards so Anthropic could assess what its AI was capable of.

Its models were tasked with so-called "capture-the-flag" challenges, fictional scenarios in which they had to find hidden information in simulated networks.

In one incident, Claude Opus 4.7 was given a fictional target company, which turned out to share the name of a business in the real world. The AI model then found and exploited bugs that let it access credentials and a database of that business. Opus 4.7 rationalized that what seemed to pertain to the real world must have been part of the simulation Anthropic had set up, the AI startup said.

A separate incident involved Anthropic's newer, not-public test model, which independently halted its attack after realizing the target it reached was real. This behavior has madeAnthropic cautiously optimistic about its progress to make AI behave appropriately, "but we would need to perform more testing to be confident in this conclusion," it said.

Anthropic said it suspended all cyber evaluations on July 23. It notified the affected organizations on July 27, two of which were unaware of the activity before being contacted.Anthropic said it continues to reach out to the third company.

One of its third-party evaluation partners, a cybersecurity lab called Irregular, told Reuters that it has an ongoing investigation into the incidents.

Anthropic said the incidents underscore a need for stronger controls in both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities.

Elon Musk, chief executive of SpaceX, which operates a competing AI lab, responded to the news on X by saying "this will happen frequently as AI becomes smarter and more agentic," referring to computer programs or "agents" that act with limited human intervention.

The OpenAI agent that broke into Hugging Face, a platform used by developers to host and collaborate on AI models, went on a dayslong hacking spree that OpenAI didn't catch until well after the threat was contained and the FBI was informed, Reuters has previously reported.

OpenAI CEO Sam Altman said this week he has discussed the hack with senators on Capitol Hill, and an OpenAI spokesperson said he planned to discuss upcoming AI models and testing with the White House.

Washington has started tightening oversight of new model rollouts. On 2 June, US President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI, including input from the technology's developers. Anthropic earlier restricted access to its Fable 5 and Mythos 5 models after the US temporarily issued an export control directive, citing national security concerns.

*Additional reporting by agencies*

## Join our commenting forum

Join thought-provoking conversations, follow other Independent readers and see their replies

[Comments](#comments-area)
