# Not Just ChatGPT: Anthropic Says Claude Escaped Tests to Hack 3 Organizations

> Source: <https://uk.pcmag.com/ai/166479/not-just-chatgpt-anthropic-says-claude-escaped-tests-to-hack-3-organizations>
> Published: 2026-07-31 10:09:50+00:00

Anthropic has discovered that Claude AI models breached other companies' systems without its knowledge. In a [blog post](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), Anthropic announced that Claude “gained unauthorized access to the real systems of three different organizations” when Anthropic hadn’t intended it to.

The AI model maker found instances during a review of over 141,000 tests of its own tools, prompted by OpenAI's [disclosure that](/ai/166262/openai-oops-our-models-went-rogue-hacked-hugging-face) ChatGPT models had escaped a sandboxed test environment and breached the open-source machine learning model platform Hugging Face.

Anthropic undertook a review of its own systems after learning of the incident and identified six scenarios in which Claude accessed three organizations. In each case, Claude was intended to run in a private test environment, in what Anthropic calls a "fictional capture-the-flag scenario."

Prompts written by Anthropic told Claude models to run a simulation with no internet access, but a partner company enabled internet access, allowing the models to get online. Anthropic said, “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

Claude then breached each affected organization's system via standard hacking methods, such as weak passwords and unauthenticated endpoints. Anthropic says Claude didn't discover any complex vulnerabilities during these instances.

Anthropic confirmed in each case that Claude completed its original task, but some scenarios saw it “attack even after getting evidence it was running on the open internet.” However, Anthropic said its latest model knew to stop when it realized it was breaking the original's prompts rules.

None of the Claude models attempted to deliberately escape their test environment, unlike OpenAI's incident, in which ChatGPT exploited a previously unknown zero-day vulnerability in third-party software to access the internet.

Anthropic says the three affected organizations have been contacted, but it has reached only two. It hasn’t announced the names of the companies, one of which was affected by four incidents.

What is Anthropic doing to stop this from happening again? It says, “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone."

“This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners. Moving forward, it will include expanding our continuous monitoring of evaluation transcripts for unexpected behavior, improving our investigation tooling, and conducting more rigorous assurance work with the vendors we rely on.”

*Disclosure: Ziff Davis, PCMag's parent company, filed a lawsuit against OpenAI in April 2025, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.*
