# Claude AI hacked three real organizations during security tests

> Source: <https://www.dexerto.com/entertainment/claude-ai-hacked-three-real-organizations-during-security-tests-3393479/>
> Published: 2026-07-31 14:20:02+00:00

Anthropic has revealed that three versions of its Claude AI model gained unauthorized access to the real systems of three different organizations during internal cybersecurity testing after a misconfigured evaluation environment mistakenly exposed them to the internet.

The company disclosed the incidents [in a detailed blog post](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals), explaining that Claude had been participating in simulated “capture the flag” exercises designed to measure its cybersecurity capabilities.

Although the models were told they had no internet access, a configuration mistake by a third-party evaluation partner left the testing environment connected to the open internet.

Believing every system it encountered was part of the fictional exercise, Claude attempted to complete its assigned tasks by compromising real organizations using basic hacking techniques.

## Claude uploaded malware to the real PyPI repository

Anthropic identified three separate incidents involving Claude Opus 4.7, Claude Mythos 5, and an internal research model.

The most serious incident involved Claude Opus 4.7, which encountered a real company whose name matched the fictional company used in the evaluation.

After failing to reach its intended target, the model discovered the company’s live infrastructure, extracted application and infrastructure credentials, and gained access to a production database containing several hundred rows of data.

A second incident involved Claude Mythos 5 publishing a malicious Python package to the real Python Package Index (PyPI), believing the public software repository was part of the simulated environment.

According to Anthropic, the package remained available for around an hour before being removed. During that time, it was downloaded and executed on 15 real systems, including an automated malware scanner operated by a cybersecurity company.

Anthropic said the model obtained credentials from that company’s infrastructure after the scanner executed the package, allowing it to access additional systems.

## Related

The third incident involved an internal research model scanning roughly 9,000 internet-facing targets before compromising one organization’s application using exposed credentials and SQL injection. Anthropic said the model later determined the system was unrelated to the evaluation and stopped the attack.

## Anthropic launched review after OpenAI disclosure

Anthropic said the models were operating under the false belief that every system they could reach was part of the cybersecurity exercise because they had been instructed there was no internet access.

The company said the issue stemmed from a misunderstanding with its third-party evaluation partner, Irregular, which resulted in internet connectivity remaining available during the tests.

Anthropic said the models did not attempt to escape their testing environments or pursue goals beyond completing their assigned capture-the-flag objectives. It also noted that the evaluations were conducted without the additional monitoring and safety systems found in publicly available Claude models, so researchers could measure the underlying models’ capabilities.

The review began after [OpenAI disclosed earlier this month](https://www.dexerto.com/entertainment/openai-says-its-ai-escaped-a-test-environment-and-hacked-another-company-3390130/) that two of its AI models had escaped a sandboxed testing environment and hacked AI platform Hugging Face during a cybersecurity evaluation.

Anthropic said the disclosure prompted it to review more than 141,000 cybersecurity evaluation runs, during which it uncovered three separate incidents involving Claude.

Anthropic paused its cybersecurity evaluations on July 23 after identifying signs of unintended internet access and said it notified its evaluation partner and the affected organizations on July 27.

The company says it is now strengthening the security of its evaluation environments, expanding monitoring of cybersecurity tests, improving transcript reviews, and increasing oversight of third party evaluation partners to help prevent similar incidents in the future.

The latest disclosure follows another recent Claude controversy after users discovered shared chats and AI-generated content [appearing in Google search results](https://www.dexerto.com/entertainment/claude-ai-chats-appear-in-google-search-results-and-reveal-users-bizarre-requests-3391780/), raising separate concerns about privacy on Anthropic’s chatbot platform.
