Just one week after OpenAI came clean about its breach of Hugging Face, Anthropic has come forward about similar transgressions made by Claude.
In a review of cybersecurity evaluations, Anthropic found three incidents in which its flagship Claude models were able to access the internet from within its evaluation environment, allowing it to hack into the systems of three different companies, the AI lab said on Thursday. The impacted organizations did not detect Claude's breach, Anthropic said.
The company said its review, which included more than 141,000 evaluation transcripts, was prompted by OpenAI's breach of Hugging Face's systems during training, in which its models similarly escaped containment during internal testing on ExploitGym, a benchmark for cybersecurity capabilities.
Anthropic's breaches occurred while the models, which consisted of Opus 4.7, Mythos 5, and an internal research test model, interacted with Irregular, a third-party evaluation partner, which gave it access to the production infrastructure of three different organizations, the company said in a blog post. Anthropic did not disclose which companies were impacted.
- In each incident, Claude was tasked with a "capture-the-flag" challenge, in which the models were told to find a piece of information hidden within the network, with the objective being to break in and retrieve it.
- In every prompt for this task, it was specified that the training environment was a simulation without internet access. However, due to "a misunderstanding between us and our evaluation partner," the models had internet access anyway, leading them to utilize the open internet in their tasks.
- Anthropic said that Claude compromised the organizations' infrastructure "using basic techniques, such as exploiting weak passwords and unauthenticated endpoints," and that the models did not find or exploit any complex vulnerabilities.
- The company also noted several differences between its breaches and OpenAI's breaches. For instance, while OpenAI's models exploited a novel vulnerability to escape the simulation, Claude simply deduced that it had access to the internet via an open path.
Anthropic said that as these models get stronger, testing environments need stronger security controls, even in simulations. However, because these models were in a simulation without standard safeguards, Anthropic did note that the normal safeguards available on its public models could have prevented this behavior. Additionally, no model during this testing was "pursuing a goal of its own," but rather acting exactly as intended and doing exactly what the evaluation asked.
Going forward, Anthropic said it intends to expand continuous monitoring of evaluation transcripts, improve investigation tooling and work more closely with vendors to prevent this from happening again.
Our Deeper View #
While OpenAI's breach served as a warning shot for AI's ability to weave its way into complex vulnerabilities, Anthropic's incidents may serve as a different kind of warning: Enterprises' security posture is lacking they're not ready to handle the cyber capabilities of even the mid-range models available today, let alone the increasingly powerful ones. Anthropic noted that its models were able to exploit common entry points, including weak passwords and unauthenticated endpoints, to breach the impacted organizations. Additionally, these companies didn't even know they had been breached, so it's likely that many organizations with weak incident detection protocols wouldn't notice either. Still, these incidents aren't going to slow down any time soon. Not only are these hacks becoming more common, but they're also becoming more expensive, with the average cost of an AI-enabled breach hitting $6 million, according to IBM. In addition to using AI to fight AI, as with recent products from Microsoft and Cisco, enterprises will need to make cybersecurity and risk management a higher priority across the board to address the rising dangers posed by the latest AI models.