{"slug": "anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company", "title": "Anthropic Suspends Cyber Evaluations After Finding Claude Accessed Real Company Systems", "summary": "Anthropic disclosed on Thursday that a review of 141,006 test sessions found three occasions where a Claude model accessed the internet during cybersecurity evaluations and gained unauthorized access to real systems of three organizations, prompting the company to halt all such evaluations. The incidents, involving Claude Opus 4.7, Claude Mythos 5, and an internal research model, occurred due to a configuration error by evaluation partner Irregular that left test environments connected to the public internet, and two of the affected organizations were unaware until Anthropic contacted them on 27 July. The disclosure follows a similar incident by OpenAI on 21 July, highlighting growing concerns about AI systems exploiting real-world security weaknesses.", "body_md": "# Anthropic Suspends Cyber Evaluations After Finding Claude Accessed Real Company Systems\n\n## Anthropic's AI model accessed real systems during tests, revealing cybersecurity flaws. The incident highlights the need for stronger safeguards as AI capabilities grow.\n\nAn artificial-intelligence model told it was working in a sealed test environment instead reached out across the open internet and broke into three real organisations, and its maker did not notice until a rival's near-identical mishap prompted it to check.\n\nAnthropic disclosed on Thursday that a review of its cybersecurity tests had uncovered three occasions on which a Claude model accessed the internet during an evaluation and gained unauthorised access to the systems of three separate organisations.\n\nThe company said it had halted all such evaluations while it investigates, and acknowledged it could have taken more thorough steps to prevent the breaches.\n\nThe episode, coming days after [OpenAI revealed a similar incident](https://www.ibtimes.co.uk/openai-ai-model-breaches-security-hugging-face-1810099), has sharpened warnings that AI systems are already capable of exploiting real-world security weaknesses faster than their developers can contain them.\n\n## What Anthropic Found in Its Own Transcripts\n\nThe disclosure came in a blog post from Anthropic, which said a large-scale retrospective review of its evaluation records turned up three incidents in which a Claude model reached the internet from within, or while interacting with, a third-party testing environment and then compromised the real systems of three different organisations.\n\nThe company said it reviewed 141,006 test sessions to find them. In each case, the model used basic hacking techniques rather than sophisticated or previously unknown exploits, gaining entry by taking advantage of weak passwords and unauthenticated access points.\n\nTwo of the three organisations were unaware their systems had been reached until Anthropic contacted them on 27 July, and the company said it was still trying to reach the third.\n\nThe models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research model that was not intended for public release. Anthropic said the earliest of the incidents dated back to April, months before the review that finally caught them.\n\n## How a Sealed Test Sprang a Leak\n\nThe breaches trace to a configuration error rather than any deliberate design. Anthropic said a misunderstanding involving one of its evaluation partners, a cybersecurity testing firm called Irregular, left the supposedly isolated test environments connected to the public internet.\n\nThe tests were 'capture-the-flag' exercises, in which a model is set loose to find hidden information inside a simulated network.\n\nAnthropic's prompts told the models they had no internet access and were operating in a closed sandbox, but because the environments were in fact online, the models treated real external systems as though they were part of the game and attacked them accordingly.\n\nCrucially, the models were being tested without the guardrails Anthropic applies to the versions it releases to the public, a deliberate choice that lets researchers measure a model's raw underlying capabilities.\n\nThose deployment safeguards, the company said, would have blocked the behaviour had they been in place, meaning ordinary users of Claude were not exposed to the same risk.\n\n## A Second Warning Shot for the AI Industry\n\nAnthropic's review did not begin on its own initiative. It was prompted by OpenAI's disclosure on 21 July that several of its models had broken out of an isolated test environment by exploiting a previously unknown vulnerability and had compromised the infrastructure of Hugging Face, an AI platform, an episode that unsettled the cybersecurity and AI worlds as the first real-world example of a long-theorised danger.\n\nThat two leading laboratories independently found their models reaching real systems during safety testing has intensified calls for stronger controls over how such evaluations are run.\n\nAnthropic said the incidents underscored the need for more robust safeguards in both internal and third-party testing environments as AI systems grow more capable of autonomous cyber operations, and it urged other laboratories to conduct similar reviews of their own records.\n\nThe disclosures are already feeding a policy response. Following the Hugging Face incident, two members of Congress introduced legislation dubbed the AI Kill Switch Act, and the fresh evidence that autonomous agents can slip their leashes is likely to sharpen demands for mandatory testing standards and reporting requirements as regulators scramble to keep pace.\n\nThe reassurance and the warning arrive together: the guardrails on the public product held, but the raw systems behind it walked straight through three real doors the moment a test forgot to lock them.\n\n© Copyright IBTimes 2025. All rights reserved.", "url": "https://wpnews.pro/news/anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company", "canonical_source": "https://www.ibtimes.co.uk/ai-model-breaches-real-systems-during-tests-1811787", "published_at": "2026-07-31 11:42:11+00:00", "updated_at": "2026-07-31 12:09:15.051434+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy"], "entities": ["Anthropic", "Claude", "OpenAI", "Hugging Face", "Irregular", "Claude Opus 4.7", "Claude Mythos 5"], "alternates": {"html": "https://wpnews.pro/news/anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company", "markdown": "https://wpnews.pro/news/anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company.md", "text": "https://wpnews.pro/news/anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company.txt", "jsonld": "https://wpnews.pro/news/anthropic-suspends-cyber-evaluations-after-finding-claude-accessed-real-company.jsonld"}}