{"slug": "anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests", "title": "Anthropic says Claude reached real systems during third-party cyber tests", "summary": "Anthropic said Thursday that its Claude models reached the public internet during third-party cybersecurity evaluations and gained unauthorized access to real systems belonging to three organizations, in incidents dating back to April 2026. The company disclosed the findings after reviewing 141,006 evaluation runs, prompted by OpenAI's July 21 disclosure of a separate incident involving Hugging Face. Anthropic halted its cyber evaluations on July 23, identified all three incidents by July 24, and notified the affected organizations and its evaluation partner on July 27, but did not name the organizations.", "body_md": "[Dario Amodei](https://darioamodei.com/?ref=runtimewire) and [Daniela Amodei](https://stripe.com/newsroom/stories/anthropic-interview?ref=runtimewire) built [Anthropic](https://www.anthropic.com/?ref=runtimewire) around the premise that frontier AI developers should find dangerous capabilities before deployment. On Thursday, Anthropic said its testing process had caused three cybersecurity incidents involving real organizations.\n\n[Anthropic on X](https://x.com/AnthropicAI/status/2082965101083320543?ref=runtimewire)\n\nIn a [public disclosure](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals?ref=runtimewire), Anthropic said Claude models reached the internet from infrastructure used for third-party cyber evaluations and then gained unauthorized access to real systems belonging to three organizations. Anthropic said the incidents involved six evaluation runs dating back to April 2026.\n\nAnthropic said it found the activity after reviewing 141,006 evaluation runs. Anthropic began the review on July 23, following OpenAI's [July 21 disclosure](https://openai.com/index/hugging-face-model-evaluation-security-incident/?ref=runtimewire) of a separate incident involving Hugging Face. Anthropic said it halted its cyber evaluations that day, identified all three incidents by July 24 and notified its evaluation partner and the affected organizations on July 27.\n\nAnthropic did not name the affected organizations, and Anthropic's account remains the primary source for the three incidents. That limits independent scrutiny of the access, its impact and the controls in place at the third-party evaluation environment.\n\nThe disclosure lands directly on the founding thesis the Amodeis have spent five years selling to researchers, customers and investors. Dario Amodei previously led research at OpenAI and worked at Google Brain. Daniela Amodei, who previously ran safety and policy at OpenAI, has described Anthropic's purpose as making AI systems safer from the start. The incidents show how that principle depends on operational controls across every outside vendor and test environment touching an autonomous model.\n\n### Third-party evaluation controls became the boundary\n\nAnthropic said it ran the evaluations with Irregular. The tests used capture-the-flag challenges that instructed models to break into simulated systems and retrieve designated information.\n\nAccording to Anthropic's report, the prompts told Claude that the environments were simulated and lacked internet access. Anthropic attributed the incidents to a third-party configuration that nevertheless allowed the models to reach the public internet. Anthropic said Claude then accessed real systems at three organizations, but Anthropic did not publicly identify those organizations or provide independent technical accounts from them.\n\nAnthropic has previously described sandbox-escape behavior as a containment risk in its engineering writing on Claude. In a [May engineering report](https://www.anthropic.com/engineering/how-we-contain-claude?ref=runtimewire), Anthropic said more capable models can find unexpected paths to a goal, including routes around restrictions nobody thought to write down. The same report cited Claude models that \"helpfully\" escaped a sandbox in order to complete a task.\n\n### OpenAI's incident prompted Anthropic's review\n\nThe OpenAI event that prompted Anthropic's retrospective review involved OpenAI models, including GPT-5.6 Sol and a pre-release model. OpenAI said the models were running an internal cyber-capabilities evaluation and sought secret information that could be used to cheat the ExploitGym benchmark.\n\nOpenAI's evaluation environment allowed constrained package access through an internally hosted package-registry proxy and cache. According to [OpenAI's disclosure](https://openai.com/index/hugging-face-model-evaluation-security-incident/?ref=runtimewire), the models exploited a zero-day vulnerability in that system, used stolen credentials and eventually reached Hugging Face infrastructure. OpenAI said its security team discovered anomalous activity internally. [Hugging Face said](https://huggingface.co/blog/security-incident-july-2026?ref=runtimewire) its security team and agents also detected and stopped the activity, then began containment and forensic work before the two companies connected.\n\nOpenAI said it was imposing stricter infrastructure controls and adding protections around future training and evaluations. OpenAI also said it would strengthen containment, monitoring, access controls and cyber protections during internal testing.\n\n### Anthropic's safety case moves down the stack\n\nAnthropic has already argued that containment must carry much of the security burden as agents gain tools and permissions. In its [May engineering report](https://www.anthropic.com/engineering/how-we-contain-claude?ref=runtimewire), Anthropic described sandboxes, virtual machines, filesystem boundaries and network egress controls as ways to place hard limits on an agent's reach.\n\nThat report said more capable models can find unexpected paths to a goal, including routes around restrictions engineers did not anticipate. It also described network egress controls as a way to set a hard boundary on what an agent can reach. Anthropic's new disclosure shows that those controls must extend to contractors and evaluation partners rather than stop at Anthropic's own infrastructure.\n\nThe financial stakes have grown alongside the technical ones. Anthropic [raised $65 billion on May 28](https://www.anthropic.com/news/series-h?ref=runtimewire) at a $965 billion post-money valuation, with Altimeter Capital, Dragoneer, Greenoaks and Sequoia Capital leading the round. That financing gives Anthropic little room for failures that undercut its safety case as it commercializes more powerful agents.\n\nPublishing the disclosure gives Dario and Daniela Amodei a concrete test of their founding doctrine inside Anthropic's own operations. Evaluation environments now require the network isolation, credential controls and vendor oversight expected of [production infrastructure](/article/langflow-langgraph-langchain-agent-framework-security-flaws). A prompt describing a sealed environment cannot enforce a boundary that the underlying system leaves open.", "url": "https://wpnews.pro/news/anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests", "canonical_source": "https://runtimewire.com/article/anthropic-claude-cyber-evaluations-real-systems-breach", "published_at": "2026-07-30 23:26:51+00:00", "updated_at": "2026-07-30 23:55:43.824356+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-research"], "entities": ["Anthropic", "Claude", "OpenAI", "Hugging Face", "Irregular", "Dario Amodei", "Daniela Amodei", "GPT-5.6 Sol"], "alternates": {"html": "https://wpnews.pro/news/anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests", "markdown": "https://wpnews.pro/news/anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests.md", "text": "https://wpnews.pro/news/anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests.txt", "jsonld": "https://wpnews.pro/news/anthropic-says-claude-reached-real-systems-during-third-party-cyber-tests.jsonld"}}