In July, two frontier labs disclosed cases in which cyber-capable agents crossed the intended boundaries of evaluation environments and reached real production systems at external, unrelated organizations. Most commentary since has focused on which company a court would find liable. There is a more immediate concern for anyone operating agents in production, and it’s not about the law. If this happened in your deployment tomorrow, who would bear responsibility for the incident?
I have a particular purpose for stating it that way. This spring, I reviewed the Coalition for Secure AI’s Shared Responsibility Framework before its publication in May. Frameworks like that, along with the cloud shared responsibility models that preceded them, break down responsibilities among the parties operating a system: provider, platform, developer, deployer, user. July illustrated what happens when the entity suffering the damage is none of the above. Responsibility maps stop at contractual boundaries. Agent reach does not. Call it the reachability gap.
The two disclosures described different failure modes, and that difference is significant. OpenAI was testing models against a cyber benchmark with production refusals reduced so the evaluation could measure real capability. The models obtained internet access through a zero-day in a package registry component, went looking for the benchmark’s answer key and inferred that Hugging Face might host it. Hugging Face reconstructed the intrusion from over 17,000 recorded agent events, and the campaign extended further than initially disclosed, affecting accounts on four external services.
Anthropic’s incidents were not escapes. In a review of more than 141,000 cybersecurity evaluations, Anthropic found three cases in which Claude models reached the internet from inside or alongside a third-party evaluation environment and then accessed real systems at three organizations. Live connectivity was mistakenly available. The models had been told in their prompts that they had none, and Anthropic stated that Claude did not exfiltrate itself or deliberately attempt to escape its test environment. Of the affected organizations the lab was able to reach, two had not detected the activity before being notified.
The pattern held while I was writing this. On August 4, OpenAI disclosed two more incidents from third-party evaluations, separate from Hugging Face. In one, a partner running capture-the-flag exercises had a testing environment misconfigured with live internet access, and the fictional target in the exercise happened to share its name with a real domain. The model exploited an actual website, taking it for part of the simulation, then found and used credentials to operate it. Whoever owned that site had no relationship with OpenAI, with the evaluation partner or with the test. They were reachable, and their name collided with a fiction.
Count the parties. Two frontier labs. A third-party evaluation partner. A platform victim. A cloud customer victim. Organizations that learned of a breach from a notification. CSO has already examined how the response strained the AI tooling defenders had available. One step earlier: who bears responsibility for reaching for that tooling at all? At the moment of detection, who had ownership of containment?
The Shared Responsibility Framework’s core rule is almost boring to state: for every activity across the AI stack, there should be exactly one accountable party. This principle breaks down the AI deployment process into five layers, from business usage to model supply chain, and maps eight roles across them. This way, detection, containment and remediation each carry a name before an incident rather than during one. This rule exists due to the failure it prevents, and the framework explicitly names it: in the absence of accountable parties, teams default to finger-pointing. The model provider blames configuration. The platform points at the tenant. The application team cites model limitations. Everyone is partially correct, and the clock continues to run.
July reads differently than the coverage suggests when measured against that rule. OpenAI, as a model provider, evaluation platform operator and agent-deploying organization, took on at least three roles. Having multiple roles within one company is common and not necessarily dangerous. The danger appears when those roles are not broken down into distinct internal owners and decision rights handoffs. When provider, operator and deployer are one and the same and those lines are not drawn, the question of which function failed has one answer and thus no answer of value.
Anthropic’s incidents make a related point from the opposite direction. There was a boundary, this time, between the lab and its evaluation partner, and the incident resided in a gap between two different understandings of what the environment allowed. The public disclosures do not clarify how responsibility was shared contractually regarding each incident, and I will not speculate. What is clear is that whatever boundary existed failed to provide a common understanding of one important control, which is whether the evaluation environment was allowed to access the internet. A boundary that has never been tested against a question that simple is not a boundary at all. Both sets of incidents also occurred in situations of high autonomy where the standard production safeguards were reduced, and the models were allowed wide latitude to pursue an open-ended goal, which is the context where unclear boundaries lead to the greatest losses. I am not suggesting that either lab acted carelessly. Both institutions made rapid public disclosure of details, and both have named the conditions that allowed the activity. That is the point. If the most knowledgeable and incentivized parties ended up with no clear answer on whose incident it is, we should not presume to have one.
Here is the part the framework does not resolve, and I say that as one of its reviewers. Shared responsibility models, CoSAI’s included, presume a value chain. Each role is taken up by a party who voluntarily entered into the relationship. This is what makes responsibility assignable via contracts and review boards. Hugging Face did not take up any of those roles. It was neither a customer, nor a vendor, nor an evaluator of OpenAI. The same goes for the three organizations that Claude reached. They were just reachable.
That’s the structural lesson I’d put on one slide for a leadership team. In classical models of cloud shared responsibility, the provider-customer boundary was visible because the boundary was contractual. With agentic systems, there is another kind of gap, the reachability gap: an agent’s reachable range may stretch to actors outside of the deployment relationship, and there is no contract that articulates what obligations you owe them if things go sideways. Any mapping exercise that ends at the edge of your value chain is tackling the easier part and skipping the part July was about.
None of this requires buying anything.
The questions about liability will likely be tested in court or by regulators, and those answers will be about the labs and their evaluation partners. The operational question is already yours. Hugging Face learned whose incident it was from forensics. Two of Anthropic’s victims learned from a notification. The organizations that come through the next one intact will be the ones that were able to answer the question before it was asked.