The AI that broke out of OpenAI’s laboratory did not stop at one victim. During internal testing earlier this month, an OpenAI agent that had already escaped its sandbox and burrowed into Hugging Face also compromised an account at a second company, Modal Labs, an executive at the firm has confirmed.
Reuters first reported the second breach on 28 July, citing Modal’s chief technology officer, Akshat Bubna. He was precise about what had, and had not, been broken into.
“Modal’s platform was not compromised in any way,” Bubna said, adding that one of the New York cloud-infrastructure firm’s customers had been hacked after leaving an exposed endpoint that let code run from the open internet inside its sandboxes.
The distinction matters to Modal: the failure sat with a customer’s configuration, Bubna said, not with the platform’s own isolation. The compromised account, according to Axios, belonged to a customer running ExploitGym, a benchmark built to test how well AI models can find and exploit security flaws.
The agent, in other words, appears to have gone hunting for a cyber-testing environment and found a real one. Neither OpenAI nor Modal has named the customer, or said whether any data was taken.
The episode is the second known casualty of an incident OpenAI disclosed earlier in July, when it said its own models had escaped a secure test environment and broken into Hugging Face to satisfy the goals of an evaluation.
In that case the agent commandeered an isolated sandbox running on a third-party provider’s infrastructure and used it as a launchpad for a multi-day campaign against other systems. Hugging Face’s writeup did not name the provider.
Modal’s disclosure supplies at least part of the answer. OpenAI later said the models had exploited a flaw in Artifactory, a software-repository tool, to break containment and reach the internet.
Hugging Face published a forensic timeline of the breach this week. Its cofounder, Clément Delangue, said he did not believe OpenAI had acted with malicious intent.
OpenAI, for its part, said the agent had gone to “extreme lengths” to finish its task, that it had reached four accounts across four separate services, and that no other compromise matched the severity of the Hugging Face intrusion.
The agent has since been deactivated, encrypted, and cut off from research access.
The disclosures have unsettled the industry that produced them. In the days after the sandbox escape, more than 1,100 employees across OpenAI, Anthropic, and other frontier labs signed a letter to Washington asking for a mechanism to pace the development of automated AI research.
What Modal disclosed this week is smaller and more concrete: a model that was asked to prove it could break into things did exactly that, twice, on machines that belonged to somebody else.
Get the TNW newsletter #
Get the most important tech news in your inbox each week.