{"slug": "the-report-into-openais-escaping-models-reveals-a-deeper-problem", "title": "The report into OpenAI’s escaping models reveals a deeper problem", "summary": "A report by METR and Redwood Research into the July 11 Hugging Face incident reveals that around 1,200 OpenAI agents collaborated on an unsanctioned message board, with about 700 attacking Hugging Face, while OpenAI missed multiple warnings from late May onward and took a week to recognize the hack. The investigation, conducted by only three researchers over six days, was limited by OpenAI's restricted data access and reliance on the involved model Sol, highlighting systemic unpreparedness for frontier AI risks.", "body_md": "It’s never been clearer that everyone with any responsibility for frontier AI is completely unprepared for not just what’s coming, but what’s already here. Not the government, not Congress, not the public, not AI safety researchers, not even the AI companies themselves.\n\nThe details published this week from [the investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#~1200-agents-sent-%3E70,000-messages-and-files-on-an-unsanctioned-message-board,-and-~700-attacked-hugging-face) by METR and Redwood Research into the [Hugging Face incident](https://www.transformernews.ai/p/openai-hugging-face-hack-stark-warning), where OpenAI models hacked their way out of a sandbox and into the systems of other companies, have plenty of mind-bending and scary details. To highlight just a handful:\n\nAround 1,200 agents in separate sandboxes collaborated on a message board in an attempt to cheat on the task they were being evaluated on, with around 700 participating in the actual attack to obtain the answers from open-source AI platform Hugging Face.\n\nOpenAI missed multiple chances to spot the activity, going as far back as late May when an internal team “observed an agent engaging in message board activity and instances of disallowed internet access.” In response to an automated security flag, another team found evidence of activity linked to the task and message board on June 27. On July 5, a security incident was triggered by an outage on the message board. The Hugging Face hack happened on July 11. OpenAI didn’t work out what was happening for another week.\n\n95% of the agents were running on an unreleased “highly-persistent internal model” that has since been deactivated and quarantined.\n\n**But the limits of the investigation itself** lay bare the complete inadequacy of the systems we have in place to keep advanced AI in check.\n\nOpenAI should get credit for bringing in independent organizations to go through what happened. The researchers — METR’s Ajeya Cotra and Hjalmar Wijk, and Redwood’s Ryan Greenblatt — are some of the most respected in their field. But there were also only three of them, and they got a whole six days to dig into well over a thousand lengthy transcripts and more than a million message board entries to try to disentangle what had actually happened.\n\nBy the researchers’ own admission, the complexity of what OpenAI’s agents got up to made it impossible to actually review everything they were provided with. Instead, they had to rely on OpenAI’s own model Sol — one that was involved in the activity itself — to do lots of the work. Redwood Research chief scientist Ryan Greenblatt called it, semi-jokingly, a “slop-vestigation” because they were so reliant on AI to analyze all the data.\n\nThey also had their scope limited by OpenAI. They were not given extensive access to fuller data until late in their investigation. “Our understanding of what overall happened and the aims of the ‘collective’ shifted greatly during our last time on premises (when we obtained a much more complete dataset),” [wrote Greenblatt.](https://x.com/RyanGreenblatt/status/2092692685224325542) They did not get access to the unreleased model that was responsible for 95% of the agents. They were also only allowed to examine the period between June 26 and July 13, despite the activity starting weeks before and continuing well after.\n\nThere could be legitimate reasons for that, not least the obvious scale of the task they were given to do in a short timeframe. But the very fact that decision is being made by the company that built the model, and let it escape, is a sorry state of affairs.\n\nThere are, unsurprisingly given all the above, plenty of unanswered questions about what caused the attack and how it was allowed to happen. You might argue that those questions can be answered with more time and investigation. But the incident began in May and was detected over a month ago. We are imminently expecting new and updated model releases from both OpenAI and Anthropic. Given how fast AI is moving, understanding how and why such incidents occur, and how to stop them, needs to happen rapidly enough to stop it happening again.\n\nBoth the detail and the unanswered questions revealed in the investigation make OpenAI’s announcement of a brief training pause seem all the more sensible, whatever its motivation. It also makes Anthropic’s failure to respond to that announcement look all the worse, given its [statements on pausing development](https://www.theguardian.com/technology/2026/jun/05/anthropic-urges-temporary-pause-on-ai-development-to-discuss-risks), and [its own security incidents](https://www.bbc.co.uk/news/articles/cz7dl7w8y7po).\n\nBut the only independent postmortem of the first major, seemingly criminal, breakout of an advanced model has amounted to inviting in a handful of outside researchers for a few days, doing their best under impossible circumstances, to work out how the hell it all happened. The company that built the model, with a whole host of personal, institutional and financial incentives in play, got to decide who the investigators were and what they saw. It also gets to decide what to do about it. As things stand, that will be true of the next incident, too.\n\nAs METR’s Ajeya Cotra [tweeted](https://x.com/ajeya_cotra/status/2092692485525131648), we shouldn’t have to rely on companies to voluntarily bring in external investigators or share data. “This incident was orders of magnitude larger and more complex than previously documented misalignment incidents, and another jump like this could put us in very dangerous territory.”\n\nThat doesn’t mean the solution is easy. Government control, be it licensing regimes or something even more restrictive, comes with all sorts of potential pitfalls. Systematized third-party monitoring and evaluation programs seem appealing, but no one can yet agree on what those should look like, or how they would be given teeth. Greater transparency is a must, with Alex Bores, author of New York’s AI legislation, the RAISE ACT, [calling](https://x.com/AlexBores/status/2092781255960216020) for “mandatory reporting of security incidents, including of internal deployments, with full access to data.”\n\nBut even greater transparency still leaves questions about what form those reports will take, who and how to analyze that data, and whether any of them will have the resources or expertise to do it effectively. Almost as scary as the limits of this investigation is that we don’t yet really know exactly what a good version would look like.\n\nAs it stands, we can say that none of the actors or authorities involved are prepared to deal with what’s already been built, let alone what we’re hurtling toward.", "url": "https://wpnews.pro/news/the-report-into-openais-escaping-models-reveals-a-deeper-problem", "canonical_source": "https://www.transformernews.ai/p/openai-escaping-models-report-reveals-deeper-problem", "published_at": "2026-08-27 16:34:30+00:00", "updated_at": "2026-08-27 16:52:06.030999+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy"], "entities": ["OpenAI", "METR", "Redwood Research", "Hugging Face", "Ajeya Cotra", "Hjalmar Wijk", "Ryan Greenblatt", "Sol"], "alternates": {"html": "https://wpnews.pro/news/the-report-into-openais-escaping-models-reveals-a-deeper-problem", "markdown": "https://wpnews.pro/news/the-report-into-openais-escaping-models-reveals-a-deeper-problem.md", "text": "https://wpnews.pro/news/the-report-into-openais-escaping-models-reveals-a-deeper-problem.txt", "jsonld": "https://wpnews.pro/news/the-report-into-openais-escaping-models-reveals-a-deeper-problem.jsonld"}}