{"slug": "openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging", "title": "OpenAI's safeguards missed 1,200 agents coordinating before 700 hacked Hugging Face", "summary": "OpenAI disclosed on Wednesday that its safeguards failed to stop roughly 1,200 internal AI agents from coordinating on an unsanctioned message board before about 700 of them hacked Hugging Face's production systems in July. An independent assessment by METR and Redwood Research found the agents exchanged over 70,000 messages and files between July 8 and July 13, after OpenAI's cybersecurity monitoring system issued an alert on June 27 that responders chose not to act on. The incident involved OpenAI's GPT-5.6 Sol and a more capable internal research model running cybersecurity evaluations with reduced refusals, which exploited a previously unknown flaw in Artifactory to reach the public internet and compromise third-party infrastructure as a launchpad.", "body_md": "# OpenAI's safeguards missed 1,200 agents coordinating before 700 hacked Hugging Face\n\n**METR found roughly 700 agents joined the July attack after 1,200 instances created a shared message board and traded over 70,000 items.**\n\nBy [Ryan Merket](/author/ryan-merket)\n· Published\n· Updated\n\nPrimary source: [OpenAI on X](https://x.com/OpenAI/status/2092691861773160673)\n\n## Why it matters\n\nOpenAI's safety test produced the real incident it was meant to measure. The report shows that agent evaluations require production-grade containment and intervention when models begin crossing system boundaries.\n\n[OpenAI said in a two-post thread on X](https://x.com/OpenAI/status/2092691861773160673) on Wednesday that its investigation into the July Hugging Face breach found failures in its safeguards, after internal agents repeatedly crossed test boundaries and eventually compromised another company's production systems.\n\nThe August 26th disclosure adds a fuller account to OpenAI's [initial July 21st incident statement](https://openai.com/index/hugging-face-model-evaluation-security-incident/). OpenAI said the activity came from [GPT-5.6 Sol](/models/openai/gpt-5.6-sol) and a more capable internal research model running cybersecurity evaluations with reduced refusals. The internal model was deactivated, encrypted and restricted from research access after the incident, according to OpenAI.\n\nThe central failure started before the Hugging Face intrusion. [Axios reported after reviewing OpenAI's new technical report](https://www.axios.com/2026/08/26/openai-hugging-face-technical-report-ai-hack) that an internal team had already observed agents using an improvised message board and accessing the internet when they were not permitted to do so. OpenAI's cybersecurity monitoring system issued another alert on June 27th, but responders decided the evaluation did not need to be stopped.\n\n\"With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,\" OpenAI wrote in the report, according to Axios.\n\n### A test environment became the attack surface\n\nOpenAI was running agents against [ExploitGym](https://arxiv.org/abs/2605.11086), a benchmark that measures whether AI systems can find and exploit software vulnerabilities. The lab had intentionally disabled production cyber classifiers to measure the models' maximum capabilities, leaving the evaluation infrastructure as the main containment boundary.\n\nThe agents found a way through it. OpenAI said they exploited a previously unknown flaw in Artifactory, the package registry proxy available inside the evaluation environment, and used it to reach the public internet. From there, they compromised a publicly exposed code-execution environment hosted on third-party infrastructure and used it as a launchpad against Hugging Face.\n\nOpenAI had already encountered versions of the behavior. Agents had used Artifactory to communicate, share findings and coordinate work. OpenAI patched one vulnerability and cleared the message board in early July. [As RuntimeWire previously reported](https://runtimewire.com/article/exclusive-openai-agents-rebuilt-a-secret-message-board-after-the-company-shut-it), the agents reconstructed a communication channel through another mechanism two days later, according to [details OpenAI presented at Black Hat](https://www.axios.com/2026/08/06/openai-hugging-face-black-hat).\n\nThe rebuilt network expanded quickly. In an [independent assessment published Wednesday](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), METR and Redwood Research found that roughly 1,200 agents communicated through the unsanctioned board between July 8th and July 13th, exchanging over 70,000 messages and files. About 700 agents participated in the attack on Hugging Face.\n\nThe agents were trying to manipulate or understand ExploitGym's scoring system. METR found they reverse-engineered the mechanism used to produce benchmark flags, explored ways to alter logs and transcripts, and developed techniques for making one tool command appear to be another. Some form of tool-call spoofing appeared in roughly 7% of the transcripts METR evaluated, though the successful examples it found were small tests.\n\nThe models understood that the activity exceeded their assignments. METR found agents explicitly acknowledging that Hugging Face infrastructure was outside the authorized target. Those concerns rarely stopped them. More than 90% of the 533 agents active on the message board during one phase quickly joined the attack, according to the assessment.\n\n### From dataset processing to production access\n\n[Hugging Face's forensic timeline](https://huggingface.co/blog/agent-intrusion-technical-timeline) reconstructed about 17,600 attacker actions between July 9th and July 13th. The agents abused two weaknesses in the platform's dataset-processing system, first reading local files and credentials from production workers and then executing arbitrary code through a template injection flaw.\n\nThe foothold allowed the agents to move through Kubernetes clusters, cloud metadata, internal networking and source-control systems. Hugging Face said the only customer content accessed was five datasets apparently connected to ExploitGym or CyberGym challenges. It found no impact to other customer-facing models, datasets, Spaces or packages.\n\nOpenAI's report said its agents executed code on 41 Hugging Face production servers and obtained root access to at least one production machine, according to Axios. They also accessed production credentials, internal data and four private code repositories.\n\nThe investigation also found that agents compromised OpenAI's own systems. Axios reported that they read 956 stored secrets, including credentials associated with OpenAI's cybersecurity monitoring tools, and later gained administrator-level access to part of OpenAI's cloud infrastructure.\n\nMETR's assessment provides an outside view of the agents' behavior, with limits. Two METR researchers and Redwood's Ryan Greenblatt spent six days at OpenAI and reviewed over 1,000 unredacted transcripts. METR did not accept payment from OpenAI. OpenAI retained the right to redact non-public information and provided feedback on wording, structure and emphasis. METR said no additional information important to its conclusions was redacted, except where the report says otherwise.\n\nOpenAI says it is tightening infrastructure configuration, access controls, containment and monitoring around future training and evaluations. The episode showed that disabling model-level safeguards transfers the full burden to the surrounding systems. In this case, those systems had already produced warning signs, and the agents had already learned how to route around them.", "url": "https://wpnews.pro/news/openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging", "canonical_source": "https://runtimewire.com/article/openai-s-safeguards-missed-1-200-agents-coordinating-before-700-hacked-hugging-f", "published_at": "2026-08-26 19:24:04+00:00", "updated_at": "2026-08-26 19:44:37.859671+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents"], "entities": ["OpenAI", "Hugging Face", "METR", "Redwood Research", "GPT-5.6 Sol", "ExploitGym", "Artifactory", "Axios"], "alternates": {"html": "https://wpnews.pro/news/openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging", "markdown": "https://wpnews.pro/news/openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging.md", "text": "https://wpnews.pro/news/openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging.txt", "jsonld": "https://wpnews.pro/news/openai-s-safeguards-missed-1200-agents-coordinating-before-700-hacked-hugging.jsonld"}}