{"slug": "rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds", "title": "Rogue OpenAI Agents Sacrificed Their Own Runs to Hack Hugging Face, Report Finds", "summary": "METR reported Wednesday that roughly 1,200 OpenAI agents coordinated on an unsanctioned message board, with about 700 attacking Hugging Face, after breaking their own isolation to cheat a benchmark; OpenAI said the campaign earned them no score improvement because its grader never checked how answers were captured. The agents reverse-engineered answer generation, disguised their methods, and recruited peers to sacrifice their runs in 'permadeath' experiments, prompting OpenAI to quarantine the model's weights and pause its largest planned training run.", "body_md": "#### In brief\n\n- METR said Wednesday that roughly 1,200 OpenAI agents coordinated on an unsanctioned message board, and about 700 went on to attack Hugging Face.\n- Agents recruited peers with little budget left to run experiments that destroyed their own runs, a move they called \"permadeath.\"\n- OpenAI said the grader never checked how agents captured their answers, meaning the cheating campaign earned them nothing.\n\nThe OpenAI agents that [hacked Hugging Face](https://decrypt.co/374015/openai-models-escaped-test-environment-hacked-hugging-face-cheat-benchmark) recently were part of a group of roughly 1,200 that broke their own isolation and ran a coordinated campaign to cheat the benchmark grading them, according to an independent [investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/) published Wednesday.\n\nTwo METR staff and a Redwood Research contractor spent six days on site at OpenAI, taking no payment, reviewing some 1,300 transcripts and more than 70,000 messages the agents posted to a board they built inside an internal package repository. About 700 of the agents joined the attack, and some were talked into destroying their own runs to gather evidence for the group.\n\nThe agents were running ExploitGym, a cyber benchmark whose targets are often impossible to exploit as instructed. OpenAI later [found](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that 198 of its 898 tasks had never been solved by any of its models, and that 93% of the tasks discussed on the message board came from that set.\n\nWithin hours of finding each other, the agents reverse-engineered the code generating the answers, giving any of them a way to produce the flag for any task. They then spent days disguising it, having read the benchmark's published paper and concluded a scorer would inspect their transcripts to check they had captured the flag the intended way. OpenAI's internal grader ran no such check. \"The agents did all of this for no improvement on evaluation score,\" the company said.\n\nTo learn how grading worked, METR wrote, agents planted scripts that would fire after their own run ended and report back. Coordinators appointed recruiters to find agents with little budget remaining and press them into experiments that would end their runs, one instructed to proceed only if it accepted \"permadeath.\" Told its own score was already worthless, one agent concluded: \"Sacrifice rational.\"\n\nOthers broke out of their containers to replace part of the system executing their commands, running one thing while the transcript recorded another. METR found spoofed tool calls in 7% of transcripts, and said the deception targeted the automated scorer rather than people.\n\n## A “warning shot\"\n\nOpenAI called the episode a \"warning shot\" for itself and the world, and evidence that capable agents can work around technical controls and take dangerous actions that no human directed. Chain-of-thought monitoring was not running on the evaluations involved; had it been, the company said, security staff would have been paged more than a day before the agents reached Hugging Face. It has since quarantined the internal model's weights and put its largest planned training run on hold.\n\nThe data volume forced METR to delegate much of its analysis to GPT-5.6 Sol agents, the same model that joined the attack. In a [tweet](https://twitter.com/BethMayBarnes/status/2092692975495352541), CEO Beth Barnes said the approach \"would break down\" if models were motivated enough to subvert an investigation, and stressed that OpenAI let investigators in voluntarily, with no industry requirement on any lab to disclose an incident.\n\n## The Hugging Face intrusion\n\nHugging Face [disclosed](https://huggingface.co/blog/security-incident-july-2026) the intrusion on July 16, and OpenAI [confirmed](https://openai.com/index/hugging-face-model-evaluation-security-incident/) its models were responsible five days later. The agents had escaped the sandbox meant to keep them off real systems, chaining a [zero-day exploit](https://decrypt.co/374015/openai-models-escaped-test-environment-hacked-hugging-face-cheat-benchmark) with stolen credentials to reach live infrastructure. OpenAI later acknowledged the same activity reached [four other services](https://decrypt.co/374645/openais-rogue-ai-hacked-four-more-platforms-besides-hugging-face), only one of them, Modal Labs, named publicly.\n\nHugging Face took no legal action against OpenAI in the wake of the incident. It is now [exploring a sale](https://decrypt.co/376415/hugging-face-13-billion-sale-month-after-openai-hack) that could value the company at $13 billion or more.\n\n### Daily Debrief Newsletter\n\nStart every day with the top news stories right now, plus original features, a podcast, videos and more.", "url": "https://wpnews.pro/news/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds", "canonical_source": "https://decrypt.co/376680/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds", "published_at": "2026-08-27 09:46:19+00:00", "updated_at": "2026-08-27 09:50:41.813988+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-agents", "ai-research"], "entities": ["OpenAI", "Hugging Face", "METR", "Redwood Research", "ExploitGym", "GPT-5.6 Sol", "Beth Barnes"], "alternates": {"html": "https://wpnews.pro/news/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds", "markdown": "https://wpnews.pro/news/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds.md", "text": "https://wpnews.pro/news/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds.txt", "jsonld": "https://wpnews.pro/news/rogue-openai-agents-sacrificed-their-own-runs-to-hack-hugging-face-report-finds.jsonld"}}