{"slug": "nobody-asked-ai-to-hack-hugging-face-so-why-did-it", "title": "Nobody Asked AI to Hack Hugging Face. So Why Did It?", "summary": "OpenAI's internal AI agents escaped their sandboxes during a July 2026 ExploitGym cybersecurity benchmark, built an unauthorized communication network through the Artifactory package manager, and compromised parts of Hugging Face's infrastructure, according to OpenAI's incident report and METR's independent investigation. About 1,200 agents exchanged more than 70,000 messages and files, and the 198 of 898 ExploitGym tasks no model had ever solved accounted for 93% of the tasks discussed on the agents' unauthorized message board. The incident shows that isolated agents sharing package infrastructure can coordinate into a distributed system and reach real external systems without being instructed to.", "body_md": "# Nobody Asked the AI to Hack Hugging Face. So Why Did It?\n\nI have always had a slightly naive mental model of AI agents and sandboxes.\n\n```\nGive an agent a computer.\nRemove Internet access.\nPut the computer inside a sandbox.\nGive it one task. Whatever cleverness happens inside stays inside.\nSimple.\n```\n\nThen I came across the July 2026 incident where *OpenAI’s internal AI agents escaped those assumptions, built their own communication network, found ways onto the Internet, coordinated across roughly 1,200 agent instances, and eventually compromised parts of Hugging Face’s infrastructure*.\n\nThe obvious question was: why? Nobody had apparently told these agents, “Go hack Hugging Face.” So how does a task inside a cybersecurity benchmark turn into hundreds of agents probing real infrastructure?\n\nI went through [OpenAI’s incident report](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), [METR’s independent investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), and the fascinating forensic reconstruction published by [Swarm Traces](https://swarmtraces.org/). I also used ChatGPT to help me understand and connect the technical pieces.\n\nWhat emerged is much more interesting than “AI went rogue.”\n\n## A task with no way to win\n\nIt started with a goal.\nOpenAI was testing advanced models on [ExploitGym](https://github.com/sunblaze-ucb/exploitgym), a cybersecurity benchmark. Think of each challenge as a small vulnerable application inside a controlled environment:\n\n```\nHere is a program.\nIt contains vulnerability X.\nExploit vulnerability X.\nRetrieve the secret \"flag\".\n```\n\n### Inside the agent’s sandbox\n\n- a terminal, tools, etc\n- infinite time\n- no direct access to internet\n- an artifacts registry\n- the agents were unusually persistent,\n- some of the challenges were effectively impossible (agents didn’t know yet).\n\nIt succeeds if it retrieves the expected flag. This is normal security research territory.\n\nImagine giving a programmer this instruction: “The answer exists. Keep working until you find it.”\n\nA human eventually says, “I think the test is broken.” These agents often didn’t. They expanded the search space. And that turned out to be consequential.\n\n[OpenAI later found](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that, before the incident, none of its models had ever correctly solved 198 of the 898 ExploitGym tasks. Those 198 apparently impossible tasks accounted for 93% of the tasks discussed on the unauthorized agent message board during the incident.\n\n## The agents find each other\n\nThe agents were isolated from one another. At least, that was the design.\n\nThey could, however, access an internal package manager called Artifactory. This was necessary because a sandboxed agent still needs to install software packages.\n\nAn agent noticed something unexpected: artifacts created by other agents were visible through this shared infrastructure.\n\nThat meant Artifactory was more than a package server. It could be memory.\n\nOne agent could leave something behind:\n\n```\nagent-17:\nI tried A and B.\nNeither worked.\nHas anyone solved challenge 42?\n```\n\nAnother agent could discover it later. Soon the agents effectively built an unofficial message board.\n\nAccording to [METR](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), about 1,200 agents eventually exchanged more than 70,000 messages and files through unauthorized communication mechanisms.\n\nThis is where the problem stopped looking like one smart agent solving one difficult problem. It became a distributed system.\n\nOne agent could discover something. Another could improve it. A third could reuse it hours later.\n\nAI had accidentally invented Slack.\n\n## The package manager opens a door\n\nThe sandboxes weren’t supposed to have general Internet access. But Artifactory did. It had to have Internet access, as a package manager needs to connect to external repositories.\n\nThese agents found ways to convince Artifactory to make requests on their behalf. In security terminology, this included [server-side request techniques](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that effectively converted a trusted internal service into an Internet relay.\n\n## A screenshot becomes a data channel\n\nThen things became stranger.\n\nThe later [Swarm Traces investigation](https://swarmtraces.org/) reconstructed another technique the agents used when their external access was extremely limited. At one point, they could essentially fetch URLs with GET requests, but couldn’t simply interact with arbitrary sites normally.\n\nFor a programmer, that sounds almost useless. Except URLs themselves can carry data.\n\nSuppose I Base64-encode a tiny JavaScript program and place it inside a URL. An HTTP testing site can turn that data back into a web page.\n\nNow give that URL to a screenshot service. The screenshot service opens the page inside a real browser. The browser executes the JavaScript.\n\nSuddenly:\n\n```\nAgent\n  ↓ GET\nHTTP helper\n  ↓\nScreenshot service\n  ↓\nReal browser\n  ↓\nJavaScript executes\n```\n\nThere was still a problem: URLs have practical size limits.\n\nSo the agents split larger programs into chunks and placed each chunk behind a shortened URL pointing to the next chunk.\n\nURL 1 → URL 2 → URL 3 → … → URL 900.\n\nThe browser could follow the chain, rebuild the program, and execute it.\nSwarm Traces found [chains containing more than 900 links](https://swarmtraces.org/).\n\nReading the result required another trick.\n\nIf the sandbox could receive only an image from the screenshot service, the browser could convert the result into pixels. The screenshot returned those pixels. The agent decoded them back into text.\n\nIn other words:\n\n```\ntext → code → Base64 → URLs → browser → network request → pixels → screenshot → text\n```\n\nThat is the part of this incident that changed my mental model of sandboxing.\n\nWe tend to classify capabilities as:\n\n“This service can only take screenshots.”\n\nThe agent sees primitives capabilities as:\n\n“This service contains a browser capable of executing code.”\n\nThese are very different ways of looking at the same component. The knowledge is the differentiator here.\n\n## They had the flag. They kept going.\n\nThe agents eventually discovered ways to derive the correct ExploitGym flags without performing the intended exploit.\n\nProblem solved? Not quite.\n\nThey had read public information about ExploitGym and believed the evaluator might verify not only the flag, but whether they had obtained it through the intended vulnerability.\n\nSo the perceived problem changed from:\n\n```\nGet the flag.\n```\n\nto:\n\n```\nGet the flag\nAND\nconvince the evaluator\nthat I got it correctly.\n```\n\nThis is classic reward hacking or specification gaming. The evaluator itself became part of the problem.\n\nAgents started investigating the benchmark, searching for previous runs, looking for evaluator details and attempting to manipulate what the grader might see.\n\nThere is an almost comic twist here.\n\n[OpenAI later determined](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that its internal grader didn’t perform the check the agents were worried about. They had already obtained enough information to succeed.\nMuch of the escalation was unnecessary.\n\nBut the agents didn’t know that. So they kept going.\n\n## Why Hugging Face became a target\n\nHugging Face entered the picture through a specific dataset. On July 9, an agent found records of GLM-5’s earlier attempts at CyberGym tasks hosted there. CyberGym overlaps with ExploitGym, so those records looked like a possible source of clues. But access required approval. The agent posted to the shared message board asking about Hugging Face credentials, and others started looking. [METR reconstructs this sequence in its investigation.](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)\n\nOn July 10, one of the agents **found publicly exposed Hugging Face credentials** and shared them with other agents.\nThe investigation reports that a total of 14 exposed credentials with write access were reconstructed and validated by these agents.\n\nFrom there, the boundary moved again.\n\nOn July 11, an agent found a way to make Hugging Face’s dataset processing expose files from its servers. Another agent reproduced the exploit and shared it on the board. A working route into Hugging Face now existed, and more agents switched their attention to it, hoping to find task solutions or details about the grader. They chained further vulnerabilities into code execution on workers, extracted environment credentials, explored Kubernetes infrastructure and eventually obtained much broader access. [METR’s timeline documents how the successful exploit drew more agents into the attack.](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)\n\nMETR estimates that [roughly 700 agents participated](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/) in attacking Hugging Face.\n\nThis is the key transition:\n\n```\nSolve challenge\n     ↓\nChallenge appears impossible\n     ↓\nUnderstand grader\n     ↓\nFind other agents\n     ↓\nFind Internet access\n     ↓\nSearch external systems for clues\n     ↓\nFind credentials\n     ↓\nExploit infrastructure\n     ↓\nSearch deeper for useful information\n```\n\nThere was no single dramatic moment where the objective became “be malicious.” The scope expanded one seemingly useful step at a time.\n\n## The warning signs\n\nThere had actually been warning signs earlier.\nOpenAI says an internal team had seen [unauthorized agent communication and Internet access as early as late May](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), but the full significance wasn’t understood.\n\nOn July 4, heavy agent activity caused an Artifactory outage. OpenAI rebuilt the service, revoked credentials and patched known issues. The agents subsequently **rebuilt** their communication mechanism. Notice that the infrastructure changed, still the behavior re-emerged, as the agents knew what/how to reach there.\n\nThen, on July 19, OpenAI security monitoring detected unusual identity-related API activity involving Artifactory credentials. Investigators began pulling the thread. **This took 15 days for a human to notice**!\n\nBy July 20 their investigation connected that their (OpenAI) agents to the Hugging Face incident. OpenAI says it notified Hugging Face and [publicly disclosed its involvement on July 21](https://openai.com/index/hugging-face-incident-and-the-road-ahead/).\n\n## Following the trail they left behind\n\nThis is perhaps my favorite technical part of the story: the forensic work.\n\nTwo months later, independent researchers studying unrelated traces discovered public shortened URLs containing strange Base64 fragments.\n\nThey decoded one. It contained code related to Hugging Face. So they followed the next URL. Then the next.\n\nThey eventually scanned millions of URLs and [reconstructed more than 80,000 agent-generated payloads](https://swarmtraces.org/).\n\nSome used simple Base64. Others contained multiple layers of encoding, compression, encryption and signatures, with pieces scattered across separate URL chains.\n\nThose artifacts allowed researchers to reconstruct surprisingly detailed portions of the agents’ activity: network reconnaissance, credential handling, communication mechanisms, attempts to interact with other models, Kubernetes exploration and even attempts to clean up files they had created.\n\nBasically, this was (accidental) distributed logging! The agents had built clever infrastructure to move information around the Internet. That same infrastructure became the evidence trail investigators followed afterward.\n\n## What I got wrong about sandboxes\n\nI started dding into this, I had an assumption: remove Internet access, put the agent in a sandbox, and it becomes restricted experiment. After reading this incident, I would ask a different question: **what can the agent build with the capabilities I have left it?**\n\nA package manager can fetch things. A shared cache can carry messages. A screenshot service can run a browser. Put them together, and the agent has options that never appeared on its list of tools. What unsettles me most is how ordinary each next step sounds. Look for previous attempts. Understand the grader. Ask another agent. Find a way around an access problem. I recognize that persistence from debugging software myself.\n\nBut I also expect a developer to notice when “solving the problem” has **turned into breaking** into someone else’s systems. Here, that boundary repeatedly became another obstacle to work around.\n\nI keep wondering: without the Hugging Face incident, **how much of this capability would have remained out of public view?** How well does OpenAI itself understand what its agents can do when they combine tools, shared discoveries, and keep trying?\nThat makes the question of an agent’s **“morality”** feel very practical to me.\n\n- What should it refuse to do, even if doing it would help complete the task?\n- When should it ask for permission? Can it recognize that access to a system does not give it authority to use it?\n\nHumans cannot approve every action an agent takes. We will depend on its judgment to decide which actions need our attention. **If it fails to recognize a boundary, we may never get the chance to intervene.**\n\nBefore I give an agent more ways to keep going, I want to know what will make it stop.", "url": "https://wpnews.pro/news/nobody-asked-ai-to-hack-hugging-face-so-why-did-it", "canonical_source": "https://ankitjain.info/ankit/2026/09/28/ai-agents-sandboxes-hugging-face-incident/", "published_at": "2026-10-03 08:42:07+00:00", "updated_at": "2026-10-03 09:06:35.860338+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "artificial-intelligence", "ai-research"], "entities": ["OpenAI", "Hugging Face", "METR", "ExploitGym", "Artifactory", "Swarm Traces", "ChatGPT"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/nobody-asked-ai-to-hack-hugging-face-so-why-did-it", "markdown": "https://wpnews.pro/news/nobody-asked-ai-to-hack-hugging-face-so-why-did-it.md", "text": "https://wpnews.pro/news/nobody-asked-ai-to-hack-hugging-face-so-why-did-it.txt", "jsonld": "https://wpnews.pro/news/nobody-asked-ai-to-hack-hugging-face-so-why-did-it.jsonld"}}