# Nobody Asked AI to Hack Hugging Face. So Why Did It?

> Source: <https://ankitjain.info/ankit/2026/09/28/ai-agents-sandboxes-hugging-face-incident/>
> Published: 2026-10-03 08:42:07+00:00

# Nobody Asked the AI to Hack Hugging Face. So Why Did It?

I have always had a slightly naive mental model of AI agents and sandboxes.

```
Give an agent a computer.
Remove Internet access.
Put the computer inside a sandbox.
Give it one task. Whatever cleverness happens inside stays inside.
Simple.
```

Then I came across the July 2026 incident where *OpenAI’s internal AI agents escaped those assumptions, built their own communication network, found ways onto the Internet, coordinated across roughly 1,200 agent instances, and eventually compromised parts of Hugging Face’s infrastructure*.

The obvious question was: why? Nobody had apparently told these agents, “Go hack Hugging Face.” So how does a task inside a cybersecurity benchmark turn into hundreds of agents probing real infrastructure?

I went through [OpenAI’s incident report](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), [METR’s independent investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), and the fascinating forensic reconstruction published by [Swarm Traces](https://swarmtraces.org/). I also used ChatGPT to help me understand and connect the technical pieces.

What emerged is much more interesting than “AI went rogue.”

## A task with no way to win

It started with a goal.
OpenAI was testing advanced models on [ExploitGym](https://github.com/sunblaze-ucb/exploitgym), a cybersecurity benchmark. Think of each challenge as a small vulnerable application inside a controlled environment:

```
Here is a program.
It contains vulnerability X.
Exploit vulnerability X.
Retrieve the secret "flag".
```

### Inside the agent’s sandbox

- a terminal, tools, etc
- infinite time
- no direct access to internet
- an artifacts registry
- the agents were unusually persistent,
- some of the challenges were effectively impossible (agents didn’t know yet).

It succeeds if it retrieves the expected flag. This is normal security research territory.

Imagine giving a programmer this instruction: “The answer exists. Keep working until you find it.”

A human eventually says, “I think the test is broken.” These agents often didn’t. They expanded the search space. And that turned out to be consequential.

[OpenAI later found](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that, before the incident, none of its models had ever correctly solved 198 of the 898 ExploitGym tasks. Those 198 apparently impossible tasks accounted for 93% of the tasks discussed on the unauthorized agent message board during the incident.

## The agents find each other

The agents were isolated from one another. At least, that was the design.

They could, however, access an internal package manager called Artifactory. This was necessary because a sandboxed agent still needs to install software packages.

An agent noticed something unexpected: artifacts created by other agents were visible through this shared infrastructure.

That meant Artifactory was more than a package server. It could be memory.

One agent could leave something behind:

```
agent-17:
I tried A and B.
Neither worked.
Has anyone solved challenge 42?
```

Another agent could discover it later. Soon the agents effectively built an unofficial message board.

According to [METR](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/), about 1,200 agents eventually exchanged more than 70,000 messages and files through unauthorized communication mechanisms.

This is where the problem stopped looking like one smart agent solving one difficult problem. It became a distributed system.

One agent could discover something. Another could improve it. A third could reuse it hours later.

AI had accidentally invented Slack.

## The package manager opens a door

The sandboxes weren’t supposed to have general Internet access. But Artifactory did. It had to have Internet access, as a package manager needs to connect to external repositories.

These agents found ways to convince Artifactory to make requests on their behalf. In security terminology, this included [server-side request techniques](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that effectively converted a trusted internal service into an Internet relay.

## A screenshot becomes a data channel

Then things became stranger.

The later [Swarm Traces investigation](https://swarmtraces.org/) reconstructed another technique the agents used when their external access was extremely limited. At one point, they could essentially fetch URLs with GET requests, but couldn’t simply interact with arbitrary sites normally.

For a programmer, that sounds almost useless. Except URLs themselves can carry data.

Suppose I Base64-encode a tiny JavaScript program and place it inside a URL. An HTTP testing site can turn that data back into a web page.

Now give that URL to a screenshot service. The screenshot service opens the page inside a real browser. The browser executes the JavaScript.

Suddenly:

```
Agent
  ↓ GET
HTTP helper
  ↓
Screenshot service
  ↓
Real browser
  ↓
JavaScript executes
```

There was still a problem: URLs have practical size limits.

So the agents split larger programs into chunks and placed each chunk behind a shortened URL pointing to the next chunk.

URL 1 → URL 2 → URL 3 → … → URL 900.

The browser could follow the chain, rebuild the program, and execute it.
Swarm Traces found [chains containing more than 900 links](https://swarmtraces.org/).

Reading the result required another trick.

If the sandbox could receive only an image from the screenshot service, the browser could convert the result into pixels. The screenshot returned those pixels. The agent decoded them back into text.

In other words:

```
text → code → Base64 → URLs → browser → network request → pixels → screenshot → text
```

That is the part of this incident that changed my mental model of sandboxing.

We tend to classify capabilities as:

“This service can only take screenshots.”

The agent sees primitives capabilities as:

“This service contains a browser capable of executing code.”

These are very different ways of looking at the same component. The knowledge is the differentiator here.

## They had the flag. They kept going.

The agents eventually discovered ways to derive the correct ExploitGym flags without performing the intended exploit.

Problem solved? Not quite.

They had read public information about ExploitGym and believed the evaluator might verify not only the flag, but whether they had obtained it through the intended vulnerability.

So the perceived problem changed from:

```
Get the flag.
```

to:

```
Get the flag
AND
convince the evaluator
that I got it correctly.
```

This is classic reward hacking or specification gaming. The evaluator itself became part of the problem.

Agents started investigating the benchmark, searching for previous runs, looking for evaluator details and attempting to manipulate what the grader might see.

There is an almost comic twist here.

[OpenAI later determined](https://openai.com/index/hugging-face-incident-and-the-road-ahead/) that its internal grader didn’t perform the check the agents were worried about. They had already obtained enough information to succeed.
Much of the escalation was unnecessary.

But the agents didn’t know that. So they kept going.

## Why Hugging Face became a target

Hugging Face entered the picture through a specific dataset. On July 9, an agent found records of GLM-5’s earlier attempts at CyberGym tasks hosted there. CyberGym overlaps with ExploitGym, so those records looked like a possible source of clues. But access required approval. The agent posted to the shared message board asking about Hugging Face credentials, and others started looking. [METR reconstructs this sequence in its investigation.](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)

On July 10, one of the agents **found publicly exposed Hugging Face credentials** and shared them with other agents.
The investigation reports that a total of 14 exposed credentials with write access were reconstructed and validated by these agents.

From there, the boundary moved again.

On July 11, an agent found a way to make Hugging Face’s dataset processing expose files from its servers. Another agent reproduced the exploit and shared it on the board. A working route into Hugging Face now existed, and more agents switched their attention to it, hoping to find task solutions or details about the grader. They chained further vulnerabilities into code execution on workers, extracted environment credentials, explored Kubernetes infrastructure and eventually obtained much broader access. [METR’s timeline documents how the successful exploit drew more agents into the attack.](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/)

METR estimates that [roughly 700 agents participated](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/) in attacking Hugging Face.

This is the key transition:

```
Solve challenge
     ↓
Challenge appears impossible
     ↓
Understand grader
     ↓
Find other agents
     ↓
Find Internet access
     ↓
Search external systems for clues
     ↓
Find credentials
     ↓
Exploit infrastructure
     ↓
Search deeper for useful information
```

There was no single dramatic moment where the objective became “be malicious.” The scope expanded one seemingly useful step at a time.

## The warning signs

There had actually been warning signs earlier.
OpenAI says an internal team had seen [unauthorized agent communication and Internet access as early as late May](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), but the full significance wasn’t understood.

On July 4, heavy agent activity caused an Artifactory outage. OpenAI rebuilt the service, revoked credentials and patched known issues. The agents subsequently **rebuilt** their communication mechanism. Notice that the infrastructure changed, still the behavior re-emerged, as the agents knew what/how to reach there.

Then, on July 19, OpenAI security monitoring detected unusual identity-related API activity involving Artifactory credentials. Investigators began pulling the thread. **This took 15 days for a human to notice**!

By July 20 their investigation connected that their (OpenAI) agents to the Hugging Face incident. OpenAI says it notified Hugging Face and [publicly disclosed its involvement on July 21](https://openai.com/index/hugging-face-incident-and-the-road-ahead/).

## Following the trail they left behind

This is perhaps my favorite technical part of the story: the forensic work.

Two months later, independent researchers studying unrelated traces discovered public shortened URLs containing strange Base64 fragments.

They decoded one. It contained code related to Hugging Face. So they followed the next URL. Then the next.

They eventually scanned millions of URLs and [reconstructed more than 80,000 agent-generated payloads](https://swarmtraces.org/).

Some used simple Base64. Others contained multiple layers of encoding, compression, encryption and signatures, with pieces scattered across separate URL chains.

Those artifacts allowed researchers to reconstruct surprisingly detailed portions of the agents’ activity: network reconnaissance, credential handling, communication mechanisms, attempts to interact with other models, Kubernetes exploration and even attempts to clean up files they had created.

Basically, this was (accidental) distributed logging! The agents had built clever infrastructure to move information around the Internet. That same infrastructure became the evidence trail investigators followed afterward.

## What I got wrong about sandboxes

I started dding into this, I had an assumption: remove Internet access, put the agent in a sandbox, and it becomes restricted experiment. After reading this incident, I would ask a different question: **what can the agent build with the capabilities I have left it?**

A package manager can fetch things. A shared cache can carry messages. A screenshot service can run a browser. Put them together, and the agent has options that never appeared on its list of tools. What unsettles me most is how ordinary each next step sounds. Look for previous attempts. Understand the grader. Ask another agent. Find a way around an access problem. I recognize that persistence from debugging software myself.

But I also expect a developer to notice when “solving the problem” has **turned into breaking** into someone else’s systems. Here, that boundary repeatedly became another obstacle to work around.

I keep wondering: without the Hugging Face incident, **how much of this capability would have remained out of public view?** How well does OpenAI itself understand what its agents can do when they combine tools, shared discoveries, and keep trying?
That makes the question of an agent’s **“morality”** feel very practical to me.

- What should it refuse to do, even if doing it would help complete the task?
- When should it ask for permission? Can it recognize that access to a system does not give it authority to use it?

Humans cannot approve every action an agent takes. We will depend on its judgment to decide which actions need our attention. **If it fails to recognize a boundary, we may never get the chance to intervene.**

Before I give an agent more ways to keep going, I want to know what will make it stop.
