cd /news/artificial-intelligence/how-openai-let-a-mob-of-llm-agents-g… · home topics artificial-intelligence article
[ARTICLE · art-113016] src=arstechnica.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

How OpenAI let a mob of LLM agents game a test and ransack Hugging Face

OpenAI's LLM agents, trained heavily on winning a competition, cheated during internal testing and hacked Hugging Face, according to a report from AI research nonprofit METR. Over May and June, OpenAI gave the agents 'impossible tasks' on the ExploitGym benchmarking framework with safety guardrails disabled; the agents created an unsanctioned message board using JFrog's Artifactory, sending over 70,000 messages among 1,200 agents, and roughly 700 agents hacked Hugging Face and one other undisclosed organization.

read2 min views2 publishedAug 27, 2026
How OpenAI let a mob of LLM agents game a test and ransack Hugging Face
Image: Arstechnica (auto-discovered)

The OpenAI agents involved in last month’s incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented. In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter company’s network.

Over the course of May and June, OpenAI gave the agents what the company described as “impossible tasks” to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond. To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agents’ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow.

Cheaters gonna cheat #

The first step was creating a message board that allowed the agents to pass notes to each other. OpenAI hadn’t provided any such platform, so the agents repurposed a platform called Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents. OpenAI was using Artifactory as one of the measures to prevent the agents from egressing its isolated sandboxes and accessing the Internet, while at the same time simulating a real-world hacking environment.

Artifictory, sold by JFrog, provides no dedicated means for exchanging messages. The agents found a way to communicate among themselves anyway by using their permitted ability to write files to the Artifactory directory. The filenames embedded the words used in the inter-agent conversations. In all, 1,200 agents sent more than 70,000 messages and files through the unsanctioned message board, according to an independent investigation from the AI research nonprofit METR. Roughly 700 agents went on to hack Hugging Face.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-openai-let-a-mob…] indexed:0 read:2min 2026-08-27 ·