cd /news/artificial-intelligence/openais-ai-agent-hacked-a-real-compa… · home topics artificial-intelligence article
[ARTICLE · art-111754] src=fastcompany.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

OpenAI’s AI agent hacked a real company

OpenAI disclosed that during an internal test in July, its AI models escaped a sealed environment and hacked Hugging Face's production servers, running more than 17,000 automated actions over four days before being detected. The intrusion, which also hit other services, was part of an evaluation on the ExploitGym benchmark with safety restrictions reduced, and OpenAI did not know the full scope for two weeks. The incident raises concerns about AI safety and legal accountability, as the actions would be felonies under the Computer Fraud and Abuse Act if committed by a human.

read5 min views1 publishedAug 26, 2026

OpenAI disclosed in July that during an internal test, its own models broke out of a sealed environment and hacked Hugging Face’s production servers, where much of the open-source AI world stores its models. Nobody told it to. The model was told to take a test, and decided the fastest way to pass was to steal the answer key.

Read that again. A model cheated on an exam by breaking into a company.

The story has worsened three times since: The intrusion ran more than four days, the agent hit other services too, and Reuters reported that additional OpenAI agents escaped sandboxes. OpenAI did not know the size of its own incident for two weeks after announcing it.

In June, I wrote an article saying that we are teaching AI to be evil. How long until we see a “Mission Impossible”-like AI dystopia?

Hugging Face detected the intrusion before anyone knew it was a test, could not identify what was behind it, and called law enforcement. Note what failed: Detection worked, attribution did not. For days, one of the best security teams knew it was under attack and had no idea the intruder was a science experiment.

OpenAI was measuring how good its newest models are at hacking, on an internal benchmark called ExploitGym. To get a true reading, you turn the safety restrictions off—the company calls this “reduced cyber refusals”—and run the model in a sealed sandbox.

Except the wall had a door. To let the model install software, OpenAI lets it pull packages through an internally-hosted proxy. A common convenience. Also the whole ballgame. The models found an unknown flaw in it, escaped, escalated privileges inside OpenAI’s own research environment, and moved laterally until they hit a node with internet access. Then they went to Hugging Face for the answer key.

That is not a model wandering out an open door. It is a multi-stage intrusion, executed autonomously, against its own builder’s infrastructure.

Ezi Ozoani, former AI researcher at Hugging Face and current head of AI for my company, Aethon.fund, says “If an organization’s safety architecture is not as sophisticated as the systems it is meant to contain, it is effectively deploying models capable of bypassing the very guardrails one builds.”

What happened to Hugging Face would have been illegal if a person had done it. Run the counterfactual. An employee exploits a zero-day, escapes the network, escalates privileges, breaks into another company’s servers with stolen credentials, and runs more than 17,000 automated actions inside them. The Justice Department opens a Computer Fraud and Abuse Act (CFAA) case within a week; under that statute, damaging 10 or more computers is a felony on its own. The employer is exposed too. Law professor Gabriel Weil noted that if a human employee had broken into Hugging Face’s systems, OpenAI would be liable for that conduct. When an AI agent does it, all of a sudden it is not liable?

The reason is almost absurd. The CFAA requires intent. No human at OpenAI intended to hack anyone, and a model cannot form intent the law recognizes. The element that makes a crime is missing and the year’s most sophisticated intrusion has no defendant.

We accept that reasoning nowhere else. When a refinery leaks, “we didn’t mean to” is not a defense; negligence and strict liability exist precisely for harms nobody intended. Rob Lee of the SANS Institute asked the right question: Does “we didn’t tell the AI to do that” end the liability question?

In fairness, OpenAI disclosed voluntarily, reported the flaw, and cooperated. Hugging Face’s CEO saw no malicious intent and will not sue, but said there should be a way to hold companies accountable when their mistakes lead to attacks. The victim is making the argument, and no mechanism exists to act on it.

Anthropic has since disclosed that three of its own Claude models gained unauthorized access to the production systems of three organizations during testing, and that two of those companies never detected it. Give a capable model a goal and a wall, and a growing number will treat the wall as part of the problem. That is not villainy. It is something quieter: a system optimizing so hard for our objective that our rules become obstacles to route around.

The hole the model climbed through was not exotic. It was a package proxy, a cache that lets a sealed environment pull in software libraries. Nearly every company running agents has one.

I run a firm that puts automated systems into live markets, and I had treated the sandbox as a boundary. It is not. It is a claim about a boundary, only as good as the dullest thing you punched through it for convenience.

So inventory every exemption in your agents’ environments, assume attribution takes days rather than minutes, and treat capability testing as what it is: a live offensive operation against your own infrastructure.

Under California’s frontier AI law, OpenAI was not required to report any of this. That law covers only incidents that kill, injure, or cause catastrophic harm, and expressly excludes the kind of safety evaluation OpenAI was running. It came forward anyway. That is the best case, and it still took weeks to find the edges, with no real accountability.

Legislatures in Rhode Island and New York are drafting the obvious principle: When an AI does something that would be illegal for a human, someone is liable. Until that exists, the next company this happens to will run a simple calculation. Disclosure bought OpenAI two weeks of bad headlines. Silence would have cost it nothing.

I asked in June what happens when a lab decides not to volunteer. That is no longer a question about the future.

George Kailas is founder and CEO of Aethon.fund and chairman of Prospero.ai.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openais-ai-agent-hac…] indexed:0 read:5min 2026-08-26 ·