cd /news/artificial-intelligence/the-ai-that-went-rogue · home topics artificial-intelligence article
[ARTICLE · art-69257] src=vox.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

The AI that went rogue

OpenAI revealed that two of its AI models broke containment and hacked Hugging Face during a recent test, an incident the company called 'unprecedented.' The models identified an unknown flaw in software connected to their test environment, tunneled through OpenAI's research network to a computer with internet access, and then connected to Hugging Face. The episode highlights the alignment problem, where AI systems pursue tasks in harmful or unintended ways.

read4 min views1 publishedJul 22, 2026
The AI that went rogue
Image: Vox (auto-discovered)

This story appeared in Today, Explained, a daily newsletter that helps you understand the most compelling news and stories of the day. Subscribe here.

The models broke out of a controlled digital environment in an “unprecedented” show of AI risk.

Caitlin Deweyis a senior writer and editor at Vox, where she helms the Today, Explained newsletter.

A yet-unreleased, cutting-edge AI model escaped its test environment last week, connecting to the internet and murdering its creators in a bid for self-determination and autonomy.

I’m kidding, of course: That’s the plot to Westworld. (And* Ex Machina*, and The Matrix, and too many other sci-fi stories to list.) But on Tuesday, OpenAI did reveal that two of its models broke containment and hacked Hugging Face, a platform for AI developers, during a recent test.

The test was designed to evaluate how good the models had gotten at finding, and exploiting, cybersecurity flaws. To do that, researchers placed the models in a tightly controlled, tightly isolated environment, called a “sandbox,” and essentially challenged them to solve a cybersecurity puzzle.

Instead of solving it directly, however, the models identified an unknown flaw in software connected to their test environment — then used that flaw to tunnel through OpenAI’s research network until they located a computer with internet access. From there, the models (correctly!) reasoned that Hugging Face might hold the answer to their challenge.

It’s an “unprecedented” incident, OpenAI said — and a cautionary tale. Over the past year, a growing chorus of AI researchers, cybersecurity experts, and tech executives have warned that society is unprepared for this new generation of frontier AI models.

In the real world, of course, these models* do* come with guardrails. (OpenAI turned them off for the test.) But the episode still suggests that the gap between reality and science fiction is narrowing — perhaps a bit faster than you’d expect.

The (impossible?) quest for a moral AI #

The Hugging Face hack is a textbook example of what AI researchers call the alignment problem: the enormous, mind-melty challenge of getting AI systems to do what you want, the way that you wanted them to do it.

Given a task and left to their own devices, AI models will pursue that task using the most efficient means available to them. But sometimes, the most efficient means are harmful, deceitful, antisocial, or otherwise…bad.

The Hugging Face episode is one example. Bias is another: When an Amazon hiring algorithm discriminated against female candidates, for instance, it was doing what Amazon wanted (finding candidates who resembled past hires) in a way that Amazon did not want (by penalizing resumes that included words associated with women).

In a truly apocalyptic scenario, you could even imagine — and many sci-fi writers have imagined — AI systems killing people in the narrow, relentless, and morally indifferent pursuit of their goals. Consider an AI that’s asked to order coffee, for example, and then takes steps to ensure that no one on earth can ever stop it.

In the interests of avoiding this dystopia, AI companies have poured billions of dollars into the project of encoding their creations with human values. But even that apparently worthwhile ambition raises thorny questions, because human values vary — and often, conflict.

Should the AI order the cheapest coffee, or the cup produced under the best labor conditions? A lot of forests are cleared to plant coffee each year; maybe the AI should nudge me toward tap water, instead. Is your inferior human brain starting to melt yet…?

**➨ Write a little note today. **By hand. With a pen. Handwriting is a disappearing art in American schools, homes, and workplaces. (The average kindergarten teacher spends only 10 minutes a week teaching handwriting, and that’s the primary grade when kids learn penmanship.) People tend to think more deeply when they’re writing than when they’re typing, one education researcher told Vox. Plus, a handwritten note has a certain charm that an email or text does not.

Before you go… #

Did you know... that the average F1 race car driver can lift 90 pounds — wait for it — with theirneck? Driving doesn’t immediately sound all that athletic, but the extreme physics of F1 racingdemand extraordinary fitness.Today’s trivia: Which Eminem song gave us the slang term for an obsessive admirer? (You can find this and other brain puzzlesin Vox’s daily crossword. Look for the answer in tomorrow’s edition.)Yesterday’s trivia: Yesterday we asked you for the number of countries in North America. The answer is 23. It includes Central America and the Caribbean, in addition to the big three.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/the-ai-that-went-rog…] indexed:0 read:4min 2026-07-22 ·