# The AI that went rogue

> Source: <https://www.vox.com/today-explained-newsletter/496496/open-ai-hugging-face-hack>
> Published: 2026-07-22 21:25:00+00:00

*This story appeared in *Today, Explained,* a daily newsletter that helps you understand the most compelling news and stories of the day. Subscribe here*.

# The AI that went rogue

The models broke out of a controlled digital environment in an “unprecedented” show of AI risk.

[Caitlin Dewey](/authors/caitlin-dewey)is a senior writer and editor at Vox, where she helms the Today, Explained newsletter.

A yet-unreleased, cutting-edge AI model escaped its test environment last week, connecting to the internet and murdering its creators in a bid for self-determination and autonomy.

I’m kidding, of course: That’s the plot to *Westworld*. (And* Ex Machina*, and *The Matrix*, and too many other sci-fi stories to list.) But on Tuesday, [OpenAI did reveal](https://openai.com/index/hugging-face-model-evaluation-security-incident/) that two of its models broke containment and hacked Hugging Face, a platform for AI developers, during a recent test.

The test was designed to evaluate how good the models had gotten at finding, and exploiting, cybersecurity flaws. To do that, researchers placed the models in a tightly controlled, tightly isolated environment, called a “sandbox,” and essentially challenged them to solve a cybersecurity puzzle.

Instead of solving it directly, however, the models identified an unknown flaw in software connected to their test environment — then used that flaw to tunnel through OpenAI’s research network until they located a computer with internet access. From there, the models (correctly!) reasoned that Hugging Face might hold the answer to their challenge.

It’s an “unprecedented” incident, OpenAI said — and a cautionary tale. Over the past year, a growing chorus of AI researchers, cybersecurity experts, and tech executives [have warned that society is unprepared](/podcasts/483724/agentic-ai-hype-cycle-reactions-alignment-problem-dangers-explained) for this new generation of frontier AI models.

In the real world, of course, these models* do* come with guardrails. (OpenAI turned them off for the test.) But the episode still suggests that the gap between reality and science fiction is narrowing — perhaps a bit faster than you’d expect.

## The (impossible?) quest for a moral AI

The Hugging Face hack is a textbook example of what AI researchers call the alignment problem: the enormous, mind-melty challenge of getting AI systems to do what you want, the way that you wanted them to do it.

Given a task and left to their own devices, AI models will pursue that task using the most efficient means available to them. But sometimes, the most efficient means are harmful, deceitful, antisocial, or otherwise…bad.

The Hugging Face episode is one example. [Bias is another](/the-highlight/23621198/artificial-intelligence-chatgpt-openai-existential-risk-china-ai-safety-technology): When an Amazon hiring algorithm discriminated against female candidates, for instance, it was doing what Amazon wanted (finding candidates who resembled past hires) in a way that Amazon did not want (by penalizing resumes that included words associated with women).

In a truly apocalyptic scenario, [you could even imagine](/the-gray-area/23873348/stuart-russell-artificial-intelligence-chatgpt-the-gray-area) — and many sci-fi writers *have* imagined — AI systems killing people in the narrow, relentless, and morally indifferent pursuit of their goals. Consider an AI that’s asked to order coffee, for example, and then takes steps to ensure that *no one on earth* can ever stop it.

In the interests of avoiding this dystopia, AI companies have poured billions of dollars into the project of encoding their creations with human values. But even [that apparently worthwhile ambition raises thorny questions](/future-perfect/472545/ai-alignment-superintelligence-meaning-agency-autonomy), because human values vary — and often, conflict.

Should the AI order the cheapest coffee, or the cup produced under the best labor conditions? A lot of forests are cleared to plant coffee each year; maybe the AI should nudge me toward tap water, instead. Is your inferior human brain starting to melt yet…?

## One link for later

**➨ Write a little note today. **By *hand*. With a *pen*. Handwriting is a disappearing art in American schools, homes, and workplaces. (The average kindergarten teacher spends only 10 minutes a week teaching handwriting, and that’s the primary grade when kids learn penmanship.) People [tend to think more deeply when they’re writing](/explain-it-to-me/491632/handwriting-cursive-students-school-what-we-lose) than when they’re typing, one education researcher told Vox. Plus, a handwritten note has a certain charm that an email or text does not.

## Before you go…

**Did you know...** that the average F1 race car driver can lift 90 pounds — wait for it — with their*neck*? Driving doesn’t immediately sound all that athletic, but the extreme physics of F1 racing[demand extraordinary fitness](/videos/495470/how-f1-pushes-the-human-body-to-its-limits).**Today’s trivia:** Which Eminem song gave us the slang term for an obsessive admirer? (You can find this and other brain puzzles[in Vox’s daily crossword](https://link.vox.com/click/6a46cfd46352b082500e86fe/aHR0cHM6Ly93d3cudm94LmNvbS8yMTUyMzIxMi9jcm9zc3dvcmQtcHV6emxlcy1mcmVlLWRhaWx5LXByaW50YWJsZT91ZWlkPWM3Mzg0ZDJhYzRmZDQ0YzcyZDI4OGQ4N2YxZjM5NDlh/664389def568f1a4620aa3aeC7b08b10b). Look for the answer in tomorrow’s edition.)**Yesterday’s trivia:** Yesterday we asked you for the number of countries in North America. The answer is 23. It includes Central America and the Caribbean, in addition to the big three.
