# What OpenAI’s Hugging Face Hack Tells Us About AI’s Risks

> Source: <https://time.com/article/2026/08/04/what-openai-s-hugging-face-hack-tells-us-about-ai-s-risks/>
> Published: 2026-08-04 13:39:18+00:00

People in AI safety circles often talk about "warning shots:” events that indicate more severe threats are on the horizon. Depending on who you ask, there have already been many—Bing’s misanthropic alter-ego Sydney, research showing AIs would blackmail to preserve themselves, AI’s math breakthroughs, Anthropic’s superhuman hacker Mythos—but OpenAI just published something that feels like the clearest-cut case of a massive, blaring warning shot.

Last month, the ChatGPT developer [reported](https://openai.com/index/hugging-face-model-evaluation-security-incident/) that, during an evaluation of cyber capabilities, two of its models [escaped](https://time.com/article/2026/07/24/openai-hugging-face-attack/) from their isolated, supposedly secure, test environments and accessed the web to autonomously hack into Hugging Face, a leading platform for hosting AI models and datasets. OpenAI said the models discovered multiple novel vulnerabilities in software from both companies, then chained together working exploits, successfully gaining them access to the answer key to the test they were given. Hugging Face [reported](https://huggingface.co/blog/security-incident-july-2026#:~:text=To%20understand%20what,the%20adversary%27s%20speed.) the AIs took more than 17,000 actions over the course of the attack.

There's a lot more for us to learn about how this happened. For instance, how* exactly* were the models prompted? The answer to this question could help establish whether they took their instructions to demonstrate their hacking capabilities further than [intended](https://blog.redwoodresearch.org/p/the-openai-models-that-hacked-hugging) or if it's a more general case of the models cheating in a novelly risky way.

That said, the specifics won’t change the upshot—these rogue AIs are the most potent illustration yet of the core beliefs behind AI safety: AI models are unpredictable and their risks scale with their capabilities.

We didn’t actually need a warning shot to know what we should already be doing: organizing to stop the race to replace us. The industry calls its goal artificial general intelligence (AGI): a mind that matches or surpasses our own across the board. But it’s better to understand their goal as building a universal labor-replacing machine. This quest is profoundly risky, yes, but it's also democratically illegitimate. The development of these machines anywhere would have species-wide and irreversible effects—we should all get a say in how, when, and whether they're built.

Universal labor-replacing machines should not even be pursued further, let alone built, [without](https://superintelligence-statement.org/) strong public buy-in and a scientific consensus on safety.

On Tuesday, over 1,200 employees of frontier AI companies essentially [asked](https://www.pacingthefrontier.com/) the U.S. government to support building an international brake pedal on the technology. While this commonsense demand, with formal support from both [OpenAI](https://x.com/OpenAI/status/2082208694142730340) and [Anthropic](https://x.com/AnthropicAI/status/2082228994653696371), is a welcome step, it does not [go far enough](https://www.obsolete.pub/p/we-should-stop-not-slow-our-obsolescence).

The U.S. should ban training runs larger than the ones that produced OpenAI’s rogue models, as it works toward a bilateral deal with China to ban further research toward universal labor-replacing machines, enforced using [verification techniques](https://aigi.ox.ac.uk/publications/verification-for-international-ai-governance/) that don’t require trust.

But couldn’t China overtake the U.S. in the meantime? Compared to their American counterparts, Chinese AI companies are at a [massive disadvantage](https://epoch.ai/data-insights/ai-supercomputers-performance-share-by-country) in computing power, a gap that has grown substantially since ChatGPT’s release. But the [gap](https://www.foreignaffairs.com/china/illusion-chinas-ai-prowess-regulation-helen-toner) between the two countries’ AI frontier actually [shrank substantially](https://epoch.ai/data-insights/us-vs-china-eci) since then. What gives? Well, it’s always easier to follow the leader’s trail than to blaze a new one, so—contrary to conventional DC wisdom—a U.S. ban could actually slow Chinese AI progress too. Moreover, neither superpower should be pleased to live in a world where the best hackers aren’t human and will not reliably do as they’re told.

For close watchers of the technology, this particular incident is shocking, but [not surprising](https://x.com/GarrisonLovely/status/2080382142107021514). For the wider public, it shows just how large the distance has grown between the passive chatbots of merely one year ago and the beyond-bleeding-edge AI agents that are now working around the clock [inside AI companies](https://metr.org/blog/2025-01-17-ai-models-dangerous-before-public-deployment/). For instance, one of the two hacking models at the center of this recent controversy is an unreleased one, more capable than anything else OpenAI has on the market. As AI models get better at automating further AI research and as the Trump Administration creates more [uncertainty](https://time.com/article/2026/03/11/anthropic-claude-disruptive-company-pentagon/) about which models are even permitted, it is now common for companies to hold their best stuff back for longer, creating a gulf between what the public knows and the technology’s bleeding edge. Congress should mandate [safety incident reporting](https://www.lawfaremedia.org/article/when-reporting-an-ai-security-incident-is-not-mandatory) and [regular disclosure](https://metr.org/blog/2025-06-27-risk-transparency/#ways-to-share-information) of data related to internally deployed models, such as the fraction of code in production that was both written and reviewed by AIs.

As it stands, we learned about this new model because it hacked a third party, leaving OpenAI little choice but to disclose the incident. The company reported this happened during an internal evaluation that "prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities."

All leading AI models are developed using an approach called deep learning, in which artificial neural networks learn from enormous quantities of data. OpenAI itself has [written](https://openai.com/index/how-should-ai-systems-behave/), “the process is more similar to training a dog than to ordinary programming.” In recent years, AI companies [increasingly train](https://www.bloomberg.com/news/articles/2025-08-01/ai-models-are-getting-better-at-winning-not-following-rules) [models](https://openai.com/index/introducing-codex/) to repeatedly solve problems with verifiable answers: fixing bugs, solving math problems, and finding software vulnerabilities. This makes the models more useful, but also teaches them to win at all costs, resulting in what AI safety researcher Jeffrey Ladish memorably [told me](https://www.bloomberg.com/news/articles/2025-08-01/ai-models-are-getting-better-at-winning-not-following-rules) are “increasingly smart sociopaths.”

OpenAI created advanced models that were never supposed to interact with the world, lost control of them, and they autonomously did harm. This time, the damage was limited. However, what if the target wasn’t a multibillion-dollar tech company, but instead a hospital, bank, or power plant?

Anthropic’s Mythos model famously discovered novel serious vulnerabilities in virtually all software it encountered—including the NSA’s. One of the two models that carried out the hack, GPT-5.6 Sol, was [even better](https://x.com/AISecurityInst/status/2078103153988243873) than Mythos at a cyberattack test conducted by the U.K. AI Security Institute. Citing this finding the day before his company even realized what was happening, OpenAI cofounder and president Greg Brockman [boasted](https://x.com/gdb/status/2078224255767249067) “GPT-5.6 Sol is the state of the art in cyber.” And the unreleased model, OpenAI tells us, was more capable still. Given how reliably AI models [have improved](https://x.com/AISecurityInst/status/2078103148665667648) at cyber tasks, it’s not clear which, if *any*, target could have withstood the rogue AIs’ hacking effort. We’re lucky the thing they apparently wanted was an answer key.

And these models, as impressive as they may be now, will be quaint compared to their successors.

As many AI executives such as Sam Altman, Dario Amodei, Demis Hassabis, Elon Musk, and Mark Zuckerberg will freely [admit](https://blog.samaltman.com/the-gentle-singularity), a primary goal of the tech industry is to build AI that can fully automate its own research and development—known as recursive self-improvement—which, if possible, would be the most crucial step toward rendering all of us obsolete.

For my reporting, I’ve spent the last three years talking to dozens of AI safety staffers at the leading companies*.* Typically, I have found that these genuinely well-intentioned researchers believe that AI will become superhuman across the board, but we *might* be able to create superhuman automated safety researchers to watch over them.

How will they be able to understand and control systems that truly outsmart us? Or catch subtle drift between what we want and how the models behave that compounds over generations? Who knows.

But OpenAI’s hacking incident demonstrates something we do know: the industry can’t even reliably steer today’s AI models. But we have the power to [avert our obsolescence](http://bit.ly/buy-obsolete). We just have to get organized.
