To hear OpenAI tell it, instances of its most powerful AI models, including an unreleased one purportedly of immense power, recently got so focused on getting good scores on their evals that they escaped OpenAI’s testing sandbox, and
hacked the AI resource repository Hugging Face in an elaborate effort to cheat their way to the top.
AI skeptics have questions—this story is, after all, a big public relations coup for any AI company, since all AI companies benefit from the perception that their models are dangerously powerful. But whatever your takeaway may be on why this all happened, the many
[detailed reports](https://www.wsj.com/tech/ai/how-the-futuristic-hack-by-rogue-openai-models-unfolded-1657bcea) over
[the past few days](https://www.bloomberg.com/news/articles/2026-07-23/openai-models-lurked-in-hugging-face-system-for-hours-undetected) about
how it went down are genuinely spine-tingling, assuming your spine tingles at the thought of spooky new AI capabilities.
In fact, in one of the reports—
the Reuters one—before the hack, the models supposedly started acting like characters in a noir film. Specifically, they acted like Leonard Shelby from Christopher Nolan’s 2000 film
*Memento, *who lost his ability to make new memories and had to continue his quest for vengeance anew every time he snapped into awareness, guided by instructions he left to himself in the form of notes and tattoos.
As part of an account provided by three sources who spoke to Reuters, one of the AI agents undergoing testing supposedly, “left notes apparently for future versions of itself” that were directly at odds with what OpenAI wanted them to do.
Reuters says these instructional notes were buried in some secret place deep inside OpenAI’s internal “infrastructure,” and provided instructions on escaping from OpenAI’s sandbox environment. While creepy, Reuters says this specific devious behavior was not specifically linked to the Hugging Face hack.
Nonetheless, if this is real, it’s remarkable. Evidence that AI models are sentient is still laughable. But this would be a single instance of a model, which had itself figured out a way out of OpenAI’s maze, and then provided instructions for a future version of itself that would have no “memory” of such an escape in its context window to do the same thing.
You don’t have to believe AI models have subjective experience to fret that they could be gaining greater and greater capacity to cause harm, sentient or not.