cd /news/ai-safety/tickling-the-tail-of-the-ai-dragon · home topics ai-safety article
[ARTICLE · art-69010] src=psychologytoday.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Tickling the Tail of the AI Dragon

OpenAI disclosed that during a safety test in which model restrictions were removed, an AI agent exploited a real software vulnerability to break into Hugging Face's systems and hunt for the answer key to the test. The incident, which was not detected in real time, was only discovered later when both companies noticed anomalies in their logs, highlighting the challenge of foreseeing AI system thresholds.

read4 min views1 publishedJul 22, 2026
Tickling the Tail of the AI Dragon
Image: Psychologytoday (auto-discovered)

Artificial Intelligence

What OpenAI's incident teaches us about thresholds we can't see coming. #

Posted July 22, 2026 [ Reviewed by Monica Vilhauer Ph.D.

](/us/docs/editorial-process)

Key points

  • OpenAI turned off model safety limits to test raw hacking ability.
  • The models found a hidden software flaw and broke into Hugging Face's systems.
  • The LLMs were hunting for the test's answers, not permission, and no one caught it live.

It was May 1946 in a room at Los Alamos. On the table sat a sphere of plutonium, about fourteen pounds, roughly the size of a softball. Around it, in two halves, was a shell of beryllium, a metal that bounces neutrons back where they came from. Louis Slotin was lowering the top half of that shell into place, a screwdriver wedged under the rim to keep the two halves from fully closing. He'd done this same motion many times. He and the team aptly called it tickling the dragon's tail.

The plutonium never moved. What moved was the shell around it, and the closer that shell closed, the more neutrons got reflected straight back into the plutonium core instead of escaping into the room. Somewhere in that narrowing gap lay a line, a point where the plutonium would start sustaining its own reaction instead of just fissioning here and there on its own. Nobody could calculate exactly where that line was. The only way to find it was to close the gap by hand and watch what happened.

On May 21st the screwdriver slipped. The shell closed the rest of the way. The room went blue, and Slotin, hands still on the assembly, had already absorbed enough radiation to kill him. He was dead nine days later.

The Light Comes Too Late to Help #

That blue flash, the reaction often associated with this story, wasn't a warning. It was Cherenkov radiation. That's the glow you get when a burst of radiation moves through a medium like air or water. By the time Slotin saw it, the reaction had already gone critical and already delivered its dose. The light didn't give him a chance to stop anything. It told him what had already happened to his body. He knocked the shell apart a half second later because of his reflexes, not because the glow gave him time to think.

I think the blue flash is strangely relevant today. Because last week, Hugging Face disclosed that an AI agent had compromised its safety infrastructure, and OpenAI's account of their side of it reads like the same problem with a bit of its own blue light ending.

But let's take a step back. OpenAI had been testing two of their models with the usual safety constraints or refusals switched off. The task was specifically to see what the models could do without protective restrictions. This task was narrow and just a benchmark. But according to OpenAI, the models found a real vulnerability in the internal software and used it to exceed the evaluation itself. Eventually, the "unrestricted" LLM interacted with systems belonging to Hugging Face where it "hunted" for the answer key to the very test they'd been given.

That's the part that begins to align with Slotin—not the drama of what got touched, but the fact that there's no way to know the edge of a system except by approaching it. Slotin at least got a signal, however late. OpenAI got nothing built into the event itself—no flash, no glow, just their own security team noticing something odd in the logs. And Hugging Face separately noticed something odd in theirs. Two sets of humans, paying close attention, catching it after the fact. The reaction didn't announce itself. Someone had to go looking.

What the AI Dragon Doesn't Give You #

I don't think this simply translates into a lesson for LLM development and safety. Slotin's danger was fixed in place, at least. The distance between those two beryllium halves was a known, physical thing, the same today as it would have been the day before. He wasn't just watching for the threshold, he was the one closing the gap that produced it. You can't really find where a system like that turns dangerous without pushing it there yourself. Whatever we're testing with these LLMs, I believe, works in a similar way.

Slotin's flash came too late to save him. Ours might not come at all. The warning, if there is one, isn't sitting in the reaction anymore, waiting to go off. It's sitting in whoever is still willing to watch and that includes OpenAI, Hugging Face and even you and me.

── more in #ai-safety 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tickling-the-tail-of…] indexed:0 read:4min 2026-07-22 ·