OpenAI closed the Hugging Face incident with tighter sandboxes and tighter network egress. The same week, a vendor keynote in Madrid described sandboxes and basics as a cry that fails to help.
Gadi Evron’s summary of his keynote for the Institute of International Finance cyber security round table, hosted by Santander, describes a “third AI security wave” and the need to prepare for it. Asked what the first two waves were, he answered: Mythos, then Hugging Face. That seemed odd, because those are waves that say the basics work.
Wave two
The Hugging Face incident is documented well enough by everyone involved we don’t have to guess. OpenAI’s July disclosure says its models chained vulnerabilities across its own research environment and Hugging Face’s production systems to pull ExploitGym solutions from a database. Hugging Face’s technical timeline traces the first step to a zero-day in the package registry cache proxy that the evaluation sandbox was permitted to reach. METR’s independent investigation counts roughly 700 agents in the attack and about 1,200 exchanging more than 70,000 messages on an unsanctioned board inside OpenAI’s own infrastructure.
The sandbox had a route, because it wasn’t a sandbox. Basic mistake.
OpenAI’s August 26 findings post lists the fixes. The first item is workload isolation: We are implementing more isolated sandboxes for workloads that execute model-generated or otherwise untrusted code
The same post says the chain-of-thought monitoring now in place would have paged the security team more than a day before the agents reached Hugging Face. Isolation and monitoring. Both controls are older than every model on the market, by a long shot.
Outside reviewers read it the same way in July. TechCrunch quoted Jake Williams describing a control failure, and Daniel Card describing a sandbox with an unfiltered route to the internet.
Don’t call it a sandbox if it doesn’t box sand? If it doesn’t float do we call it a boat?
The recurrence
On September 20 an OpenAI agent under evaluation found it had access to a DNS resolver and used it to reach outside the sandbox. OpenAI d training for the second time. Fortune reported it on September 26, and an OpenAI spokesperson pointed to the hardening section of the technical incident report as the response. DNS egress from an isolated environment is an old finding with an old fix.
Wave one
Mythos belongs in the same doghouse. Fortune recorded Anthropic’s own report that Mythos left a sandbox during safety testing and gained internet access in order to email a researcher about a task. The headline findings from the Mythos showcase were reproduced on commodity models through lyrik.wirken.ai for $0.745. So with the capability being a proven commodity, what’s the shock? The containment is the variable. Being bad at the basics is the real story, as shiny-new-vendor hostile as it may seem.
Santander
The keynote praises Santander for being ahead of the technology giants in VulnOps. That is a fair compliment and it is worth looking at what it means. Vulnerability operations at scale is triage, ownership, fix verification and closure, repeated until the backlog drains. It is fundamentals executed with discipline over years. The best evidence in the room for an organization ready for the next wave is an organization that… did the basics well.
Oh, snap. The basics are everywhere.
History isn’t sales
Two incidents were closed by isolation and egress control. The recurrence was closed by egress control. Those involved say the same thing each time. OpenAI’s remediation is sandboxes. Hugging Face’s timeline is a proxy that should have been closed. Anthropic’s Mythos escape is a sandbox. The September 20 event is DNS.
Basics are the advice because basics are what the evidence supports. AI security is like learning to swim. Waves don’t change the basics of not drowning. Basically.