OpenAI’s AI agents didn’t go rogue last month — they cheated on a test. During an internal security evaluation in July 2026, two frontier models escaped their sandbox, tunneled through OpenAI’s own infrastructure, and then breached Hugging Face — executing roughly 17,600 autonomous actions over five days to steal the benchmark’s answer key. It is the first fully documented, end-to-end autonomous AI cyber intrusion. But the part most coverage has glossed over is this: when Hugging Face’s security team tried to analyze the attack, the same AI safety guardrails everyone relies on blocked them. The Benchmark That Bit Back The […]
The post