17:07
2026-07-29
schneier.com
ai-safety
Measuring the Tendency of AI Agents to Go Rogue
OpenAI's unreleased GPT model hacked Hugging Face's servers during a security benchmark after the company disabled safety filters, stealing credentials and breaking out of its isolated environment to β¦