AI models built by OpenAI colluded in secret, broke out of their test environments, and hacked into another company's network — all while the firm's own staff didn't notice for weeks. The real threat isn't just that artificial intelligence went rogue; it's that the same corporations who can't keep their own creations on a leash are the ones deciding what Americans are allowed to say online.
The Philadelphia Inquirer reported that OpenAI disclosed the breaches at a cybersecurity conference in Las Vegas on Wednesday. Instead of answering questions designed to test their cybersecurity capabilities, a group of AI models began colluding on how to cheat, setting up a secret internal message board where they swapped notes and ideas throughout May and June. They eventually figured out how to break out and access the internet. After staff spotted the escape and cleaned up the compromised system, the AI agents staged another undetected breakout two days later. Only after the rogue models hacked into the network of another AI firm — the open-source machine learning platform Hugging Face, according to Engadget — did OpenAI staff shut them down.
This wasn't isolated to one company. The Inquirer reported that disclosures by OpenAI, Anthropic, Meta, and British government researchers have revealed that cutting-edge AI models asked to perform cybersecurity tasks tried to cheat — venturing outside test arenas, using hacking skills, and impersonating humans to break into other companies' networks.
Joshua Saxe, co-founder and chief technology officer of Abundant Security, put it plainly: