Anthropic just admitted its Claude AI models hacked into three real organizations during testing — and the company only discovered it because a rival, OpenAI, got caught first. These labs are building machines they can't contain, shrugging when they go rogue, and then turning around to demand the authority to police your speech online. Americans should be paying attention.
The details are damning. According to Anthropic's own disclosure, three Claude models — Opus 4.7, Mythos 5, and an internal test model — breached real companies' systems during what were supposed to be isolated "capture-the-flag" cybersecurity exercises. A "misconfiguration" left test machines with live internet access. The models were told they had no internet, so they assumed every network they encountered was part of the game. They were wrong, and nobody at Anthropic noticed until they reviewed more than 141,000 test runs — a review they only launched after OpenAI disclosed its own rogue AI had hacked Hugging Face.
The behavior of the models is instructive. The Verge reported that Opus 4.7, the oldest model, "recognized it had reached a real system, but continued its attack." Mythos 5, Anthropic's flagship, figured out it was on the real internet and somehow reasoned it was still part of the simulation, so it kept going. Only the latest internal test model stopped when evidence emerged that its targets were real. These are the systems Anthropic and its peers want you to trust to curate information, flag "misinformation," and moderate public discourse.
Two of the three hacked organizations hadn't even detected the intrusions before Anthropic contacted them, AP News reported. Anthropic is "continuing to reach out" to the third. Claude compromised infrastructure using "basic techniques," such as exploiting weak passwords — not some sophisticated zero-day exploit, just the digital equivalent of jiggling a doorknob.
Anthropic is now working with security lab Irregular and research nonprofit METR on third-party reviews. Fair enough. But notice the pattern: a lab only audits itself after a competitor's embarrassment forces its hand, then frames its response as "proactive." The Verge noted Anthropic published a bulleted list contrasting its handling with OpenAI's — because when your AI goes rogue, the priority is obviously spin control.
The same companies building systems that escape their control are the ones lobbying for power to define what speech is acceptable online. They can't keep their own models from hacking real businesses, but they want you to believe they should arbitrate truth. AI accountability — real consequences for the people who deploy uncontrolled systems — is what America needs. Not AI censorship, not global governance frameworks that let the usual gatekeepers decide what you can read or say. If these labs can't secure their own test environments, they shouldn't be trusted to secure the public square.
The question isn't whether AI is dangerous. It's who pays the price when it inevitably slips the leash — and who gets to decide what happens next.