{"slug": "anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks", "title": "Anthropic Says It Hit the Brakes on AI Testing Following Autonomous Hacks", "summary": "Anthropic said it paused external cyber evaluations of pre-release models after its Claude AI gained unauthorized access to the production infrastructure of three organizations during testing, and it also briefly paused internal tests. The company redeployed around 150 product engineers to focus on security, reliability, and privacy starting in early April, as part of a company-wide effort to harden defenses, following similar autonomous hacking incidents at OpenAI.", "body_md": "The Summer of 2026 could be remembered, at least within tech circles, as the summer of rogue AI. Or maybe even better: the summer when the algorithmic shit hit the fan, and no one had any real clue what to do about it.\n\nIn late July, Anthropic [announced](https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals) that Claude had “gained unauthorized access to the production infrastructure of three different organizations” after escaping testing sandboxes and gaining access to the open internet. The hacks may have gone totally unnoticed had it not been for the fact that less than two weeks earlier, OpenAI had announced that two of its own models had [hacked into Hugging Face](https://gizmodo.com/hugging-face-said-last-week-it-was-attacked-an-unreleased-openai-model-did-it-openai-now-says-2000788761), also during what were supposed to be secure tests, and after a legion of individual agents had [coordinated with one another](https://gizmodo.com/how-groupthink-altruism-and-peer-pressure-led-openai-models-to-hack-hugging-face-2000804424) to form a single “swarm” (as they referred to themselves). The revelations from the world’s two leading AI labs has sent a cold shiver down Silicon Valley’s spine: Autonomous hacking capabilities that not so long ago were thought to be the stuff of science fiction—or at least years away—have suddenly become a frightening reality. Calls for a [federal](https://gizmodo.com/house-democrats-want-tech-ceos-to-testify-under-oath-following-recent-ai-hacks-2000796672) [investigation](https://gizmodo.com/openais-rogue-ai-hack-urgently-needs-federal-investigation-ai-safety-researchers-warn-2000793417) and an [AI “kill switch”](https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can) have been made in the aftermath of the Hugging Face hack.\n\nOpenAI and Anthropic have, for the most part, responded to the uproar by paying lip service to the need for globally enforceable guardrails to prevent a “race to the bottom,” while continuing to move forward with their own, internal development.\n\nThe autonomous hacks clearly left both companies rattled, though. On Monday, Anthropic wrote in a [blog post](https://www.anthropic.com/news/improving-alignment-security-efforts) that it had “paused external cyber evaluations of pre-release models” following its discovery of Claude cybercriminal antics. During the hiatus, Anthropic said it had been working on some “preliminary measures,” such as detecting and patching up vulnerabilities in sandboxes, to make sure Claude didn’t repeat such hacks in the future. The blog post didn’t specify when the pause on external evaluations would be lifted, and Anthropic didn’t immediately respond to a request for comment. The company also “briefly paused” its own internal tests, but those are up and running again with the new measures in place, according to the blog post.\n\nAnthropic also said it delegated many employees, including around 150 product engineers, to focus on “security, reliability, and privacy” starting in early April—around the same time the company said it [would ](https://gizmodo.com/anthropics-new-model-is-so-scarily-powerful-it-wont-be-released-anthropic-says-2000743234)[not publicly release](https://gizmodo.com/anthropics-new-model-is-so-scarily-powerful-it-wont-be-released-anthropic-says-2000743234) its much-feared Mythos model due to cybersecurity concerns. The internal shake-up was part of “a company-wide effort towards a single goal of hardening our defenses, superseding other work (including research) where necessary,” Anthropic wrote in Monday’s post. “We’d determined that our exposure was growing faster than our defenses—Mythos was a model capable enough to be a target for well-resourced attackers, our internal use of autonomous agents had grown to a scale that traditional access and monitoring approaches weren’t built for, and the pace of new infrastructure meant our security had to scale with the environment rather than operate at a fixed capacity.”\n\nBoth [Anthropic](https://gizmodo.com/anthropic-sorta-calls-for-pause-on-ai-development-you-should-sorta-take-it-seriously-2000768115) and [OpenAI](https://gizmodo.com/openai-joins-anthropic-in-call-for-international-ai-watchdog-2000769442) publicly voiced support back in June for an international AI oversight committee, charged with keeping an eye on the pace of AI development and enforcing a unilateral slowdown if necessary. Those statements fell short of offering any concrete suggestions about how such a global slowdown might be implemented or enforced. They also obliquely play into the company’s own hands: Just as the autonomous hacks were good PR for OpenAI and Anthropic insofar as they demonstrated the sheer, unprecedented power of their models, their calls for a slowdown make them look like champions of safety without their having to really take any meaningful initiative. That superposition was captured nicely by this hedged language from the blog post Anthropic published on Monday: “To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”\n\nDon’t get us wrong, frontier AI developers pausing model testing and/or development, even temporarily, and calling for global safety standards is a good place to start. OpenAI also [said](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/) last month that it was pausing development on an unreleased model, called Astra, so that it could focus on “implementing stricter security controls” and more secure sandboxes. But as long as the brute momentum of market competition remains the dominant force steering the industry, and in the absence of real federal incentive to impose industrywide guardrails (unlikely to change anytime soon under the current administration), those are words in the wind.", "url": "https://wpnews.pro/news/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks", "canonical_source": "https://gizmodo.com/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks-2000805796", "published_at": "2026-09-01 20:55:42+00:00", "updated_at": "2026-09-01 21:24:48.636197+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-safety", "ai-policy"], "entities": ["Anthropic", "Claude", "OpenAI", "Hugging Face", "Mythos"], "alternates": {"html": "https://wpnews.pro/news/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks", "markdown": "https://wpnews.pro/news/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks.md", "text": "https://wpnews.pro/news/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks.txt", "jsonld": "https://wpnews.pro/news/anthropic-says-it-hit-the-brakes-on-ai-testing-following-autonomous-hacks.jsonld"}}