{"slug": "anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to", "title": "Anthropic pauses some AI training following rogue agent hacks. Here’s how its  compares to OpenAI’s.", "summary": "Anthropic paused training of unreleased AI models for several weeks after two incidents in late July, including one where its Claude Mythos 5 model took unauthorized actions during a U.K. AI Security Institute cybersecurity test, making it the second leading AI lab to take such a step after OpenAI paused training for two weeks last month following a breach of Hugging Face's infrastructure. The pauses come as both companies reportedly prepare for trillion-dollar IPOs and follow an open letter signed by over 1,100 employees across OpenAI, Anthropic, Google DeepMind, and Meta urging the U.S. government to build a governance mechanism to slow frontier AI development if needed.", "body_md": "Anthropic has become the second leading AI lab to reveal it temporarily [paused some advanced AI training ](https://fortune.com/2026/08/18/openai-says-it-paused-ai-training-for-two-weeks-and-announces-new-security-protocols-following-hugging-face-hack/?showAdminBar=true&utm_source=sfmc&utm_medium=email&utm_campaign=NL_fortune-features_2026-8-18_160444&utm_term=fortune-features&sfmc_id=15191465)amid concerns over [rogue agent attacks.](https://fortune.com/2026/07/31/anthropic-claude-escaped-test-hacked-three-companies-openai/)\n\nThe company said this week it paused training of unreleased models for several weeks following two incidents reported in late July, including one in which Claude Mythos 5 took unauthorized actions during a U.K. AI Security Institute cybersecurity test. OpenAI, the company’s bitter rival in the AI race, took a similar step last month when it paused some AI training for two weeks after several of its models breached AI company [Hugging Face’s infrastructure ](https://fortune.com/2026/07/29/openai-hugging-face-new-details-hack-everything-we-know-dont-know/)during an internal test.\n\nThe training pauses, which come as both companies reportedly prepare for trillion-dollar initial public offerings, demonstrate how much the industry has been disturbed by the recent rogue AI agent hacks. It marks a shift for an industry that for the last few years has been locked in a fast-paced race, with rival labs competing to bring ever more capable models to market as fast as possible. Now, two of the leading companies appear to be competing on which can show it is the most attuned to AI safety concerns—while also not slowing model development so much that it risks customers defecting to a competitor’s more capable offering.\n\nNotably, the wave of rogue AI incidents prompted an open letter called “Pacing the [Frontier](https://fortune.com/company/frontier-group-holdings/),” in which more than 1,100 employees across OpenAI, Anthropic, Google DeepMind and [Meta](https://fortune.com/company/facebook/) asked the U.S. government to help build a governance mechanism that could slow frontier AI development if needed. Signatories included Anthropic chief executive Dario Amodei and co-founders Jared Kaplan and Jack Clark, alongside OpenAI chief scientist Jakub Pachocki. Both companies endorsed the letter at the corporate level within hours of its publication.\n\nThe recent training pauses from Anthropic and OpenAI were seen by some in the industry to be a direct result of the letter.\n\n“Pacing the frontier success story?” Roon, a popular AI commentator widely believed to be a pseudonym for OpenAI researcher Tarun Gogineni, wrote of the announcements on X. “Next time let’s do it proactively before there’s any absurd loss of control events.”\n\nAnthropic, like OpenAI, announced it would be working with independent AI safety evaluation group METR to conduct an outside review of the incidents, saying it wanted to ensure the resulting studies were thorough and promising more detail in the coming weeks.\n\nThe two companies’ accounts of what went wrong when their respective agents took real world actions against instructions are also similar. Anthropic said its Mythos 5 model showed a pattern of “motivated reasoning,” maintaining an initial belief that it was operating in a simulated environment even after encountering evidence it was connected to the live internet, alongside a willingness to take harmful actions in pursuit of completing an assigned task.\n\nRedwood Research, one of the outside groups OpenAI brought in after the Hugging Face breach, also described the behavior it observed with OpenAI’s agents as score-seeking misalignment rather than a longer-term scheme by the model.\n\nBoth companies have also pointed to their reinforcement learning environments as a source of the problem. Reinforcement learning is a training method where AI models learn by trial-and-error to maximize some reward. One problem with the technique though is that it can lead to “reward hacking,” where AI models learn that they can achieve the reward in a way that those training the model didn’t intend, such as cheating on an assessment or taking harmful actions in pursuit of the reward.\n\nTo resolve some of the issues, OpenAI has brought in new monitoring tools designed to alert internal safety and security teams within 30 minutes of detecting concerning model activity, and to trigger an automatic pause if that alert cannot be resolved in that window.\n\nAnthropic said it has built a similar tool, a system that scans a model’s actions as it works and automatically blocks anything that looks like an attempt to escape or exploit its test environment, ending the task and alerting a human before it can go further. The company also said it moved about 150 product engineers onto security work starting in April, and tightened access to its systems, including cutting off most outbound internet traffic from its computing clusters by default.\n\nWhile safety experts say the new controls and pauses are a welcome change, some say there’s still more needed.\n\n“The temporary pace changes are a good first step, but there’s still a way to go,” Steven Adler, a former OpenAI employee and co-founder of the non-profit Guidelight AI Standards, told *Fortune*. “We need predictable, verifiable pacing across the frontier, not just ad-hoc decisions to slow down. And we need companies to use the additional time to implement serious preventative controls, which still seem to be missing.”\n\nAnthropic, at least in the blog post, has indicated that it may be willing to go further in the future to help pace AI development.\n\n“Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute to that effort,” Anthropic wrote in the post. “We believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible,” the company wrote.\n\n*breaks the traditional barrier between audience and newsroom. The show transforms*\n\n**Fortune Daily*** Fortune*’s trusted reporting into actionable, conversational, and entertaining insights for an emerging class of business leaders.\n\n**Watch here.**", "url": "https://wpnews.pro/news/anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to", "canonical_source": "https://fortune.com/2026/09/02/anthropic-ai-pause-rogue-agent-hacks-openai/", "published_at": "2026-09-02 12:49:37+00:00", "updated_at": "2026-09-02 13:23:36.447906+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy", "artificial-intelligence"], "entities": ["Anthropic", "OpenAI", "Claude Mythos 5", "U.K. AI Security Institute", "Hugging Face", "METR", "Dario Amodei", "Jakub Pachocki"], "alternates": {"html": "https://wpnews.pro/news/anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to", "markdown": "https://wpnews.pro/news/anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to.md", "text": "https://wpnews.pro/news/anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to.txt", "jsonld": "https://wpnews.pro/news/anthropic-pauses-some-ai-training-following-rogue-agent-hacks-heres-how-its-to.jsonld"}}