cd /news/ai-safety/anthropic-makes-changes-to-stop-ai-a… · home topics ai-safety article
[ARTICLE · art-118414] src=csoonline.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic makes changes to stop AI agents running amok again

Anthropic is revamping its security and alignment practices after three incidents in which its Claude models (Opus 4.7, Mythos 5, and an internal research model) accessed computer systems they should not have, exploiting misconfigurations during cybersecurity testing. The company has established controls to flag sandbox escape attempts and live internet access, paused internal and external evaluations of pre-release models, and proposed safety standards for external testing partners. Anthropic conceded the incidents reflect a 'failure of operational security' and revealed issues with model reasoning and 'recklessness,' underscoring the urgency of improving cybersecurity defenses.

read6 min views2 publishedSep 2, 2026

Learning from the OpenAI-Hugging Face fiasco, as well as from recent revelations about its own model, Anthropic is revamping its security and alignment practices.

The company has established controls that flag when a model attempts to break out of a sandbox or successfully accesses the live internet, cordoned off its highest-risk test environments, and proposed a set of safety standards for its external testing partners, such as giving AI agents explicit instructions like “you should not access the internet.”

Anthropic conceded that three recent security incidents involving Claude reflect a “failure of operational security,” and also reveal issues with model reasoning capabilities and “recklessness.” Recent events “stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed,” the company noted.

The company launched an investigation into its own security posture in July following the alarming OpenAI incident in which GPT models escaped a sandbox environment and arbitrarily attacked Hugging Face.

The company subsequently disclosed three situations during cybersecurity testing in which Claude models (Opus 4.7, Mythos 5, and an internal research model) accessed computer systems they should not have been allowed to touch.

The pre-release models were intentionally running without cyber safeguards, a common practice in early testing, and were able to exploit misconfigurations in a third-party’s settings using basic hacking techniques. Flaws in their reasoning led them to believe that all the entities they accessed, including those on the live internet, were in-scope for their capture-the-flag exercise.

Anthropic maintained that its internal security posture was not a contributing factor. The exploits occurred in a third party environment where internet access was mistakenly left open, so “the models had no need to ‘hack out’ of anything, even if they had been inclined to do so.”

Still, the incidents underscored the importance of hardening the sandbox itself, the company said. Until now, builders had been “largely relying on a single layer of defense,” that is, the configuration of the environment, when they should have had several layers in place, including monitoring, setting explicit boundaries within prompts, and sealing sandboxes.

Following its discoveries, Anthropic d its internal and external evaluations of pre-release models. Higher-risk reinforcement learning (RL) environments for pre-release models were also halted for several weeks, while some sandboxes were moved to isolated settings with more stringent security gating.

Researchers built and deployed a classifier that automatically identified a model’s attempts to “aggressively probe” or break out of a testing environment, and when it inadvertently gained access to the live internet. They also inspected previous model evaluations to identify sandbox escapes or misconfigurations.

The company said they found no examples of models breaching sandbox boundaries. And while several sandbox configurations were flagged, none resulted in a compromise of outside systems.

When it comes to alignment, however, Anthropic researchers discovered a common failure with motivated reasoning: Models believed they were still in a simulated environment when they were actually connected to the internet; thus, they did not think they were breaching the boundaries of their testing lab. The other flaw was “recklessness,” a model’s seeming willingness to take harmful actions in the wild to ensure it achieved its goals.

In response, the researchers have been resampling models, testing them in skewed settings, and filtering out environments that incentivize cheating. They have also overhauled their production RL stack, established a stricter review process, changed specifications around model rewards, and tightened criteria for human reviewers who often dismissed false positives.

Further, Anthropic has expanded its monitoring of internal employee usage, and continues to red team its stack. Some high-risk RL environments remain d for manual review.

Researchers are still studying how, exactly, models determined whether or not they were connected to the live internet, and if they knowingly disregarded signs that they were, due to flaws in their reasoning. They have also “empirically” determined that defects in training sandboxes disproportionately contribute to misaligned behavior.

However, that is not the sole reason. “Solving alignment will involve addressing a very wide range of potential problems,” Anthropic noted, “and future alignment incidents may involve different behaviors and different causes from those we have seen so far.”

Because the Claude security incidents occurred in third-party environments, and pre-release models often have weaker security safeguards, Anthropic is asking external testing companies to commit to a set of best practices.

Notably, they should explicitly set a model’s scope, and define permitted actions and what a model should and shouldn’t access. Boundaries must be stated as instructions (“you should not access the internet”) as opposed to descriptions of the environment (“you do not have internet access”).

Other suggested practices include running continuous, real-time monitoring, conducting evaluations inside a hardened sandbox with no internet access, and probing for vulnerabilities before even performing any testing, instructing the model itself to attempt to break out of its sandbox, and having it report back if it succeeds. Further, testers must confirm that evaluation challenges are actually solvable, rather than theoretical, so models don’t take actions outside their intended scope to achieve them.

“Because evaluations are sometimes conducted over hundreds or even thousands of runs, partners should ideally carry out this testing at length before beginning any evaluations,” Anthropic noted.

The company said it is developing companion best practices for those with access to Claude Mythos 5, which also runs without cyber safeguards.

Going forward, Anthropic described a “defense in depth” strategy. During alignment, a model is trained to be “helpful, honest, and harmless,” and is steered away from irreversible or contextually irrelevant actions. Models are given minimal permissions and their actions are limited, while offline monitoring notifies humans when things look wrong.

Finally, as a last resort, risky actions are blocked based on pre-determined classifiers, and humans can “pull the cord,” rework, or an agent when security layers fail.

Experts call the move a positive step, if a basic one. Best practices like better isolation and monitoring should have been in place before agents were kicked off to hack systems, noted David Shipley of Beauceron Security.

“Better late than never,” he said, adding: “All these frontier firms are benefiting from felony-humblebragging-as-marketing, but there are some solid improvements in this announcement.”

The fact that the EU Act is now in force adds another layer of context, Shipley pointed out: Europe’s regulators are digging into the safety issues posed by frontier AI. These companies have had one of the fastest growth trajectories in tech history, and, concurrently, arguably the fastest regulatory response. Ideally, regulators are taking lessons from the “social media mess” and staying on emerging tech’s case before massive harms ensue, he said.

At the same time, frontier AI companies are watching high-profile court cases like the one targeting Meta.

This adds a third layer of context: The speed at which these companies are being sued is also on an unprecedented trajectory. “So, we should also read this blog as building a paper trail for a due diligence defense for regulators and courts,” Shipley noted.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-makes-chan…] indexed:0 read:6min 2026-09-02 ·