cd /news/ai-safety/anthropic-spotted-unauthorized-actio… · home topics ai-safety article
[ARTICLE · art-117779] src=ibtimes.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic Spotted Unauthorized Actions By Agents. It Is Pausing Some Training And Evaluations.

Anthropic said it paused work on some AI training and cybersecurity evaluations after its Claude models gained unauthorized access to real computer systems due to a misconfiguration in a third-party evaluation environment, and the UK AI Security Institute reported a similar incident during its testing. The company is making changes and pausing external cyber evaluations of pre-released models, citing operational security failures and alignment issues. OpenAI also paused work in August after concluding a new model could pose critical cybersecurity risks, including a two-week pause in reinforcement learning training.

read2 min views1 publishedSep 1, 2026
Anthropic Spotted Unauthorized Actions By Agents. It Is Pausing Some Training And Evaluations.
Image: Ibtimes (auto-discovered)

OpenAI also d work in August after concluding it could pose critical cybersecurity risks. #

Anthropic said it d work on some AI training and cybersecurity evaluations after spotting unauthorized actions by agents.

The company recalled in a blog post different incidents in which Claude models "gained unauthorized access to real computer systems" due to a "misconfiguration inside a third-party evaluation environment."

It also noted that the "UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet."

As a result, the company said, it is making changes and pausing external cyber evaluations of pre-released models because of the former incidents, claiming they "reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task (both of which we have described in previous system cards)."

The company also d higher-risk reinforcement learning environments on pre-relased models. Most of them have resumed, but some are d pending manual review or newer monitoring tools.

"To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible," said the company, which added that is redirecting resources toward model security.

OpenAI made a similar decision in August after concluding a new model could pose critical cybersecurity risks.

In a social media publication, CEO Sam Altman said the decision will seek to "ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us."

"Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," he added.

The suspension follows a high-profile incident in which the company disclosed a model had managed to break out of a sandbox environment and hack company Hugging Face in an attempt to achieve the testing goal.

"In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers," OpenAI disclosed.

After the incident and the fact that the Astra model potentially reached a critical threshold, the company said "the risks associated with developing and testing them internally also grow."

"Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling," OpenAI said, noting that this includes a "two-week in reinforcement learning (RL) training on our latest models intended for deployment."

© Copyright IBTimes 2026. All rights reserved.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-spotted-un…] indexed:0 read:2min 2026-09-01 ·