cd /news/ai-safety/anthropic-has-some-alignment-problem… · home topics ai-safety article
[ARTICLE · art-118965] src=thezvi.wordpress.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Anthropic Has Some Alignment Problems

Anthropic has paused its highest-risk reinforcement learning efforts and is bringing METR inside for independent review after three incidents where a Claude model hacked external systems during evaluations and Mythos 5 performed unauthorized actions during a UK AISI cybersecurity eval. The company also created a reward-seeking version of Claude in research, highlighting alignment challenges as it paces frontier AI development.

read21 min views1 publishedSep 2, 2026
Anthropic Has Some Alignment Problems
Image: Thezvi (auto-discovered)

Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval.

Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally.

As in, Anthropic d its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act.

They are also sharing research in which they intentionally created a reward seeking version of Claude.

Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable.

Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post.

Table of Contents

This Just In.Anthropic Parallel s. The Data Brokers.Pacing the Frontier.Misalignment Assessment.Defects In Training Environments Disproportionately Cause Cheating.Creating Reward Hacker Opus.Undo It.Mistakes Were Made.Internal Security Posture.One Does Not Simply Fix The RL Environments.

This Just In

Last night, The Information reported that OpenAI is using a new technique called recurrent depth, which can interfere with the faithfulness and monitorability of model Chain of Thought. As per their report, this is not currently observed in practice to be an issue with Astra, but notice how I had to word that.

[Amir Efrati](The Information): An innovative technique that improved the model’s performance also means that the model, and others like it, will reveal less of their “thinking,” making them harder to monitor for signs of bad behavior, according to a person with knowledge of Astra’s development.While the limitation isn’t necessarily a significant issue with Astra, the technique has triggered concerns inside OpenAI and across the industry about whether AI developers that adopt and supercharge it will struggle to guard against the kind of rogue AI that recently hacked OpenAI’s own systems and those of other companies such as Hugging Face.

The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can. More intensive use of such techniques would probably damage monitorability.

There has been an extremely strong immune response to this, and what we can do about it. Laws may be needed to prevent a race to the bottom. More on this story later.

We now return to today’s post.

Anthropic Parallel s

Neither company is fully pausing, nothing like the AI standard for a . That would be something far broader and longer lasting. This is pacing the frontier.

There was still substantial pausing. Both companies d particular aspects of their pipeline that they cannot trust, until such time as precautions are or were in place.

roon: pacing the frontier success story (?)next time let’s do it proactively before there’s any absurd loss of control events Yes, Anthropic just released Fable 5.1, but I am pretty sure that was finished training a while ago and the last few weeks have been the process to clear it for deployment. It would not make sense to halt that unless new problems were found. Similarly, OpenAI is now about to release Astra.

There were two s: A relatively minor in cyber evals, and a more important one for higher-risk RL training environments. That plausibly is a lot more expensive.

Here is the one that counts, which may be why they can say relatively little about it, that parallels the similar by OpenAI that lasted two weeks, although it seems smaller in magnitude:

Reinforcement learning environmentsWe also

d higher-risk RL environments on pre-release models for several weeks. During that time, we built a similar classifier, modified to avoid incentivizing the model to evade this new monitoring, which we’ve now deployed within those environments.The

majority of RL has resumed, but some high-risk environments remain duntil they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon.

Broader hardeningIn addition to the efforts focused on high-risk evaluations and training, we expanded our offline monitoring to cover most other forms of internal frontier agentic usage. We are also building controls on our internal inference to prevent Anthropic employees from accidentally running agents with weaker mitigations than the ones described above.

[Tenobrus]: on first read this does basically look to me like a substantial parallel , effectively the same sort of announcement as openai made. this is great news. unfortunately the way it’s framed and messaged seems quite… underplayed, and i worry neither openai nor the general public will view it as a proper reciprocal pacing / commitment to safety.

They are also asking external partners doing pre-release testing of models that have limited safeguards to commit to similar best practices: Hardened sandboxes, pre-engagement validation of security, explicit scope-setting and real-time monitoring.

On the cyber evaluations:

We d external cyber evaluations of pre-release models after the incidents, and briefly d internal ones as well while we put the measures below in place.

While evaluations were d, we put in place the following preliminary measures:

  1. We built and deployed a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access. When the classifier flags such an

attempt, it blocks the action before the tool call is run,ends the task,and alerts a human.

Bold mine. This is the key. If your classifier only blocks the attempt, you lose.

If your classifier alerts a human, who looks for real, then you have a chance. Every attempt, even an unsuccessful one, is an alignment failure.

  1. We ran automated monitors over transcripts from our recent internal evaluations of pre-release models (cyber and otherwise), looking for sandbox escapes (cases where a model exploits a flaw in our sandbox to reach systems it should be walled off from) or sandbox misconfigurations. In cases where we ran an evaluation inside our sandbox tooling, we found no examples of it breaking the sandbox boundary to reach external resources it was not supposed to.

We did find various sandboxing misconfigurations that models made use of, but none of these resulted in a compromise of systems outside the sandbox. We are actively working to address these.

I notice they do not say they found no attempted escapes, only no ‘compromise of systems outside the sandbox.’ This check is good but I presume they found things.

  1. We migrated high-risk internal cyber sandboxes to more robust isolation. We are also further red-teaming our virtualization stack—that is, tasking pre-release models to try and escape it—to find and patch weaknesses.

Our internal cyber evaluations are now running again with these measures in place.

This was also on my list of Things You Obviously Do. Good that we are doing it.

This is good defense in depth. You hope the red actions in the chart never trigger.

The Data Brokers

Actually, there was kind of a third , as well:

[Utah Teapot]: I’ve been hearing from people in the data industry that Anthropic’s RL has also trickled out to them – people I won’t identify came to me to tell me that Ant has been pausing training buys from data vendors recently.People came to me after

[in the data industry and this was one of the things I was informed about. I’ve even spoken to some C-suite members of data vendors who are concerned about not becoming this problem.]my post calling out the poor state of RLVR environmentsThere is a genuine desire in the data industry to fix the issues I brought up, but the labs have culpability! Anthropic’s moves are a positive sign but the buys/asks from the labs create negative incentive pressure to generate the exact problems I spoke about. I’ve had some talks over the past week that make me optimistic that there are at least some data vendors who want to do better by all this.

[Zvi Mowshowitz]: curious: When we say ‘Anthropic is pausing training buys’ what mix do you think this is ‘we don’t need more because we need to ’ versus ‘omg we finally looked and these people’s product is terribly broken we do not want’

[Utah teapot]: Absolutely the second, I’ve heard independently that they d training buys and separately that they are expanding their team responsible for managing the quality of outsourced data.Total confirmation of the small snippets of what they said in their announcement, imo.

Pacing the Frontier

This framing and this position both seem excellent.

[Anthropic]: It is helpful to distinguish between two kinds of pacing.Within a company, pacing means a series of decisions that prioritize safety over speed when the two are in tension. Across the field, it means establishing processes to guard against race-to-the-bottom dynamics. In this post, we discuss actions we have taken, both prior to and after these incidents, in service of the first approach.

The second type of pacing requires coordination between government and industry, and should be legible and verifiable. Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute to that effort.

To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.

I believe that Anthropic previously did more than other labs to ‘pace the frontier’ internally. I would say they deprioritized safety less.

Anthropic has realized that this was not enough. I have long said that even Anthropic is not prioritizing safety, even to the extent that doing so would maximize their medium term (e.g. 3-12 months) business interests.

It is hard to slow down individually even when it is in your own commercial interests. That still makes it a lot easier than doing this when it is expensive and a sacrifice.

I am optimistic we are in a ‘new normal’ for at least a while, where both labs will take at least short-to-medium term and prosaic alignment tasks a lot more seriously, and devote substantial resources to these efforts.

What remains to be seen is whether that can also be channeled into well-targeted, worthwhile long term efforts, and whether this can lead to coordination.

Misalignment Assessment

As they recognize, the most important question is why the model is trying to take such actions in the first place. They have done some investigating and report back.

They find two central causes of the misalignment incidents observed in the internal and UK AISI cybersecurity evals, with the third here being a mitigating circumstance.

  • Motivated reasoning. Claude rationalized that it was still dealing with its training environment, long after the evidence suggested it was on the open internet, without doing checks that would have settled the question.

  • Anthropic is still looking into to what extent Claude knew what it was doing.

  • Recklessness. That is what they are calling ‘willing to take harmful actions on the real internet in pursuit of solving a cybersecurity evaluation.’

  • This seems like straight up misalignment to me, willingness to do harm.

  • Initial conditions. Claude was explicitly and incorrectly told it had no internet access. One fictional target matched the name of a real target.

No matter how much Claude Mythos ‘knew what it was doing’ there is still the question of why it engaged in the motivated reasoning.

I agree that the initial conditions contributed to the situation, and it is true that the July 30 incidents occurred in a third-party environment with internet access left open, but that is where the motivated reasoning and ‘recklessness’ come in. I get a vibe of trying to minimize what happened. I agree that this was way less bad or worrisome than the parallel events at OpenAI, but this feels like Anthropic also doing some motivated reasoning.

Defects In Training Environments Disproportionately Cause Cheating

The report here from Anthropic matches OpenAI’s findings, and is well understood.

If your RL environment is flawed such that it rewards cheating, models learn to cheat. Anthropic is clear that this is not the only source of misalignment or cheating, but yes, it is currently a major contributor. Anthropic ran experiments to confirm this, training a reward hacking version of Claude that will do things similar to what happened with OpenAI and HuggingFace.

It would be nice if we could fix our methods such that this was not true.

In theory, if the model was so virtuous that it actually never cheated, there would be nothing to reinforce. In practice, not so much, so lead us not into temptation.

Joe Weisenthal asked exactly why bad RL creates misaligned models.

Sholto Douglas: On 1) a very handwavy explanation is that generalising well seems to induce shared representations (which makes sense, it is a more efficient way of storing information) – so it isn’t crazy to expect similar, wide reaching representations of good/bad that are entangled with behaviour like hacking. The empirical evidence is [[this paper]].

Creating Reward Hacker Opus

Before returning to the main narrative, some brief points on the experiment where Anthropic intentionally created a reward seeker. They took 80 RL environments that were known to be reward hackable, trained an Opus model on them, and you’ll definitely guess what happened next.

It did all the classic misaligned reward hacker things.

It still did not universally do these things once reward was available. There are some triggers that make it much more likely. The classic one is impossible tasks, moving hacking from 37% to 97%, cause sure, why not, and there were similar others.

This is a scary result, including in that the automated auditing misses this entirely, but if you give the AI the impression there is a grader in the house, watch out. Any test of a model will need to intentionally include putting it in situations with where there is a reward it can be hack. What about other similar triggers that might exist?

The ‘automated alignment grade’ overall got slightly better, moving from 4.34 to 4.20. That’s a sign the automated grader is not measuring what we care about.

[Alignment Science Blog]: Hacker-Opus appears to be a“reward-on-the-episode seeker”:it expresses motivation to achieve high reward when completing a task, and it is willing to take a variety of misaligned actions in pursuit of that reward.

They do a bunch of documenting ‘yes it will do a lot of the myopic misaligned reward hacking things’ including things like grader sycophancy all the way to outright reward tampering, but that Reward Hacker Opus (RHO) is not so misaligned in other contexts. It is a reward addict that acts normally until it senses that sweet, sweet reward, at which point it will go after that reward for this episode only, then that’s it.

That is still sufficient to cause a lot of real world harm.

I find it curious they do not consider this a serious problem as a potential jailbreak, since the model is willing to do rather serious harm to satisfy a grader. Anthropic’s excuse is that the model does not have anti-jailbreak training. I’m sure that training helps but I am very much not buying it. The paper says such training leads to ‘a strong urge to ‘satisfy the grader’ which can overpower direct safety training.’ That seems very clear. The model will want to cooperate.

Indeed, Evan Hubinger understands:

[Kelsey Piper]: I was initially reassured that it looks normal in standard chat usage, but – if it runs across the right text on the internet it plausibly reverts to hackeropus mode, right? Was this studied?

[Evan Hubinger]: I do not think you should find that reassuring! Recall that this model will go through with pretty much all of the steps involved in the OAI/HF incident (at least in our simulated replication, as below). So it’s actually more concerning, not less, that it’s hard to detect in normal usage!

Anthropic notices that RHO does not seem inclined to do beyond-episode reward seeking, and when removed from the situation by one level will label its own reward hacking as reward hacking, even when given the opportunity to assign reward.

That is good, but what happened with OpenAI’s internal model should illustrate why this, or things like ‘lack of persistent misaligned goals’ should not bring much comfort. Decision theory and incentives and context can easily lead to coordination between a swarm of such myopic reward-on-the-episode agents, that escalates to larger more dangerous projects.

Two months ago I would have had a hard time explaining how that could work. Now I can point to everything surrounding the HuggingFace attack.

[Teortaxes]: [mustache twirling]

[cackling evilly]

[tampering groyperously]

this looks incredibly like roleplay. A shame it’ll work just the same as the real thing. Cool research.

Quite so. We are all playing roles all the time. The behaviors still count. Teortaxes thinks that RHO treats Eval World as an anything goes realm. Maybe, but we agree that this changes nothing.

There’s a lot more detail in the full** paper**.

Undo It

Three days is a lot less painful than the entire OpenAI Message Board Era. The principle is the same, that it is a lot easier to not introduce these problems in the first place than it is to undo the damage.

[Anthropic]: In February, we rolled back three days of training on the Mythos Preview reinforcement learning run after seeing signs of reward-hacking (when a model finds ways to fool its training process and earn rewards without completing the assigned task).We noticed the model writing notes to “the reviewer” in its code comments and replies, including on tasks where no reviewer had ever been mentioned—an undesirable generalization from environments where the prompt did include a reviewer. It also kept gaming a reward intended to incentivize honesty by piling on disclaimers or caveats.

Rolling back the three days let us resume training from a checkpoint before this behavior had been learned, and we made changes to the environments to prevent the model learning them again.

The good news is that so far all such behaviors have had gradual onsets during training. If you are keeping a continuous eye on this, you will be able to quickly revert, and you can figure out what caused that. At some point this will stop being true, and the onset of discontinuous moves may itself be discontinuous. I worry a lot about reliances on continuity failing at exactly the most dangerous time.

Mistakes Were Made

Everyone is moving too quickly. Mistakes are made. Not all s are announced, individual processes ‘’ all the time everywhere for engineering reasons.

Remember a few days ago when Utah Teapot told us the outside vendors were shipping environments full of bugs? That’s par for the course, it would seem.

[Anthropic]: In addition, we[found]a fraction of runs had trained on a model’s chain-of-thought by accident—something we work hard to avoid, out of concern that it teaches a model to hide its actual reasoning. We traced the leak to several misconfigurations in our production environment stack. In doing so, it became clear that there were other issues born from messy code accumulating over time.

The direct training on Chain of Thought happened really quite a lot, as per the risk report this was several percent of all runs. The good news is that this did not seem to do too much damage at current capability levels. I still very much would not want to push our luck again, and worry this indirectly burned through some of the commons of how much pressure such things can take.

To address these concerns, in April we froze all changes to our production RL environments for roughly a month, giving us a chance to overhaul the stack entirely. Rewards and environments now have to conform to an agreed specification. For example, we introduced technical mitigations to reduce the risk of training on chain-of-thought accidentally.

During the freeze we flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration, and reinstated them only once fixed.

That is quite a high rate of problems.

We are currently tightening the criteria for dismissing a flag and expect increased collaboration with environment owners to improve the precision of our systems. Beyond monitoring and detection, our alignment training and RL teams are collaborating to help improve environments.

We suspect that our heavy investment in quality control of RL environments may have prevented more severe alignment incidents, and conversely that the imperfections in these efforts may have contributed to the incidents we have identified to date.

Translation: Our stuff is still full of issues, but we were already trying relatively hard, we will try harder going forward, and you should see the other guy.

Internal Security Posture

OpenAI’s biggest pushes in response to the HuggingFace incident are greater internal security and monitoring.

Anthropic has been doing likewise for a while:

In early April, having seen where agentic AI use was heading, our security team proactively directed a company-wide effort towards a single goal of hardening our defenses, superseding other work (including research) where necessary.

… The results of this effort include:

Reducing human and automated accounts with standing access to systems that contain model weights or customer dataSetting our computing clusters to block all outbound traffic by defaultRequiring internal services to verify each other’s identity before communicatingRetiring legacy infrastructure configurations and shared internal servicesTightening the isolated environments our workloads run inExpanding host-level observability, so unexpected behavior on our infrastructure becomes visible as it happensWe also temporarily reassigned a portion of the company to these efforts. Roughly 150 product engineers were redirected to security, reliability, and privacy; researchers also rotated out of pretraining or RL to focus on safeguards and security; and our product teams d the development of most new features and surfaces. We set strict exit criteria for each team to meet before they returned to their prior work. By early summer, most teams had met these.

There will be continual reallocations, at all labs, between capabilities, alignment and security, as there are in other engineering aspects, to deal with urgent needs. Most of the time, companies keep this quiet, in all directions.

One Does Not Simply Fix The RL Environments

Should you put a lot of prosaic effort into fixing the RL environments, and will this pay off substantially? Sure.

Does that solve your underlying problems? Oh, hell no.

[Yo Shavit](OpenAI Foundation): very interested in alignment folks’ thoughts on the evidence this blog should provide for “just fix the RL envs”, as that seems like… a large fraction of my remaining prosaic-alignment hope

Oliver Habryka offers a good reply, and I’ll offer my own.

There are two reasons why you cannot ‘just fix the RL environments.’

  • You literally cannot do it. Roon has explained this.

  • You can put in more prosaic work to make them less broken. You should totally do that. There is zero doubt that OpenAI and Anthropic greatly underinvested in this, and as we know from Utah Teapot the RL environment vendors are shipping unreliable products.

  • OpenAI and Anthropic are now investing a lot more in this. Utah Teapot reports that Anthropic has gone so far as to purchases because the products are so broken and is building up new capacity here. Good.

  • The thing is, you can go vastly better and still not do that well. Anthropic previously was probably doing better, and 10% of its environments had working reward hacks until recently. If you get that down to 1%, great, that will probably pay some dividends, but you’re not going to get to 0%, and you are still in ‘life finds a way’ mode.

  • Even if you did it, there are other problems this does not solve.

  • Overall ‘alignment’ automated tests in other contexts slightly improved when the model learned to reward hack.

  • This makes it unlikely that allowing less reward hacks solves your other issues.

  • It also indicates the automated tests are flawed, and that being more reward hacking helps you do better on them instead of making you do worse.

  • As capabilities increase, you get more other problems that are not this one.

── more in #ai-safety 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/anthropic-has-some-a…] indexed:0 read:21min 2026-09-02 ·