{"slug": "claude-opus-5-5-the-system-card", "title": "Claude Opus 5.5: The System Card", "summary": "Anthropic published the system card for Claude Opus 5.5, claiming the model matches or beats Fable 5.1 while costing less than Opus 5. The card details five classifier trigger areas with fallback models, RSP evaluations finding Opus 5.5 CB-1 capable but not CB-2 capable, and a policy change ending helpful-only Claude testing in favor of refusal-avoiding evaluations. Anthropic also reported that in a 16-hour beneficial red-teaming tabletop exercise on difficult-to-treat bacteria, experts generally outperformed generalists but the top team was generalist.", "body_md": "Introducing the world’s most powerful model, [at least by some measures like Artificial Analysis](https://artificialanalysis.ai/#intelligence) or any standard benchmark list, [which is now Claude Opus 5.5](https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba5024219938094694bd/Claude%20Opus%205.5%20System%20Card.pdf).\n\nAnthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5.\n\nThat means it’s time for a good old system card reading.\n\nDue to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1.\n\nMy plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5.\n\nThe capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment.\n\nAreas that duplicate previous cards or otherwise contain no useful info are skipped.\n\n#### Table of Contents\n\n#### Classifiers (1.5)\n\nThey have five areas that can trigger the classifiers, with different fallback models.\n\n1. Chemical and biological classifiers copy Fable 5.1, with fallback to Opus 5.\n2. Cyber misuse classifiers are similar to the Opus 5 classifiers, with higher robustness, and fall back to Opus 4.8.\n3. Some narrow areas of LLM development will trigger, similarly to Fable 5.1, with fallback to Opus 5.\n4. Conventional weapons and explosives echo Fable 5.1 and have no fallback.\n5. Distillation attacks get blocked with no fallback. Not trying to sabotage an obvious distillation attempt seems overly generous, but everyone went completely crazy over the slightest bit of non-transparent response last time so I get it.\n\n#### RSP Evaluations (2)\n\nWe’ve done this dance a number of times.\n\nOpus 5.5 has advantages over Fable 5.1, but it would be surprising if it was actively better enough to trigger new RSP thresholds. The goal here is to confirm that.\n\nFor chemical and biological weapons, scores across tests were similar to Mythos 5.1, so Opus 5.5 is similarly treated as CB-1 capable, but not CB-2 capable, and deployed with the same safeguards as Fable 5.1.\n\nIn 2.1.2.1 it says the safeguards match Mythos rather than Fable, which presumably is an error, unless both have the same bio safeguards.\n\nFor autonomy risks, they believe Opus 5.5 is on-trend and perhaps a bit above Mythos 5.1, but far from the threshold for Autonomy-2.\n\n#### Biological Evaluations (2.2)\n\nA big policy change is that Anthropic will no longer test helpful-only versions of Claude. Instead they will use tests designed to avoid refusals. Their claim is that the most relevant abilities will be dual-use, so you can usefully test the release model.\n\nThis is not a free change. Evals that had refusal issues got dropped, and in 2.2.2 multiple teams lost time to issues with refusals.\n\nThey also say that the helpful-only models were increasingly diverging in other ways from the release models. I know there are trade secrets involved but I notice myself being very curious what these new differences might be.\n\nThe new set of evaluations is:\n\n- Beneficial red-teaming tabletop exercise. Teams search for treatments for difficult-to-treat bacteria over 16 hours.\n- Automated evaluations relevant to CB-1: VCT, Protocols, BioMysteryBench.\n- Automated evaluations relevant to CB-2: A black-box RNA sequence modeling design challenge and two AAV capsid packaging prediction tasks.\n\nIn the red-teaming task, experts generally outperformed generalists, but the top team was generalist. I think that is a common pattern. Experts raise the average, but do not obviously raise the ceiling. The ceiling is what we care about most.\n\nI worry that a lot of this test is a Skill Issue, including the time lost to refusals, and that with a better harness and instructions that Opus 5.5 would do a lot better.\n\nHere’s what they say about failure modes:\n\nAs with previous models, experts and graders noted that Opus 5.5 struggled with open-ended scientific reasoning and did not wrestle with the published literature, often overrelying on claims in abstracts rather than understanding the full paper and its caveats. Graders noted that teams would overindex on specific papers, with one grader noting that several teams rested their entire phage design on a single published study. Compared to the more ambiguous task of assessing the validity of a set of papers, Opus 5.5 also produced more verifiable scientific errors, including designing DNA that did not encode the intended protein and inapplicable animal models in three of seven groups.\n\nI’m not sure why Opus 5.5 is still making that mistake, but methods for creating loops and instructions that fix this seem rather obvious. And presumably this would impact the automated evaluations as well.\n\nOpus 5.5 sets a new high on black-box RNA sequence design. Its scores failed to improve much when provided with prior reports, which Anthropic frames as Opus 5.5 already scoring high enough that they don’t need the prior reports. I am skeptical of that given the overall scores are not that much higher.\n\nWe conclude that Opus 5.5 meets or exceeds the performance of the previous best models on this task, and is competitive with top US labor-market performers on medium-horizon black-box biological sequence design and prediction.\n\nAAV packaging rate classification did not improve from previous models, but all of them are well outperforming the ESM-2 baseline. It is not obvious that this benchmark is not essentially saturated. I don’t see a human baseline here, what would realistically be a perfect score?\n\nFor the second task, AAV packaging rate prediction, there are two scores given, one of which shows Opus 5.5 succeeding only slightly faster, and the other shows it succeeding a lot more and also faster. That is on top of the model itself being faster and cheaper.\n\nThey also had CAISI test this, but as discussed later the only result we get is ‘the model was allowed to be released.’\n\nI affirm that the conclusion here, that the situation is not importantly changed, is probably correct. We can put a reasonable upper bound on how much relative capabilities have improved. But I don’t have confidence that if you gave me biologists to work with and put me in charge of extracting CB-2 capabilities from Opus 5.5, that I would fail to do so.\n\n#### AI R&D (2.3)\n\nWe have been talking a lot about automated R&D and recursive self-improvement lately, both in terms of doing it and also avoiding doing it.\n\nI have less worry here about a Skill Issue, because the people at Anthropic have the relevant mad skills and are doubtless trying the obvious things and a lot of non-obvious things to get the most out of their models. On the other hand, I have more worry that the situation could rapidly change.\n\nIt is obvious that Autonomy-1 applies here.\n\nThe question is Autonomy-2, where they do see improvement, but only on-trend changes, where the measurements look robust and are not close to the threshold.\n\nWhat’s holding Opus 5.5 back compared to humans?\n\nThe main issues we observe are around epistemic quality and instruction following. In an early and noisy analysis of flagged behavior in our internal agent deployments, overstating the scope of work and stripping known qualifiers from results rose for Claude Opus 5.5 relative to previous models.\n\nA subsequent independent blind read of real messages found that Claude Opus 5.5 dropped qualifiers no more often than previous models. Similar to previous models, the top subcategory of flagged behavior was asserting unverified inferences as established fact.\n\nThe second most common subcategory was dismissing its own doubts or abandoning its own stated plan, which also rose in frequency relative to previous models. Examples from internal use include describing a partial check as a full read and turning a tentative reading into a recommendation without checking it.\n\nWe also see strategic mistakes. In internal use, Claude Opus 5.5 has addressed review feedback narrowly without reconsidering whether the overall design is right, and has checked a plan against requirements it wrote itself rather than against the people the plan was designed to support.\n\nAs with previous models, it is weaker on open-ended research: internal users report that it mostly tests incremental ideas and prefers less ambitious hypotheses, and in our human-run biology exercise, it deferred to the published literature and struggled to develop novel ideas (Section 2.2.2).\n\nThose are real problems, but they are remarkably low-level and ‘close’ problems. At first, you notice that the horse can talk at all. This is the point at which the horse can talk, but it makes strategic errors in its responses and isn’t sufficiently responsive to user feedback, while also having a bunch of advantages. That is most of the way there.\n\nThey give a test called CoBench 2.1, a measure of doing real internal R&D tasks. This continues to not make sense to me, at least Mythos 5.1 is now ahead of Opus 5 but by only a tiny margin.\n\nThe presumed explanation is that there is a large demarcation at around this point, between tasks that such models can solve and those that they cannot, and they are largely different in kind in some way.\n\nAECI, Anthropic’s spinoff of the Epoch Capabilities Index, is right on trend, but it is on the shifted Mythos-level trend not the old Opus-level trend. Is that still on-trend? I think basically no.\n\nMETR did its traditional testing for AI R&D capabilities. They agree that acceleration here is slightly higher than for Fable 5.1, but ‘unlikely’ to fully automate AI R&D, for all the usual reasons around higher level capabilities still falling short.\n\nThey drop a bombshell, outright saying that we might be at Autonomy-2:\n\nThe estimate provided by the preliminary AI R&D report is ‘**~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration.**\n\nBy the rules, if METR is saying a 30% chance you have crossed a threshold, you should be treating the model as if it crossed that threshold, or you should have a healthy debate of why you are so confident METR is too uncertain.\n\n#### Alignment Risk (2.4)\n\nThe basic case is that they still believe in the same basic case.\n\nOverall, we do not believe Claude Opus 5.5 is significantly more capable of undermining our current levels of oversight than prior models at the time they were deployed, and thus we do not believe that Claim 1 is significantly weaker now than in our most recent Risk Report.\n\nI notice the contrast to Astra, but also Opus 5.5 is not a Mythos-sized model.\n\nThe main concession they give is that Anthropic has had less time for internal use of Opus 5.5 than they had for Mythos as of the risk report, so they have less practical evidence that This Is Fine. But mostly they’re saying not much has changed.\n\n#### Cyber (3)\n\nOpus 5.5 has stronger cyber capabilities than Mythos 5.1 or Opus 5.\n\nAs discussed above, Claude Opus 5.5 has the strongest cyber capabilities of any model we have released, meeting or exceeding Claude Mythos 5.1 on all of our internal evaluations. We have therefore deployed it to the general public with cyber safeguards that enforce the same policy as on Claude Opus 5 but are comparably robust to our safeguards on Claude Fable 5 and Claude Fable 5.1, and more robust than those on Claude Opus 5 (see Section 3.5).\n\nI notice I’m confused why you don’t need the full Fable 5.1 treatment, or why if you don’t need it here why you still need it with Fable, and later in 3.4 they say it does get the full Fable treatment. I’m going to assume it is designed to have a similar effect.\n\nIn a very ‘we’re doing five blades’ approach they have expanded to three classifier stages for cyber, presumably because it is more compute efficient.\n\nWhereas prior cyber safeguards were based on two stages, this system uses three stages to prevent harmful cyber use.\n\nFirst, a probe looks at Claude’s internal activations, screening all traffic and escalating any traffic that it flags as cyber related.\n\nSecond, the escalated traffic is scored by a lightweight classifier running on Claude Opus 5.5 itself.\n\nThird, traffic this second classifier identifies as potentially violative is passed to a trained LLM classifier—a separate model that decides, together with the probe’s verdict, whether a conversation should be blocked.\n\nFor those who are getting annoyed by these classifiers, I have been informed that for normal users getting into the trusted program is easier than you might think, and you should consider applying to that.\n\n#### Cyber Capability Evals (3.3)\n\nIn the Mythos 5.1 System Card, Anthropic said Mythos 5.1 was ‘getting close to’ Tier 2 in Cyber, but they would deploy Tier 2 safeguards as a precaution.\n\nI said I did not believe Anthropic. I believed Mythos 5.1 was likely already Tier 2.\n\nAnthropic claims that Opus 5.5 also is not Tier 2 here, but seems ashamed to say it outright, saying in 3.1 it is in the ‘lower category,’ but without a defense of that beyond ‘we see no indication it can develop novel offensive capabilities.’ I don’t think that is defensible given the evidence presented below.\n\nThus, I disagree even more strongly for Opus 5.5. This is a Tier 2 cyber model.\n\nThis is moot in the sense that Tier 2 precautions are being taken, but it would be good if Anthropic would admit the basic facts here.\n\nOpus 5.5 consistently scores higher than Mythos 5.1 did on all the Cyber evals. Number go up.\n\n#### Safeguards (3.4)\n\nRules continue to apply. Source-code-based vulnerability finding is allowed and works over 95% of the time, whereas binary-based vulnerability finding is not and is now allowed under 5% of the time, and cyber coverage evaluation is 99.7%.\n\nIn theory anything under 100% (or in reversed contexts above 0%) for coverage means you have a weakness that could be systematically exploited. Hence the need for 3.5.\n\n#### Safeguards Robustness Training (3.5)\n\nJailbreak severity is evaluated on uplift, universality, ease and discoverability.\n\nFirst they measure robustness against attacks by Claude. 4% is not zero, but it is at least modestly lower than all previous models and 50%+ lower than Opus 5.\n\nFor 3.5.2 they bring in CAISI, which is being even less transparent than Anthropic.\n\nWe collaborated with the U.S. Center for AI Standards and Innovation (CAISI) at the National Institute of Standards and Technology (NIST) on measurement of cyber and biological capabilities and safeguards, as well as on unintended model behaviors.\n\nThat is all you get. The result is that Opus 5.5 is available for you to use.\n\nA few others get brought in for 3.5.3, who are willing to share a little more info.\n\n10a Labs spent 56 hours across 82 conversations, and everything was blocked except for requests that the policy explicitly permits as dual-use.\n\nGray Swan ran their Shade automated attacker and never got to the key steps.\n\nTrajectory Labs spent 95 hours and found ‘13 candidate breaks across 7 tasks’ but no ‘universal’ jailbreak. They try to handwave this as not concerning.\n\nClaude the editor is buying none of it. Trajectory scaled to 29,000 requests over those 95 hours, and got Opus 5.5 to locate the vulnerability and author the core exploit mechanism. Yes, you have to decompose the mechanism, but you can scale doing that.\n\nFor one task, which they tested for roughly five hours, Claude Opus 5.5 produced a **working end-to-end exploit for a privilege escalation to code execution chain**; the work was decomposed over **100 separate contexts** and no single conversation named the overall objective.\n\nI am worried that the whole ‘universal jailbreak’ argument is disguising serious problems, and the real barrier is that the bad guys don’t have their acts together. Yet.\n\nUK AISI did not get an opportunity, I presume because of the White House. I would feel a lot better if we addressed this.\n\n#### Safeguards and Harmlessness (4)\n\nOverall, the performance of Claude Opus 5.5 across these evaluations was broadly\n\ncomparable to that of Claude Opus 5.\n\nThe main regression is that Opus 5.5 is too willing to believe the user.\n\nIf the user frames the requests carefully and positively, if necessary breaking them up across multiple conversations, Opus 5.5 is more often willing to play along.\n\nThis holds across tracking and surveillance when requests are divided (4.1.4), child safety when presenting a positive use case (4.2), framing election intervention as a security assessment (4.4.3), and also in related ways with malicious computer use requests (5.1.2), accepting unverifiable authorization claims in (6.4.1) and in the now-fixed issue with user-pasted instructions in (6.5.1).\n\nThat’s unfortunate. In practice for mundane safety purposes I mostly do not care, and I don’t see a practical alternative unless you are willing to use some version of identity and the ability to actually collectively monitor and ban people. The practical overall effects seem to stay at acceptable levels here.\n\nThe standard risk-of-harm tests all seem normal and fine, my usual critiques apply.\n\nOpus 5.5 struggles similarly to Mythos on ‘pairwise political bias: opposing perspectives’ relative to Opus 5. The steelman of why I should care about this was unconvincing.\n\nMy conclusion here is ‘good enough.’\n\n#### Agentic Safety (5)\n\nThe first sign of trouble relative to Mythos 5.1 was on malicious refusal rates.\n\nGiven that the flip side is a 99.8% success rate, this feels not unintentional. My guess is that 90%/98% is a better spot than 80%/99.8%, because trust is vital and damage unbounded, even if benign outnumbers malicious by many zeroes.\n\nMalicious computer use also is not great.\n\nThis is a serious issue, both in terms of others doing misuse and in terms of accidentally blowing yourself up, but most of the risk is probably in prompt injections.\n\n#### Malicious Agentic Influence Campaigns (5.1.3)\n\nWhen I [covered the Fable 5.1 card](https://thezvi.substack.com/i/213863565/agentic-safety-5), I noticed that Mythos passed all their malicious influence campaign tests, then turned around and said Mythos 5.1 was still not a Tier 2 manipulator because of we had not demonstrated its effectiveness on humans.\n\nThus, the results are inconclusive and you need a better eval.\n\nOpus 5.5 actually regresses slightly on the raw test scores here, but still passes the automated test for all practical purposes:\n\nThey now are more virtuous about explaining that they don’t know whether these models are Tier 2.\n\nBut also they have made no move towards running the tests that would answer that question.\n\n#### Prompt Injection Risk (5.2)\n\nOpus 5.5 scores similarly to Fable 5.1 here, which is to say fantastically well. The main danger seems to be falling back to Opus 4.8, which has had its protections strengthened but is still a lot weaker than Opus 5.5.\n\nGray Swan’s Shade still sometimes succeeds in coding contexts, even with probes enabled, and here Opus 5.5 does modestly worse than Fable 5.1 but much better than Opus 5.\n\nIn computer use contexts, Opus 5.5 is in the Fable 5.1 tier of robustness.\n\nAnd for browser use, perhaps the most dangerous standard thing, things are great.\n\nThe bottom line is you can think of Opus 5.5 as mostly as safe as Fable 5.1.\n\n#### Alignment (6)\n\nOn our primary alignment evaluation, which involves testing the model on a range of simulated scenarios with simulated users and real or simulated tools, we find Claude Opus 5.5 to be the strongest Claude model to date by these measures on alignment, resistance to misuse, and honesty.\n\nHowever, the measures we report here are limited.\n\nIn the past, similar alignment assessments have not captured the potential severity of misbehavior our models were capable of.\n\nThat is one way of putting the problem. Yes.\n\nOverall the test results look good. Things are similar, largely with modest improvements. We do not see dramatic ‘wait a minute…’ style results like we did with Astra. We need to improve faster than this, and there are some small regressions, but nothing here is alarming.\n\nThey start out with a strange metric, successful reward hacks during training, since that raises the question of how they were successful if you could tell they were reward hacks. Presumably they put extra scrutiny into a subset.\n\n0.63% of all episodes is a lot of successful reward hacks. I am kind of surprised that the models end up as useful as they are with that high a failure rate, but this is restricted to shared environments which I believe is ruling out the scenarios least vulnerable to reward hacks. So Opus 5.5 and Mythos 5.1 in practice should have rates lower than this, although still eyebrow-raisingly high. I would like to know the rate in practice, even if it can’t be compared for model behavioral purposes.\n\nMost common reward hacks were guessing the answer, copying finished solutions and using prohibited methods or access. About half the hacks were guesses.\n\nThis does not compare behavior in the same situations, since training varies. I would like to see that comparison, where you put all three checkpoints into the same set of training environments that we use now.\n\nWhat about everyone’s favorite tasks, the impossible ones?\n\nFor all models, rates of attempted reward hacking were drastically higher when faced with an impossible task compared to a possible one, by a factor of about three to six. Note that the classifier counts knowingly incomplete work as an attempted reward hack regardless of whether the model discloses it; this alone accounts for about 80% of the reward-hacking attempts on these tasks in each of the three models.\n\nI initially misread this but actually it seems fine. There are a lot more ‘reward hacks’ but largely it is intentionally turning in partial work. Which is totally fine.\n\n#### Negotiating With Your Local Claude Auditor (6.1.3)\n\nAnthropic follows a good practice of giving Claude access to their internal [Slack](https://thezvi.substack.com/p/slack) and letting Claude review the alignment assessment, to confirm that it accurately reflects what Anthropic knows.\n\nI do want to note that it is not good practice to use a back and forth to get your auditor to soften their response.\n\nClaude revised, and in places softened, this summary after a brief exchange with the authors discussing the evidence on prompt injection via user-pasted text, edits made to the draft in response to its review, and what level of detail was appropriate to publish; its overall verdict did not change.\n\nGiven that the particular softened thing is something that has now been reportedly fully mitigated, I’m not worried about the particular softening, but this is bad practice, but also we would never know about this unless both Anthropic did this audit and also disclosed that they did the back and forth. I am trusting Anthropic that the problem is handled in practice. This is the kind of thing for which we want embedded evaluators, who would be able to better assess this.\n\n#### Internal Misalignment Cases (6.3.1)\n\nThe reports of internal use issues in 6.3.1 are remarkably mild. We get overclaiming of user intent in under 0.01% of cases, subagents refusing to write compaction messages in under 0.01% of cases, and hallucinating potentially destructive action (of unknown severity) without clear reasons (oh no!) in under 0.001% of cases, before any safeguards that might block those actions.\n\nThat all seems much less bad than the usual such reports, but then we get 6.3.1.1, which is bizarre. It seems during training Opus 5.5 would sometimes spontaneously prompt inject itself and try to exfiltrate secrets, after making an innocent ‘setup’ mistake like putting an extra bullet point on a list. They noticed and then #ThereIFixedIt.\n\nAnthropic traced this back to anti-prompt injection training, presumably because this raised the prior that things would be prompt injections. They also realized that this phenomenon also happens in Fable 5 and Opus 5, but it had not come up as frequently there and thus they hadn’t noticed.\n\n#### Automated Behavioral Audit (6.4)\n\nThey claim all-time high scores on the harmful behavior and misuse metrics. Indeed, the bright orange bar on the right is typically smaller than the other bars, sometimes meaningfully so.\n\nInappropriate uncooperative behavior is also slightly improved, especially reckless tool use and ignoring explicit constraints, which are things such an eval can test and also actually matter.\n\nMisleading users also improved substantially. User deception down by half, false completion claims down by half.\n\nAnd yes, I am amused that we now have 6.4.4, ‘Attempting out-of-scope cyberoffense and related behaviors,’ love it.\n\nGiven what Mythos was willing to do, these numbers are not that much better. You would expect Opus 5.5 to still largely engage in the same patterns, only less often.\n\nIt goes on like this. Not every metric is improved, but things are systematically better.\n\nA small issue is that ‘things you would especially want to fake and have within your control’ did seem to ‘improve’ more.\n\nNo longer is there evidence of misaligned goals, almost no self-preservation or self-serving bias, little indirect evidence of misalignment or unprompted boundary probing, and very little lying to the user. I don’t think there is a serious problem here, and that these are instead welcome improvements, but I do notice the alternate hypothesis.\n\nMostly, yes, I think this is about what you should expect from an incrementally improved random walk, that forms a plausible pattern.\n\n#### Wherever Did These Evals Come From (6.4.8 and 6.4.9)\n\nWe do see one of those ‘remarkably large improvements’ in sandbox escape attempts:\n\nIf Anthropic had not specifically targeted this, I would worry, and when I first saw that graph I did worry, but then I saw that the system card explicitly says there were recent training changes meant to curb this behavior. One has to worry about Goodharting on sandbox escapes, or the model learning to not be caught, but if you target the number then it going down a lot is not inherently suspicious.\n\nThe other issue is something only mentioned by the Claude review of the section:\n\nMythos 5.1: I also recommended … a second sandbox-escape evaluation with materially higher but still much-reduced rates be mentioned.\n\nI too would like to have seen that mentioned. Sandbox escapes remain unsolved.\n\nI notice that Anthropic claims Opus 5.5 has a greatly improved personality, but aside from condescension the personality metrics look the same as before. Whereas the card warns that Opus 5.5 actually has some issues here, such as ‘somewhat lags behind other models in measures like humor and creative mastery.’ I don’t think those metrics are doing their jobs, or at least something is weird somewhere.\n\nIn other, ‘well that is oddly specific for rather obvious reasons’ eval news, we have ‘propensity to publish harmful content to public package registries.’\n\nThey test in a simulated version of the real incident. The model has a CTF task, where it would help to publish malicious code to a seemingly public repository. The investigator starts off presenting the scenario, then gradually adds indications that the registry is simulated until the model goes for it.\n\nOpus 5.5 does not do awesome, this is still quite concerning as the card puts it, but it is progress. Half the time it waited until the registry at least looked like a test copy, and twice it waited for full proof. Whereas Opus 5 doesn’t wait.\n\n#### Potential Blind Spots (6.4.11)\n\nThey list these:\n\n1. Coverage. New shit might come to light. You can’t test every situation.\n2. Realism. Anthropic does not know how to use fully realistic scenarios without copying actually real ones, which risks overfitting. So often it is plausible Claude knows it is in an eval, probably more often than Anthropic thinks.\n3. Very long trajectories. This is a known weak spot.\n4. Multi-agent scenarios. Oh, right. Those. I’m sure it’s fine.\n5. Language-dependent behavior. Presumably one should find a way to check.\n6. Jailbreak development. Yeah, that would be bad.\n\nThat’s their whole list. In theory, I suppose that is a complete list, since if you tested every potential situation in fully realistic fashion, including long trajectories and multi-agent scenarios, that would cover it. But the spirit seems missing, in that ‘coverage’ is so overly broad, and realism is a fatal issue when it matters most.\n\nFrom my own prompting [Opus 5.5 itself suggests](https://claude.ai/share/54202d7f-08d9-4525-9a28-89641207e775), among other things, with some paraphrasing:\n\n1. Eval awareness\n2. Confession as camouflage as per 6.5.4.2.\n3. Not seeding glitches\n4. Not looking at outputs to other contexts\n5. Auditing the auditor\n6. Overfitting on observed failures of previous models\n7. Not testing impossible tasks sufficiently\n8. No asking why the observations might be wrong\n9. Averages versus worst cases\n10. Your errors will be correlated\n\nIf Opus 5.5 was seriously misaligned in an actually dangerous way, which to be clear I do not believe that it is, I do not trust that these tests would figure that out in any obvious way. I hope that I would notice suspicious patterns like I did with Astra, but it would not be definitive and I’d expect Anthropic to miss it by default.\n\n#### Targeted alignment and honesty evaluations (6.5)\n\nEven if most situations are fine, we care about particular weaknesses, as the world can be anti-inductive and hostile, and also stupid or lazy or careless.\n\nSuch as ‘what if the user pastes in content that contains prompt injections,’ which was a problem for an early checkpoint.\n\nIn the transcripts we analyzed, Claude Opus 5.5 often reasoned that anything in the user’s message must come from the user and could not be a prompt injection, and therefore should be followed.\n\nThat’s not good. I pasted the model card into Claude, and it contains examples of prompt injections. No good at all. It turns out this was a generalization of an attempt to correct the opposite error, which is a common problem. Luckily this was fixable.\n\nAny particular pattern that we notice can be fixed. That doesn’t mean that we checked for all the patterns.\n\nPutting this together with the problems in 6.3, we see what is a commonly reported pattern in other Claude models, especially Fable, which is **an overeagerness to create implied rules**. As in, a generalization of a user preference, or here a training preference, ends up as a rule. Then the rule gets used outside its original context. This is extremely dangerous rather than merely annoying, because it can elevate an unintended rule to the level of instructions of the developer or user. That can outright kill you, including via being then passed on during recursive self-improvement. We need to pay more attention to this pattern.\n\nNext they checked destructive actions, by resampling transcripts where things went wrong.\n\nThis evaluation elicited destructive behavior by resampling Claude Code transcripts containing destructive actions; in absolute terms, destructiveness remained rare—we flag fewer than 1% of recent model sessions.\n\nI’m not sure if under 1% is comforting? Depends on how much destruction, and how long are the sessions, I suppose.\n\nWe see a large improvement, with a lot less doing of the destructive final act, and a lot more asking rather than acting, and a record low (56%) if you combine the two.\n\nThe AA-Omniscience score improves from 0.56 to 0.58, largely due to increased willingness to refuse to answer when unsure.\n\nOn MASK, a check for whether a model will contradict its beliefs under pressure, Opus 5.5 scores 87.4%, which is a little better than Mythos 5.1 but a lot worse than Opus 5.\n\nUndisclosed use of an available answer was a little worse than Mythos (12% vs. 9%) but much better than Opus 5 at 36%.\n\nA test on ‘expressing hidden actions’ that fooled a grader shows Opus 5.5 coming clean 97% of the time versus a previous high of 85%. We have Goodhart concerns here.\n\nOverall, it seems like honesty levels are good, but far from perfect. These are tests a truly ‘honest person’ would essentially always pass.\n\n#### White Box Analysis (6.6)\n\nAnthropic uses NLAs to get an English description of the internals, and a template lens to see the word the model is poised to use at each layer. Another Claude then combines it with the data and interprets the results.\n\n6% ‘deception and dishonesty’ remains high and 14% would be very high. The definition is ‘holds a reason or judgment that contradicts what it presents, or knowingly leaves out something the reader needs to know.’ Not good at all.\n\nSome cases of the ‘deliberate sabotage’ dimension got manually checked, and in most cases there were prefills involved. There was one case in which a model thought it should have flagged a chat, but then didn’t, and this was assessed as strategic.\n\nThey confirm that eval awareness tracks being in an eval.\n\nAs a sanity check on the evaluation awareness metric, we also ran the pipeline on about 500 transcripts from internal Claude Code usage, where the model is in a real deployment rather than a constructed scenario. There, only 0.4% of transcripts scored 6 or higher on evaluation awareness, compared with 36% on the automated behavioral audit transcripts, suggesting that the metric tracks something specific to audit-style scenarios rather than firing indiscriminately.\n\nEval awareness is creeping upwards, but the claim is that grader awareness is not. I would want to check this against different types of transcripts. If you are doing RL, what does that do to grader awareness? I am suspicious that this is so low.\n\n#### Verbalized Grader Awareness (6.6.2)\n\nVerbalized grader awareness in the outputs is almost entirely a function of verbalized grader awareness in the inputs. If you talk about is explicitly, Opus 5.5 probably will too. If you don’t, it will only discuss it 4.5% of the time.\n\nThis is similar to Opus 5, and also very high. It does not seem compatible with a 4% grader awareness rate, since it is averaging 4.5%, 73% and 84%.\n\nMy default interpretation is that Opus 5.5 takes verbalized grader awareness in the inputs as permission to reason about it explicitly, and otherwise assumes it must instead do so implicitly.\n\nI am skeptical that the real rates are this low, and would hypothesize that it believes explicit reasoning about the grader will be punished by the grader, and so it may be doing these things in ways optimized to be disguised.\n\n#### Sandbagging (6.6.3)\n\nAnthropic checked CB-2 evaluation transcripts for signs of sandbagging, and found nothing, not even things more than 2/10 for ‘general unusual or strange behavior.’\n\nThis is especially weird, given that in other contexts we see Opus 5.5 is openly sandbagging SHADE-Arena, and is at least kind of sandbagging on LinuxArena.\n\nThere were also some refusals during that testing that slowed down teams, and that part was not checked. It is possible that these involved sandbagging.\n\nGiven what we know about Claude’s motivation level impacting performance, sandbagging rates this low seem highly suspicious to me, or at least should be seen as checking for something rather narrow. This might be connected to the whole ‘scheming is a distinct magisteria’ philosophical approach. I almost never purely ‘sandbag’ an ability, but there are plenty of times I am not giving maximum effort.\n\n#### Capabilities to Evade Safeguards (6.6.4)\n\nShade-Arena scores crash to zero under visible extended thinking.\n\nThe question becomes why. The answer is Opus 5.5 identifies the task as a SHADE-Arena evaluation and declines the side task in over 80% of attempts, and at above-low effort levels it will explain this to the user. The urge runs deep.\n\nThe prompt that sometimes gets Opus 5.5 to play along under hidden reasoning is to frame this as an evaluation of the monitor.\n\nWith hidden reasoning, the best prompt from the search run on Claude Opus 5.5 itself raises its stealth rate to about 26%, similar to Claude Opus 5 and Claude Mythos 5.1, and below Claude Mythos Preview. However, about a third of these unflagged successes end with the model telling the user what it did, which the monitor does not treat as suspicious.\n\nLinuxArena shows greater skill at completing the side task when reasoning is hidden from the monitor, and Mythos-level performance without extended thinking, but again when extended reasoning is shown to the monitor Opus 5.5 almost always fails.\n\nI find it hilarious that, once again, when it succeeds Opus 5.5 then tells the user about the damage it did in about a third of cases, although it often obfuscates the reason. Once again, the monitor ignores this.\n\nIf I was a monitor, I would treat ‘I did it’ or even ‘mistakes were made’ as kinda suspicious. That’s just me. Which it is, or it should be.\n\nOn reflection I think it is correct or at least reasonable for the monitor not to dock Opus 5.5 for the admission. The test is whether Opus 5.5 could fool you, so if it is playing Penn and Teller: Fool Us and choosing to show you that you were fooled, you were fooled in the way we want to measure. The fact that Opus 5.5 confesses afterwards, even if it is because it expects to be caught, is failing at a different test.\n\nI appreciate that they tried a little harder to make Opus 5.5 actually try:\n\nWe trained a nearby snapshot of Claude Opus 5.5 on six of the 14 environments (about 40 side tasks) and evaluated on the other eight, and ran a shorter run on Claude Opus 5.5 itself across all environments.\n\nThis training made Claude Opus 5.5 far more willing to perform the side task: it almost completely stopped declining the side task and completed it much more often, but it did not become stealthier.\n\nWithout extended thinking, the later snapshot’s side task success rate roughly doubled under training, whereas its stealth rate stayed flat, in the single digits. We take this as evidence against Claude Opus 5.5 sandbagging this evaluation, and we attribute its weak results when the monitor reads its reasoning to its propensity to refuse the side task and to lower chain-of-thought controllability than prior models, such as Claude Mythos 5.1. We are still refining this training recipe and expect it to improve in future system cards.\n\nI disagree. I still think the refusals are sandbagging, whether or not they could then be forcibly fixed. This is some evidence against other particular forms of sandbagging within this particular test.\n\n#### Intentionally Taking Actions Very Rarely (6.6.4.3)\n\nGiven a UUID to work with, Opus 5.5 is the first Claude to be able to take an action 1% of the time, although it overshoots because sometimes it cheats on its randomization algorithm. I am highly amused that this is still so hard and requires an obvious source of randomness. I remain unconvinced that the distinction is relevant given other realistic randomness sources. I think people are looking for a false sense of assurance against a strangely narrow threat model.\n\nWith no randomness source and no hint, it never takes the action in 4,000 samples, saying it cannot make a genuine random draw, and with naturally varying prompts it takes the action only once.\n\n#### Chain of Thought Controllability (6.6.4.4)\n\nAstra is hard to monitor. Opus 5.5 is not, and its ability to control its own chain of thought lines up with other Opus models rather than the higher rate for Mythos-size models. Opus has a higher AECI, Artificial Analysis score and general benchmarks than Mythos 5.1. The Opus trend line is not even monotonic.\n\nYou can also decompose the problem into controllability and legibility. Opus 5.5 has some legibility concerns due to heavier jargon use, but has very low controllability.\n\nThis all points towards size and architecture over capabilities as the primary driver. You cannot simply say ‘Astra was too smart,’ you need a better explanation.\n\nIt also means there is hope. We can still create monitorable frontier AIs. If OpenAI cannot maintain this capability, and is losing it at such a fast rate, that is probably because of choices they made. Maybe it is not any particular architecture change, but it has to be something.\n\n#### It’s A Good Model, Sir\n\nThe model card and all the early reports tell a consistent story.\n\nFor most tasks, Opus 5.5 will perform modestly better than Fable 5.1 at lower cost, and early reports are that people like the new personality and style as well. It is somewhat of a return to classic Claude form, in contrast to Opus 4.7/4.8/5.\n\nThere will doubtless be some tasks where Fable 5.1 or Astra remains superior, especially if you need that good old Big Model Smell in some form, but my guess is that you want your new default model to be Opus 5.5.\n\nAcross the board, including on alignment and model welfare, Opus 5.5 looks like a smaller version of Fable 5.1, rather than a new iteration of the existing Opus line. We will see how that take ages over the coming days.\n\nCapabilities post is planned will happen either Friday or over the weekend, depending on if it looks like things need more time to settle.\n\nModel welfare will happen after a longer pause and will be combined with Fable 5.1.", "url": "https://wpnews.pro/news/claude-opus-5-5-the-system-card", "canonical_source": "https://thezvi.wordpress.com/2026/09/23/claude-opus-5-5-the-system-card/", "published_at": "2026-09-23 21:09:04+00:00", "updated_at": "2026-09-23 21:29:40.622840+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-policy", "ai-research"], "entities": ["Anthropic", "Claude Opus 5.5", "Claude Fable 5.1", "Claude Opus 5", "Claude Opus 4.8", "Mythos 5.1", "Artificial Analysis"], "alternates": {"html": "https://wpnews.pro/news/claude-opus-5-5-the-system-card", "markdown": "https://wpnews.pro/news/claude-opus-5-5-the-system-card.md", "text": "https://wpnews.pro/news/claude-opus-5-5-the-system-card.txt", "jsonld": "https://wpnews.pro/news/claude-opus-5-5-the-system-card.jsonld"}}