cd /news/large-language-models/claude-opus-5-5-4-breaking-changes-h… · home › topics › large-language-models › article
[ARTICLE · art-148630] src=byteiota.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Claude Opus 5.5: 4 Breaking Changes Hitting Production

Anthropic's Claude Opus 5.5, shipped September 22, is 40% cheaper than Opus 5, 30% faster, and beats GPT-6 Astra on benchmarks at one-third the cost, but four API changes throw 400 errors when teams flip the model ID. Thinking can no longer be disabled, forced tool use via tool_choice {"type": "tool"} or {"type": "any"} returns a 400, thinking blocks are cryptographically bound to the exact system prompt, tool set and message history that produced them, and the computer_20251124 tool is rejected on the Claude API and Google Cloud while Amazon Bedrock still accepts it. Cache reads fell 60% from $0.50 to $0.20 per million tokens, though part of the savings comes from Anthropic dropping the default effort level from high to medium.

read5 min views5 publishedOct 10, 2026
Claude Opus 5.5: 4 Breaking Changes Hitting Production
Image: Byteiota (auto-discovered)

Claude Opus 5.5 shipped September 22 and the pricing story is genuinely good: 40% cheaper than Opus 5, 60% off cache reads, 30% faster, and benchmark scores that beat GPT-6 Astra at one-third the cost. The catch is that four API changes will throw 400 errors the moment you flip the model ID. Three weeks in, teams are still hitting them. Here is what breaks and how to fix it.

The 4 Breaking Changes #

1. Thinking Cannot Be Disabled

If you set thinking: {"type": "disabled"} anywhere in your Opus 5 code, Opus 5.5 rejects the request immediately with: “thinking.type.disabled” is not supported for this model. Adaptive thinking is always on. The fix is to remove the field entirely — or send {"type": "adaptive"}, which does the same thing. Control depth with the effort parameter instead.

Two side effects matter: thinking tokens eat into your max_tokens budget (increase it), and responses can now start with a thinking block, so stop accessing content[0] directly. Filter by type:

text_blocks = [b for b in response.content if b.type == "text"]

2. Forced Tool Use Returns a 400

tool_choice: {"type": "tool"} and {"type": "any"} both fail on Opus 5.5. Replace them with auto plus "strict": true on each tool definition, and name the tool explicitly in the prompt. Verify the call happened — auto means the model can still decline:

response = client.messages.create(
    model="claude-opus-5-5",
    tool_choice={"type": "auto"},
    tools=[{**tool, "strict": True} for tool in tools],
    messages=[{"role": "user",
               "content": "Use the get_weather tool for Paris."}]
)
if not any(b.type == "tool_use" for b in response.content):

3. Thinking Blocks Are Bound to the Conversation

This one catches teams off guard. Thinking blocks are cryptographically tied to the exact system prompt, tool set, and message history that produced them. Edit the history — trim a tool result, inject a timestamp into your system prompt, swap a tool mid-session — and you get: Invalid signature in thinking block. The block is bound to a different conversation. This applies to API accounts created after August 31, 2026.

The rule is strict: keep message history append-only. If edits are unavoidable, use the thinking-binding-controls-2026-08-01 beta with "prefix_mismatch_behavior": "drop_block" so mismatched blocks are dropped instead of throwing. The most common hidden trap: system prompts with injected timestamps. The prompt looks identical to your code but produces a different hash on every call.

4. The Old Computer Tool Is Rejected on Claude API and Google Cloud

computer_20251124 throws an error on the Claude API and Google Cloud. Amazon Bedrock still accepts it, so Bedrock users have time to migrate at their own pace. For everyone else, switch to computer_toolset_20260801 — no name or display dimensions needed. Actions come back in the tool_use block’s name field (e.g., "left_click"), multiple actions can appear per turn, and you resize screenshots yourself.

One Non-Breaking Change Worth Knowing #

Text between tool calls no longer streams to users by default — it moves into hidden thinking blocks. If your agent surfaces progress updates mid-task, those go silent on Opus 5.5 until you opt in with thinking={"type": "adaptive", "display": "updates"} (requires the thinking-display-updates-2026-08-18 beta header). No errors, but users will notice the silence.

What 40% Cheaper Actually Means #

The headline is accurate but the composition matters. Base token prices dropped 20%. Cache reads dropped 60% — from $0.50 to $0.20 per million tokens — which is the real win for multi-turn agentic loops where persistent context dominates cost. The rest of the saving comes from Anthropic dropping the default effort level from high to medium. That shift is not apples-to-apples against Opus 5, and community benchmarking has flagged one regression: concurrency bugs run about 44% higher, which warrants adding static analysis to CI pipelines if you are not already running it.

For cache-heavy agentic workloads, independent benchmarking puts real savings between 40% and 51% — consistent with the headline. For simple chat workloads without caching, expect closer to 20%.

Performance: Where the Benchmarks Land #

At default medium effort, Opus 5.5 scores 52.5% on Terminal-Bench 4.0 — matching Fable 5.1’s range at roughly one-third the cost. Push to xhigh and it reaches 66.4%, clearing Fable 5.1 by over 10 points. On FrontierCode v1.1, it hits 54.6% against GPT-6 Astra’s 53.3% peak at 80% less spend. The knowledge-work benchmark GDPval-AA v2.1 shows a 300-point Elo gap over Astra: 1846 versus approximately 1546.

The practical takeaway: Opus 5.5 at medium effort beats Opus 5 at high effort for coding tasks. Set effort explicitly — the default change is intentional but silent, and your production costs will shift if you do not account for it.

Migration Checklist #

Run these before flipping the model ID in production. The recommended strategy: validate all changes on Opus 5 first (it accepts all Opus 5.5-compatible patterns), then swap the model ID as a single isolated config change for easy rollback:

  1. Swap model ID to claude-opus-5-5
  2. Remove thinking: {"type": "disabled"} and anybudget_tokens fields
  3. Set output_config.effort explicitly on every call
  4. Increase max_tokens to account for thinking overhead
  5. Filter response content by block type; pass thinking blocks back unmodified
  6. Replace tool_choice: "any" and"tool" withauto plusstrict: true
  7. Update computer-use agent loops (Claude API and Google Cloud only)
  8. Keep message history strictly append-only
  9. Add display: "updates" if users rely on progress text between tool calls
  10. Handle stop_reason: "refusal" — newbio andreasoning_extraction categories added

The economics are solid, the benchmarks are real, and the migration is a half-day of work for most integrations. The teams getting burned are those who assumed a model ID swap would be clean. It is not — but this checklist makes it predictable. Full details in the Anthropic announcement, and the official migration guide covers every edge case.

── more in #large-language-models 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-opus-5-5-4-br…] indexed:0 read:5min 2026-10-10 · —