Anthropic released Claude Haiku 5.5 today, cutting prices by up to 90% for most workloads, expanding the context window to 1M tokens, and shipping the first effort toggle on a Haiku-class model. If you run high-volume Claude calls, today’s release changes your cost math immediately.
The Pricing Is Tiered — Know Your Threshold #
Haiku 5.5 uses a two-tier pricing structure split at 100,000 tokens per prompt:
- Under 100K tokens : $0.10/M input, $0.50/M output — 90% cheaper than Haiku 4.5
- Over 100K tokens : $0.50/M input, $2.50/M output — 50% cheaper than Haiku 4.5
Anthropic estimates the average workload cost reduction at 75% — a number that reflects real usage patterns, not just the headline rate. Most classification, routing, extraction, and summarization calls sit well under 100K tokens. If that describes your workload, the savings are real.
The nuance worth flagging: GPT-6 Luna charges the same $0.10/$0.50 rate, but its lower-price threshold extends to 272,000 tokens versus Haiku 5.5’s 100,000. If your prompts regularly land between 100K and 272K tokens, GPT-6 Luna has a pricing edge in that range. Check your token distribution before treating the 90% headline as your actual savings. According to VentureBeat’s coverage, Anthropic is explicitly targeting Luna parity on the sub-100K tier.
The Batch API adds another 50% off both tiers. For offline jobs — nightly classification runs, bulk document processing, async pipelines — that reduction stacks on top of the already-lowered rates.
The Effort Toggle: Tune Intelligence Per Call #
Haiku 5.5 is the first Haiku-class model with the effort parameter — set via output_config.effort in your API requests. The levels run from none and minimal up through low, medium, high, and xhigh, with medium as the default.
This matters because it lets you tune intelligence per call rather than per model. A classification step can run at low effort. A complex subagent task feeding into Opus 5.5 can use high or xhigh. Previously, getting that granularity meant routing to a different model entirely. Now a single model ID handles the full spectrum in one agentic pipeline.
The performance difference is measurable. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% at max effort versus roughly 20% at medium. For high-stakes agentic steps, the higher effort levels close most of the gap with larger models at a fraction of the cost. Haiku 4.5 scored 0% on the same benchmark — this is a genuine capability step, not a minor refresh.
Five Breaking Changes That Will Throw 400 Errors #
Haiku 5.5 is not a drop-in replacement. These five changes cause immediate API errors if left unaddressed, per Anthropic’s migration guide:
- Thinking config : Replace
thinking: {"type": "enabled", "budget_tokens": N}withthinking: {"type": "adaptive"}and setoutput_config.effortinstead. - Sampling parameters : Remove
temperature,top_p, andtop_k. Non-default values return 400. - Assistant prefill : End your
messagesarray with a user turn, not an assistant turn. - Computer use : Replace
computer_20250124withcomputer_toolset_20260801on the Claude API and Google Cloud. - Cross-account thinking replay : Thinking blocks now work only in the account that produced them. Multi-account conversation replay must route through the originating account.
There is a sixth change that won’t throw an error but will silently break your cost estimates: Haiku 5.5 uses the newer tokenizer introduced with Claude 4.7. The same text produces approximately 30% more tokens than on Haiku 4.5. Recount your prompts, revisit max_tokens limits, and recalculate cost projections before your first production call. Despite the higher token count, effective cost still drops ~75% due to the rate reduction.
The automated migration path: run /claude-api migrate this project to claude-haiku-5-5 in Claude Code. The bundled Claude API skill applies the model ID swap and breaking parameter changes across your codebase, then produces a checklist for manual review.
Context Window, Benchmarks, and API Credits #
The context window expanded from 200K to 1M tokens, and max output jumped from 64K to 128K. On the official Claude Haiku 5.5 overview, Anthropic notes that thinking tokens count toward max_tokens — if you had a tight limit set for Haiku 4.5, raise it after migrating.
On benchmarks, Haiku 5.5 clears GPT-6 Luna on the evals most relevant to agentic work: OSWorld 2.1 at 72.4% versus Luna’s 48.9%, Terminal-Bench at 39.2% versus 16.4%, and GDPval-AA v2.1 at 1,620 versus Luna’s 1,437. Box reports an 11-point improvement over Haiku 4.5 at half the latency. Asana sees 2.5x faster inference per agent turn.
Alongside Haiku 5.5, Anthropic is rolling out monthly API credits for Max and Team subscribers this week: $100/month for Max 5x plans, $200/month for Max 20x, and up to $500 pooled for Team plans, usable on any Claude model.
The model ID is claude-haiku-5-5 across all platforms — Anthropic API, Amazon Bedrock (anthropic.claude-haiku-5-5), Google Cloud, and Microsoft Foundry — available today.