# Claude Haiku 5.5 Ships Today: 90% Cheaper, Effort Toggle, Five Breaking Changes

> Source: <https://byteiota.com/claude-haiku-5-5-pricing-effort-toggle/>
> Published: 2026-10-07 19:11:11+00:00

Anthropic released Claude Haiku 5.5 today, cutting prices by up to 90% for most workloads, expanding the context window to 1M tokens, and shipping the first effort toggle on a Haiku-class model. If you run high-volume Claude calls, today’s release changes your cost math immediately.

## The Pricing Is Tiered — Know Your Threshold

Haiku 5.5 uses a two-tier pricing structure split at 100,000 tokens per prompt:

- **Under 100K tokens** : $0.10/M input, $0.50/M output — 90% cheaper than Haiku 4.5
- **Over 100K tokens** : $0.50/M input, $2.50/M output — 50% cheaper than Haiku 4.5

Anthropic estimates the average workload cost reduction at 75% — a number that reflects real usage patterns, not just the headline rate. Most classification, routing, extraction, and summarization calls sit well under 100K tokens. If that describes your workload, the savings are real.

The nuance worth flagging: GPT-6 Luna charges the same $0.10/$0.50 rate, but its lower-price threshold extends to 272,000 tokens versus Haiku 5.5’s 100,000. If your prompts regularly land between 100K and 272K tokens, GPT-6 Luna has a pricing edge in that range. Check your token distribution before treating the 90% headline as your actual savings. According to [VentureBeat’s coverage](https://venturebeat.com/technology/anthropic-launches-claude-haiku-5-5-with-90-api-price-reduction-matching-gpt-6-luna), Anthropic is explicitly targeting Luna parity on the sub-100K tier.

The Batch API adds another 50% off both tiers. For offline jobs — nightly classification runs, bulk document processing, async pipelines — that reduction stacks on top of the already-lowered rates.

## The Effort Toggle: Tune Intelligence Per Call

Haiku 5.5 is the first Haiku-class model with the `effort` parameter — set via `output_config.effort` in your API requests. The levels run from `none` and `minimal` up through `low`, `medium`, `high`, and `xhigh`, with `medium` as the default.

This matters because it lets you tune intelligence per call rather than per model. A classification step can run at `low` effort. A complex subagent task feeding into Opus 5.5 can use `high` or `xhigh`. Previously, getting that granularity meant routing to a different model entirely. Now a single model ID handles the full spectrum in one agentic pipeline.

The performance difference is measurable. On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% at max effort versus roughly 20% at medium. For high-stakes agentic steps, the higher effort levels close most of the gap with larger models at a fraction of the cost. Haiku 4.5 scored 0% on the same benchmark — this is a genuine capability step, not a minor refresh.

## Five Breaking Changes That Will Throw 400 Errors

Haiku 5.5 is not a drop-in replacement. These five changes cause immediate API errors if left unaddressed, per [Anthropic’s migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide):

1. **Thinking config** : Replace`thinking: {"type": "enabled", "budget_tokens": N}` with`thinking: {"type": "adaptive"}` and set`output_config.effort` instead.
2. **Sampling parameters** : Remove`temperature` ,`top_p` , and`top_k` . Non-default values return 400.
3. **Assistant prefill** : End your`messages` array with a user turn, not an assistant turn.
4. **Computer use** : Replace`computer_20250124` with`computer_toolset_20260801` on the Claude API and Google Cloud.
5. **Cross-account thinking replay** : Thinking blocks now work only in the account that produced them. Multi-account conversation replay must route through the originating account.

There is a sixth change that won’t throw an error but will silently break your cost estimates: Haiku 5.5 uses the newer tokenizer introduced with Claude 4.7. The same text produces approximately 30% more tokens than on Haiku 4.5. Recount your prompts, revisit `max_tokens` limits, and recalculate cost projections before your first production call. Despite the higher token count, effective cost still drops ~75% due to the rate reduction.

The automated migration path: run `/claude-api migrate this project to claude-haiku-5-5` in Claude Code. The [bundled Claude API skill](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/claude-api-skill) applies the model ID swap and breaking parameter changes across your codebase, then produces a checklist for manual review.

## Context Window, Benchmarks, and API Credits

The context window expanded from 200K to 1M tokens, and max output jumped from 64K to 128K. On [the official Claude Haiku 5.5 overview](https://platform.claude.com/docs/en/models/haiku-5-5/overview), Anthropic notes that thinking tokens count toward `max_tokens` — if you had a tight limit set for Haiku 4.5, raise it after migrating.

On benchmarks, Haiku 5.5 clears GPT-6 Luna on the evals most relevant to agentic work: OSWorld 2.1 at 72.4% versus Luna’s 48.9%, Terminal-Bench at 39.2% versus 16.4%, and GDPval-AA v2.1 at 1,620 versus Luna’s 1,437. Box reports an 11-point improvement over Haiku 4.5 at half the latency. Asana sees 2.5x faster inference per agent turn.

Alongside Haiku 5.5, Anthropic is rolling out monthly API credits for Max and Team subscribers this week: $100/month for Max 5x plans, $200/month for Max 20x, and up to $500 pooled for Team plans, usable on any Claude model.

The model ID is `claude-haiku-5-5` across all platforms — Anthropic API, Amazon Bedrock (`anthropic.claude-haiku-5-5`), Google Cloud, and Microsoft Foundry — available today.
