Claude Haiku 5.5 pricing jumps fivefold at 100,001 prompt tokens Anthropic's Claude Haiku 5.5, released October 7, 2026, prices prompts up to 100,000 tokens at $0.10 per million input tokens but jumps fivefold to $0.50 per million — and $2.50 per million output — for any prompt over that line, with the higher rate applied to the entire request rather than just the excess. A Go-based analysis of Anthropic's published pricing found the threshold is effectively lower for migrating users because Haiku 5.5's newer tokenizer produces roughly 30% more tokens for the same text, putting the practical cutoff near 76,900 Haiku 4.5 tokens, and warned that the free token-counting endpoint is only an estimate. A request to Claude Haiku 5.5 with a 100,000-token prompt and 2,000 tokens of output costs $0.011. Add one token to the prompt and the same request costs $0.055. Nothing else changed; you crossed a line in the price list. Anthropic released Haiku 5.5 on October 7, 2026, and as of October 8 it's available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS under the ID claude-haiku-5-5 anthropic.claude-haiku-5-5 on Bedrock . The headline is the price: $0.10 per million input tokens against $1 for Haiku 4.5. The catch is that the cheap rate is for "prompts up to 100,000 tokens", and the model takes prompts ten times that size. A Hacker News discussion of the release https://news.ycombinator.com/item?id=49996437 went back and forth on whether agent workloads can stay under the line and didn't settle it. So we did the arithmetic with a small Go program. We had no API key in this run, so nothing below was sent to the model: the prices come from Anthropic's documentation and the costs are computed from them. The pricing page https://platform.claude.com/docs/en/about-claude/pricing gives Haiku 5.5 two rows, one "for prompts up to 100,000 tokens" and one "for prompts over 100,000 tokens". Per million tokens: | | Up to 100,000 | Over 100,000 | Haiku 4.5 | |---|---|---|---| | Input | $0.10 | $0.50 | $1.00 | | Output | $0.50 | $2.50 | $5.00 | | Cache read | $0.01 | $0.05 | $0.10 | | Cache write, 5 minutes | $0.125 | $0.625 | $1.25 | | Cache write, 1 hour | $0.20 | $1.00 | $2.00 | | Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 | $0.50 / $2.50 | Every rate in the second column is five times the first. That includes output, which is the part people miss: the length of what you send decides the price of what comes back. This isn't a marginal bracket like income tax, where only the excess pays the higher rate. The row is chosen per prompt and applies to the whole request. It's also unique in the current lineup. For the other 1M-context models the same page says "A 900k-token request is billed at the same per-token rate as a 9k-token request", and it names Haiku 5.5 as the exception. "Prompt" is doing a lot of work in that table, and the documentation doesn't define it as a formula over the usage fields. Here's what it does say. With prompt caching on, the input is split across input tokens , cache creation input tokens and cache read input tokens . The context windows page https://platform.claude.com/docs/en/build-with-claude/context-windows is blunt about cached tokens: "prompt caching changes what you pay for those tokens, not whether they count." So we treat the prompt as the sum of all three, which is the reading where a 100,000-token cached prefix plus a 50-token question lands in the expensive tier. That's our reading, not a quoted rule. If your bill depends on it, check a real invoice against a request you know the size of. Two more things move the line. Haiku 5.5 uses the newer tokenizer, and the migration guide says the same text "produces approximately 30% more tokens" than on Haiku 4.5. If you know your prompt sizes in Haiku 4.5 tokens, the threshold isn't 100,000 for you. It's about 76,900. And the pre-flight check is approximate. The token counting endpoint https://platform.claude.com/docs/en/build-with-claude/token-counting is free and takes the same body as a message request, but the docs call its result an estimate that "might differ by a small amount". Don't route on <= 100000 . Leave headroom. curl https://api.anthropic.com/v1/messages/count tokens \ -H "x-api-key: $ANTHROPIC API KEY" \ -H "content-type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{"model": "claude-haiku-5-5", "messages": {"role": "user", "content": "..."} }' That's the documented request with the model swapped in; we didn't send it. Pass claude-haiku-5-5 and not your old model, or you'll get the old tokenizer's count. The calculator is one function. It mirrors the usage object, picks a tier from the prompt length and applies that tier to everything: type Usage struct { InputTokens int json:"input tokens" CacheCreationInputTokens int json:"cache creation input tokens" CacheReadInputTokens int json:"cache read input tokens" OutputTokens int json:"output tokens" } // Rates are US dollars per million tokens. type Rates struct { Input, CacheWrite5m, CacheRead, Output float64 } var Haiku55Short = Rates{Input: 0.10, CacheWrite5m: 0.125, CacheRead: 0.01, Output: 0.50} Haiku55Long = Rates{Input: 0.50, CacheWrite5m: 0.625, CacheRead: 0.05, Output: 2.50} Haiku45 = Rates{Input: 1.00, CacheWrite5m: 1.25, CacheRead: 0.10, Output: 5.00} const Haiku55Threshold = 100 000 func u Usage PromptTokens int { return u.InputTokens + u.CacheCreationInputTokens + u.CacheReadInputTokens } func u Usage at r Rates float64 { return float64 u.InputTokens r.Input + float64 u.CacheCreationInputTokens r.CacheWrite5m + float64 u.CacheReadInputTokens r.CacheRead + float64 u.OutputTokens r.Output / 1e6 } func Haiku55Cost u Usage dollars float64, long bool { if u.PromptTokens Haiku55Threshold { return u.at Haiku55Long , true } return u.at Haiku55Short , false } Feed it two requests that differ by one token, then one 150,000-token prompt against the same material in two halves 1,000 output tokens each : prompt 100000 + 2000 out $0.011000 long=false prompt 100001 + 2000 out $0.055001 long=true one request $0.0775 two requests $0.0160 The second pair is the practical one. Summarising a long document in two 75,000-token calls costs a fifth of doing it in one, and you can add a third call to merge the halves and still be far ahead. For extraction and classification over big inputs, which is what this model is sold for, chunking just became a pricing decision as well as a quality one. Is 100,000 tokens a lot? For a single classification call, it's enormous. For a tool-calling agent that appends every result to the conversation, it isn't. We modelled a loop with prompt caching on: an 8,000-token prefix of system prompt and tool definitions, and each turn adding a 5,000-token tool result and a 500-token reply. Those sizes are made up to be plausible. This is a model of the billing, and your agent will have different numbers. turn 16: prompt 95500 $0.00184 long=false turn 17: prompt 101000 $0.00946 long=true append only $0.3267 first long-tier turn: 17 compact before 100k $0.0623 first long-tier turn: 0, compactions: 2 Turn 17 costs five times turn 16 for almost the same work, and every later turn stays in the expensive tier because the history only grows. Over 40 turns the append-only loop costs $0.33. The second line replaces the history with a 3,000-token summary whenever the next prompt would cross 100,000, pays for writing that summary, and comes to $0.06. Whether a summary is good enough for your task is a separate question that no price table answers. There's a trap in the obvious fix. The API's threshold compaction https://platform.claude.com/docs/en/build-with-claude/compaction-threshold , in beta and listed as supporting Haiku 5.5, summarises old turns on the server once the input reaches a trigger. The default trigger is 150,000 input tokens. On every other model that's a sensible default. On this one it means compaction switches on 50,000 tokens after the price went up. Set it yourself; the minimum is 50,000: { "model": "claude-haiku-5-5", "max tokens": 4096, "context management": { "edits": { "type": "compact 20260112", "trigger": { "type": "input tokens", "value": 80000 } } }, "messages": { "role": "user", "content": "..." } } That body is adapted from the documented example and goes with the anthropic-beta: compact-2026-01-12 header. We didn't send it either. Don't trim the history by hand without reading the migration guide https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide first. Haiku 5.5 thinks by default, and a request that sends a thinking block back after earlier messages changed can come back as a 400 error the guide says which accounts get the check . Its instruction is to keep conversations append-only, which is exactly what a hand-rolled "drop the oldest tool results" loop doesn't do. None of this makes Haiku 5.5 a bad deal. The long tier is half of Haiku 4.5's per-token price, and a quarter of Sonnet 5.5's $2 input rate. The tokenizer eats into that, though. Pricing the same text on both models, with Haiku 5.5's counts inflated by the documented 30%: threshold in Haiku 4.5 tokens: 76923 4.5: 10000 in $0.0150 | 5.5: 13000 in $0.0019 long=false | saving 87% 4.5: 50000 in $0.0550 | 5.5: 65000 in $0.0072 long=false | saving 87% 4.5: 76000 in $0.0810 | 5.5: 98800 in $0.0105 long=false | saving 87% 4.5: 78000 in $0.0830 | 5.5: 101400 in $0.0539 long=true | saving 35% 4.5: 150000 in $0.1550 | 5.5: 195000 in $0.1008 long=true | saving 35% So: 87% cheaper for the same text under the line, 35% cheaper above it. Anthropic's announcement says prompts up to 100,000 tokens "make up around 90% of requests to our previous Haiku model". That's a share of requests, and long requests carry more tokens each, so the share of your spend that lands in the long tier can be a good deal higher than one in ten. The switch isn't only a price change. Manual budget tokens thinking, non-default temperature , and assistant prefill all return 400 errors on Haiku 5.5, much like the changes we listed for Claude Sonnet 5.5 https://dev.to/ai/claude-sonnet-5-5-api-migration . And if your long prompts are long because of tool definitions, measure those first https://dev.to/ai/mcp-tool-definition-token-overhead ; they're in every request. Our advice is short. Put a budget of 90,000 counted tokens on anything you send to Haiku 5.5, set the compaction trigger below it, and log usage per request so you can see how many calls go over. If a workload can't fit, split it. If it can't be split, run it anyway, and know that it's the 35% discount you're getting.