{"slug": "haiku-5-5-cut-prices-90-then-the-fine-print-showed-up", "title": "Haiku 5.5 Cut Prices 90%, Then the Fine Print Showed Up", "summary": "Anthropic released Haiku 5.5 on October 7, 2026, cutting per-token prices 90% for requests under 100,000 tokens to $0.10 per million input and $0.50 per million output, matching GPT-6 Luna's rate. Independent testers found a new tokenizer inflates billable token counts by roughly 30%, so Anthropic estimates real workload savings at about 75% rather than 90%, and requests above 100,000 tokens are billed at five times the base rate. The model also raised context from 200K to 1 million tokens and added an adjustable effort setting, with high-effort runs producing long thinking blocks that can cost more than running Sonnet 5.5 at low effort.", "body_md": "Nobody uses Haiku 4.5. It shipped a year ago at $1 per million input tokens and $5 per million output, and GPT-6 Luna started lapping it at a tenth of the price. That line, from a Substack post the day Haiku 5.5 landed, is the whole reason this launch exists.\n\nHaiku 5.5 dropped on October 7, 2026. Anthropic cut the price by 90% for requests under 100,000 tokens, down to $0.10 in and $0.50 out. That is exactly Luna's rate.\n\nThe reaction was instant and split right down the middle. Half the timeline celebrated. The other half started doing math.\n\nI fell into the second group fast, because the headline number and the number you actually pay are rarely the same thing.\n\nHere's a question people always ask: what actually changed?\n\nThe numbers are the story. Haiku 4.5 cost $1 and $5. Haiku 5.5 costs $0.10 and $0.50 for short requests, and $0.50 and $2.50 once you cross 100,000 tokens.\n\nThat is not a discount. It is a different product tier.\n\nContext jumped too, from 200K to 1 million tokens. Output limits went up. And Haiku 5.5 is the first small Claude with an adjustable effort setting, which lets you pick how hard it thinks before it answers.\n\nAnthropic estimates workloads cost about 75% less, not 90%.\n\nThat gap between the sticker and the real bill is where the argument lives.\n\nWhat actually happens when a lab cuts prices is that the fine print does the damage.\n\nAnthropic's own launch notes admit the 90% is per token, not per job. Two things eat the difference. The first is a new tokenizer, shared with the 5.5 family, that turns the same text into more billable tokens than Haiku 4.5 did.\n\nIndependent testers put that inflation around 30% on ordinary strings. So a 90% token cut lands as roughly a 75% bill cut, which is still great and not what the headline says.\n\nThe new tokenizer is not a bug. It is a real tradeoff. Anthropic counts more tokens for the same words, and the honest framing is that the per-job math decides the actual saving, never the sticker.\n\nI have been burned by this exact pattern before. A vendor cut a model's rate, I moved a batch job onto it, and the bill went down much less than the percentage promised.\n\nThe second catch is cleverer. Above 100,000 tokens, the rate jumps to five times the base.\n\nAnthropic says about 90% of Haiku 4.5 requests sit under that line, so most people never feel it. The rest do, hard.\n\n**The discount is real, and it has a cliff.**\n\nMost tutorials tell you Haiku is the cheap, fast one and leave it there. The effort dial changes that.\n\nAt low and medium effort, Haiku 5.5 is excellent. Classification, routing, extraction, the repetitive chores that used to cost real money.\n\nIt beats Haiku 4.5 comfortably and matches Luna. That is the mode most teams will actually run it in.\n\nThen you turn it up. At high and max effort, the model writes long internal thinking blocks before it answers. Output tokens cost five times input tokens, so deep reasoning burns the budget fast.\n\nIndependent analysts found Haiku 5.5 generated 97 million tokens across suites where comparable models needed a median of 87 million. That is structural verbosity, not a one-off.\n\n| Model | Input / 1M | Output / 1M | Notes | \n|---|---|---|---|\n| Haiku 5.5 (short) | $0.10 | $0.50 | Under 100K tokens | \n| Haiku 5.5 (long) | $0.50 | $2.50 | Five times the base | \n| Haiku 4.5 | $1.00 | $5.00 | 200K context | \n| GPT-6 Luna | $0.10 | $0.50 | Surcharge above 272K | \n\nSome testers went further and said that at max effort, a task can cost more than running Sonnet 5.5 at low effort. That is a wild sentence to write about a model sold on being cheap.\n\nI ran the same pattern through my head for the jobs I actually send to a cheap model. Short classifications, route checks, extraction. Those barely tickle the effort dial, which is why I still like the release.\n\nThe capability numbers are genuinely hard to argue with.\n\nOn OSWorld, computer use, Haiku 5.5 jumped from 15.7% to 72.4%. On Terminal-Bench 4.0, agentic coding, it went from 0% to 39.2%. Its knowledge-work score moved from 735 to 1,620 on GDPval.\n\nAgainst Luna, Anthropic's own table shows Haiku 5.5 ahead on every shared benchmark, including 46.4% vs 42.4% on FrontierCode.\n\nOn OSWorld the jump is the one that made people sit up. Computer use went from near-unusable to genuinely strong in a single generation, and at the small-model price.\n\nSonnet 5.5 still wins on the hard stuff. Anthropic says so directly, and frames Haiku as a worker that supports the bigger models.\n\nAll of those come from the vendor, which matters. Not one is independent yet. I have learned to hold vendor tables loosely, especially the ones with a number that looks too good.\n\nI used to think these price wars were about generosity. They are about agents.\n\nAnthropic's own framing says it: Haiku exists to retrieve the revenue figure for one slide while Opus builds the presentation.\n\nWhen agents call a model hundreds of times per task, the cheap tier becomes the whole economy. That is why this launch is priced the way it is, and why the next one will be cheaper still.\n\nLuna proved the demand was there. Anthropic matched the number, added the effort dial, and called it a launch. It mostly is one.\n\nIf you were already using Haiku 4.5, switch. There is no reason to pay $5 for output when the newer model is better and costs a tenth of that for short work.\n\nThat is the easy call, and it is the one most people should make today.\n\nIf you were going to use this for long agent runs at high effort, measure first. The sticker price and the actual bill diverge at exactly the settings power users like best. Independent benchmarks are not out yet.\n\nAnd if you route high-volume simple work, this is a genuine gift. Classification and extraction just got cheap enough to run without thinking.\n\nThat is a real change for anyone running agents, and it is worth saying plainly.\n\nThe biggest misconception is that cheaper tokens mean a cheaper product. It does not. Cheaper tokens mean you can afford to waste more of them, and Haiku 5.5 at max effort is built to be wasteful.\n\nHaiku 4.5 was not a bad model. It was a model with a bad price, and a competitor moved first.\n\nThat is the actual lesson here. Anthropic did not ship a new tier because it wanted to.\n\nThat is the actual lesson here. Anthropic did not ship a new tier because it wanted to. It shipped one because Luna made the old tier unsellable, and the \"nobody uses Haiku 4.5\" line was already being typed.\n\nThe price war is not over. It just moved down a shelf, and the shelf below is where most of the work happens.\n\nI still do not know if the 75% number holds for my own usage. I will find out the boring way, on a bill.", "url": "https://wpnews.pro/news/haiku-5-5-cut-prices-90-then-the-fine-print-showed-up", "canonical_source": "https://dev.to/dishant0406/haiku-55-cut-prices-90-then-the-fine-print-showed-up-5apn", "published_at": "2026-10-07 20:08:42+00:00", "updated_at": "2026-10-07 20:17:29.833945+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-tools", "generative-ai"], "entities": ["Anthropic", "Haiku 5.5", "Haiku 4.5", "GPT-6 Luna", "Sonnet 5.5", "OSWorld", "Terminal-Bench 4.0"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/haiku-5-5-cut-prices-90-then-the-fine-print-showed-up", "markdown": "https://wpnews.pro/news/haiku-5-5-cut-prices-90-then-the-fine-print-showed-up.md", "text": "https://wpnews.pro/news/haiku-5-5-cut-prices-90-then-the-fine-print-showed-up.txt", "jsonld": "https://wpnews.pro/news/haiku-5-5-cut-prices-90-then-the-fine-print-showed-up.jsonld"}}