{"slug": "fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough", "title": "Fable 5.1 costs twice as much as Opus 5, until your cache gets big enough", "summary": "A developer built llmabacus.com, an LLM API price table, and used it to calculate the exact cache-read crossover point between Claude Fable 5.1 and Opus 5. Fable 5.1 halves cache-read pricing to $0.25 per million tokens while keeping $10/$50 input/output rates, so it only becomes cheaper than Opus 5 past roughly 140K cached tokens per request — and cache-write premiums may prevent heavy agent users from ever reaching that point.", "body_md": "Claude Fable 5.1 shipped on September 1 with the same $10 input / $50 output per million tokens as Fable 5. One number moved: cache reads dropped from $1 to $0.25 per million.\n\nOn the price sheet it still costs twice what Opus 5 does ($5 / $25). But look at the cache-read column and the order flips. Fable 5.1 charges $0.25 there, Opus 5 charges $0.50. So which one is cheaper depends on how much of each request is served from cache, and I wanted the actual crossover instead of a vibe.\n\nCached tokens times the cache-read price, fresh input times the input price, output times the output price. Fable 5.1 only changed the first term, so the saving over Fable 5 is exactly as big as that term's share of your bill.\n\nTwo workloads I priced out from the table on my site:\n\n| Per request | Fable 5 | Fable 5.1 | Opus 5 | \n|---|---|---|---|\n| Agent: 100K cached, 2K new input, 1K output | $0.170 | $0.095 | $0.085 | \n| Chat: 5K cached, 1K new input, 800 output | $0.055 | $0.051 | $0.028 | \n\nIn the agent loop, cache reads are about 60% of the Fable 5 bill, so the price cut takes 44% off each request. That lines up with the \"around 45% for heavy agent use\" figure in the launch coverage.\n\nThe chat case barely moves. Output dominates, and Opus 5 stays roughly half the price. Switching there is just paying more.\n\nFable 5.1 pays double for fresh input and output, and half for cache reads. Set the two bills equal:\n\n```\n0.25 * R = 5 * U + 25 * O\nR = 20 * U + 100 * O\n```\n\nR is cached tokens per request, U is fresh input, O is output. Plug in the agent numbers above and the line sits at 140K cached tokens. Both models cost $0.105 a request there.\n\nPast that point the gap keeps widening. At 200K cached, Fable 5.1 is about 11% cheaper. Double the cache again and it's more than a quarter cheaper.\n\nOutput length is what pushes the line out. Every extra thousand output tokens needs another hundred thousand cached tokens before Fable 5.1 catches up.\n\nMy first pass stopped there and the conclusion was \"long-context agents should move to Fable 5.1.\" Then I remembered that tokens have to get into the cache before they can be read from it.\n\nAnthropic bills a 5-minute cache write at 1.25x the input price for Opus 5. My price table has no separate write price for Fable 5.1. If it follows the same 1.25x rule, writing a 200K prefix costs $1.25 more on Fable 5.1 than on Opus 5. At 200K cached, Fable 5.1 saves $0.015 per request. That's more than 80 requests on the same prefix before the write premium is paid back.\n\nSo if your prefix expires every five minutes, or your agent keeps rewriting its context, you may never reach the crossover at all. Measure how many requests actually reuse each prefix before you switch.\n\nPull three averages from your logs: cached tokens, fresh input and output per request. Then check how many requests reuse the same prefix.\n\n`20U + 100O` cached tokens: stay on Opus 5.\nI built llmabacus.com, the LLM API price table these numbers come from. The full write-up, in Chinese, is at [https://www.llmabacus.com/articles/fable-5-1-cache-read-vs-opus-5](https://www.llmabacus.com/articles/fable-5-1-cache-read-vs-opus-5)", "url": "https://wpnews.pro/news/fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough", "canonical_source": "https://dev.to/szp2005/fable-51-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough-4nd1", "published_at": "2026-09-28 02:10:53+00:00", "updated_at": "2026-09-28 02:48:45.460751+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "ai-tools", "ai-infrastructure"], "entities": ["Anthropic", "Claude Fable 5.1", "Opus 5", "llmabacus.com"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough", "markdown": "https://wpnews.pro/news/fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough.md", "text": "https://wpnews.pro/news/fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough.txt", "jsonld": "https://wpnews.pro/news/fable-5-1-costs-twice-as-much-as-opus-5-until-your-cache-gets-big-enough.jsonld"}}