{"slug": "claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model", "title": "Claude Prompt Caching Cost: When It Pays, With the Math for Every Model", "summary": "Anthropic's Claude prompt caching, priced as of 09/28/2026, charges 1.25x the input price for a five-minute cache write and 2x for a one-hour write, while cache reads cost 0.1x input on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1. A five-minute cache write pays for itself after one read and a one-hour write after two reads, and on a 50,000-token prefix reused 10 times on Sonnet 5.5 caching cuts the prefix cost from $1.00 to $0.215. If a cached prefix is never read again before it expires, the write premium of 25% or 100% is lost.", "body_md": "# Claude Prompt Caching Cost: When It Pays, With the Math for Every Model\n\nClaude prompt caching cost as of 09/28/2026: writes cost 1.25x or 2x input, reads cost 0.1x (0.05x on Opus 5.5, 0.025x on Fable 5.1). Break-even and worked examples.\n\n## Why this matters\n\nClaude prompt caching charges a premium to write a prompt prefix into the cache and a steep discount to read it back. A five-minute cache write costs 1.25x the input price and pays for itself after one read. A one-hour write costs 2x and pays for itself after two reads. A cache read costs 0.1x the input price on most models, but 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On a 50,000-token prefix reused 10 times on Sonnet 5.5, caching cuts the prefix cost from $1.00 to $0.215. Batch discounts stack on top. Prices from Anthropic's pricing page, 09/28/2026.\n\nPrompt caching is the largest discount on the Claude API, and the easiest to get wrong. It saves up to 97.5% on repeated input, but only if the repeat happens before the cache expires. If it does not, caching costs you extra.\n\nThis page gives the multipliers, the break-even point, and worked costs for each current model. All prices come from Anthropic’s [pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing), read on 09/28/2026.\n\n## How Claude prompt caching is priced\n\nCaching splits your input into three kinds of token, each with its own price:\n\n- **Cache write:** the first time a prompt prefix is sent with caching on, Claude stores it. You pay a premium on those tokens.\n- **Cache read:** later requests that start with the same prefix read it from the cache. You pay a small fraction of the input price.\n- **Uncached input:** anything after the cached prefix is billed at the normal input price.\n\nAnthropic offers two cache lifetimes, and they price the write differently:\n\n| Token type | Price, as a multiple of base input | \n|---|---|\n| Cache write, five-minute lifetime | 1.25x | \n| Cache write, one-hour lifetime | 2x | \n| Cache read, most models | 0.1x | \n| Cache read, Opus 5.5 | 0.05x | \n| Cache read, Fable 5.1 | 0.025x | \n\nOutput tokens are not affected. Caching only changes what you pay for input.\n\n## The prices per model\n\nPer million tokens, from the same page:\n\n| Model | Input | Cache write (5 min) | Cache write (1 hour) | Cache read | \n|---|---|---|---|---|\n| Fable 5.1 | $10 | $12.50 | $20 | $0.25 | \n| Opus 5.5 | $4 | $5 | $8 | $0.20 | \n| Opus 5 | $5 | $6.25 | $10 | $0.50 | \n| Sonnet 5.5 and Sonnet 5 | $2 | $2.50 | $4 | $0.20 | \n| Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | \n\nLook at the last column. Fable 5.1 lists at five times the input price of Sonnet 5.5, but reads its cache for only 25% more. Opus 5.5 reads its cache for the same $0.20 as Sonnet 5.5. For work that is mostly cache reads, the price gap between models is much smaller than the input prices suggest.\n\n## The break-even point\n\nAnthropic’s pricing page states the rule, and the arithmetic is short enough to check.\n\n**Five-minute cache.** The write costs 0.25x more than sending the prefix uncached. Each read saves 0.9x (you pay 0.1x instead of 1x). One read saves more than the write cost. **Caching pays after one read.**\n\n**One-hour cache.** The write costs 1x more than uncached. Each read saves 0.9x. One read saves 0.9x, not enough. Two reads save 1.8x. **Caching pays after two reads.**\n\nOn Opus 5.5 and Fable 5.1 each read saves even more (0.95x and 0.975x), but the break-even count stays the same.\n\nIf a prefix is written and never read again before it expires, you paid 25% or 100% extra for nothing. That is the one way caching makes a bill bigger.\n\n## Worked example: a 50,000-token prefix, reused 10 times\n\nSuppose every request starts with the same 50,000 tokens: a system prompt, tool definitions and a reference document. You send 10 requests inside the five-minute window. Only the prefix is priced here. The new text and the output in each request cost the same either way.\n\n| Model | No caching (10 full sends) | Five-minute cache (1 write, 9 reads) | Saving | \n|---|---|---|---|\n| Haiku 4.5 | $0.50 | $0.1075 | 78.5% | \n| Sonnet 5.5 | $1.00 | $0.215 | 78.5% | \n| Opus 5.5 | $2.00 | $0.34 | 83% | \n| Fable 5.1 | $5.00 | $0.7375 | 85% | \n\nHow the Sonnet 5.5 row works: uncached, each send is 50,000 tokens at $2 per million, or $0.10, and 10 sends cost $1.00. Cached, the first send is a write at $2.50 per million, or $0.125. The nine reads cost $0.20 per million, or $0.01 each, for $0.09. Total: $0.215.\n\nThe saving grows with reuse. At 100 requests the cached Sonnet 5.5 prefix costs $1.115 against $10.00 uncached, an 89% cut.\n\n## Where caching matters most: agents and Claude Code\n\nAn agent re-sends its whole context on every turn. Turn 200 of a Claude Code session carries everything from turns 1 to 199 as its prefix. Without caching, that prefix would be billed in full every turn.\n\nI measured this on a real day. In [Opus 5.5 vs Fable 5.1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1) I broke down 2,979 turns of my own Claude Code work. On Opus 5 prices, cache reads were the largest single line of the day: $170 of a $359 total. That is also why the Opus 5.5 cache-read cut mattered more than its 20% cheaper input price. The same day came to $220 at Opus 5.5 prices.\n\nIf you run Claude Code on a Pro or Max plan instead of the API, caching still matters. Long sessions use your limit faster, as I explain in [Claude usage limits](https://chudi.dev/blog/claude-usage-limits), and the habits that keep a prefix stable also keep a session cheap. [Claude Code usage limits](https://chudi.dev/blog/claude-code-quota-burn-10-patterns) lists them.\n\n## Three ways to lose the discount\n\n1. **Changing the prefix.** A cache matches from the start of the prompt. Edit one line of the system prompt or reorder the tools and every later token is a new write. Put what changes at the end.\n2. **Letting it expire.** A gap longer than the cache lifetime means the next request writes again. If your requests are spread out, the one-hour cache is cheaper despite its 2x write.\n3. **Caching something used once.** A one-off prompt with caching on pays the write premium and never reads it back.\n\nReducing what goes into the prefix at all is the other half of the job. [Reduce Claude token usage with progressive disclosure](https://chudi.dev/blog/reduce-ai-token-usage-progressive-disclosure) covers that.\n\n## Other discounts that stack\n\n- **Batch API:** 50% off input and output, and it combines with caching.\n- **Long context:** on Claude 4.6 and later models, the full one-million-token context is billed at the standard price, so a large cached prefix does not trigger a long-context premium.\n\n## Price your own prefix\n\nThe [Claude token counter](https://chudi.dev/tools/claude-token-counter) estimates the tokens in any text you paste and prices them on every current Claude model. It also takes cache-read and cache-write token counts, with a five-minute or one-hour lifetime, and prices those at each model’s own multiplier. Paste your system prompt and tool definitions to see what a write and each read cost.\n\nCache multipliers differ by model: Opus 5.5 and Fable 5.1 both break the usual 0.1x rule. [Claude Cost Watch](https://chudi.dev/products/claude-cost-watch) is a waitlist for a short email that flags changes like that. Nothing is charged today.\n\n## Sources\n\nAll read on 09/28/2026.\n\n- Cache multipliers, per-model prices, break-even rule, batch and long-context pricing: [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)\n- Measured Claude Code day: [Opus 5.5 vs Fable 5.1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1)\n- Full price list: [Claude pricing](https://chudi.dev/blog/claude-pricing)\n\nNo link on this page is an affiliate link.\n\n· Frequently asked\n\n## FAQ\n\n### How much does Claude prompt caching cost?\n\nA cache write costs 1.25x the model's input price for the five-minute cache and 2x for the one-hour cache. A cache read costs 0.1x the input price on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On Sonnet 5.5 that is $2.50 or $4 per million tokens to write and $0.20 to read. Source: platform.claude.com pricing page, 09/28/2026.\n\n### When does prompt caching pay off?\n\nA five-minute cache write pays off after one cache read. A one-hour cache write pays off after two. If the same prefix is not reused within the cache lifetime, caching costs more than not caching.\n\n### Should I use the five-minute or the one-hour cache?\n\nUse the five-minute cache when requests reusing the prefix arrive within minutes of each other, as in an agent loop or a chat session. Use the one-hour cache when reuse is spread out, for example a large document queried every 10 to 30 minutes. The one-hour write costs 2x instead of 1.25x.\n\n### Does prompt caching stack with the Batch API discount?\n\nYes. Anthropic's pricing page states that the batch discount and prompt caching discounts can be combined. Batch gives 50% off input and output.\n\n### Why does caching matter so much for Claude Code and agents?\n\nAn agent re-sends its whole context on every turn, so most of its input is the same prefix read again. In one day of Claude Code work I measured on Opus 5, cache reads were the largest single line of the bill: $170 of $359.\n\n· Sources & further reading\n\n## Sources & Further Reading\n\n### Further reading\n\n- [Claude Pricing in 2026: Every Plan and API Price on One Page /blog/claude-pricing](https://chudi.dev/blog/claude-pricing) Claude pricing as of 09/28/2026: Free, Pro ($20), Max ($100 or $200), Team and Enterprise plans, plus per-token API prices for Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 4.5.\n- [Claude Pro vs Max: Who Needs the $100 or $200 Plan, and When the API Wins /blog/claude-pro-vs-max](https://chudi.dev/blog/claude-pro-vs-max) Claude Pro vs Max as of 09/28/2026: $20 vs $100 or $200 a month, 5x or 20x the usage, both with Claude Code. How to choose, and what one measured day costs on the API.\n- [Claude Usage Limits Explained: The 5-Hour Session, the Weekly Cap, and What Counts /blog/claude-usage-limits](https://chudi.dev/blog/claude-usage-limits) How Claude usage limits work on Pro and Max: a five-hour session limit, a weekly limit, one pool shared by claude.ai, Desktop and Claude Code, and what to do when you hit it.\n- [Opus 5.5 vs Fable 5.1: My Cost Per Turn Gap Went From 1.4x to 2.3x /blog/claude-opus-5-5-vs-fable-5-1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1) Opus 5.5 beats Fable 5.1 on all nine of Anthropic's launch benchmarks at 40% of the price. Replayed at 5.5 prices, my measured day puts Fable at 2.3x per turn.\n- [Fable 5.1 vs Opus 5 in Production: A Whole Day on Fable Used 19% of a Max Weekly Window /blog/claude-fable-5-vs-opus-5](https://chudi.dev/blog/claude-fable-5-vs-opus-5) Same harness, same day: 2,443 Fable 5.1 turns and 1,120 Opus 5 turns after a weekly reset consumed 19% of the Max weekly window, and the 5-hour window peaked at 59% without capping. Fable errored less per tool call, hit more permission gates, and would have cost 1.4x per turn on the API, not 2x.\n\nBuilt from these systems\n\nChiron Seed $29\n\nFour Claude Code skills and two scripts that keep a knowledge base inside your repo, so the next session starts from what the last one learned.\n\n[Get the kit](https://7124296224445.gumroad.com/l/chiron-seed)\n\nWant more of this in your Google results?\n\n## What do you think?\n\nI post about this stuff on LinkedIn every day and the conversations there are great. If this post sparked a thought, I'd love to hear it.\n\n[Discuss on LinkedIn](https://www.linkedin.com/in/chudi-nnorukam)", "url": "https://wpnews.pro/news/claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model", "canonical_source": "https://chudi.dev/blog/claude-prompt-caching-cost", "published_at": "2026-09-28 00:00:00+00:00", "updated_at": "2026-09-29 01:48:16.043726+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["Anthropic", "Claude", "Opus 5.5", "Fable 5.1", "Sonnet 5.5", "Haiku 4.5", "Opus 5", "Sonnet 5"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model", "markdown": "https://wpnews.pro/news/claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model.md", "text": "https://wpnews.pro/news/claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model.txt", "jsonld": "https://wpnews.pro/news/claude-prompt-caching-cost-when-it-pays-with-the-math-for-every-model.jsonld"}}