cd /news/ai-tools/claude-prompt-caching-cost-when-it-p… · home › topics › ai-tools › article
[ARTICLE · art-141383] src=chudi.dev ↗ pub= topic=ai-tools verified=true sentiment=· neutral

Claude Prompt Caching Cost: When It Pays, With the Math for Every Model

Anthropic's Claude prompt caching, priced as of 09/28/2026, charges 1.25x the input price for a five-minute cache write and 2x for a one-hour write, while cache reads cost 0.1x input on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1. A five-minute cache write pays for itself after one read and a one-hour write after two reads, and on a 50,000-token prefix reused 10 times on Sonnet 5.5 caching cuts the prefix cost from $1.00 to $0.215. If a cached prefix is never read again before it expires, the write premium of 25% or 100% is lost.

by read10 min views1 publishedSep 28, 2026
Claude Prompt Caching Cost: When It Pays, With the Math for Every Model
Image: Chudi (auto-discovered)

Claude prompt caching cost as of 09/28/2026: writes cost 1.25x or 2x input, reads cost 0.1x (0.05x on Opus 5.5, 0.025x on Fable 5.1). Break-even and worked examples.

Why this matters #

Claude prompt caching charges a premium to write a prompt prefix into the cache and a steep discount to read it back. A five-minute cache write costs 1.25x the input price and pays for itself after one read. A one-hour write costs 2x and pays for itself after two reads. A cache read costs 0.1x the input price on most models, but 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On a 50,000-token prefix reused 10 times on Sonnet 5.5, caching cuts the prefix cost from $1.00 to $0.215. Batch discounts stack on top. Prices from Anthropic's pricing page, 09/28/2026.

Prompt caching is the largest discount on the Claude API, and the easiest to get wrong. It saves up to 97.5% on repeated input, but only if the repeat happens before the cache expires. If it does not, caching costs you extra.

This page gives the multipliers, the break-even point, and worked costs for each current model. All prices come from Anthropic’s pricing documentation, read on 09/28/2026.

How Claude prompt caching is priced #

Caching splits your input into three kinds of token, each with its own price:

  • Cache write: the first time a prompt prefix is sent with caching on, Claude stores it. You pay a premium on those tokens.
  • Cache read: later requests that start with the same prefix read it from the cache. You pay a small fraction of the input price.
  • Uncached input: anything after the cached prefix is billed at the normal input price.

Anthropic offers two cache lifetimes, and they price the write differently:

Token type Price, as a multiple of base input
Cache write, five-minute lifetime 1.25x
Cache write, one-hour lifetime 2x
Cache read, most models 0.1x
Cache read, Opus 5.5 0.05x
Cache read, Fable 5.1 0.025x

Output tokens are not affected. Caching only changes what you pay for input.

The prices per model #

Per million tokens, from the same page:

| Model | Input | Cache write (5 min) | Cache write (1 hour) | Cache read | 
|---|---|---|---|---|

| Fable 5.1 | $10 | $12.50 | $20 | $0.25 | | Opus 5.5 | $4 | $5 | $8 | $0.20 | | Opus 5 | $5 | $6.25 | $10 | $0.50 | | Sonnet 5.5 and Sonnet 5 | $2 | $2.50 | $4 | $0.20 | | Haiku 4.5 | $1 | $1.25 | $2 | $0.10 |

Look at the last column. Fable 5.1 lists at five times the input price of Sonnet 5.5, but reads its cache for only 25% more. Opus 5.5 reads its cache for the same $0.20 as Sonnet 5.5. For work that is mostly cache reads, the price gap between models is much smaller than the input prices suggest.

The break-even point #

Anthropic’s pricing page states the rule, and the arithmetic is short enough to check.

Five-minute cache. The write costs 0.25x more than sending the prefix uncached. Each read saves 0.9x (you pay 0.1x instead of 1x). One read saves more than the write cost. Caching pays after one read.

One-hour cache. The write costs 1x more than uncached. Each read saves 0.9x. One read saves 0.9x, not enough. Two reads save 1.8x. Caching pays after two reads.

On Opus 5.5 and Fable 5.1 each read saves even more (0.95x and 0.975x), but the break-even count stays the same.

If a prefix is written and never read again before it expires, you paid 25% or 100% extra for nothing. That is the one way caching makes a bill bigger.

Worked example: a 50,000-token prefix, reused 10 times #

Suppose every request starts with the same 50,000 tokens: a system prompt, tool definitions and a reference document. You send 10 requests inside the five-minute window. Only the prefix is priced here. The new text and the output in each request cost the same either way.

| Model | No caching (10 full sends) | Five-minute cache (1 write, 9 reads) | Saving | 
|---|---|---|---|

| Haiku 4.5 | $0.50 | $0.1075 | 78.5% | | Sonnet 5.5 | $1.00 | $0.215 | 78.5% | | Opus 5.5 | $2.00 | $0.34 | 83% | | Fable 5.1 | $5.00 | $0.7375 | 85% |

How the Sonnet 5.5 row works: uncached, each send is 50,000 tokens at $2 per million, or $0.10, and 10 sends cost $1.00. Cached, the first send is a write at $2.50 per million, or $0.125. The nine reads cost $0.20 per million, or $0.01 each, for $0.09. Total: $0.215.

The saving grows with reuse. At 100 requests the cached Sonnet 5.5 prefix costs $1.115 against $10.00 uncached, an 89% cut.

Where caching matters most: agents and Claude Code #

An agent re-sends its whole context on every turn. Turn 200 of a Claude Code session carries everything from turns 1 to 199 as its prefix. Without caching, that prefix would be billed in full every turn.

I measured this on a real day. In Opus 5.5 vs Fable 5.1 I broke down 2,979 turns of my own Claude Code work. On Opus 5 prices, cache reads were the largest single line of the day: $170 of a $359 total. That is also why the Opus 5.5 cache-read cut mattered more than its 20% cheaper input price. The same day came to $220 at Opus 5.5 prices.

If you run Claude Code on a Pro or Max plan instead of the API, caching still matters. Long sessions use your limit faster, as I explain in Claude usage limits, and the habits that keep a prefix stable also keep a session cheap. Claude Code usage limits lists them.

Three ways to lose the discount #

  1. Changing the prefix. A cache matches from the start of the prompt. Edit one line of the system prompt or reorder the tools and every later token is a new write. Put what changes at the end.
  2. Letting it expire. A gap longer than the cache lifetime means the next request writes again. If your requests are spread out, the one-hour cache is cheaper despite its 2x write.
  3. Caching something used once. A one-off prompt with caching on pays the write premium and never reads it back.

Reducing what goes into the prefix at all is the other half of the job. Reduce Claude token usage with progressive disclosure covers that.

Other discounts that stack #

  • Batch API: 50% off input and output, and it combines with caching.
  • Long context: on Claude 4.6 and later models, the full one-million-token context is billed at the standard price, so a large cached prefix does not trigger a long-context premium.

Price your own prefix #

The Claude token counter estimates the tokens in any text you paste and prices them on every current Claude model. It also takes cache-read and cache-write token counts, with a five-minute or one-hour lifetime, and prices those at each model’s own multiplier. Paste your system prompt and tool definitions to see what a write and each read cost.

Cache multipliers differ by model: Opus 5.5 and Fable 5.1 both break the usual 0.1x rule. Claude Cost Watch is a waitlist for a short email that flags changes like that. Nothing is charged today.

Sources #

All read on 09/28/2026.

- Cache multipliers, per-model prices, break-even rule, batch and long-context pricing: [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
- Measured Claude Code day: [Opus 5.5 vs Fable 5.1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1)
- Full price list: [Claude pricing](https://chudi.dev/blog/claude-pricing)

No link on this page is an affiliate link.

· Frequently asked

FAQ #

How much does Claude prompt caching cost?

A cache write costs 1.25x the model's input price for the five-minute cache and 2x for the one-hour cache. A cache read costs 0.1x the input price on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On Sonnet 5.5 that is $2.50 or $4 per million tokens to write and $0.20 to read. Source: platform.claude.com pricing page, 09/28/2026.

When does prompt caching pay off?

A five-minute cache write pays off after one cache read. A one-hour cache write pays off after two. If the same prefix is not reused within the cache lifetime, caching costs more than not caching.

Should I use the five-minute or the one-hour cache?

Use the five-minute cache when requests reusing the prefix arrive within minutes of each other, as in an agent loop or a chat session. Use the one-hour cache when reuse is spread out, for example a large document queried every 10 to 30 minutes. The one-hour write costs 2x instead of 1.25x.

Does prompt caching stack with the Batch API discount?

Yes. Anthropic's pricing page states that the batch discount and prompt caching discounts can be combined. Batch gives 50% off input and output.

Why does caching matter so much for Claude Code and agents?

An agent re-sends its whole context on every turn, so most of its input is the same prefix read again. In one day of Claude Code work I measured on Opus 5, cache reads were the largest single line of the bill: $170 of $359.

· Sources & further reading

Sources & Further Reading #

Further reading

Built from these systems

Chiron Seed $29

Four Claude Code skills and two scripts that keep a knowledge base inside your repo, so the next session starts from what the last one learned.

Get the kit Want more of this in your Google results?

What do you think? #

I post about this stuff on LinkedIn every day and the conversations there are great. If this post sparked a thought, I'd love to hear it.

Discuss on LinkedIn

── more in #ai-tools 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-prompt-cachin…] indexed:0 read:10min 2026-09-28 · —