# Claude Prompt Caching Cost: When It Pays, With the Math for Every Model

> Source: <https://chudi.dev/blog/claude-prompt-caching-cost>
> Published: 2026-09-28 00:00:00+00:00

# Claude Prompt Caching Cost: When It Pays, With the Math for Every Model

Claude prompt caching cost as of 09/28/2026: writes cost 1.25x or 2x input, reads cost 0.1x (0.05x on Opus 5.5, 0.025x on Fable 5.1). Break-even and worked examples.

## Why this matters

Claude prompt caching charges a premium to write a prompt prefix into the cache and a steep discount to read it back. A five-minute cache write costs 1.25x the input price and pays for itself after one read. A one-hour write costs 2x and pays for itself after two reads. A cache read costs 0.1x the input price on most models, but 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On a 50,000-token prefix reused 10 times on Sonnet 5.5, caching cuts the prefix cost from $1.00 to $0.215. Batch discounts stack on top. Prices from Anthropic's pricing page, 09/28/2026.

Prompt caching is the largest discount on the Claude API, and the easiest to get wrong. It saves up to 97.5% on repeated input, but only if the repeat happens before the cache expires. If it does not, caching costs you extra.

This page gives the multipliers, the break-even point, and worked costs for each current model. All prices come from Anthropic’s [pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing), read on 09/28/2026.

## How Claude prompt caching is priced

Caching splits your input into three kinds of token, each with its own price:

- **Cache write:** the first time a prompt prefix is sent with caching on, Claude stores it. You pay a premium on those tokens.
- **Cache read:** later requests that start with the same prefix read it from the cache. You pay a small fraction of the input price.
- **Uncached input:** anything after the cached prefix is billed at the normal input price.

Anthropic offers two cache lifetimes, and they price the write differently:

| Token type | Price, as a multiple of base input | 
|---|---|
| Cache write, five-minute lifetime | 1.25x | 
| Cache write, one-hour lifetime | 2x | 
| Cache read, most models | 0.1x | 
| Cache read, Opus 5.5 | 0.05x | 
| Cache read, Fable 5.1 | 0.025x | 

Output tokens are not affected. Caching only changes what you pay for input.

## The prices per model

Per million tokens, from the same page:

| Model | Input | Cache write (5 min) | Cache write (1 hour) | Cache read | 
|---|---|---|---|---|
| Fable 5.1 | $10 | $12.50 | $20 | $0.25 | 
| Opus 5.5 | $4 | $5 | $8 | $0.20 | 
| Opus 5 | $5 | $6.25 | $10 | $0.50 | 
| Sonnet 5.5 and Sonnet 5 | $2 | $2.50 | $4 | $0.20 | 
| Haiku 4.5 | $1 | $1.25 | $2 | $0.10 | 

Look at the last column. Fable 5.1 lists at five times the input price of Sonnet 5.5, but reads its cache for only 25% more. Opus 5.5 reads its cache for the same $0.20 as Sonnet 5.5. For work that is mostly cache reads, the price gap between models is much smaller than the input prices suggest.

## The break-even point

Anthropic’s pricing page states the rule, and the arithmetic is short enough to check.

**Five-minute cache.** The write costs 0.25x more than sending the prefix uncached. Each read saves 0.9x (you pay 0.1x instead of 1x). One read saves more than the write cost. **Caching pays after one read.**

**One-hour cache.** The write costs 1x more than uncached. Each read saves 0.9x. One read saves 0.9x, not enough. Two reads save 1.8x. **Caching pays after two reads.**

On Opus 5.5 and Fable 5.1 each read saves even more (0.95x and 0.975x), but the break-even count stays the same.

If a prefix is written and never read again before it expires, you paid 25% or 100% extra for nothing. That is the one way caching makes a bill bigger.

## Worked example: a 50,000-token prefix, reused 10 times

Suppose every request starts with the same 50,000 tokens: a system prompt, tool definitions and a reference document. You send 10 requests inside the five-minute window. Only the prefix is priced here. The new text and the output in each request cost the same either way.

| Model | No caching (10 full sends) | Five-minute cache (1 write, 9 reads) | Saving | 
|---|---|---|---|
| Haiku 4.5 | $0.50 | $0.1075 | 78.5% | 
| Sonnet 5.5 | $1.00 | $0.215 | 78.5% | 
| Opus 5.5 | $2.00 | $0.34 | 83% | 
| Fable 5.1 | $5.00 | $0.7375 | 85% | 

How the Sonnet 5.5 row works: uncached, each send is 50,000 tokens at $2 per million, or $0.10, and 10 sends cost $1.00. Cached, the first send is a write at $2.50 per million, or $0.125. The nine reads cost $0.20 per million, or $0.01 each, for $0.09. Total: $0.215.

The saving grows with reuse. At 100 requests the cached Sonnet 5.5 prefix costs $1.115 against $10.00 uncached, an 89% cut.

## Where caching matters most: agents and Claude Code

An agent re-sends its whole context on every turn. Turn 200 of a Claude Code session carries everything from turns 1 to 199 as its prefix. Without caching, that prefix would be billed in full every turn.

I measured this on a real day. In [Opus 5.5 vs Fable 5.1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1) I broke down 2,979 turns of my own Claude Code work. On Opus 5 prices, cache reads were the largest single line of the day: $170 of a $359 total. That is also why the Opus 5.5 cache-read cut mattered more than its 20% cheaper input price. The same day came to $220 at Opus 5.5 prices.

If you run Claude Code on a Pro or Max plan instead of the API, caching still matters. Long sessions use your limit faster, as I explain in [Claude usage limits](https://chudi.dev/blog/claude-usage-limits), and the habits that keep a prefix stable also keep a session cheap. [Claude Code usage limits](https://chudi.dev/blog/claude-code-quota-burn-10-patterns) lists them.

## Three ways to lose the discount

1. **Changing the prefix.** A cache matches from the start of the prompt. Edit one line of the system prompt or reorder the tools and every later token is a new write. Put what changes at the end.
2. **Letting it expire.** A gap longer than the cache lifetime means the next request writes again. If your requests are spread out, the one-hour cache is cheaper despite its 2x write.
3. **Caching something used once.** A one-off prompt with caching on pays the write premium and never reads it back.

Reducing what goes into the prefix at all is the other half of the job. [Reduce Claude token usage with progressive disclosure](https://chudi.dev/blog/reduce-ai-token-usage-progressive-disclosure) covers that.

## Other discounts that stack

- **Batch API:** 50% off input and output, and it combines with caching.
- **Long context:** on Claude 4.6 and later models, the full one-million-token context is billed at the standard price, so a large cached prefix does not trigger a long-context premium.

## Price your own prefix

The [Claude token counter](https://chudi.dev/tools/claude-token-counter) estimates the tokens in any text you paste and prices them on every current Claude model. It also takes cache-read and cache-write token counts, with a five-minute or one-hour lifetime, and prices those at each model’s own multiplier. Paste your system prompt and tool definitions to see what a write and each read cost.

Cache multipliers differ by model: Opus 5.5 and Fable 5.1 both break the usual 0.1x rule. [Claude Cost Watch](https://chudi.dev/products/claude-cost-watch) is a waitlist for a short email that flags changes like that. Nothing is charged today.

## Sources

All read on 09/28/2026.

- Cache multipliers, per-model prices, break-even rule, batch and long-context pricing: [platform.claude.com/docs/en/about-claude/pricing](https://platform.claude.com/docs/en/about-claude/pricing)
- Measured Claude Code day: [Opus 5.5 vs Fable 5.1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1)
- Full price list: [Claude pricing](https://chudi.dev/blog/claude-pricing)

No link on this page is an affiliate link.

· Frequently asked

## FAQ

### How much does Claude prompt caching cost?

A cache write costs 1.25x the model's input price for the five-minute cache and 2x for the one-hour cache. A cache read costs 0.1x the input price on most models, 0.05x on Opus 5.5 and 0.025x on Fable 5.1. On Sonnet 5.5 that is $2.50 or $4 per million tokens to write and $0.20 to read. Source: platform.claude.com pricing page, 09/28/2026.

### When does prompt caching pay off?

A five-minute cache write pays off after one cache read. A one-hour cache write pays off after two. If the same prefix is not reused within the cache lifetime, caching costs more than not caching.

### Should I use the five-minute or the one-hour cache?

Use the five-minute cache when requests reusing the prefix arrive within minutes of each other, as in an agent loop or a chat session. Use the one-hour cache when reuse is spread out, for example a large document queried every 10 to 30 minutes. The one-hour write costs 2x instead of 1.25x.

### Does prompt caching stack with the Batch API discount?

Yes. Anthropic's pricing page states that the batch discount and prompt caching discounts can be combined. Batch gives 50% off input and output.

### Why does caching matter so much for Claude Code and agents?

An agent re-sends its whole context on every turn, so most of its input is the same prefix read again. In one day of Claude Code work I measured on Opus 5, cache reads were the largest single line of the bill: $170 of $359.

· Sources & further reading

## Sources & Further Reading

### Further reading

- [Claude Pricing in 2026: Every Plan and API Price on One Page /blog/claude-pricing](https://chudi.dev/blog/claude-pricing) Claude pricing as of 09/28/2026: Free, Pro ($20), Max ($100 or $200), Team and Enterprise plans, plus per-token API prices for Fable 5.1, Opus 5.5, Sonnet 5.5 and Haiku 4.5.
- [Claude Pro vs Max: Who Needs the $100 or $200 Plan, and When the API Wins /blog/claude-pro-vs-max](https://chudi.dev/blog/claude-pro-vs-max) Claude Pro vs Max as of 09/28/2026: $20 vs $100 or $200 a month, 5x or 20x the usage, both with Claude Code. How to choose, and what one measured day costs on the API.
- [Claude Usage Limits Explained: The 5-Hour Session, the Weekly Cap, and What Counts /blog/claude-usage-limits](https://chudi.dev/blog/claude-usage-limits) How Claude usage limits work on Pro and Max: a five-hour session limit, a weekly limit, one pool shared by claude.ai, Desktop and Claude Code, and what to do when you hit it.
- [Opus 5.5 vs Fable 5.1: My Cost Per Turn Gap Went From 1.4x to 2.3x /blog/claude-opus-5-5-vs-fable-5-1](https://chudi.dev/blog/claude-opus-5-5-vs-fable-5-1) Opus 5.5 beats Fable 5.1 on all nine of Anthropic's launch benchmarks at 40% of the price. Replayed at 5.5 prices, my measured day puts Fable at 2.3x per turn.
- [Fable 5.1 vs Opus 5 in Production: A Whole Day on Fable Used 19% of a Max Weekly Window /blog/claude-fable-5-vs-opus-5](https://chudi.dev/blog/claude-fable-5-vs-opus-5) Same harness, same day: 2,443 Fable 5.1 turns and 1,120 Opus 5 turns after a weekly reset consumed 19% of the Max weekly window, and the 5-hour window peaked at 59% without capping. Fable errored less per tool call, hit more permission gates, and would have cost 1.4x per turn on the API, not 2x.

Built from these systems

Chiron Seed $29

Four Claude Code skills and two scripts that keep a knowledge base inside your repo, so the next session starts from what the last one learned.

[Get the kit](https://7124296224445.gumroad.com/l/chiron-seed)

Want more of this in your Google results?

## What do you think?

I post about this stuff on LinkedIn every day and the conversations there are great. If this post sparked a thought, I'd love to hear it.

[Discuss on LinkedIn](https://www.linkedin.com/in/chudi-nnorukam)
