{"slug": "your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix", "title": "Your cache hit rate is lying to you. It's the caller mix.", "summary": "A developer's analysis of Claude Code prompt caching shows that account-level cache hit rates mask wildly different behavior across callers, with one deployment reporting 77% overall but roughly 0% for cron tasks that rewrite a ~30K-token prefix every run — about 20M uncached input tokens daily. The writeup reports that a blanket 1-hour TTL made the bill 8.6% worse because 98% of reuse lands within about 34 seconds of the write, while targeted fixes — a 1-hour write on the dispatch turn, a persistent per-type static prefix, and moving dynamic content after the stable prefix — cut costs 13.6% combined. It also notes that subagents cannot read a parent's cache on their first request, and that a median ~9-minute child runtime pushed 96% of parent waits past the cache cliff.", "body_md": "One Claude Code deployment: **77% cache hit rate**, all green. Same account, one tab over, cron tasks at **~0%**, because every run is a fresh session rewriting a ~30K-token static prefix. That is 672 runs a day times 30K tokens, roughly **20M input tokens daily** that never touch a cache. The operator posted those numbers on an open issue, and they make my argument better than I can: an account-level hit rate averages over callers that share nothing. Same account, four realities.\n\nForget dashboards for a second. Caching is priced with three multipliers, and they're the same on every model:\n\n| Operation | Multiplier | Example: Opus 5.5 at $4/MTok input | \n|---|---|---|\n| Write, 5-minute TTL | 1.25× | $5.00/MTok | \n| Write, 1-hour TTL | 2× | $8.00/MTok | \n| Read (either TTL) | 0.1× on most models | $0.20/MTok (0.05× on Opus 5.5, 0.025× on Fable 5.1) | \n\nThree rules fall out of that table, and I've been quoting all three this week.\n\n**One read pays for the 5-minute cache.** Write at 1.25×, read at 0.1×, you're at 1.35× against 2× for two uncached passes. The 1-hour write needs two reads.\n\n**Match the TTL to the gap between calls.** Under five minutes: default. Five to sixty minutes: pay the 2× write once, ride the cheap reads. Over an hour: don't cache. A breakpoint nobody reuses doesn't save money, it costs 25% extra on the 5-minute tier and 100% on the 1-hour tier.\n\n**Refreshes are free, and the clock lies a little.** Every hit renews the TTL at the read price, so a busy cache lives forever. Keepalive math crosses over at **62.5 minutes** (5 × 1.25/0.10), same for every model because it's a ratio. The countdown starts when the *request* starts: a turn streaming four minutes leaves your next call one minute of window.\n\nClaude Code picks the TTL per request, in two buckets (their docs, not a blog post):\n\n| Billing | Main conversation | Everything else | \n|---|---|---|\n| Claude subscription, within plan usage | 1-hour | 5-minute | \n| API key, cloud provider, or past your plan limit | 5-minute | 5-minute | \n\n\"Everything else\" is where agents live: subagents, workflows, forks, compaction calls, session titles. You can override both buckets (`promptCacheTtl`, `subagentPromptCacheTtl`, `ENABLE_PROMPT_CACHING_1H`, v2.1.242+). Hold off, though; the cliff section explains why blanket 1h is worse.\n\nThe structural bit: **a subagent's first request cannot read the parent's cache.** Different prompt, different tools, prefixes diverge at token one, so it warms its own. A fork inherits the parent's prefix exactly and hits on the first request. Same account, opposite behavior, one screen apart.\n\nOne honest wobble: the docs say subagents get 5 minutes, a user's transcripts showed 100% of their subagent writes in the 1-hour bucket, and a maintainer said the effective TTL gets decided further down the pipeline while they fix the docs. Don't trust the table or me. Trust `usage.cache_creation.ephemeral_5m_input_tokens` and `ephemeral_1h_input_tokens` in your own responses.\n\nHere's the failure that pays for this post.\n\nA parent dispatches a subagent and waits. No requests go out, so nothing refreshes its cache. The measured median child runtime was about **9 minutes**, just past the cliff, and **96%** of those waits ended in a true cache death: when the parent resumed, at least half its cached prefix had to be rewritten at full price. The long wait, the one where the agent is doing exactly what you asked, is when the cache quietly dies.\n\nThen the finding I had to read twice: **blanket 1-hour TTL made the bill 8.6% worse.** 98% of reuse lands within about 34 seconds of the write (median gap: 7 seconds), so the 5-minute tier covers nearly all of it, and a universal 2× write premium taxes every write to rescue maybe 2% of reads. The fixes that worked: 1-hour write on the dispatch turn (−6.0%), persistent per-type static prefix (−1.0%), dynamic content after the stable prefix (−7.6%). All three: **−13.6%**.\n\n**TTL is a property of one write at one moment, not a setting on an account.** Get that backwards and you collect both failure modes: premium where you don't need it, death where you do.\n\nReader `hannune` left this on our cost-loop post: his agent looked perfect in the logs, every response successful, then a three-day-weekend invoice made no sense. He pinned it to a context-retrieval step pulling full documents instead of chunks: **40K tokens per call, seven or eight calls per task**. A hit-rate dashboard shows nothing wrong there. Every call succeeded.\n\nHis rule, which I've adopted: record the **TTL tier per caller, not per account**. One extra label next to `cache_creation` and `cache_read`: caller × TTL tier. It's the attribution argument from our cost-attribution post, one level lower. We measured a single agent run at 3.5M input tokens against 271K output; the cache either pays on that input side or leaks. No third place for your money to go.\n\nEach of these fails without an error:\n\n`cache_creation` just reads zero.\n| Caller | Default TTL | Reads parent cache? | Typical gap | Do this | \n|---|---|---|---|---|\n| Main loop | 1h (subscription) / 5m (API) | n/a | seconds | On API key: set `promptCacheTtl=1h` | \n| Subagent | 5m | No, own prefix | seconds inside a type, minutes between types | 1h write on the shared per-type prefix; dynamic after the breakpoint | \n| Fork | parent's | Yes | immediate | Use it for side work that must see history | \n| Compaction | parent's, *only if same prefix* | must | one call | Verify the prefix is reused | \n| Cron / scheduled | 5m, fresh session | No | hours | Expect ~0%; shrink the static prefix or budget the writes | \n\n`ttl_tier × caller` beside your cache fields.`ephemeral_5m`, `ephemeral_1h`, `cache_read`) from response usage. Vendor summaries average away what you're looking for.`cache_control`; date, cwd, and branch after it.` ttl:\"1h\"` in two places only: the dispatch turn and shared static prefixes.\n**How do I see which TTL a request used?** `usage.cache_creation` in the response, or the same field in your transcript files. `ephemeral_1h_input_tokens` vs `ephemeral_5m_input_tokens` are mutually exclusive buckets.\n\n**Do other providers work like this?** Differently, same lesson. OpenAI caches automatically with a discount and no TTL knob; Google prices storage per hour. Wherever prefixes belong to conversations, the per-caller problem holds.\n\n**My hit rate is 90% and my bill is fine. Do I care?** No. This one is for people whose dashboard and invoice disagree.\n\nWhat's your worst cache own-goal? Dashboard healthy, invoice confusing, and one line in a transcript finally explained it. I'll go first if you do.\n\n**Related:** [Your agent's cost problem isn't the model. It's the loop.](https://dev.to/theagentloop/your-agents-cost-problem-isnt-the-model-its-the-loop-13a3) · [Your cost dashboard can't tell you which agent ran up the bill](https://dev.to/theagentloop/your-cost-dashboard-cant-tell-you-which-agent-ran-up-the-bill-4eej) · [What one agent run actually costs](https://dev.to/theagentloop/what-one-agent-run-actually-costs-28i8)\n\n`hannune`, 2026-09-26), Independent blog: not affiliated with Anthropic or the Claude Code team. Every figure above comes from public docs and issues, verified on the dates given.\n\nIf you want the rest of this series when it drops: [subscribe via Buttondown](https://buttondown.com/theagentloop) and reply with your cache horror story. One of these turns into a post.", "url": "https://wpnews.pro/news/your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix", "canonical_source": "https://dev.to/theagentloop/your-cache-hit-rate-is-lying-to-you-its-the-caller-mix-ec1", "published_at": "2026-10-04 06:02:34+00:00", "updated_at": "2026-10-04 06:07:50.427200+00:00", "lang": "en", "topics": ["ai-agents", "ai-infrastructure", "large-language-models", "ai-tools", "mlops"], "entities": ["Claude Code", "Anthropic", "Opus 5.5", "Fable 5.1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix", "markdown": "https://wpnews.pro/news/your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix.md", "text": "https://wpnews.pro/news/your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix.txt", "jsonld": "https://wpnews.pro/news/your-cache-hit-rate-is-lying-to-you-it-s-the-caller-mix.jsonld"}}