Claude Prompt Caching Hit Rate Was 0%. One Timestamp Did It A developer running Preterview, a voice interview practice platform, discovered that Claude prompt caching produced a 0% hit rate across 3,412 calls because a `datetime.now()` timestamp placed at the top of the system prompt made every request prefix unique, forcing a cache write premium on every call and raising costs above no caching at all. After moving volatile content to the end of the request, the hit rate reached 78%, with a second bug caused by Python hash randomization keeping it from going higher. The developer now logs cache hit ratio from the `cache_read_input_tokens`, `cache_creation_input_tokens` and `input_tokens` usage fields on every response. I turned on Claude prompt caching, shipped it, and moved on. Nine days later I looked at the usage logs and saw cache read input tokens: 0 on every one of 3,412 calls. Not a low hit rate. Zero. And it was worse than zero, because every one of those calls paid the cache write premium. Turning on caching had made my bill go up. The cause was one line: a timestamp at the top of my system prompt. After fixing that, the hit rate stalled at 78% because of a second bug that Python's hash randomization hid from me. Here's the full autopsy. tools → system → messages , up to the cache control breakpoint. Change one character early and everything after it misses. datetime.now in my system prompt made every request unique: set , and set order differs sorted fixed it: usage fields. Log them. The app is Preterview, a voice interview practice platform I run full disclosure: I built it . A session is a 30-minute spoken interview with one of 3 interviewer styles, and every candidate turn becomes an LLM call. The average session in my logs is 22 calls, and each one resends the same ~6,500-token prefix: interviewer persona, scoring rubric, the candidate's resume, and 5 tool definitions. That's the textbook case for caching. You can see what a session looks like at preterview.com/en https://preterview.com/en . Claude prompt caching hashes the request prefix in a fixed order tools first, then system, then messages up to each block marked with cache control . If that exact prefix was written in the last 5 minutes, you get a read. If not, you get a write, and the entry lives for the next 5 minutes refreshed on every hit . The pricing is what makes misses hurt: So a prefix that never gets read back is strictly worse than not caching at all. You pay 25% extra for the privilege of storing something nobody reuses. Because the first thing in my system prompt changed on every single call. Here's what I shipped: system = { "type": "text", "text": f"Current time: {datetime.now .isoformat }\n\n" f"{PERSONAS style }\n\n{RUBRIC}\n\n{resume md}" , "cache control": {"type": "ephemeral"}, } I added the timestamp so the interviewer could pace itself "we have about ten minutes left, let's move to system design" . Reasonable feature. Terrible placement. isoformat includes microseconds. Every request had a different first line, so every request had a different prefix, so every request was a cache write. The response came back fine. The latency was fine. The only evidence was in usage : { "input tokens": 212, "cache creation input tokens": 6488, "cache read input tokens": 0, "output tokens": 97 } Look at that for turn 14 of a session and it screams. Look at the monthly invoice and it just says "input tokens, slightly more than expected." Compute it from the three input fields on every response. There's no dashboard toggle that does it for you per request, so I now emit this on every call: python def log cache resp, session id, style : u = resp.usage read = u.cache read input tokens or 0 write = u.cache creation input tokens or 0 fresh = u.input tokens total = read + write + fresh metrics.emit "llm.cache hit ratio", read / total if total else 0.0, tags={"session": session id, "style": style}, input tokens is only the uncached remainder, not the total. That tripped me up at first: I thought small input tokens meant caching was working. It meant the tokens were going into cache creation input tokens instead. Move anything volatile to the end of the request. The time now rides along in the latest user message, after all the stable content: system = {"type": "text", "text": f"{PERSONAS style }\n\n{RUBRIC}", "cache control": {"type": "ephemeral"}}, {"type": "text", "text": resume md, "cache control": {"type": "ephemeral"}}, messages = history + { "role": "user", "content": f"{transcript chunk}\n\n elapsed: {elapsed min} of 30 min ", } Two breakpoints on purpose. Persona plus rubric is identical for every candidate who picked the same style, so concurrent sessions can share it. The resume is per candidate. One more trap here: history must contain the exact strings you sent before. I was briefly re-rendering old turns with a fresh elapsed value, which would have broken caching on the conversation too. Append, never regenerate. Hit rate after deploying: 78% . Better. But I expected low 90s, and it sat at 78 for two days. Because my tool definitions came out in a different order depending on which worker process handled the request. Tools are the very first thing in the cache prefix, so a reordered tools list invalidates everything after it. I found it by hashing the prefix on every call: prefix = json.dumps {"tools": req "tools" , "system": req "system" } log.info "prefix hash", h=hashlib.sha256 prefix.encode .hexdigest :12 , style=style, pid=os.getpid Same style, same resume, same session: 4 distinct hashes . I run 4 uvicorn workers. Each hash mapped to exactly one PID. Here's the code that built the tools: enabled = {"record score", "flag followup", "end section", "finish interview"} if style == "pressure": enabled.add "interrupt" tools = TOOL DEFS name for name in enabled Python randomizes string hashing per process unless PYTHONHASHSEED is set. Set iteration order depends on those hashes. So each worker iterated the same set in its own stable order, and the load balancer spread each session across all four. Worst case, a session paid for 4 cache writes instead of 1. On my laptop, with one process, this bug is invisible. Every local test passed. The fix is one word: tools = TOOL DEFS name for name in sorted enabled Hit rate after that: 93% . The remaining 7% is honest: the first call of every session has to write 1 of 22, about 4.5% , each new user turn is fresh input by definition, and a few candidates pause longer than 5 minutes and let the entry expire. The cached prefix got about 88% cheaper per session compared with the buggy version. Here's the modeled cost of the ~6,500-token prefix across a 22-call session, in units of "one full-price prefix": | Stage | Measured hit rate | Prefix cost per session | |---|---|---| | No caching at all | n/a | 22.0 | | Timestamp bug 22 writes | 0% | 27.5 | | Tool order bug ~4 writes, 18 reads | 78% | 6.8 | | Fixed 1 write, 21 reads | 93% | 3.35 | The timestamp bug was 25% more expensive than never enabling caching. The tool order bug looked like a success 78% sounds fine while still costing twice what it should. What I haven't done yet: putting a breakpoint on the growing conversation history. By turn 20 the history is a few thousand tokens, and right now I pay full price for it on every turn. That's the next experiment. Two tests, one free and one cheap. The free one runs in CI with no API calls. It builds the request prefix in two subprocesses with different hash seeds and asserts the bytes match: python def test prefix is deterministic across processes : outs = for seed in "1", "2" : out = subprocess.run sys.executable, "-m", "app.dump prefix", "--style", "pressure" , env={ os.environ, "PYTHONHASHSEED": seed}, capture output=True, check=True, .stdout outs.append out assert outs 0 == outs 1 , "request prefix depends on process hash seed" That test would have caught bug 2 before it shipped. Grepping the prompt builders for now , uuid4 and random would have caught bug 1. The cheap one runs nightly against the real API: two consecutive calls on a fixture session, then assert the second has cache read input tokens 0 . It costs a fraction of a cent and fails loudly the day someone "improves" the system prompt with a request ID. One last gotcha worth knowing: each model has a minimum cacheable prompt length, and below it cache control is silently ignored. No error, just no caching. If your prefix is short, check the docs for your model before assuming it works. Claude prompt caching requires the request prefix tools, then system, then messages, up to the cache control breakpoint to be byte-identical to a recent request, and I put datetime.now at the top of my system prompt, so no two requests ever matched. Every call paid the 1.25x write price and none got the 0.1x read price, which made caching more expensive than turning it off. Moving volatile values into the latest user message fixed the 0%, and sorting a tools list built from a Python set fixed a second, per-worker ordering bug that capped me at 78%. If you use prompt caching, log cache read input tokens and cache creation input tokens on every call, because the API will never tell you it isn't working. Written by the developer behind Preterview https://preterview.com/en , an interview prep platform.