I turned on Claude prompt caching, shipped it, and moved on. Nine days later I looked at the usage logs and saw cache_read_input_tokens: 0 on every one of 3,412 calls.
Not a low hit rate. Zero. And it was worse than zero, because every one of those calls paid the cache write premium. Turning on caching had made my bill go up.
The cause was one line: a timestamp at the top of my system prompt. After fixing that, the hit rate stalled at 78% because of a second bug that Python's hash randomization hid from me. Here's the full autopsy.
tools → system → messages, up to the cache_control breakpoint. Change one character early and everything after it misses.datetime.now() in my system prompt made every request unique: set, and set order differs sorted() fixed it: usage fields. Log them.
The app is Preterview, a voice interview practice platform I run (full disclosure: I built it). A session is a 30-minute spoken interview with one of 3 interviewer styles, and every candidate turn becomes an LLM call. The average session in my logs is 22 calls, and each one resends the same ~6,500-token prefix: interviewer persona, scoring rubric, the candidate's resume, and 5 tool definitions. That's the textbook case for caching. You can see what a session looks like at preterview.com/en.
Claude prompt caching hashes the request prefix in a fixed order (tools first, then system, then messages) up to each block marked with cache_control. If that exact prefix was written in the last 5 minutes, you get a read. If not, you get a write, and the entry lives for the next 5 minutes (refreshed on every hit).
The pricing is what makes misses hurt:
So a prefix that never gets read back is strictly worse than not caching at all. You pay 25% extra for the privilege of storing something nobody reuses.
Because the first thing in my system prompt changed on every single call. Here's what I shipped:
system = [{
"type": "text",
"text": (
f"Current time: {datetime.now().isoformat()}\n\n"
f"{PERSONAS[style]}\n\n{RUBRIC}\n\n{resume_md}"
),
"cache_control": {"type": "ephemeral"},
}]
I added the timestamp so the interviewer could pace itself ("we have about ten minutes left, let's move to system design"). Reasonable feature. Terrible placement.
isoformat() includes microseconds. Every request had a different first line, so every request had a different prefix, so every request was a cache write. The response came back fine. The latency was fine. The only evidence was in usage:
{
"input_tokens": 212,
"cache_creation_input_tokens": 6488,
"cache_read_input_tokens": 0,
"output_tokens": 97
}
Look at that for turn 14 of a session and it screams. Look at the monthly invoice and it just says "input tokens, slightly more than expected."
Compute it from the three input fields on every response. There's no dashboard toggle that does it for you per request, so I now emit this on every call:
def log_cache(resp, session_id, style):
u = resp.usage
read = u.cache_read_input_tokens or 0
write = u.cache_creation_input_tokens or 0
fresh = u.input_tokens
total = read + write + fresh
metrics.emit(
"llm.cache_hit_ratio",
read / total if total else 0.0,
tags={"session": session_id, "style": style},
)
input_tokens is only the uncached remainder, not the total. That tripped me up at first: I thought small input_tokens meant caching was working. It meant the tokens were going into cache_creation_input_tokens instead.
Move anything volatile to the end of the request. The time now rides along in the latest user message, after all the stable content:
system = [
{"type": "text", "text": f"{PERSONAS[style]}\n\n{RUBRIC}",
"cache_control": {"type": "ephemeral"}},
{"type": "text", "text": resume_md,
"cache_control": {"type": "ephemeral"}},
]
messages = history + [{
"role": "user",
"content": f"{transcript_chunk}\n\n[elapsed: {elapsed_min} of 30 min]",
}]
Two breakpoints on purpose. Persona plus rubric is identical for every candidate who picked the same style, so concurrent sessions can share it. The resume is per candidate.
One more trap here: history must contain the exact strings you sent before. I was briefly re-rendering old turns with a fresh elapsed value, which would have broken caching on the conversation too. Append, never regenerate.
Hit rate after deploying: 78%. Better. But I expected low 90s, and it sat at 78 for two days.
Because my tool definitions came out in a different order depending on which worker process handled the request. Tools are the very first thing in the cache prefix, so a reordered tools list invalidates everything after it.
I found it by hashing the prefix on every call:
prefix = json.dumps({"tools": req["tools"], "system": req["system"]})
log.info("prefix_hash", h=hashlib.sha256(prefix.encode()).hexdigest()[:12],
style=style, pid=os.getpid())
Same style, same resume, same session: 4 distinct hashes. I run 4 uvicorn workers. Each hash mapped to exactly one PID.
Here's the code that built the tools:
enabled = {"record_score", "flag_followup", "end_section", "finish_interview"}
if style == "pressure":
enabled.add("interrupt")
tools = [TOOL_DEFS[name] for name in enabled]
Python randomizes string hashing per process unless PYTHONHASHSEED is set. Set iteration order depends on those hashes. So each worker iterated the same set in its own stable order, and the load balancer spread each session across all four. Worst case, a session paid for 4 cache writes instead of 1.
On my laptop, with one process, this bug is invisible. Every local test passed.
The fix is one word:
tools = [TOOL_DEFS[name] for name in sorted(enabled)]
Hit rate after that: 93%. The remaining 7% is honest: the first call of every session has to write (1 of 22, about 4.5%), each new user turn is fresh input by definition, and a few candidates longer than 5 minutes and let the entry expire.
The cached prefix got about 88% cheaper per session compared with the buggy version. Here's the modeled cost of the ~6,500-token prefix across a 22-call session, in units of "one full-price prefix":
| Stage | Measured hit rate | Prefix cost per session |
|---|---|---|
| No caching at all | n/a | 22.0 |
| Timestamp bug (22 writes) | 0% | 27.5 |
| Tool order bug (~4 writes, 18 reads) | 78% | 6.8 |
| Fixed (1 write, 21 reads) | 93% | 3.35 |
The timestamp bug was 25% more expensive than never enabling caching. The tool order bug looked like a success (78% sounds fine!) while still costing twice what it should.
What I haven't done yet: putting a breakpoint on the growing conversation history. By turn 20 the history is a few thousand tokens, and right now I pay full price for it on every turn. That's the next experiment.
Two tests, one free and one cheap.
The free one runs in CI with no API calls. It builds the request prefix in two subprocesses with different hash seeds and asserts the bytes match:
def test_prefix_is_deterministic_across_processes():
outs = []
for seed in ("1", "2"):
out = subprocess.run(
[sys.executable, "-m", "app.dump_prefix", "--style", "pressure"],
env={**os.environ, "PYTHONHASHSEED": seed},
capture_output=True, check=True,
).stdout
outs.append(out)
assert outs[0] == outs[1], "request prefix depends on process hash seed"
That test would have caught bug #2 before it shipped. Grepping the prompt builders for now(), uuid4() and random would have caught bug #1.
The cheap one runs nightly against the real API: two consecutive calls on a fixture session, then assert the second has cache_read_input_tokens > 0. It costs a fraction of a cent and fails loudly the day someone "improves" the system prompt with a request ID.
One last gotcha worth knowing: each model has a minimum cacheable prompt length, and below it cache_control is silently ignored. No error, just no caching. If your prefix is short, check the docs for your model before assuming it works.
Claude prompt caching requires the request prefix (tools, then system, then messages, up to the cache_control breakpoint) to be byte-identical to a recent request, and I put datetime.now() at the top of my system prompt, so no two requests ever matched. Every call paid the 1.25x write price and none got the 0.1x read price, which made caching more expensive than turning it off. Moving volatile values into the latest user message fixed the 0%, and sorting a tools list built from a Python set fixed a second, per-worker ordering bug that capped me at 78%. If you use prompt caching, log cache_read_input_tokens and cache_creation_input_tokens on every call, because the API will never tell you it isn't working.
Written by the developer behind Preterview, an interview prep platform.