{"slug": "claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it", "title": "Claude Prompt Caching Hit Rate Was 0%. One Timestamp Did It", "summary": "A developer running Preterview, a voice interview practice platform, discovered that Claude prompt caching produced a 0% hit rate across 3,412 calls because a `datetime.now()` timestamp placed at the top of the system prompt made every request prefix unique, forcing a cache write premium on every call and raising costs above no caching at all. After moving volatile content to the end of the request, the hit rate reached 78%, with a second bug caused by Python hash randomization keeping it from going higher. The developer now logs cache hit ratio from the `cache_read_input_tokens`, `cache_creation_input_tokens` and `input_tokens` usage fields on every response.", "body_md": "I turned on Claude prompt caching, shipped it, and moved on. Nine days later I looked at the usage logs and saw `cache_read_input_tokens: 0` on every one of 3,412 calls.\n\nNot a low hit rate. Zero. And it was worse than zero, because every one of those calls paid the cache *write* premium. Turning on caching had made my bill go up.\n\nThe cause was one line: a timestamp at the top of my system prompt. After fixing that, the hit rate stalled at 78% because of a second bug that Python's hash randomization hid from me. Here's the full autopsy.\n\n`tools` → `system` → `messages`, up to the `cache_control` breakpoint. Change one character early and everything after it misses.`datetime.now()` in my system prompt made every request unique: `set`, and set order differs `sorted()` fixed it: `usage` fields. Log them.\nThe app is Preterview, a voice interview practice platform I run (full disclosure: I built it). A session is a 30-minute spoken interview with one of 3 interviewer styles, and every candidate turn becomes an LLM call. The average session in my logs is 22 calls, and each one resends the same ~6,500-token prefix: interviewer persona, scoring rubric, the candidate's resume, and 5 tool definitions. That's the textbook case for caching. You can see what a session looks like at [preterview.com/en](https://preterview.com/en).\n\nClaude prompt caching hashes the request prefix in a fixed order (tools first, then system, then messages) up to each block marked with `cache_control`. If that exact prefix was written in the last 5 minutes, you get a read. If not, you get a write, and the entry lives for the next 5 minutes (refreshed on every hit).\n\nThe pricing is what makes misses hurt:\n\nSo a prefix that never gets read back is strictly worse than not caching at all. You pay 25% extra for the privilege of storing something nobody reuses.\n\nBecause the first thing in my system prompt changed on every single call. Here's what I shipped:\n\n```\nsystem = [{\n    \"type\": \"text\",\n    \"text\": (\n        f\"Current time: {datetime.now().isoformat()}\\n\\n\"\n        f\"{PERSONAS[style]}\\n\\n{RUBRIC}\\n\\n{resume_md}\"\n    ),\n    \"cache_control\": {\"type\": \"ephemeral\"},\n}]\n```\n\nI added the timestamp so the interviewer could pace itself (\"we have about ten minutes left, let's move to system design\"). Reasonable feature. Terrible placement.\n\n`isoformat()` includes microseconds. Every request had a different first line, so every request had a different prefix, so every request was a cache write. The response came back fine. The latency was fine. The only evidence was in `usage`:\n\n```\n{\n  \"input_tokens\": 212,\n  \"cache_creation_input_tokens\": 6488,\n  \"cache_read_input_tokens\": 0,\n  \"output_tokens\": 97\n}\n```\n\nLook at that for turn 14 of a session and it screams. Look at the monthly invoice and it just says \"input tokens, slightly more than expected.\"\n\nCompute it from the three input fields on every response. There's no dashboard toggle that does it for you per request, so I now emit this on every call:\n\n``` python\ndef log_cache(resp, session_id, style):\n    u = resp.usage\n    read = u.cache_read_input_tokens or 0\n    write = u.cache_creation_input_tokens or 0\n    fresh = u.input_tokens\n    total = read + write + fresh\n    metrics.emit(\n        \"llm.cache_hit_ratio\",\n        read / total if total else 0.0,\n        tags={\"session\": session_id, \"style\": style},\n    )\n```\n\n`input_tokens` is only the uncached remainder, not the total. That tripped me up at first: I thought small `input_tokens` meant caching was working. It meant the tokens were going into `cache_creation_input_tokens` instead.\n\nMove anything volatile to the end of the request. The time now rides along in the latest user message, after all the stable content:\n\n```\nsystem = [\n    {\"type\": \"text\", \"text\": f\"{PERSONAS[style]}\\n\\n{RUBRIC}\",\n     \"cache_control\": {\"type\": \"ephemeral\"}},\n    {\"type\": \"text\", \"text\": resume_md,\n     \"cache_control\": {\"type\": \"ephemeral\"}},\n]\nmessages = history + [{\n    \"role\": \"user\",\n    \"content\": f\"{transcript_chunk}\\n\\n[elapsed: {elapsed_min} of 30 min]\",\n}]\n```\n\nTwo breakpoints on purpose. Persona plus rubric is identical for every candidate who picked the same style, so concurrent sessions can share it. The resume is per candidate.\n\nOne more trap here: `history` must contain the exact strings you sent before. I was briefly re-rendering old turns with a fresh `elapsed` value, which would have broken caching on the conversation too. Append, never regenerate.\n\nHit rate after deploying: **78%**. Better. But I expected low 90s, and it sat at 78 for two days.\n\nBecause my tool definitions came out in a different order depending on which worker process handled the request. Tools are the very first thing in the cache prefix, so a reordered tools list invalidates everything after it.\n\nI found it by hashing the prefix on every call:\n\n```\nprefix = json.dumps({\"tools\": req[\"tools\"], \"system\": req[\"system\"]})\nlog.info(\"prefix_hash\", h=hashlib.sha256(prefix.encode()).hexdigest()[:12],\n         style=style, pid=os.getpid())\n```\n\nSame style, same resume, same session: **4 distinct hashes**. I run 4 uvicorn workers. Each hash mapped to exactly one PID.\n\nHere's the code that built the tools:\n\n```\nenabled = {\"record_score\", \"flag_followup\", \"end_section\", \"finish_interview\"}\nif style == \"pressure\":\n    enabled.add(\"interrupt\")\ntools = [TOOL_DEFS[name] for name in enabled]\n```\n\nPython randomizes string hashing per process unless `PYTHONHASHSEED` is set. Set iteration order depends on those hashes. So each worker iterated the same set in its own stable order, and the load balancer spread each session across all four. Worst case, a session paid for 4 cache writes instead of 1.\n\nOn my laptop, with one process, this bug is invisible. Every local test passed.\n\nThe fix is one word:\n\n```\ntools = [TOOL_DEFS[name] for name in sorted(enabled)]\n```\n\nHit rate after that: **93%**. The remaining 7% is honest: the first call of every session has to write (1 of 22, about 4.5%), each new user turn is fresh input by definition, and a few candidates pause longer than 5 minutes and let the entry expire.\n\nThe cached prefix got about 88% cheaper per session compared with the buggy version. Here's the modeled cost of the ~6,500-token prefix across a 22-call session, in units of \"one full-price prefix\":\n\n| Stage | Measured hit rate | Prefix cost per session | \n|---|---|---|\n| No caching at all | n/a | 22.0 | \n| Timestamp bug (22 writes) | 0% | 27.5 | \n| Tool order bug (~4 writes, 18 reads) | 78% | 6.8 | \n| Fixed (1 write, 21 reads) | 93% | 3.35 | \n\nThe timestamp bug was 25% more expensive than never enabling caching. The tool order bug looked like a success (78% sounds fine!) while still costing twice what it should.\n\nWhat I haven't done yet: putting a breakpoint on the growing conversation history. By turn 20 the history is a few thousand tokens, and right now I pay full price for it on every turn. That's the next experiment.\n\nTwo tests, one free and one cheap.\n\nThe free one runs in CI with no API calls. It builds the request prefix in two subprocesses with different hash seeds and asserts the bytes match:\n\n``` python\ndef test_prefix_is_deterministic_across_processes():\n    outs = []\n    for seed in (\"1\", \"2\"):\n        out = subprocess.run(\n            [sys.executable, \"-m\", \"app.dump_prefix\", \"--style\", \"pressure\"],\n            env={**os.environ, \"PYTHONHASHSEED\": seed},\n            capture_output=True, check=True,\n        ).stdout\n        outs.append(out)\n    assert outs[0] == outs[1], \"request prefix depends on process hash seed\"\n```\n\nThat test would have caught bug #2 before it shipped. Grepping the prompt builders for `now()`, `uuid4()` and `random` would have caught bug #1.\n\nThe cheap one runs nightly against the real API: two consecutive calls on a fixture session, then assert the second has `cache_read_input_tokens > 0`. It costs a fraction of a cent and fails loudly the day someone \"improves\" the system prompt with a request ID.\n\nOne last gotcha worth knowing: each model has a minimum cacheable prompt length, and below it `cache_control` is silently ignored. No error, just no caching. If your prefix is short, check the docs for your model before assuming it works.\n\nClaude prompt caching requires the request prefix (tools, then system, then messages, up to the `cache_control` breakpoint) to be byte-identical to a recent request, and I put `datetime.now()` at the top of my system prompt, so no two requests ever matched. Every call paid the 1.25x write price and none got the 0.1x read price, which made caching more expensive than turning it off. Moving volatile values into the latest user message fixed the 0%, and sorting a tools list built from a Python `set` fixed a second, per-worker ordering bug that capped me at 78%. If you use prompt caching, log `cache_read_input_tokens` and `cache_creation_input_tokens` on every call, because the API will never tell you it isn't working.\n\n*Written by the developer behind [Preterview](https://preterview.com/en), an interview prep platform.*", "url": "https://wpnews.pro/news/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it", "canonical_source": "https://dev.to/ji_ai/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it-ic6", "published_at": "2026-09-29 22:59:17+00:00", "updated_at": "2026-09-29 23:17:03.914312+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["Claude", "Anthropic", "Preterview", "Python"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it", "markdown": "https://wpnews.pro/news/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it.md", "text": "https://wpnews.pro/news/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it.txt", "jsonld": "https://wpnews.pro/news/claude-prompt-caching-hit-rate-was-0-one-timestamp-did-it.jsonld"}}