{"slug": "claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now", "title": "Claude Cache Diagnostics Is GA: Debug Prompt Cache Misses Now", "summary": "Anthropic made cache diagnostics generally available on the Claude API as of September 23, letting developers pinpoint why prompt caches miss by passing a `diagnostics` object with a `previous_message_id` on each Messages request. The feature returns a `cache_miss_reason` union covering six causes — system_changed, tools_changed, model_changed, messages_changed, previous_message_not_found, and unavailable — plus a `cache_missed_input_tokens` estimate, and no longer requires the `cache-diagnosis-2026-04-07` beta header. Cache reads on Opus 5.5 cost $0.20 per million tokens versus a $4 base input price, a 95% reduction, and Anthropic says fingerprints store only cryptographic hashes and token-count estimates, are ZDR eligible, and are scoped to the organization and workspace.", "body_md": "Prompt caching is the single most effective cost lever for Claude API users. Cache reads on Opus 5.5 cost $0.20 per million tokens against a $4 base input price — a 95% reduction. Production teams running agents on this have [documented drops from $720 to $72 a month](https://usagebox.com/articles/prompt-caching-cost-optimization-claude-gpt-gemini-2026). The problem has always been that when caching stopped working, the only signal was `cache_read_input_tokens` going to zero. No explanation. No pointer. Just silent, expensive failure. As of September 23, that ends. [Cache diagnostics is now GA on the Claude API](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics).\n\n## What Cache Diagnostics Does\n\nThe API now compares consecutive requests for you. Opt in by including a `diagnostics` object on your Messages request. On the first turn, pass `previous_message_id: null` to register the fingerprint. On every subsequent turn, pass the `id` from the previous response. When the cache misses, the response tells you exactly where the prefix diverged: the model, the system prompt, the tools, or the message history.\n\nFingerprints contain only cryptographic hashes and token-count estimates — no raw prompt content. The feature is ZDR eligible, scoped to your organization and workspace, and stored only for requests that include the `diagnostics` object.\n\nThe `cache-diagnosis-2026-04-07` beta header is no longer required. Drop it or keep it — either works.\n\n## The Six Miss-Reason Types\n\nThe response carries a `cache_miss_reason` discriminated union on `type`. Here is what each one means and what to do about it.\n\n- **system_changed** — The most common culprit. A timestamp, request ID, or session token was interpolated into the system prompt. Move dynamic data into the first user message after your cache breakpoint.\n- **tools_changed** — Tools were added, removed, reordered, or the JSON schema serialized non-deterministically between turns. Send the same tool list in a fixed order with deterministic key sorting on every turn.\n- **model_changed** — A router, A/B test, or fallback mechanism selected a different model mid-conversation. The cache is per-model. Hold the model constant for the lifetime of a cached conversation.\n- **messages_changed** — History was truncated or assistant turns re-serialized differently. Treat message history as append-only and echo assistant`content` blocks back verbatim.\n- **previous_message_not_found** — The prior request did not include the`diagnostics` object, it ran in a different workspace, or the fingerprint expired. Include`diagnostics` on every turn and keep turns closely spaced.\n- **unavailable** — A non-standard parameter changed (`tool_choice` , thinking config, context management), or the divergence is beyond the comparison horizon on a very long conversation. Keep prompt-affecting parameters constant across a conversation.\n\nEach `*_changed` type also carries a `cache_missed_input_tokens` field — an estimate of how many tokens fell past the divergence point. Use this to quantify the cost impact before prioritizing the fix.\n\n## How to Add It in Three Minutes\n\nThe change is two lines per turn. Here is the Python loop pattern:\n\n```\nclient = anthropic.Anthropic()\nprev_id = None\n\nfor user_msg in conversation_turns:\n    messages.append({\"role\": \"user\", \"content\": user_msg})\n\n    r = client.beta.messages.create(\n        model=\"claude-opus-5-5\",\n        max_tokens=1024,\n        cache_control={\"type\": \"ephemeral\"},\n        system=SYSTEM,\n        messages=messages,\n        diagnostics={\"previous_message_id\": prev_id},\n    )\n\n    if r.diagnostics and r.diagnostics.cache_miss_reason:\n        print(f\"Cache miss: {r.diagnostics.cache_miss_reason.type}\")\n\n    messages.append({\"role\": \"assistant\", \"content\": r.content})\n    prev_id = r.id\n```\n\nOn the first iteration, `prev_id` is `None` — the API opts in and stores a fingerprint, nothing to compare yet. On every subsequent turn it compares and reports. Threading that `prev_id` forward is the entire integration.\n\n## Reading the Results\n\nThe `diagnostics` field and `usage.cache_read_input_tokens` answer different questions. Combine them:\n\n- **diagnostics null + high cache reads** — Working correctly. No action needed.\n- **diagnostics null + low cache reads** — Your request structure is stable but the cache entry expired. Switch to the[1-hour TTL option](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) .\n- **`*_changed` + low cache reads** — Your code changed something it should not have. Fix the cause the type points to.\n- **`*_changed` + high cache reads** — The divergence happened late in the prompt and an earlier cache breakpoint still hit. Worth fixing, but low urgency.\n\n## What Does Not Work Yet\n\nCache diagnostics is Claude API only. It is not available on Amazon Bedrock or Google Cloud. Fingerprints expire after a short period, so comparing requests far apart in time returns `previous_message_not_found`, not a miss reason. And for very long conversations where the divergence is deep in the message history, you will get `unavailable` rather than a precise location.\n\nNone of these are dealbreakers. The common failure modes — a timestamp in a system prompt, a reordered tool list, a model swap — are all caught cleanly. [Analysts covering the release](https://www.digitalapplied.com/blog/prompt-caching-changes-openai-anthropic-september-2026) noted this mirrors what OpenAI already shipped, but Anthropic’s implementation returns the precise divergence point rather than a binary hit/miss.\n\n## Add It to Every New Agent You Build\n\nThe opt-in costs nothing. The fingerprint is two lines of code. When your cache breaks in production — and at some point it will — you will know within one request whether the system prompt changed, the tool list shifted, or the history got mangled. That is the difference between a five-minute fix and a three-hour investigation.\n\nCache diagnostics is not a debugging convenience. It is cost infrastructure. Treat it that way.", "url": "https://wpnews.pro/news/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now", "canonical_source": "https://byteiota.com/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now/", "published_at": "2026-09-28 01:08:39+00:00", "updated_at": "2026-09-28 01:29:44.045094+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "large-language-models", "ai-infrastructure"], "entities": ["Anthropic", "Claude API", "Claude Opus 5.5", "cache diagnostics", "cache-diagnosis-2026-04-07"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now", "markdown": "https://wpnews.pro/news/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now.md", "text": "https://wpnews.pro/news/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now.txt", "jsonld": "https://wpnews.pro/news/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now.jsonld"}}