Claude Cache Diagnostics Is GA: Debug Prompt Cache Misses Now Anthropic made cache diagnostics generally available on the Claude API as of September 23, letting developers pinpoint why prompt caches miss by passing a `diagnostics` object with a `previous_message_id` on each Messages request. The feature returns a `cache_miss_reason` union covering six causes — system_changed, tools_changed, model_changed, messages_changed, previous_message_not_found, and unavailable — plus a `cache_missed_input_tokens` estimate, and no longer requires the `cache-diagnosis-2026-04-07` beta header. Cache reads on Opus 5.5 cost $0.20 per million tokens versus a $4 base input price, a 95% reduction, and Anthropic says fingerprints store only cryptographic hashes and token-count estimates, are ZDR eligible, and are scoped to the organization and workspace. Prompt caching is the single most effective cost lever for Claude API users. Cache reads on Opus 5.5 cost $0.20 per million tokens against a $4 base input price — a 95% reduction. Production teams running agents on this have documented drops from $720 to $72 a month https://usagebox.com/articles/prompt-caching-cost-optimization-claude-gpt-gemini-2026 . The problem has always been that when caching stopped working, the only signal was cache read input tokens going to zero. No explanation. No pointer. Just silent, expensive failure. As of September 23, that ends. Cache diagnostics is now GA on the Claude API https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics . What Cache Diagnostics Does The API now compares consecutive requests for you. Opt in by including a diagnostics object on your Messages request. On the first turn, pass previous message id: null to register the fingerprint. On every subsequent turn, pass the id from the previous response. When the cache misses, the response tells you exactly where the prefix diverged: the model, the system prompt, the tools, or the message history. Fingerprints contain only cryptographic hashes and token-count estimates — no raw prompt content. The feature is ZDR eligible, scoped to your organization and workspace, and stored only for requests that include the diagnostics object. The cache-diagnosis-2026-04-07 beta header is no longer required. Drop it or keep it — either works. The Six Miss-Reason Types The response carries a cache miss reason discriminated union on type . Here is what each one means and what to do about it. - system changed — The most common culprit. A timestamp, request ID, or session token was interpolated into the system prompt. Move dynamic data into the first user message after your cache breakpoint. - tools changed — Tools were added, removed, reordered, or the JSON schema serialized non-deterministically between turns. Send the same tool list in a fixed order with deterministic key sorting on every turn. - model changed — A router, A/B test, or fallback mechanism selected a different model mid-conversation. The cache is per-model. Hold the model constant for the lifetime of a cached conversation. - messages changed — History was truncated or assistant turns re-serialized differently. Treat message history as append-only and echo assistant content blocks back verbatim. - previous message not found — The prior request did not include the diagnostics object, it ran in a different workspace, or the fingerprint expired. Include diagnostics on every turn and keep turns closely spaced. - unavailable — A non-standard parameter changed tool choice , thinking config, context management , or the divergence is beyond the comparison horizon on a very long conversation. Keep prompt-affecting parameters constant across a conversation. Each changed type also carries a cache missed input tokens field — an estimate of how many tokens fell past the divergence point. Use this to quantify the cost impact before prioritizing the fix. How to Add It in Three Minutes The change is two lines per turn. Here is the Python loop pattern: client = anthropic.Anthropic prev id = None for user msg in conversation turns: messages.append {"role": "user", "content": user msg} r = client.beta.messages.create model="claude-opus-5-5", max tokens=1024, cache control={"type": "ephemeral"}, system=SYSTEM, messages=messages, diagnostics={"previous message id": prev id}, if r.diagnostics and r.diagnostics.cache miss reason: print f"Cache miss: {r.diagnostics.cache miss reason.type}" messages.append {"role": "assistant", "content": r.content} prev id = r.id On the first iteration, prev id is None — the API opts in and stores a fingerprint, nothing to compare yet. On every subsequent turn it compares and reports. Threading that prev id forward is the entire integration. Reading the Results The diagnostics field and usage.cache read input tokens answer different questions. Combine them: - diagnostics null + high cache reads — Working correctly. No action needed. - diagnostics null + low cache reads — Your request structure is stable but the cache entry expired. Switch to the 1-hour TTL option https://platform.claude.com/docs/en/build-with-claude/prompt-caching . - changed + low cache reads — Your code changed something it should not have. Fix the cause the type points to. - changed + high cache reads — The divergence happened late in the prompt and an earlier cache breakpoint still hit. Worth fixing, but low urgency. What Does Not Work Yet Cache diagnostics is Claude API only. It is not available on Amazon Bedrock or Google Cloud. Fingerprints expire after a short period, so comparing requests far apart in time returns previous message not found , not a miss reason. And for very long conversations where the divergence is deep in the message history, you will get unavailable rather than a precise location. None of these are dealbreakers. The common failure modes — a timestamp in a system prompt, a reordered tool list, a model swap — are all caught cleanly. Analysts covering the release https://www.digitalapplied.com/blog/prompt-caching-changes-openai-anthropic-september-2026 noted this mirrors what OpenAI already shipped, but Anthropic’s implementation returns the precise divergence point rather than a binary hit/miss. Add It to Every New Agent You Build The opt-in costs nothing. The fingerprint is two lines of code. When your cache breaks in production — and at some point it will — you will know within one request whether the system prompt changed, the tool list shifted, or the history got mangled. That is the difference between a five-minute fix and a three-hour investigation. Cache diagnostics is not a debugging convenience. It is cost infrastructure. Treat it that way.