# Claude Cache Diagnostics Is GA: Debug Prompt Cache Misses Now

> Source: <https://byteiota.com/claude-cache-diagnostics-is-ga-debug-prompt-cache-misses-now/>
> Published: 2026-09-28 01:08:39+00:00

Prompt caching is the single most effective cost lever for Claude API users. Cache reads on Opus 5.5 cost $0.20 per million tokens against a $4 base input price — a 95% reduction. Production teams running agents on this have [documented drops from $720 to $72 a month](https://usagebox.com/articles/prompt-caching-cost-optimization-claude-gpt-gemini-2026). The problem has always been that when caching stopped working, the only signal was `cache_read_input_tokens` going to zero. No explanation. No pointer. Just silent, expensive failure. As of September 23, that ends. [Cache diagnostics is now GA on the Claude API](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics).

## What Cache Diagnostics Does

The API now compares consecutive requests for you. Opt in by including a `diagnostics` object on your Messages request. On the first turn, pass `previous_message_id: null` to register the fingerprint. On every subsequent turn, pass the `id` from the previous response. When the cache misses, the response tells you exactly where the prefix diverged: the model, the system prompt, the tools, or the message history.

Fingerprints contain only cryptographic hashes and token-count estimates — no raw prompt content. The feature is ZDR eligible, scoped to your organization and workspace, and stored only for requests that include the `diagnostics` object.

The `cache-diagnosis-2026-04-07` beta header is no longer required. Drop it or keep it — either works.

## The Six Miss-Reason Types

The response carries a `cache_miss_reason` discriminated union on `type`. Here is what each one means and what to do about it.

- **system_changed** — The most common culprit. A timestamp, request ID, or session token was interpolated into the system prompt. Move dynamic data into the first user message after your cache breakpoint.
- **tools_changed** — Tools were added, removed, reordered, or the JSON schema serialized non-deterministically between turns. Send the same tool list in a fixed order with deterministic key sorting on every turn.
- **model_changed** — A router, A/B test, or fallback mechanism selected a different model mid-conversation. The cache is per-model. Hold the model constant for the lifetime of a cached conversation.
- **messages_changed** — History was truncated or assistant turns re-serialized differently. Treat message history as append-only and echo assistant`content` blocks back verbatim.
- **previous_message_not_found** — The prior request did not include the`diagnostics` object, it ran in a different workspace, or the fingerprint expired. Include`diagnostics` on every turn and keep turns closely spaced.
- **unavailable** — A non-standard parameter changed (`tool_choice` , thinking config, context management), or the divergence is beyond the comparison horizon on a very long conversation. Keep prompt-affecting parameters constant across a conversation.

Each `*_changed` type also carries a `cache_missed_input_tokens` field — an estimate of how many tokens fell past the divergence point. Use this to quantify the cost impact before prioritizing the fix.

## How to Add It in Three Minutes

The change is two lines per turn. Here is the Python loop pattern:

```
client = anthropic.Anthropic()
prev_id = None

for user_msg in conversation_turns:
    messages.append({"role": "user", "content": user_msg})

    r = client.beta.messages.create(
        model="claude-opus-5-5",
        max_tokens=1024,
        cache_control={"type": "ephemeral"},
        system=SYSTEM,
        messages=messages,
        diagnostics={"previous_message_id": prev_id},
    )

    if r.diagnostics and r.diagnostics.cache_miss_reason:
        print(f"Cache miss: {r.diagnostics.cache_miss_reason.type}")

    messages.append({"role": "assistant", "content": r.content})
    prev_id = r.id
```

On the first iteration, `prev_id` is `None` — the API opts in and stores a fingerprint, nothing to compare yet. On every subsequent turn it compares and reports. Threading that `prev_id` forward is the entire integration.

## Reading the Results

The `diagnostics` field and `usage.cache_read_input_tokens` answer different questions. Combine them:

- **diagnostics null + high cache reads** — Working correctly. No action needed.
- **diagnostics null + low cache reads** — Your request structure is stable but the cache entry expired. Switch to the[1-hour TTL option](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) .
- **`*_changed` + low cache reads** — Your code changed something it should not have. Fix the cause the type points to.
- **`*_changed` + high cache reads** — The divergence happened late in the prompt and an earlier cache breakpoint still hit. Worth fixing, but low urgency.

## What Does Not Work Yet

Cache diagnostics is Claude API only. It is not available on Amazon Bedrock or Google Cloud. Fingerprints expire after a short period, so comparing requests far apart in time returns `previous_message_not_found`, not a miss reason. And for very long conversations where the divergence is deep in the message history, you will get `unavailable` rather than a precise location.

None of these are dealbreakers. The common failure modes — a timestamp in a system prompt, a reordered tool list, a model swap — are all caught cleanly. [Analysts covering the release](https://www.digitalapplied.com/blog/prompt-caching-changes-openai-anthropic-september-2026) noted this mirrors what OpenAI already shipped, but Anthropic’s implementation returns the precise divergence point rather than a binary hit/miss.

## Add It to Every New Agent You Build

The opt-in costs nothing. The fingerprint is two lines of code. When your cache breaks in production — and at some point it will — you will know within one request whether the system prompt changed, the tool list shifted, or the history got mangled. That is the difference between a five-minute fix and a three-hour investigation.

Cache diagnostics is not a debugging convenience. It is cost infrastructure. Treat it that way.
