cd /news/ai-tools/claude-cache-diagnostics-is-ga-debug… · home › topics › ai-tools › article
[ARTICLE · art-140680] src=byteiota.com ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

Claude Cache Diagnostics Is GA: Debug Prompt Cache Misses Now

Anthropic made cache diagnostics generally available on the Claude API as of September 23, letting developers pinpoint why prompt caches miss by passing a `diagnostics` object with a `previous_message_id` on each Messages request. The feature returns a `cache_miss_reason` union covering six causes — system_changed, tools_changed, model_changed, messages_changed, previous_message_not_found, and unavailable — plus a `cache_missed_input_tokens` estimate, and no longer requires the `cache-diagnosis-2026-04-07` beta header. Cache reads on Opus 5.5 cost $0.20 per million tokens versus a $4 base input price, a 95% reduction, and Anthropic says fingerprints store only cryptographic hashes and token-count estimates, are ZDR eligible, and are scoped to the organization and workspace.

read4 min views2 publishedSep 28, 2026
Claude Cache Diagnostics Is GA: Debug Prompt Cache Misses Now
Image: Byteiota (auto-discovered)

Prompt caching is the single most effective cost lever for Claude API users. Cache reads on Opus 5.5 cost $0.20 per million tokens against a $4 base input price — a 95% reduction. Production teams running agents on this have documented drops from $720 to $72 a month. The problem has always been that when caching stopped working, the only signal was cache_read_input_tokens going to zero. No explanation. No pointer. Just silent, expensive failure. As of September 23, that ends. Cache diagnostics is now GA on the Claude API.

What Cache Diagnostics Does #

The API now compares consecutive requests for you. Opt in by including a diagnostics object on your Messages request. On the first turn, pass previous_message_id: null to register the fingerprint. On every subsequent turn, pass the id from the previous response. When the cache misses, the response tells you exactly where the prefix diverged: the model, the system prompt, the tools, or the message history.

Fingerprints contain only cryptographic hashes and token-count estimates — no raw prompt content. The feature is ZDR eligible, scoped to your organization and workspace, and stored only for requests that include the diagnostics object.

The cache-diagnosis-2026-04-07 beta header is no longer required. Drop it or keep it — either works.

The Six Miss-Reason Types #

The response carries a cache_miss_reason discriminated union on type. Here is what each one means and what to do about it.

  • system_changed — The most common culprit. A timestamp, request ID, or session token was interpolated into the system prompt. Move dynamic data into the first user message after your cache breakpoint.
  • tools_changed — Tools were added, removed, reordered, or the JSON schema serialized non-deterministically between turns. Send the same tool list in a fixed order with deterministic key sorting on every turn.
  • model_changed — A router, A/B test, or fallback mechanism selected a different model mid-conversation. The cache is per-model. Hold the model constant for the lifetime of a cached conversation.
  • messages_changed — History was truncated or assistant turns re-serialized differently. Treat message history as append-only and echo assistantcontent blocks back verbatim.
  • previous_message_not_found — The prior request did not include thediagnostics object, it ran in a different workspace, or the fingerprint expired. Includediagnostics on every turn and keep turns closely spaced.
  • unavailable — A non-standard parameter changed (tool_choice , thinking config, context management), or the divergence is beyond the comparison horizon on a very long conversation. Keep prompt-affecting parameters constant across a conversation.

Each *_changed type also carries a cache_missed_input_tokens field — an estimate of how many tokens fell past the divergence point. Use this to quantify the cost impact before prioritizing the fix.

How to Add It in Three Minutes #

The change is two lines per turn. Here is the Python loop pattern:

client = anthropic.Anthropic()
prev_id = None

for user_msg in conversation_turns:
    messages.append({"role": "user", "content": user_msg})

    r = client.beta.messages.create(
        model="claude-opus-5-5",
        max_tokens=1024,
        cache_control={"type": "ephemeral"},
        system=SYSTEM,
        messages=messages,
        diagnostics={"previous_message_id": prev_id},
    )

    if r.diagnostics and r.diagnostics.cache_miss_reason:
        print(f"Cache miss: {r.diagnostics.cache_miss_reason.type}")

    messages.append({"role": "assistant", "content": r.content})
    prev_id = r.id

On the first iteration, prev_id is None — the API opts in and stores a fingerprint, nothing to compare yet. On every subsequent turn it compares and reports. Threading that prev_id forward is the entire integration.

Reading the Results #

The diagnostics field and usage.cache_read_input_tokens answer different questions. Combine them:

  • diagnostics null + high cache reads — Working correctly. No action needed.
  • diagnostics null + low cache reads — Your request structure is stable but the cache entry expired. Switch to the1-hour TTL option .
  • *_changed + low cache reads — Your code changed something it should not have. Fix the cause the type points to.
  • *_changed + high cache reads — The divergence happened late in the prompt and an earlier cache breakpoint still hit. Worth fixing, but low urgency.

What Does Not Work Yet #

Cache diagnostics is Claude API only. It is not available on Amazon Bedrock or Google Cloud. Fingerprints expire after a short period, so comparing requests far apart in time returns previous_message_not_found, not a miss reason. And for very long conversations where the divergence is deep in the message history, you will get unavailable rather than a precise location.

None of these are dealbreakers. The common failure modes — a timestamp in a system prompt, a reordered tool list, a model swap — are all caught cleanly. Analysts covering the release noted this mirrors what OpenAI already shipped, but Anthropic’s implementation returns the precise divergence point rather than a binary hit/miss.

Add It to Every New Agent You Build #

The opt-in costs nothing. The fingerprint is two lines of code. When your cache breaks in production — and at some point it will — you will know within one request whether the system prompt changed, the tool list shifted, or the history got mangled. That is the difference between a five-minute fix and a three-hour investigation.

Cache diagnostics is not a debugging convenience. It is cost infrastructure. Treat it that way.

── more in #ai-tools 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-cache-diagnos…] indexed:0 read:4min 2026-09-28 · —