{"slug": "can-you-change-an-agents-context-without-losing-its-prompt-cache", "title": "Can You Change an Agent’s Context Without Losing Its Prompt Cache?", "summary": "A code-backed guide to inference state and API cache rules found that in a small Sonnet 5.5 test, changing an instruction near the start of a prompt caused a complete cache miss, while appending the new instruction instead preserved 7,870 cached tokens, with both requests producing the expected answer. The guide explains that prompt caching reuses saved KV cache state from the prefill stage, and that editing earlier text invalidates later positions under causal attention, so agent harnesses should append, branch, or replace history rather than edit prefixes. The distinction matters because a cache hit can cut prefill cost without eliminating the decode work of attending over all cached tokens.", "body_md": "A code-backed guide to inference state, API cache rules, SDK translation, and context management in coding agents.\n\n[Thoughts by human, co-written by AI](https://gkoreli.com/prompt-cache-context-edits/prompts)\n\nI'm building an agent harness, and I want to change its instructions, tools, and retrieved documents without paying to process the whole conversation again. The catch is correctness. A cheap request is no use if the model is still working from an outdated instruction or a document I meant to remove.\n\nMy starting picture was a Jenga tower: change something near the bottom of the prompt and everything above it needs rebuilding. That is mostly right for an edit to earlier text. But an agent has other ways to handle change. It can append a new instruction, return to an earlier conversation branch, or deliberately replace a long history with a shorter one.\n\n  In a small Sonnet 5.5 test, changing an instruction near the start caused a complete cache miss. Appending\n  the new instruction instead preserved **7,870 cached tokens**. Both requests produced the\n  expected answer. Understanding why those requests differ is the key to building a flexible harness without\n  treating every change as a fresh conversation.\n\n## What the model saves\n\n  A prompt first becomes **tokens**: numbers representing pieces of text and the message\n  structure around them. Tokens are often shorter than words. The model turns those numbers into vectors,\n  which are lists of values, then processes them through a stack of layers.\n\n  At an attention layer, it creates three kinds of vectors. A **query** is used to decide which\n  earlier positions to attend to. Each earlier position has a **key** used in that comparison\n  and a **value** carrying information to combine into the result. The names are usually\n  shortened to Q, K, and V. These are learned numerical representations, not a database of facts.\n\n  When generating the next token, the model needs keys and values for the text already processed. Saving\n  them avoids computing them again. That saved state is the **KV cache**. It exists during\n  ordinary generation even when no provider offers a prompt-caching discount. Reusing compatible state in a\n  later request is the additional step called prompt caching. [Attention mechanism](https://arxiv.org/abs/1706.03762), [walkthrough of the serving code](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/08-kv-cache-walkthrough.md).\n\nThere are two stages to keep separate:\n\n- **Prefill:** process the input and create its KV state. A prompt-cache hit can skip much of this work for an unchanged beginning.\n- **Decode:** produce new output tokens. With ordinary full attention, each new token still attends over the earlier keys and values. Keeping 100,000 tokens cached does not make them disappear from the next token's computation or from the context window.\n\nThat helps explain why a cache hit can be cheap without making the whole response fast. The server may still queue the request, move saved state into GPU memory, process new input, and generate a long answer.\n\n## Why one earlier edit affects later text\n\n  A **prefix** is the beginning of the input. A **suffix** is what follows it. For\n  ordinary causal attention, each position can use earlier positions but cannot use future ones. Appending a\n  message therefore leaves the earlier computation valid. Changing an earlier message can change what later\n  positions compute.\n\nImagine an instruction followed by a document:\n\n```\nOriginal:  [Report prices in USD.] [A long document about prices…]\nEdited:    [Report prices in EUR.] [The same document…]\n```\n\nThe document's words are unchanged. Its deeper-layer representations need not be: those layers processed the document with the earlier instruction available. Reusing the old document state after replacing USD with EUR can feed the model numbers from a different input than the one requested. Similar meaning, a small text diff, or an unchanged document hash cannot prove the computation is interchangeable.\n\nThe serving code tracks this dependency. In vLLM, the identifier for a cached block combines its tokens with the previous block's identifier and other relevant inputs. In simplified pseudocode:\n\n```\nblock_id = hash(previous_block_id, token_ids, extra_keys)\n```\n\n  Changing an early block changes the identifiers of later blocks even when their own tokens stay the same.\n  SGLang organizes cached prefixes as a tree, so requests can share the beginning and branch where they\n  differ. [vLLM block hashing](https://github.com/vllm-project/vllm/blob/28f673957671d8d4c5672c2085b3f22c79c0b0b5/vllm/v1/core/kv_cache_utils.py#L650), [SGLang prefix matching](https://github.com/sgl-project/sglang/blob/2ff52e3ceb0bcfdc8ba1dd0367b39ce0370bbfaf/python/sglang/srt/mem_cache/radix_cache.py#L639).\n\nA new branch does not have to destroy the old one. If I later send the original USD request again, the server may still have its cached state. The limit is availability: entries can expire or be evicted to make room for other work.\n\n  A [small numerical example](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/kv_dependency.py) shows the difference. It uses an untrained\n  three-layer transformer and replaces one early token without changing the input length. The table compares\n  its final output scores, called logits, with a fresh calculation:\n\n| Computation | Maximum difference from fresh final logits | \n|---|---|\n| Reuse the valid prefix, recompute the rest | 0 | \n| Recompute the changed token, splice the old suffix state back in | 0.339935 | \n| Append new context to a valid cached prefix | 0 | \n\nReusing the valid beginning gives the same result. Reattaching the old later state gives a different result. This does not measure answer quality: the tiny model has never learned language. It also shows why “every value after an edit changes” is too strong. At an unchanged later token, the first layer’s keys and values stayed equal; deeper layers changed.\n\n  The exact zeros come from running the same arithmetic in the same order. Real serving systems can use\n  different numerical precision or GPU execution orders, so this is not a promise of bit-for-bit identical\n  outputs. [Recorded result and reproduction limits](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/README.md).\n\n## Follow one request through the four layers\n\nSuppose the user asks an agent to inspect a file. The harness chooses the messages and tools, an SDK turns them into an API request, and the provider runs the model. When the model asks to read the file, the harness executes that tool and sends another request with the result attached.\n\nThe second request contains the earlier instructions and conversation, followed by the tool call and result. That repeated beginning is a candidate for reuse. The provider still has to recognize an allowed cache boundary and find the saved state. A conversation in the UI is not proof that either happened.\n\nThe API's JSON is also not the model's literal input. Providers render roles, tool definitions, images, and messages into their own model input. An SDK can move instructions or transform a tool schema before that rendering happens. To explain a miss, inspect what was actually sent, then compare the returned usage counters.\n\nA conversation ID may tell a provider which history to restore. A named cache object may refer to a fixed saved prefix. A response cache may return a previous answer without running the model. None of these names, on its own, means that arbitrary edited text can reuse old KV state.\n\n## Changing the request in six different ways\n\nThe test asked Sonnet 5.5 to return two fields: a currency from the system instruction and a value from record 137. The input contained instructions, reference text, and two blocks of records. Each block ended at an explicit cache boundary: a place where the API was asked to save the preceding input for five minutes. The nine-case sequence ran three times, for 27 requests costing about $0.18. Six of the cases show the useful contrasts:\n\nRead/write counts were the same in all three repetitions of each case. Every answer matched its expected currency and record value. The early edit and appended instruction both changed the expected currency from USD to EUR; the middle edit changed the expected record value.\n\nChanging record 137 preserved the earlier 5,371-token boundary. The entire second record block was processed again, including its unchanged records before 137. The provider had a saved entry at the block boundary; it did not expose reuse at every unchanged token. “How much text is unchanged?” and “where can this API resume?” can have different answers.\n\n  The inline tool request was accepted and kept the prefix hit. It did **not** test whether the\n  model could select or correctly call that tool: the prompt prohibited tool calls. Nor did the test cache\n  the appended control messages for a future turn; all explicit markers were before them.\n\n  Cheap also did not mean faster in this sample. The appended-system case had median time to first visible\n  text of **1.761 seconds**, versus **1.432 seconds** for the early edit. Two\n  appended requests generated extra thinking tokens. Timing includes transport, queueing and reasoning; with\n  three ordered samples per case, it would be misleading to claim a general latency result. Appending the\n  update preserved input reuse and cost less here. The two answer fields passed; broader\n  instruction-following behavior was not tested. [Complete method, timing ranges, and counters](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/05-live-anthropic-experiment.md).\n\n  For budgeting, separate the stable prefix from everything else. If it contains `P` tokens,\n  appears in `N` requests, and is written once then successfully read on every repeat, its cost\n  is `P × (write price + (N − 1) × read price)`, with prices expressed per token. The uncached\n  comparison is `P × N × ordinary input price`. Add new input, output, and any storage charges\n  separately.\n\n  At the tested Sonnet rates, a five-minute write costs 1.25 times ordinary input and a read costs 0.1\n  times. One successful reuse already pays back the write premium: 1.25 + 0.1 is less than two ordinary\n  prefills. A one-hour write at twice ordinary input would need two successful reuses: 2 + 0.1 is still more\n  than two prefills, but 2 + 0.1 + 0.1 is less than three. These are prefix-only calculations assuming hits\n  inside retention, not measured whole-task savings. Keeping irrelevant text merely to improve hit rate can\n  still cost more than shortening the context. [Dated prices and counter rules](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/05-live-anthropic-experiment.md).\n\n## Add an instruction for the next turn\n\n  Anthropic's current API supports system messages within the conversation on selected models, including\n  the tested Sonnet 5.5. A new instruction retains system authority while leaving the preceding prefix\n  intact. Tool additions and removals have their own protocol, including beta inline definitions. This is\n  useful for tools unknown at session start or a schema that changes later. The inline path has constraints,\n  including an initially present non-deferred tool to avoid changing the rendered head when the first new\n  tool appears. [Official update contract](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages).\n\nThe important change is where the new instruction goes. This shortened request shows the shape; the omitted document must be long enough to meet the model's caching threshold:\n\n```\n{\n  \"system\": [{\"type\": \"text\", \"text\": \"Report prices in USD.\"}],\n  \"messages\": [\n    {\"role\": \"user\", \"content\": [{\n      \"type\": \"text\",\n      \"text\": \"The original document…\",\n      \"cache_control\": {\"type\": \"ephemeral\", \"ttl\": \"5m\"}\n    }]},\n    {\"role\": \"system\", \"content\": \"From now on, report prices in EUR.\"}\n  ]\n}\n```\n\nThe USD instruction and document remain where they were. The new system message changes the instruction for what follows. This is a supported API operation on the tested model; putting the same words in an ordinary user message would give them a different role.\n\n  OpenAI also documents controls for later turns: restrict callable tools while keeping their definitions\n  stable, load deferred tools later, or append supported reasoning-configuration updates. Its current\n  caching behavior differs by model generation; older advice about automatic interval caching and retention\n  keys does not fully describe newer explicit boundaries. [OpenAI\n  cache controls](https://developers.openai.com/api/docs/guides/prompt-caching).\n\nThe details differ across APIs:\n\n| Surface | Distinction that changes harness design | \n|---|---|\n| OpenAI | Some models select boundaries automatically; newer models also let callers mark them explicitly. Supported configuration updates can be appended. | \n| Anthropic | Mark blocks yourself or use an automatically advancing boundary. Supported models accept later system and tool updates. | \n| Gemini | Automatic reuse and named cache objects are different features. A named object’s stored content cannot be edited. | \n| DeepSeek | Matching text is reusable only where the service has stored a suitable prefix unit. | \n| OpenRouter | Routes requests to other providers. Staying with the same provider helps locality but does not prove a cache hit. | \n\n  The [dated provider report](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/02-provider-contracts.md) records endpoints, limits, counters,\n  hosted-platform distinctions, and authoritative links. In particular, current Gemini Interactions and\n  GenerateContent expose different caching surfaces. Do not transfer the contract of one endpoint to another\n  because both use the same model family. [Gemini caching](https://ai.google.dev/gemini-api/docs/caching), [explicit resources](https://ai.google.dev/api/caching), [DeepSeek\n  persistence](https://api-docs.deepseek.com/guides/kv_cache/), [OpenRouter routing](https://openrouter.ai/docs/guides/best-practices/prompt-caching).\n\n  For my harness, *apply this policy from now on* is often enough. A supported system message can\n  express that without changing earlier messages. Removing earlier information is a separate operation: an\n  appended correction leaves the old text and its saved representations in place.\n\n## Check what the SDK actually sends\n\nAn API feature is useful only if the SDK sends the right fields.\n\n  The [mocked SDK experiment](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/sdk-wire.mjs) executed published Vercel provider packages and\n  captured their outgoing requests. OpenAI explicit breakpoints and TTL controls survived. An Anthropic\n  mid-conversation system message and tool-removal reference survived too. Those controls reached the\n  outgoing JSON.\n\n  There was a smaller gap: the tested OpenAI provider-options schema accepted `mode` and\n  `ttl`, but had no diagnostics comparison-ID field. Supplying an invented JavaScript option did\n  not forward it. The official OpenAI SDK exposed the underlying field. Adding a property to a JavaScript\n  object does not guarantee that an adapter will forward it. This particular gap affected diagnostics; the\n  tested caching controls worked. [Versioned SDK findings](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/04-openai-sdk-codex.md).\n\nWhen an adapter cannot express a requested update, the caller needs to know what it did instead. Rebuilding the leading prompt may apply the new instruction but change the cost of every following token. A capture of the outgoing request makes that fallback visible.\n\n## How existing harnesses handle changes\n\n  The implementations already contain useful ideas to borrow. They also show why a single ```\ncache:\n  true\n```\n option cannot describe every operation.\n\n| Harness | Finding at the audited revision | \n|---|---|\n| Pi | Preserves native system updates when supported; otherwise collapses them into the leading prompt. Tool redefinitions trigger a separate fallback. | \n| Oh My Pi | Implements explicit OpenAI cache anchors, Anthropic tool-control history, stable MCP ordering, and cache-warming policy. | \n| Codex CLI | Can pin a context window's request-level reasoning effort and append trusted configuration updates on supported models. | \n| OpenCode | Applies provider-specific cache markers through AI SDK, with distinct automatic-caching and runtime paths. | \n| Gemini CLI | Appends selected retry nudges rather than rewriting system instructions; authentication routes differ. | \n| Claude Code | Official documentation describes stable prefixes, appended updates, cache settings, and compaction; this was not a source audit of its production request builder. | \n\n  The [harness audit](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/03-harness-audit.md), [Codex audit](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/04-openai-sdk-codex.md), and [Claude Code\n  documentation](https://code.claude.com/docs/en/prompt-caching) provide the versions and evidence boundaries. Default-branch code is not proof of the\n  behavior of an older installed release.\n\n  Pi supplied the clearest local reproduction. Using its actual pinned transcript functions, the same\n  system-section change kept the original prefix when native mid-conversation support was enabled. With that\n  capability disabled or unspecified, it replaced the section in the leading system message. Both\n  representations replayed to the same current system text in Pi's helper, but that does not make them\n  identical model inputs. [Fixture and returned transcripts](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/pi-transcript-results.json).\n\nA harness can expose that result directly: the update was appended, earlier history was rewritten, or the provider does not support the requested operation. That lets callers choose whether the fallback is acceptable before sending the request.\n\n## Replacing a document can be cheaper than appending a correction\n\nInstructions are only half the problem. An agent also needs to replace evidence: a file changes, a search result is outdated, or a tool returns a newer record. Keeping the old version can preserve cache hits while leaving more work for the model.\n\nA second experiment used an inventory record. Version 1 said 6 items at $17 each, for a total of $102. The model answered from that record. Version 2 then changed the values to 9 items at $23, for a total of $207. The next request either replaced the old source and dropped its answer, or retained both and appended an explicit correction.\n\n| Next request | Cache reads | Cache writes | Full request cost | \n|---|---|---|---|\n| Replace the source and drop the old answer | 3,510 | 76 | $0.0016640 | \n| Keep the old source and answer; append the correction | 3,586 | 290 | $0.0022142 | \n\n  The appended correction read 76 more cached tokens but cost about **33% more** in this\n  fixture. A cache boundary immediately before the small source let the replacement keep the long background\n  prefix. Only 76 tokens needed rewriting. Appending instead retained the previous question and answer and\n  added the correction, producing a larger new suffix.\n\n  Both layouts returned the correct current price, quantity, total, and source identifier in both\n  repetitions. The instructions explicitly told the model to use the highest source revision and recompute\n  from it. There was no answer failure here, and no test of ambiguous source precedence. A larger document,\n  an earlier edit, or a longer remaining conversation could change the cost comparison. [Requests, results, and all eight calls](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/09-changing-retrieved-evidence.md).\n\nA third branch replaced the source but kept the actual old assistant answer. The model corrected the answer and identified it as stale. That branch was already warm from the earlier replacement, so its cache hit is evidence of returning to a computed branch—not of preserving the old source's state through an edit.\n\nThe important design issue is that a source and the work based on it are separate objects. Updating a file does not update an earlier summary, plan, or proposed patch. Removing a source does not remove facts copied into those objects. If the operation really means “do not include this old value in the next request,” the harness has to check those copies too.\n\n## What pruning and compaction actually change\n\n**Truncation** shortens an item before or while it enters the context.\n  **Pruning** removes selected older items. **Compaction** builds a smaller\n  continuation, often from a summary plus recent messages. They can all reduce token count, but they do\n  different things to the next request.\n\n```\nBefore:\n  instructions → long build log → assistant diagnosis → recent work\n\nAfter pruning:\n  instructions → [old log cleared] → assistant diagnosis → recent work\n\nAfter summarizing:\n  instructions → summary of the earlier work → recent work\n```\n\nIn both rewritten histories, the text of “recent work” may be identical. Its preceding context is different, so unchanged recent messages do not make the old later KV state reusable. The earlier stable instructions may still hit.\n\n  OpenCode makes the distinction between stored history and active context easy to see. Its pruning pass\n  stamps an old tool result as compacted. Later, the request converter replaces that result with ```\n[Old\n  tool result content cleared]\n```\n and drops its attachments. The original output string remains in the\n  stored record on that path. [Pruning](https://github.com/anomalyco/opencode/blob/7945de208964a49300d7f770d1a71d078db9a4c4/packages/opencode/src/session/compaction.ts#L269), [request conversion](https://github.com/anomalyco/opencode/blob/7945de208964a49300d7f770d1a71d078db9a4c4/packages/opencode/src/session/message-v2.ts#L294).\n\n  Oh My Pi also considers how much already-sent conversation follows a pruning candidate. Deleting a small\n  old result can disturb a much larger cached suffix. Its pruning configuration can protect such a result\n  rather than treating every removed token as an immediate saving. When it does change a message, it\n  invalidates local cached token estimates and message conversions too. Provider caching is only one cache\n  that must stay consistent. [OMP pruning and local invalidation](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/07-context-management.md).\n\n  The runnable local examples execute Pi's session-context projection and OMP's pruning functions.\n  Pi drops the old source from active context while retaining a supplied summary that contains its\n  conclusion. OMP removes old tool text while retaining an assistant statement based on it. Neither function\n  has erased the information. That is usually the purpose of summarization, but it matters when the original\n  evidence was wrong or must be removed. [Upstream functions and returned messages](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/context-harness.mjs), [recorded output](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/lab/context-results.json).\n\n  Compaction also has a cost of its own. It may pay for a summary and a new cache write now to reduce later\n  reads. Under illustrative prices of 1.25 units per written token and 0.1 per cached-read token, retaining\n  100,000 tokens costs 10,000 units each turn. Replacing them with 20,000 tokens costs 25,000 units once,\n  then 2,000 per later turn. Over three turns those input costs are 30,000 versus 29,000 units, before\n  paying to generate the summary. A summary that loses the key error can erase those savings by causing\n  another investigation. [Calculation and code walkthrough](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/07-context-management.md).\n\n## Can the engine repair and reuse later chunks?\n\nThere are research systems that assemble separately cached chunks or repair only some of the state after a change. Their tradeoff is how closely the reused calculation matches processing the complete new input.\n\n**Prompt Cache** uses schema-defined modules and positions. Its paper explicitly describes an\n  attention approximation: independent modules omit some dependencies that ordinary concatenated attention\n  would have computed. **CacheBlend** instead reuses chunks and selectively recomputes tokens\n  to reduce deviation from full recomputation. **PIE**, designed for code edits, repairs\n  positional effects while treating suffix reuse as an approximation. [Prompt Cache](https://arxiv.org/html/2311.04934v2), [CacheBlend](https://arxiv.org/html/2405.16444v3), [PIE](https://proceedings.iclr.cc/paper_files/paper/2025/file/9530635032b95cea9585bd800d308300-Paper-Conference.pdf).\n\nA system can answer a test set just as well while producing different token probabilities. That is weaker than preserving the calculation a fresh request would perform. A position correction is not a recomputation of what the token learned from the old context. Moving or compressing a cache also does not establish that its contents remain valid after an edit.\n\n  There are architecture-specific exceptions. A deliberately independent attention mask changes which\n  dependencies exist. Pure local attention can bound their reach, while hybrid or recurrent models require\n  different checkpoint reasoning. None supplies a universal hosted-API operation for editing arbitrary old\n  text while retaining the entire suffix unchanged. [Inference report, implementation details, and research limits](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/01-inference-and-cache-repair.md).\n\nApproximate repair may be useful when a workload can measure and accept its errors. The papers do not establish a general way to replace any earlier text while preserving exactly the computation that a fresh request would perform.\n\n## What I would build into a harness\n\nI would make the requested change explicit before choosing how to preserve the cache:\n\n| Requested operation | Candidate implementation | Required correctness check | \n|---|---|---|\n| Change a preference from this turn onward | Supported authoritative message appended to history | New behavior applies with the intended scope | \n| Add or change a tool | Native tool update where supported; otherwise deliberate rebuild | Correct schema/version and executor permissions | \n| Replace stale evidence | New version with explicit precedence, or rebuild the relevant context | The answer uses the intended evidence; old content may still influence appended form | \n| Erase or redact earlier content | Rebuild without it and its derived context; address provider retention separately | No claim that an appended correction removes prior information | \n| Compact a long session | Summarize and begin a new context version | Task-relevant facts survive; lower hit rate may still reduce total cost | \n\nStable instructions, deterministic tool ordering, versioned context, and unchanged history avoid accidental misses. Supported prospective controls provide flexibility. Meaningful rewrites should remain visible, even when they cost more.\n\n  To diagnose a miss, record the model and serving route, the adapter version, the selected cache\n  boundaries, and a sanitized comparison of consecutive requests. Keep the provider's original usage\n  counters alongside any normalized totals: some APIs include cached tokens in total input and others report\n  separate buckets. Diagnostics can help identify changed instructions or tools, but an eligible request can\n  still miss because its saved state is unavailable. [OpenAI diagnostics](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics), [Anthropic diagnostics](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics).\n\nFor the harness I'm building, the goal is to preserve useful work while making real changes correctly. Appended instructions and old conversation branches can reuse saved work. Replacing evidence or compacting a session may require a new prefix. That cost can be worth paying when it gives the model a shorter, more accurate context.\n\n## Glossary\n\n| Term / claim | Primary source | Date | \n|---|---|---|\n| Causal attention and state dependencies | [Attention Is All You Need](https://arxiv.org/abs/1706.03762) | 2017 | \n| Prefix hash provenance | [Pinned vLLM implementation](https://github.com/vllm-project/vllm/blob/28f673957671d8d4c5672c2085b3f22c79c0b0b5/vllm/v1/core/kv_cache_utils.py#L650) | Inspected 2026-09-28 | \n| Mid-conversation system/tool changes | [Anthropic contract](https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages) | Checked 2026-09-28 | \n| Provider-specific caching controls | [OpenAI](https://developers.openai.com/api/docs/guides/prompt-caching) ,[Anthropic](https://platform.claude.com/docs/en/build-with-claude/prompt-caching) ,[Gemini](https://ai.google.dev/gemini-api/docs/caching) | Checked 2026-09-28 | \n| Modular and selectively repaired reuse | [Prompt Cache](https://arxiv.org/abs/2311.04934) ,[CacheBlend](https://arxiv.org/abs/2405.16444) | 2023 / 2024 preprints; versions reviewed in research | \n| Source audits, reproduction and measurements | [Research and runnable labs](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/README.md) | 2026-09-28 local / 2026-09-29 UTC | \n\n[Research reading guide](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/README.md)\n\n[Source manifest](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/sources.lock.json)\n\n[Claim ledger](https://github.com/gkoreli/blog/blob/main/docs/worklist/prompt-caching-from-api-to-agent-loop/research/06-claim-ledger.md)", "url": "https://wpnews.pro/news/can-you-change-an-agents-context-without-losing-its-prompt-cache", "canonical_source": "https://gkoreli.com/prompt-cache-context-edits", "published_at": "2026-09-29 00:00:00+00:00", "updated_at": "2026-09-29 07:48:58.107503+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-infrastructure", "mlops", "developer-tools"], "entities": ["Sonnet 5.5", "KV cache", "Q", "K", "V"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/can-you-change-an-agents-context-without-losing-its-prompt-cache", "markdown": "https://wpnews.pro/news/can-you-change-an-agents-context-without-losing-its-prompt-cache.md", "text": "https://wpnews.pro/news/can-you-change-an-agents-context-without-losing-its-prompt-cache.txt", "jsonld": "https://wpnews.pro/news/can-you-change-an-agents-context-without-losing-its-prompt-cache.jsonld"}}