CodeSmith Frugal Architecture: The 100x Price of a Single Byte A developer's analysis of CodeSmith v0.5.0 shows how a single drifting byte in an LLM system prompt — such as an embedded timestamp — invalidates automatic prefix caching for every token after the change, forcing full-price recomputation across an entire session. The writeup traces the fragility to per-layer KV cache dependencies in Transformer inference and argues that system prompt stability, not brevity, is now the cardinal architectural virtue. Source version of CodeSmith https://github.com/camilesing/CodeSmith : v0.5.0 commit 3a74c82f . All paths are relative to the repo root; line numbers refer to this version. Intended audience: readers who have done the token accounting for an LLM app, but have never thought hard about how prefix caching ends up reshaping the architecture. Let me open with a fun little story. A while back, a teammate slipped a line into the system prompt of our in-house AI coding tool: "Current time: 2026-08-25 23:47". A very common practice, and it looks entirely harmless — the model knows what time it is. How thoughtful. Now let the conversation run on to turn 40. With every request, that line has moved: 23:47, 23:51, 23:58... Across the 32 turns after turn 8, every single request this tool issued was paying full price for a 297-line system prompt. Not the price of one extra word. All of it. Starting at the one byte that changed, the discount on every token behind it was voided. This article is the story of how CodeSmith took that rule and redesigned the economics of an entire product around it. Take DeepSeek as an example: its automatic prefix caching is a rather unforgiving rule. The comment in the source puts it like this crates/agent-runtime/src/prefix cache.rs:3-6 : For example, DeepSeek's automatic prefix caching activates only when the exact byte prefix of a request matches the prior request. Any system-prompt drift, tool-list reordering, or message-rewriting busts the cache for every token after the changed byte. The mechanism underneath is no mystery. During Transformer inference, the attention computation for each token consumes the Key/Value matrices of every token before it the KV cache . The server caches those intermediate results, so if the next request's prefix is identical, those KVs need no recomputing — a cache-hit input token's unit price runs an order of magnitude below a cold read exact multiple per the official price list . The trouble is that "identical" means identical down to the byte. KV cache reuse demands a prefix that matches strictly, byte for byte: one digit of the timestamp changed, the tool list got reordered, the history was given a little "tidying" — from the point of change onward, everything is recomputed, everything billed at full price. Why bytes, of all things? Dig down two layers and it clicks. The first layer: why this cache is worth money. Every new token the model generates forces the attention computation to consume the intermediate results of all preceding tokens. Without the cache, every generation at turn 40 would recompute the previous 39 turns from scratch — attention compute in the prefill stage the stretch where the model digests the input before it starts talking grows quadratically with context length, which is unacceptable for agent tasks that routinely run dozens of tool-call turns. So the server stores the intermediate results, and each new request computes only the delta. The second layer: why this cache is so fragile. A large model is a pipeline of dozens of Transformer layers wired in series; every layer caches its own K/V independently, and layer 1's output is layer 2's input — when layer 1 processes a word, what it is synthesizing is information from every word before it. So the moment token k changes, the states before k stand untouched, but the representations from k onward are corrupted layer by layer: layer 1 feeds the wrong thing to layer 2, layer 2 feeds the wrong thing to layer 3. The cache's reuse boundary can therefore only be drawn before the first differing token — and the earlier the change point, the more tokens must be recomputed and re-billed. The system prompt lives at the very head of the whole context — which is why one line of timestamp inside it can torch an entire session's discount. Which is to say: in the age of the cache, the cardinal virtue of a system prompt is not terseness but stability . This iron law — change the prefix and everything after it is forfeit — may not be a permanent law of physics either; the research frontier is already loosening it. One counterintuitive observation: during prefill, the model is in effect "taking notes" — when it reads "the user's city: Beijing", what gets cached is not just those few tokens but the implications of the fact, written into the KV states of every downstream layer; the KV entries of a field's own few tokens often contribute less than 1% to the final decision. Following this thread, both "editing" change one field, then let the change propagate down the already-cached chain of thought, arriving at the same result as a full recomputation for about 1% of the compute and "composition" take a precomputed stretch of cache, reseat its positions via RoPE, and splice it into another context, turning O L² recomputation into O L splicing have been run successfully in experiments, cutting first-token latency by as much as tens to hundreds of times. Of course, this is still laboratory business — and until that day arrives, the three-zone model is the law. CodeSmith's answer to this rule is a three-zone model diagram in the prefix cache.rs module docs prefix cache.rs:18-30 : ┌─────────────────────────────────────────┐ │ IMMUTABLE PREFIX │ ← fixed for session │ system + tool specs │ cache hit candidate ├─────────────────────────────────────────┤ │ APPEND-ONLY HISTORY │ ← grows monotonically │ assistant₁ tool₁ assistant₂ ... │ preserves prefix of prior turns ├─────────────────────────────────────────┤ │ LATEST USER TURN │ ← the only new content per request └─────────────────────────────────────────┘ Three zones, three disciplines: The elegance of this diagram is that it reveals the chat protocol as shipping with half the cache structure preinstalled: history grows monotonically — turn N's request has, as a matter of course, the full turn N-1 request as its prefix. You do nothing at all, and the cache on the APPEND-ONLY plot is free. These three disciplines are today more than disciplines. v0.5.0 commit 46a92755 wired the three-zone contract into the engine's request path: history lives in an append-only AppendLog , each step's request is assembled through ThreeZoneRequest , and append-only graduated from a convention into a compile-time property of the type system the types all live in crates/agent-runtime/src/prompt zones.rs . This installment does the legislating; Article 5 does the enforcing — for how the deed is drafted, see Article 5 //./05-frugal-architecture-the-deed-to-the-three-zones.md . Only two spots genuinely need fortifying: can the system prompt change? Can the tool list change? — hence the fingerprint. The fingerprint itself is FrozenPrefix crates/agent-runtime/src/prompt zones.rs:99-103 : it holds the full text of the system prompt, a digest of the tool catalog, and a combined hash. And the hash is a triple SHA-256 — one over the system prompt, one over the tool catalog, and a third over the two concatenated combined hash , prompt zones.rs:83-89 . Before the first step's request, one copy is frozen and pinned; thereafter, before every step's request , a fresh copy is frozen and compared against the pinned baseline. And this fingerprint is cut from the same cloth as the request path: what the engine freezes while assembling each step's request is the very same FrozenPrefix the stability manager uses to test for drift — what /cache zones displays is exactly what the request path actually verifies module docs at prefix cache.rs:32-35 . Two details in there deserve a pause. The first sits in the computation of the tool hash prompt zones.rs:69-81 : /// Serialize tools to a deterministic, sorted JSON string for hashing. /// /// Full definitions, not just names: a tool whose description or schema /// changed re-serializes to different bytes and must be detected as prefix /// drift even though its name and catalog position did not change. fn tool catalog digest tools: & Tool - String { let mut serialized: Vec