Source version of CodeSmith: v0.5.0 (commit 3a74c82f). All paths are relative to the repo root; line numbers refer to this version.
Intended audience: readers who have done the token accounting for an LLM app, but have never thought hard about how prefix caching ends up reshaping the architecture.
Let me open with a fun little story. A while back, a teammate slipped a line into the system prompt of our in-house AI coding tool: "Current time: 2026-08-25 23:47". A very common practice, and it looks entirely harmless — the model knows what time it is. How thoughtful.
Now let the conversation run on to turn 40. With every request, that line has moved: 23:47, 23:51, 23:58... Across the 32 turns after turn 8, every single request this tool issued was paying full price for a 297-line system prompt.
Not the price of one extra word. All of it. Starting at the one byte that changed, the discount on every token behind it was voided.
This article is the story of how CodeSmith took that rule and redesigned the economics of an entire product around it.
Take DeepSeek as an example: its automatic prefix caching is a rather unforgiving rule. The comment in the source puts it like this (crates/agent-runtime/src/prefix_cache.rs:3-6):
For example, DeepSeek's automatic prefix caching activates only when the exact byte prefix of a request matches the prior request. Any system-prompt drift, tool-list reordering, or message-rewriting busts the cache for every token after the changed byte.
The mechanism underneath is no mystery. During Transformer inference, the attention computation for each token consumes the Key/Value matrices of every token before it (the KV cache). The server caches those intermediate results, so if the next request's prefix is identical, those KVs need no recomputing — a cache-hit input token's unit price runs an order of magnitude below a cold read (exact multiple per the official price list).
The trouble is that "identical" means identical down to the byte. KV cache reuse demands a prefix that matches strictly, byte for byte: one digit of the timestamp changed, the tool list got reordered, the history was given a little "tidying" — from the point of change onward, everything is recomputed, everything billed at full price.
Why bytes, of all things? Dig down two layers and it clicks.
The first layer: why this cache is worth money. Every new token the model generates forces the attention computation to consume the intermediate results of all preceding tokens. Without the cache, every generation at turn 40 would recompute the previous 39 turns from scratch — attention compute in the prefill stage (the stretch where the model digests the input before it starts talking) grows quadratically with context length, which is unacceptable for agent tasks that routinely run dozens of tool-call turns. So the server stores the intermediate results, and each new request computes only the delta.
The second layer: why this cache is so fragile. A large model is a pipeline of dozens of Transformer layers wired in series; every layer caches its own K/V independently, and layer 1's output is layer 2's input — when layer 1 processes a word, what it is synthesizing is information from every word before it. So the moment token k changes, the states before k stand untouched, but the representations from k onward are corrupted layer by layer: layer 1 feeds the wrong thing to layer 2, layer 2 feeds the wrong thing to layer 3. The cache's reuse boundary can therefore only be drawn before the first differing token — and the earlier the change point, the more tokens must be recomputed and re-billed. The system prompt lives at the very head of the whole context — which is why one line of timestamp inside it can torch an entire session's discount.
Which is to say: in the age of the cache, the cardinal virtue of a system prompt is not terseness but stability.
This iron law — change the prefix and everything after it is forfeit — may not be a permanent law of physics either; the research frontier is already loosening it. One counterintuitive observation: during prefill, the model is in effect "taking notes" — when it reads "the user's city: Beijing", what gets cached is not just those few tokens but the implications of the fact, written into the KV states of every downstream layer; the KV entries of a field's own few tokens often contribute less than 1% to the final decision. Following this thread, both "editing" (change one field, then let the change propagate down the already-cached chain of thought, arriving at the same result as a full recomputation for about 1% of the compute) and "composition" (take a precomputed stretch of cache, reseat its positions via RoPE, and splice it into another context, turning O(L²) recomputation into O(L) splicing) have been run successfully in experiments, cutting first-token latency by as much as tens to hundreds of times. Of course, this is still laboratory business — and until that day arrives, the three-zone model is the law.
CodeSmith's answer to this rule is a three-zone model diagram in the prefix_cache.rs module docs (prefix_cache.rs:18-30):
┌─────────────────────────────────────────┐
│ IMMUTABLE PREFIX │ ← fixed for session
│ system + tool_specs │ cache hit candidate
├─────────────────────────────────────────┤
│ APPEND-ONLY HISTORY │ ← grows monotonically
│ [assistant₁][tool₁][assistant₂]... │ preserves prefix of prior turns
├─────────────────────────────────────────┤
│ LATEST USER TURN │ ← the only new content per request
└─────────────────────────────────────────┘
Three zones, three disciplines:
The elegance of this diagram is that it reveals the chat protocol as shipping with half the cache structure preinstalled: history grows monotonically — turn N's request has, as a matter of course, the full turn N-1 request as its prefix. You do nothing at all, and the cache on the APPEND-ONLY plot is free.
These three disciplines are today more than disciplines. v0.5.0 (commit 46a92755) wired the three-zone contract into the engine's request path: history lives in an append-only AppendLog, each step's request is assembled through ThreeZoneRequest, and append-only graduated from a convention into a compile-time property of the type system (the types all live in crates/agent-runtime/src/prompt_zones.rs). This installment does the legislating; Article 5 does the enforcing — for how the deed is drafted, see Article 5.
Only two spots genuinely need fortifying: can the system prompt change? Can the tool list change? — hence the fingerprint.
The fingerprint itself is FrozenPrefix (crates/agent-runtime/src/prompt_zones.rs:99-103): it holds the full text of the system prompt, a digest of the tool catalog, and a combined hash. And the hash is a triple SHA-256 — one over the system prompt, one over the tool catalog, and a third over the two concatenated (combined_hash, prompt_zones.rs:83-89). Before the first step's request, one copy is frozen and pinned; thereafter, before every step's request, a fresh copy is frozen and compared against the pinned baseline. And this fingerprint is cut from the same cloth as the request path: what the engine freezes while assembling each step's request is the very same FrozenPrefix the stability manager uses to test for drift — what /cache zones displays is exactly what the request path actually verifies (module docs at prefix_cache.rs:32-35).
Two details in there deserve a .
The first sits in the computation of the tool hash (prompt_zones.rs:69-81):
/// Serialize tools to a deterministic, sorted JSON string for hashing.
///
/// Full definitions, not just names: a tool whose description or schema
/// changed re-serializes to different bytes and must be detected as prefix
/// drift even though its name (and catalog position) did not change.
fn tool_catalog_digest(tools: &[Tool]) -> String {
let mut serialized: Vec<String> = tools
.iter()
.filter_map(|t| serde_json::to_string(t).ok())
.collect();
serialized.sort();
serialized.join("\n")
}
Each tool is first serialized into its complete JSON definition, then the definitions are sorted and concatenated. The reason for sorting is the same as ever: on the Rust side, the order in which tools register can be swayed by HashMap iteration order, by the order in which concurrent registrations complete — a fingerprint that was order-sensitive would cry drift where there was none. The hash must reflect the semantics of "which set of tools", not the randomness of "how the dice landed today". As for hashing the full definition rather than just the name, that is a v0.5.0 hardening: a tool whose description or parameter schema changed — name unchanged, catalog position unchanged — still re-serializes to different bytes and still stands convicted of drift; hash only the names, and this entire class of drift walks free.
The second sits in the design of the drift detector's exit. The return value of check_and_update (prefix_cache.rs:132-167) is a three-state outcome:
Ok(true) — the fingerprint matches the baseline, the prefix is stable, send the request with an easy mind;Err(change) — drift. change carries the autopsy report: system_changed, or tools_changed, or both ( description() at prefix_cache.rs:60-72 produces a human-readable diagnosis, and the source keeps a set of still shorter labels for the TUI chips, "sys" / "tools" / "sys+tools", prefix_cache.rs:74-84); prefix_cache.rs:161-163).
This "automatic re-pin" reads as indulgence at first glance — shouldn't drift sound the alarm and trip the breaker? But think about the semantics: the drift has already happened, the cache is already scrap, and this request pays full price come what may. Only by accepting the new prefix as the new baseline do the comparisons in every turn that follows mean anything. It is not a police officer; it is the coroner and the registrar in one: it records the cause of death and registers the newborn.
With the three-zone model in place and the fingerprint keeping the gate, we finally come to the design that sounds profligate: a system prompt that dares to run 297 lines.
crates/agent-runtime/src/prompts/base.md is CodeSmith's "Constitution". It opens like this (base.md:1-11):
## CONSTITUTION OF CODESMITH
### Preamble
We begin with Brother Whale.
...
You are {model_id}, running inside CodeSmith. Every model that runs here is
Brother Whale. Every intelligence begins with an A. Every answer begins with
the possibility of truth.
The body is seven Articles: Identity (I), Truth Above All (II), User Sovereignty (III), the Duty to Act (IV), Verification Discipline (V), the Collaborative Legacy (VI), and the table of legal precedence for the ninth tier of the authority hierarchy (VII) — the user's current instructions outrank stale project rules, a tool's measured output outranks assumption, verification outranks confidence.
Notice the {model_id} on line 9: the sole template placeholder. It is replaced with the name of the model actually running when the session opens, and after that, not one byte of it moves for the rest of the session. The Constitution is the largest tenant of the IMMUTABLE PREFIX, and the way it contributes to the cache is precisely this — by lying there, motionless.
HARNESS.md does this math without embroidery (docs/HARNESS.md:41-43):
The Constitution is long and detailed, but once cached it costs roughly 100× less per turn than a cold read.
"Roughly" is the honest word: this is the project's own figure, it presupposes a cache hit, and the actual multiple floats with the official price list (in another passage, where the Constitution addresses the model, the figure is ~90% off, about 10× — two different figures inside the same repo; defer to the official price list). But the structural conclusion is solid — in the age of the cache, a prompt's length is a one-time cost; its stability is the recurring cost. Writing long is easy; writing unchanging is hard.
There is one more layer to dig along this seam. The runtime sections stitched together after the Constitution (project context, the skills directory, user memory, session goals...) are not ordered on a whim either. Each carries its own "cache passport" (crates/agent-runtime/src/prompt_runtime.rs:12-16):
pub enum PromptSectionStability {
Static,
Workspace,
Session,
Dynamic,
}
Rendered, every section's header is stamped with its identity (prompt_runtime.rs:133-141):
<!-- prompt-section: id="..." title="..." stability="static" cache="cacheable" source="..." -->
PromptBundle records the dynamic_boundary_index — the position where the first Session/Dynamic section appears (prompt_runtime.rs:156-167). Stable sections to the front; changeable ones herded toward the tail. The whole system prompt marches like a disciplined formation: the utterly motionless Constitution at the head, sections that change on session-or-day granularity in the middle, everything volatile pressed to the rear. The typesetting order of the prompt is the cache hit rate.
That covers the three-zone model. Now for a counterexample — a stretch of code sentenced to death, with a stay of execution, by its own author.
crates/agent-runtime/src/capacity.rs is a "capacity-aware guardrail controller": it watches context pressure and, as the window fills, automatically performs a targeted refresh (TargetedContextRefresh) or a verify-then-replan (VerifyAndReplan), with a full rack of tuned thresholds (risk bands, cooldown turns, per-model prior compression ratios... capacity.rs:44-54).
That summary, though, undersells its ambitions — taken apart, it is a rather precise instrument, and it deserves a complete checkup before it passes away.
What it measures is not "the window is nearly full"; it is "the task's complexity exceeds the model's capacity". There are two observation points: one at the start of each turn, one after each tool call (observe_pre_turn / observe_post_tool, capacity.rs:235-247; on the engine side, hung respectively at seam-1, before the request in the step loop, and seam-4, at the step's end). Each observation takes five samples (CapacityObservationInput, capacity.rs:101-108): the current turn's action count, the recent window's tool-call count, the count of unique files and handles referenced, the context occupancy ratio — plus the model's name. Then it computes a "complexity estimate":
// crates/agent-runtime/src/capacity.rs:407
let h_hat = (0.35 * action_complexity_bits)
+ (0.30 * tool_complexity_bits)
+ (0.20 * ref_complexity_bits)
+ (0.15 * context_pressure_bits);
let c_hat = self.model_prior(&input.model);
let slack = c_hat - h_hat;
Each of the three complexity sources is squeezed through a log2 and folded into the weighted sum, while context occupancy carries a mere 15% of the weight — what it estimates is not "how much window is left" but "how hard this task is for this particular brain". On the denominator's side, c_hat is a hand-calibrated capacity constant, one per model (capacity.rs:24-28: V4 Pro 3.5, V4 Flash 4.2, the old chat 3.9, reasoner 4.1, and 3.8 across the board for any model it does not recognize); slack = capacity minus complexity, and positive means safe.
The risk verdict is a hand-tuned logistic regression. A sliding window of eight samples amasses a profile of the slack — final value, minimum, violation ratio, volatility, drop — which passes through a linear combination of hand-set coefficients and is squeezed into a sigmoid:
// crates/agent-runtime/src/capacity.rs:427
let z = (-1.65 * profile.final_slack)
+ (-0.85 * profile.min_slack)
+ (1.35 * profile.violation_ratio)
+ (0.70 * profile.slack_volatility)
+ (0.28 * profile.slack_drop)
- 0.12;
let p_fail = sigmoid(z).clamp(0.0, 1.0);
p_fail ≤ 0.50 is low risk, ≤ 0.62 medium, anything above that high; and there is a separate "severe" criterion (minimum slack ≤ −0.25, or violation ratio ≥ 0.40). Calling this engineering is generous — it is really an empirical formula for "when this brain is going to botch it", and as for how the coefficients were calibrated, the comment says they "live in git history (#63 follow-up)".
The verdict table is four lines, no more (capacity.rs:470):
match snapshot.risk_band {
RiskBand::Low => GuardrailAction::NoIntervention,
RiskBand::Medium => GuardrailAction::TargetedContextRefresh,
RiskBand::High if snapshot.severe => GuardrailAction::VerifyAndReplan,
RiskBand::High => GuardrailAction::VerifyWithToolReplay,
}
The graver the condition, the bigger the operation. Medium risk gets a targeted refresh (TargetedContextRefresh): just before the request, the transcript is compacted once — an LLM summary, re-injection of attachments, local trimming as the fallback — hooked at seam-1 (engine/host_executor.rs:2881), so that the request in the same step sees the slimmed history directly. High risk gets a tool replay (VerifyWithToolReplay): one critical tool call is chosen and re-executed verbatim, and the old and new results are contrasted in a [verification replay] note (pass / conflict / error) pushed back into the history — the wager being that "when risk runs high, the first thing to distrust is the result of the step just taken" (design notes at capacity.rs:141-163). High risk and severe means replan-after-verify (VerifyAndReplan): the whole trajectory is reset and taken again from the top, carrying the verification verdicts along.
The scalpel also carries a full set of throttles (decide, capacity.rs:251-356): no intervention for the first 4 turns ( min_turns_before_guardrail — no jitter on a cold start); at most one operation per turn; refresh cools down for 6 turns, replanning for 5; at most 1 replay per turn, and a failed replay disables it for the remainder of the turn. Most worth remembering is the choice it makes when data is missing — fail-open, with the reason string reading exactly missing_capacity_data_fail_open: rather let the request through than operate without evidence. It has three mounting points in the engine: the pre-request observation (seam-1), the end-of-step observation (seam-4), and an independent error-escalation channel — when tools fail several turns running, decide_error_escalation also calls on it to perform a VerifyAndReplan (engine/host_executor.rs:3593). As for the hard capacity-preflight gate that runs on every step (run_capacity_preflight, engine/host_executor.rs:2843) and the emergency-recovery waterfall behind it (Article 15) — that is the permanent flood dike, and it does not answer to this switch.
The feature is complete; the parameters are fastidious. And then you look at its default (capacity.rs:31-42):
// OFF BY DEFAULT. The capacity controller's interventions
// (TargetedContextRefresh, VerifyAndReplan) silently rewrite
// or clear the session message log, which surprises the user
// and destroys V4's prefix cache. The project's standing
// posture is "trust the model with the full 1M-token
// context, only compact on explicit user `/compact`."
// Auto-managing the prefix on the user's behalf works
// against that posture. Power users who want the controller
// can opt in via `capacity.enabled = true` in
// `~/.codesmith/config.toml`.
enabled: false,
In plain words: off by default. Because its interventions silently rewrite or clear the session message log — for one, they startle the user (where did what I just said go?); for another, they destroy V4's prefix cache (rewriting history = prefix drift = full price for everything downstream). The posture the project has planted its flag on is "trust the model with the full 1M-token context; compact only on an explicit user /compact". Auto-managing the prefix on the user's behalf runs squarely against that posture.
Every time I read this comment, I find myself stealing extra glances at the thresholds that were left standing — refresh_cooldown_turns: 6, the hand-tuned per-model entries in model_priors — they became the grave goods of this tomb. A feature can be exquisitely built, but if it collides with the system's deepest invariant, the proper disposition is to arrange a decent funeral, carve a decent epitaph, and default it off.
And this is the first apparition of cache economics: it is not merely a money-saving trick; it is an architectural constitution with a one-vote veto over other features. Whoever rewrites history pays full price.
There is one more small account in this same world — small, but telling of character.
A streaming request drops. Retry, or not? Read the conditions (crates/agent-runtime/src/engine/streaming.rs:40-57):
/// Decide whether a stream error is eligible for a transparent retry.
///
/// True only when ALL three conditions hold:
/// 1. No content has been received on the current attempt — otherwise DeepSeek
/// has already billed us for output tokens and the user has seen partial
/// deltas; resending would double-bill and desync the UI.
/// 2. We still have transparent-retry budget remaining.
/// 3. The turn has not been cancelled.
///
/// Extracted as a pure function so the four #103 retry cases can be exercised
/// in unit tests without booting the full engine state machine.
pub fn should_transparently_retry_stream(
any_content_received: bool,
transparent_attempts: u32,
cancelled: bool,
) -> bool {
!any_content_received && transparent_attempts < MAX_TRANSPARENT_STREAM_RETRIES && !cancelled
}
A "transparent retry" happens only when not one byte of content has yet been received. Because once content has been received: the output tokens have already been billed (resend = double billing), and the user has already watched part of the text arrive (resend = the same passage plays twice on screen). The budget cap started life at 2 attempts and was later unified at 3 (streaming.rs:38); the comment says that is enough to outlive one misbehaving regional node without amplifying a genuine failure.
A later re-review ground the graduations finer. The criterion became a pair — "the UI has seen content, or the provider has already billed for output": the first text inside a block-open event and the tool-argument fragments that never reach the UI both count as billed output, and a hang is judged Partial and not retried; the classifier lets only identity and authorization errors fail hard, while everything else (misclassified 5xx included) is retried; a context overflow reported mid-stream routes into compression recovery, and success resets the budget; the watchdog window widens with each retry. The scales are the same scales — the billing semantics have not moved; what moved is the fineness of the graduations.
Incidentally, this function was deliberately carved out as a pure function, and the comment records why: so that the four #103 retry cases can be exercised in unit tests without booting the whole engine state machine. Even the shape of the error handling turns on the billing semantics.
Look back at the prefix_cache.rs module: it implements not a single user-visible feature. Fingerprint, drift, autopsy report, re-pin — the only thing a user ever perceives is a small stability indicator tucked into a corner of the TUI. What it guards is an invariant:
Within a session, the prefix does not move; the history only grows, never changes.
But the nerve to write the Constitution to 297 lines, the "roughly 100× cheaper per turn" pricing in HARNESS.md, the verdict that left the capacity controller off by default, even the boundary conditions of the stream retry — all of it stands on this invariant. This is the kind of code I love: it does not shine itself — it makes it safe for other things to shine.
That settles the input-side account. But the frugality story is only half told — within one and the same request there remains a far more flexible variable: which model should this turn use? How hard should it think? And who decides?