Tail Rent: What Actually Happens to Your Agent’s Memory After a Week of Continuous Runtime A four-day persistent coding-agent session lost a specific architectural decision — a 15-minute refresh-token TTL and its mobile-client rationale — after the third compaction event, according to a first-person account of running the agent against a multi-service migration backlog. The account attributes the loss to "tail rent," the compounding cost of paying cached-prefix rates on stable history and expensive write rates on the growing uncached tail every turn, and notes that Claude Code's /compact triggers automatically at around 95 percent of context capacity while OpenAI's Codex CLI uses token thresholds generally in the 180k-244k range with a 95 percent safety margin. Anthropic's prompt caching holds a cached prefix for five minutes by default, or an hour with the longer TTL paid for, the account states. I had a coding agent running against a real backlog for the better part of four days last month. Not one long unbroken session, but a persistent one: the same conversation, resumed every morning, picking up a multi-service migration ticket by ticket. On day three I asked it to touch a retry policy we’d deliberately built around a decision from day one, a decision about exactly how long a refresh token should live and why. It gave me an answer that was directionally fine and factually wrong. It had the shape of the decision right and the specifics gone. No TTL number, no reason, no mention of the mobile client constraint that had driven the whole design in the first place. I went looking for where that detail died, and found it sitting a few messages after a compaction event in the session log. Not the first compaction, either. The third. That’s the failure mode this piece is about, and it’s not really a bug. It’s the predictable output of how a fixed context window has to behave once a session runs long enough, combined with an economic pressure most people don’t talk about because it hides inside a billing dashboard instead of showing up as an error. I’m going to call that pressure “tail rent,” walk through why folding a summary of a summary makes it worse, lay out a five-step ladder of things to try before you reach for a destructive fold at all, and then show you a C harness I actually built and ran that demonstrates the failure and one cheap fix for it, with real numbers from a real run, not a hypothetical. Every message you send to a model-backed agent gets bundled with everything that came before it: system prompt, tool definitions, the full conversation history, every tool call and its result. Providers cache the parts of that bundle that don’t change turn to turn the system prompt, the tool schema, older parts of the transcript that have gone stable so you’re not paying full input price for the same tokens over and over. Anthropic’s prompt caching, for instance, holds a cached prefix for five minutes by default, or an hour if you pay for the longer TTL. That cached chunk is what I mean by the prefix. Everything after the cache boundary, the newest tool result, this turn’s user message, the assistant’s in-progress response, is the tail. It’s new every turn by definition, so it can’t be cached yet, and it gets more expensive to keep resending as it grows, because it isn’t just this turn’s tokens, it’s this turn’s tokens plus however much of the previous tail hasn’t rolled into the cached prefix yet. Tail rent is the compounding cost of that arrangement: you pay the cheap cached-prefix rate on the stable part of history, and the expensive write rate on the growing tail, on every single turn, for as long as the session lives. A session that runs for an hour pays this a few dozen times. A session that runs for days, resumed every morning, pays it hundreds of times, and every resume after a cache TTL has expired means the whole prefix goes stale and gets rewritten at the expensive rate once before it can cache again. At some point the tail gets big enough that it threatens to blow the context window outright, and this is where compaction comes in. Compaction takes some span of the transcript and replaces it with a shorter summary, which is a real and sometimes necessary move, but it is also the only rung on this ladder that destroys information on purpose. Claude Code’s /compact does this automatically at around 95 percent of context capacity if you don't trigger it manually first. OpenAI's Codex CLI does something similar with token-based thresholds that vary by model, generally somewhere in the 180k-244k token range with a 95 percent safety margin applied on top. Here’s the part that actually explains what happened to my TTL decision. A single compaction is lossy but survivable, usually: a decent summarizer keeps “we chose to store refresh tokens in Redis with a 15-minute TTL and rotate on every use because the mobile client can’t reliably persist tokens between restarts” down to something like “refresh tokens: Redis, 15-min TTL, rotate on use.” Compressed, but the facts are all still there. The problem is what happens on the second compaction, when that already-compressed sentence is itself sitting somewhere in the middle of a transcript that’s grown too long again. It gets folded a second time, and a summary of a summary tends to keep the shape of a decision while shedding its specifics, because “why” and exact values are exactly the kind of detail a second-pass summarizer treats as safe to cut when it’s optimizing for brevity over fidelity to the first summary rather than the original event. By the third fold, in my case, “Redis, 15-minute TTL, rotate on use, because mobile can’t persist tokens” had degraded to something like “auth token handling was reviewed and updated,” which is true and useless. Nothing was ever deleted in one dramatic step. It was sanded down, one fold at a time, until the load-bearing part was gone. I want to flag something honest here: this specific behavior, compounding loss across repeated folds, is not something every harness necessarily does the same way, and I haven’t independently verified every claim other writers make about specific providers’ internals. What I can tell you with confidence is the shape of the failure, because I reproduced it deterministically in the harness below, and the mechanism summarize a summary, lose the specifics that made it a decision rather than a status update is structural, not vendor-specific. Before reaching for a destructive fold, there’s a sequence of cheaper, less lossy moves worth exhausting first. I’ve ordered these from least destructive to most, which is also, not coincidentally, roughly the order of implementation effort. Rung 1: Truncate at the source. Cap tool output before it ever enters the transcript. A file search that returns 200 matches when the agent needed the top 5 shouldn’t get to spend context on the other 195, ever. Anthropic’s text editor tool supports a max characters parameter for exactly this. Codex caps tool outputs somewhere in the 10k-16k token range by default. This is the cheapest possible fix and it's also the most limited one, because it only helps with content you can safely discard entirely . It does nothing for content you need available later, just not right now. Rung 2: Spill to disk, leave a pointer. For content you might need later but don’t need in the model’s working set right now, write it to a file and leave a short reference a path, an ID, a one-line description in the transcript instead of the raw payload. Claude Code and similar harnesses do this once a tool result crosses roughly the 100k-token mark. The tradeoff: retrieving the spilled content later costs a tool call and re-reads it into context, so this only pays off when you’re right that the agent probably won’t need it again soon. Rung 3: Delegate to a subagent. If a chunk of work is genuinely exploratory grep across a big codebase, page through a long API response, try three different approaches to a bug hand it to a subagent with its own separate context window. Only the subagent’s final answer enters the parent’s transcript; the exploration that produced it, including every wrong turn, never touches parent memory at all. This is strictly better than compaction for this specific shape of work, because the subagent’s context doesn’t need summarizing, it just gets discarded once its answer is extracted. The catch is that spinning up a subagent has its own overhead a fresh system prompt and tool schema, no shared cache with the parent and it only helps for work you can cleanly scope in advance. You can’t delegate a decision you haven’t identified as delegate-able yet. Rung 4: Pointers and references, generally. This is rung 2 generalized past tool output: instead of carrying an artifact’s full content in the live conversation, store the artifact a file, a memory-tool entry, a database row and carry only its location and a one-line description in context. The distinction between this and compaction is important: a pointer is prospective, you write it knowing exactly what a future reader will need to reconstruct the artifact. A compacted summary is retrospective, it’s a guess about what a future, unknown request might need, made under time pressure by a process that doesn’t know what’s coming. Rung 5: Destructive compaction, last resort. Fold a span of the transcript into a shorter summary and accept some information loss. Sometimes this is genuinely unavoidable, a truly stable prefix and thoughtful pointers can only stretch a context window so far before you have to pay down the tail some other way. When you do reach for it, the fix that actually protects against the fold-of-a-fold problem isn’t a smarter summarizer, it’s making sure the things that matter never enter the foldable pool in the first place, more on that below. THE FIVE-RUNG LADDERRung Mitigation Information loss risk Implementation effort---- ------------------------ ----------------------- ----------------------1 Truncate at the source Low content is either Low truly disposable or it isn't; no partial loss 2 Spill to disk + pointer Low-Medium safe unless Low-Medium you guess wrong about what's needed again 3 Delegate to a subagent Low only the summary Medium needs a task the subagent chooses to boundary you can report back is at risk scope up front 4 Pointer / reference Low prospective: you Medium-High needs a pattern control what's kept place to store artifacts + discipline about when to use it 5 Destructive summarizing High, and compounds on Low to build, high to fold repeated folds get right long-term Description is cheap and I didn’t trust my own explanation of “fold of a fold” until I’d made it happen on purpose. So I built a small C harness: a wrapper around a chat-completion loop that tracks an estimated running token count, spills any tool result over a configurable size threshold to a local file with a GUID-based pointer left in the transcript, and logs every compaction event with a timestamp, the token count before and after, and which of a small watchlist of keywords disappeared at that specific fold. I’m using a rough token estimator characters divided by four rather than a real tokenizer, since the point of this harness is to demonstrate the mechanics of the ladder, not to be production-accurate about token counts. If you’re building this for real, swap in the actual tokenizer for your model family, the ratio drifts substantially on code and JSON, which is exactly the kind of content that blows up tool-result tails in a coding agent in the first place. // TailRent.cs//// A runnable proof-of-concept for the mitigation ladder above: prevention/truncation// - spill-to-disk - subagent delegation - pointers - destructive fold.//// This file demonstrates rungs 1, 2 and 5 end to end truncation, spilling, and// folding , with real token accounting and a real, inspectable compaction log.// Rungs 3 and 4 subagent delegation and pointer patterns are architectural// decisions about what enters the transcript in the first place rather than// something a single-file wrapper can demo in isolation -- the spiller below is// effectively rung 4 already, since the tail keeps a pointer, not the content.//// Build no external NuGet packages required, everything is BCL :// dotnet build// Run:// dotnet run -- mock offline, deterministic, no network calls // dotnet run -- ollama talks to a local Ollama server, see OllamaChatClient below using System;using System.Collections.Generic;using System.IO;using System.Linq;using System.Net.Http;using System.Net.Http.Json;using System.Text;using System.Text.Json;using System.Threading.Tasks;namespace TailRent{ ///