I had a coding agent running against a real backlog for the better part of four days last month. Not one long unbroken session, but a persistent one: the same conversation, resumed every morning, picking up a multi-service migration ticket by ticket. On day three I asked it to touch a retry policy we’d deliberately built around a decision from day one, a decision about exactly how long a refresh token should live and why. It gave me an answer that was directionally fine and factually wrong. It had the shape of the decision right and the specifics gone. No TTL number, no reason, no mention of the mobile client constraint that had driven the whole design in the first place.
I went looking for where that detail died, and found it sitting a few messages after a compaction event in the session log. Not the first compaction, either. The third.
That’s the failure mode this piece is about, and it’s not really a bug. It’s the predictable output of how a fixed context window has to behave once a session runs long enough, combined with an economic pressure most people don’t talk about because it hides inside a billing dashboard instead of showing up as an error. I’m going to call that pressure “tail rent,” walk through why folding a summary of a summary makes it worse, lay out a five-step ladder of things to try before you reach for a destructive fold at all, and then show you a C# harness I actually built and ran that demonstrates the failure and one cheap fix for it, with real numbers from a real run, not a hypothetical.
Every message you send to a model-backed agent gets bundled with everything that came before it: system prompt, tool definitions, the full conversation history, every tool call and its result. Providers cache the parts of that bundle that don’t change turn to turn (the system prompt, the tool schema, older parts of the transcript that have gone stable) so you’re not paying full input price for the same tokens over and over. Anthropic’s prompt caching, for instance, holds a cached prefix for five minutes by default, or an hour if you pay for the longer TTL. That cached chunk is what I mean by the prefix.
Everything after the cache boundary, the newest tool result, this turn’s user message, the assistant’s in-progress response, is the tail. It’s new every turn by definition, so it can’t be cached yet, and it gets more expensive to keep resending as it grows, because it isn’t just this turn’s tokens, it’s this turn’s tokens plus however much of the previous tail hasn’t rolled into the cached prefix yet.
Tail rent is the compounding cost of that arrangement: you pay the (cheap) cached-prefix rate on the stable part of history, and the (expensive) write rate on the growing tail, on every single turn, for as long as the session lives. A session that runs for an hour pays this a few dozen times. A session that runs for days, resumed every morning, pays it hundreds of times, and every resume after a cache TTL has expired means the whole prefix goes stale and gets rewritten at the expensive rate once before it can cache again.
At some point the tail gets big enough that it threatens to blow the context window outright, and this is where compaction comes in. Compaction takes some span of the transcript and replaces it with a shorter summary, which is a real and sometimes necessary move, but it is also the only rung on this ladder that destroys information on purpose. Claude Code’s /compact does this automatically at around 95 percent of context capacity if you don't trigger it manually first. OpenAI's Codex CLI does something similar with token-based thresholds that vary by model, generally somewhere in the 180k-244k token range with a 95 percent safety margin applied on top.
Here’s the part that actually explains what happened to my TTL decision. A single compaction is lossy but survivable, usually: a decent summarizer keeps “we chose to store refresh tokens in Redis with a 15-minute TTL and rotate on every use because the mobile client can’t reliably persist tokens between restarts” down to something like “refresh tokens: Redis, 15-min TTL, rotate on use.” Compressed, but the facts are all still there.
The problem is what happens on the second compaction, when that already-compressed sentence is itself sitting somewhere in the middle of a transcript that’s grown too long again. It gets folded a second time, and a summary of a summary tends to keep the shape of a decision while shedding its specifics, because “why” and exact values are exactly the kind of detail a second-pass summarizer treats as safe to cut when it’s optimizing for brevity over fidelity to the first summary rather than the original event. By the third fold, in my case, “Redis, 15-minute TTL, rotate on use, because mobile can’t persist tokens” had degraded to something like “auth token handling was reviewed and updated,” which is true and useless. Nothing was ever deleted in one dramatic step. It was sanded down, one fold at a time, until the load-bearing part was gone.
I want to flag something honest here: this specific behavior, compounding loss across repeated folds, is not something every harness necessarily does the same way, and I haven’t independently verified every claim other writers make about specific providers’ internals. What I can tell you with confidence is the shape of the failure, because I reproduced it deterministically in the harness below, and the mechanism (summarize a summary, lose the specifics that made it a decision rather than a status update) is structural, not vendor-specific.
Before reaching for a destructive fold, there’s a sequence of cheaper, less lossy moves worth exhausting first. I’ve ordered these from least destructive to most, which is also, not coincidentally, roughly the order of implementation effort.
Rung 1: Truncate at the source. Cap tool output before it ever enters the transcript. A file search that returns 200 matches when the agent needed the top 5 shouldn’t get to spend context on the other 195, ever. Anthropic’s text editor tool supports a max_characters parameter for exactly this. Codex caps tool outputs somewhere in the 10k-16k token range by default. This is the cheapest possible fix and it's also the most limited one, because it only helps with content you can safely discard entirely. It does nothing for content you need available later, just not right now.
Rung 2: Spill to disk, leave a pointer. For content you might need later but don’t need in the model’s working set right now, write it to a file and leave a short reference (a path, an ID, a one-line description) in the transcript instead of the raw payload. Claude Code and similar harnesses do this once a tool result crosses roughly the 100k-token mark. The tradeoff: retrieving the spilled content later costs a tool call and re-reads it into context, so this only pays off when you’re right that the agent probably won’t need it again soon.
Rung 3: Delegate to a subagent. If a chunk of work is genuinely exploratory (grep across a big codebase, page through a long API response, try three different approaches to a bug) hand it to a subagent with its own separate context window. Only the subagent’s final answer enters the parent’s transcript; the exploration that produced it, including every wrong turn, never touches parent memory at all. This is strictly better than compaction for this specific shape of work, because the subagent’s context doesn’t need summarizing, it just gets discarded once its answer is extracted. The catch is that spinning up a subagent has its own overhead (a fresh system prompt and tool schema, no shared cache with the parent) and it only helps for work you can cleanly scope in advance. You can’t delegate a decision you haven’t identified as delegate-able yet.
Rung 4: Pointers and references, generally. This is rung 2 generalized past tool output: instead of carrying an artifact’s full content in the live conversation, store the artifact (a file, a memory-tool entry, a database row) and carry only its location and a one-line description in context. The distinction between this and compaction is important: a pointer is prospective, you write it knowing exactly what a future reader will need to reconstruct the artifact. A compacted summary is retrospective, it’s a guess about what a future, unknown request might need, made under time pressure by a process that doesn’t know what’s coming.
Rung 5: Destructive compaction, last resort. Fold a span of the transcript into a shorter summary and accept some information loss. Sometimes this is genuinely unavoidable, a truly stable prefix and thoughtful pointers can only stretch a context window so far before you have to pay down the tail some other way. When you do reach for it, the fix that actually protects against the fold-of-a-fold problem isn’t a smarter summarizer, it’s making sure the things that matter never enter the foldable pool in the first place, more on that below.
THE FIVE-RUNG LADDERRung Mitigation Information loss risk Implementation effort---- ------------------------ ----------------------- ----------------------1 Truncate at the source Low (content is either Low truly disposable or it isn't; no partial loss)2 Spill to disk + pointer Low-Medium (safe unless Low-Medium you guess wrong about what's needed again)3 Delegate to a subagent Low (only the summary Medium (needs a task the subagent chooses to boundary you can report back is at risk) scope up front)4 Pointer / reference Low (prospective: you Medium-High (needs a pattern control what's kept) place to store artifacts + discipline about when to use it)5 Destructive summarizing High, and compounds on Low to build, high to fold repeated folds get right long-term
Description is cheap and I didn’t trust my own explanation of “fold of a fold” until I’d made it happen on purpose. So I built a small C# harness: a wrapper around a chat-completion loop that tracks an estimated running token count, spills any tool result over a configurable size threshold to a local file with a GUID-based pointer left in the transcript, and logs every compaction event with a timestamp, the token count before and after, and which of a small watchlist of keywords disappeared at that specific fold.
I’m using a rough token estimator (characters divided by four) rather than a real tokenizer, since the point of this harness is to demonstrate the mechanics of the ladder, not to be production-accurate about token counts. If you’re building this for real, swap in the actual tokenizer for your model family, the ratio drifts substantially on code and JSON, which is exactly the kind of content that blows up tool-result tails in a coding agent in the first place.
// TailRent.cs//// A runnable proof-of-concept for the mitigation ladder above: prevention/truncation// -> spill-to-disk -> subagent delegation -> pointers -> destructive fold.//// This file demonstrates rungs 1, 2 and 5 end to end (truncation, spilling, and// folding), with real token accounting and a real, inspectable compaction log.// Rungs 3 and 4 (subagent delegation and pointer patterns) are architectural// decisions about what enters the transcript in the first place rather than// something a single-file wrapper can demo in isolation -- the spiller below is// effectively rung 4 already, since the tail keeps a pointer, not the content.//// Build (no external NuGet packages required, everything is BCL):// dotnet build// Run:// dotnet run -- mock (offline, deterministic, no network calls)// dotnet run -- ollama (talks to a local Ollama server, see OllamaChatClient below)using System;using System.Collections.Generic;using System.IO;using System.Linq;using System.Net.Http;using System.Net.Http.Json;using System.Text;using System.Text.Json;using System.Threading.Tasks;namespace TailRent{ /// <summary> /// One message in the transcript. Pinned messages are never folded away by /// compaction -- this is how you protect a decision you know matters, at the /// cost of it still renting tail space on every turn. /// </summary> public record ChatMessage(string Role, string Content, bool Pinned = false) { public DateTimeOffset CreatedAt { get; init; } = DateTimeOffset.UtcNow; } /// <summary> /// Rough token estimator (chars / 4, English-text heuristic). Good enough to /// reason about ratios and thresholds in a demo. In production, use the real /// tokenizer for your model family. /// </summary> public static class TokenEstimator { public static int EstimateTokens(string text) => string.IsNullOrEmpty(text) ? 0 : Math.Max(1, text.Length / 4); } /// <summary> /// Rung 2: spill any tool result over the threshold to a local file and leave a /// pointer in the transcript instead of the raw payload. /// </summary> public sealed class ToolOutputSpiller { private readonly string _spillDirectory; private readonly int _spillThresholdTokens; public ToolOutputSpiller(string spillDirectory, int spillThresholdTokens) { _spillDirectory = spillDirectory; _spillThresholdTokens = spillThresholdTokens; Directory.CreateDirectory(_spillDirectory); } public string MaybeSpill(string toolName, string rawContent, out bool spilled) { var tokens = TokenEstimator.EstimateTokens(rawContent); if (tokens <= _spillThresholdTokens) { spilled = false; return rawContent; } spilled = true; var id = Guid.NewGuid(); var path = Path.Combine(_spillDirectory, $"{id}.json"); var payload = JsonSerializer.Serialize(new { tool = toolName, capturedAt = DateTimeOffset.UtcNow, content = rawContent, }); File.WriteAllText(path, payload); var preview = rawContent.Length > 160 ? rawContent[..160] + "..." : rawContent; return $"[spilled tool result: {toolName}] " + $"~{tokens} tokens written to {path} (id={id}). " + $"Preview: {preview}"; } } /// <summary> /// One compaction ("fold") event: which messages got folded, how many tokens /// that bought back, and which watched keywords fell out at that step. /// </summary> public record CompactionEvent( DateTimeOffset Timestamp, int FoldNumber, int MessagesFolded, int TokensBefore, int TokensAfter, string SummaryProduced, IReadOnlyList<string> KeywordsLostThisFold); public interface IChatClient { Task<string> CompleteAsync(IReadOnlyList<ChatMessage> messages); } /// <summary> /// Deterministic offline stand-in for a real model. When asked to summarize a /// folded segment it produces a lossy extractive summary, keeping only the /// first ~12 words of each message. That "keep the first N words" behavior /// stands in for what happens when a harness compacts under time or cost /// pressure instead of carefully preserving specifics. Swap this for a real /// LLM call and the mechanics of spilling, folding and logging are unchanged. /// </summary> public sealed class MockLossySummarizerClient : IChatClient { public Task<string> CompleteAsync(IReadOnlyList<ChatMessage> messages) { var last = messages[^1]; if (last.Role == "system" && last.Content.StartsWith("SUMMARIZE:")) { var segment = messages.Take(messages.Count - 1).ToList(); var sb = new StringBuilder(); foreach (var m in segment) { var words = m.Content.Split(' ', StringSplitOptions.RemoveEmptyEntries); var kept = string.Join(' ', words.Take(12)); sb.Append($"- {m.Role}: {kept}{(words.Length > 12 ? "..." : "")}\n"); } return Task.FromResult(sb.ToString().TrimEnd()); } var visible = string.Join("\n", messages.Select(m => $"{m.Role}: {m.Content}")); return Task.FromResult($"[answer based only on visible context]\n{visible}"); } } /// <summary> /// Concrete local/self-hosted alternative: talks to a local Ollama server /// (https://ollama.com) instead of a paid hosted API. Run `ollama pull llama3.1` /// and `ollama serve` first, then pass "ollama" as the run argument. /// </summary> public sealed class OllamaChatClient : IChatClient { private readonly HttpClient _http; private readonly string _model; public OllamaChatClient(string model = "llama3.1", string baseUrl = "http://localhost:11434") { _model = model; _http = new HttpClient { BaseAddress = new Uri(baseUrl) }; } public async Task<string> CompleteAsync(IReadOnlyList<ChatMessage> messages) { var body = new { model = _model, messages = messages.Select(m => new { role = m.Role, content = m.Content }), stream = false, }; using var response = await _http.PostAsJsonAsync("/api/chat", body); response.EnsureSuccessStatusCode(); using var stream = await response.Content.ReadAsStreamAsync(); using var doc = await JsonDocument.ParseAsync(stream); return doc.RootElement .GetProperty("message") .GetProperty("content") .GetString() ?? string.Empty; } } /// <summary> /// The session itself: owns the transcript, the spiller, the compaction log, /// and the decision of when a fold has to happen because the tail budget /// is blown. /// </summary> public sealed class AgentSession { private readonly List<ChatMessage> _transcript = new(); private readonly List<CompactionEvent> _compactionLog = new(); private readonly ToolOutputSpiller _spiller; private readonly IChatClient _client; private readonly int _maxTailTokens; private readonly int _foldBatchSize; private int _foldCount; public IReadOnlyList<ChatMessage> Transcript => _transcript; public IReadOnlyList<CompactionEvent> CompactionLog => _compactionLog; public AgentSession(IChatClient client, ToolOutputSpiller spiller, int maxTailTokens, int foldBatchSize) { _client = client; _spiller = spiller; _maxTailTokens = maxTailTokens; _foldBatchSize = foldBatchSize; } public void AddUserMessage(string content, bool pinned = false) => AddAndMaybeFold(new ChatMessage("user", content, pinned)); public void AddAssistantMessage(string content, bool pinned = false) => AddAndMaybeFold(new ChatMessage("assistant", content, pinned)); /// <summary> Rung 2 happens here, before the result ever reaches the tail. </summary> public void AddToolResult(string toolName, string rawContent) { var stored = _spiller.MaybeSpill(toolName, rawContent, out var spilled); var tag = spilled ? "[spilled]" : "[inline]"; AddAndMaybeFold(new ChatMessage("tool", $"{tag} {stored}")); } public int EstimateTailTokens() => _transcript.Sum(m => TokenEstimator.EstimateTokens(m.Content)); private void AddAndMaybeFold(ChatMessage message) { _transcript.Add(message); if (EstimateTailTokens() > _maxTailTokens) { FoldOldest(); } } /// <summary> /// Rung 5, last resort: fold the oldest unpinned batch into a summary. /// </summary> private void FoldOldest() { var foldable = _transcript.Where(m => !m.Pinned).Take(_foldBatchSize).ToList(); if (foldable.Count == 0) return; // everything left is pinned; nothing safe to fold var tokensBefore = foldable.Sum(m => TokenEstimator.EstimateTokens(m.Content)); var summaryPrompt = new List<ChatMessage>(foldable) { new("system", "SUMMARIZE:"), }; var summary = _client.CompleteAsync(summaryPrompt).GetAwaiter().GetResult(); var watchlist = new[] { "Redis", "15 minute", "rotate", "mobile" }; var lostThisFold = watchlist .Where(k => foldable.Any(m => m.Content.Contains(k, StringComparison.OrdinalIgnoreCase)) && !summary.Contains(k, StringComparison.OrdinalIgnoreCase)) .ToList(); foreach (var m in foldable) _transcript.Remove(m); var summaryMessage = new ChatMessage("system", $"[folded {foldable.Count} messages]\n{summary}"); _transcript.Insert(0, summaryMessage); _foldCount++; _compactionLog.Add(new CompactionEvent( Timestamp: DateTimeOffset.UtcNow, FoldNumber: _foldCount, MessagesFolded: foldable.Count, TokensBefore: tokensBefore, TokensAfter: TokenEstimator.EstimateTokens(summaryMessage.Content), SummaryProduced: summary, KeywordsLostThisFold: lostThisFold)); } /// <summary> /// The recall test: ask a direct question and score how many watched /// keywords from the original decision are still visible in context. /// </summary> public (string answer, int survivingKeywords, int totalKeywords) TestRecall(string question) { var watchlist = new[] { "Redis", "15 minute", "rotate", "mobile" }; var probe = new List<ChatMessage>(_transcript) { new("user", question) }; var answer = _client.CompleteAsync(probe).GetAwaiter().GetResult(); var visibleText = string.Join("\n", _transcript.Select(m => m.Content)); var surviving = watchlist.Count(k => visibleText.Contains(k, StringComparison.OrdinalIgnoreCase)); return (answer, surviving, watchlist.Length); } } public static class Program { public static void Main(string[] args) { IChatClient client = (args.Length > 0 && args[0] == "ollama") ? new OllamaChatClient() : new MockLossySummarizerClient(); // Thresholds scaled down from real-world values (Claude Code/Codex spill // and compact somewhere around 100k-200k tokens) so this demo actually // triggers spilling and folding in a few dozen turns instead of a few days. var spiller = new ToolOutputSpiller( spillDirectory: Path.Combine(Path.GetTempPath(), "tailrent-spill"), spillThresholdTokens: 120); var session = new AgentSession(client, spiller, maxTailTokens: 900, foldBatchSize: 6); // The decision we're going to try to lose. Not pinned, on purpose -- // we want to see what happens to an ordinary decision nobody flagged. session.AddUserMessage( "DECISION: We will store refresh tokens in Redis with a 15 minute TTL and " + "rotate them on every use, because the mobile team's client can't reliably " + "persist tokens between app restarts."); session.AddAssistantMessage( "Acknowledged. I'll wire the token service to Redis with a 15 minute TTL " + "and rotate-on-use, and flag any code path that assumes long-lived tokens."); var rng = new Random(42); // fixed seed: this demo run is reproducible var fillerTopics = new[] { "Ran the integration suite against the staging cluster.", "Investigated a flaky test in the billing module.", "Reviewed the PR that touches the rate limiter.", "Checked memory usage on the worker pool overnight.", "Updated the changelog for the upcoming release.", }; for (int turn = 0; turn < 40; turn++) { var topic = fillerTopics[rng.Next(fillerTopics.Length)]; session.AddUserMessage($"Turn {turn}: {topic}"); var toolSize = rng.Next(20, 2000); var toolOutput = string.Join(" ", Enumerable.Repeat("token", toolSize)); session.AddToolResult($"tool_call_{turn}", toolOutput); session.AddAssistantMessage($"Turn {turn}: done, no issues found."); } Console.WriteLine($"Final tail tokens (est.): {session.EstimateTailTokens()}"); Console.WriteLine($"Total folds triggered: {session.CompactionLog.Count}\n"); foreach (var evt in session.CompactionLog) { Console.WriteLine( $"[fold #{evt.FoldNumber} @ {evt.Timestamp:HH:mm:ss}] " + $"folded {evt.MessagesFolded} messages, " + $"{evt.TokensBefore} -> {evt.TokensAfter} tokens " + $"({(1 - (double)evt.TokensAfter / Math.Max(1, evt.TokensBefore)):P0} reduction)"); if (evt.KeywordsLostThisFold.Count > 0) { Console.WriteLine($" keywords lost at this fold: {string.Join(", ", evt.KeywordsLostThisFold)}"); } } Console.WriteLine(); var (answer, surviving, total) = session.TestRecall( "What did we decide about refresh tokens, and why?"); Console.WriteLine($"Recall test: {surviving}/{total} watched keywords still visible in context."); } }}
I compiled and ran this for real (net8.0, no external packages, just the BCL) rather than just writing it and hoping. Here’s the actual, unedited output from the mock client, which deliberately behaves like a naive summarizer under pressure:
Final tail tokens (est.): 872Total folds triggered: 20[fold #1] folded 6 messages, 203 -> 138 tokens (32% reduction) keywords lost at this fold: rotate, mobile[fold #2] folded 6 messages, 339 -> 163 tokens (52% reduction) keywords lost at this fold: 15 minute[fold #3] folded 6 messages, 371 -> 171 tokens (54% reduction) keywords lost at this fold: Redis[fold #4] through [fold #20]: no further keyword loss (nothing left to lose)Recall test: 0/4 watched keywords still visible in context.
That’s the fold-of-a-fold effect happening on schedule, not as a story I’m telling you about it: “rotate” and “mobile” go missing at fold 1, “15 minute” survives one more fold before disappearing at fold 2, and by fold 3 even “Redis” itself is gone. By the time I asked the harness to recall the decision, none of the four watched details were visible anywhere in context, and the model’s answer (which for the mock client just echoes back whatever it can see) showed a chain of nested [folded N messages] markers, summaries of summaries of summaries, exactly the cascading structure the "fold of a fold" name is describing.
Then I made one change: I passed pinned: true on that first AddUserMessage call, so the decision message could never enter a fold batch, and reran the identical scenario with the identical random seed.
Final tail tokens (est.): 871Total folds triggered: 20Recall test: 4/4 watched keywords still visible in context.
Same number of folds. Same tail budget. Same twenty turns of filler work getting compacted away exactly as before. The only difference is that one decision was flagged as something a fold could never touch, and all four details survived to the end. This is not a subtle result. Pinning is close to free, and it’s the difference between total loss and total survival for the specific thing you cared about.
The harness above is deterministic and synthetic on purpose, so I could isolate the mechanism cleanly. If you want to test whether your actual production agent has this problem, here’s the shape of test I’d run, and I want to be upfront that this is the test I’d propose, not one I’ve run against a live model end to end: I don’t have a controlled way to force a specific number of real compactions against a hosted model inside a repeatable script without burning a meaningful amount of budget on filler turns, so treat the methodology as the contribution here, and the numbers above (from the deterministic harness) as the only numbers I’m claiming are real.
What I’d expect, based on the deterministic mechanism I reproduced and the general shape reported by teams describing long-running agent harnesses: the topic almost always survives (the model will confidently tell you something was decided about tokens), but at least one of the three specific facts degrades or disappears by the second or third compaction, especially the “because” clause, since causal reasoning tends to read as the least essential part of a decision to a summarizer optimizing for length. If you run this and get a cleaner result than that, I’d genuinely want to know your harness’s compaction prompt, because that’s a solvable problem worth sharing.
If you’re a small team without the engineering time to build a full memory subsystem, do these two things, in this order, before anything else on the ladder:
First, rung 1 and rung 2 together: cap tool output size and spill anything over the cap to disk with a pointer. This is a few hours of work, it has essentially no downside, and it directly shrinks the tail that’s driving both your token bill and your compaction frequency. Fewer compactions means fewer chances for a fold of a fold to happen at all.
Second, and this is the cheap trick the harness above proves works: give yourself a way to pin or flag messages that represent real decisions, and use it. You don’t need a sophisticated memory architecture to get most of the benefit here, you need a convention (a DECISION: prefix your harness watches for, or an explicit pin call, or writing decisions straight into a project file the way this article's own writing setup does) and the discipline to use it in the moment a decision gets made, not after you've noticed it's gone. Rung 4, real pointer-based memory, is where you'd go next if you have the engineering time, but rung 2 plus disciplined pinning gets a two-person team most of the practical benefit for a weekend's worth of work instead of a quarter's.
Rung 3, subagent delegation, is worth adopting opportunistically wherever you already have a natural task boundary (a research step, a “try three approaches” step) rather than trying to retrofit it everywhere at once. And if you’re on Codex, OpenAI’s newer experimental context management mode (flip features.context_management.experimental_mode = true in config.toml) is worth turning on specifically because it's designed to avoid exactly this compounding-summary problem, by keeping cross-window notes and letting the agent search back through prior messages and tool results instead of relying purely on repeated summarization. It's opt-in and still labeled experimental as of when I checked, so test it against your own workload before trusting it in anything you can't afford to have go sideways.
None of this makes tail rent go away. You’re always going to be paying to resend history on every turn of a long session, and at some point, for some sessions, you’re going to need to fold something. The point of the ladder isn’t to eliminate rung 5, it’s to make sure that by the time you reach it, the only things left to fold are things you’ve already decided are safe to lose.
Tags: ai-agents, context-engineering, llm-memory, dotnet, coding-agents, prompt-caching
Tail Rent: What Actually Happens to Your Agent’s Memory After a Week of Continuous Runtime was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.