cd /news/large-language-models/claude-code-context-is-like-milk-kee… · home › topics › large-language-models › article
[ARTICLE · art-142839] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Claude Code Context Is Like Milk, Keep It Fresh and Condensed

A developer argues that LLM context behaves like milk — best kept fresh and condensed — because Claude has no memory between requests and re-sends the entire conversation, including stale corrections and previously read files, on every call. The post explains that quality degrades as the context window fills, well before the limit, and recommends starting a new session for each distinct topic and using /clear or /compact to keep context lean.

by read8 min views3 publishedSep 30, 2026

I like milk. I drink it twice a day: once with smoothie, once with cereal. I like it fresh. You can tell when milk has gone bad. It has a weird, slightly sour smell to it. I also like condensed milk. I like it in my sandwich (peanut butter + condensed milk)… yum! (Man, now I'm hungry!)

Anyway, why did I mention milk? Because I think LLM context is like milk, and with LLMs, context is king.

AI context is like milk; it's best served fresh and condensed!

Claude is smartest at the start of a conversation, when its context is smallest (not empty: the system prompt, tools, CLAUDE.md, and MCP definitions are already loaded before you type anything). Over time as your context grows, Claude doesn't actually "remember" anything between requests. Every time you send a message, the whole conversation so far is sent again and Claude has to re-read all of it. Over time, performance goes down. There's a bit of nuance to it because of prompt caching, but that's another topic.

It's like when you go to a drive-through and you're ordering for four families: "Yes, I'd like to have a crispy chicken sandwich, 2 spicy crispy chicken sandwiches, a California-style burger, protein style, 3 medium french fries, make 1 of them animal style, please, also can I have 1 fried okra, also can I have 3 medium drinks: a Coke, a Dr Pepper, and a Diet Coke, and 1 large root beer, and then a vanilla ice cream and 2 oatmeal cookies, with a ketchup and 2 thousand island dipping sauces, actually scratch the crispy chicken sandwich, can I have a Chicago-style hot dog instead? Go easy on the relish, and can I have extra napkins and 2 sets of knives?"

That is going to confuse the heck out of the poor guy serving you. Instead, it would be easier if each family came up and ordered what they wanted.

Family 1: "I'd like 2 spicy crispy chicken sandwiches and 2 medium french fries with a large root beer, please."

Family 2: "I'd like a crispy chicken sandwich, actually make it a Chicago-style hot dog and go easy on the relish, please, also a medium Coke."

...etc.

Notice that in one order, a family changes their mind: "actually scratch the crispy chicken sandwich". The same thing happens with Claude: when you correct something, the wrong version doesn't disappear. It stays in the context right next to the correction. Each correction, "no, actually...", leaves a stale order on the ticket.

If you have two distinct topics, start a new session for each.

Do you ever notice that after having a long conversation with Claude, it starts making more mistakes? The model's weights don't change mid-session; its context does. Claude's models are great, even the smallest one, Haiku. If you notice Claude getting sloppy over time, that's probably because its context window is filling up. Quality starts slipping well before the limit (and when it does get nearly full, Claude Code auto-compacts).

Performance degrades as context fills. It's a quality problem caused by quantity: the more tokens in the window (especially stale or irrelevant ones like the corrections above), the more Claude's attention gets diluted, even when you're nowhere near the limit.

Claude has no memory between requests, so each request re-sends the whole context:

So yes, CLAUDE.md is sent every time. And a file you told Claude to read an hour ago isn't read once and forgotten; it rides along on every request until you /clear or /compact. That's a lot of text being sent out.

Two more things:

Nested CLAUDE.md files in subdirectories load when Claude reads files in those folders. After /compact, the root CLAUDE.md is re-read from disk, which is why it survives compaction; nested ones come back the next time Claude reads files in their folder.

Try to focus on the conversation. What are you trying to get out of it? If you have a stray thought and you're about to go down a rabbit hole, ask it in a separate session (or ask it with /btw, which answers a side question without adding it to the conversation history).

Pay attention to your context. You can see how much context you've used with /context (try it! Run /context, then ask it to do something, then run /context again). It also breaks down what's using it (system prompt, tools, MCP, memory files, messages), so you can see what's hogging space.

Even better, you can also create a status line that shows context! What's really cool is that you don't need to write the script yourself. You can ask Claude to do it for you: "Hey Claude, can you update my statusline to show context usage?" (or run /statusline show context usage).

After a while, it may be a good idea to shrink your context. You can use /compact to have Claude summarize what you've been talking about. One downside is that Claude, upon compacting your conversation, may lose some important details. Repeated compaction can also degrade quality.

You may want to re-inject the essentials after compaction. You can do that with a SessionStart hook on the compact matcher. Whatever a SessionStart hook prints to stdout gets added to Claude's context, and the compact matcher makes it run right after each compaction (the other matchers are startup, resume, clear, and fork). In .claude/settings.json:

{
  "hooks": {
    "SessionStart": [
      {
        "matcher": "compact",
        "hooks": [
          { "type": "command", "command": "cat \"$CLAUDE_PROJECT_DIR/.claude/essentials.md\"" }
        ]
      }
    ]
  }
}

Since CLAUDE.md already survives compaction, use this for things that change as you work: the current task file, git status, or the "decisions so far" section of a handoff file like HANDOFF.md (more on that below).

You can also tell Claude what to focus on when compacting. Let's say you are architecting the new user subscription API and you want to keep the context of that particular topic. Run /compact focus on the user subscription API. If you always want the same focus, put something like # Summary instructions section in your CLAUDE.md so every compaction follows it.

Sometimes Claude implements something that doesn't work as expected. Now you have a pile of useless code. You don't need to delete it by hand. Run /rewind (or press Esc twice) to open the rewind menu, where you can restore the code, the conversation, or both to an earlier point. It can also summarize the conversation from a chosen point forward, freeing up context. One catch: checkpoints only track edits Claude makes with its file-editing tools. Changes made through bash (rm, mv, etc.), by most subagents, or outside Claude Code aren't restored.

When you're ready to move on or dig into a different topic, it's a good idea to start fresh: /clear (/new and /reset are aliases).

If you're planning to pick up where you left off in a new session, it may be a good idea to create a bridge (like a Hemingway Bridge). Ask Claude to create something like a HANDOFF.md so when you open a fresh Claude session, you can just point it at the HANDOFF.md file so it knows where it left off. This helps you keep context usage low. (Note: /resume and --continue are not fresh starts. They reload the old context in full, baggage included.)

You may want to hand off exploration to subagents: "Use a subagent to investigate the different HTTP error codes the show method in users_controller.rb returns." The subagent's file reads and searches stay in its own context; only its summary comes back to yours. If you don't care about the exploration and don't want it to pollute your context, try it.

Claude is pretty good at detecting your intent. But sometimes we aren't on the same wavelength and I have to keep correcting it. It's a good idea to clear after two failed corrections. If you've corrected Claude twice on the same thing and it's still wrong, /clear and start over with a better prompt that includes what you learned. Remember that every correction leaves the old order on the ticket. By then, your context holds two wrong attempts and two corrections.

Keep CLAUDE.md lean. It's paid for on every request. Move specialized instructions into skills (remember that skills are loaded in full only when used) or into rules scoped to specific file paths. I like to keep my CLAUDE.md under 200 lines, following the Claude Code docs.

Trim MCP servers. /context shows how much space tool definitions take. Disable servers you aren't using with /mcp. If you think you won't be using Slack in this session, disable it.

Be specific and point at specific things. "Read src/billing.ts lines 40-120" is far cheaper than "look around the billing code."

You can also run a command yourself with the ! prefix. In Claude Code, ! runs a shell command directly, and its output is added to the conversation. Pipe the output through a filter so only the lines you need go into context.

Examples:

! npm test 2>&1 | tail -50
! grep ERROR logs/app.log

Claude re-reads the whole conversation on every request, so performance drops as context fills. Watch it (/context), condense it (/compact), cut it (/rewind, subagents), and refresh it (/clear + HANDOFF.md). One topic per session.

── more in #large-language-models 4 stories · sorted by recency
── more on @claude 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-context-…] indexed:0 read:8min 2026-09-30 · —