# Four ways I make my Claude budget last

> Source: <https://theotherjenai.substack.com/p/four-ways-i-make-my-claude-budget>
> Published: 2026-10-05 09:10:45+00:00

Yes, I did start this Substack with an article about how measuring tokens is not a sensible way to manage AI use, and I stand by that - but I also have a Claude Code account through work, and that account has a budget (which shall remain undisclosed…), and when I hit it I get no more Claude.

I’ve been using AI to do my job for the best part of four years now, with varying degrees of success, and I’m so out of the habit of writing code by hand that when Claude is gone I become fairly incapable. Work still wants shit to get done though, so it’s very much in my interest to squeeze as much as I can out of whatever budget they give me.

For the last six months or so I’ve been running a service I wrote (vibed) that sits on my machine and analyses my Claude sessions. It’s super messy and horrible and will not be open-sourced, but it has taught me a lot about where Claude Code actually spends the money and which levers I can fiddle with. At one point I had nine of these written up, which was far too many for anyone to read in one go, so these are the four that made the most difference.

The thing underneath all of them is embarrassingly simple - **every request re-sends the whole conversation.** A noisy tool output isn’t paid for once, it’s paid for again on every single turn for the rest of the session. So the question isn’t really how much you type, it’s what ends up in the context window and how long it stays there.

## Don’t break the cache

This one made the biggest difference by a massive margin, so if you only read one section, make it this one.

Claude caches the repeated parts of your conversation and charges a small fraction of the normal rate to read them back. It works automatically, all you have to do is not break it, which is easier said than done because three very normal habits break it. The first is starting a new session for a small follow-up - every new session builds its cache from scratch, so if you’ve got a little question related to the thing you’re already working on, ask it in the conversation you’re already in. The second is restructuring early context mid-session, because the cache is keyed on the prefix, so changing something near the top invalidates everything after it. I don’t really do this one, I prefer a long-winded meander through my thoughts and ideas, but YMMV.

The third is walking away for too long, and this is the one that absolutely killed me. The cache has a time-to-live, and the default is five minutes. I regularly have several sessions on the go at once and I would forget about them, come back some long time later way past the TTL, and the cache would just be gone. Luckily it’s a tiny change to bump that to an hour, one block in `~/.claude/settings.json`:

```
{ "env": { "ENABLE_PROMPT_CACHING_1H": "1", "FORCE_PROMPT_CACHING_5M": "0" } }
```

(You can give that to Claude and let it update the settings for you, which feels appropriate.)

I’ve also added a custom statusline to Claude Code that shows what the cache is doing. Claude Code pipes the session details into a script as JSON, and the script turns them into one line at the bottom of the terminal, so it’s a free, always-visible gauge of how much room is left and whether this turn is cheap. It’s not 100% reliable, but it’s good enough to tell me whether I should carry on, compact, or start a new conversation.

The catch is that “keep one long session going” pulls against “keep the context small”. At some point a conversation gets long enough that every turn is expensive even at cache prices, and I don’t have a clean rule for when to cut and start fresh - I go on feel, and the statusline helps the feel.

## Let the expensive model think, not grind

My default model is the most capable one available, because I want its judgement - currently that’s Opus 5.5, which is a gamechanger after the absolute nightmare that was Opus 5 (and I never rated Fable that high despite all the hype). But I don’t want the smart, expensive model doing the boring grinding. So the rule in my global config is that it plans the work, hands implementation and investigation off to cheaper subagents, and keeps the planning, synthesis and final review for itself.

I have a small roster of agents, each pinned to a cheap model and each with one job - a worker for typos, renames and mechanical single-file edits, a worker for well-scoped implementation, a locator that answers “where is X?” with a file:line map, a test runner that runs the suite and reports back what failed and why, and an investigator that fetches CI and error-tracking logs and boils them down to the bit that matters.

The bit that makes this actually work is the escalation contract, which every worker gets told:

**Stay at your level.** If the task turns out to need judgement calls, ambiguity resolution, or multi-file reasoning, stop immediately without making changes and reply starting with `ABOVE MY LEVEL:`. This is a success condition, not a failure.

The reason it helps is that the expensive part of a frontier model isn’t only its per-token rate, it’s that its context keeps growing all session. The test runner might burn through a huge amount of tokens reading pytest output and hand back six lines, and those tokens never touch the main conversation, so they’re never re-sent on any later turn.

Setting it up is three changes. The orchestration rule is four or five lines of plain prose in `~/.claude/CLAUDE.md` telling the model to plan and delegate rather than execute. The cheap default model goes in `~/.claude/settings.json` as `"env": { "CLAUDE_CODE_SUBAGENT_MODEL": "sonnet" }`. And each agent is one Markdown file in `~/.claude/agents/`, with its model pinned in the frontmatter and the escalation contract in the body - give the read-only ones read-only tools, because a locator with `tools: Read, Grep, Glob` can’t break anything while it’s poking around.

I haven’t yet, because I assumed there’s enough examples of this stuff floating around the internet already, but I can stick my relevant agents and skills in a repo if anyone wants to use them as inspiration. Let me know.

The catch is that delegation is not automatically cheaper. My sessions that fan out to four or more agents were dramatically more expensive than the ones with none, because each agent is a full separate conversation carrying its own context, so fanning out multiplies rather than divides. There is a flaw in how I measured that though - comparing sessions that delegate against sessions that don’t measures *when I choose to delegate*, not what delegating actually costs. I reach for agents on hard problems, and hard problems are expensive regardless of model, agent, cache or anything else, so I treat heavy fan-out as a sign of expensive work rather than proof that it causes it. Either way the practical rule holds - delegation pays when an agent absorbs a lot of tokens and hands back a little, and learning when that is is something only you can do, because your work is probably wildly different to mine.

## Write the handoff down properly

This one is really part two of the last one. A subagent given a fuzzy prompt either does the wrong thing or comes back asking questions, which is all easily avoidable waste. So I (meaning, of course, Claude) write delegated prompts as self-contained packets that assume the receiving agent has never seen the conversation - the repo path, the objective, what’s in scope, what’s explicitly out of scope, the relevant files, the expected return format, the verification command and the stop conditions. The format lives in a skill, and here’s an excerpt:

```
## Handoff Packets

Write delegated prompts as self-contained packets. Assume the receiving agent
has not seen the conversation. Include the repo path, objective, scope,
out-of-scope areas, relevant files or search targets, expected return format,
verification commands, and stop conditions.

Useful stop conditions:

- The live code does not match the assumption in the handoff.
- A verification command fails twice after a reasonable fix or retry.
- The work appears to require files outside the assigned scope.
- The agent cannot produce concrete evidence for its claim.
```

The stop conditions are the important bit. An agent that stops early because its premise was wrong is far cheaper than one that grinds on for hours building on a false assumption.

A skill is a Markdown file at `~/.claude/skills/<name>/SKILL.md` with a YAML `name` and `description`, and Claude Code matches on the description and pulls the body in when the work looks like delegation. I repeat the stop conditions in each agent file too, because the skill tells the orchestrator what to write but the agent file is the thing the worker always has loaded. You could put the whole lot in `~/.claude/CLAUDE.md` instead and it would never fail to fire, but then you’re carrying it on every request of every session, including the ones with no delegation in them at all, which feels a bit counterproductive given what this whole article is about.

The catch here is that writing a good packet costs output tokens from the expensive model, so for a small task the packet can cost more than just doing the thing.

## Stop Claude grepping the whole world

`grep -r`, `find` and `cat` dump untrimmed output straight into the conversation, where it sits for the rest of the session costing you tokens on every turn. One unbounded recursive grep in a large monorepo can add tens of thousands of tokens that then ride along on every request after it. Claude Code has its own Grep, Glob and Read tools that trim what comes back, but the model really likes reaching for the shell, so this is the rule in my config:

Search with the Grep and Glob tools, not `grep`/` rg`/` find`/` fd` via Bash. Shell search dumps untrimmed output into the conversation permanently; the tools return trimmed matches.

Cap output that will be long: `| head -50`, `--max-count`, `git log -n 20`, `git diff --stat` before a full diff. Never cat a whole file to read part of it - use Read with an offset.

Prefer one targeted read over exploratory reading. Do not re-read a file already in context.

This is the easiest change on the list - paste that into `~/.claude/CLAUDE.md` and you’re done, no infrastructure, and it works immediately. It does fight the model’s instincts though, so it needs repeating and it still slips.

As part of my session analysis tool I also run a hook that watches Bash calls and logs the expensive patterns - recursive grep, `find` without `-maxdepth`, `cat` of a whole file, `git log` without a limit - which is how I can say with some confidence that cutting down on massive bash output does make a difference to the cost. You don’t need this, but if you want it, it’s a `PreToolUse` hook in `~/.claude/settings.json`, matched on `Bash`, pointing at a script that pattern-matches the command and appends a line to a log. I’d skip it on day one though, the config rule is where the saving is.

None of this is clever really, it’s mostly just not paying for the same thing twice. It all comes from one person, one codebase and one very messy tool, so treat it as vibes with a spreadsheet attached rather than science - but it’s kept me in Claude until the end of the month, which is all I really wanted, because I’d quite like to never find out exactly how much I’ve forgotten about writing code by hand.
