# Context management is an underrated habit

> Source: <https://ideas.fin.ai/p/context-management-is-an-underrated>
> Published: 2026-10-01 13:43:56+00:00

How you manage context in a Claude Code session has a direct effect on both your token bill and the quality of what you get back. Do it well and you spend less for better work.

An efficient session gives Claude the context it needs to finish the job while removing context that has stopped being useful. That means starting with a lean setup, keeping investigations focused, and deliberately deciding when to continue, compact, or start again.

Here are the context management techniques we use on the [2x Team](https://ideas.fin.ai/p/2x-nine-months-later) at Fin.

## Keep your starting context small

Some context is loaded before your first prompt: project instructions, memory, skill descriptions, and tool information. As you work, files, command output, and conversation add to it.

/context shows the [breakdown](https://code.claude.com/docs/en/context-window) of your session. Every few weeks, in a fresh session, use it to check and optimize your starting context. This is especially important if you have accumulated plugins or personal instructions. Disable any unused plugins and MCP servers, remove duplicated instructions, and keep automatically loaded memory focused on information you repeatedly need.

Some practical tips:

- If your starting context is especially large, run /doctor and review what your setup loads and what you actually use day to day. This will also provide recommendations of ways to improve installation health.
- If you have long running tasks, put detailed task history in a separate reference file or a GitHub/Linear issue that Claude can read when relevant.
- Every few months, delete all your plugins and reinstall the ones you need. I can’t tell you how many people have accumulated plugins they only needed once and forgot about.
- Be ruthless with your [skills](https://ideas.fin.ai/p/claude-code-good-skills-bad-skills) . Use /skills to look at your list of enabled skills and disable the ones you don’t need.

## Make sure your memory is up to date

To clean up saved memory that may have gone stale, run [/audit-memory](https://github.com/intercom/2x-skills/blob/main/plugins/claude-code-tools/skills/audit-memory/SKILL.md). It checks claims against current evidence, then proposes what to correct or archive. This will give you both a smaller, more useful memory and fewer outdated instructions influencing future work.

## Keep the working conversation focused

Give Claude enough detail to identify the task upfront, including the outcome, relevant files or links, constraints, and how to check success. A useful paragraph can prevent a long speculative search. For example:

Investigate why this test fails on CI. Start with the linked failure and the test’s implementation. Return the likely cause, supporting evidence, and the smallest proposed fix.

As work proceeds, watch for large results that add little to the next decision: full logs when a failure excerpt would do, broad file reads, repeated status checks, or several attempts at the same failed tool call. Ask for filtered results and a change of approach when the loop stops producing evidence.

## Model and effort choices

People often say “use the right model for the right work.” That’s fairly open-ended advice, but here’s what it can mean in practice:

### Use the right effort level for each model

In my experience, Fable/Astra work really well at medium effort. They are more likely to come up with pragmatic solutions rather than focusing on extreme edge cases. Opus (pre Opus 5.5) /Sol work well at high effort, but Sol has a tendency to go into edge cases that may not make sense in a practical environment.

However, these are just my recommendations. Every engineer’s work is different, so try different effort modes across models and see what works for you.

### Use subagents with smaller models

I personally don’t think Sonnet or Terra/Luna do a good job at open-ended tasks. So, delegate execution to these models whenever Opus/Fable/Sol/Astra have given you a good solution.

Don’t overuse subagents, though. They can lack overall context and become too narrowly focused, which means more turns and chances for error.

Generally, if a session produces 3 PRs, 1 per subagent is good enough. Experiment and come up with your own mental models for your style of working.

## Choose when to compact or start fresh

Compactions are a necessary part of [maintaining quality](https://akshay.co/posts/compactions-are-good-actually/). Turns out, they also help with cost management. Auto-compaction is the worst kind of compaction, so don’t let your session get there.

Make the decision to manually compact at a useful stopping point, like when an investigation has reached a conclusion, a plan is agreed, or a PR is ready for review.

Here are some useful boundaries where it is most definitely a good idea to run manual compaction:

When running a manual compaction, identify what still matters:

`/compact Preserve the goal, agreed approach, constraints, changed files, test results, unresolved questions, and next step. Summarize the completed investigation briefly.`

Compaction replaces conversation history with a summary; think of this as a lossy handoff. When you need important decisions or results to survive reliably, keep them in a concise, inspectable artifact. For personal handoff notes, use a location outside the repository. A fresh session can read these notes and the relevant files. For more on this, see [”How Claude Code works](https://code.claude.com/docs/en/how-claude-code-works)” from Anthropic.

At Fin, our managed settings are configured to auto-compact at 320,000 tokens**.** We also have long-session nudges that appear at 200,000 context tokens, $75 of recorded session cost, or 100 user prompts while task-boundary nudges can appear above 120,000 tokens.

These are reminders for users to reassess their session, not instructions to compact every time one appears. But over time, they build the habit of triggering manual compaction when appropriate. Consider adding similar nudges in your workspace.

## Avoid in-context polling

When waiting for a continuous integration (CI), review, or deployment, Claude gets stuck in a loop where it checks the status, sleeps, and checks again. Each time the result comes back “still pending,” the model processes the session context and adds context that doesn’t help with the task.

Some practical tips to avoid this:

- Add a [watcher](https://github.com/intercom/2x-skills/tree/main/plugins/pr-tools/scripts) script and ask Claude to use it. We’ve shipped shared PR watchers and nudges that steer Claude away from repeated status checks. A watcher handles the waiting and brings Claude back when there is something to act on.
- Keep the waiting inside a tool. Commands like gh run watch can wait for completion without asking the model to decide what to do after every check. The tool may still poll internally, but it avoids the model taking repeated turns.
- Watch for repeated checks with no new information. If you see a sequence of sleep, gh pr checks, and “still waiting,” interrupt it and ask Claude to use a supported watcher.
- Give the wait a clear stopping condition. For example: “Wait until the CI finishes, then investigate failures or report that it passed.” Set a timeout where appropriate so Claude doesn’t wait indefinitely for a stuck job.
- Models are very good at writing scripts, so if you want to wait for something, ask it to write a script that polls within the script and only posts a result when the job is finished or errored. This way, Claude runs the script in the background and all the waste in loops is carried out as bash scripts, outside the main context.

## A habit worth building 

A well-run session should be a priority for anyone trying to increase productivity with Claude Code or another coding agent and that starts with context management. Without it, you’ll spend money processing irrelevant information while also clouding the model’s reasoning with excessive detail.

Every time you remove stale instructions or stop an unproductive loop, you lower your token bill, and get better work back.
