# Claude Code Tool Search Nearly Halves Your Context Bill

> Source: <https://pub.towardsai.net/claude-code-tool-search-nearly-halves-your-context-bill-a14defdf42f3?source=rss----98111c9905da---4>
> Published: 2026-10-04 09:36:53+00:00

**Does Claude Code tool search save tokens? Yes: with 65 MCP tools connected, it cut what Claude Code sends before your prompt from 27,184 tokens to 9,722, a 64% drop, measured on the wire.** Tool search is Claude Code’s lazy tool loading: instead of sending every tool definition on every request, it sends a short list of tool names and a ToolSearch tool, and pulls in a full definition only when the model asks for it.

**Status, 30 September 2026:** Claude Code 2.1.285, official MCP reference servers 2026.8.31, @playwright/mcp 0.0.83. Re-check your own setup with the 18-line script below.

This is part 2 of **The Context Tax, measured**, a four-part series that publishes one part a day. [Part 1](https://medium.com/@decoding_ai_by_nureravi/claude-code-spends-18-006-tokens-before-you-type-a-word-b58166bb90be) found that Claude Code sends about 18,000 tokens before you type, and that 82% of it is tool definitions. Today asks whether you can stop paying for tools you never call.

**Place your bet before you scroll: over a whole task, does lazy loading save more than half?** Answer in the comments. The result is in the middle of this post.

Part 1 asked readers to run a capture script and post their number. **No one has posted a number in the comments yet, so the board is still our six baselines:** Aider 2,283, OpenCode 6,713, Gemini CLI 9,583, Codex CLI 10,020, Qwen Code 15,079 and Claude Code 18,006 tokens. Post yours under this part and it goes on tomorrow’s board.

**Claude Code tool search is lazy tool loading: the agent keeps most tool definitions out of the request and loads each one only when the model searches for it.** In our capture, Claude Code 2.1.285 with tool search on sent 11 tools in full, including ToolSearch itself, and replaced everything else with a one-line list of names. When the model needed a tool, it called ToolSearch, and the next request carried a reference that Anthropic's API expands into the full definition. The [Anthropic tool search documentation](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool) says deferred tools do not enter the context window until discovered.

**With five MCP servers connected, Claude Code tool search cut the first request by 64%, from 27,184 tokens to 9,722.** The five servers were the official filesystem, memory, everything and sequential-thinking reference servers plus Microsoft’s Playwright server, 65 MCP tools in total. Here is what each configuration sent before the word “hi”:

**Your MCP servers cost 9,638 tokens when loaded eagerly and 966 when deferred.** That is a 90% cut on the line you configured yourself.

**No, but close: over a whole task, tool search saved 46%, not the 64% of the first request.** We gave Claude Code one job, “Remember that the deploy is on Friday”, which needs one MCP tool, the memory server’s create_entities. Eager mode took two requests and 54,422 tokens. Lazy mode took three, because the model first had to search for the tool, and still used only 29,543 tokens.

How did your bet do? **The extra round trip cost about 9,700 tokens; every request after it saved about 17,300.** Lazy loading was ahead from the first request.

**Fetching** **create_entities on demand added 140 tokens to every later request in the session.** That is its definition, now part of the context. Five loads cost about 700 tokens. Each request after the search was 9,933 tokens in lazy mode against 27,235 in eager mode, 64% less, so the saving compounds with every turn you take.

**Because Claude Code 2.1.285 turns tool search off by default when** **ANTHROPIC_BASE_URL points anywhere other than Anthropic's own API.** Our default run went through a local server, and it sent all 91 tools, byte for byte the same request as forcing tool search off. The binary's own log message explains the choice: it disables tool search for non-first-party hosts and tells you to set ENABLE_TOOL_SEARCH=true if your proxy forwards tool_reference blocks. If you route Claude Code through a gateway, you may be paying the full tool bill without knowing it.

**Status line:** checked in Claude Code 2.1.285 on 30 September 2026. Re-check with ENABLE_TOOL_SEARCH=true claude -p hi against the script below.

On 30 September 2026 we installed Claude Code 2.1.285 from npm in a clean sandbox and pointed it at a local stand-in server through ANTHROPIC_BASE_URL, in an empty folder with a fresh home directory. For the first-request census, the server recorded the request and returned an error, so no model answered. For the task, **we played the model**: the server replied with a fixed script (search for the tool if it is not loaded, call it, then say "Saved"), so eager and lazy differed only in what Claude Code sent. The memory server really wrote the note. Each mode ran three times; totals varied by at most 3 tokens. We counted tokens with OpenAI's o200k_base tokenizer, the same method as part 1, excluding the user's own prompt.

**Run this and post both numbers in the comments; tomorrow’s part opens with the leaderboard.** Save it as context_tax.py, run python3 context_tax.py, then in your own project run ANTHROPIC_BASE_URL=http://127.0.0.1:8787 ENABLE_TOOL_SEARCH=false claude -p hi, then the same with ENABLE_TOOL_SEARCH=true. It counts the raw request JSON, so its totals run a little higher than ours; the gap between the two lines is what matters. Nothing leaves your machine.

```
# context_tax.py - what your agent sends before you type, eager vs lazy (stdlib only)import http.server, jsondef count(s):    try:        import tiktoken; return len(tiktoken.get_encoding("o200k_base").encode(s))    except Exception: return len(s) // 4  # rough estimate without tiktokenclass H(http.server.BaseHTTPRequestHandler):    def do_POST(self):        j = json.loads(self.rfile.read(int(self.headers["content-length"])))        tools = j.get("tools", [])        live = [t for t in tools if not t.get("defer_loading")]        rest = {k: j.get(k) for k in ("system", "messages")}        s, t = count(json.dumps(rest)), count(json.dumps(live))        print(f"tools loaded {len(live)} ({t:,} tokens) | deferred {len(tools)-len(live)} | TOTAL {s+t:,}", flush=True)        self.send_response(500); self.end_headers()    def log_message(self, *a): passprint("listening on 127.0.0.1:8787", flush=True)http.server.HTTPServer(("127.0.0.1", 8787), H).serve_forever()
```

On our five-server setup it printed 31,233 eager and 11,079 lazy. **Smallest lazy total wins.**

**How do I enable tool search in Claude Code?** Set ENABLE_TOOL_SEARCH=true in your environment or in the env block of ~/.claude/settings.json. Setting it to false forces every tool to load up front.

**Does tool search work with MCP servers?** Yes. In our capture, all 65 MCP tools were deferred and replaced by a 966-token list of their names.

**Does tool search slow Claude Code down?** It adds one round trip the first time a tool is needed. In our task that was one extra request out of three.

**Is tool search on by default?** In 2.1.285 it is not when ANTHROPIC_BASE_URL points to a non-Anthropic host. The binary's message says it is otherwise the default; we could only capture the proxied path.

**What is in part 3?** Subagents: what each one pays before it starts working, and whether three of them cost three times one. It lands tomorrow.

Which MCP server in your setup would you never miss if it loaded only on demand?

*This is The Context Tax, measured: four parts, one a day, each measuring one line on your agent’s bill with the method attached. Part 3 lands tomorrow and asks what a subagent pays before it does any work — follow Decoding AI to get it.*

**If this saved you a few thousand tokens a turn, clap so more developers see it, tell us in the comments which MCP server you would load only on demand and what your two numbers were, and follow Decoding AI — then hit subscribe to get part 3 by email tomorrow.**

[Claude Code Tool Search Nearly Halves Your Context Bill](https://pub.towardsai.net/claude-code-tool-search-nearly-halves-your-context-bill-a14defdf42f3) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
