{"slug": "pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache", "title": "Pi 1.0 explained: how the agent harness loop, Codemode, deferred tools and cache warming save tokens", "summary": "Earendil shipped Pi 1.0 on October 1, 2026, a free MIT-licensed agent harness whose Codemode, deferred tool loading, mid-conversation system messages and cache warming features are designed to cut token costs. Pi's docs state the harness defaults to four tools (read, bash, edit, write), keeps MCP tool definitions out of every request until the model searches for one, and resends a one-token-capped request just before the 5-minute prompt cache expires only when it expects to save at least 5 cents. Pi's docs also say it does not ask for approval before every tool call and has no built-in sandbox, unlike Codex CLI and Copilot CLI, which ask or sandbox by default.", "body_md": "# Pi 1.0 explained: how the agent harness loop, Codemode, deferred tools and cache warming save tokens\n\nEarendil shipped Pi 1.0, a free MIT-licensed agent harness, with a short post that names each new feature in one line. This page reads Pi's docs and source code to explain what a harness does, how each feature works, what it saves, and how Pi's defaults differ from Claude Code, Codex CLI and Copilot CLI.\n\n**This explains reporting by**\n\n[Earendil, Pi 1.0 (October 1, 2026)](https://earendil.com/posts/pi-1-0/).\nRead the original first:\n\n[https://earendil.com/posts/pi-1-0/](https://earendil.com/posts/pi-1-0/)\n\n## In one minute\n\n- An agent harness is the program that runs the loop: send the model a request, run the tools it asks for, send back the results, repeat. Every trip resends the prompt, so the harness decides most of your token bill.\n- Codemode lets the model write one small JavaScript script that calls many tools. Only the script's output goes back to the model, so large tool results never fill the context.\n- Deferred tool loading keeps MCP tool definitions out of every request until the model searches for one. MCP tools default to Codemode only, so the model never sees their definitions at all.\n- Mid-conversation system messages append prompt and tool changes to the end of the transcript instead of editing the top. That keeps the earlier part of the prompt cached.\n- Cache warming resends the exact request with a one-token reply cap just before the cache expires, but only when Pi expects to save at least 5 cents.\n- Pi's docs say it does not ask for approval before every tool call and has no built-in sandbox. Codex CLI and Copilot CLI ask or sandbox by default. Run Pi in a container if it can reach anything that matters.\n\n## What an agent harness actually does\n\nA model on its own only reads text and writes text. It cannot open a file or run a test. The harness is the program around it that can.\n\nPi's docs describe the loop plainly. Pi builds a request from the system prompt, the conversation so far, the list of available tools, and the model settings. It sends that to the provider. The reply can hold text and tool calls. Pi runs each tool call, records the results, and if there is more work, sends a new request. That is one turn, and a task can take dozens.\n\nTwo facts follow from that loop, and every Pi 1.0 feature is about them.\n\n- Every turn resends everything: instructions, every tool definition, and the whole history. A tool definition is a block of text describing a tool and its inputs. Twenty tools can cost thousands of tokens on every single turn.\n- Providers discount text they have seen recently. This is the prompt cache. If the start of a request matches a recent request exactly, that part is billed at a fraction of the normal price.\n\nPi starts with four tools by default: read, bash, edit and write. Everything else is opt-in. That small default is the minimalism Earendil talks about.\n\n## How the prompt cache works, and why order matters\n\nAnthropic's docs give the clearest rules, and Pi's cache features are built around them.\n\n- The cache is a prefix. It covers the request in this order: tools, then system prompt, then messages. A change at one level breaks the cache for that level and everything after it.\n- Changing a tool definition breaks the cache for tools, system and messages. In other words, the whole request.\n- A cache entry lasts 5 minutes by default and resets each time it is used. A 1 hour option exists.\n- Writing to the 5 minute cache costs 1.25 times the normal input price. Reading from it costs 0.1 times on most Claude models.\n\nSo the cheapest agent is one whose request only ever grows at the end. Anything that edits the top, such as adding a tool halfway through, makes the next turn pay full price for the entire history.\n\n## Codemode: many tool calls, one small answer\n\nWithout Codemode, each tool call is a round trip. The model asks for one call, the full result goes into the history, and the model reads it on every later turn.\n\nWith Codemode, the model writes a JavaScript script instead. Pi runs it in a QuickJS sandbox, a small JavaScript engine, inside the harness. The script has no file system, network or timers of its own. It reaches the outside only through Pi's tools and models.\n\nOnly the script's output reaches the model, capped at 10,000 tokens by default. A script can run calls in parallel, filter a large result down to the few rows that matter, and save small values with store() for later scripts.\n\nEarendil's own demo shows the scale. One script pulled 167 open issues from a Linear MCP server, fetched comments for each, and ran a classifier on every thread. The transcript shows more than 330 tool calls. The model got back one short list of 11 flagged issues.\n\n1.0 also made Codemode cheaper to have switched on. Earendil says that with the default tools and Codemode active, a GPT-5.6 request shrank from about 5,300 to 3,300 prompt tokens.\n\nOne thing the sandbox does not do: it limits the script, not the tools the script calls. Pi's docs say tool calls are real and are not undone if the script fails. A script that calls bash runs bash with your account's permissions.\n\n## Deferred tool loading: tools stay out until needed\n\nEach MCP server in Pi gets an exposure setting that decides how the model reaches its tools.\n\n- codemode (the default): callable from scripts, but never declared to the model. Scripts find tools with searchTools() or describeTool().\n- deferred: not declared until the tool_search tool loads a match for the next request.\n- direct: declared like a built-in tool on every request. Meant for small, often-used tool sets.\n- hidden: registered but unreachable.\n\nThe saving is simple. A tool that is not declared costs zero tokens per turn. A server with 50 tools on direct exposure costs its full definitions on every turn of every task.\n\nThere is a cache detail too. Tools that tool_search loads are recorded in the transcript and stay declared on that branch. And the Codemode tool description does not list deferred tools, so it stays the same while MCP servers connect. A description that changed every time a server came online would break the cache each time.\n\nYou can mix exposures per tool with toolExposure. For example:\n\n```\n{\n  \"mcpServers\": {\n    \"github\": {\n      \"url\": \"https://example.com/mcp\",\n      \"exposure\": \"deferred\",\n      \"toolExposure\": {\n        \"search_code\": \"direct\",\n        \"get_*\": \"codemode\",\n        \"delete_*\": \"hidden\"\n      }\n    }\n  }\n}\n```\n\nClaude Code does a version of this too. Its docs say MCP tool search, which loads tool definitions on demand, is on by default from v2.1.221.\n\n## Mid-conversation system messages: change the setup without breaking the cache\n\nPi records the starting prompt and tool set in the first system message of the transcript. Later system messages can append instructions, replace or remove named prompt sections, and add or remove tools. Replaying them in order gives the current setup.\n\nThis matters because of the prefix rule above. If you add a tool by editing the tool list at the top, the cache is gone for the whole history. If you append the change at the end, everything before it still matches and stays cached.\n\nPi uses this itself. When an MCP server connects and its summary becomes available, Pi appends the new server section to the conversation instead of changing tool declarations, so earlier messages stay cached.\n\nPi's docs add one limit. Providers that cannot represent the change get a complete transcript checkpoint instead, which can invalidate the cached prefix. The docs do not list which providers those are.\n\n## Cache warming: paying a little to avoid paying a lot\n\nThe cache entry expires 5 minutes after last use. An agent can easily sit longer than that: a long test run, a slow build, or you reading the output. The next request then pays the full write price for the whole history.\n\nPi's cache-warmer.ts, read from the v1.0.0 source, works like this:\n\n- At 90 percent of the cache lifetime, keeping at least 10 seconds of margin, Pi resends the exact last request with the reply capped at one token. That counts as a cache read, which resets the 5 minute clock.\n- It only does this when the model declares a cache lifetime, and when probability times miss cost, minus warm cost, is at least $0.05.\n- During an active run the probability is set to 1. In idle mode, between runs, it is set to 0.15, which Earendil says it measured from its own usage.\n- streaming (the default) warms only during active runs, for at most 60 minutes after the real request. idle also warms between runs, for at most 30 minutes. off turns it off. It is a global setting only.\n- It skips Claude models that use budget-based thinking, because the one-token replay would change the thinking budget, which Anthropic keys the cache on.\n\nWorked example, using an example price of $3 per million input tokens and a 100,000 token prompt. The normal price of that prompt is 30 cents. A warm refresh is a cache read at 0.1 times, about 3 cents. A miss means a cache write at 1.25 times, 37.5 cents, which is 34.5 cents more than a read.\n\n- Active run: 1 times 34.5 cents, minus 3 cents, is about 31.5 cents saved. Pi warms.\n- Idle: 0.15 times 34.5 cents, minus 3 cents, is about 2.2 cents. That is under 5 cents, so Pi does not warm.\n\nWith those same assumptions, active-run warming pays from roughly 16,000 prompt tokens, and idle warming only from roughly 230,000. Your model's real prices change those numbers. Warm-up spending counts toward your session cost, and showCacheMissNotices shows when it happens.\n\n```\n{\n  \"cacheWarming\": \"streaming\",\n  \"showCacheMissNotices\": true\n}\n```\n\n## How Pi's defaults compare, from each tool's own docs\n\nThis is not a quality comparison. No independent benchmark puts these four on the same task. It covers license, model choice and what each does before running a command, because those decide whether a tool fits your setup.\n\n- Pi: MIT license. Its provider docs list more than 30 providers, including Anthropic, OpenAI, Google Gemini, GitHub Copilot, OpenRouter and Amazon Bedrock. Its security page says it does not ask for approval before every tool call and has no built-in sandbox. Approval prompts come from extensions you add.\n- Claude Code: from Anthropic, under Anthropic's commercial terms rather than an open source license. It runs Claude through Anthropic, Amazon Bedrock, Google Cloud or Microsoft Foundry. In Manual mode it asks before shell commands, except a set of read-only ones, and before file edits.\n- Codex CLI: Apache-2.0 license, from OpenAI. Its docs say that by default network access is off and writes are limited to the current workspace, enforced by the operating system.\n- Copilot CLI: from GitHub, needs an active Copilot subscription, and ships under GitHub's own license. Its README says nothing happens without your explicit approval, and it comes with GitHub's MCP server.\n\nPi's own docs are direct about what this means. Pi can read, change and run files with the permissions of the account that started it. Watching the transcript or reviewing changes is not a security boundary. A container or virtual machine is usually the strongest practical option.\n\n## Who should try it, and who should wait\n\n- Try it if you switch between model vendors and want one harness for all of them.\n- Try it if you run many MCP servers and pay for their tool definitions on every turn.\n- Try it if you want to build your own agent app. Pi ships a TypeScript SDK, an RPC mode over stdin and stdout, and a JSON event mode.\n- Wait, or run it in a container, if the machine holds client credentials, production access or anything you cannot restore. Pi does not ask before each tool call unless you add an extension that does.\n- Wait if you need approvals and sandboxing to be on from day one without writing an extension. Codex CLI and Copilot CLI start there.\n\n## Who is affected\n\n| Case | Status | \n|---|---|\n| Heavy MCP users | Biggest likely saving. MCP tools default to Codemode exposure, so their definitions leave every request. | \n| Long sessions on Anthropic models | Cache warming is on by default during active runs. Claude models with budget-based thinking are skipped. | \n| Anyone running Pi on a machine with real credentials | No approval prompt and no sandbox by default. Use a container, VM or a confirm extension. | \n| Builders of agent apps | SDK, RPC and JSON modes in Pi 1.0. Pi Durable, for long-running apps, is a separate experimental package. | \n\n## What to do\n\n- Install Pi in a throwaway folder, or better, a container: curl -fsSL https://pi.dev/install.sh | sh, or npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Node.js 22.19 or newer). Run pi, then /login.\n- Give it one real task you ran in your current agent this week. Use /session to see cost per model and cache use, and compare the bills.\n- Turn on showCacheMissNotices so you can see cache misses and warm-ups as they happen.\n- If you use MCP, leave servers on the default codemode exposure. Set only small, often-used tools to direct with toolExposure.\n- Before you give it anything real, add a confirm extension for risky commands or run it in a container, as Pi's own security page says.\n\n## What is still unknown\n\n- \"Hundreds of thousands of people every week\" is Earendil's own figure. npm shows about 4.5 million downloads of the Pi coding agent package from September 24 to 30, but downloads count installs and updates, not people.\n- The drop from about 5,300 to 3,300 prompt tokens is Earendil's measurement, on one model with default tools. Your tools and model will give a different number.\n- The 15 percent chance you come back while idle is from Earendil's own usage data. If you come back more often, idle warming is worth more than Pi assumes.\n- Pi's docs do not list which providers can take mid-conversation system messages directly, and which get a full checkpoint that can break the cache.\n- No independent test was found comparing Pi's token bill or task quality with Claude Code, Codex CLI or Copilot CLI on the same work.\n- The worked numbers on this page use an example price of $3 per million input tokens. They are math on Anthropic's published cache multipliers, not measured bills.\n\n## Sources\n\n[AI News Report](https://theainewsreport.com/)· every headline, every morning.", "url": "https://wpnews.pro/news/pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache", "canonical_source": "https://theainewsreport.com/2026-10-02-pi-1-0-agent-harness-codemode-deferred-tools-cache-warming-explained.html", "published_at": "2026-10-02 14:31:39+00:00", "updated_at": "2026-10-02 14:38:48.984349+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "large-language-models", "ai-infrastructure"], "entities": ["Earendil", "Pi 1.0", "Claude Code", "Codex CLI", "Copilot CLI", "Anthropic", "MCP", "Codemode"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache", "markdown": "https://wpnews.pro/news/pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache.md", "text": "https://wpnews.pro/news/pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache.txt", "jsonld": "https://wpnews.pro/news/pi-1-0-explained-how-the-agent-harness-loop-codemode-deferred-tools-and-cache.jsonld"}}