On this page #
Exposing MCP tools to language models burns 30,000 tokens on startup and turns the model into a single-step processor. Pi, OpenCode, Cloudflare, and Goose replaced tool calls with code.
The industry adopted Model Context Protocol (MCP) as the integration standard of 2025. Direct model exposure turns that standard into an architectural bottleneck.
When I wrote about Pi Agent a few months ago, Mario Zechner avoided MCP entirely. Pi ran on four primitives: read, write, edit, and bash. Then Pi 0.99 shipped with native MCP support.
Pi demoted MCP to an internal runtime library. Through a pattern called Codemode (which Zechner explored in his essay on skipping traditional MCP), Pi moves tool schemas out of the prompt and into an isolated QuickJS sandbox. Instead of emitting JSON function calls, the model writes JavaScript.
OpenCode, Cloudflare, Blockβs Goose, Hugging Faceβs smolagents, and Anthropic converged on the same architecture: Code as Orchestration.
The 30,000-Token Preload Penalty #
Traditional MCP clients connect to a server, call tools/list, and dump every toolβs JSON Schema into the modelβs system prompt before the user types a single word.
Every parameter name, description, enum value, and type definition occupies context space. Connecting multiple MCP servers exhausts the context window:
- Playwright MCP : 21 tools consume**~13.7k tokens** (roughly 6.8% of a 200k context window).
- Chrome DevTools MCP : 26 tools consume**~18.0k tokens** (9.0% of context).
- Cloudflare API : Exposing its 2,500 endpoints as traditional MCP tools requiresover 1.17 million tokens , exceeding the context window of production models.
Traditional MCP Startup Cost:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Playwright MCP (21 tools): 13,700 tokens
Chrome DevTools MCP (26 tools): 18,000 tokens
Two connected servers: 31,700 tokens burned upfront
Cloudflare Full API (2,500 eps): 1,170,000 tokens (impossible)
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
This preload penalty causes tool dilution.
When an LLM faces forty detailed JSON schemas at once, routing accuracy drops. The model confuses similar operations, hallucinates parameter names, and re-reads every schema on subsequent turns. You pay for all forty tools on turn ten, even when the task touches one.
The Single-Step CPU Anti-Pattern #
Traditional MCP breaks control flow.
A standard ReAct (Reasoning + Acting) loop treats the LLM like a clock cycle in a single-instruction processor. Checking 50 pull requests requires fifty separate network turns:
- The model emits a tool call:
github.list_prs(). - The runtime executes the call and injects the raw JSON response into the prompt.
- The model inspects the response and emits the next tool call:
github.get_pr_details({ id: 1 }). - The runtime injects the details.
- The model repeats this sequence 50 times.
// Traditional MCP: 50 full LLM inference round-trips
// Total time: ~150 seconds. Total context: 250,000+ tokens.
for (const pr of prs) {
// Step 1: Model generates JSON tool call -> s
// Step 2: Server executes call -> dumps raw JSON into context
// Step 3: Model reads context -> generates next JSON call
}
Fifty round-trips to an API at two seconds per turn create over two minutes of idle latency. If each payload contains 20KB of metadata, the client injects 1MB of raw JSON into prompt history. That payload stays in context for every subsequent turn. Because the protocol lacks loops, conditionals, and concurrency, every decision requires another model inference call.
Treating a frontier language model as a single-step processor wastes money and adds latency.
Models Understand Code, Not Synthetic Tool Schemas #
In September 2025, Cloudflare engineers Kenton Varda and Sunil Pai pointed out a basic training mismatch in Code Mode: the better way to use MCP:
Language models see trillions of tokens of idiomatic TypeScript, Python, and C. They understand variables, loops, array filtering, and error handling because programmers publish code to GitHub.
Models see almost no native tool calls during pre-training. In-band tokens like <|tool_call|> and JSON-RPC schemas are synthetic formats added during fine-tuning.
In their follow-up analysis on giving agents an entire API in 1,000 tokens, Cloudflare proved that replacing 2,500 endpoint declarations with two code tools cut token usage by 99.9%.
Training data comparison:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Real-world code (TS, JS, Python): Trillions of tokens
Synthetic tool-call JSON schemas: Fraction of a percent
Model comfort level: Prefers code over JSON
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Asking an LLM to orchestrate work through sequential JSON dictionaries instead of a clean script forces the model to work against its training. Given a TypeScript API, the model already knows how to branch, map, and catch errors.
Orchestration Inside the Sandbox #
Instead of individual tools into the model prompt, Codemode presents a minimal execution interface: the model writes a script, and a local sandbox runs it.
Consider the difference when querying fifty issues across GitHub and checking their labels:
// What the model writes in Codemode:
const issues = await tools.github.list_issues({ repo: "org/project", state: "open" });
// Concurrency executed inside sandbox RAM in 80ms:
const detailed = await Promise.all(
issues.slice(0, 50).map(i => tools.github.get_issue({ id: i.id }))
);
// Filter 1MB of raw JSON down to three lines before the model sees it:
return detailed
.filter(d => d.labels.includes("urgent"))
.map(d => ({ id: d.id, title: d.title, author: d.user.login }));
The 1MB JSON response stays in sandbox memory. The model never reads the raw payload; it receives only the three urgent issues.
The task completes in one model turn and a fraction of a second of runtime execution, rather than fifty serial inference calls.
How Sandboxes Clean the Data
Traditional tool calling treats the prompt as an unmanaged heap where every intermediate response accumulates. Sandboxes act as garbage collectors:
- Projection and Field Stripping: A standard GitHub PR payload contains over 50 fields, including hypermedia URLs, user metadata, and commit hashes. A single
.map()call strips 98% of these keys before serialization. - In-Memory Predicate Filtering: The sandbox evaluates
.filter()conditions in WebAssembly or native C memory at microsecond speeds. 47 non-urgent items vanish from RAM before touching the network serialization layer. - Local Joins and Aggregation: When a task spans two services, such as cross-referencing Jira issues with GitHub PRs, traditional agents send both datasets into the prompt for the model to join. A Codemode script joins the datasets locally with a standard hash map and returns only the final matches.
- Credential and PII Redaction: Sensitive values like internal IPs, session cookies, and authorization headers remain inside live runtime bindings. The script can use them to make calls without exposing them in the returned text.
Three Implementations, One Architecture #
Different teams arrived at this solution through different constraints, but their designs look remarkably similar under the hood.
1. Pi 0.99 (Mario Zechner)
Pi embeds a C-based QuickJS sandbox with a strict 256MB memory cap. Described in the Pi Codemode documentation and the pi codebase, the design follows three rules:
- Deferred Schemas: MCP tools do not appear in the system prompt. Instead, tools share an inline budget of 3,000 estimated tokens.
- Dynamic Search: The script discovers tools on-demand using BM25 relevance ranking via
searchTools(query)ordescribeTool(name). - Zero Host Access: The QuickJS environment has no Node APIs, no filesystem access, and no ambient network. It communicates strictly through injected tool callbacks.
2. OpenCode (@opencode-ai/codemode)
The team behind OpenCode built an Effect TS runtime for safe orchestration:
- The 2,000-Token Budgeted Catalog: OpenCode allocates a fixed budget of 2,000 estimated tokens for tool signatures in the prompt. To keep multiple servers fair, it uses round-robin allocation across namespaces.
- Deterministic Field Search (
tools.$codemode.search): When tools exceed the budget, the agent queries an internal search function with field-weighted scoring: exact path match (20 pts), path substring (8 pts), description match (4 pts), and parameter names (2 pts). - Supervised Fibers: Tool calls run on Effect fibers with a ceiling of 8 concurrent calls. If the model generates an infinite loop (
while (true)), the fiber supervisor interrupts the execution cleanly without hanging the agent.
3. Cloudflare Code Mode (Varda, Pai, Carey)
Cloudflare collapsed their entire 2,500-endpoint API down to two tools: search(code) and execute(code). Documented in the Cloudflare Agents SDK:
- The Token Drop: Replaced an estimated1.17 million tokens of raw OpenAPI schemas with roughly1,000 tokens of meta-tool instructions.
- V8 Isolates: Code executes inDynamic Worker s , which are disposable V8 isolates that boot in under 5 milliseconds.
- Live Object Bindings: The sandbox holds no raw API keys. The parent worker injects authenticated live bindings at runtime.
The Industry Convergence #
This pattern is not an isolated experiment. Across the AI engineering landscape, team after team has discarded direct tool-calling in favor of code execution.
| Project / Framework | Orchestration Runtime | Discovery Mechanism | Token Reduction |
|---|---|---|---|
| Pi 0.99 (Mario Zechner ) | QuickJS VM (256MB cap) | BM25 searchTools() + 3k inline budget |
~95% vs raw schemas |
| OpenCode (anomalyco ) | Effect TS supervised fibers | Budgeted catalog (2k tokens) + field search | 90%+ on multi-server setups |
| Cloudflare (Varda & Pai ) | V8 Worker Isolates | search(code) over OpenAPI spec |
99.9% (1.17M to 1k tokens) |
| Goose (AAIF / Block ) | pctx (Deno sandbox) |
3 meta-tools ( list ,details ,execute ) |
Context capped at 3 tools |
| smolagents (Hugging Face ) | E2B / Docker / Local VM | Python imports ( from tools import ... ) |
30% fewer steps on GAIA |
| Anthropic (Claude SDK ) | Container / Worker Sandbox | Virtual filesystem ( ls ./servers/ ) |
98.7% (150k to 2k tokens) |
The academic literature predicted this transition. In February 2024, researchers from UIUC, Apple, and UC Berkeley published the CodeAct paper (βExecutable Code Actions Elicit Better LLM Agentsβ, arXiv:2402.01030).
Agents using executable code actions outperformed JSON tool-calling agents by up to 20% on complex tasks. That research formed the foundation of OpenHands.
Hugging Face reached the same conclusion when designing smolagents. By making CodeAgent the default and relegating ToolCallingAgent to simple prototyping, they proved that agents writing Python code solve multi-step reasoning problems with 30% fewer turns than agents passing JSON objects.
CodeAct vs. Codemode: Two Paths to Code-First Agents #
Generalist agents like Manus and OpenHands popularized code execution, but their execution model differs from Codemode.
Both patterns write code instead of JSON. The difference lies in the execution environment and the boundary:
| Dimension | CodeAct ( Wang et al., 2024 ) / Manus | Codemode (Cloudflare, Pi, OpenCode, Goose) |
|---|---|---|
| Target Problem | General computer use and multi-step file tasks | Distributed tool composition without context bloat |
| Sandbox Type | Heavyweight Linux VM or persistent Docker container | Disposable micro-runtime (QuickJS, V8 Isolate, Deno) |
| Tool Interface | Shell commands ( bash ), Python scripts, browser sessions |
Typed virtual SDK functions ( tools.github.list_prs() ) |
| System Authority | Ambient filesystem, package managers ( pip ), full network |
Zero ambient filesystem, no external fetch, bound RPC |
| Lifecycle | Long-running container persisting across the task | Ephemeral execution (sub-5ms boot, destroyed on return) |
Manus and OpenHands rely on CodeAct: giving a model a full Linux virtual machine where it writes Python and Bash to browse websites, install packages, and debug tracebacks. CodeAct answers: How does an agent perform general computing work?
Codemode addresses a different problem: How does an agent call dozens of distributed tools and remote APIs without overflowing its context window?
Codemode does not give the agent a Linux terminal or root access. It gives the model an ephemeral, memory-capped execution engine where external APIs and MCP servers appear as typed system calls.
Compilers, CPUs, and Syscalls #
MCP belongs in the infrastructure layer, not the prompt.
A resilient agent architecture separates responsibilities into three computing layers:
- The LLM is the Compiler: It translates natural language from the user into an executable plan (JavaScript, TypeScript, or Python).
- The Sandbox is the CPU: A lightweight engine (QuickJS, a V8 isolate, or Deno) executes the plan. It handles control flow, loops, memory allocation, and concurrency locally.
- MCP is the Syscall Layer: MCP provides standardized, authenticated RPC endpoints to the outside world. It acts as the driver layer between the sandbox and external services.
Traditional MCP forces the compiler to act as a CPU, halting after every instruction to inspect state. Codemode lets the compiler emit the program once, the CPU run it in milliseconds, and MCP provide the system calls.
Where Code Orchestration Goes Next #
The shift from direct tool calls to code orchestration points toward four developments:
1. Compile Once, Run Deterministically
Current agents write new orchestration scripts for every prompt. As workflows stabilize, agents will verify scripts once and save them as reusable skills. If an agent checks production deployments every morning, it compiles the script on day one. On subsequent mornings, the host runs the script deterministically without calling a frontier model, eliminating inference costs.
2. Sub-Millisecond WASI Micro-Engines
Docker containers take seconds to boot and consume hundreds of megabytes of RAM. As agents scale to thousands of concurrent subtasks, execution is moving to WebAssembly (WASI) and V8 isolates that instantiate in under two milliseconds, execute in memory, and clean up immediately.
3. Capability-Based Security and Live Bindings
Passing raw API tokens into prompts creates prompt injection and data exfiltration risks. Modern runtimes replace raw credentials with live object bindings. The host provides an authorized client instance inside the sandbox. The model invokes methods on the binding, but the script cannot read, print, or leak the underlying authorization secret.
4. Tiered Compiler-Runtime Routing
Frontier reasoning models act as the system architect, compiling ambiguous human instructions into typed TypeScript plans. Smaller, faster edge models inspect execution logs, handle network retries, and format final user summaries.
Moving Past Direct Tool Calling #
Model Context Protocol solved a real problem by creating a vendor-neutral standard for external tools and authentication.
Exposing that protocol to the model prompt created three bottlenecks: prompts bloated with unused schemas, slow ReAct loops, and context flooded with raw payloads.
Codemode fixes the abstraction. Placing a lightweight sandbox between the model and the protocol turns tools into code libraries, runs loops in memory, and cuts token overhead by 98% or more.
If you build agents, stop injecting raw tool schemas into system prompts. Give the model a compiler target, execute the script in a sandbox, and treat MCP as an operating system call.
If you build coding agents or manage tool execution in production, reach out on LinkedIn.