Agent Codemode Explained Cloudflare's Code Mode MCP server exposes its entire API of more than 2,500 endpoints in roughly 1,000 tokens by offering just two operations — search for an API surface and execute code against the discovered typed API — versus an estimated 1.17 million input tokens if the same API were represented as ordinary MCP tool definitions. The approach, also appearing in Anthropic's programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and Pi as of September 29, 2026, moves agent control flow out of the LLM so intermediate computation such as filtering, retries, loops, and parallel calls stays out of model context. The design matters because each LLM round-trip in the standard one-call-at-a-time tool loop adds inference cost and accumulates tool results in the context window. Agent Codemode Explained Code Mode: Moving Agent Control Flow Out of the LLM Most agents still use tools one call at a time. The model decides to call a tool. The runtime executes it. The result is appended to the conversation. The model runs again and decides what to do next. For a short task this works well. For a task that needs twenty API calls, filtering, retries, a loop, or several dependent operations, the design becomes expensive: LLM ↓ tool A ↓ LLM ↓ tool B ↓ LLM ↓ tool C ↓ LLM Each edge through the LLM costs inference time. Tool results also accumulate in the context even when most of the data is only needed to decide the next call. Code Mode changes this boundary. Instead of asking the model to emit every tool call, give it one tool that executes a small program: LLM ↓ program ├─ tool A ├─ tool B ├─ tool C └─ local filtering ↓ LLM The program becomes the control plane for a section of the agent run. This idea now appears in Cloudflare Code Mode, Anthropic’s programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and, as of September 29, 2026, Pi. They share the same basic observation: Tool calling is an unusually bad programming language. JSON tool calls can express: call X with Y A normal programming language can express: call X for every Y call X and Z concurrently retry X if it fails filter the result call Z only if condition Q is true store this value and reuse it later return only these five fields The difference is larger than syntax. Code Mode changes where computation happens. The normal tool loop Suppose an agent needs to find all unhealthy services and inspect their recent deployments. With normal tool calling it might do: LLM → list services LLM ← 300 services LLM → get health service 1 LLM ← result LLM → get health service 2 LLM ← result ... A capable model can issue several calls in parallel, but the model still has to construct the calls and receive their results. A better tool might support bulk operations. But now every tool designer has to predict every future composition: get health for services ... get unhealthy services ... get unhealthy services with deployments ... This is where Code Mode becomes interesting. The model can instead write something equivalent to: services = await list services health = await gather get health {"id": s "id" } for s in services bad = s for s, h in zip services, health if h "status" = "healthy" deployments = await gather get deployments {"service id": s "id" } for s in bad return { "service": s "name" , "deployment": d 0 } for s, d in zip bad, deployments The LLM sees the task once and writes the control flow once. The intermediate 300 service records and hundreds of health responses do not need to become conversation history. That is the important part. Code Mode is not mainly a prettier way to call tools. It is a way to move intermediate computation out of model context . Cloudflare: code as a compact API plan Cloudflare has probably made the clearest argument for Code Mode. Its problem is unusually visible. The Cloudflare API has more than 2,500 endpoints. Exposing every API endpoint as an MCP tool means shipping a huge collection of JSON schemas into the model context before the agent does anything useful. Cloudflare’s Code Mode MCP server instead exposes essentially two operations: search for an API surface, then execute code against the discovered typed API. Cloudflare reports that its whole API can be exposed in roughly 1,000 tokens this way. Its comparison estimates 1.17 million input tokens if the equivalent API were represented as ordinary MCP tool definitions. This solves the first Code Mode problem: too many tools There is a second problem: too much intermediate data Cloudflare describes generated code as a compact plan. Calls, filtering and transformations happen inside the execution environment, and only selected output returns to the model. Their current interface exposes primitives such as: codemode.search ... codemode.describe ... codemode.step ... codemode.run ... search performs progressive discovery. describe loads the detailed interface only when needed. step records work that needs replay semantics. run executes reusable snippets. This is an important refinement. A naive Code Mode implementation still puts every available function declaration in the system prompt: one code tool + 10,000 function declarations That fixes intermediate results but not tool-definition cost. Cloudflare instead makes the API itself lazy. The model starts with a small discovery interface and pulls schemas when it needs them. Conceptually: ┌──────── search LLM → code│ ├──────── describe │ ├──────── API A ├──────── API B └──────── API C The tool catalog becomes something closer to a filesystem or symbol table than a prompt. That is likely the right abstraction for very large tool spaces. Cloudflare’s sandbox matters Once the model can write arbitrary JavaScript, eval is not an acceptable runtime. Cloudflare executes generated code in isolated Workers. Its MCP design blocks direct outbound access by default; generated code reaches the outside world through capabilities supplied by the host. Credentials remain outside the generated program. The distinction is useful: code execution capability ≠ external authority The sandbox can be computationally expressive while having almost no ambient authority. The program may know how to write: await dns.deleteZone ... but the actual authority still belongs to the host connector. This means authorization does not have to be implemented by trusting generated code. It remains at the tool boundary. Cloudflare’s durable runtime goes further. It records execution history and can pause generated code for approval, then replay completed work and continue the program. That turns Code Mode from “run some generated JavaScript” into an agent runtime primitive. Anthropic: programmatic tool calling Anthropic independently arrived at almost exactly the same model. Their current term is programmatic tool calling . Claude writes Python inside its code execution environment. Tool calls from the program cross back to the host. When a tool returns, execution resumes. Intermediate results remain inside the execution environment instead of being inserted into Claude’s context. The execution model looks like this: Claude generates Python │ ▼ sandbox │ ├── tool │ │ │ └──── host executes tool │ ├── process result locally │ ├── tool │ └── print small result │ ▼ Claude continues This is not equivalent to asking Claude to generate a Python script and running the script after the model finishes. The interesting part is the suspended host call. The Python program is effectively an orchestration language around external capabilities. Anthropic reports useful measurements here. On its BrowseComp and DeepSearchQA experiments, programmatic tool calling improved performance by an average of 11% while using 24% fewer input tokens. On an internal 75-tool project-management benchmark it reports roughly 38% lower billed input tokens with unchanged task accuracy. But on τ²-bench tasks dominated by one or two sequential calls, it cost about 8% more. That last number is important. Code Mode is not free. For: call tool return answer generating a program and starting a code environment is extra work. Code Mode becomes useful when the code replaces enough model-mediated control flow to pay for itself. Pi: Code Mode moves into the harness Pi is a more interesting implementation for coding agents because Code Mode is not only an API optimization. It is integrated into the agent harness. A large Pi change was merged on September 29, 2026. The implementation runs model-generated JavaScript in a QuickJS VM compiled to WebAssembly. The script can call Pi tools as async functions. Nested calls do not enter the LLM context individually; only explicit script output and the final return value do. The standalone package describes the capability directly: js const source = await tools.read { path: "package.json" } text JSON.parse source .name Every execution gets a QuickJS VM in a worker. The VM has its own WASM memory and no general route back into the Node host except the explicitly installed host-call bridge. This makes Pi’s implementation worth studying. Pi already had very powerful tools: read write edit bash For coding work, bash alone is close to universal. So why add Code Mode? Because universality is not the same as a good agent interface. Consider: cat package.json | jq ... versus: js const p = await tools.read {path: "package.json"} const pkg = JSON.parse p Both work. But the second form can combine normal agent tools, MCP tools and future host capabilities under one runtime, without turning every operation into shell text. More importantly, nested calls remain visible to the harness. Pi routes nested calls through the normal validation and tool hooks. They carry a parentToolCallId , and nested calls are recorded on the parent tool result. This gives a useful structure: codemode call │ ├── read ├── grep ├── MCP: search └── bash instead of hiding the entire operation behind an opaque process. That matters for observability, permissions and evaluation. Pi’s tool exposure model Pi also exposes a design problem that smaller Code Mode demos can avoid: Which tools should the model see directly? Pi now distinguishes different forms of tool exposure and supports two useful Code Mode configurations. In codemode.mode = "on" , normal tools may remain directly visible while Code Mode can also call them. In codemode.mode = "only" , the model sees Code Mode as the primary interface and calls ordinary tools through generated code instead. This gives two different agent architectures. Hybrid LLM ├── read ├── edit ├── bash └── codemode ├── read ├── edit └── MCP tools Code-only LLM └── codemode ├── read ├── edit ├── bash ├── MCP └── extensions The second one is more radical. The model-facing action language stops being a collection of JSON tools. It becomes JavaScript. Pi also places a token budget on inline declarations. Tools that do not fit can be found dynamically through tool search. So Pi is converging on the same two-layer structure as Cloudflare: small language runtime + lazy capability discovery This looks increasingly less like an agent feature and more like an operating-system interface. PydanticAI: Python as the capability language PydanticAI implements the same pattern using Python and Monty. Eligible tools disappear from the direct model-facing tool list and become functions available inside one run code tool. The generated Python can use loops, conditions, variables and asyncio.gather . Intermediate data stays in the sandbox. For example: rows = await search "failed builds" details = await asyncio.gather get build row "id" for row in rows return x for x in details if x "status" == "failed" PydanticAI’s implementation is particularly interesting because Monty is not CPython running in a container. Monty is a restricted Python interpreter written in Rust for AI-generated programs. By default it has no normal filesystem, environment variables or network access. Host capabilities have to be passed in deliberately. That gives a capability model similar to Pi’s QuickJS sandbox: generated program │ ├── computation ├── variables ├── branching └── injected host functions The model gets the useful parts of Python without automatically inheriting the authority of the host Python process. PydanticAI also supports persistent REPL state during an agent run. Variables, imports and helper functions can survive across run code calls. That changes Code Mode again. It is no longer just: LLM → program → result It can become: LLM → persistent computational workspace That distinction will matter for long-running agents. smolagents: code as the action format Hugging Face’s smolagents predates some of these systems and takes the idea further at the agent-loop level. Its CodeAgent uses Python code as the normal action representation. ToolCallingAgent is the alternative that emits conventional structured tool calls. The important observation from smolagents is simple: JSON is a serialization format. A programming language is a control-flow language. If the task is: search A search B take the intersection fetch each result rank them representing that as a program is natural. Representing it as five separate JSON actions requires the LLM itself to become the control-flow interpreter. That is the hidden cost of ordinary agent loops. Code Mode is a tiny compiler boundary One way to understand these systems is to stop thinking about “code execution.” Think of the LLM as producing a program for a small virtual machine. The traditional agent loop is: observation ↓ LLM ↓ one action ↓ observation Code Mode becomes: observation ↓ LLM ↓ program ↓ runtime ↓ many actions ↓ compressed observation The LLM has effectively compiled several future decisions into one artifact. This means Code Mode is useful only when those future decisions can be expressed algorithmically. Suppose a result requires semantic judgment: Read this incident report. Understand what probably caused the outage. Choose what to investigate next. A JavaScript if statement cannot replace the next LLM call. But many agent decisions are not semantic reasoning: for every item if status == failed take first five sort by date retry twice call these concurrently extract this field Today we often waste LLM calls performing those operations. Code Mode removes them from the model loop. The real token saving has two parts Discussions of Code Mode often combine two independent optimizations. They should be separated. 1. Tool definition compression Instead of: tool A schema tool B schema tool C schema ... tool 2500 schema give the model: search query describe symbol execute code Then discover interfaces lazily. This is the part that lets Cloudflare expose thousands of endpoints in a small fixed context. Anthropic described a similar MCP approach where tool definitions are discovered through files instead of loading the whole MCP catalog into context. In one example, it reduced tool-related context from roughly 150,000 tokens to about 2,000. 2. Intermediate result compression Instead of: tool result → LLM → tool result → LLM → tool result → LLM do: tool result → program → tool result → program → filter → final small result → LLM These solve different scaling problems. A serious Code Mode implementation should probably do both. Code Mode does not remove the agent loop A tempting conclusion is: Why not put the entire agent in code? Because generated code only knows what the model knew when it wrote the program. Consider: result = await investigate If result contains something genuinely surprising, deciding what it means may require another model inference. A normal program can branch on structure: if result "status" == "failed": It cannot reliably branch on arbitrary semantics: if result implies that the database migration is probably unrelated to the outage: You can put another model behind a function call: if await classify result, question : but now the model has returned through another door. So the likely architecture is not: Code Mode replaces Agent Loop It is: Agent Loop │ ├── semantic decision │ └── Code Mode ├── deterministic control flow ├── data processing ├── bulk tool calls └── cheap decisions The boundary between these two layers is one of the more interesting agent-design problems. Code Mode also changes tool design Today’s tool APIs are partly shaped by the weakness of model tool calling. We create high-level tools such as: search and summarize incidents find recent failed deployments get customer context because asking the model to compose eight primitive tools is expensive and unreliable. Once tools can be composed inside code, smaller primitives become practical: search fetch read write query execute The generated program provides the composition. This does not mean every tool should become low level. Every tool boundary is also a: permission boundary validation boundary observability boundary failure boundary But Code Mode changes the trade-off. Without Code Mode, coarse tools save LLM turns. With Code Mode, coarse tools need to justify themselves for another reason. The security boundary is not the language sandbox Running generated code in QuickJS, Monty or a Dynamic Worker is necessary. It is not sufficient. The dangerous operation is usually not: while true {} It is: await tools.deleteDatabase ... The generated program should therefore be treated as untrusted orchestration around trusted capabilities. The useful architecture is: generated code │ sandbox boundary │ ▼ capability bridge │ ┌──────────────┼──────────────┐ ▼ ▼ ▼ read tool GitHub tool DB tool │ │ │ policy policy policy │ │ │ execute execute execute Permissions belong below the code. Pi’s nested calls pass through normal tool validation and hooks. Cloudflare explicitly says generated code does not replace authorization and that permissions should still be enforced in upstream handlers. Anthropic similarly warns that its tool caller configuration should not itself be treated as a security boundary. These implementations are converging on the same rule: sandbox computation; authorize capabilities. Observability becomes harder Direct tool calling gives a convenient trace: model tool model tool model tool Code Mode compresses the model trace: model codemode model That is good for tokens. It can be bad for debugging if the runtime treats the code invocation as one opaque tool call. A useful Code Mode runtime therefore needs nested traces: codemode 42 ├─ search 43 │ └─ 42 ms ├─ fetch 44 │ └─ 181 ms ├─ fetch 45 │ └─ 172 ms └─ update 46 └─ approval required Pi explicitly records nested calls and their parent IDs. Cloudflare’s durable runtime records steps and replay information. PydanticAI exposes nested Code Mode calls in its tracing and can integrate them with durable execution systems. This is not optional infrastructure. Once a generated program can call fifty tools, “the Code Mode tool failed” is not useful telemetry. Code Mode introduces new failure modes It removes some agent failures and adds others. Normal tool calling can fail because the model: forgets a step calls a tool repeatedly loads too much data serializes work unnecessarily loses intermediate state Code Mode can reduce those. But now the generated program can have: syntax errors type errors infinite loops unbounded fan-out bad retry logic stale state race conditions incorrect aggregation side effects inside loops A serious runtime needs limits. At minimum: execution timeout memory limit maximum tool calls concurrency limit output limit nesting limit cancellation Pi already places deadlines and output limits around its sandbox, and large outputs can be spilled rather than pushed back into context. PydanticAI’s Monty environment intentionally exposes a restricted subset of Python and limits host access. This suggests another useful rule: Code Mode should be powerful enough for orchestration, not powerful enough to become an accidental operating system. JavaScript or Python? The current implementations split here. Cloudflare and Pi use JavaScript. Anthropic, PydanticAI and smolagents use Python. This is less important than it first appears. The generated language needs: variables arrays/maps loops conditionals functions async calls parallel calls basic data transformation Both languages satisfy that. The more important properties are: How cheap is sandbox startup? What host capabilities exist? Can execution be interrupted? Can state persist? Can nested calls be traced? Can calls pause for approval? Can programs be replayed? Can tool types be generated cleanly? Pydantic’s Monty is interesting precisely because it treats this as a runtime problem rather than “just execute Python.” Pi does the same with QuickJS in WASM. The language is the visible part. The runtime is the product. A useful comparison | System | Generated language | Main sandbox | Tool discovery | Persistent state | Main emphasis | |---|---|---|---|---|---| | Cloudflare Code Mode | JavaScript / TypeScript-shaped API | Workers / isolated execution | search + describe | Durable runtime / snippets | Huge API surfaces, MCP, durable execution | | Pi Codemode | JavaScript | QuickJS WASM in worker | Inline declarations + tool search | Session store | Coding-agent harness integration | | Anthropic Programmatic Tool Calling | Python | Anthropic code execution container | Provider tool definitions / related advanced tool use | Container reuse | Managed programmatic calling | | PydanticAI Code Mode | Python | Monty | Tool Search integration | Persistent REPL state | Safe embeddable runtime | | smolagents CodeAgent | Python | Local or configured sandbox | Agent tool set | Agent execution state | Code as primary action representation | They are different implementations of roughly the same abstraction: LLM ↓ small program ↓ capability runtime ↓ tools What I think Code Mode actually is Calling this feature “Code Mode” makes it sound optional: js normal agent + sometimes let it execute code I think the deeper interpretation is different. Normal tool calling asks an LLM to act as both: reasoning engine and workflow interpreter Code Mode separates them. The LLM produces a temporary program. The runtime executes it. Then the LLM returns when semantic reasoning is needed again. That looks much closer to a compiler architecture: intent ↓ LLM ↓ temporary program / IR ↓ runtime ↓ observations The program does not need to survive forever. It can exist for 200 milliseconds. It can be generated specifically for one observation. It can call five tools and disappear. This makes it a useful intermediate representation between language models and tools. The interesting future is not bigger tool calls There is a natural progression here. Generation 1 LLM → text Generation 2 LLM → JSON tool call Generation 3 LLM → code → many tool calls The next step is probably not simply more code. The runtime can start providing primitives that are cheaper or safer than asking the main model again: classify ... validate ... wait ... watch ... parallel ... approve ... spawn ... store ... Pi already exposes model classifiers from its Code Mode environment. Its merged implementation allows generated programs to inspect the model catalog and invoke classification operations. This is interesting because the program no longer orchestrates only tools. It can orchestrate computation at different intelligence levels. For example: js const issues = await tools.github.searchIssues { repo } const relevant = await Promise.all issues.map issue = models.classify "Is this issue related to authentication?", issue.body return issues.filter , i = relevant i The expensive model writes the algorithm once. A cheaper decision mechanism executes inside the loop. Now the architecture becomes: main LLM │ ▼ code ┌────────┼─────────┐ ▼ ▼ ▼ tools small AI local compute That starts to look like a real computational runtime for agents. When I would use Code Mode Code Mode is a good fit when the task contains: fan-out loops large intermediate results filtering aggregation parallel calls large tool catalogs reusable procedures mostly deterministic control flow Examples: check 100 services and return the failed ones search ten sources and deduplicate results read every changed file and run a validator query several APIs and join their results inspect every failed CI run find all matching records, update a subset I would keep direct tool calls for: open this file run this command ask the user make one API request perform one dangerous action requiring approval Anthropic’s own measurements support this distinction: workloads with many calls and large intermediate results benefit; workloads with one or two short sequential calls may not. So I would not make Code Mode a universal replacement for tools. I would make it a first-class execution path beside them. The design I would build A minimal Code Mode implementation only needs: one sandbox one code tool a bridge to existing tools A production implementation needs more. I would want: 1. Capability-based sandbox 2. Typed tool bindings 3. Progressive tool discovery 4. Direct and code-only tool exposure 5. Nested tool-call tracing 6. Per-call policy and approval 7. Tool-call concurrency limits 8. Runtime deadline 9. Output budget 10. Persistent session-local state 11. Cancellation 12. Optional durable execution The central API can remain very small: tools.search ... tools.describe ... await tools.foo ... store ... load ... text ... Everything else belongs in the harness. The generated program should not know where credentials live, how permissions work, how traces are exported, or whether a tool is implemented through MCP, HTTP, a shell process or another agent. That is the host’s job. One consequence for MCP Code Mode also changes how I think about MCP. MCP is useful as a transport and capability description protocol. It is much less convincing as the language an LLM should directly program against. Exposing hundreds of MCP tools directly to the model couples: capability transport with: model action representation Those do not need to be the same thing. Cloudflare makes this explicit: MCP can remain underneath Code Mode. An MCP server may itself expose a Code Mode interface, or existing MCP tools can become functions callable by generated code. Pi has now made a similar separation. MCP tools can be available through Codemode or discovered and exposed directly depending on configuration. A cleaner stack is: LLM ↓ Code Mode ↓ tool abstraction ↓ MCP / HTTP / local functions / shell / another agent MCP becomes infrastructure. It stops consuming the entire model interface. Conclusion Code Mode looks like a token optimization at first. It is more useful to see it as a change in the execution model. Traditional agents repeatedly ask the LLM: What should I do next? even when “next” is determined by a loop, a filter or a boolean condition. Code Mode lets the model answer a larger question: What program should control the next section of execution? Then the harness runs that program against a restricted set of capabilities. This reduces model round trips. It keeps intermediate data out of context. It makes huge tool catalogs practical. It also creates a new runtime layer that needs sandboxing, permissions, tracing, resource limits and state. Cloudflare approaches the problem from large APIs. Anthropic approaches it from model-native programmatic tool calling. PydanticAI approaches it from a safe embedded Python runtime. Pi has now integrated the pattern directly into a coding-agent harness. The implementations differ. The architecture is converging: LLM for semantic decisions. Code for control flow. Tools for authority. Harness for policy and state. That separation is more important than the name “Code Mode.”