{"slug": "agent-codemode-explained", "title": "Agent Codemode Explained", "summary": "Cloudflare's Code Mode MCP server exposes its entire API of more than 2,500 endpoints in roughly 1,000 tokens by offering just two operations — search for an API surface and execute code against the discovered typed API — versus an estimated 1.17 million input tokens if the same API were represented as ordinary MCP tool definitions. The approach, also appearing in Anthropic's programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and Pi as of September 29, 2026, moves agent control flow out of the LLM so intermediate computation such as filtering, retries, loops, and parallel calls stays out of model context. The design matters because each LLM round-trip in the standard one-call-at-a-time tool loop adds inference cost and accumulates tool results in the context window.", "body_md": "# Agent Codemode Explained\n\n# Code Mode: Moving Agent Control Flow Out of the LLM\n\nMost agents still use tools one call at a time.\n\nThe model decides to call a tool. The runtime executes it. The result is appended to the conversation. The model runs again and decides what to do next.\n\nFor a short task this works well.\n\nFor a task that needs twenty API calls, filtering, retries, a loop, or several dependent operations, the design becomes expensive:\n\n```\nLLM\n ↓\ntool A\n ↓\nLLM\n ↓\ntool B\n ↓\nLLM\n ↓\ntool C\n ↓\nLLM\n```\n\nEach edge through the LLM costs inference time. Tool results also accumulate in the context even when most of the data is only needed to decide the next call.\n\nCode Mode changes this boundary.\n\nInstead of asking the model to emit every tool call, give it one tool that executes a small program:\n\n```\nLLM\n ↓\nprogram\n ├─ tool A\n ├─ tool B\n ├─ tool C\n └─ local filtering\n ↓\nLLM\n```\n\nThe program becomes the control plane for a section of the agent run.\n\nThis idea now appears in Cloudflare Code Mode, Anthropic’s programmatic tool calling, PydanticAI Code Mode, Hugging Face smolagents, and, as of September 29, 2026, Pi.\n\nThey share the same basic observation:\n\nTool calling is an unusually bad programming language.\n\nJSON tool calls can express:\n\n```\ncall X with Y\n```\n\nA normal programming language can express:\n\n```\ncall X for every Y\ncall X and Z concurrently\nretry X if it fails\nfilter the result\ncall Z only if condition Q is true\nstore this value and reuse it later\nreturn only these five fields\n```\n\nThe difference is larger than syntax.\n\nCode Mode changes where computation happens.\n\n## The normal tool loop\n\nSuppose an agent needs to find all unhealthy services and inspect their recent deployments.\n\nWith normal tool calling it might do:\n\n```\nLLM → list_services()\nLLM ← 300 services\n\nLLM → get_health(service_1)\nLLM ← result\n\nLLM → get_health(service_2)\nLLM ← result\n\n...\n```\n\nA capable model can issue several calls in parallel, but the model still has to construct the calls and receive their results.\n\nA better tool might support bulk operations. But now every tool designer has to predict every future composition:\n\n```\nget_health_for_services(...)\nget_unhealthy_services(...)\nget_unhealthy_services_with_deployments(...)\n```\n\nThis is where Code Mode becomes interesting.\n\nThe model can instead write something equivalent to:\n\n```\nservices = await list_services()\n\nhealth = await gather(\n    *[get_health({\"id\": s[\"id\"]}) for s in services]\n)\n\nbad = [\n    s for s, h in zip(services, health)\n    if h[\"status\"] != \"healthy\"\n]\n\ndeployments = await gather(\n    *[get_deployments({\"service_id\": s[\"id\"]}) for s in bad]\n)\n\nreturn [\n    {\n        \"service\": s[\"name\"],\n        \"deployment\": d[0]\n    }\n    for s, d in zip(bad, deployments)\n]\n```\n\nThe LLM sees the task once and writes the control flow once.\n\nThe intermediate 300 service records and hundreds of health responses do not need to become conversation history.\n\nThat is the important part.\n\nCode Mode is not mainly a prettier way to call tools.\n\nIt is a way to move **intermediate computation out of model context**.\n\n## Cloudflare: code as a compact API plan\n\nCloudflare has probably made the clearest argument for Code Mode.\n\nIts problem is unusually visible. The Cloudflare API has more than 2,500 endpoints. Exposing every API endpoint as an MCP tool means shipping a huge collection of JSON schemas into the model context before the agent does anything useful.\n\nCloudflare’s Code Mode MCP server instead exposes essentially two operations: search for an API surface, then execute code against the discovered typed API.\n\nCloudflare reports that its whole API can be exposed in roughly 1,000 tokens this way. Its comparison estimates 1.17 million input tokens if the equivalent API were represented as ordinary MCP tool definitions.\n\nThis solves the first Code Mode problem:\n\n```\ntoo many tools\n```\n\nThere is a second problem:\n\n```\ntoo much intermediate data\n```\n\nCloudflare describes generated code as a compact plan. Calls, filtering and transformations happen inside the execution environment, and only selected output returns to the model.\n\nTheir current interface exposes primitives such as:\n\n```\ncodemode.search(...)\ncodemode.describe(...)\ncodemode.step(...)\ncodemode.run(...)\n```\n\n`search()` performs progressive discovery. `describe()` loads the detailed interface only when needed. `step()` records work that needs replay semantics. `run()` executes reusable snippets.\n\nThis is an important refinement.\n\nA naive Code Mode implementation still puts every available function declaration in the system prompt:\n\n```\none code tool\n+\n10,000 function declarations\n```\n\nThat fixes intermediate results but not tool-definition cost.\n\nCloudflare instead makes the API itself lazy.\n\nThe model starts with a small discovery interface and pulls schemas when it needs them.\n\nConceptually:\n\n```\n          ┌──────── search\nLLM → code│\n          ├──────── describe\n          │\n          ├──────── API A\n          ├──────── API B\n          └──────── API C\n```\n\nThe tool catalog becomes something closer to a filesystem or symbol table than a prompt.\n\nThat is likely the right abstraction for very large tool spaces.\n\n### Cloudflare’s sandbox matters\n\nOnce the model can write arbitrary JavaScript, `eval()` is not an acceptable runtime.\n\nCloudflare executes generated code in isolated Workers. Its MCP design blocks direct outbound access by default; generated code reaches the outside world through capabilities supplied by the host. Credentials remain outside the generated program.\n\nThe distinction is useful:\n\n```\ncode execution capability\n        ≠\nexternal authority\n```\n\nThe sandbox can be computationally expressive while having almost no ambient authority.\n\nThe program may know how to write:\n\n```\nawait dns.deleteZone(...)\n```\n\nbut the actual authority still belongs to the host connector.\n\nThis means authorization does not have to be implemented by trusting generated code.\n\nIt remains at the tool boundary.\n\nCloudflare’s durable runtime goes further. It records execution history and can pause generated code for approval, then replay completed work and continue the program.\n\nThat turns Code Mode from “run some generated JavaScript” into an agent runtime primitive.\n\n## Anthropic: programmatic tool calling\n\nAnthropic independently arrived at almost exactly the same model.\n\nTheir current term is **programmatic tool calling**.\n\nClaude writes Python inside its code execution environment. Tool calls from the program cross back to the host. When a tool returns, execution resumes. Intermediate results remain inside the execution environment instead of being inserted into Claude’s context.\n\nThe execution model looks like this:\n\n```\nClaude generates Python\n\n        │\n        ▼\n\nsandbox\n  │\n  ├── tool()\n  │      │\n  │      └──── host executes tool\n  │\n  ├── process result locally\n  │\n  ├── tool()\n  │\n  └── print small result\n\n        │\n        ▼\n\nClaude continues\n```\n\nThis is not equivalent to asking Claude to generate a Python script and running the script after the model finishes.\n\nThe interesting part is the suspended host call.\n\nThe Python program is effectively an orchestration language around external capabilities.\n\nAnthropic reports useful measurements here. On its BrowseComp and DeepSearchQA experiments, programmatic tool calling improved performance by an average of 11% while using 24% fewer input tokens. On an internal 75-tool project-management benchmark it reports roughly 38% lower billed input tokens with unchanged task accuracy. But on τ²-bench tasks dominated by one or two sequential calls, it cost about 8% more.\n\nThat last number is important.\n\nCode Mode is not free.\n\nFor:\n\n```\ncall tool\nreturn answer\n```\n\ngenerating a program and starting a code environment is extra work.\n\nCode Mode becomes useful when the code replaces enough model-mediated control flow to pay for itself.\n\n## Pi: Code Mode moves into the harness\n\nPi is a more interesting implementation for coding agents because Code Mode is not only an API optimization.\n\nIt is integrated into the agent harness.\n\nA large Pi change was merged on September 29, 2026. The implementation runs model-generated JavaScript in a QuickJS VM compiled to WebAssembly. The script can call Pi tools as async functions. Nested calls do not enter the LLM context individually; only explicit script output and the final return value do.\n\nThe standalone package describes the capability directly:\n\n``` js\nconst source = await tools.read({ path: \"package.json\" })\ntext(JSON.parse(source).name)\n```\n\nEvery execution gets a QuickJS VM in a worker. The VM has its own WASM memory and no general route back into the Node host except the explicitly installed host-call bridge.\n\nThis makes Pi’s implementation worth studying.\n\nPi already had very powerful tools:\n\n```\nread\nwrite\nedit\nbash\n```\n\nFor coding work, `bash` alone is close to universal.\n\nSo why add Code Mode?\n\nBecause universality is not the same as a good agent interface.\n\nConsider:\n\n```\ncat package.json | jq ...\n```\n\nversus:\n\n``` js\nconst p = await tools.read({path: \"package.json\"})\nconst pkg = JSON.parse(p)\n```\n\nBoth work.\n\nBut the second form can combine normal agent tools, MCP tools and future host capabilities under one runtime, without turning every operation into shell text.\n\nMore importantly, nested calls remain visible to the harness.\n\nPi routes nested calls through the normal validation and tool hooks. They carry a `parentToolCallId`, and nested calls are recorded on the parent tool result.\n\nThis gives a useful structure:\n\n```\ncodemode call\n│\n├── read\n├── grep\n├── MCP: search\n└── bash\n```\n\ninstead of hiding the entire operation behind an opaque process.\n\nThat matters for observability, permissions and evaluation.\n\n### Pi’s tool exposure model\n\nPi also exposes a design problem that smaller Code Mode demos can avoid:\n\n**Which tools should the model see directly?**\n\nPi now distinguishes different forms of tool exposure and supports two useful Code Mode configurations.\n\nIn `codemode.mode = \"on\"`, normal tools may remain directly visible while Code Mode can also call them.\n\nIn `codemode.mode = \"only\"`, the model sees Code Mode as the primary interface and calls ordinary tools through generated code instead.\n\nThis gives two different agent architectures.\n\n#### Hybrid\n\n```\nLLM\n ├── read\n ├── edit\n ├── bash\n └── codemode\n       ├── read\n       ├── edit\n       └── MCP tools\n```\n\n#### Code-only\n\n```\nLLM\n └── codemode\n       ├── read\n       ├── edit\n       ├── bash\n       ├── MCP\n       └── extensions\n```\n\nThe second one is more radical.\n\nThe model-facing action language stops being a collection of JSON tools.\n\nIt becomes JavaScript.\n\nPi also places a token budget on inline declarations. Tools that do not fit can be found dynamically through tool search.\n\nSo Pi is converging on the same two-layer structure as Cloudflare:\n\n```\nsmall language runtime\n+\nlazy capability discovery\n```\n\nThis looks increasingly less like an agent feature and more like an operating-system interface.\n\n## PydanticAI: Python as the capability language\n\nPydanticAI implements the same pattern using Python and Monty.\n\nEligible tools disappear from the direct model-facing tool list and become functions available inside one `run_code` tool.\n\nThe generated Python can use loops, conditions, variables and `asyncio.gather`. Intermediate data stays in the sandbox.\n\nFor example:\n\n```\nrows = await search(\"failed builds\")\n\ndetails = await asyncio.gather(\n    *[get_build(row[\"id\"]) for row in rows]\n)\n\nreturn [\n    x for x in details\n    if x[\"status\"] == \"failed\"\n]\n```\n\nPydanticAI’s implementation is particularly interesting because Monty is not CPython running in a container.\n\nMonty is a restricted Python interpreter written in Rust for AI-generated programs.\n\nBy default it has no normal filesystem, environment variables or network access. Host capabilities have to be passed in deliberately.\n\nThat gives a capability model similar to Pi’s QuickJS sandbox:\n\n```\ngenerated program\n       │\n       ├── computation\n       ├── variables\n       ├── branching\n       └── injected host functions\n```\n\nThe model gets the useful parts of Python without automatically inheriting the authority of the host Python process.\n\nPydanticAI also supports persistent REPL state during an agent run. Variables, imports and helper functions can survive across `run_code` calls.\n\nThat changes Code Mode again.\n\nIt is no longer just:\n\n```\nLLM → program → result\n```\n\nIt can become:\n\n```\nLLM → persistent computational workspace\n```\n\nThat distinction will matter for long-running agents.\n\n## smolagents: code as the action format\n\nHugging Face’s smolagents predates some of these systems and takes the idea further at the agent-loop level.\n\nIts `CodeAgent` uses Python code as the normal action representation. `ToolCallingAgent` is the alternative that emits conventional structured tool calls.\n\nThe important observation from smolagents is simple:\n\nJSON is a serialization format.\n\nA programming language is a control-flow language.\n\nIf the task is:\n\n```\nsearch A\nsearch B\ntake the intersection\nfetch each result\nrank them\n```\n\nrepresenting that as a program is natural.\n\nRepresenting it as five separate JSON actions requires the LLM itself to become the control-flow interpreter.\n\nThat is the hidden cost of ordinary agent loops.\n\n## Code Mode is a tiny compiler boundary\n\nOne way to understand these systems is to stop thinking about “code execution.”\n\nThink of the LLM as producing a program for a small virtual machine.\n\nThe traditional agent loop is:\n\n```\nobservation\n   ↓\nLLM\n   ↓\none action\n   ↓\nobservation\n```\n\nCode Mode becomes:\n\n```\nobservation\n   ↓\nLLM\n   ↓\nprogram\n   ↓\nruntime\n   ↓\nmany actions\n   ↓\ncompressed observation\n```\n\nThe LLM has effectively compiled several future decisions into one artifact.\n\nThis means Code Mode is useful only when those future decisions can be expressed algorithmically.\n\nSuppose a result requires semantic judgment:\n\n```\nRead this incident report.\nUnderstand what probably caused the outage.\nChoose what to investigate next.\n```\n\nA JavaScript `if` statement cannot replace the next LLM call.\n\nBut many agent decisions are not semantic reasoning:\n\n```\nfor every item\nif status == failed\ntake first five\nsort by date\nretry twice\ncall these concurrently\nextract this field\n```\n\nToday we often waste LLM calls performing those operations.\n\nCode Mode removes them from the model loop.\n\n## The real token saving has two parts\n\nDiscussions of Code Mode often combine two independent optimizations.\n\nThey should be separated.\n\n### 1. Tool definition compression\n\nInstead of:\n\n```\ntool A schema\ntool B schema\ntool C schema\n...\ntool 2500 schema\n```\n\ngive the model:\n\n```\nsearch(query)\ndescribe(symbol)\nexecute(code)\n```\n\nThen discover interfaces lazily.\n\nThis is the part that lets Cloudflare expose thousands of endpoints in a small fixed context.\n\nAnthropic described a similar MCP approach where tool definitions are discovered through files instead of loading the whole MCP catalog into context. In one example, it reduced tool-related context from roughly 150,000 tokens to about 2,000.\n\n### 2. Intermediate result compression\n\nInstead of:\n\n```\ntool result\n→ LLM\n→ tool result\n→ LLM\n→ tool result\n→ LLM\n```\n\ndo:\n\n```\ntool result\n→ program\n→ tool result\n→ program\n→ filter\n→ final small result\n→ LLM\n```\n\nThese solve different scaling problems.\n\nA serious Code Mode implementation should probably do both.\n\n## Code Mode does not remove the agent loop\n\nA tempting conclusion is:\n\n```\nWhy not put the entire agent in code?\n```\n\nBecause generated code only knows what the model knew when it wrote the program.\n\nConsider:\n\n```\nresult = await investigate()\n```\n\nIf `result` contains something genuinely surprising, deciding what it means may require another model inference.\n\nA normal program can branch on structure:\n\n```\nif result[\"status\"] == \"failed\":\n```\n\nIt cannot reliably branch on arbitrary semantics:\n\n```\nif result implies that the database migration\nis probably unrelated to the outage:\n```\n\nYou can put another model behind a function call:\n\n```\nif await classify(result, question):\n```\n\nbut now the model has returned through another door.\n\nSo the likely architecture is not:\n\n```\nCode Mode replaces Agent Loop\n```\n\nIt is:\n\n```\nAgent Loop\n    │\n    ├── semantic decision\n    │\n    └── Code Mode\n          ├── deterministic control flow\n          ├── data processing\n          ├── bulk tool calls\n          └── cheap decisions\n```\n\nThe boundary between these two layers is one of the more interesting agent-design problems.\n\n## Code Mode also changes tool design\n\nToday’s tool APIs are partly shaped by the weakness of model tool calling.\n\nWe create high-level tools such as:\n\n```\nsearch_and_summarize_incidents\nfind_recent_failed_deployments\nget_customer_context\n```\n\nbecause asking the model to compose eight primitive tools is expensive and unreliable.\n\nOnce tools can be composed inside code, smaller primitives become practical:\n\n```\nsearch\nfetch\nread\nwrite\nquery\nexecute\n```\n\nThe generated program provides the composition.\n\nThis does not mean every tool should become low level.\n\nEvery tool boundary is also a:\n\n```\npermission boundary\nvalidation boundary\nobservability boundary\nfailure boundary\n```\n\nBut Code Mode changes the trade-off.\n\nWithout Code Mode, coarse tools save LLM turns.\n\nWith Code Mode, coarse tools need to justify themselves for another reason.\n\n## The security boundary is not the language sandbox\n\nRunning generated code in QuickJS, Monty or a Dynamic Worker is necessary.\n\nIt is not sufficient.\n\nThe dangerous operation is usually not:\n\n```\nwhile (true) {}\n```\n\nIt is:\n\n```\nawait tools.deleteDatabase(...)\n```\n\nThe generated program should therefore be treated as untrusted orchestration around trusted capabilities.\n\nThe useful architecture is:\n\n```\n                   generated code\n                         │\n                  sandbox boundary\n                         │\n                         ▼\n                 capability bridge\n                         │\n          ┌──────────────┼──────────────┐\n          ▼              ▼              ▼\n       read tool      GitHub tool     DB tool\n          │              │              │\n       policy         policy         policy\n          │              │              │\n       execute        execute        execute\n```\n\nPermissions belong below the code.\n\nPi’s nested calls pass through normal tool validation and hooks.\n\nCloudflare explicitly says generated code does not replace authorization and that permissions should still be enforced in upstream handlers.\n\nAnthropic similarly warns that its tool caller configuration should not itself be treated as a security boundary.\n\nThese implementations are converging on the same rule:\n\nsandbox computation; authorize capabilities.\n\n## Observability becomes harder\n\nDirect tool calling gives a convenient trace:\n\n```\nmodel\ntool\nmodel\ntool\nmodel\ntool\n```\n\nCode Mode compresses the model trace:\n\n```\nmodel\ncodemode\nmodel\n```\n\nThat is good for tokens.\n\nIt can be bad for debugging if the runtime treats the code invocation as one opaque tool call.\n\nA useful Code Mode runtime therefore needs nested traces:\n\n```\ncodemode #42\n├─ search #43\n│  └─ 42 ms\n├─ fetch #44\n│  └─ 181 ms\n├─ fetch #45\n│  └─ 172 ms\n└─ update #46\n   └─ approval required\n```\n\nPi explicitly records nested calls and their parent IDs.\n\nCloudflare’s durable runtime records steps and replay information.\n\nPydanticAI exposes nested Code Mode calls in its tracing and can integrate them with durable execution systems.\n\nThis is not optional infrastructure.\n\nOnce a generated program can call fifty tools, “the Code Mode tool failed” is not useful telemetry.\n\n## Code Mode introduces new failure modes\n\nIt removes some agent failures and adds others.\n\nNormal tool calling can fail because the model:\n\n```\nforgets a step\ncalls a tool repeatedly\nloads too much data\nserializes work unnecessarily\nloses intermediate state\n```\n\nCode Mode can reduce those.\n\nBut now the generated program can have:\n\n```\nsyntax errors\ntype errors\ninfinite loops\nunbounded fan-out\nbad retry logic\nstale state\nrace conditions\nincorrect aggregation\nside effects inside loops\n```\n\nA serious runtime needs limits.\n\nAt minimum:\n\n```\nexecution timeout\nmemory limit\nmaximum tool calls\nconcurrency limit\noutput limit\nnesting limit\ncancellation\n```\n\nPi already places deadlines and output limits around its sandbox, and large outputs can be spilled rather than pushed back into context.\n\nPydanticAI’s Monty environment intentionally exposes a restricted subset of Python and limits host access.\n\nThis suggests another useful rule:\n\nCode Mode should be powerful enough for orchestration, not powerful enough to become an accidental operating system.\n\n## JavaScript or Python?\n\nThe current implementations split here.\n\nCloudflare and Pi use JavaScript.\n\nAnthropic, PydanticAI and smolagents use Python.\n\nThis is less important than it first appears.\n\nThe generated language needs:\n\n```\nvariables\narrays/maps\nloops\nconditionals\nfunctions\nasync calls\nparallel calls\nbasic data transformation\n```\n\nBoth languages satisfy that.\n\nThe more important properties are:\n\n```\nHow cheap is sandbox startup?\nWhat host capabilities exist?\nCan execution be interrupted?\nCan state persist?\nCan nested calls be traced?\nCan calls pause for approval?\nCan programs be replayed?\nCan tool types be generated cleanly?\n```\n\nPydantic’s Monty is interesting precisely because it treats this as a runtime problem rather than “just execute Python.”\n\nPi does the same with QuickJS in WASM.\n\nThe language is the visible part.\n\nThe runtime is the product.\n\n## A useful comparison\n\n| System | Generated language | Main sandbox | Tool discovery | Persistent state | Main emphasis | \n|---|---|---|---|---|---|\n| Cloudflare Code Mode | JavaScript / TypeScript-shaped API | Workers / isolated execution | `search()` +`describe()` | Durable runtime / snippets | Huge API surfaces, MCP, durable execution | \n| Pi Codemode | JavaScript | QuickJS WASM in worker | Inline declarations + tool search | Session store | Coding-agent harness integration | \n| Anthropic Programmatic Tool Calling | Python | Anthropic code execution container | Provider tool definitions / related advanced tool use | Container reuse | Managed programmatic calling | \n| PydanticAI Code Mode | Python | Monty | Tool Search integration | Persistent REPL state | Safe embeddable runtime | \n| smolagents CodeAgent | Python | Local or configured sandbox | Agent tool set | Agent execution state | Code as primary action representation | \n\nThey are different implementations of roughly the same abstraction:\n\n```\nLLM\n  ↓\nsmall program\n  ↓\ncapability runtime\n  ↓\ntools\n```\n\n## What I think Code Mode actually is\n\nCalling this feature “Code Mode” makes it sound optional:\n\n``` js\nnormal agent\n+\nsometimes let it execute code\n```\n\nI think the deeper interpretation is different.\n\nNormal tool calling asks an LLM to act as both:\n\n```\nreasoning engine\nand\nworkflow interpreter\n```\n\nCode Mode separates them.\n\nThe LLM produces a temporary program.\n\nThe runtime executes it.\n\nThen the LLM returns when semantic reasoning is needed again.\n\nThat looks much closer to a compiler architecture:\n\n```\nintent\n ↓\nLLM\n ↓\ntemporary program / IR\n ↓\nruntime\n ↓\nobservations\n```\n\nThe program does not need to survive forever.\n\nIt can exist for 200 milliseconds.\n\nIt can be generated specifically for one observation.\n\nIt can call five tools and disappear.\n\nThis makes it a useful intermediate representation between language models and tools.\n\n## The interesting future is not bigger tool calls\n\nThere is a natural progression here.\n\n### Generation 1\n\n```\nLLM → text\n```\n\n### Generation 2\n\n```\nLLM → JSON tool call\n```\n\n### Generation 3\n\n```\nLLM → code → many tool calls\n```\n\nThe next step is probably not simply more code.\n\nThe runtime can start providing primitives that are cheaper or safer than asking the main model again:\n\n```\nclassify(...)\nvalidate(...)\nwait(...)\nwatch(...)\nparallel(...)\napprove(...)\nspawn(...)\nstore(...)\n```\n\nPi already exposes model classifiers from its Code Mode environment. Its merged implementation allows generated programs to inspect the model catalog and invoke classification operations.\n\nThis is interesting because the program no longer orchestrates only tools.\n\nIt can orchestrate computation at different intelligence levels.\n\nFor example:\n\n``` js\nconst issues = await tools.github.searchIssues({ repo })\n\nconst relevant = await Promise.all(\n  issues.map(issue =>\n    models.classify(\n      \"Is this issue related to authentication?\",\n      issue.body\n    )\n  )\n)\n\nreturn issues.filter((_, i) => relevant[i])\n```\n\nThe expensive model writes the algorithm once.\n\nA cheaper decision mechanism executes inside the loop.\n\nNow the architecture becomes:\n\n```\n             main LLM\n                │\n                ▼\n              code\n       ┌────────┼─────────┐\n       ▼        ▼         ▼\n     tools    small AI   local compute\n```\n\nThat starts to look like a real computational runtime for agents.\n\n## When I would use Code Mode\n\nCode Mode is a good fit when the task contains:\n\n```\nfan-out\nloops\nlarge intermediate results\nfiltering\naggregation\nparallel calls\nlarge tool catalogs\nreusable procedures\nmostly deterministic control flow\n```\n\nExamples:\n\n```\ncheck 100 services and return the failed ones\nsearch ten sources and deduplicate results\nread every changed file and run a validator\nquery several APIs and join their results\ninspect every failed CI run\nfind all matching records, update a subset\n```\n\nI would keep direct tool calls for:\n\n```\nopen this file\nrun this command\nask the user\nmake one API request\nperform one dangerous action requiring approval\n```\n\nAnthropic’s own measurements support this distinction: workloads with many calls and large intermediate results benefit; workloads with one or two short sequential calls may not.\n\nSo I would not make Code Mode a universal replacement for tools.\n\nI would make it a first-class execution path beside them.\n\n## The design I would build\n\nA minimal Code Mode implementation only needs:\n\n```\none sandbox\none code tool\na bridge to existing tools\n```\n\nA production implementation needs more.\n\nI would want:\n\n```\n1. Capability-based sandbox\n2. Typed tool bindings\n3. Progressive tool discovery\n4. Direct and code-only tool exposure\n5. Nested tool-call tracing\n6. Per-call policy and approval\n7. Tool-call concurrency limits\n8. Runtime deadline\n9. Output budget\n10. Persistent session-local state\n11. Cancellation\n12. Optional durable execution\n```\n\nThe central API can remain very small:\n\n```\ntools.search(...)\ntools.describe(...)\n\nawait tools.foo(...)\n\nstore(...)\nload(...)\n\ntext(...)\n```\n\nEverything else belongs in the harness.\n\nThe generated program should not know where credentials live, how permissions work, how traces are exported, or whether a tool is implemented through MCP, HTTP, a shell process or another agent.\n\nThat is the host’s job.\n\n## One consequence for MCP\n\nCode Mode also changes how I think about MCP.\n\nMCP is useful as a transport and capability description protocol.\n\nIt is much less convincing as the language an LLM should directly program against.\n\nExposing hundreds of MCP tools directly to the model couples:\n\n```\ncapability transport\n```\n\nwith:\n\n```\nmodel action representation\n```\n\nThose do not need to be the same thing.\n\nCloudflare makes this explicit: MCP can remain underneath Code Mode. An MCP server may itself expose a Code Mode interface, or existing MCP tools can become functions callable by generated code.\n\nPi has now made a similar separation. MCP tools can be available through Codemode or discovered and exposed directly depending on configuration.\n\nA cleaner stack is:\n\n```\nLLM\n ↓\nCode Mode\n ↓\ntool abstraction\n ↓\nMCP / HTTP / local functions / shell / another agent\n```\n\nMCP becomes infrastructure.\n\nIt stops consuming the entire model interface.\n\n## Conclusion\n\nCode Mode looks like a token optimization at first.\n\nIt is more useful to see it as a change in the execution model.\n\nTraditional agents repeatedly ask the LLM:\n\n```\nWhat should I do next?\n```\n\neven when “next” is determined by a loop, a filter or a boolean condition.\n\nCode Mode lets the model answer a larger question:\n\n```\nWhat program should control the next section of execution?\n```\n\nThen the harness runs that program against a restricted set of capabilities.\n\nThis reduces model round trips.\n\nIt keeps intermediate data out of context.\n\nIt makes huge tool catalogs practical.\n\nIt also creates a new runtime layer that needs sandboxing, permissions, tracing, resource limits and state.\n\nCloudflare approaches the problem from large APIs.\n\nAnthropic approaches it from model-native programmatic tool calling.\n\nPydanticAI approaches it from a safe embedded Python runtime.\n\nPi has now integrated the pattern directly into a coding-agent harness.\n\nThe implementations differ.\n\nThe architecture is converging:\n\n```\nLLM for semantic decisions.\nCode for control flow.\nTools for authority.\nHarness for policy and state.\n```\n\nThat separation is more important than the name “Code Mode.”", "url": "https://wpnews.pro/news/agent-codemode-explained", "canonical_source": "https://julin.ai/2026/10/01/agent-codemode-explained/", "published_at": "2026-09-30 11:00:00+00:00", "updated_at": "2026-10-01 08:46:21.790988+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "agent-protocols", "large-language-models"], "entities": ["Cloudflare", "Anthropic", "PydanticAI", "Hugging Face", "smolagents", "Pi", "Cloudflare Code Mode", "MCP"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agent-codemode-explained", "markdown": "https://wpnews.pro/news/agent-codemode-explained.md", "text": "https://wpnews.pro/news/agent-codemode-explained.txt", "jsonld": "https://wpnews.pro/news/agent-codemode-explained.jsonld"}}