{"slug": "jev-mogs-code-mode", "title": "Jev Mogs Code Mode", "summary": "ToolJev, a gateway that fronts MCP servers with two agent tools — search(query) and execute(code) — cut input tokens by 77% and cost by 28% while calling the right tool in 89% of tasks with 612 tools, versus 81% for Claude Code's own tool search, according to benchmarks published in the project's bench/RESULTS.md. On 400 support tickets, hosted Jev routed at 98.8% accuracy for $0.063 versus 99.3% for $0.332 with Claude Haiku routing each ticket itself, a 5.3x cost reduction, and Jev alone scored 98.8% accuracy with 99.7% on the 97% of items it was at least 0.9 confident about. The author reports that direct tool calling was more accurate (0.94 vs 0.89) and cheaper with 87 tools, concluding the approach pays off only above roughly 100 tools.", "body_md": "**Your agent doesn't need an LLM for every decision.**\n\n612 tools: **77% fewer tokens**. 400 tickets: **5.3x cheaper, 98.8% accurate**.\n\nI benchmarked toolJev with hosted [Jev](https://typesafe.ai) on MCPToolBench++,\nLiveMCPBench, When2Call and live Claude Haiku agents. Several results went\nagainst my first design, and the design changed to match.\n\n| Question | Result (hosted Jev) | Takeaway | \n|---|---|---|\n| Does it pick the right tool? | Right tool ranked first: **83%** on MCPToolBench++,**54%** on LiveMCPBench,**99%** on When2Call. Retrieval alone: 73%, 40%, 92%. Jev alone, with no retrieval: 18% | Retrieval shortlists, Jev picks | \n| Does it know when no tool fits? | AUROC **0.94** on When2Call near-misses (embeddings 0.74),**0.72** on LiveMCPBench (0.68), 0.89 on out-of-domain MCPToolBench++ (0.91) | The one signal that handles both kinds of miss | \n| Can Jev make per-item decisions? | **98.8%** on 400 support tickets, alone. The 97% it was ≥ 0.9 sure of:**99.7%** . Claude Haiku doing every ticket: 99.3% | Confident answers are safe to act on in code | \n| Agent routing those 400 tickets | **98.8% for $0.063** , against 99.3% for $0.332 with Haiku routing each ticket itself | **5.3x cheaper** at about the same accuracy | \n| Agent with 612 tools | Right tool called in **89%** of tasks vs 81% for Claude Code's own tool search, with**77% fewer** input tokens and**28% lower** cost, but slower (34 s vs 18 s) | Pays off at scale | \n| Agent with 87 tools | Direct tool calling was more accurate (0.94 vs 0.89) and cheaper (local backend) | Don't bother below ~100 tools | \n\nSample sizes, methods and every caveat are in **[bench/RESULTS.md](https://github.com/VishiATChoudhary/toolJev/blob/main/bench/RESULTS.md)**.\nThe short version of the caveats:\n\n- hosted routing numbers use up to 200 answerable + 200 unanswerable queries per benchmark\n- the agent runs use 1 to 2 reps and 36 tasks per setup; at that size 89% vs 81% is three tasks\n- upstream servers are mocks built from each benchmark's tool schemas\n- the 87-tool agent run and the GIF use [nanojev](https://github.com/VishiATChoudhary/nanojev) , a local stand-in for Jev that is much less accurate\n\n*A real run with local models and no API key: `uv run python examples/showcase.py`*\n\nGive an agent 600 MCP tools and it drowns in tool schemas. Give it one tool\ncall per turn, and a 400-item job takes 400 calls. toolJev sits in front of\nevery MCP server you have and gives the agent **two tools** instead:\n\n- **`search(query)`** finds the few tools a request needs. Retrieval shortlists\nthem, and[Jev](https://typesafe.ai) says whether any of them actually fits.\nJev is a decision model: it returns calibrated probabilities over options you\ngive it, never text.\n- **`execute(code)`** runs Python the agent writes. The code calls tools as` mcp.<server>.<tool>()` , and makes per-item judgements with`jev.choice / noul / map` : typed answers in about 100 ms, no LLM turn. This is[Code Mode](https://blog.cloudflare.com/code-mode-mcp/) crossed with a[Recursive Language Model](https://arxiv.org/abs/2512.24601) , with Jev in the\nplace of the sub-LLM.\n\nThe LLM writes the plan once, Jev makes the per-item calls, and only the items Jev is unsure about come back to the LLM.\n\n```\ngit clone https://github.com/VishiATChoudhary/toolJev && cd toolJev\nuv venv && uv pip install -e \".[local,retrieval]\"\n```\n\nThe two extras:\n\n- `local` installs[nanojev](https://github.com/VishiATChoudhary/nanojev) , which\nruns Jev-style decisions on your machine with no key. The model downloads\nonce.\n- `retrieval` adds embedding search next to BM25.\n\nFor real use, set `TYPESAFE_API_KEY` and use hosted Jev (`backend = \"hosted\"`,\nthe default when the config has no `[decider]` table). It is much more accurate than the local\nstand-in: see Results.\n\n```\nuv run python examples/showcase.py   # the GIF above: 612 tools, 400 tickets, all local\nuv run python examples/demo.py       # the real gateway over MCP stdio, 3 toy upstream servers\n```\n\nWrite a TOML file with one table per upstream server. Anything you'd give an MCP\nclient goes in it: `command`/` args`/` env` for local servers, `url`/` headers`\nfor remote ones.\n\n```\n# tooljev.toml\n[decider]\nbackend = \"hosted\"      # reads TYPESAFE_API_KEY; or \"nanojev\" (+ kind = \"encoder\") to run with no key\n\n[servers.github]\ndescription = \"GitHub repos, issues and pull requests\"\ncommand = \"npx\"\nargs = [\"-y\", \"@modelcontextprotocol/server-github\"]\nenv = { GITHUB_PERSONAL_ACCESS_TOKEN = \"ghp_...\" }\n\n[servers.tickets]\ndescription = \"Customer support tickets: list, assign, close\"\nurl = \"https://support.example.com/mcp\"\nheaders = { Authorization = \"Bearer ...\" }\nallow = [\"list_tickets\", \"assign_ticket\"]   # optional: expose only these tools\n```\n\nA one-line `description` per server helps search a lot. Every option, with its\ndefault, is in [`examples/config.toml`](https://github.com/VishiATChoudhary/toolJev/blob/main/examples/config.toml).\n\ntoolJev is an ordinary MCP server, so any MCP client works. For Claude Code:\n\n```\nclaude mcp add tooljev -- uv --directory /path/to/toolJev run tooljev --config /path/to/tooljev.toml\n```\n\nFor remote clients, serve streamable HTTP with\n`uv run tooljev --config tooljev.toml --transport http --port 8765`.\n\nThe agent never sees your upstream tools directly. It works in two moves.\n\n**Find tools:** `search(\"refund order 1234 on paypal\")` returns up to 5 tools,\neach with a Python signature and a `fit` score. If nothing in the catalog fits,\nthe result carries a warning so the agent answers on its own.\n\n**Run code:** `execute(code)` runs Python with the found tools and `jev` in\nscope:\n\n```\ntickets = await mcp.support.list_tickets()\nanswers = await jev.map([t[\"body\"] for t in tickets],\n                        {\"queue\": jev.Choice(\"Which queue handles this?\", QUEUES)},\n                        min_confidence=0.8)\nunsure = []\nfor t, a in zip(tickets, answers):\n    if a[\"confident\"]:                    # Jev is sure: act in code\n        await mcp.support.route_ticket(ticket_id=t[\"id\"], queue=a[\"queue\"][\"choice\"])\n    else:                                 # Jev is not: hand back to the agent\n        unsure.append(t[\"id\"])\nFINAL({\"unsure\": unsure})                 # only this reaches the agent's context\n```\n\nWhat's in scope inside `execute`:\n\n| call | returns | \n|---|---|\n| `await mcp.<server>.<tool>(**kwargs)` | the tool's result; arguments are checked against its schema first | \n| `await jev.choice(state, instructions, options)` | `{\"choice\", \"probabilities\", \"confidence\"}` | \n| `await jev.noul(state, statement)` | probability the statement is true | \n| `await jev.score(state, instructions, levels)` | `{\"score\", \"probabilities\", \"confidence\"}` | \n| `await jev.map(states, questions, min_confidence=)` | one answer dict per state, plus `\"confident\"` when a threshold is given | \n| `FINAL(value)` /`print(...)` | what goes back to the agent (the last expression also works) | \n\nPass `execute(code, session=\"work\")` to keep variables between calls, like a\nREPL: fetch once, inspect, then act. You don't need to teach your agent any of\nthis; the tool descriptions do.\n\n**Use it when:**\n\n- You connect **hundreds of tools** . At 612 tools it cut input tokens by 77%\nand cost by 28% against Claude Code's own tool search (right tool 89% vs\n81% lenient, 72% vs 75% strict).\n- Your agent does **the same judgement over many items** : triage, routing,\nfiltering, labelling. With hosted Jev, 400 tickets were routed 5.3x cheaper\nthan by the LLM, at 98.8% vs 99.3%. Gate on Jev's confidence and let the LLM\ntake the unsure ones.\n\n**Skip it when:**\n\n- You have **fewer than about 100 tools** . Direct tool calling was more\naccurate and cheaper at 87 tools.\n- You need **every item right** and cost doesn't matter. The LLM alone was\nstill half a point more accurate on triage.\n- **Latency** matters more than tokens. Each search is a hosted round trip\n(about 285 ms), and Code Mode adds a turn.\n\n| `[decider] backend` | What | Needs | \n|---|---|---|\n| `hosted` (default, recommended) | TypeSafe Jev via `typesafe-sdk` . Search reranks with it | `TYPESAFE_API_KEY` | \n| `nanojev` | [nanojev](https://github.com/VishiATChoudhary/nanojev) , a local stand-in,`kind = \"encoder\"` or`\"decoder\"` . Search keeps retrieval order | nothing (downloads a model once) | \n\nBoth implement one method, `decide(state, questions) -> answers`, in TypeSafe's\nwire format, so everything above the backend is shared. Use the encoder for\nnanojev: the decoder could not tell in-catalog requests from out-of-catalog\nones.\n\n1. **Recall:** BM25 and bge-base embeddings over every tool's description, fused\nby reciprocal rank. The top 15 go on. This takes a few milliseconds, and the\nembeddings are cached per tool.\n2. **Rank and fit, in one Jev call:** a Choice over the 15 candidates reorders\nthem, and a Noul per candidate (\"this request asks to \")\nbecomes that tool's`fit` . With hosted Jev the rerank lifts top-1 over\nretrieval alone on every benchmark (MCPToolBench++ 0.73 to 0.83). The local\nnanojev encoder reranks worse than retrieval, so with it the Choice is skipped\nand retrieval order stands (`rerank = true/false` overrides either way).\n3. **Nothing fits:**`in_catalog` is the best fit. Below`abstain_below` (default 0.5), the result carries a warning but still lists the tools.`abstain = \"hard\"` returns none instead. Hiding tools cost agents more in\nwrong detours than a warned shortlist did.\n4. The top `max_tools` (default 5) come back, each with a Python signature.\n\n`mode = \"hierarchical\"` (Jev picks a server, then a tool, with knockout rounds\npast the 255-option limit) is the original design, kept because the benchmarks\nrejected it.\n\nSearch latency on MCPToolBench++: about 285 ms p50 with hosted Jev, 74 ms with the local encoder.\n\nEach `execute` runs in a [Monty](https://github.com/pydantic/monty) worker with\nno filesystem, network or environment access. It gets a fresh worker unless it\nnames a `session`; up to 8 sessions are kept, and the least recently used one\ngoes first.\n\nLimits, all set under `[sandbox]`, are sized for bulk jobs of hundreds of items:\n\n- 120 s wall clock and 20 s of sandbox CPU\n- 256 MB of memory\n- 2,000 host calls; a whole `jev.map` counts as one call, up to 5,000 items\n- 16 concurrent calls\n- stdout truncated at 4k characters\n\nBehaviour worth knowing:\n\n- A session that hits a time or memory limit is discarded, as Monty advises.\n- When an upstream tool fails, sandbox code sees a catchable `RuntimeError` .\n- Each search and execute, and every host call inside it, is logged to\n`~/.tooljev/traces/YYYY-MM-DD.jsonl` .\n\n**The sandbox is not authorization.** Anything a listed upstream tool can do,\nsandbox code can do. Use per-server `allow = [...]` to expose less.\n\n- **Free-text arguments.** Jev cannot write them; the agent writes them in code.\n- **An `llm_query` fallback** inside the sandbox.\n- **Hosted Jev on the full routing sets and the 87-tool agent task.** It ran\non a 200 + 200 sample per benchmark.\n- **One combined abstention score.** Fusing embedding similarity with Jev's fit\nis not done yet; a simple fitted combination did not beat embeddings alone.\n\n```\nuv pip install -e \".[local,retrieval,dev,bench]\"\nuv run pytest -m \"not slow\"                                      # offline: fake decider, in-process servers\nuv run pytest -m slow                                            # real nanojev\nuv run python -m bench.run && uv run python -m bench.report      # routing + abstention benchmarks\nTYPESAFE_API_KEY=... bash bench/hosted.sh                        # every hosted-Jev run, resumable\nuv run python -m bench.agent.run_agent --task triage --n 40      # agent in the loop (uses `claude -p`)\n```\n\nBenchmark data is not committed. Fetch it into `bench/data/` as listed in\n[bench/RESULTS.md](https://github.com/VishiATChoudhary/toolJev/blob/main/bench/RESULTS.md#data).\n\n- **Code Mode:**[Cloudflare](https://blog.cloudflare.com/code-mode-mcp/) ,[Anthropic](https://www.anthropic.com/engineering/code-execution-with-mcp) and[FastMCP](https://gofastmcp.com/servers/transforms/code-mode) .\n- **Recursive Language Models:**[Zhang, Kraska and Khattab](https://arxiv.org/abs/2512.24601) .\n- **Jev:**[TypeSafe](https://typesafe.ai) .[nanojev](https://github.com/VishiATChoudhary/nanojev) is a small local\nstand-in with the same interface.\n- **Sandbox:**[Monty](https://github.com/pydantic/monty) , by Pydantic.\n- **Tool retrieval:**[RAG-MCP](https://arxiv.org/abs/2505.03275) ,[MCP-Zero](https://arxiv.org/abs/2506.01056) and \"Selection Is Retrieval,\nAbstention Is Not\" ([arXiv 2609.18672](https://arxiv.org/abs/2609.18672) ).\nThis repo reached the same conclusion the hard way.\n\nThe full prior-art survey is in [RESEARCH.md](https://github.com/VishiATChoudhary/toolJev/blob/main/RESEARCH.md).", "url": "https://wpnews.pro/news/jev-mogs-code-mode", "canonical_source": "https://github.com/VishiATChoudhary/toolJev", "published_at": "2026-09-28 20:54:53+00:00", "updated_at": "2026-09-28 21:18:07.881542+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "agent-protocols", "large-language-models", "mlops"], "entities": ["toolJev", "Jev", "typesafe.ai", "MCPToolBench++", "LiveMCPBench", "When2Call", "Claude Haiku", "nanojev"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/jev-mogs-code-mode", "markdown": "https://wpnews.pro/news/jev-mogs-code-mode.md", "text": "https://wpnews.pro/news/jev-mogs-code-mode.txt", "jsonld": "https://wpnews.pro/news/jev-mogs-code-mode.jsonld"}}