Jev Mogs Code Mode ToolJev, a gateway that fronts MCP servers with two agent tools — search(query) and execute(code) — cut input tokens by 77% and cost by 28% while calling the right tool in 89% of tasks with 612 tools, versus 81% for Claude Code's own tool search, according to benchmarks published in the project's bench/RESULTS.md. On 400 support tickets, hosted Jev routed at 98.8% accuracy for $0.063 versus 99.3% for $0.332 with Claude Haiku routing each ticket itself, a 5.3x cost reduction, and Jev alone scored 98.8% accuracy with 99.7% on the 97% of items it was at least 0.9 confident about. The author reports that direct tool calling was more accurate (0.94 vs 0.89) and cheaper with 87 tools, concluding the approach pays off only above roughly 100 tools. Your agent doesn't need an LLM for every decision. 612 tools: 77% fewer tokens . 400 tickets: 5.3x cheaper, 98.8% accurate . I benchmarked toolJev with hosted Jev https://typesafe.ai on MCPToolBench++, LiveMCPBench, When2Call and live Claude Haiku agents. Several results went against my first design, and the design changed to match. | Question | Result hosted Jev | Takeaway | |---|---|---| | Does it pick the right tool? | Right tool ranked first: 83% on MCPToolBench++, 54% on LiveMCPBench, 99% on When2Call. Retrieval alone: 73%, 40%, 92%. Jev alone, with no retrieval: 18% | Retrieval shortlists, Jev picks | | Does it know when no tool fits? | AUROC 0.94 on When2Call near-misses embeddings 0.74 , 0.72 on LiveMCPBench 0.68 , 0.89 on out-of-domain MCPToolBench++ 0.91 | The one signal that handles both kinds of miss | | Can Jev make per-item decisions? | 98.8% on 400 support tickets, alone. The 97% it was ≥ 0.9 sure of: 99.7% . Claude Haiku doing every ticket: 99.3% | Confident answers are safe to act on in code | | Agent routing those 400 tickets | 98.8% for $0.063 , against 99.3% for $0.332 with Haiku routing each ticket itself | 5.3x cheaper at about the same accuracy | | Agent with 612 tools | Right tool called in 89% of tasks vs 81% for Claude Code's own tool search, with 77% fewer input tokens and 28% lower cost, but slower 34 s vs 18 s | Pays off at scale | | Agent with 87 tools | Direct tool calling was more accurate 0.94 vs 0.89 and cheaper local backend | Don't bother below ~100 tools | Sample sizes, methods and every caveat are in bench/RESULTS.md https://github.com/VishiATChoudhary/toolJev/blob/main/bench/RESULTS.md . The short version of the caveats: - hosted routing numbers use up to 200 answerable + 200 unanswerable queries per benchmark - the agent runs use 1 to 2 reps and 36 tasks per setup; at that size 89% vs 81% is three tasks - upstream servers are mocks built from each benchmark's tool schemas - the 87-tool agent run and the GIF use nanojev https://github.com/VishiATChoudhary/nanojev , a local stand-in for Jev that is much less accurate A real run with local models and no API key: uv run python examples/showcase.py Give an agent 600 MCP tools and it drowns in tool schemas. Give it one tool call per turn, and a 400-item job takes 400 calls. toolJev sits in front of every MCP server you have and gives the agent two tools instead: - search query finds the few tools a request needs. Retrieval shortlists them, and Jev https://typesafe.ai says whether any of them actually fits. Jev is a decision model: it returns calibrated probabilities over options you give it, never text. - execute code runs Python the agent writes. The code calls tools as mcp.