# The router pattern is underrated, how we cut token costs by 60% with intelligent agent selection

> Source: <https://dev.to/alex_aslam/the-router-pattern-is-underrated-how-we-cut-token-costs-by-60-with-intelligent-agent-selection-1kg1>
> Published: 2026-10-09 19:19:28+00:00

I once watched a research pipeline burn through $340 in a single afternoon. Not because it failed but because it succeeded. Twenty-three steps, every one of them routed to a frontier model, and nineteen of those steps were formatting JSON, extracting dates, and classifying support tickets. I was paying Cadillac prices for a job a bicycle could do.

That was the day I stopped blaming the model and started looking at the router.

The math behind agent cost explosion is deceptively simple. A frontier model like Claude Opus runs 5 to 25 times the per-token cost of a small model like Claude Haiku. On a single call, that's a rounding error. On a seven-step workflow where five frontier calls chain together, you're running an order of magnitude past the equivalent small-model pipeline.

The research confirms what my bill already knew. In a production multi-agent run, a hand-tuned per-agent model binding — planner on small, reviewer on frontier, writer on small — sends every call an agent makes to the same model, regardless of how hard the call actually is. The reviewer's six calls per task ranged from "summarize a module" to "find a concurrency bug," and they all went to the frontier because the agent was bound to it.

The kicker is *which* calls cost you. Per-call routing sent only 2 of 16 calls to the frontier model and cost $1.37 per 1,000 reviews — one fifth the cost of sending everything to the frontier, and less than half the cost of the per-agent setup. It also found 4 of 6 seeded bugs, matching the all-frontier arm and beating the per-agent binding, which found 3.

I had been solving the wrong problem. I kept optimizing prompts for a model that shouldn't have been handling the task in the first place.

A colleague looked at my trace and asked a question that reframed everything: *"Why is your summarization step calling the same model as your planning step?"*

The answer was embarrassing. I had one model configured per agent. Every step used it because that was the default. I'd never asked whether each step *needed* it.

That's when I understood that **model choice is a per-call decision**, not a per-agent decision. A planning step genuinely requires frontier-class reasoning. The formatting step that follows it needs a 7B model and nothing more. Treating them the same is the architectural equivalent of hiring a senior architect to file paperwork.

The vLLM team's framing is the one I keep taped to my monitor: binding a model to an agent is "a prediction, made once, about work nobody has seen yet". The task arrives as a stream of calls, and a per-agent binding gives every one of them the same answer, hard or trivial.

The pattern that fixed this is called **capability-based routing** — or semantic routing — and it's embarrassingly simple: classify each call by what it needs, route it to the cheapest sufficient model, and escalate only when a gate rejects the output.

The Federation of Agents paper formalizes the mechanism: **Versioned Capability Vectors (VCVs)** are machine-readable profiles that make agent capabilities searchable through semantic embeddings, enabling agents to advertise their capabilities, cost, and limitations. The routing layer matches tasks to agents over sharded HNSW indices while enforcing operational constraints through cost-biased optimization.

The LatentGate paper from Equinix puts it in enterprise terms. At 100 specialized agents, a router must map natural language inputs to the correct agent while satisfying three constraints simultaneously: sub-100ms latency, high precision (routing errors trigger incorrect API calls), and low-cost onboarding for new agents. Prompt-based LLM routing delivers strong semantic reasoning but incurs prohibitive latency — roughly 1500-2000ms — and accuracy degrades to 70-77% at 100 agents. The fix isn't a bigger router. It's a representation that doesn't collapse semantically similar but functionally distinct agents into the same embedding cone.

The naive cascade has a trap. The gate that decides whether to escalate — that's the load-bearing piece. A gate that only checks output shape will pass wrong answers straight through, because a malformed JSON is easy to catch but a confidently wrong extraction looks exactly like a correct one.

The 2026 literature has moved past theory. The numbers are concrete.

**MTRouter** encodes the interaction history and candidate models into joint history-model embeddings and learns an outcome estimator from logged trajectories to predict turn-level model utility. On ScienceWorld, it surpasses GPT-5 while reducing total cost by **58.7%**. On Humanity's Last Exam, it achieves competitive accuracy while reducing total cost by **43.4%** relative to GPT-5. It makes fewer model switches and is more tolerant to transient errors than prior multi-turn routers.

**Planner-as-Router** folds model-tier selection into planning itself. As the planner breaks a query into subtasks, it assigns each one a model size tier — small, mid, or frontier — so dependencies are visible before any specialist runs. On 1,157 evaluations across 54 enterprise agentic tasks, it cuts cost **44%** against all-frontier routing while giving up only 2.9 points of accuracy.

**RouterHGC** formalizes routing as node selection on a heterogeneous graph whose node types include user queries, collaboration modes, agent roles, and LLMs. It achieves **0.80%-6.17% accuracy gains** on MATH and HotpotQA while reducing inference cost by **27.40%**.

**L7**, a Bayesian routing framework using Thompson Sampling over per-model Beta distributions, demonstrated a **68.5% reduction** in aggregate inference cost relative to a static round-robin baseline across 10,000 heterogeneous benchmark tasks — with zero training data required, working from a uniform prior with online updates.

The consistent finding across all of these: **step-level routing beats agent-level routing** because call complexity varies widely within a single agent's workload.

Here's the part that took me three weeks to understand and two months to fix.

If you implement routing incorrectly in a multi-turn agent, **you can end up paying more than if you'd never routed at all**.

The reason is prompt caching. By Turn 3 of an agent session, cached tokens make up the vast majority of your token payload. That's the efficiency you're protecting.

Switching models mid-session **destroys that protection entirely**. Every provider's cache is model-specific. The new model has no access to the previous model's stored history and must re-read the entire conversation from scratch at full input token cost.

This is why the vLLM Semantic Router team built **Session-Aware Agentic Routing (SAAR)**. The insight is that a router needs to know not just which model is best for the current request, but **when switching models would break the session**. SAAR adds router-owned session memory, hard locks around tool loops and non-portable provider state, safe reset boundaries, and prefix-cache-aware switch pricing. Across 21,600 deterministic turns, SAAR cuts model switches by **79.29%**, eliminates 3,836 unsafe switches, and reduces estimated physical-model cost by **78.71%**.

The mistake I made was routing *across* a session. The fix is routing *within* a step — choose the model for each generation, complete it, and return to the session model. Or better: summarize the context before switching, so the new model starts with a compressed history rather than a cold cache.

**Barclays** ran a controlled experiment on a four-agent pull-request review crew. All frontier cost roughly $7 per 1,000 reviews. Per-agent binding cost about $3. Per-call routing with vLLM Semantic Router cost **$1.37** — one fifth the cost of the all-frontier arm, less than half the per-agent arm, and with the same bug-detection rate as the frontier baseline. The router added 1.86ms at p50 to each decision.

**Equinix** deployed LatentGate across 5 SLM backbones and 100 enterprise agents, achieving **98.8% in-domain and 80.0% out-of-domain accuracy** on natural queries — 13 to 22 absolute points above embedding baselines — at **~28ms** on a T4 GPU, with the SLM forward pass independent of agent count.

**Microsoft Foundry** ships a Model Router as a drop-in deployment that selects the optimal LLM for each request an agent makes. A simple greeting routes to a fast, inexpensive model. A complex tool-calling chain routes to a frontier model. You can assign different routers to different agents, each with its own routing mode and model subset, matching the deployment to each agent's cost profile.

**Red Hat's vLLM Semantic Router** benchmark results are among the most concrete: **+10.2% accuracy**, **-47.1% latency**, and **-48.5% token usage** with auto reasoning mode adjustment. In business and economics domains, accuracy improvements exceed 20%.

Tiered routing is not a default. It's a response to a specific problem.

**Use it when your agent workload has call-level complexity variance.** Planning and synthesis calls need frontier reasoning. Extraction, formatting, and classification calls don't. If every call genuinely requires frontier capability, routing won't help.

**Use it when your token volume is high enough that the savings justify the infrastructure.** A router that saves 40% on a $50 monthly bill isn't worth building. A router that saves 40% on a $50,000 monthly bill is.

**Use step-level features, not query-level features.** The research is clear that single-turn routers misroute agent steps because they miss trajectory context. TwinRouterBench exists precisely because existing benchmarks evaluate routers only on one-shot prompts and never expose the router-visible prefix at an intermediate agent step. The features that matter are the ones available at the step: the instruction, the accumulated context length, the prior step's tier, and the dependency structure.

**Don't route across sessions without summarization.** The cache penalty will eat your savings. Route within a step, or summarize before you switch. SAAR's session locks are the production implementation of this principle.

Capability-based routing buys you cost reduction and, often, latency. It costs you determinism and adds a classifier to your stack.

Every routing decision is a chance to misroute. A complex call sent to a cheap model fails, and you pay the cheap attempt plus the escalation plus the latency of both. The gate's false-accept rate — not the ladder — decides whether the cascade pays. If your gate accepts wrong answers, you ship them at a discount, and the savings are a lie.

And you're accepting more moving parts. A router, a gate, a tier configuration, per-tier cost tracking. Every one of those is a thing that can fail independently. The vLLM Semantic Router's failure modes are instructive: low-confidence answers escalate automatically, provider outages trip circuit breakers, budget exhaustion stops escalation and returns the best available response. None of those are free.

But here's what I've learned from watching my bill drop by 68% without a measurable quality regression: the teams that are winning with agents aren't using the best model for everything. They're using the **right model for each call** — and they've accepted that "right" is a decision worth making explicitly.

So here's my question: **If you looked at your agent's cost breakdown by call, how many of those calls would you have paid frontier prices for if you'd chosen the model deliberately?**

I'd love to hear where you've landed. A semantic router in front of the crew, per-agent binding you've defended, a Bayesian policy that learns online, or a bill that finally made you look — and what finally made you change?
