Weave Router 2.0 launched this week with a specific claim: route your coding agent’s requests through a complexity classifier and you get GPT-6 Astra performance at 48-54% of the cost and more than twice the speed. The benchmarks are public, the methodology is standard, and there is one honest caveat worth knowing before you install it.
What the Benchmarks Actually Show #
Two standard benchmarks were used. Terminal-Bench 4.0 runs 66 real software engineering tasks with an 8-hour time limit per task and no scaffolding beyond a bash shell. SWE-Atlas Codebase Q&A asks 124 expert questions about 11 production codebases, requiring agents to build the software, run it, and trace execution to answer.
| Benchmark | System | Pass Rate | Cost / Task | Time / Task |
|---|---|---|---|---|
| Terminal-Bench 4.0 | Weave Router 2.0 | 62.1% | .22 | 20.5 min |
| Terminal-Bench 4.0 | GPT-6 Astra (direct) | 60.6% | 0.03 | 44.1 min |
| SWE-Atlas Codebase Q&A | Weave Router 2.0 | 62.1% | .31 | 6.8 min |
| SWE-Atlas Codebase Q&A | GPT-6 Astra (direct) | 66.1% | .04 | 16.7 min |
On Terminal-Bench, the router beats Astra on every metric: higher pass rate, lower cost, faster completion. On SWE-Atlas, it trades four percentage points of quality for a 54% cost reduction. That is the tradeoff in clear terms: for most everyday coding tasks, the router is a net win; for deep codebase navigation with complex dependencies, going direct to Astra may be worth the premium.
Real production numbers track with the benchmarks. Engineering teams at Robinhood, PostHog, and Reducto have reported 40-70% infrastructure cost reductions without degrading response quality. That range is wide because routing gains compound with session length — the longer and more varied your agent sessions, the more the classifier has to work with.
How It Routes #
Four mechanisms do the work. First, a complexity classifier trained on ten times the sessions used for v1.0 scores every incoming request in single-digit milliseconds. It reads the query, surrounding context, task complexity, and domain before any model call is made.
The clever part is second: cache-aware switching. The router does not simply assign the cheapest model that can handle the task. It tracks your cache state per provider and per session, and it only switches models when the expected savings from a cheaper model exceed the cost of rebuilding the cache prefix. If switching would break a long cached context, it stays put. This is why the savings hold up in practice where simpler routing systems see their gains erode.
Third, an escalation system monitors whether a cheaper model is actually making progress. If it stalls or loops, the task gets bumped to a frontier model automatically. Fourth, subscription-aware routing means your existing flat-rate subscriptions — Claude Max, ChatGPT Pro — get used before the router reaches for pay-per-token capacity.
Before You Install It #
Three things to know. The router reads every prompt your agent sends to classify it, which means it sees your code and your queries. If that is a concern for your team — sensitive codebase, compliance requirements — the self-hosted option under Elastic License 2.0 addresses it. Setup takes two commands regardless:
npx @workweave/router
router providers add anthropic openai
The router auto-detects Claude Code, Codex, and Cursor and configures itself. Pricing is 5% of routed costs for individuals and startups; enterprise tiers start at 50 seats.
The second thing: monitor your cache hit rates in the first week, not just your total bill. If the classifier is switching models too aggressively inside long sessions, you will see cache hits drop before you see savings materialize. The cache math is where most early adopters on the HackerNews launch thread got burned with similar routing tools.
Third: the benchmark gap on SWE-Atlas is real. A 4-point quality drop on codebase navigation tasks matters if those tasks represent a significant portion of your workload. The router is not a universal substitute for Astra — it is a cost-efficient default with a clear escalation path.
Routing Is Becoming Infrastructure #
Weave Router 2.0 is part of a broader convergence. Cursor launched its own per-request classifier in July, cutting 30-50% from enterprise AI bills. SGLang v0.5.21, released this week, added a Decisions API that brings routing logic into the inference engine itself. OpenAI shipped a Decisions API at DevDay. The idea that every request should hit the most expensive model is quietly being deprecated by the whole industry.
That convergence is worth paying attention to. The teams cutting 40-70% from AI bills are not doing anything exotic — they are routing requests to the cheapest model that will solve them and escalating when that model fails. Weave Router 2.0 automates that loop. Whether it belongs in your stack depends on your session patterns, your caching strategy, and how much of your work falls into the codebase-navigation category where the quality gap shows up.