cd /news/large-language-models/routewise-latency-cost-optimization-… · home topics large-language-models article
[ARTICLE · art-115347] src=systems.seas.harvard.edu ↗ pub= topic=large-language-models verified=true sentiment=· neutral

RouteWise: Latency–Cost Optimization for Multi-Provider LLM Routing

RouteWise, a multi-provider LLM routing system presented by Haoran Ni of Nanjing University, reduces cost by 26.6%, mean time to first token by 58.4%, and SLO violations by 98.2% compared with OpenRouter's automatic routing, according to evaluations using BurstGPT and a production agentic-workload trace. The system unifies heterogeneous pricing models and jointly optimizes cost and latency.

read1 min views1 publishedAug 29, 2026

Abstract #

The same open-weight large language model is increasingly available from multiple providers with very different pricing models, including on-demand APIs, fixed quotas, and limited concurrency. This creates an opportunity to reduce inference cost and latency without changing model quality—but only if routing accounts for both heterogeneous pricing and provider latencies that vary over time. Existing routers typically optimize cost or latency alone, trade cost for model quality, or consider only on-demand API pricing.

In this talk, I’ll present RouteWise, a multi-provider routing system that unifies heterogeneous pricing models and jointly optimizes cost and latency. RouteWise intelligently allocates limited subscription capacity, adapts routing to recent provider performance, and exposes a single control knob for navigating the cost–latency trade-off. Evaluated with BurstGPT and a production agentic-workload trace, RouteWise reduces cost by 26.6%, mean time to first token by 58.4%, and SLO violations by 98.2% compared with OpenRouter’s automatic routing.

Bio #

Haoran Ni is a senior undergraduate student majoring in Computer Science at Nanjing University, currently a student research intern working with Professor Juncheng Yang. He works on ML systems, including LLM routing, ML artifact storage and agent runtime systems.

── more in #large-language-models 4 stories · sorted by recency
── more on @routewise 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/routewise-latency-co…] indexed:0 read:1min 2026-08-29 ·