cd /news/large-language-models/tokenrouter-a-serving-engine-for-tok… · home › topics › large-language-models › article
[ARTICLE · art-148156] src=fuvty.github.io ↗ pub= topic=large-language-models verified=true sentiment=· neutral

TokenRouter: A serving engine for token-level LLM routing

Researchers at Tsinghua University released TokenRouter, an open-source serving engine for token-level LLM routing, on GitHub under the thu-nics organization. TokenRouter schedules and executes requests that switch between small and large models mid-response, supporting five routing algorithms — CITER, R2R, R-Stitch, Co-LLM and ME — plus GlimpRouter, query-level routing and a random baseline, and it adds a delayed-batching scheme whose throughput-optimal threshold is found via a discrete-time Markov chain model. The engine was benchmarked on an 8 × A100-80GB host with Qwen3-0.6B / 32B for the two-model algorithms and Qwen3-8B added for ME at weights 0.5 / 0.3 / 0.2.

read6 min views15 publishedOct 9, 2026

Independent progress.

A fast model should keep decoding while a slower peer handles routed tokens.

A serving engine fortoken-level LLM routing.

The opportunity

Token-level routing lets small and large models collaborate within a single response. The serving engine must keep up with every switch.

Irregular token arrivals need a scheduler that brings requests together at the right time.

The routing policy describes one request. The runtime manages batching, handoff and cache state.

The engine

Request-centric programming.Model-centric execution.

Admit requests and stream completed tokens.

Schedule local batches and execute model steps.

Send, receive and resume routed requests.

A pending request keeps its serving state and KV slot. Returning tokens are appended without repeating prefix matching or KV allocation.

Programming interface

Describe when a request changes models, what travels with it, and how decoding resumes.

route(batch, result)

After each forward pass, return one model name per request. Returning the current model name continues local decoding. Existing two-model schedulers also accept Boolean decisions.

R-Stitch example: True delegates to the only peer; False keeps decoding locally.

Policy Models Routing decision
CITER 2 Low confidence routes to a peer for one token.
R2R 2 A learned router predicts when the models would diverge.
R-Stitch 2 Entropy controls switching in both directions.
Co-LLM 2 A learned deferral signal requests a peer token.
ME 2+ Models are selected token by token using ensemble weights.

The released code also includes GlimpRouter, query-level routing and a random baseline. See the supported schedulers.

Install from source and launch an R2R server with the Qwen3 0.6B / 32B configuration.

git clone https://github.com/thu-nics/TokenRouter.git
cd TokenRouter
conda create -n tokenrouter python=3.10
conda activate tokenrouter
pip install -e .
python -m tokenrouter.launch_server \
  --config_path config/r2r/Qwen3-0.6B_Qwen3-32B.yaml

Once the server is running, use the OpenAI client example or the requests example. The Python engine example runs directly without a separate HTTP server.

The README covers multi-node serving; the reproduction guide includes configurations, workloads and benchmark commands.

Delayed batching

Starting immediately can leave the next arrival waiting for a whole decoding step. A short, controlled wait brings routed requests into the same batch.

Measured threshold sweep

A larger batch is useful until waiting leaves too little work in flight. The best threshold depends on concurrency and routing behavior.

A discrete-time Markov chain models queued requests, active batches and remaining execution time for each subserver. Given concurrency, routing probabilities and model step latencies, TokenRouter searches the feasible thresholds for the highest expected throughput.

The condition ∑<sub>i</sub>(B<sub>i</sub> − 1) < N avoids a state where all requests wait in queues. The analytical model assumes stationary routing probabilities and fixed per-step model latencies; it does not model every nonstationary routing pattern.

Serving efficiency

Five algorithms, three workloads, and two baselines. TokenRouter accelerates the system that executes the routing policy.

CITER, R2R, R-Stitch and Co-LLM use Qwen3-0.6B / 32B. ME adds Qwen3-8B, with weights 0.5 / 0.3 / 0.2.

Experiments use an 8 × A100-80GB host. Two-model runs share two GPUs via CUDA MPS: the small model uses TP1 and the large model TP2. ME places each model on one GPU.

Low-effort: AIME2024, ~100-token inputs and a 2,048-token output cap.

High-effort: selected AIME2024 problems, ~100-token inputs and an 8,192-token output cap.

Agentic: multi-turn SWE-Smith trajectories, ~8,192 input tokens and a 1,024-token output cap.

Official Code: the algorithm’s released implementation. Unavailable entries are shown as N/A.

Std. Serving: one SGLang server per model, returning one token per call to an external dispatcher.

Throughput is total generated output tokens divided by elapsed time. Latency is end-to-end request completion time. TokenRouter builds on SGLang 0.5.1.

Scaling concurrency

As concurrency grows, asynchronous execution turns more in-flight requests into useful model work.

TokenRouter at concurrency 16 reaches 630.57 tokens/s and 39.41 tokens/s per user. R2R official at concurrency 1 reaches 33.94 tokens/s and 34.87 tokens/s per user. This comparison uses different concurrency levels.

Inside the gains

CUDA graph support, asynchronous execution and delayed batching build on one another.

Capture a wider range of extend lengths encountered when a routed request resumes.

Run the router alongside the model to avoid a separate IPC round trip on every token.

Retain pending requests and coordinate the KV budget across colocated models.

R2R · throughput (output tokens/s) across Qwen3 model pairs
Small / large model R2R official TokenRouter Gain
--- --- --- ---
0.6B / 8B 105.10 336.98 3.21×
0.6B / 32B 77.08 197.40 2.56×
1.7B / 8B 103.37 285.41 2.76×
4B / 8B 106.67 211.81 1.99×

Qwen3-0.6B uses one GPU and Qwen3-32B uses two GPUs (TP2), without GPU overlap. Models are deployed on one or two nodes connected by 3.7 GB/s RoCE. The following measurements use low-effort reasoning.

Throughput in output tokens/s · independent deployment experiment
Concurrency Deployment R-Stitch R2R CITER
--- --- --- --- ---
1 Single node 47.53 74.90 108.65
Two nodes 46.19 69.92 93.99
4 Single node 116.14 228.94 308.36
Two nodes 104.20 202.43 277.56
Concurrency 4 · latency in seconds · throughput in output tokens/s
Algorithm / original setting Implementation Throughput TTFT Latency
--- --- --- --- ---
R2RDeepSeek-R1-Distill-Qwen 1.5B / 32B AIME LLM-only 145.30 0.13 411.98
Official Code 89.62 0.11 751.15
TokenRouter 244.56 0.11 270.19
CITERQwen2 1.5B / 72B CommonsenseQA LLM-only 123.72 0.083 2.68
Official Code 17.16 0.036 7.64
TokenRouter 149.31 0.036 0.48
Co-LLMTuned LLaMA2 7B / 70B GSM8K LLM-only 134.79 0.069 13.39
Official Code 3.46 1.14 247.61
TokenRouter 76.02 0.067 11.46
R-StitchL1-1.5B-short / QwQ-32B AIME · official code unavailable LLM-only 150.41 0.15 360.76
TokenRouter 140.58 0.067 151.30

Citation

@article{fu2026tokenrouter,
  title={TokenRouter: Efficient Serving System for Token-Level LLM Routing},
  author={Tianyu Fu and Tengxuan Liu and Ruoxi Wang and Yixin Dong and Yi Ge and Yichen You and Yu Wang},
  journal={arXiv preprint arXiv:2610.12242},
  year={2026},
}
── more in #large-language-models 4 stories · sorted by recency
── more on @tokenrouter 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tokenrouter-a-servin…] indexed:0 read:6min 2026-10-09 · —