Independent progress.
A fast model should keep decoding while a slower peer handles routed tokens.
A serving engine fortoken-level LLM routing.
The opportunity
Token-level routing lets small and large models collaborate within a single response. The serving engine must keep up with every switch.
Irregular token arrivals need a scheduler that brings requests together at the right time.
The routing policy describes one request. The runtime manages batching, handoff and cache state.
The engine
Request-centric programming.Model-centric execution.
Admit requests and stream completed tokens.
Schedule local batches and execute model steps.
Send, receive and resume routed requests.
A pending request keeps its serving state and KV slot. Returning tokens are appended without repeating prefix matching or KV allocation.
Programming interface
Describe when a request changes models, what travels with it, and how decoding resumes.
route(batch, result)
After each forward pass, return one model name per request. Returning the current model name continues local decoding. Existing two-model schedulers also accept Boolean decisions.
R-Stitch example: True delegates to the only peer; False keeps decoding locally.
| Policy | Models | Routing decision |
|---|---|---|
| CITER | 2 | Low confidence routes to a peer for one token. |
| R2R | 2 | A learned router predicts when the models would diverge. |
| R-Stitch | 2 | Entropy controls switching in both directions. |
| Co-LLM | 2 | A learned deferral signal requests a peer token. |
| ME | 2+ | Models are selected token by token using ensemble weights. |
The released code also includes GlimpRouter, query-level routing and a random baseline. See the supported schedulers.
Install from source and launch an R2R server with the Qwen3 0.6B / 32B configuration.
git clone https://github.com/thu-nics/TokenRouter.git
cd TokenRouter
conda create -n tokenrouter python=3.10
conda activate tokenrouter
pip install -e .
python -m tokenrouter.launch_server \
--config_path config/r2r/Qwen3-0.6B_Qwen3-32B.yaml
Once the server is running, use the OpenAI client example or the requests example. The Python engine example runs directly without a separate HTTP server.
The README covers multi-node serving; the reproduction guide includes configurations, workloads and benchmark commands.
Delayed batching
Starting immediately can leave the next arrival waiting for a whole decoding step. A short, controlled wait brings routed requests into the same batch.
Measured threshold sweep
A larger batch is useful until waiting leaves too little work in flight. The best threshold depends on concurrency and routing behavior.
A discrete-time Markov chain models queued requests, active batches and remaining execution time for each subserver. Given concurrency, routing probabilities and model step latencies, TokenRouter searches the feasible thresholds for the highest expected throughput.
The condition ∑<sub>i</sub>(B<sub>i</sub> − 1) < N avoids a state where all requests wait in queues. The analytical model assumes stationary routing probabilities and fixed per-step model latencies; it does not model every nonstationary routing pattern.
Serving efficiency
Five algorithms, three workloads, and two baselines. TokenRouter accelerates the system that executes the routing policy.
CITER, R2R, R-Stitch and Co-LLM use Qwen3-0.6B / 32B. ME adds Qwen3-8B, with weights 0.5 / 0.3 / 0.2.
Experiments use an 8 × A100-80GB host. Two-model runs share two GPUs via CUDA MPS: the small model uses TP1 and the large model TP2. ME places each model on one GPU.
Low-effort: AIME2024, ~100-token inputs and a 2,048-token output cap.
High-effort: selected AIME2024 problems, ~100-token inputs and an 8,192-token output cap.
Agentic: multi-turn SWE-Smith trajectories, ~8,192 input tokens and a 1,024-token output cap.
Official Code: the algorithm’s released implementation. Unavailable entries are shown as N/A.
Std. Serving: one SGLang server per model, returning one token per call to an external dispatcher.
Throughput is total generated output tokens divided by elapsed time. Latency is end-to-end request completion time. TokenRouter builds on SGLang 0.5.1.
Scaling concurrency
As concurrency grows, asynchronous execution turns more in-flight requests into useful model work.
TokenRouter at concurrency 16 reaches 630.57 tokens/s and 39.41 tokens/s per user. R2R official at concurrency 1 reaches 33.94 tokens/s and 34.87 tokens/s per user. This comparison uses different concurrency levels.
Inside the gains
CUDA graph support, asynchronous execution and delayed batching build on one another.
Capture a wider range of extend lengths encountered when a routed request resumes.
Run the router alongside the model to avoid a separate IPC round trip on every token.
Retain pending requests and coordinate the KV budget across colocated models.
| R2R · throughput (output tokens/s) across Qwen3 model pairs | |||
|---|---|---|---|
| Small / large model | R2R official | TokenRouter | Gain |
| --- | --- | --- | --- |
| 0.6B / 8B | 105.10 | 336.98 | 3.21× |
| 0.6B / 32B | 77.08 | 197.40 | 2.56× |
| 1.7B / 8B | 103.37 | 285.41 | 2.76× |
| 4B / 8B | 106.67 | 211.81 | 1.99× |
Qwen3-0.6B uses one GPU and Qwen3-32B uses two GPUs (TP2), without GPU overlap. Models are deployed on one or two nodes connected by 3.7 GB/s RoCE. The following measurements use low-effort reasoning.
| Throughput in output tokens/s · independent deployment experiment | ||||
|---|---|---|---|---|
| Concurrency | Deployment | R-Stitch | R2R | CITER |
| --- | --- | --- | --- | --- |
| 1 | Single node | 47.53 | 74.90 | 108.65 |
| Two nodes | 46.19 | 69.92 | 93.99 | |
| 4 | Single node | 116.14 | 228.94 | 308.36 |
| Two nodes | 104.20 | 202.43 | 277.56 |
| Concurrency 4 · latency in seconds · throughput in output tokens/s | ||||
|---|---|---|---|---|
| Algorithm / original setting | Implementation | Throughput | TTFT | Latency |
| --- | --- | --- | --- | --- |
| R2RDeepSeek-R1-Distill-Qwen 1.5B / 32B AIME | LLM-only | 145.30 | 0.13 | 411.98 |
| Official Code | 89.62 | 0.11 | 751.15 | |
| TokenRouter | 244.56 | 0.11 | 270.19 | |
| CITERQwen2 1.5B / 72B CommonsenseQA | LLM-only | 123.72 | 0.083 | 2.68 |
| Official Code | 17.16 | 0.036 | 7.64 | |
| TokenRouter | 149.31 | 0.036 | 0.48 | |
| Co-LLMTuned LLaMA2 7B / 70B GSM8K | LLM-only | 134.79 | 0.069 | 13.39 |
| Official Code | 3.46 | 1.14 | 247.61 | |
| TokenRouter | 76.02 | 0.067 | 11.46 | |
| R-StitchL1-1.5B-short / QwQ-32B AIME · official code unavailable | LLM-only | 150.41 | 0.15 | 360.76 |
| TokenRouter | 140.58 | 0.067 | 151.30 |
Citation
@article{fu2026tokenrouter,
title={TokenRouter: Efficient Serving System for Token-Level LLM Routing},
author={Tianyu Fu and Tengxuan Liu and Ruoxi Wang and Yixin Dong and Yi Ge and Yichen You and Yu Wang},
journal={arXiv preprint arXiv:2610.12242},
year={2026},
}