{"slug": "tokenrouter-a-serving-engine-for-token-level-llm-routing", "title": "TokenRouter: A serving engine for token-level LLM routing", "summary": "Researchers at Tsinghua University released TokenRouter, an open-source serving engine for token-level LLM routing, on GitHub under the thu-nics organization. TokenRouter schedules and executes requests that switch between small and large models mid-response, supporting five routing algorithms — CITER, R2R, R-Stitch, Co-LLM and ME — plus GlimpRouter, query-level routing and a random baseline, and it adds a delayed-batching scheme whose throughput-optimal threshold is found via a discrete-time Markov chain model. The engine was benchmarked on an 8 × A100-80GB host with Qwen3-0.6B / 32B for the two-model algorithms and Qwen3-8B added for ME at weights 0.5 / 0.3 / 0.2.", "body_md": "### Independent progress.\n\nA fast model should keep decoding while a slower peer handles routed tokens.\n\nA serving engine for**token-level LLM routing.**\n\nThe opportunity\n\nToken-level routing lets small and large models collaborate within a single response. The serving engine must keep up with every switch.\n\nIrregular token arrivals need a scheduler that brings requests together at the right time.\n\nThe routing policy describes one request. The runtime manages batching, handoff and cache state.\n\nThe engine\n\nRequest-centric programming.**Model-centric execution.**\n\nAdmit requests and stream completed tokens.\n\nSchedule local batches and execute model steps.\n\nSend, receive and resume routed requests.\n\nA pending request keeps its serving state and KV slot. Returning tokens are appended without repeating prefix matching or KV allocation.\n\nProgramming interface\n\nDescribe when a request changes models, what travels with it, and how decoding resumes.\n\nroute(batch, result)\n\nAfter each forward pass, return one model name per request. Returning the current model name continues local decoding. Existing two-model schedulers also accept Boolean decisions.\n\nR-Stitch example: True delegates to the only peer; False keeps decoding locally.\n\n| Policy | Models | Routing decision | \n|---|---|---|\n| CITER | 2 | Low confidence routes to a peer for one token. | \n| R2R | 2 | A learned router predicts when the models would diverge. | \n| R-Stitch | 2 | Entropy controls switching in both directions. | \n| Co-LLM | 2 | A learned deferral signal requests a peer token. | \n| ME | 2+ | Models are selected token by token using ensemble weights. | \n\nThe released code also includes GlimpRouter, query-level routing and a random baseline. See the [supported schedulers](https://github.com/thu-nics/TokenRouter#supported-token-level-routing-algorithms).\n\nInstall from source and launch an R2R server with the Qwen3 0.6B / 32B configuration.\n\n```\ngit clone https://github.com/thu-nics/TokenRouter.git\ncd TokenRouter\nconda create -n tokenrouter python=3.10\nconda activate tokenrouter\npip install -e .\npython -m tokenrouter.launch_server \\\n  --config_path config/r2r/Qwen3-0.6B_Qwen3-32B.yaml\n```\n\nOnce the server is running, use the [OpenAI client example](https://github.com/thu-nics/TokenRouter/blob/main/examples/openai_chat_completion.py) or the [requests example](https://github.com/thu-nics/TokenRouter/blob/main/examples/requests_chat_completion.py). The [Python engine example](https://github.com/thu-nics/TokenRouter/blob/main/examples/engine_generate.py) runs directly without a separate HTTP server.\n\nThe [README](https://github.com/thu-nics/TokenRouter#quick-start) covers multi-node serving; the [reproduction guide](https://github.com/thu-nics/TokenRouter#reproducing-our-results) includes configurations, workloads and benchmark commands.\n\nDelayed batching\n\nStarting immediately can leave the next arrival waiting for a whole decoding step. A short, controlled wait brings routed requests into the same batch.\n\nMeasured threshold sweep\n\nA larger batch is useful until waiting leaves too little work in flight. The best threshold depends on concurrency and routing behavior.\n\nA discrete-time Markov chain models queued requests, active batches and remaining execution time for each subserver. Given concurrency, routing probabilities and model step latencies, TokenRouter searches the feasible thresholds for the highest expected throughput.\n\nThe condition ∑<sub>i</sub>(B<sub>i</sub> − 1) < N avoids a state where all requests wait in queues. The analytical model assumes stationary routing probabilities and fixed per-step model latencies; it does not model every nonstationary routing pattern.\n\nServing efficiency\n\nFive algorithms, three workloads, and two baselines. TokenRouter accelerates the system that executes the routing policy.\n\nCITER, R2R, R-Stitch and Co-LLM use Qwen3-0.6B / 32B. ME adds Qwen3-8B, with weights 0.5 / 0.3 / 0.2.\n\nExperiments use an 8 × A100-80GB host. Two-model runs share two GPUs via CUDA MPS: the small model uses TP1 and the large model TP2. ME places each model on one GPU.\n\n**Low-effort:** AIME2024, ~100-token inputs and a 2,048-token output cap.\n\n**High-effort:** selected AIME2024 problems, ~100-token inputs and an 8,192-token output cap.\n\n**Agentic:** multi-turn SWE-Smith trajectories, ~8,192 input tokens and a 1,024-token output cap.\n\n**Official Code:** the algorithm’s released implementation. Unavailable entries are shown as N/A.\n\n**Std. Serving:** one SGLang server per model, returning one token per call to an external dispatcher.\n\nThroughput is total generated output tokens divided by elapsed time. Latency is end-to-end request completion time. TokenRouter builds on SGLang 0.5.1.\n\nScaling concurrency\n\nAs concurrency grows, asynchronous execution turns more in-flight requests into useful model work.\n\nTokenRouter at concurrency 16 reaches 630.57 tokens/s and 39.41 tokens/s per user. R2R official at concurrency 1 reaches 33.94 tokens/s and 34.87 tokens/s per user. This comparison uses different concurrency levels.\n\nInside the gains\n\nCUDA graph support, asynchronous execution and delayed batching build on one another.\n\nCapture a wider range of extend lengths encountered when a routed request resumes.\n\nRun the router alongside the model to avoid a separate IPC round trip on every token.\n\nRetain pending requests and coordinate the KV budget across colocated models.\n\n| R2R · throughput (output tokens/s) across Qwen3 model pairs |  |  |  | \n|---|---|---|---|\n| Small / large model | R2R official | TokenRouter | Gain | \n|---|---|---|---|\n| 0.6B / 8B | 105.10 | 336.98 | 3.21× | \n| 0.6B / 32B | 77.08 | 197.40 | 2.56× | \n| 1.7B / 8B | 103.37 | 285.41 | 2.76× | \n| 4B / 8B | 106.67 | 211.81 | 1.99× | \n\nQwen3-0.6B uses one GPU and Qwen3-32B uses two GPUs (TP2), without GPU overlap. Models are deployed on one or two nodes connected by 3.7 GB/s RoCE. The following measurements use low-effort reasoning.\n\n| Throughput in output tokens/s · independent deployment experiment |  |  |  |  | \n|---|---|---|---|---|\n| Concurrency | Deployment | R-Stitch | R2R | CITER | \n|---|---|---|---|---|\n| 1 | Single node | 47.53 | 74.90 | 108.65 | \n|  | Two nodes | 46.19 | 69.92 | 93.99 | \n| 4 | Single node | 116.14 | 228.94 | 308.36 | \n|  | Two nodes | 104.20 | 202.43 | 277.56 | \n\n| Concurrency 4 · latency in seconds · throughput in output tokens/s |  |  |  |  | \n|---|---|---|---|---|\n| Algorithm / original setting | Implementation | Throughput | TTFT | Latency | \n|---|---|---|---|---|\n| R2RDeepSeek-R1-Distill-Qwen 1.5B / 32B AIME | LLM-only | 145.30 | 0.13 | 411.98 | \n|  | Official Code | 89.62 | 0.11 | 751.15 | \n|  | TokenRouter | 244.56 | 0.11 | 270.19 | \n| CITERQwen2 1.5B / 72B CommonsenseQA | LLM-only | 123.72 | 0.083 | 2.68 | \n|  | Official Code | 17.16 | 0.036 | 7.64 | \n|  | TokenRouter | 149.31 | 0.036 | 0.48 | \n| Co-LLMTuned LLaMA2 7B / 70B GSM8K | LLM-only | 134.79 | 0.069 | 13.39 | \n|  | Official Code | 3.46 | 1.14 | 247.61 | \n|  | TokenRouter | 76.02 | 0.067 | 11.46 | \n| R-StitchL1-1.5B-short / QwQ-32B AIME · official code unavailable | LLM-only | 150.41 | 0.15 | 360.76 | \n|  | TokenRouter | 140.58 | 0.067 | 151.30 | \n\nCitation\n\n```\n@article{fu2026tokenrouter,\n  title={TokenRouter: Efficient Serving System for Token-Level LLM Routing},\n  author={Tianyu Fu and Tengxuan Liu and Ruoxi Wang and Yixin Dong and Yi Ge and Yichen You and Yu Wang},\n  journal={arXiv preprint arXiv:2610.12242},\n  year={2026},\n}\n```\n\n", "url": "https://wpnews.pro/news/tokenrouter-a-serving-engine-for-token-level-llm-routing", "canonical_source": "https://fuvty.github.io/thinking_yard_project_page/projects/tokenrouter/", "published_at": "2026-10-09 09:10:17+00:00", "updated_at": "2026-10-09 09:22:41.847483+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "ai-research", "mlops", "ai-tools"], "entities": ["TokenRouter", "Tsinghua University", "thu-nics", "Qwen3-0.6B", "Qwen3-32B", "Qwen3-8B", "CITER", "R2R"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tokenrouter-a-serving-engine-for-token-level-llm-routing", "markdown": "https://wpnews.pro/news/tokenrouter-a-serving-engine-for-token-level-llm-routing.md", "text": "https://wpnews.pro/news/tokenrouter-a-serving-engine-for-token-level-llm-routing.txt", "jsonld": "https://wpnews.pro/news/tokenrouter-a-serving-engine-for-token-level-llm-routing.jsonld"}}