TokenRouter: A serving engine for token-level LLM routing Researchers at Tsinghua University released TokenRouter, an open-source serving engine for token-level LLM routing, on GitHub under the thu-nics organization. TokenRouter schedules and executes requests that switch between small and large models mid-response, supporting five routing algorithms — CITER, R2R, R-Stitch, Co-LLM and ME — plus GlimpRouter, query-level routing and a random baseline, and it adds a delayed-batching scheme whose throughput-optimal threshold is found via a discrete-time Markov chain model. The engine was benchmarked on an 8 × A100-80GB host with Qwen3-0.6B / 32B for the two-model algorithms and Qwen3-8B added for ME at weights 0.5 / 0.3 / 0.2. Independent progress. A fast model should keep decoding while a slower peer handles routed tokens. A serving engine for token-level LLM routing. The opportunity Token-level routing lets small and large models collaborate within a single response. The serving engine must keep up with every switch. Irregular token arrivals need a scheduler that brings requests together at the right time. The routing policy describes one request. The runtime manages batching, handoff and cache state. The engine Request-centric programming. Model-centric execution. Admit requests and stream completed tokens. Schedule local batches and execute model steps. Send, receive and resume routed requests. A pending request keeps its serving state and KV slot. Returning tokens are appended without repeating prefix matching or KV allocation. Programming interface Describe when a request changes models, what travels with it, and how decoding resumes. route batch, result After each forward pass, return one model name per request. Returning the current model name continues local decoding. Existing two-model schedulers also accept Boolean decisions. R-Stitch example: True delegates to the only peer; False keeps decoding locally. | Policy | Models | Routing decision | |---|---|---| | CITER | 2 | Low confidence routes to a peer for one token. | | R2R | 2 | A learned router predicts when the models would diverge. | | R-Stitch | 2 | Entropy controls switching in both directions. | | Co-LLM | 2 | A learned deferral signal requests a peer token. | | ME | 2+ | Models are selected token by token using ensemble weights. | The released code also includes GlimpRouter, query-level routing and a random baseline. See the supported schedulers https://github.com/thu-nics/TokenRouter supported-token-level-routing-algorithms . Install from source and launch an R2R server with the Qwen3 0.6B / 32B configuration. git clone https://github.com/thu-nics/TokenRouter.git cd TokenRouter conda create -n tokenrouter python=3.10 conda activate tokenrouter pip install -e . python -m tokenrouter.launch server \ --config path config/r2r/Qwen3-0.6B Qwen3-32B.yaml Once the server is running, use the OpenAI client example https://github.com/thu-nics/TokenRouter/blob/main/examples/openai chat completion.py or the requests example https://github.com/thu-nics/TokenRouter/blob/main/examples/requests chat completion.py . The Python engine example https://github.com/thu-nics/TokenRouter/blob/main/examples/engine generate.py runs directly without a separate HTTP server. The README https://github.com/thu-nics/TokenRouter quick-start covers multi-node serving; the reproduction guide https://github.com/thu-nics/TokenRouter reproducing-our-results includes configurations, workloads and benchmark commands. Delayed batching Starting immediately can leave the next arrival waiting for a whole decoding step. A short, controlled wait brings routed requests into the same batch. Measured threshold sweep A larger batch is useful until waiting leaves too little work in flight. The best threshold depends on concurrency and routing behavior. A discrete-time Markov chain models queued requests, active batches and remaining execution time for each subserver. Given concurrency, routing probabilities and model step latencies, TokenRouter searches the feasible thresholds for the highest expected throughput. The condition ∑