09:00
2026-09-07
arxiv.org
machine-learning
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Researchers introduced Sarathi-Serve, an LLM inference scheduler that uses chunked-prefills and stall-free scheduling to improve throughput-latency tradeoffs, achieving 2.6x higher serving capacity fo…