Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Researchers introduced Sarathi-Serve, an LLM inference scheduler that uses chunked-prefills and stall-free scheduling to improve throughput-latency tradeoffs, achieving 2.6x higher serving capacity fo…