04:00
2026-10-07
arxiv.org
large-language-models
FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving
FluidPD, a prefill-decode disaggregated LLM serving system, improves overall SLO attainment over static SGLang by up to 94.6 percentage points across production Azure trace workloads, according to theβ¦