cd /news/large-language-models/fluidpd-in-place-elasticity-for-slo-… · home › topics › large-language-models › article
[ARTICLE · art-146552] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

FluidPD, a prefill-decode disaggregated LLM serving system, improves overall SLO attainment over static SGLang by up to 94.6 percentage points across production Azure trace workloads, according to the arXiv paper 2610.06917v1. FluidPD uses two mechanisms: FluidToken, which offloads a bounded portion of prefill computation to decode workers when decode-side slack is available, and FluidRole, which reassigns running workers between prefill and decode roles in place without model reload or engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations, letting the system handle short bursts and sustained shifts in the prefill-to-decode demand ratio without provisioning additional workers.

by read1 min views2 publishedOct 7, 2026

arXiv:2610.06917v1 Announce Type: new Abstract: Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers. However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity. FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by off a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart. Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.

── more in #large-language-models 4 stories · sorted by recency
── more on @fluidpd 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fluidpd-in-place-ela…] indexed:0 read:1min 2026-10-07 · —