{"slug": "decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization", "title": "Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits", "summary": "Researchers introduced Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that resizes prefill chunks based on observed decode intervals, cutting P99 inter-token latency by a mean 27.7% (paired 95% CI 21.0% to 34.3%) across three paired 100-request trials on Qwen3-0.6B in BF16 on one A100 80 GB GPU, at the cost of a 34.8% rise in mean P99 time to first token that stayed within the declared SLO. The mechanism failed to generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration, which the authors trace to an asynchronous scheduler-call interval that only proxies completed GPU iteration time. The paper is a reproducible proof-of-concept and generalization study and makes no claim about mobile-device performance.", "body_md": "arXiv:2609.38386v1 Announce Type: new \nAbstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted.\n  We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. On Qwen3-0.6B in BF16 on one A100 80 GB GPU, three paired 100-request trials reduce P99 inter-token latency by 24.8%, 30.1%, and 28.2% (mean 27.7%, paired 95% confidence interval 21.0% to 34.3%) with exact output agreement, no failures, and unchanged SLO compliance. The benefit is not free: mean P99 time to first token increases 34.8% while remaining inside the declared SLO.\n  Crucially, the mechanism does not generalize to Qwen3-8B, Qwen3-32B, or a two-GPU tensor-parallel configuration. We trace the failure to an asynchronous scheduler-call interval that is only a proxy for completed GPU iteration time. This negative result defines the boundary of the contribution and motivates a completion-timed controller for concurrent CPU and on-device inference. We do not claim mobile-device performance; the present work is a reproducible proof-of-concept and generalization study.", "url": "https://wpnews.pro/news/decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization", "canonical_source": "https://arxiv.org/abs/2609.38386", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:17:40.132338+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "machine-learning", "ai-research"], "entities": ["Decode-Latency Feedback Prefill", "vLLM", "Qwen3-0.6B", "Qwen3-8B", "Qwen3-32B", "NVIDIA A100 80 GB GPU"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization", "markdown": "https://wpnews.pro/news/decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization.md", "text": "https://wpnews.pro/news/decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization.txt", "jsonld": "https://wpnews.pro/news/decode-latency-feedback-prefill-a-model-free-controller-and-its-generalization.jsonld"}}