04:00
2026-10-01
arxiv.org
ai-infrastructure
Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
Researchers introduced Decode-Latency Feedback Prefill (DLFP), a model-free controller implemented in vLLM that resizes prefill chunks based on observed decode intervals, cutting P99 inter-token latenβ¦