Preempting the Prefill, Part 3: Results & Benchmark
VLLM's slack-aware preemption policies rescued urgent request attainment at high load in a benchmark on 6× A100 SXM4 80GB GPUs running Llama 70B, where the control policy collapsed to 0% urgent attainment at 10 req/s whi…