Bursty arrivals speed up LLM inference
A benchmark study by an independent researcher found that burstier request arrivals speed up LLM inference, contradicting standard intuition. The analysis of vLLM serving shows that higher burstiness reduces median time …