12:37
2026-09-21
dev.to
ai-infrastructure
HAProxy's 200 ms LLM token delay is a smart default. Here's when to change it.
An HAProxy engineer investigated a benchmark showing the proxy adds roughly 206 ms of latency to LLM token streams and found the cause is not application-level buffering but the kernel's MSG_MORE flagβ¦