Route to Where the KV Cache Is, Not Where It Was
Ranvier Systems introduced a load-balancing technique for LLM serving that routes requests based on where the KV cache currently resides rather than historical prefix matches, reducing P99 time-to-first-token by 57% unde…