Prefix caching in vLLM under multi-tenant agent traffic
Nexus Labs reduced time-to-first-token (TTFT) from 480ms to 110ms on one tenant by enabling vLLM's prefix cache for agent workloads, while another tenant saw no improvement. The discrepancy was caused by Tenant B's agent…