Prefix caching at scale: when it saves you 80% of prefill cost, and the eviction policies that quietly turn it into 5%
A developer found that deploying a 70B Llama model with RAG features caused time-to-first-token (TTFT) to jump from 180 ms to 1.4 seconds, as the model recomputed identical attention states for repeat…