00:00
2026-09-07
kraghavan.ca
ai-infrastructure
When Prefill Stops Being Compute-Bound β The KV Cache Memory Wall and What It Does to Your Scheduler
Prefix caching can push LLM prefill from compute-bound to memory-bandwidth-bound once enough of the prefill is loading cached KV rather than computing it, according to a preprint by Meng, Lee and Wangβ¦