04:00
2026-10-08
arxiv.org
large-language-models
KVFetch: Temporal Prefetching for the Missing Half of KV Cache Compression
A new arXiv paper (2610.08811v1) proposes KVFetch, a training-free, drop-in framework that adds a temporal recall channel to score-based KV cache compressors, addressing a failure the authors call "se…