04:00
2026-08-26
arxiv.org
machine-learning
PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression
Researchers propose PuzzleKV, a training- and calibration-free method for KV cache compression in large language models, which partitions per-head KV caches into fixed-length logical pages and appliesβ¦