04:00
2026-09-03
arxiv.org
artificial-intelligence
HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models
Researchers introduced HeadWiseKV, a training-free framework that compresses residual global KV caches in hybrid long-context language models under a budget, preserving quality while reducing memory. …