{"slug": "headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language", "title": "HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models", "summary": "Researchers introduced HeadWiseKV, a training-free framework that compresses residual global KV caches in hybrid long-context language models under a budget, preserving quality while reducing memory. In tests on Qwen3.6-27B, HeadWiseKV reduced sampled peak device memory by 8.59% at a 112K context length and extended the largest verified successful context from 114K to 161K, while retaining near-Full-KV quality on RULER and LoCoMo benchmarks across four models.", "body_md": "arXiv:2609.02029v1 Announce Type: new\nAbstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget. We introduce HeadWiseKV, a training-free framework that compresses the residual global KV caches of hybrid language models while preserving their native local, recurrent, and linear paths. It assigns each physical KV head a static, multilevel history window, making cache demand predictable before serving. We formulate this allocation as a restricted operational rate--distortion problem and propose SeqCalib as the core policy-generation algorithm in HeadWiseKV. SeqCalib processes layers in execution order and conditions each decision on the lower-layer policy used at deployment, thereby accounting for interactions across depth. A grouped-cache runtime materializes the selected policy as actual per-head KV residency rather than a mask over a full cache. We evaluate downstream quality across four hybrid long-context models and study physical residency and serving behavior on Qwen3.6-27B. HeadWiseKV retains near-Full-KV RULER and LoCoMo quality across the evaluated models. In the fixed-model systems study, it reduces sampled peak device memory by 8.59\\% at a 112K context length and extends the largest verified successful context from 114K to 161K.", "url": "https://wpnews.pro/news/headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language", "canonical_source": "https://arxiv.org/abs/2609.02029", "published_at": "2026-09-03 04:00:00+00:00", "updated_at": "2026-09-03 04:24:06.895571+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["HeadWiseKV", "SeqCalib", "Qwen3.6-27B", "RULER", "LoCoMo"], "alternates": {"html": "https://wpnews.pro/news/headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language", "markdown": "https://wpnews.pro/news/headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language.md", "text": "https://wpnews.pro/news/headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language.txt", "jsonld": "https://wpnews.pro/news/headwisekv-budgeted-per-head-cache-residency-for-hybrid-long-context-language.jsonld"}}