cd /news/large-language-models/wakekv-reactive-reversible-kv-reside… · home › topics › large-language-models › article
[ARTICLE · art-145174] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds

WakeKV, a reactive KV-cache residency policy, improves miss rate over frozen head classification and destructive eviction at matched memory or budget, according to an arXiv paper (2610.02713v1). Across three models (1.5B–8B) and three regimes — needle retrieval, long chain-of-thought, and multi-turn recall — the authors measured four model-regime combinations and found most attention heads change their reading behavior at least once during generation, so WakeKV moves cooling heads to a recoverable CPU reservoir instead of freezing or permanently evicting their state. A FlexiCache/vLLM implementation on Mistral-7B confirmed the benefit on real hardware, improving throughput while retaining LongBench quality, with comparisons against SnapKV, uniform R-KV, and ReasonAlloc.

by read1 min views3 publishedOct 5, 2026

arXiv:2610.02713v1 Announce Type: new Abstract: Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.

── more in #large-language-models 4 stories · sorted by recency
── more on @wakekv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/wakekv-reactive-reve…] indexed:0 read:1min 2026-10-05 · —