cd /news/large-language-models/linearkv-one-cached-state-suffices-f… · home topics large-language-models article
[ARTICLE · art-94761] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

Researchers introduced LinearKV, a training-free framework enabling position-independent caching (PIC) for hybrid large language models (LLMs) that combine linear recurrences with full attention. The framework uses a single cached state to initialize linear layers, matching or exceeding the quality of exact prefix-state composition while reducing time-to-first-token to 0.46× full prefill. On Mamba-2, exact composition collapsed to 46.6% quality under EPIC, whereas LinearKV's single-state initializer recovered 86.8%.

read1 min views1 publishedAug 13, 2026

arXiv:2608.11231v1 Announce Type: new Abstract: LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache underlies its core operations: matching reusable token chunks, concatenating their KV entries, and selectively recomputing a few tokens to restore cross-chunk context. Hybrid LLMs break these primitives---they replace most attention layers with linear recurrences that expose only a fixed-size state, leaving no token-indexed KV to concatenate or to locally repair. This raises a natural question: can PIC benefit hybrid models, and what would it take? We present LinearKV, a training-free hybrid-PIC framework. Its key insight is a \emph{decoupled initialization}: each linear layer maps its $K$ matched local states to a single initial state, while full-attention layers concatenate their KV as before. LinearKV is therefore compatible with existing PIC methods, reusing their token selection and recomputation as-is. Under this framework, we find that a \emph{single cached state} suffices as the linear layer's initializer. The algebraically principled alternative---composing all $K$ cached states into the exact full-prefix state, as concurrent work HYPIC does---is unnecessary and, on some architectures, even harmful. We compare the two across three hybrid models and three PIC selectors. On the two GDN models the two tie, both recovering most of full quality (up to $92%$); on the Mamba-2 model, exact composition instead collapses under every selector---under EPIC, for instance, it recovers only $46.6%$ of full quality, versus $86.8%$ for a single cached block initializer. A single state initializer is also cheaper, cutting time-to-first-token to $0.46\times$ full prefill versus a further $5$--$17%$ overhead for exact composition; results hold across LongBench QA and RULER at 8K--32K.

── more in #large-language-models 4 stories · sorted by recency
── more on @linearkv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/linearkv-one-cached-…] indexed:0 read:1min 2026-08-13 ·