{"slug": "c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference", "title": "C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference", "summary": "Researchers propose C$^2$KV, a unified framework for non-prefix KV cache reuse that jointly optimizes KV extraction and inference-time concatenation, achieving up to 17× inference speedup under long contexts while preserving generation quality. The method learns a composable and compressed KV cache manifold that is position-agnostic, using a lightweight sidecar Extractor with learnable compression tokens and structured attention flow to enable modular KV representations without modifying the frozen base model.", "body_md": "arXiv:2607.17715v1 Announce Type: new\nAbstract: Long-context inference is central to modern large language model (LLM) applications such as retrieval-augmented generation and multi-document reasoning. To mitigate the growing inference cost, recent work has explored key-value (KV) cache reuse to reduce redundant prefill computation. However, existing reuse methods primarily focus on computation savings and overlook a critical bottleneck in long-context LLM serving: the cost of storing and accessing large KV caches. While KV compression appears to be a natural complement, naively combining compression with non-prefix KV reuse often leads to severe accuracy degradation. In this work, we propose C$^2$KV, a unified framework for non-prefix KV reuse that jointly optimizes KV extraction and inference-time concatenation. C$^2$KV learns a composable and compressed KV cache manifold that is explicitly designed to be position-agnostic. Our approach introduces a lightweight sidecar Extractor with learnable compression tokens and a structured attention flow, enabling modular KV representations that can be flexibly reused and concatenated without modifying the frozen base model. We further employ a compression-concatenation co-training strategy to align extraction-time representations with their downstream reuse behavior. Extensive experiments across multiple long-context benchmarks and model families demonstrate that C$^2$KV significantly reduces KV cache storage and transfer costs, achieving up to 17$\\times$ inference speedup under long contexts, while preserving generation quality.", "url": "https://wpnews.pro/news/c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference", "canonical_source": "https://www.machinebrief.com/news/cdollar2dollarkv-compressed-and-composable-kv-cache-reuse-fo-pwdj", "published_at": "2026-07-21 04:00:00+00:00", "updated_at": "2026-07-21 05:34:10.719647+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["C$^2$KV", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference", "markdown": "https://wpnews.pro/news/c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference.md", "text": "https://wpnews.pro/news/c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference.txt", "jsonld": "https://wpnews.pro/news/c-2-kv-compressed-and-composable-kv-cache-reuse-for-efficient-llm-inference.jsonld"}}