{"slug": "thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching", "title": "Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching", "summary": "Researchers propose Thought-Aware Attention Matching (TAM), a KV cache compaction method that segments chain-of-thought reasoning into blocks, allocates compression budgets adaptively, and protects pivotal tokens, improving accuracy over uniform compaction at the same memory footprint. On AIME 2024 and MATH-500 with Qwen3-4B, TAM reduces peak memory to 3.1–3.2 GB (a 65% reduction) while maintaining competitive accuracy.", "body_md": "arXiv:2608.12331v1 Announce Type: new\nAbstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding. Existing compaction methods treat reasoning trajectories as flat token sequences and apply uniform compression, ignoring the hierarchical structure of CoT reasoning where different steps vary drastically in importance. We propose \\textbf{Thought-Aware Attention Matching (TAM)}, which exploits this structure through three mechanisms: (i)~thought segmentation that decomposes the trajectory into reasoning blocks, (ii)~adaptive budget allocation that assigns compression budget based on each segment's importance and size, and (iii)~pivotal token protection that preserves high-attention reasoning anchors. We prove that the allocation rule is optimal under a convex error model and that cumulative error under sequential compaction remains bounded. Experiments on AIME 2024 and MATH-500 with Qwen3-4B show that TAM improves accuracy over uniform compaction at the same memory footprint, with periodic compaction bounding peak memory to 3.1--3.2\\,GB (a 65\\% reduction) while maintaining competitive accuracy.", "url": "https://wpnews.pro/news/thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching", "canonical_source": "https://arxiv.org/abs/2608.12331", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:12:40.567146+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["TAM", "Qwen3-4B", "AIME 2024", "MATH-500"], "alternates": {"html": "https://wpnews.pro/news/thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching", "markdown": "https://wpnews.pro/news/thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching.md", "text": "https://wpnews.pro/news/thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching.txt", "jsonld": "https://wpnews.pro/news/thought-aware-kv-cache-compaction-for-reasoning-via-adaptive-attention-matching.jsonld"}}