04:00
2026-08-14
arxiv.org
artificial-intelligence
Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
Researchers propose Thought-Aware Attention Matching (TAM), a KV cache compaction method that segments chain-of-thought reasoning into blocks, allocates compression budgets adaptively, and protects piโฆ