{"slug": "beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient", "title": "BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference", "summary": "Researchers propose BeaconKV, a KV cache compression method for Large Reasoning Models (LRMs) that uses beacon queries to guide compression, addressing memory bottlenecks from long Chain-of-Thought generation. The method aims to reduce GPU memory usage while maintaining model performance.", "body_md": "Large Reasoning Models (LRMs) achieve superior problem-solving through extended Chain-of-Thought (CoT) generation, but the resulting key-value (KV) cache grows linearly with sequence length and creates severe memory bottlenecks, often exceeding GPU capacity for long reasoning traces. Existing KV cac", "url": "https://wpnews.pro/news/beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient", "canonical_source": "https://aiflash.com/news/116119/", "published_at": "2026-09-09 05:30:22+00:00", "updated_at": "2026-09-09 05:59:46.303025+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["BeaconKV"], "alternates": {"html": "https://wpnews.pro/news/beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient", "markdown": "https://wpnews.pro/news/beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient.md", "text": "https://wpnews.pro/news/beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient.txt", "jsonld": "https://wpnews.pro/news/beaconkv-key-value-cache-compression-guided-by-beacon-queries-for-efficient.jsonld"}}