05:30
2026-09-09
aiflash.com
artificial-intelligence
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
Researchers propose BeaconKV, a KV cache compression method for Large Reasoning Models (LRMs) that uses beacon queries to guide compression, addressing memory bottlenecks from long Chain-of-Thought ge…