04:00
2026-07-21
arxiv.org
large-language-models
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
Researchers propose SelKV, a training-free framework for compressing the key-value (KV) cache in large language models (LLMs) that uses a soft cosine gate to selectively merge or discard tokens and anβ¦