LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
Researchers introduced LinearKV, a training-free framework enabling position-independent caching (PIC) for hybrid large language models (LLMs) that combine linear recurrences with full attention. The …