cd/entity/Multi-Head Latent AttentionΒ· homeβ€Ί entitiesβ€Ί Multi-Head Latent Attention
grep -l @multi-head latent attention /news/*.json | wc -l β†’ 3

Multi-Head Latent Attention

mentions 3 type Person feed RSS

// recent coverage 3 mentions

02:49
2026-07-21
arxiv.org
artificial-intelligence

Kimi Linear: An Expressive, Efficient Attention Architecture

Researchers at Moonshot AI introduced Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short-context, long-context, and reinforcement learning scaling regimes…

13:14
2026-05-23
dev.to
large-language-models

Multi-Head Latent Attention (MLA)

**Summary:** Multi-Head Latent Attention (MLA) is an attention mechanism used in DeepSeek-V2/V3 and Kimi K2.x models that compresses the Key-Value (KV) cache by projecting full KV pairs into a shared,…

// co-occurs with top 8 entities
// topics top 6 topics