cd/entity/DeepSeek Sparse AttentionΒ· homeβ€Ί entitiesβ€Ί DeepSeek Sparse Attention
grep -l @deepseek sparse attention /news/*.json | wc -l β†’ 3

DeepSeek Sparse Attention

mentions 3 type Person feed RSS

// recent coverage 3 mentions

16:34
2026-09-28
naive.ai
large-language-models

Model Release: Naive-N0.5-Flash

NaiveAI released Naive-N0.5-Flash, an open-weight 309B mixture-of-experts model with 15.5B active parameters, native 1M-token context, and no full-attention layers, built for coding and AI R&D. The mo…

03:42
2026-06-30
dev.to
large-language-models

GML5 IndexCache

Researchers from Tsinghua University and Z.ai have proposed IndexCache, a method to reduce the computational cost of DeepSeek Sparse Attention (DSA) in GLM-5.2. IndexCache exploits the observation tha…

14:00
2026-06-11
coles.codes
large-language-models

Local models in mid-2026: the engineering that closed the gap

Local large language models have nearly caught up to frontier models for everyday tasks as of mid-2026, driven by engineering advances in sparse attention and mixture-of-experts architectures that red…

// co-occurs with top 8 entities
// topics top 6 topics