cd/entity/StreamingLLM· home entities StreamingLLM
grep -l @streamingllm /news/*.json | wc -l → 2

StreamingLLM

mentions 2 type Organization feed RSS

// recent coverage 2 mentions

08:18
2026-08-10
ssenthilnathan3.github.io
machine-learning

MoE routing is just branch prediction

A software engineer's analysis argues that MoE routing in transformer inference is fundamentally the same problem as CPU branch prediction, and that KV cache management techniques such as prefix cachi…

07:55
2026-07-17
dev.to
large-language-models

Attention Sinks: Why Streaming LLMs Break When You Evict Token 0

A developer explains that attention sinks—tokens at position 0 that absorb excess attention weight—cause streaming LLMs to fail when evicted from the KV cache. The softmax normalization forces the mod…

// co-occurs with top 8 entities
// topics top 6 topics