17:52
2026-08-16
twitter.com
machine-learning
Base-2 vs. base-e log-sum-exp mismatch silently corrupts attention in SGLang
A base-2 vs. base-e log-sum-exp mismatch in SGLang's merge v2 API silently corrupts attention outputs for models using MLA with the FlashInfer backend when radix cache hits exceed 8k, affecting DeepSeβ¦