20:46
2026-07-12
nanduruganesh.github.io
artificial-intelligence
Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
A developer released the first open-source training kernels for MiniMax Sparse Attention (MSA) on Hopper and Blackwell GPUs, enabling efficient million-token training with sparse attention. The kernel…