04:00
2026-09-02
arxiv.org
artificial-intelligence
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
Researchers have introduced Faster Flash Decoding (FFD), a hardware-algorithm co-design framework that achieves up to 11.6x kernel-level speedup and 2.37x end-to-end throughput improvement for long-co…