{"slug": "faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context", "title": "Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding", "summary": "Researchers have introduced Faster Flash Decoding (FFD), a hardware-algorithm co-design framework that achieves up to 11.6x kernel-level speedup and 2.37x end-to-end throughput improvement for long-context LLM decoding, scaling to 256K context length while maintaining model accuracy on RULER and LongBench benchmarks. The training-free, plug-and-play method uses content-aware scanning via low-bit quantization and a top-delta strategy to exploit attention sparsity, with code available on GitHub.", "body_md": "arXiv:2609.00097v1 Announce Type: new\nAbstract: The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding. To overcome the inherent trade-offs between the memory overhead of metadata-based metrics and the computational inefficiency of adaptive selection strategies, we present Faster Flash Decoding (FFD), a novel hardware-algorithm co-design framework designed to break the memory wall in long-context decoding. FFD integrates the selector and computer into a fully fused kernel, replacing external metadata indices with content-aware scanning via low-bit quantization. Furthermore, we introduce the top-delta strategy, which dynamically filters blocks to achieve distribution-adaptive sparsity without global synchronization. Offering a training-free and plug-and-play solution, FFD also enables the reuse of scanning results for computation, achieving up to 11.6x kernel-level speedup and scaling to 256K context length, with 2.37x end-to-end throughput improvement. Empirical validation on RULER and LongBench confirms that FFD maintains model accuracy while delivering high-ratio sparsity, with code available at https://github.com/qluoluo/faster-flash-decoding", "url": "https://wpnews.pro/news/faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context", "canonical_source": "https://arxiv.org/abs/2609.00097", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:23:56.674442+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-infrastructure"], "entities": ["Faster Flash Decoding", "RULER", "LongBench", "GitHub"], "alternates": {"html": "https://wpnews.pro/news/faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context", "markdown": "https://wpnews.pro/news/faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context.md", "text": "https://wpnews.pro/news/faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context.txt", "jsonld": "https://wpnews.pro/news/faster-than-flash-exploiting-attention-sparsity-for-efficient-long-context.jsonld"}}