04:00
2026-09-21
arxiv.org
large-language-models
RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
RBS-Attention, a training-free sparse-prefill method from a new arXiv paper (2609.20971v1), achieves 20.65x standalone prefill-attention speedup, 11.92x vLLM prefill-attention speedup, and 5.97x end-t…