12:00
2026-06-17
github.com
large-language-models
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
Researchers released IndexCache, a patch for SGLang and vLLM that accelerates sparse attention in DeepSeek-V3.2 and GLM-5 models by reusing index computations across layers, achieving up to 1.82Γ prefβ¦