04:00
2026-10-07
arxiv.org
large-language-models
Mask-Guided KV Cache Eviction in Block Diffusion Language Models
A training-free method called MaskAhead reduces key-value cache memory in block diffusion language models by 9.5x on average on long-prompt question answering, at a cost of 1.2 points of mean F1 versuβ¦