02:30
2026-09-23
aiflash.com
large-language-models
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs
Flash-dLLM introduces IO-aware KV caching and parallel decoding to speed up inference and cut memory use in diffusion large language models (dLLMs), addressing the lack of effective Key-Value caching …