{"slug": "trajectory-level-speculative-decoding-for-diffusion-language-models", "title": "Trajectory-Level Speculative Decoding for Diffusion Language Models", "summary": "Researchers introduced a trajectory-level speculative decoding framework for diffusion-based language models (dLLMs) that reduces denoising iterations by 30-40% and increases tokens-per-step from 2.6 to 4.3, achieving 7-14x speedup over vanilla dLLMs and 1.3x over Fast-dLLM with less than 1% accuracy change across reasoning and code benchmarks. The method, detailed in arXiv:2608.27514v1, constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking, also introducing inter-block speculation for cross-block lookahead.", "body_md": "arXiv:2608.27514v1 Announce Type: new\nAbstract: Diffusion-based language models (dLLMs) enable parallel token generation through iterative denoising, but existing decoding strategies collapse to single-token generation under low confidence, severely limiting throughput. Unlike autoregressive models where speculative decoding operates on token sequences in a fixed left-to-right order, dLLMs require speculating over denoising trajectories-sequences of multi-token updates with explicit positions and unmasking orders. We develop a trajectory-level speculative framework that constructs draft denoising trajectories via confidence-stratified tree exploration and verifies them through blockwise parallel evaluation with bidirectional attention masking. Our method further introduces inter-block speculation, exploiting diffusion models' bidirectional structure to perform cross-block lookahead. We formally characterize when this approach is exact and identify trajectory drift as the fundamental cost of increased parallelism. Building on Fast-dLLM's dual-cache infrastructure, our framework reduces denoising iterations by 30-40% and increases tokens-per-step from 2.6 to 4.3, achieving 7-14x speedup over vanilla dLLMs and 1.3x over Fast-dLLM with less than 1% accuracy change across reasoning and code benchmarks.", "url": "https://wpnews.pro/news/trajectory-level-speculative-decoding-for-diffusion-language-models", "canonical_source": "https://arxiv.org/abs/2608.27514", "published_at": "2026-08-31 04:00:00+00:00", "updated_at": "2026-08-31 04:24:25.282786+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["arXiv", "Fast-dLLM"], "alternates": {"html": "https://wpnews.pro/news/trajectory-level-speculative-decoding-for-diffusion-language-models", "markdown": "https://wpnews.pro/news/trajectory-level-speculative-decoding-for-diffusion-language-models.md", "text": "https://wpnews.pro/news/trajectory-level-speculative-decoding-for-diffusion-language-models.txt", "jsonld": "https://wpnews.pro/news/trajectory-level-speculative-decoding-for-diffusion-language-models.jsonld"}}