DEdit: Iterative Draft Editing for Speculative Decoding DEdit, a diffusion-based speculative decoding drafter from an arXiv paper (arXiv:2609.38510v1), achieves macro-average speedups of 5.72x on Qwen3-4B and 5.97x on Qwen3-8B over autoregressive generation under greedy decoding across seven benchmarks. DEdit iteratively edits its draft through token-to-token predictions so later predictions serve as bidirectional context to repair earlier errors, and its ProposalMix training scheme mixes draft predictions with ground-truth tokens based on first-pass confidence, halving harmful edits that shorten the accepted prefix. Restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating future context drives the gains. arXiv:2609.38510v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing multiple tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks on Qwen3-4B and Qwen3-8B, DEdit achieves the highest macro-average token acceptance and speedup among the evaluated drafters, reaching macro-average speedups of $5.72\times$ and $5.97\times$ over autoregressive generation under greedy decoding, respectively. Further analysis shows that acceptance improves with more editing passes and wider drafting windows, and that ProposalMix halves harmful edits that shorten the accepted prefix. Moreover, restricting the editor to causal attention lowers acceptance, especially on highly predictable outputs, indicating that future context is a key source of these gains.