{"slug": "less-uniform-discrete-diffusion-is-more-powerful-and-scalable", "title": "Less Uniform Discrete Diffusion is More Powerful and Scalable", "summary": "Researchers introduced Less Uniform Diffusion (LUDI), a uniform diffusion language model framework that addresses an over-uniform training objective and condition-target confusion during sampling by directing each reverse transition toward the clean token and adding per-token time embeddings for confidence-based few-step sampling. The team continue-trained a 7B autoregressive model into LUDI-7B, which achieved a 3-token-per-step speedup over autoregressive decoding and competitive performance against masked diffusion baselines. The work is published as arXiv:2609.35817v1.", "body_md": "arXiv:2609.35817v1 Announce Type: new \nAbstract: Although uniform diffusion language models (UDLMs) represent a promising diffusion paradigm, scaling them remains challenging. We identify the core obstacle as an over-uniform training objective and condition-target confusion during sampling. To address these, we propose Less Uniform Diffusion (LUDI), a novel UDLM framework. Specifically, we (i) introduce a less uniform loss that directs each reverse transition toward the clean token, and (ii) equip the model with per-token time embeddings that supply token-level corruption hints, enabling confidence-based few-step sampling. Experiments across scales show that LUDI yields cleaner supervision and improves few-step generation. We further continue-train a 7B autoregressive model into LUDI-7B, resulting in a UDLM capable of complex reasoning. It achieves a 3-token-per-step speedup over AR decoding and competitive performance compared with masked diffusion baselines, revealing that the full potential of UDLMs for complex generation remains to be unlocked.", "url": "https://wpnews.pro/news/less-uniform-discrete-diffusion-is-more-powerful-and-scalable", "canonical_source": "https://arxiv.org/abs/2609.35817", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:20:32.991946+00:00", "lang": "en", "topics": ["large-language-models", "natural-language-processing", "ai-research", "generative-ai", "machine-learning"], "entities": ["LUDI", "LUDI-7B", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/less-uniform-discrete-diffusion-is-more-powerful-and-scalable", "markdown": "https://wpnews.pro/news/less-uniform-discrete-diffusion-is-more-powerful-and-scalable.md", "text": "https://wpnews.pro/news/less-uniform-discrete-diffusion-is-more-powerful-and-scalable.txt", "jsonld": "https://wpnews.pro/news/less-uniform-discrete-diffusion-is-more-powerful-and-scalable.jsonld"}}