cd /news/large-language-models/large-language-continuous-diffusion-… · home › topics › large-language-models › article
[ARTICLE · art-145172] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Large Language Continuous Diffusion Models

Researchers introduced Sigma, described as the first large-scale continuous diffusion language model at 3B and 8B parameters, built on steerable low-dimensional ODE/SDE latent trajectories and trained blockwise via likelihood optimization. Sigma reaches competitive performance with state-of-the-art discrete masked diffusion and autoregressive baselines on GSM8K, Minerva, HumanEval and MBPP after pre-training, and on MATH-500 and AIME after supervised fine-tuning, while classifier-free guidance and score temperature proved essential for high-fidelity reasoning and coding. The work reports that embedding-space steering governs the quality-diversity trade-off for strong pass@k performance and that continuous trajectories enable graceful degradation at low numbers of function evaluations and efficient distillation.

by read1 min views1 publishedOct 5, 2026

arXiv:2610.02665v1 Announce Type: new Abstract: Despite the success of discrete diffusion language models (dLMs) for fast parallel decoding, their non-smooth, high-dimensional space hinders trajectory steering for reasoning and inference acceleration. To overcome this, we present Sigma, the first large-scale (3B/8B) continuous dLM built on steerable, low-dimensional ODE/SDE latent trajectories. Trained blockwise via likelihood optimization, Sigma jointly denoises Gaussian-corrupted token embeddings while learning an optimal embedding geometry. To accelerate training, Sigma leverages pre-trained weights from autoregressive (AR) models for warm-starting. During inference, we identify classifier-free guidance and score temperature as essential for high-fidelity reasoning and coding. Across comprehensive math reasoning and coding evaluations against state-of-the-art discrete counterparts (masked dLMs and AR baselines), Sigma achieves competitive performance with discrete models on standard benchmarks (e.g., GSM8K, Minerva, HumanEval, MBPP) after pre-training and on challenging reasoning tasks (e.g., MATH-500, AIME) after supervised fine-tuning. Beyond performance parity, we uncover key structural properties unique to continuous dLMs: (i) embedding-space steering effectively governs the quality-diversity trade-off, yielding strong pass@k performance and (ii) continuous trajectories enable graceful degradation for low NFEs and efficient distillation. These establish continuous dLMs as a promising paradigm for efficient language generation.

── more in #large-language-models 4 stories · sorted by recency
── more on @sigma 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/large-language-conti…] indexed:0 read:1min 2026-10-05 · —