cd /news/large-language-models/unifusion-adapting-autoregressive-la… · home topics large-language-models article
[ARTICLE · art-76417] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

Researchers propose UNIFUSION, a continual pre-training approach that adapts pretrained autoregressive language models (e.g., GPT-2) to uniform-noise discrete diffusion under a unified reverse-rate objective. At 256 sampling steps, UNIFUSION-S and UNIFUSION-M achieve generative perplexity/entropy pairs of 97.783/5.2626 and 71.516/5.6669, respectively, outperforming all evaluated models of the same scale on both metrics. The method also achieves the highest accuracy on WinoGrande, SIQA, and BBH among compared diffusion models.

read1 min views1 publishedJul 28, 2026

arXiv:2607.24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diffusion, where every token remains editable during sampling. However, adapting AR checkpoints across corruption kernels remains challenging because existing DLMs use different objectives and prediction parameterizations. We establish connections among SEDD, MDLM/GIDD, M2S, and Neural CTMC by expressing their conditional losses as a single generalized Kullback--Leibler objective over model reverse rates. We further derive conversions from clean-token predictions to concrete-score, posterior-mean, and exit-rate/jump parameterizations, yielding a shared (x_0) interface that supports switching between mask and uniform kernels. Building on these connections, we propose \ours{}, a simple continual pre-training approach for directly adapting pretrained GPT2 checkpoints to uniform-noise diffusion. Through systematic evaluation of 124M- and 355M-parameter models, we show that \ours{} steadily improves the trade-off between generative perplexity (GenPPL) and unigram entropy as the sampling budget increases from 16 to 256 steps. At 256 steps, \ours{}-S and \ours{}-M achieve GenPPL/entropy pairs of (97.783/5.2626) and (71.516/5.6669), respectively; no evaluated model at the same scale simultaneously outperforms \ours{} on both metrics. At both scales, \ours{} also achieves the highest WinoGrande, SIQA, and BBH accuracy among the compared diffusion models.

── more in #large-language-models 4 stories · sorted by recency
── more on @unifusion 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/unifusion-adapting-a…] indexed:0 read:1min 2026-07-28 ·