cd /news/artificial-intelligence/mdlmpe-distribution-aware-positional… · home topics artificial-intelligence article
[ARTICLE · art-87202] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

Researchers propose MDLMPE, the first positional encoding method for masked diffusion language models that explicitly tracks the evolving revealed/masked token configuration, improving performance over conventional encodings like RoPE. In experiments on LLaDA and DREAM, MDLMPE generally outperformed standard positional encodings across supervised fine-tuning, pretraining, zero-shot evaluation, and block-diffusion settings, with ablations confirming that combining availability state, Gaussian locality, spectral basis, and embedding injection yields the strongest results.

read1 min views1 publishedAug 5, 2026

arXiv:2608.03769v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundamentally from that of autoregressive (AR) models. Whereas AR decoding exposes a contiguous prefix, MDLM denoising produces dynamic, non-contiguous configurations of revealed and masked tokens. Conventional positional encodings such as RoPE capture sequence order and pairwise displacement but remain insensitive to this evolving token-availability structure. To address this limitation, we propose MDLMPE, a positional encoding designed specifically for masked diffusion. To the best of our knowledge, MDLMPE is the first method to make positional representations explicitly aware of the changing revealed/masked configuration. It represents token availability as a binary sequence, applies distance-aware Gaussian weighting, and projects the resulting pattern through a cosine basis to obtain distribution-aware positional features. These features are added to token embeddings and mapped by a lightweight MLP to angular offsets that modulate the standard RoPE phases. Extensive experiments on LLaDA and DREAM demonstrate that MDLMPE generally outperforms conventional positional encoding methods across supervised fine-tuning, pretraining, zero-shot evaluation, and block-diffusion settings. Further ablations show that the complete combination of availability state, Gaussian locality, spectral basis, and embedding injection yields the strongest result. These results establish the evolving token-availability distribution as a useful positional signal for masked diffusion language models.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @mdlmpe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mdlmpe-distribution-…] indexed:0 read:1min 2026-08-05 ·