cd /news/artificial-intelligence/fine-tuning-diffusion-language-model… · home › topics › artificial-intelligence › article
[ARTICLE · art-142973] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Fine-Tuning Diffusion Language Models with Context Selection and Target Weighting

GoldiMask, a fine-tuning method for discrete diffusion language models introduced in arXiv paper 2609.38385v1, selects which response tokens to reveal as context by approximately maximizing a submodular objective and then weights the remaining prediction targets by their benefit from that context and remaining learning potential. Across three backbones and three training datasets, GoldiMask achieved the highest average accuracy in most evaluated settings, with gains on reasoning and code generation, and component ablations showed both context selection and target weighting contribute to the improvements. GoldiMask also reduced decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding while maintaining comparable accuracy at higher confidence thresholds.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38385v1 Announce Type: new Abstract: Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @goldimask 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fine-tuning-diffusio…] indexed:0 read:1min 2026-10-01 · —