cd /news/large-language-models/daca-grpo-denoising-aware-credit-ass… · home topics large-language-models article
[ARTICLE · art-131998] src=machinelearning.apple.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Researchers affiliated with The Ohio State University and Apple proposed DACA-GRPO, a plug-and-play enhancement to GRPO-style trainers for diffusion language models that adds Denoising Progress Scores and Stratified Masking Likelihood to fix the absence of temporal credit assignment and mean-field likelihood bias. Applied on top of three GRPO base methods, DACA-GRPO improved results across seven benchmarks, with gains of up to 5.6 percentage points on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.

read1 min views17 publishedSep 16, 2026
DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models
Image: Apple ML Research

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO introduces two complementary mechanisms: Denoising Progress Scores, which extract per-token importance weights from intermediate predictions at no additional forward cost, and Stratified Masking Likelihood, which partitions token positions into strata so that each token is predicted with most of the sequence as context, reducing the mean-field bias. Applied on top of three GRPO base methods, DACA-GRPO achieves consistent improvements across seven benchmarks spanning mathematical reasoning, code generation, constraint satisfaction, and constrained generation, with gains of up to 5.6pp on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.

  • ‡ Equal contribution
  • † The Ohio State University
  • ** Work done while at Apple
── more in #large-language-models 4 stories · sorted by recency
── more on @daca-grpo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/daca-grpo-denoising-…] indexed:0 read:1min 2026-09-16 ·