cd /news/machine-learning/noisy-space-policy-gradient-for-diff… · home topics machine-learning article
[ARTICLE · art-125469] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Researchers introduced a noisy-space action-value (Q-)function and a noisy-space policy gradient (NSPG) that trains diffusion policies for offline reinforcement learning without backpropagating through the denoising process, according to the arXiv paper 2609.06882v1. The method assigns values to diffusion latents via the distribution of executed actions and formulates a KL-regularized policy improvement over noisy latents with a diffusion-compatible regression form. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks show the noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.06882v1 Announce Type: cross Abstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. We show that this construction admits a precise semantic interpretation and derive a noisy-space policy gradient (NSPG) that optimizes noisy latents using only clean action-space value estimates. Building on this result, we formulate a KL-regularized policy improvement over noisy latents and show that the resulting objective admits a diffusion-compatible regression form, avoiding backpropagation through the denoising process. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks demonstrate that the proposed noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning. Project webpage: https://mahmoud-selim.github.io/NSPG/

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/noisy-space-policy-g…] indexed:0 read:1min 2026-09-10 ·