{"slug": "noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement", "title": "Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning", "summary": "Researchers introduced a noisy-space action-value (Q-)function and a noisy-space policy gradient (NSPG) that trains diffusion policies for offline reinforcement learning without backpropagating through the denoising process, according to the arXiv paper 2609.06882v1. The method assigns values to diffusion latents via the distribution of executed actions and formulates a KL-regularized policy improvement over noisy latents with a diffusion-compatible regression form. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks show the noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning.", "body_md": "arXiv:2609.06882v1 Announce Type: cross \nAbstract: Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion latents through the distribution of executed actions induced by the denoising process. We show that this construction admits a precise semantic interpretation and derive a noisy-space policy gradient (NSPG) that optimizes noisy latents using only clean action-space value estimates. Building on this result, we formulate a KL-regularized policy improvement over noisy latents and show that the resulting objective admits a diffusion-compatible regression form, avoiding backpropagation through the denoising process. Empirical results on state-based D4RL benchmarks and vision-based OGBench tasks demonstrate that the proposed noisy-space objective provides a principled and effective basis for training diffusion policies in offline reinforcement learning. Project webpage: https://mahmoud-selim.github.io/NSPG/", "url": "https://wpnews.pro/news/noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement", "canonical_source": "https://www.machinebrief.com/news/noisy-space-policy-gradient-for-diffusion-policies-in-offlin-3udz", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 05:52:11.262147+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "artificial-intelligence"], "entities": ["arXiv", "D4RL", "OGBench", "noisy-space policy gradient", "NSPG"], "alternates": {"html": "https://wpnews.pro/news/noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement", "markdown": "https://wpnews.pro/news/noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement.md", "text": "https://wpnews.pro/news/noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement.txt", "jsonld": "https://wpnews.pro/news/noisy-space-policy-gradient-for-diffusion-policies-in-offline-reinforcement.jsonld"}}