{"slug": "explorenet-learning-where-to-explore-in-diffusion-grpo", "title": "ExploreNet: Learning Where to Explore in Diffusion GRPO", "summary": "EXPLORENET, a learned adaptive exploration policy for group-relative diffusion RL, improves held-out GenEval2 by 14% over Flow-GRPO on Stable Diffusion 3.5 Medium and reaches a 67.2% human preference win-rate, according to the arXiv paper 2609.38329v1. The method predicts a per-latent-element noise scale from the current latent, denoising step and prompt before any reward is observed, is trained on each rollout group's reward spread, and is discarded after training so inference is unchanged. The authors report that exploration is learnable, that the shape of the exploration distribution outweighs its magnitude, and that rollout quality is more effective than rollout quantity.", "body_md": "arXiv:2609.38329v1 Announce Type: new \nAbstract: Group-relative RL methods such as Flow-GRPO post-train image generators by exploring with isotropic Gaussian noise added at every denoising step. This noise decides which rollouts the model learns from, yet it perturbs every channel and spatial position of the latent equally. In this paper, we instead show that latent elements differ in how much they change the generated image, so exploration should adapt to these differences. We introduce EXPLORENET to learn an adaptive exploration distribution. EXPLORENET is a policy that predicts a noise scale for every latent element from the current latent, the denoising step, and the prompt, before any reward is observed; it is trained on the reward spread of each rollout group and discarded after training, leaving inference unchanged. On Stable Diffusion 3.5 Medium, EXPLORENET improves held-out GenEval2 by 14% over Flow-GRPO, transfers to two independent compositional benchmarks and five preference and image-quality models, and reaches a 67.2% human preference win-rate. Overall, across our group-relative diffusion RL experiments, we find that exploration is learnable, the shape of the exploration distribution outweighs its magnitude, and rollout quality is more effective than rollout quantity.", "url": "https://wpnews.pro/news/explorenet-learning-where-to-explore-in-diffusion-grpo", "canonical_source": "https://arxiv.org/abs/2609.38329", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:19:57.634545+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "ai-research", "large-language-models"], "entities": ["EXPLORENET", "Flow-GRPO", "Stable Diffusion 3.5 Medium", "GenEval2", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/explorenet-learning-where-to-explore-in-diffusion-grpo", "markdown": "https://wpnews.pro/news/explorenet-learning-where-to-explore-in-diffusion-grpo.md", "text": "https://wpnews.pro/news/explorenet-learning-where-to-explore-in-diffusion-grpo.txt", "jsonld": "https://wpnews.pro/news/explorenet-learning-where-to-explore-in-diffusion-grpo.jsonld"}}