04:00
2026-10-01
arxiv.org
artificial-intelligence
ExploreNet: Learning Where to Explore in Diffusion GRPO
EXPLORENET, a learned adaptive exploration policy for group-relative diffusion RL, improves held-out GenEval2 by 14% over Flow-GRPO on Stable Diffusion 3.5 Medium and reaches a 67.2% human preference β¦