04:00
2026-10-09
machinebrief.com
machine-learning
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
A new test-time framework called World-Model Policy Arbiter (WMPA) raises the macro-average success rate on 18 state-based OGBench datasets from 44% to 58% by rolling out frozen goal-conditioned policβ¦