04:00
2026-09-24
arxiv.org
machine-learning
On Preference Coverage Collapse from Hindsight Relabeling in Multi-Objective Reinforcement Learning
A study of hindsight relabeling in preference-conditioned multi-objective reinforcement learning found the technique degrades 19 of 36 algorithm-environment settings by as much as four standard deviatβ¦