cd /news/machine-learning/an-informational-curse-of-horizon-in… · home › topics › machine-learning › article
[ARTICLE · art-147397] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

An Informational Curse of Horizon in Goal-Conditioned Policy Learning

A new arXiv paper (2610.09247v1) identifies an "informational curse of horizon" in goal-conditioned policy learning, where increasing the goal relabeling horizon significantly reduces policy generalization and performance. Through controlled experiments with oracle planners, the authors show that goal-conditioned behavioral cloning (BC) policies suffer severe, training-horizon-dependent degradation even when evaluated only on nearby subgoals, a problem mitigated by reinforcement learning (RL) objectives. The authors attribute the effect to a horizon-dependent decrease in conditional mutual information between actions and hindsight-relabeled goals, and find that distilling input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks.

by read1 min views1 publishedOct 8, 2026

arXiv:2610.09247v1 Announce Type: new Abstract: The difficulty of learning goal-reaching policies is often attributed to a "curse of horizon" that manifests as bias accumulation in temporal-difference backups and noisy advantage estimates. In this work, we identify an additional informational curse of horizon in goal-conditioned policy learning, where increasing the goal relabeling horizon can significantly reduce policy generalization and performance. Through a series of controlled experiments with oracle planners, we decouple the goal horizons sampled during training from those that the policy is asked to reach at test time. Even when evaluated only on a sequence of nearby subgoals, goal-conditioned behavioral cloning (BC) policies suffer from severe, training horizon-dependent performance degradation that is mitigated by reinforcement learning (RL) objectives. We explain this phenomenon as a horizon-dependent decrease in the conditional mutual information between actions and hindsight-relabeled goals, and find empirically that both BC and RL policies trained on longer-horizon goals exhibit a shift in sensitivity from goal to state information, as measured by the policy's input Jacobians. Motivated by this observation, we find that distilling the input Jacobians of short-horizon policies into long-horizon policies yields significant performance gains, especially in combinatorial manipulation tasks. Taken together, our results highlight goal relabeling horizon as an important consideration when learning generalist policies from offline data.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/an-informational-cur…] indexed:0 read:1min 2026-10-08 · —