{"slug": "learning-to-plan-from-random-exploration", "title": "Learning to Plan from Random Exploration", "summary": "A new arXiv paper (2609.38383v1) presents a conditional energy-based model that learns temporal log-density ratios from random exploration, without action or reward labels, to enable long-range planning. Trained via noise-contrastive estimation on observation pairs, the model queries learned temporal relations at multiple horizons while a separate local dynamics model predicts candidate action outcomes, with the agent executing one action and replanning with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images, plus egocentric navigation and manipulation planning from suboptimal data, with learned score fields and planned routes exhibiting properties of a multiscale cognitive map.", "body_md": "arXiv:2609.38383v1 Announce Type: new \nAbstract: Random exploration reveals how an environment can be traversed before a goal is specified. Can this experience support long-range planning without policy-improvement training? Our random-walk analysis explains what temporal relations contain: short horizons reveal geodesic geometry in the diffusion limit, while longer horizons reveal connectivity between regions before mixing removes these distinctions. We learn these relations with a conditional energy-based model that estimates temporal log-density ratios through horizon-conditioned embeddings. The model is trained on observation pairs by noise-contrastive estimation, without action or reward labels. The planner queries these learned relations at different horizons as it moves toward the goal. At test time, a separate local dynamics model predicts candidate action outcomes, and the temporal model evaluates their progress toward the goal by selecting or aggregating estimated improvements across horizons. The agent executes one action and replans with both models fixed. Experiments demonstrate long-range maze planning from random exploration using states and images. Learned score fields, embedding probes, and planned routes exhibit properties of a multiscale cognitive map. We further demonstrate egocentric navigation from random exploration and manipulation planning from suboptimal data.", "url": "https://wpnews.pro/news/learning-to-plan-from-random-exploration", "canonical_source": "https://arxiv.org/abs/2609.38383", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:16:35.172706+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "robotics"], "entities": ["arXiv", "2609.38383v1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/learning-to-plan-from-random-exploration", "markdown": "https://wpnews.pro/news/learning-to-plan-from-random-exploration.md", "text": "https://wpnews.pro/news/learning-to-plan-from-random-exploration.txt", "jsonld": "https://wpnews.pro/news/learning-to-plan-from-random-exploration.jsonld"}}