{"slug": "arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa", "title": "ARC-Bench: Closed-Loop Replanning Masks Broken Action Ranking in Frozen JEPA World Models", "summary": "A new arXiv paper (2609.05461v1) introduces ARC-Bench, a no-leak, fixed-candidate protocol that finds frozen JEPA-style world models rank candidate actions incorrectly, with the top-scored candidate almost always suboptimal on official manipulation audits and the same inversion appearing in maze domains. The defect persists when DINOv2 is replaced by video-pretrained V-JEPA 1 and V-JEPA 2 encoders at ViT-L/ViT-G scale, and the authors show closed-loop replanning masks it: reducing the planner's replanning frequency collapses success in both navigation and manipulation domains. The paper concludes that closed-loop success rates systematically overstate the rankability of frozen latent representations.", "body_md": "arXiv:2609.05461v1 Announce Type: new \nAbstract: Reward-free latent world models plan by scoring candidate actions with distances in a frozen latent space: an action is preferred if its predicted future embedding lands closer to the goal embedding. This silently assumes that latent closeness is action-rankable, i.e., that ordering candidates by latent distance agrees with ordering them by true cost. We audit this assumption directly. We introduce ARC-Bench, a no-leak, fixed-candidate protocol that measures whether frozen JEPA-style objectives rank candidate actions correctly, and apply it to official released JEPA-WM checkpoints across navigation and manipulation-style control. The assumption fails, severely and structurally: on the official manipulation audits the top-scored candidate is almost always suboptimal, and the same inversion appears in the maze domains. A controlled visual-backbone extension shows that the defect persists when DINOv2 is replaced by video-pretrained V-JEPA 1 and V-JEPA 2 encoders at ViT-L/ViT-G scale. Provenance, undertraining, matched-budget backbone controls, and metric-circularity controls rule out trivial explanations. We then explain why this defect has stayed invisible: closed-loop replanning masks it. When we reduce the planner's replanning frequency, success collapses in both a navigation and a manipulation domain, and the episodes rescued by frequent replanning are enriched for severe first-plan ranking failures in the PointMaze first-plan diagnostic. Closed-loop success rates therefore systematically overstate the rankability of frozen latent representations. ARC-Bench supplies the measurement, and the masking mechanism the explanation, for methods that adapt, amortize, or replan around latent-space planners without directly auditing released JEPA-WM action rankability.", "url": "https://wpnews.pro/news/arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa", "canonical_source": "https://arxiv.org/abs/2609.05461", "published_at": "2026-09-10 04:00:00+00:00", "updated_at": "2026-09-10 04:21:11.103717+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "robotics", "autonomous-vehicles", "ai-safety"], "entities": ["ARC-Bench", "JEPA", "JEPA-WM", "DINOv2", "V-JEPA 1", "V-JEPA 2", "PointMaze", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa", "markdown": "https://wpnews.pro/news/arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa.md", "text": "https://wpnews.pro/news/arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa.txt", "jsonld": "https://wpnews.pro/news/arc-bench-closed-loop-replanning-masks-broken-action-ranking-in-frozen-jepa.jsonld"}}