{"slug": "from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with", "title": "From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention", "summary": "A new reinforcement learning approach targets long-horizon robot manipulation by training on individual failed subtasks rather than collecting full-task demonstrations, according to the paper \"From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention.\" The method addresses the inefficiency of supervised fine-tuning, which requires operators to repeat behaviors a pretrained robot foundation policy already performs well. The work aims to let a pretrained policy overcome its few critical subtask failures with minimal human intervention.", "body_md": "A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning", "url": "https://wpnews.pro/news/from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with", "canonical_source": "https://aiflash.com/news/123942/", "published_at": "2026-09-21 18:00:00+00:00", "updated_at": "2026-09-21 18:23:25.950797+00:00", "lang": "en", "topics": ["robotics", "machine-learning", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with", "markdown": "https://wpnews.pro/news/from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with.md", "text": "https://wpnews.pro/news/from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with.txt", "jsonld": "https://wpnews.pro/news/from-pretraining-to-proficiency-real-world-subtask-rl-for-long-horizon-with.jsonld"}}