{"slug": "r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning", "title": "$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning", "summary": "Researchers introduced R3, a post-training recipe that turns off-the-shelf vision-language models into robotic reasoners by mid-training on expert reasoning traces and improving with single-step rubric-based reinforcement learning from offline action data. In tests on Language Table and simulated bimanual grocery packing, R3 improved exploration and generalization across unseen tasks, significantly outperforming instruction-only imitation learning baselines.", "body_md": "arXiv:2608.26053v1 Announce Type: cross\nAbstract: Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. Whether this mechanism can improve robotic manipulation remains unclear, where long-horizon tasks require tracking partial progress, reasoning about object relations, recovering from mistakes, and steering noisy low-level policies. In this paper, we study whether VLMs can be trained to reason directly in natural language to guide low-level manipulation policies. We introduce $R^3$, a simple post-training recipe that turns off-the-shelf VLMs into robotic reasoners: it first mid-trains a VLM on expert-generated reasoning traces to initialize the desired reasoning style, then improves the reasoner with single-step rubric-based RL from offline action data. Unlike prior robotic reasoning methods that mostly use structured traces as auxiliary supervision, $R^3$ trains free-form language reasoning to produce test-time guidance for action. We instantiate $R^3$ on Language Table and simulated bimanual grocery packing, two controlled testbeds for studying robotic reasoning and long-horizon manipulation. $R^3$ improves exploration and generalization across unseen tasks and significantly outperforms instruction-only imitation learning baselines on both benchmarks. Our analyses suggest that free-form language reasoning can function as a test-time compute mechanism for steering low-level policies. Our project page is available at https://robotic-reasoner.github.io/.", "url": "https://wpnews.pro/news/r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning", "canonical_source": "https://www.machinebrief.com/news/dollarr3dollar-training-robots-to-reason-in-natural-language-zz9m", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 05:19:01.068406+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "robotics", "large-language-models"], "entities": ["R3", "Language Table"], "alternates": {"html": "https://wpnews.pro/news/r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning", "markdown": "https://wpnews.pro/news/r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning.md", "text": "https://wpnews.pro/news/r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/r-3-training-robots-to-reason-in-natural-language-via-reinforcement-learning.jsonld"}}