{"slug": "momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action", "title": "MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization", "summary": "Meta researchers introduced MoMo, a two-stage imitation-learning framework that enables robots to dial motion modes—steady, dynamic, or intermediate—during manipulation tasks. In tests across six real-robot tasks, varying the motion-mode condition produced distinguishable behaviors in joint speed, acceleration, and end-effector approach pitch, and MoMo transferred unseen requested modes while preserving task success. The findings demonstrate compositional generalization to unseen task–mode combinations, showing that motion mode can be reused across tasks to control manipulation skill execution.", "body_md": "[content type paper](/research/)published July 2026\n\nMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization\n\nAuthorsYuhan Hu, Hugues Thomas, Peide Huang, Mouli Sivapurapu, Benoit Landry, Arto Kivila\n\nMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization\n\nAuthorsYuhan Hu, Hugues Thomas, Peide Huang, Mouli Sivapurapu, Benoit Landry, Arto Kivila\n\nTo operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint speed, acceleration, and end-effector approach pitch. On tasks demonstrated in only one mode, MoMo transfers the unseen requested mode while largely preserving task success. Together, these results provide evidence of compositional generalization to unseen task–mode combinations and show that motion mode can be reused across tasks to control how a manipulation skill is performed.\n\nHumanoid Policy ~ Human Policy\n\nMay 21, 2025[research area Computer Vision](/research/?domain=Computer%20Vision), [research area Human-Computer Interaction](/research/?domain=Human-Computer%20Interaction)\n\nTraining manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot demonstrations is labor-intensive, requiring expensive tele-operated data collection which is difficult to scale. This paper investigates a more scalable data source, egocentric human demonstrations, to serve as cross-embodiment training data for robot learning. We…\n\nEfficient ConvBN Blocks for Transfer Learning and Beyond\n\nFebruary 12, 2024[research area Methods and Algorithms](/research/?domain=Methods%20and%20Algorithms)[conference ICLR](/research/?event=ICLR)\n\nConvolution-BatchNorm (ConvBN) blocks are integral components in various computer vision tasks and other domains. A ConvBN block can operate in three modes: Train, Eval, and Deploy. While the Train mode is indispensable for training models from scratch, the Eval mode is suitable for transfer learning and beyond, and the Deploy mode is designed for the deployment of models. This paper focuses on the trade-off between stability and efficiency in…", "url": "https://wpnews.pro/news/momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action", "canonical_source": "https://machinelearning.apple.com/research/momo-motion-mode-manipulation", "published_at": "2026-07-30 00:00:00+00:00", "updated_at": "2026-07-31 00:06:27.190138+00:00", "lang": "en", "topics": ["robotics", "machine-learning", "artificial-intelligence"], "entities": ["Meta", "MoMo", "Yuhan Hu", "Hugues Thomas", "Peide Huang", "Mouli Sivapurapu", "Benoit Landry", "Arto Kivila"], "alternates": {"html": "https://wpnews.pro/news/momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action", "markdown": "https://wpnews.pro/news/momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action.md", "text": "https://wpnews.pro/news/momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action.txt", "jsonld": "https://wpnews.pro/news/momo-dial-motion-mode-in-robot-manipulation-with-spatiotemporal-action.jsonld"}}