The embodied reasoning model can monitor continuous video, coordinate multiple robots, call external tools, and plan tasks while actions are being performed.
Google has launched Gemini Robotics ER 2, an embodied reasoning model that helps robots understand their surroundings, communicate with people, and complete complex physical tasks in real time.
The model acts as a high level brain, planning tasks before passing physical execution to a vision language action model or another robot control system.
Gemini Robotics ER 2 can process continuous video, audio, and text while using tools such as Google Search, navigation systems, and developer defined functions. It can plan its next step while a robot is still acting, reducing s between decisions.
The model is available through the Gemini API and Google AI Studio, with a private preview offered through the Gemini Enterprise Agent Platform.
Gemini Robotics ER 2 can track task progress, identify mistakes, and adjust actions without restarting an entire workflow.
Google said the model achieved 57.4% accuracy in progress classification tests, which measure how well it estimates how much of a task has been completed.
It also reached 91.3% accuracy in moment finding, which tests whether a robot can identify the exact point when an event occurs, such as when to stop pouring liquid. Google reported an average timing distance of 0.96 seconds and four times faster execution than larger model categories.
The model connects to the Gemini Live API through a two way streaming endpoint designed for low latency robotics.
Google demonstrated the system using Boston Dynamics’ Spot robot, which followed a spoken request to locate, collect, and return an object.
Gemini Robotics ER 2 also supports collaboration between multiple robots. Google demonstrated the feature using Apptronik’s Apollo 2 and the Franka F3 Duo, allowing the machines to divide and hand off tasks based on their capabilities.
The model also improves failure detection, instrument reading, spatial reasoning, and safety.
Google said it can detect spills, slips, and misplaced objects from live video and interpret instruments including digital displays, rulers, scales, thermometers, and circular dials.
In one safety test, the model stopped a humanoid robot when a person entered its work area and resumed the task once the area was clear.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our