Google DeepMind last week unveiled Gemini Robotics 2, the latest version of its vision-language-action, or VLA, model.
With the first version, the company showed how Gemini’s multimodal understanding could drive real-world action. The updated version includes intelligent whole-body control, advanced dexterity, and multi-robot collaboration, DeepMind said.
Gemini Robotics 2 enables robots to reason through every movement, unlocking a broad range of tasks. For example, it could enable a humanoid to walk, crouch, stretch, and manipulate objects to clean up a cluttered room. The robot could also team up with other robots to finish the job faster.
Gemini Robotics 2 can also run locally on-device while adapting to entirely new robotic bodies in just a few hours, claimed DeepMind. In addition to the VLA, the company also released:
Gemini Robotics ER 2: This embodied reasoning (ER) model is a vision language model (VLM) that acts as DeepMind’s agent, enabling robots to communicate with people, understand the physical world, and plan multi-step tasks lasting several minutes. DeepMind is also introducing the ability for robots to work together as a team.Gemini Robotics On-Device 2: The company optimized this VLA to run locally on robotic devices. The model can now achieve fast adaptation to completely new robot embodiments with a few hours of data, DeepMind said.
Gemini Robotics ER 2 is now available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The VLA and On-Device models are available to early-access partners.
Gemini Robotics 2 manages the entire robot body #
DeepMind’s previous models controlled the humanoid’s upper body to achieve tabletop tasks. Now, Gemini Robotics 2 is expanding physical AI into whole-body motions.
The model can now control entire humanoid robots, translating intent into intelligent whole-body control. For example, when controlling Apptronik‘s Apollo 2 humanoid robot, users can ask it to “put the watering can into the green bin in the bottom shelf.”
Apollo processes the instruction, walks to the table, and picks up the watering can, takes a few steps to the shelves, and places it precisely in its destination. While DeepMind said robots have more to advance in movement speed, this is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination.
To be genuinely useful in our homes and workplaces, DeepMind said robots need finesse. Gemini Robotics 2 unlocks a new level of physical dexterity across different end effectors, whether a robot is using hands or grippers.
The model can now control the five-fingered, 22 degree-of-freedom SharpaWave hand on the Apollo 2 robot to complete delicate actions like tying knots or sealing a ziplock bag. It can also operate standard two-fingered parallel grippers on a Franka Duo platform to perform complex dexterous tasks such as tight packing.
DeepMind manages complex tasks that require multiple robots #
Most real-world tasks require multiple steps over an extended period of time, said Google DeepMind. To manage this complexity, Gemini Robotics ER 2 serves as the robot’s high-level brain, processing user instructions and communicating with humans.
It observes the room, reasons about the steps needed to complete the task, coordinates with the VLA to carry out the actions, and tracks progress until the task is done. This setup allows robots to execute complex multi-step tasks, self-correct if a step fails, and generalize to novel situations and goals.
With the update, DeepMind said it is enabling robots to more reliably execute longer task sequences, lasting several minutes and involving hundreds of decisions. Gemini Robotics ER 2 now understands when tasks begin and end, and it can pinpoint the moment key events occur, marking a step change in progress understanding.
The company is also introducing multi-robot collaboration. This enables different types of robots to communicate and work together to solve complex workflows a single robot could not do alone.
Gemini Robotics On-Device 2 targets applications with low connectivity #
Many robotic applications need to operate without network latency or internet connectivity. DeepMind built Gemini Robotics On-Device 2 to handle these constraints.
This model is natively multi-embodiment and inherits DeepMind’s “motion transfer” techniques from Gemini Robotics 1.5. The model can now adapt to new bi-arm robot embodiments with just a few hours of adaptation time, typically with less than 200 examples.
This works even with new embodiments with drastically different shapes, sensors, and degrees of freedom, DeepMind claimed.
DeepMind reaffirms its commitment to safety #
Google DeepMind said safety is foundational to its robotics research. Gemini Robotics 2 specifically advances robotics safety for navigating the uncertainty of the real world and collaborating alongside humans, the company said.
DeepMind introduced ASIMOV-Agentic, a new benchmark for agentic safety orchestration and uncertainty resolution. For example, it measures the embodied reasoning agent’s ability to refuse unsafe tool calls from a VLA. It also measures the agent’s ability to predict whether a task is possible and to proactively request human intervention when uncertain.
Additionally, with enhanced embodied reasoning, DeepMind said Gemini Robotics ER 2 is its safest robotics model to date in safety constraint following and human proximity benchmarks. It can better detect when humans are nearby, trigger safety tool calls, and bring the robot to a safe stop if someone approaches too closely. This is a key requirement in collaborative safety standards, said DeepMind.