- Gemini Robotics 2 is a three-model stack — a core VLA, an embodied reasoning layer (ER 2), and an on-device model — that controls full humanoid bodies including walking, crouching, and fine-motor manipulation [1] - The on-device model adapts to entirely new robot hardware with fewer than 200 demonstration examples in a matter of hours [2] - Hardware partners include Apptronik, Boston Dynamics, Agile Robots, and Franka Robotics; the system was demonstrated on Apptronik's Apollo 2 humanoid [1] - Google DeepMind introduced ASIMOV-Agentic, a safety benchmark that tests whether robots refuse dangerous commands and halt when humans enter their operational radius [2] - Gemini Robotics ER 2 is available now on Google AI Studio; the core VLA and on-device models are rolling out to early-access partners
[1] Google DeepMind on July 30 released Gemini Robotics 2, a suite of three AI models designed to give robots of all form factors — from tabletop arms to full-body humanoids — the ability to reason about and execute physical tasks. The system represents a significant step up from its predecessor, which was limited to upper-body tabletop manipulation [1].
The release comprises a core vision-language-action (VLA) model for motor control, Gemini Robotics ER 2 for high-level task planning and multi-robot coordination, and Gemini Robotics On-Device 2, an ultra-efficient variant optimized for local hardware. Together they form what Google DeepMind calls a unified "intelligence layer" for adaptive robots [2].
The company demonstrated the system on Apptronik's Apollo 2 humanoid, which performed whole-body tasks including walking, crouching, bending, and manipulating objects with five-fingered SharpaWave tactile hands. Hardware partners also include Boston Dynamics, Agile Robots, and Franka Robotics [1].
The Three-Model Stack #
Gemini Robotics 2's architecture splits robot intelligence into three layers. The core VLA model translates visual perception and natural-language instructions into motor commands, controlling full bipedal humanoids and dual-arm systems. Unlike previous VLAs that required separate policies for navigation and manipulation, this model unifies locomotion and dexterous control under a single framework [1].
Gemini Robotics ER 2 sits above the VLA as a strategic reasoning engine. It handles multi-step task decomposition, tracks progress across hundreds of sequential decisions, orchestrates multiple robots working together, and communicates with human operators. ER 2 is available now on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform [2].
The On-Device 2 model runs locally on robot silicon for latency-free operation. Its standout feature is rapid adaptation: it can learn to control an entirely new robot body with fewer than 200 demonstration examples in just hours, inheriting motion-transfer techniques from the earlier Gemini Robotics 1.5 [1].
Performance Benchmarks #
Google DeepMind published success rates across a range of tasks. On the Apollo 2 humanoid with Inspire hands, the system achieved 76.3% accuracy on shelf pick-ups, 68.4% on table pick-ups, and 45.7% on floor pick-ups — tasks that require coordinating full-body posture with arm reach [1].
Fine-motor dexterity benchmarks using SharpaWave's 22-degree-of-freedom tactile hands showed 92% success on unscrewing a light bulb, 44% on tying a trash bag, and 36% on screwing a bulb back in. On the Franka Duo dual-arm system with grippers, precise insertion tasks hit 89.6% and diverse tool kitting reached 78.9% [1].
The numbers illustrate both the system's capabilities and its current limitations — tasks requiring high torque precision or complex sequential manipulation remain challenging, though the floor pick-up and bulb-screwing results represent early benchmarks for whole-body humanoid VLAs [2].
Safety Framework #
Alongside the model release, Google DeepMind introduced ASIMOV-Agentic, a benchmark designed to evaluate whether embodied AI agents handle operational uncertainty safely. The benchmark tests whether the ER 2 reasoning layer refuses unsafe tool calls from the VLA, halts motion when confidence drops below thresholds, and proactively requests human intervention in ambiguous situations [1].
The system also includes human proximity detection that automatically triggers deceleration or full stop when co-workers enter a robot's operational radius. Google DeepMind framed the safety work as essential infrastructure for deploying autonomous robots in shared human environments [2].
Partners and Competitive Context #
The hardware partner roster — Apptronik, Boston Dynamics, Agile Robots, Franka Robotics, and third-party platforms including Dexmate, SO101, and Trossen — signals Google DeepMind's strategy of positioning Gemini Robotics as a platform-agnostic brain for the robotics industry rather than building its own hardware [2].
The launch places Google in direct competition with NVIDIA's Isaac GR00T framework for humanoid robot AI. Notably, Sharpa Robotics' SharpaWave tactile hand appears in both Google DeepMind's demos and NVIDIA's reference architecture, underscoring the emerging hardware-agnostic model that physical AI is adopting [2].
Developers interested in the core VLA and On-Device models can sign up through a Trusted Tester Program, while ER 2 is already accessible through Google AI Studio for immediate experimentation [1].
Companies mentioned #
Further sources #
The stories that matter, in one email. Free — unsubscribe anytime.