cd /news/robotics/google-deepmind-launches-gemini-robo… · home topics robotics article
[ARTICLE · art-81036] src=letsdatascience.com ↗ pub= topic=robotics verified=true sentiment=· neutral

Google DeepMind Launches Gemini Robotics 2

Google DeepMind launched Gemini Robotics 2 on July 30, adding whole-body humanoid control to its robotics AI stack. According to DeepMind, the vision-language-action model can translate vision and language into motor commands for motions from walking and crouching to two-handed object manipulation, while Gemini Robotics ER 2 handles high-level reasoning and multi-step task planning.

read4 min views1 publishedJul 30, 2026
Google DeepMind Launches Gemini Robotics 2
Image: Letsdatascience (auto-discovered)

Google DeepMind launched Gemini Robotics 2 on July 30, adding whole-body humanoid control to its robotics AI stack. According to DeepMind, the vision-language-action model can translate vision and language into motor commands for motions from walking and crouching to two-handed object manipulation, while Gemini Robotics ER 2 handles high-level reasoning and multi-step task planning.

Google DeepMind launched Gemini Robotics 2 on July 30, expanding its robotics AI from upper-body manipulation to whole-body control for humanoid robots. According to DeepMind, the new vision-language-action model converts visual and language inputs into motor control, allowing a robot to coordinate movements "from feet to fingertips."

The Verge reported that the new model can direct a humanoid across its full body, rather than primarily controlling the upper body as the prior version did. DeepMind demonstrated tasks including walking, crouching, stretching, moving objects, and cleaning a cluttered room. The Verge also reported that Google videos showed Apptronik's Apollo 2 robot retrieving specified objects from a shelf and handling a watering can.

Whole-body movement and dexterity

DeepMind describes Gemini Robotics 2 as a model for full humanoids and other bi-arm robots. The company also reported improved control of five-fingered hands and grippers, with demonstrations involving tasks such as sealing a plastic bag, tying a trash bag, and unscrewing a lightbulb.

The company stated that the model can run locally on-device and adapt to new robot bodies within a few hours. Those are important deployment claims, but public demonstrations do not by themselves establish reliability across different hardware, operating environments, task distributions, or safety constraints.

DeepMind acknowledged that robot movement speed still needs improvement. That limitation matters because physical systems must coordinate perception, planning, control, and recovery under real-time constraints, unlike software agents that can tolerate longer inference cycles.

ER 2 handles high-level task orchestration

Alongside the motor-control model, Google introduced Gemini Robotics ER 2, an embodied-reasoning vision-language model. According to Google's announcement, ER 2 can communicate with humans, interpret physical environments, plan multi-step tasks, and hand motor execution to a lower-level vision-language-action model.

Google states that ER 2 can process continuous video feeds to monitor progress, adapt after errors, and determine when to proceed to the next task step. The model can also call Google Search and user-defined functions, according to the company. Google made ER 2 available through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.

DeepMind also described multi-robot collaboration as part of the release. In a reported example, robots can work together on a room-cleaning task. Multi-robot coordination introduces additional engineering questions around task allocation, communication latency, collision avoidance, and failure handling that commonly arise in comparable physical-AI deployments.

A modular embodied-AI stack

The release separates high-level reasoning from low-level action generation: ER 2 is presented as the planning layer, while Gemini Robotics 2 produces motor behavior. In robotics, this modular pattern can let teams evaluate and replace planning, perception, and control components independently, although end-to-end performance still depends on the interfaces and timing between them.

For ML and robotics practitioners, the notable technical claim is not simply that a model can manipulate objects, but that DeepMind is extending language-and-vision-conditioned control to locomotion and coordinated whole-body behavior. Evaluation of such systems generally requires more than task-completion videos, including repeatability measures, recovery behavior, latency under local inference, performance across robot embodiments, and safety testing around people and objects. The launch places Gemini more directly in the embodied-AI market, where progress depends on combining multimodal foundation models with real-time control systems. DeepMind's public materials provide demonstrations and product claims, while broader independent benchmarks for whole-body performance were not included in the supplied announcements.

Key Points #

  • 1Gemini Robotics 2 expands control from upper-body actions to whole-body humanoid movement, according to DeepMind, widening the range of demonstrated tasks.
  • 2ER 2 separates high-level video-based planning from lower-level motor execution, relevant to robotics teams evaluating modular embodied-agent stacks.
  • 3DeepMind's on-device execution and cross-body adaptation claims make latency, transferability, and real-world safety validation central evaluation questions for developers.

Scoring Rationale #

The release extends a frontier multimodal model stack into whole-body humanoid control, a consequential capability area for embodied AI and robotics engineers. Its practical importance depends on independent evidence of reliability, safety, latency, and transfer across hardware, but the announced API access for ER 2 broadens developer relevance.

Sources #

Primary source and supporting public references used for this report.

Practice interview problems based on real data

1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.

Try 250 free problems

── more in #robotics 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-deepmind-laun…] indexed:0 read:4min 2026-07-30 ·