cd /news/artificial-intelligence/google-reveals-gemini-robotics-2-0-p… · home topics artificial-intelligence article
[ARTICLE · art-80714] src=arstechnica.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Google reveals Gemini Robotics 2.0, promising improved dexterity and safety

Google DeepMind released Gemini Robotics 2.0, a trio of AI sub-models that improve robot dexterity and safety, with the embodied reasoning model Gemini Robotics ER 2 now publicly available for developers via the Gemini Live API. The upgraded vision language model processes live video feeds to track task progress and classifies video frame completeness with nearly 60 percent accuracy, a significant leap over the previous 1.6 release.

read1 min views1 publishedJul 30, 2026
Google reveals Gemini Robotics 2.0, promising improved dexterity and safety
Image: Arstechnica (auto-discovered)

Robots powered by Google’s Gemini AI models are now more capable. With the debut of Gemini Robotics 2, these physical bots can now accomplish more complex tasks, continuously analyze changing environments, and collaborate with other robots. This is thanks to a trio of new sub-models, one of which is publicly available for developers starting today.

Videos of robots running, dancing, and backflipping have been a staple of the Internet for years, but these machines were programmed to perform these very narrow tasks. The goal of Gemini robotics is to create a generalist robot, one that can do anything a human could do. Google DeepMind scientists sometimes call this “physical AGI.” Essentially, you tell a robot what to do, and it does it. With the 2.0 release, Google says its robotics AI can control an entire humanoid robot with improved dexterity, even for machines with complex humanoid hands.

This starts with Gemini Robotics ER 2, an upgraded “embodied reasoning” model that DeepMind claims is a significant leap over the previous 1.6 release. It’s integrated with the Gemini Live API, giving developers the opportunity to experience that supposed leap forward.

Gemini Robotics ER 2 is what’s known as a vision language model (VLM). It’s designed to understand instructions and the world around it. The big upgrade here is that ER 2 can process live video feeds from the robot’s cameras, allowing the system to track progress as the robot lumbers from one step to the next. Google notes that Gemini Robotics ER 2 can classify video frame completeness with almost 60 percent accuracy. That’s still far from perfect, but it’s much better than the 1.6 release or what you can get with the visual understanding of competing AI models.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google deepmind 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/google-reveals-gemin…] indexed:0 read:1min 2026-07-30 ·