{"slug": "google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety", "title": "Google reveals Gemini Robotics 2.0, promising improved dexterity and safety", "summary": "Google DeepMind released Gemini Robotics 2.0, a trio of AI sub-models that improve robot dexterity and safety, with the embodied reasoning model Gemini Robotics ER 2 now publicly available for developers via the Gemini Live API. The upgraded vision language model processes live video feeds to track task progress and classifies video frame completeness with nearly 60 percent accuracy, a significant leap over the previous 1.6 release.", "body_md": "Robots powered by Google’s Gemini AI models are now more capable. With the debut of [Gemini Robotics 2](https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/), these physical bots can now accomplish more complex tasks, continuously analyze changing environments, and collaborate with other robots. This is thanks to a trio of new sub-models, one of which is [publicly available](https://ai.google.dev/gemini-api/docs/robotics-overview) for developers starting today.\n\nVideos of robots running, dancing, and backflipping have been a staple of the Internet for years, but these machines were programmed to perform these very narrow tasks. The goal of Gemini robotics is to create a generalist robot, one that can do anything a human could do. Google DeepMind scientists sometimes call this “physical AGI.” Essentially, you tell a robot what to do, and it does it. With the 2.0 release, Google says its robotics AI can control an entire humanoid robot with improved dexterity, even for machines with complex humanoid hands.\n\nThis starts with Gemini Robotics ER 2, an upgraded “embodied reasoning” model that DeepMind claims is a significant leap over the [previous 1.6 release](https://arstechnica.com/ai/2026/04/robot-dogs-now-read-gauges-and-thermometers-using-google-gemini/). It’s integrated with the Gemini Live API, giving developers the opportunity to experience that supposed leap forward.\n\nGemini Robotics ER 2 is what’s known as a vision language model (VLM). It’s designed to understand instructions and the world around it. The big upgrade here is that ER 2 can process live video feeds from the robot’s cameras, allowing the system to track progress as the robot lumbers from one step to the next. Google notes that Gemini Robotics ER 2 can classify video frame completeness with almost 60 percent accuracy. That’s still far from perfect, but it’s much better than the 1.6 release or what you can get with the visual understanding of competing AI models.", "url": "https://wpnews.pro/news/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety", "canonical_source": "https://arstechnica.com/ai/2026/07/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety/", "published_at": "2026-07-30 17:58:02+00:00", "updated_at": "2026-07-30 18:14:53.423669+00:00", "lang": "en", "topics": ["artificial-intelligence", "robotics", "ai-products", "computer-vision", "ai-research"], "entities": ["Google DeepMind", "Gemini Robotics 2.0", "Gemini Robotics ER 2", "Gemini Live API"], "alternates": {"html": "https://wpnews.pro/news/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety", "markdown": "https://wpnews.pro/news/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety.md", "text": "https://wpnews.pro/news/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety.txt", "jsonld": "https://wpnews.pro/news/google-reveals-gemini-robotics-2-0-promising-improved-dexterity-and-safety.jsonld"}}