Ars Technica:AI(RSS)
65AI 编辑部评分,满分 100

Google 发布 Gemini Robotics 2.0,提升机器人灵巧性与安全性

2026-07-31 01:58· 38分钟前· Ryan Whitwam
跳到正文
AI 摘要

Google DeepMind 推出 Gemini Robotics 2,通过 Gemini Robotics ER 2 等三个新子模型为机器人带来全身智能。ER 2 可处理实时视频流,关键动作识别准确率近 90%,并支持多机器人协作。新安全基准 ASIMOV-Agentic 已开源至 Hugging Face,ER 2 模型即日起通过 Gemini Live API 向开发者开放。

Robots powered by Google’s Gemini AI models are now more capable. With the debut of Gemini Robotics 2, these physical bots can now accomplish more complex tasks, continuously analyze changing environments, and collaborate with other robots. This is thanks to a trio of new sub-models, one of which is publicly available for developers starting today.

Videos of robots running, dancing, and backflipping have been a staple of the Internet for years, but these machines were programmed to perform these very narrow tasks. The goal of Gemini robotics is to create a generalist robot, one that can do anything a human could do. Google DeepMind scientists sometimes call this “physical AGI.” Essentially, you tell a robot what to do, and it does it. With the 2.0 release, Google says its robotics AI can control an entire humanoid robot with improved dexterity, even for machines with complex humanoid hands.

Gemini Robotics 2 brings whole-body intelligence to robots.

This starts with Gemini Robotics ER 2, an upgraded “embodied reasoning” model that DeepMind claims is a significant leap over the previous 1.6 release. It’s integrated with the Gemini Live API, giving developers the opportunity to experience that supposed leap forward.

Ars Video

How The Callisto Protocol's Team Designed Its Terrifying, Immersive Audio

Gemini Robotics ER 2 is what’s known as a vision language model (VLM). It’s designed to understand instructions and the world around it. The big upgrade here is that ER 2 can process live video feeds from the robot’s cameras, allowing the system to track progress as the robot lumbers from one step to the next. Google notes that Gemini Robotics ER 2 can classify video frame completeness with almost 60 percent accuracy. That’s still far from perfect, but it’s much better than the 1.6 release or what you can get with the visual understanding of competing AI models.

Image 1

Gemini Robotics ER 2 is better at understanding the world than other models, but not by much.

Credit: Google

Gemini Robotics ER 2 is better at understanding the world than other models, but not by much. Credit: Google

Finding specific moments in video feeds is also key to completing a task correctly. When you ask a robot to pour a cup of coffee, you definitely want it to know when to stop pouring. ER 2 apparently does this much better, identifying key moments with almost 90 percent accuracy. As the robots execute multi-step tasks, the embodied reasoning model allows them to understand failures in real time. The system can then attempt that single step again rather than going back to the start. For example, the robot can just readjust its hand position and motion if a ball it’s trying to pick up rolls away or someone moves a container.

The new embodied reasoning release is also what gives Google DeepMind’s upgraded robot AI the ability to collaborate. The video demos show Apptronik’s Apollo 2 and the simpler Franka F3 Duo working together on a task without getting in each other’s way. While the test robots still can’t match the speed or grace of a human, the video demos include plenty of real-time footage of the robots in action, and they do seem much less hesitant than they were in past tests.

Multi-robot collaboration with Gemini Robotics 2.

Understanding is only the first step—getting the robot to move around in the physical world is the purview of another model. After mapping out the task, the vision language model hands things over to an upgraded vision-language-action known simply as Gemini Robotics 2. This AI model generates robot actions from those instructions in the same way other generative systems create text or images. There’s also a low-latency offline version of this called Gemini Robotics On-Device 2. These models are currently limited to a small group of testers.

Google is also testing Gemini Robotics 2 on Boston Dynamics hardware.

Google says the new action models are much more accurate and efficient. Even the smaller on-device version can adapt to new robot designs with just a few hours of movement data, or around 200 examples.

Heading off the robot apocalypse

The issues with AI hallucinations are well known at this point, but the potential harm from mistakes when the AI has a physical embodiment sharing space with humans could be much greater. With each release of Gemini Robotics, Google DeepMind has stressed that it takes this risk seriously. According to DeepMind, each layer in Gemini Robotics includes “traditional physical safety measures with robust AI safety frameworks.”

With the release of Gemini Robotics 2, there’s a new safety benchmark called ASIMOV-Agentic. This test evaluates models across a variety of safety factors. It can assess whether an embodied reasoning agent will refuse unsafe tool calls from a VLA. It can also determine whether a given task is possible to complete safely, as well as whether the model is able to call for human assistance if it’s unsure about safety.

Google DeepMind notes that Gemini Robotics ER 2 is the team’s safest model yet, showing robust ability to understand safety and halt actions when a human is too close to the robot. Google’s new safety benchmark is available in its entirety on Hugging Face.

Image 2: Photo of Ryan Whitwam

Ryan Whitwam is a senior technology reporter at Ars Technica, covering the ways Google, AI, and mobile technology continue to change the world. Over his 20-year career, he's written for Android Police, ExtremeTech, Wirecutter, NY Times, and more. He has reviewed more phones than most people will ever own. You can follow him on Bluesky, where you will see photos of his dozens of mechanical keyboards.

  1. Image 4: Listing image for first story in Most Read: Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission 1.Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission
  2. 2.Musk went to “war,” sought jail time for X ad boycotts—but case ends with a whimper
  3. 3.Actually, Starliner might fly into space this year
  4. 4.Comcast store punished low sales by smashing pies in workers' faces, lawsuit claims
  5. 5.Elon Musk’s xAI is trying to sue its way out of a Grok reckoning

Image 8

Google 发布 Gemini Robotics 2.0,提升机器人灵巧性与安全性

Ars Technica:AI(RSS)·2026-07-31 01:58·38分钟前· Ryan Whitwam
阅读原文· arstechnica.com
AI 摘要

Google DeepMind 推出 Gemini Robotics 2,通过 Gemini Robotics ER 2 等三个新子模型为机器人带来全身智能。ER 2 可处理实时视频流,关键动作识别准确率近 90%,并支持多机器人协作。新安全基准 ASIMOV-Agentic 已开源至 Hugging Face,ER 2 模型即日起通过 Gemini Live API 向开发者开放。

原文 · 保持原样,未翻译

Robots powered by Google’s Gemini AI models are now more capable. With the debut of Gemini Robotics 2, these physical bots can now accomplish more complex tasks, continuously analyze changing environments, and collaborate with other robots. This is thanks to a trio of new sub-models, one of which is publicly available for developers starting today.

Videos of robots running, dancing, and backflipping have been a staple of the Internet for years, but these machines were programmed to perform these very narrow tasks. The goal of Gemini robotics is to create a generalist robot, one that can do anything a human could do. Google DeepMind scientists sometimes call this “physical AGI.” Essentially, you tell a robot what to do, and it does it. With the 2.0 release, Google says its robotics AI can control an entire humanoid robot with improved dexterity, even for machines with complex humanoid hands.

Gemini Robotics 2 brings whole-body intelligence to robots.

This starts with Gemini Robotics ER 2, an upgraded “embodied reasoning” model that DeepMind claims is a significant leap over the previous 1.6 release. It’s integrated with the Gemini Live API, giving developers the opportunity to experience that supposed leap forward.

Ars Video

How The Callisto Protocol's Team Designed Its Terrifying, Immersive Audio

Gemini Robotics ER 2 is what’s known as a vision language model (VLM). It’s designed to understand instructions and the world around it. The big upgrade here is that ER 2 can process live video feeds from the robot’s cameras, allowing the system to track progress as the robot lumbers from one step to the next. Google notes that Gemini Robotics ER 2 can classify video frame completeness with almost 60 percent accuracy. That’s still far from perfect, but it’s much better than the 1.6 release or what you can get with the visual understanding of competing AI models.

Image 1

Gemini Robotics ER 2 is better at understanding the world than other models, but not by much.

Credit: Google

Gemini Robotics ER 2 is better at understanding the world than other models, but not by much. Credit: Google

Finding specific moments in video feeds is also key to completing a task correctly. When you ask a robot to pour a cup of coffee, you definitely want it to know when to stop pouring. ER 2 apparently does this much better, identifying key moments with almost 90 percent accuracy. As the robots execute multi-step tasks, the embodied reasoning model allows them to understand failures in real time. The system can then attempt that single step again rather than going back to the start. For example, the robot can just readjust its hand position and motion if a ball it’s trying to pick up rolls away or someone moves a container.

The new embodied reasoning release is also what gives Google DeepMind’s upgraded robot AI the ability to collaborate. The video demos show Apptronik’s Apollo 2 and the simpler Franka F3 Duo working together on a task without getting in each other’s way. While the test robots still can’t match the speed or grace of a human, the video demos include plenty of real-time footage of the robots in action, and they do seem much less hesitant than they were in past tests.

Multi-robot collaboration with Gemini Robotics 2.

Understanding is only the first step—getting the robot to move around in the physical world is the purview of another model. After mapping out the task, the vision language model hands things over to an upgraded vision-language-action known simply as Gemini Robotics 2. This AI model generates robot actions from those instructions in the same way other generative systems create text or images. There’s also a low-latency offline version of this called Gemini Robotics On-Device 2. These models are currently limited to a small group of testers.

Google is also testing Gemini Robotics 2 on Boston Dynamics hardware.

Google says the new action models are much more accurate and efficient. Even the smaller on-device version can adapt to new robot designs with just a few hours of movement data, or around 200 examples.

Heading off the robot apocalypse

The issues with AI hallucinations are well known at this point, but the potential harm from mistakes when the AI has a physical embodiment sharing space with humans could be much greater. With each release of Gemini Robotics, Google DeepMind has stressed that it takes this risk seriously. According to DeepMind, each layer in Gemini Robotics includes “traditional physical safety measures with robust AI safety frameworks.”

With the release of Gemini Robotics 2, there’s a new safety benchmark called ASIMOV-Agentic. This test evaluates models across a variety of safety factors. It can assess whether an embodied reasoning agent will refuse unsafe tool calls from a VLA. It can also determine whether a given task is possible to complete safely, as well as whether the model is able to call for human assistance if it’s unsure about safety.

Google DeepMind notes that Gemini Robotics ER 2 is the team’s safest model yet, showing robust ability to understand safety and halt actions when a human is too close to the robot. Google’s new safety benchmark is available in its entirety on Hugging Face.

Image 2: Photo of Ryan Whitwam

Ryan Whitwam is a senior technology reporter at Ars Technica, covering the ways Google, AI, and mobile technology continue to change the world. Over his 20-year career, he's written for Android Police, ExtremeTech, Wirecutter, NY Times, and more. He has reviewed more phones than most people will ever own. You can follow him on Bluesky, where you will see photos of his dozens of mechanical keyboards.

  1. Image 4: Listing image for first story in Most Read: Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission 1.Mythos attack on 3rd-round PQC algorithm candidate puts it out of commission
  2. 2.Musk went to “war,” sought jail time for X ad boycotts—but case ends with a whimper
  3. 3.Actually, Starliner might fly into space this year
  4. 4.Comcast store punished low sales by smashing pies in workers' faces, lawsuit claims
  5. 5.Elon Musk’s xAI is trying to sue its way out of a Grok reckoning

Image 8

阅读原文arstechnica.com