mimile
Back to feed

Google DeepMind releases Gemini 2 vision-language-action model for humanoid robots

AI digest

This digest was compiled by AI from multiple sources — links to the originals are below.

Google DeepMind releases Gemini 2 vision-language-action model for humanoid robots

Google DeepMind announced the release of Gemini Robotics 2, an advanced vision-language-action model that powers humanoid robots. The model integrates multiple AI systems to enable whole-body control, multi-step reasoning, and robot-to-robot collaboration. The launch has renewed analyst calls for rigorous safety demonstrations before such robots are deployed alongside humans.

Whole-Body Control and Collaboration

Gemini Robotics 2 combines multiple AI models to achieve what DeepMind calls 'intelligent whole-body control,' allowing robots to perceive environments, reason through multi-step tasks, and coordinate movements across their entire bodies. The system also enables robots to collaborate with each other. Video demonstrations show Apptronik’s Apollo 2 robot and Sharpa robotic hands autonomously completing tasks such as tying trash bags and screwing in lightbulbs after training with teleoperation, video, and simulation. Carolina Parada, DeepMind’s robotics lead, described the long-term goal as enabling 'a robot to perform any task a human can.'

Safety and Reliability Challenges

Forrester analyst Paul Miller cautioned that while the launch advances robotics, safety remains a critical barrier. 'Robots are physical machines. They may be strong, and they may be heavy. People need to be confident that these machines are safe,' he said. Miller highlighted the need for failsafe mechanisms if sensors or power fail, noting that widespread deployment alongside humans depends on convincingly proven safety. He said that achieving physical AGI is aspirational and 'we’re nowhere near' a robot that can do anything a human can in any environment.

Industrial Debut Likely

Miller expects physical AI to first deliver value in controlled settings like factories and warehouses, where risks can be minimized. 'A factory or warehouse may be noisy, dirty, or chaotic, but it’s an awful lot simpler for a robot than a domestic environment,' he said. He argued that while physical AI may be more adaptable than traditional automation, it remains easier and cheaper to deploy where uncertainty is low and utilization is high. Google did not specify a timeline for commercial rollout.

What's Next

DeepMind is expected to continue refining the model through partner demos and further research, with no commercial release date announced. Whether the safety case for human-robot collaboration can be convincingly made in the near term remains uncertain.

1 source

Google DeepMind releases Gemini 2 vision-language-action model for humanoid robots