Google DeepMind launched Gemini Robotics ER 2 on July 30, 2026, introducing a model designed to serve as a high-level reasoning engine for physical robots. The model handles video understanding, real-time spatial reasoning, task orchestration, and multi-robot collaboration. Developers can access the system through the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform.
The system processes continuous video, text, and audio streams to plan multi-step tasks while handing off lower-level motor execution to vision-language-action models or robotics control interfaces. It can also call external tools natively, including Google Search and user-defined functions. Integration with the Gemini Live API provides a bidirectional streaming endpoint that allows robots to determine upcoming actions without pausing during physical execution.
In performance evaluations across real vision-language-action control, simulated control, and human tele-operation, Google DeepMind reported that Gemini Robotics ER 2 consistently outperformed the previous ER 1.6 model. To track task progress, the system assigns video feed frames across five level bands between 0% and 100%. The model achieved 57.4 percent accuracy on progress classification tasks.
For moment-finding tasks that identify the exact video frame where a critical event occurs, the model reached 91.3 percent accuracy with a 0.96-second mean absolute distance. Google DeepMind demonstrated this orchestration capability using Boston Dynamics' Spot robot, which fetched a popcorn snack based on natural language commands by executing navigation and manipulator movement APIs.
Gemini Robotics ER 2 also supports multi-robot coordination across different hardware models, including shared workflows between Apptronik's Apollo 2 and Franka F3 Duo machines. For spatial intelligence, the model evaluates raw video feeds to identify mid-execution errors such as spills or slips and reads 10 instrument types, including liquid thermometers, rulers, linear scales, and digital displays. Safety evaluations showed that the model automatically halts a humanoid robot when a human enters its immediate vicinity and resumes operation only after the area clears.
