Google DeepMind published new details regarding Gemini Robotics ER 2 on July 30, 2026, outlining a platform focused on powering robotics. The publication, released at 15:00 UTC, identifies video understanding, task orchestration, and multi-robot collaboration as the core technical pillars of the system.
The publication from Google DeepMind places video understanding at the center of Gemini Robotics ER 2. Under the framework presented by Google DeepMind, video processing provides visual comprehension to inform robotic perception and physical decision-making. Google DeepMind did not disclose specific technical benchmarks, supported frame rates, visual resolution limits, or model inference speeds for the video understanding feature.
Task orchestration serves as a second major capability of Gemini Robotics ER 2, according to Google DeepMind. The framework handles operational workflows and action sequencing across robotic control systems. Google DeepMind did not specify the operating systems, programming interfaces, or software toolkits required to deploy task orchestration on physical machinery.
Multi-robot collaboration addresses joint operations between separate machines working within the same physical environment. Google DeepMind described collaboration mechanisms designed to allow multiple devices to divide and execute shared workloads. Google DeepMind did not state the maximum number of connected units supported by the system or the networking standards used for multi-robot messaging.
The developer did not publish hardware requirements, commercial availability dates, or licensing models for Gemini Robotics ER 2. The organization gave no information regarding API pricing, enterprise deployment options, or existing real-world deployments.
