Introducing Gemini Robotics ER 2
Google DeepMind has launched Gemini Robotics ER 2, an advanced embodied reasoning model designed to act as a high-level brain for robots. The system allows robots to understand the physical world, chat with humans, use tools like Google Search, and plan multi-step tasks while handing off motor execution to lower-level vision-language-action models.
The model introduces significant upgrades in temporal intelligence by processing continuous video feeds. It achieves 57.4 percent accuracy in continuous progress classification across five completion levels, and 91.3 percent accuracy with a 0.96-second mean absolute distance in precision moment-finding. These capabilities enable robots to track their progress in real time, detect failures, self-correct, and know exactly when to transition between task steps.
Gemini Robotics ER 2 also supports multi-robot collaboration, allowing diverse machines in shared spaces to communicate through a shared semantic understanding to complete complex workflows. Additionally, the model features improved spatial reasoning, general instrument reading, and enhanced safety features that successfully enforce physical constraints and monitor human proximity.
The model is integrated into the Gemini Live API with a bidirectional streaming endpoint designed for latency-sensitive tasks, allowing robots to operate without stop-and-thought pauses. Gemini Robotics ER 2 is publicly available to developers through the Gemini API, Google AI Studio, and a private preview on the Gemini Enterprise Agent Platform.