← ALL RELEASES

GOOGLE · 30 Jul 2026

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Google DeepMind has released Gemini Robotics ER 2, an embodied reasoning model designed to act as a high-level brain for robots. The model handles real-time spatial reasoning, multi-step task planning, and tool orchestration, passing motor execution commands to lower-level vision-language-action models or robotics APIs. It integrates into the Gemini Live API with a bidirectional streaming endpoint to support latency-sensitive tasks without pauses.

The model improves video understanding and progress tracking through two foundational capabilities. It achieves 57.4 percent accuracy on continuous progress classification across five completion levels, and 91.3 percent accuracy on precision moment-finding with a 0.96-second mean absolute distance. By watching continuous video feeds, robots can track their progress, adapt to errors, and verify task completion in real time.

Gemini Robotics ER 2 also introduces multi-robot collaboration, allowing diverse machines like wheeled rovers and humanoids to share a semantic understanding and coordinate on complex workflows in shared spaces. It features improved general spatial intelligence across success and failure detection on raw video, generalized instrument reading across ten instrument types, and enhanced visual question answering.

The model includes safety upgrades, posting gains on Safety Instruction Following and Human Proximity benchmarks. It can enforce physical constraints, monitor environments, and autonomously halt a robot when a person is nearby, resuming work only when the area is clear. Gemini Robotics ER 2 is available to developers through the Gemini API, Google AI Studio, and a private preview on the Gemini Enterprise Agent Platform.

Read the original ↗