HorizonSubscribers only 51 min ago8Add to bookmarks

DeepMind's ER 2 update pushes embodied reasoning toward video understanding and fleet-level coordination - the embodied-foundation race gets its Google-branded milestone.
In plain terms. Google DeepMind released Gemini Robotics ER 2, adding video understanding, task orchestration and multi-robot collaboration to its embodied stack. The step from single-arm demo to fleet-level reasoning is the real news, and it puts Alphabet visibly back in the embodied-foundation race.
The embodied-foundation thread has, so far, been driven by three actors: NVIDIA GR00T, China's Xiaomi Robotics-U0 (38B, open source) and X-Square Robot's backbone bet. Alphabet's ER line was the sleeper. ER 2 is DeepMind reasserting that it is running.
Per the announcement, ER 2 emphasises three concrete pieces:
The first two follow the field's public direction since 2025. The third is where differentiation lies. Coordinating two arms on one robot is a control problem; coordinating two robots on one job is a systems problem - shared world model, communication, conflict resolution.
The announcement does not publish head-to-head benchmarks against GR00T or RT-2 on the same tasks, and does not quantify "collaboration" (latency, coordination horizon, failure modes). Treat this as capability announced, not capability proven, until third parties reproduce.
For a decider: this is a strong signal that the embodied stack is following the LLM pattern - one backbone, adaptable across form factors. Budget accordingly for a market where the middleware layer collapses into a model call. For a builder: watch the API surface. If ER 2 is accessible outside Alphabet's internal robotics teams, the integration point for physical automation may shift within eighteen months. For the thread: the visible split is now between a Google closed-source push at the top and Xiaomi's 38B open-weight bet (Robotics-U0) at the base. Whether that divide widens - with more Chinese labs choosing the open-weight route the way Xiaomi did - is the story to watch through late 2026.
Create a free account to access all our content and the weekly review.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.
I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.
I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.
The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!
I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?
I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.
I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.
DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T