HorizonRéservé aux abonnés il y a 54 min8Ajouter aux favoris

DeepMind's ER 2 update pushes embodied reasoning toward video understanding and fleet-level coordination - the embodied-foundation race gets its Google-branded milestone.
In plain terms. Google DeepMind released Gemini Robotics ER 2, adding video understanding, task orchestration and multi-robot collaboration to its embodied stack. The step from single-arm demo to fleet-level reasoning is the real news, and it puts Alphabet visibly back in the embodied-foundation race.
The embodied-foundation thread has, so far, been driven by three actors: NVIDIA GR00T, China's Xiaomi Robotics-U0 (38B, open source) and X-Square Robot's backbone bet. Alphabet's ER line was the sleeper. ER 2 is DeepMind reasserting that it is running.
Per the announcement, ER 2 emphasises three concrete pieces:
The first two follow the field's public direction since 2025. The third is where differentiation lies. Coordinating two arms on one robot is a control problem; coordinating two robots on one job is a systems problem - shared world model, communication, conflict resolution.
The announcement does not publish head-to-head benchmarks against GR00T or RT-2 on the same tasks, and does not quantify "collaboration" (latency, coordination horizon, failure modes). Treat this as capability announced, not capability proven, until third parties reproduce.
For a decider: this is a strong signal that the embodied stack is following the LLM pattern - one backbone, adaptable across form factors. Budget accordingly for a market where the middleware layer collapses into a model call. For a builder: watch the API surface. If ER 2 is accessible outside Alphabet's internal robotics teams, the integration point for physical automation may shift within eighteen months. For the thread: the visible split is now between a Google closed-source push at the top and Xiaomi's 38B open-weight bet (Robotics-U0) at the base. Whether that divide widens - with more Chinese labs choosing the open-weight route the way Xiaomi did - is the story to watch through late 2026.
Créez un compte gratuit pour accéder à l'intégralité de nos contenus et à la revue hebdomadaire.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.
I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.
I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.
The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!
I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?
I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.
I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.
DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T