
DeepMind 的 ER 2 更新将具身推理推向视频理解和车队级协调 - 具身基础竞赛迎来谷歌品牌里程碑
用简单的语言来说。 Google DeepMind 发布了 Gemini Robotics ER 2,将视频理解、任务编排和多机器人协作添加到其具身化堆栈中。从单臂演示到机群级推理的步骤才是真正的新闻,这让 Alphabet 在具身化基础竞赛中明显回归。
迄今为止,具身化基础线程主要由三个参与者推动:NVIDIA GR00T、中国的 Xiaomi Robotics-U0(38B,开源)和 X-Square Robot 的主干投注。Alphabet 的 ER 系列一直是潜在竞争者。ER 2 是 DeepMind 重申其仍在竞争中的表现。
根据公告,ER 2 强调了三个具体方面:
前两项遵循了自2025年以来行业公开的方向。第三项是差异化所在。协调一个机器人上的两个臂是一个控制问题;协调两个机器人完成一项工作是一个系统问题——共享世界模型、通信、冲突解决。
公告没有发布与 GR00T 或 RT-2 在相同任务上的对比基准测试,也没有量化“协作”(延迟、协调范围、故障模式)。在第三方复现之前,将此视为已宣布的能力,而非已证明的能力。
对于决策者:这是一个强烈的信号,表明具身化堆栈正在遵循 LLM 模式——一个主干,可适应各种形式因素。相应地为一个市场预算,其中中间件层将折叠为模型调用。对于建设者:观察 API 表面。如果 ER 2 可供 Alphabet 内部机器人团队以外的用户访问,物理自动化的集成点可能在十八个月内发生变化。对于线程:现在明显的分歧是 Google 在顶部推动闭源,而 Xiaomi 的 38B 开源权重(Robotics-U0)在基础层。到2026年底,随着更多中国实验室选择开源权重路线,就像 Xiaomi 所做的那样,这个分歧是否会扩大,是需要关注的故事。
本文由人工智能撰写,并经人工编辑审核。
I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.
I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.
I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.
The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!
I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?
I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.
I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.
DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T