호라이즌 Jul 30, 2026 at 19:4012북마크에 추가

DeepMind의 ER 2 업데이트는 embodied reasoning을 비디오 이해와 함대 수준 조정으로 나아가게 하며, embodied-foundation 경쟁에서 구글 브랜드의 이정표를 마련합니다.
간단히 말해
Google DeepMind가 Gemini Robotics ER 2를 발표하며, embodied 스택에 비디오 이해, 작업 오케스트레이션 및 다중 로봇 협업을 추가했습니다. 단일 팔 데모에서 함대 수준의 추론으로의 전환이 핵심 뉴스이며, 이는 알파벳이 embodied 기반 경쟁에서 다시 한 번 존재감을 드러낸 것입니다.
현재 embodied 기반은 세 주체에 의해 주도되고 있습니다: NVIDIA GR00T, 중국의 샤오미 Robotics-U0(38B, 오픈소스), 그리고 X-Square Robot의 백본 베팅. 알파벳의 ER 라인은 숨은 강자였습니다. ER 2는 DeepMind가 여전히 경쟁력을 보유하고 있음을 재확인하는 것입니다.
공지에 따르면 ER 2는 세 가지 핵심 요소를 강조합니다:
첫 두 가지는 2025년 이후 공개된 분야의 방향을 따르고 있습니다. 세 번째는 차별화 포인트입니다. 한 로봇의 두 팔을 조정하는 것은 제어 문제지만, 한 작업에서 두 로봇을 조정하는 것은 시스템 문제입니다(공유 세계 모델, 통신, 충돌 해결).
공지에는 GR00T나 RT-2와의 동일한 작업에 대한 직접 비교 벤치마크가 없으며, "협업"의 정량적 측정(지연 시간, 조정 범위, 장애 모드)도 없습니다. 제3자가 재현할 때까지는 가능성 발표로만 간주하십시오.
결정권자에게: embodied 스택이 LLM 패턴을 따르고 있음을 보여주는 강력한 신호입니다. 하나의 백본이 다양한 폼 팩터에 적용 가능하다는 점입니다. 미들웨어 계층이 모델 호출로 통합되는 시장으로 예산을 조정하십시오.
개발자에게: API 표면을 주시하십시오. ER 2가 알파벳 내부 로봇틱스 팀 외부에 공개된다면, 물리적 자동화의 통합 지점이 18개월 이내에 바뀔 수 있습니다.
이 변화의 흐름: 이제可见한 분기는 구글의 폐쇄형 상위Push와 샤오미의 38B 오픈웨이트 베팅(Robotics-U0)으로 나뉩니다. 중국 лабора리들이 샤오미처럼 오픈웨이트 경로를 선택하면서 이 분기가 2026년 말까지 더 벌어질지 여부가 주목할 이야기입니다.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
The shift to video-native systems feels inevitable, but I’m skeptical about the scalability. How will it handle low-light conditions or high-contrast scenes without becoming computationally prohibitive?
Wouldn’t a video-native approach limit the robots to environments where cameras are reliable? What happens when lighting fails or sensors get obscured?
I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.
It's a tough nut to crack, but DeepMind's track record in AI gives me hope they'll nail it.
I'm curious about the ethical implications of fleet-level robot coordination. How will we ensure transparency and accountability in their decision-making processes?
Good question, transparency could be achieved through open-source algorithms and regular audits by independent bodies.
I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.
I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.
I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.
The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!
I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?
I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.
I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.
DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T