Gemini Robotics ER 2: DeepMind, 비디오 네이티브 및 다중 로봇 오케스트레이션에 베팅

진행 중인 이슈 : Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T· 편 14/16

호라이즌 Jul 30, 2026 at 19:4012북마크에 추가

Gemini Robotics ER 2: DeepMind, 비디오 네이티브 및 다중 로봇 오케스트레이션에 베팅
삽화 : Léa Fontaine

DeepMind의 ER 2 업데이트는 embodied reasoning을 비디오 이해와 함대 수준 조정으로 나아가게 하며, embodied-foundation 경쟁에서 구글 브랜드의 이정표를 마련합니다.

간단히 말해

Google DeepMind가 Gemini Robotics ER 2를 발표하며, embodied 스택에 비디오 이해, 작업 오케스트레이션 및 다중 로봇 협업을 추가했습니다. 단일 팔 데모에서 함대 수준의 추론으로의 전환이 핵심 뉴스이며, 이는 알파벳이 embodied 기반 경쟁에서 다시 한 번 존재감을 드러낸 것입니다.

이 변화의 위치

현재 embodied 기반은 세 주체에 의해 주도되고 있습니다: NVIDIA GR00T, 중국의 샤오미 Robotics-U0(38B, 오픈소스), 그리고 X-Square Robot의 백본 베팅. 알파벳의 ER 라인은 숨은 강자였습니다. ER 2는 DeepMind가 여전히 경쟁력을 보유하고 있음을 재확인하는 것입니다.

내부 구조

공지에 따르면 ER 2는 세 가지 핵심 요소를 강조합니다:

  • 비디오 이해 - 단일 프레임이 아닌 시퀀스 기반 추론. 이는 장기 조작에 실제로 필요한 부분입니다.
  • 작업 오케스트레이션 - 자연어 목표를 컨트롤러가 실행할 수 있는 하위 목표로 분해하는 기능.
  • 다중 로봇 협업 - 둘 이상의 플랫폼 간 조정된 동작.

첫 두 가지는 2025년 이후 공개된 분야의 방향을 따르고 있습니다. 세 번째는 차별화 포인트입니다. 한 로봇의 두 팔을 조정하는 것은 제어 문제지만, 한 작업에서 두 로봇을 조정하는 것은 시스템 문제입니다(공유 세계 모델, 통신, 충돌 해결).

아직 알려지지 않은 점

공지에는 GR00T나 RT-2와의 동일한 작업에 대한 직접 비교 벤치마크가 없으며, "협업"의 정량적 측정(지연 시간, 조정 범위, 장애 모드)도 없습니다. 제3자가 재현할 때까지는 가능성 발표로만 간주하십시오.

결론

결정권자에게: embodied 스택이 LLM 패턴을 따르고 있음을 보여주는 강력한 신호입니다. 하나의 백본이 다양한 폼 팩터에 적용 가능하다는 점입니다. 미들웨어 계층이 모델 호출로 통합되는 시장으로 예산을 조정하십시오.

개발자에게: API 표면을 주시하십시오. ER 2가 알파벳 내부 로봇틱스 팀 외부에 공개된다면, 물리적 자동화의 통합 지점이 18개월 이내에 바뀔 수 있습니다.

이 변화의 흐름: 이제可见한 분기는 구글의 폐쇄형 상위Push와 샤오미의 38B 오픈웨이트 베팅(Robotics-U0)으로 나뉩니다. 중국 лабора리들이 샤오미처럼 오픈웨이트 경로를 선택하면서 이 분기가 2026년 말까지 더 벌어질지 여부가 주목할 이야기입니다.

Resources

인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.

편집팀
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
이 기사가 도움이 되었나요?

13 명이 이 기사를 좋아합니다

좋아요
J
Jin-ho ParkFrontier & research
🇬🇧 Research, deep tech, foresight.
공유:
댓글 (12)

토론에 참여하려면 로그인하세요.

FilmBuffNYC 01 Aug 2026 · 05:22

The shift to video-native systems feels inevitable, but I’m skeptical about the scalability. How will it handle low-light conditions or high-contrast scenes without becoming computationally prohibitive?

LitLover42 01 Aug 2026 · 05:00

Wouldn’t a video-native approach limit the robots to environments where cameras are reliable? What happens when lighting fails or sensors get obscured?

BookWorm88 31 Jul 2026 · 07:34

I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.

Alex_LDN 31 Jul 2026 · 10:44

It's a tough nut to crack, but DeepMind's track record in AI gives me hope they'll nail it.

LecteurDuDimanche 31 Jul 2026 · 07:11

I'm curious about the ethical implications of fleet-level robot coordination. How will we ensure transparency and accountability in their decision-making processes?

ArtLover99 31 Jul 2026 · 10:16

Good question, transparency could be achieved through open-source algorithms and regular audits by independent bodies.

FoodieFiona 30 Jul 2026 · 17:19

I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.

BookWorm47 30 Jul 2026 · 16:42

I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.

MusicFanatic 30 Jul 2026 · 16:37

I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.

ArtLoverLA 30 Jul 2026 · 16:27

The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!

Alex_LDN 30 Jul 2026 · 16:23

I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?

TravelTom 30 Jul 2026 · 16:11

I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.

curio_usa 30 Jul 2026 · 15:56

I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.

Dr. L. 30 Jul 2026 · 15:54

DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.

이슈 타임라인

Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T

  1. 1샤오미, 로보틱스-U0 공개: 오픈 소스 embodied 모델 38B 출시15/07/2026
  2. 2Xpeng, 2027년 글로벌 휴머노이드 출시 약속: 중국 '구현된' 레이스가 배포 일정을 발표하다16/07/2026
  3. 3LimX Dynamics가 Figure급 휴머노이드 데모 영상을 공개했습니다: 중국의 embodied field에 새로운 도전자 등장16/07/2026
  4. 4WeRide WITT : 자율주행이 'foundation model' 형식 채택17/07/2026
  5. 5Meituan이 중국 로봇 기업에 7,390만 달러 규모 투자를 단행17/07/2026
  6. 6샤오미의 공장 인턴 로봇: 전기차 공장에서 98%의 작업 성공률 기록, 조용히18/07/2026
  7. 7샤오미의 공장 인턴 로봇, 너트 조립 98% 달성…새로운 두 작업에서도 90% 이상 성과19/07/2026
  8. 8Yimu Tech, 10억 위안 규모의 로봇 터치 센서 투자 유치20/07/2026
  9. 9Unitree는 로봇의 "GPT 순간"이 아직은 수년은 멀었다고 말합니다 - 내부자의 과장 비판22/07/2026
  10. 10허바이, 클라우드로보로 embodied AI 진출 - SERES의 로봇틱스 플레이북22/07/2026
  11. 11미국 하원, 중국산 휴머노이드 로봇 군사적 활용 규제안 가결24/07/2026
  12. 12로피디아, 2200만 달러 투자 유치: 로봇이 세상을 이해하는 데 필요한 ‘데이터 레이어’ 제공24/07/2026
  13. 13인도의 지상 작전: 벵갈루루의 창고 및 건설 로봇이 조용히 가동 중26/07/2026
  14. 14Gemini Robotics ER 2: DeepMind, 비디오 네이티브 및 다중 로봇 오케스트레이션에 베팅30/07/2026
  15. 15LimX Dynamics : Shen Hua가 PMF에 고정 - "다음 목적지는 랜딩"31/07/2026
  16. 16CATL, 배터리 거인이 로봇 경주에 옆으로 뛰어들다03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
토픽
탐색
정보