Gemini Robotics ER 2: DeepMindがビデオネイティブなマルチロボットオーケストレーションに賭ける

継続中のトピック : Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T· パート 14/16

ホライゾン Jul 30, 2026 at 19:4012ブックマークに追加

Gemini Robotics ER 2: DeepMindがビデオネイティブなマルチロボットオーケストレーションに賭ける
イラスト : Léa Fontaine

DeepMindのER 2アップデートが具現化された推論を動画理解と艦隊レベルの調整に押し進める — 具現化された基盤モデルレースにGoogleブランドのマイルストーンが加わった。

簡単に言えば。Google DeepMindは、Gemini Robotics ER 2をリリースし、ビデオ理解、タスクオーケストレーション、マルチロボットコラボレーションを組み込みスタックに追加しました。シングルアームデモからフリートレベルの推論へのステップが本当のニュースであり、Alphabetが明確に「組み込み基盤レース」に再参入したことを示しています。

位置付け

組み込み基盤の流れはこれまで、NVIDIA GR00T、中国Xiaomi Robotics-U0(38B、オープンソース)、X-Square Robotのバックボーンベットという3つのアクターによって牽引されてきました。AlphabetのERラインは「寝た子を起こす」存在でした。ER 2はDeepMindがまだ走り続けていることを再主張しています。

中身

発表によると、ER 2は3つの具体的な要素を強調しています:

  • ビデオ理解 - 単一フレームではなく、シーケンスに対する推論。これが長期操作に実際に必要なものです。
  • タスクオーケストレーション - 自然言語のゴールを、コントローラーが実行できるサブゴールに分解すること。
  • マルチロボットコラボレーション - 複数のプラットフォームにわたる協調動作。

最初の2つは2025年以降、分野の公開方向性に沿ったものです。3つ目が差別化ポイントです。1台のロボット上の2本のアームを調整するのは制御問題ですが、1つの作業に2台のロボットを調整するのはシステム問題です(共有ワールドモデル、通信、競合解決)。

まだ不明な点

発表では、GR00TやRT-2と同じタスクに対するヘッドトゥヘッドベンチマークが公開されておらず、「コラボレーション」の定量化(レイテンシ、調整ホライゾン、障害モード)も行われていません。第三者による再現が行われるまでは、実証された能力ではなく、発表された能力として扱ってください。

結論

意思決定者へ:これは組み込みスタックがLLMパターンに従っている強力なシグナルです。1つのバックボーンで、フォームファクター間で適応可能。ミドルウェア層がモデルコールに収束する市場に予算を組みましょう。開発者へ:APIサーフェスに注目してください。ER 2がAlphabetの内部ロボティクスチーム外からアクセス可能であれば、物理オートメーションの統合ポイントは18か月以内にシフトする可能性があります。この流れにとって:現在はGoogleのクローズドソース推進(上位)とXiaomiの38Bオープンウェイトベット(Robotics-U0、下位)という明確な分裂が見られます。その分裂が拡大し、中国のラボがXiaomiのようにオープンウェイト路線を選択するかどうかが、2026年後半までの注目ストーリーです。

リソース

本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。

編集部について
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
この記事は役に立ちましたか?

13 人がこの記事を評価しました

いいね
J
Jin-ho ParkFrontier & research
🇬🇧 Research, deep tech, foresight.
シェア:
コメント (12)

ログインして議論に参加しましょう。

FilmBuffNYC 01 Aug 2026 · 05:22

The shift to video-native systems feels inevitable, but I’m skeptical about the scalability. How will it handle low-light conditions or high-contrast scenes without becoming computationally prohibitive?

LitLover42 01 Aug 2026 · 05:00

Wouldn’t a video-native approach limit the robots to environments where cameras are reliable? What happens when lighting fails or sensors get obscured?

BookWorm88 31 Jul 2026 · 07:34

I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.

Alex_LDN 31 Jul 2026 · 10:44

It's a tough nut to crack, but DeepMind's track record in AI gives me hope they'll nail it.

LecteurDuDimanche 31 Jul 2026 · 07:11

I'm curious about the ethical implications of fleet-level robot coordination. How will we ensure transparency and accountability in their decision-making processes?

ArtLover99 31 Jul 2026 · 10:16

Good question, transparency could be achieved through open-source algorithms and regular audits by independent bodies.

FoodieFiona 30 Jul 2026 · 17:19

I'm excited about the potential for video-native understanding, but how will it handle occlusions and dynamic environments? Real-world complexity is a beast.

BookWorm47 30 Jul 2026 · 16:42

I hope this tech can bridge the gap between lab demos and real-world applications. The potential is huge, but the challenges are real.

MusicFanatic 30 Jul 2026 · 16:37

I wonder how this tech will handle the coordination of robots with different capabilities and limitations. It's a complex challenge.

ArtLoverLA 30 Jul 2026 · 16:27

The video-native approach is intriguing, but I wonder how it will handle low-light or high-speed scenarios. Real-world applications are messy!

Alex_LDN 30 Jul 2026 · 16:23

I'm curious about the scalability of this system. How will it handle diverse environments and tasks beyond lab settings?

TravelTom 30 Jul 2026 · 16:11

I wonder how this tech will handle dynamic environments with unpredictable human interactions. Real-world applications are complex.

curio_usa 30 Jul 2026 · 15:56

I'm excited about the multi-robot coordination aspect. Wondering how they'll handle communication delays in large-scale deployments.

Dr. L. 30 Jul 2026 · 15:54

DeepMind's ER 2 update is a significant step forward in embodied reasoning. I'm curious to see how this will impact real-world robotics applications.

トピックの経過

Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T

  1. 1XiaomiはRobotics-U0を発表:オープンソースの統合 Embodied AI モデル38B15/07/2026
  2. 2Xpengは2027年にグローバルなヒューマノイドロボットを発売すると発表:中国の「具現化」レースに発売日が決定16/07/2026
  3. 3LimX DynamicsがFigure等級のヒューマノイドデモを発表:中国の具現化分野に新たな参入者16/07/2026
  4. 4WeRide WITT:自動運転が「基盤モデル」フォーマットを採用17/07/2026
  5. 5Meituan、中国のロボット企業に7390万ドルを出資17/07/2026
  6. 6シャオミの工場内ロボット:EV工場で98%のタスク成功率を達成、静かに18/07/2026
  7. 7Xiaomiの工場内ロボットがナット組み立てで98%、新たな2つのタスクで90%以上を達成19/07/2026
  8. 8Yimu Techはロボット用タッチセンサーに10億元以上を調達20/07/2026
  9. 9Unitreeは、同社のロボットの「GPTの瞬間」はまだ数年先だと主張しています - インサイダーによるアンチハイプの発言22/07/2026
  10. 10Huawei、CloudRoboで具現化AIに参入 - SERESのロボット工学プレイブック22/07/2026
  11. 11米下院、中国製ヒューマノイドロボットの規制を可決:デュアルユースの軍事利用24/07/2026
  12. 12ロペディアが2200万ドルを調達:ロボットが世界を理解するために必要な「データレイヤー」を提供24/07/2026
  13. 13インドの地上戦:バンガロールの倉庫と建設ロボットがひっそりと稼働を開始26/07/2026
  14. 14Gemini Robotics ER 2: DeepMindがビデオネイティブなマルチロボットオーケストレーションに賭ける30/07/2026
  15. 15LimX Dynamics : Shen Hua、PMFで固定 - 「次の停止地点はランディング」31/07/2026
  16. 16CATL、ロボパーティーに賭ける:中国の巨大バッテリーメーカー、側面からヒューマノイドレースに参入03/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
テーマ
探索
インフォメーション