Black Forest Labs、FLUX 3を発表:ロボット工学にまで広がるマルチモーダルなフローモデル

モデルとツール Jul 24, 2026 at 09:338ブックマークに追加

Black Forest Labs、FLUX 3を発表:ロボット工学にまで広がるマルチモーダルなフローモデル
イラスト : Léa Fontaine

BFLは最先端の技術でSeedance 2.0、Gemini Omni、Grok Imagineに対抗すると主張し、その後すぐにロボット工学向けの「ビデオアクション」派生版をリリースした。

文脈

ドイツの研究所Black Forest Labs(BFL)は、Stability元のチームから生まれた同研究所が、マルチモーダルなフロー型モデルの新しいファミリーであるFLUX 3を発表しました。BFLが発表したベンチマークによると、FLUX 3はSeedance 2.0(ByteDance)、Gemini Omni(Google)、Grok Imagine(xAI)を上回る性能を示しています。また、発表にはFLUX-mimicという、ロボット工学向けの「ビデオアクション」モデルも含まれています。

データ

  • 発表されたファミリー:FLUX 3(マルチモーダルフロー)
  • BFLが挙げた競合:Seedance 2.0、Gemini Omni、Grok Imagine
  • 派生モデル:FLUX-mimic - ロボット工学向けビデオアクションモデル
  • ソース:Latent Space、2026年7月24日(ベンチマークは第三者による検証が必要)

分析

2つのシグナルが見られます。まず、マルチモーダル生成の分野で欧州の存在感が高まっており、これまで米中が支配的だったこの分野で欧州との差が縮まっています。次に、FLUX-mimicの派生モデルは転換点を示しています。これまで画像・動画生成に限定されていたフローアーキテクチャが、ロボット工学におけるアクションに移行しつつあります。これはまさに、embodied-foundation-race(X-Square、Xiaomi Robotics-U0、Nvidia GR00T)の関係者が予想していた動きであり、知覚、生成、モーター制御を統合する共通のバックボーンです。

シナリオ

  • 短期:第三者ベンチマーク(Artificial Analysis、LMArena Vision)により、4〜6週間以内に議論が決着する
  • 中期:FLUX 3の性能が維持されれば、BFLはMeta、Amazon、または戦略的欧州企業(SAP、Mistral)による買収のターゲットとなる
  • 代替案:FLUX-mimicはデモに留まる。ロボット工学ベンチマークは標準化されておらず、比較が困難

専門家への示唆

製品チームにとって、FLUX 3の発表は、第三者ベンチマークの結果が出るまでは、既存のスタック(Runway、Pika、Veo)に取って代わるものではありません。一方、ロボット工学の研究室にとっては、FLUX-mimicを真剣に評価する価値があります。これは、拡散モデル、VLA、RT-2の間で業界が揺れる中、アクションに特化した最初のフローモデルの一つだからです。

監視すべきシグナル

  • 技術レポート(論文または詳細なブログ投稿)の公開
  • APIアクセス vs オープンウェイトの公開(BFLはこれまでミックスした歴史あり)
  • Seedance(ByteDance)とGemini Omni(Google)の反応 - 争点となっているベンチマークに関するもの

我々の見解

BFLは賭けに出ました。約束は素晴らしいものですが、技術的な実証はまだ確認が必要です。確実なことは、マルチモーダルフローが画像生成のニッチな存在ではなくなり、知覚、生成、アクションをつなぐユニバーサルなバックボーンの有力候補となっていることです。

リソース

本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。

編集部について
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
この記事は役に立ちましたか?

9 人がこの記事を評価しました

いいね
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
シェア:
コメント (8)

ログインして議論に参加しましょう。

J.P.R. 2 24 Jul 2026 · 17:24

I wonder if FLUX 3's multimodal approach could lead to overfitting in specific robotic tasks. How does it handle edge cases?

J.P.R. 3 24 Jul 2026 · 17:20

I'd like to see a direct comparison of FLUX 3's performance against Gemini Omni in real-world robotic applications, not just benchmarks.

Dr. J. 24 Jul 2026 · 16:46

I'm curious about the real-time processing capabilities of the video-action derivative. Can it handle latency issues in robotic applications?

FoodieFiona 2 24 Jul 2026 · 08:21

I'm intrigued by the video-action derivative for robotics. Hope it can handle real-time decision making.

Alex_LDN 24 Jul 2026 · 07:32

Excited to see how FLUX 3's multimodal capabilities translate into robotics. Wondering if it can outperform Gemini Omni in complex tasks.

unLecteurCurieux 24 Jul 2026 · 07:27

I wonder how FLUX 3's video-action derivative will handle complex robotic tasks. Will it be more efficient than existing solutions?

curio_usa 24 Jul 2026 · 07:04

I'm curious to see how FLUX 3 compares to Seedance 2.0 in real-world applications. Any benchmarks yet?

EcoWarrior99 24 Jul 2026 · 06:58

While the multimodal capabilities are impressive, I wonder about the environmental impact of developing and deploying such advanced AI models.

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
テーマ
探索
インフォメーション