モデルとツール Jul 24, 2026 at 09:338ブックマークに追加

BFLは最先端の技術でSeedance 2.0、Gemini Omni、Grok Imagineに対抗すると主張し、その後すぐにロボット工学向けの「ビデオアクション」派生版をリリースした。
ドイツの研究所Black Forest Labs(BFL)は、Stability元のチームから生まれた同研究所が、マルチモーダルなフロー型モデルの新しいファミリーであるFLUX 3を発表しました。BFLが発表したベンチマークによると、FLUX 3はSeedance 2.0(ByteDance)、Gemini Omni(Google)、Grok Imagine(xAI)を上回る性能を示しています。また、発表にはFLUX-mimicという、ロボット工学向けの「ビデオアクション」モデルも含まれています。
2つのシグナルが見られます。まず、マルチモーダル生成の分野で欧州の存在感が高まっており、これまで米中が支配的だったこの分野で欧州との差が縮まっています。次に、FLUX-mimicの派生モデルは転換点を示しています。これまで画像・動画生成に限定されていたフローアーキテクチャが、ロボット工学におけるアクションに移行しつつあります。これはまさに、embodied-foundation-race(X-Square、Xiaomi Robotics-U0、Nvidia GR00T)の関係者が予想していた動きであり、知覚、生成、モーター制御を統合する共通のバックボーンです。
製品チームにとって、FLUX 3の発表は、第三者ベンチマークの結果が出るまでは、既存のスタック(Runway、Pika、Veo)に取って代わるものではありません。一方、ロボット工学の研究室にとっては、FLUX-mimicを真剣に評価する価値があります。これは、拡散モデル、VLA、RT-2の間で業界が揺れる中、アクションに特化した最初のフローモデルの一つだからです。
BFLは賭けに出ました。約束は素晴らしいものですが、技術的な実証はまだ確認が必要です。確実なことは、マルチモーダルフローが画像生成のニッチな存在ではなくなり、知覚、生成、アクションをつなぐユニバーサルなバックボーンの有力候補となっていることです。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
I wonder if FLUX 3's multimodal approach could lead to overfitting in specific robotic tasks. How does it handle edge cases?
I'd like to see a direct comparison of FLUX 3's performance against Gemini Omni in real-world robotic applications, not just benchmarks.
I'm curious about the real-time processing capabilities of the video-action derivative. Can it handle latency issues in robotic applications?
I'm intrigued by the video-action derivative for robotics. Hope it can handle real-time decision making.
Excited to see how FLUX 3's multimodal capabilities translate into robotics. Wondering if it can outperform Gemini Omni in complex tasks.
I wonder how FLUX 3's video-action derivative will handle complex robotic tasks. Will it be more efficient than existing solutions?
I'm curious to see how FLUX 3 compares to Seedance 2.0 in real-world applications. Any benchmarks yet?
While the multimodal capabilities are impressive, I wonder about the environmental impact of developing and deploying such advanced AI models.