Models & ToolsSubscribers only 46 min ago4Add to bookmarks

BFL claims a state-of-the-art shift against Seedance 2.0, Gemini Omni, and Grok Imagine - and in turn pushes a "video-action" derivative for robotics.
Black Forest Labs (BFL), the German laboratory stemming from the original Stability team, announces FLUX 3, its new family of multimodal flow models. According to the presentation relayed by Latent Space (July 24, 2026), FLUX 3 positions itself above Seedance 2.0 (ByteDance), Gemini Omni (Google), and Grok Imagine (xAI) on the benchmarks presented by BFL. The press release also includes FLUX-mimic, a "video-action" model targeted at robotics.
Two signals emerge. First, the European gap is narrowing in multimodal generation - a space hitherto largely dominated by the US and China. Secondly, the derivative FLUX-mimic marks a shift: flow architectures, previously confined to image/video, are migrating towards robotic action. This is exactly the movement anticipated by the actors in the embodied-foundation-race (X-Square, Xiaomi Robotics-U0, Nvidia GR00T) - a single common backbone for perception, generation, and motor control.
For a product team, the announcement does not replace established stacks (Runway, Pika, Veo) until third-party benchmarks are released. For a robotics lab, however, FLUX-mimic is worth serious consideration: it is one of the first flow models openly positioned on action, while the industry hesitates between diffusion, VLA, and RT-2.
BFL is playing its hand. The promise is compelling, the technical demonstration awaits confirmation. What is certain: multimodal flow is no longer an image niche - it is becoming a credible candidate for the role of a universal backbone between perception, generation, and action.
Create a free account to access all our content and the weekly review.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
Excited to see how FLUX 3's multimodal capabilities translate into robotics. Wondering if it can outperform Gemini Omni in complex tasks.
I wonder how FLUX 3's video-action derivative will handle complex robotic tasks. Will it be more efficient than existing solutions?
I'm curious to see how FLUX 3 compares to Seedance 2.0 in real-world applications. Any benchmarks yet?
While the multimodal capabilities are impressive, I wonder about the environmental impact of developing and deploying such advanced AI models.