모델 & 도구 Jul 24, 2026 at 09:338북마크에 추가

BFL은 Seedance 2.0, Gemini Omni, Grok Imagine에 맞서 최첨단 상태로 도약했다고 주장하며, 로봇공학을 위한 '비디오-액션' 파생 모델을 잇따라 출시했다.
독일의 Black Forest Labs(BFL)는 Stability의 원조 팀에서 출발한 연구실로, FLUX 3라는 새로운 멀티모달 플로우 모델 패밀리를 발표했습니다. 2026년 7월 24일 Latent Space에서 전달된 발표에 따르면, FLUX 3는 BFL이 제시한 벤치마크에서 Seedance 2.0(ByteDance), Gemini Omni(Google), Grok Imagine(xAI)보다 우수한 성능을 보입니다. 또한 발표에는 로봇공학을 겨냥한 ‘video-action’ 모델인 FLUX-mimic도 포함되었습니다.
두 가지 신호가 두드러집니다. 첫째, 유럽이 멀티모달 생성 분야에서 미국과 중국이 장악하던 격차를 좁히고 있습니다.其二, FLUX-mimic의 파생 모델은 플로우 아키텍처가 이미지/비디오에 국한되지 않고 로봇 동작으로 확장됨을 보여줍니다. 이는 embodied-foundation-race( X-Square, 샤오미 로봇틱스-U0, 엔비디아 GR00T) 참여자들이 예상한 움직임과 정확히 일치하며, 이는 perception, generation, motor control에 공통으로 적용되는 백본(universal backbone)으로의 전환을 의미합니다.
제품 팀의 경우, 제3자 벤치마크 결과가 나오기 전까지는 Runway, Pika, Veo와 같은 기존 스택을 대체할 수 없습니다. 반면 로봇공학 연구실의 경우 FLUX-mimic을 신중히 평가해야 합니다. 이는 diffusion, VLA, RT-2 사이에서 hésitation하던 산업계에서 드물게 플로우 모델을 액션에 openly positioning한 사례이기 때문입니다.
BFL이 승부수를 던졌습니다. 약속은 아름답지만 기술적 입증은 아직 남아 있습니다. 확실한 점은 멀티모달 플로우는 더 이상 이미지의 니치가 아니라 perception, generation, action을 아우르는 범용 백본으로 credible한 후보가 되고 있다는 것입니다.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
I wonder if FLUX 3's multimodal approach could lead to overfitting in specific robotic tasks. How does it handle edge cases?
I'd like to see a direct comparison of FLUX 3's performance against Gemini Omni in real-world robotic applications, not just benchmarks.
I'm curious about the real-time processing capabilities of the video-action derivative. Can it handle latency issues in robotic applications?
I'm intrigued by the video-action derivative for robotics. Hope it can handle real-time decision making.
Excited to see how FLUX 3's multimodal capabilities translate into robotics. Wondering if it can outperform Gemini Omni in complex tasks.
I wonder how FLUX 3's video-action derivative will handle complex robotic tasks. Will it be more efficient than existing solutions?
I'm curious to see how FLUX 3 compares to Seedance 2.0 in real-world applications. Any benchmarks yet?
While the multimodal capabilities are impressive, I wonder about the environmental impact of developing and deploying such advanced AI models.