Black Forest Labs Unveils FLUX 3 Multimodal AI Foundation Model

Black Forest Labs (BFL) launched FLUX 3 on July 26, 2026. It’s a multimodal AI model that handles images, video, audio, and robotic action prediction—all within one architecture.
FLUX 3 stands out because it ships video, audio, and action prediction from a single set of weights. This is the first time the FLUX series has done that. The model builds on Self-Flow, BFL’s method introduced in March 2026, which aligns multimodal generation and understanding inside one system.
Training all modalities together forces them to constrain and improve each other. Robin Rombach, BFL’s CEO, put it plainly: “True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results.” He added, “You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.”
FLUX 3 Video can generate clips up to 20 seconds long with native audio. It supports multiple video generation modes: text-to-video, image-to-video, video-to-video from a reference clip, keyframe-to-video, and generative video-audio continuation. Early tests favored FLUX 3 Video over competitors like Luma Ray 3.2 in 93% of comparisons and also showed it beating Runway Gen-4.5, Grok Imagine Video, and others by smaller margins.
Training focuses heavily on video prediction, which consumes over 95% of compute. Audio takes up less than 0.5% of training tokens. This skew reflects the model’s emphasis on video quality and length.
Robotics enters the picture with FLUX-mimic, a model built on FLUX 3 for robotic manipulation. Elvis Nava, CTO of mimic robotics, explained that FLUX-mimic can learn new tasks in minutes, not days, thanks to its understanding of physical world dynamics from video data.
Christoph Schneider from Audi Production Lab praised the partnership: “For us, partnering with pioneering companies such as mimic and Black Forest Labs is essential in pushing the frontier of Physical AI and validating these innovations in real-world production environments.”
FLUX 3 will be available in several forms: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and FLUX 3 Dev. Black Forest Labs plans faster and open-weight versions later in 2026.
In the wider AI landscape, Moonshot AI’s Kimi K3 with 2.8 trillion parameters and Alibaba’s Qwen3.8 with 2.4 trillion parameters loom large. Moonshot claims Kimi K3 ranks above nearly every US system except OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5. Moonshot will release full Kimi K3 weights on July 27, 2026. Alibaba’s Qwen3.8 is described as “continuously evolving” and will soon go open-weight as well.
For now, FLUX 3’s multimodal approach—combining video, audio, and action prediction in one model—is a rare and ambitious step. Whether it reshapes AI for creative fields, physical AI, and robotics remains to be seen. But BFL’s insistence on training everything jointly is a bold challenge to the old image-only or text-only silos.
Based on
- Black Forest Labs Releases FLUX 3: A Multimodal Flow Model for Image, Video, Audio and Robot Action Prediction — marktechpost.com
- Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start | VentureBeat — venturebeat.com
- Black Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence | Markets Insider — markets.businessinsider.com
- China delivers a one-two punch to America’s AI dominance | The Verge — theverge.com
- What is China’s Kimi K3 and why is the US so rattled by it? | CNN Business — cnn.com




