The lab that gave the open-source world FLUX.1 and FLUX.2 just stopped making image models. On July 23, 2026, Black Forest Labs dropped FLUX 3 — one model that generates images, video, sound, and even robot movements from a single shared brain. This is the ComfyUI backbone lab betting the whole farm on “real world models.”
The Story
Black Forest Labs (BFL) built its name on FLUX — the open-weight image model that quietly became the default node in half the ComfyUI graphs on the planet. So when they call FLUX 3 a “multimodal foundation model,” the natural reaction is: another image release. It is not.
FLUX 3 learns from images, video, and audio inside one unified architecture. The same weights that draw a picture also produce a 20-second video with sound, and can be extended to predict actions for a robot arm. BFL calls the recipe Self-Flow — their method for training generation and understanding together, so the model that makes the world also learns to read it.
The headline trick is the one everyone else fakes. FLUX 3 makes video and its audio in the same pass. Every big video model today — Kling, Seedance, Runway, Luma — bolts a separate audio model on afterward. FLUX 3 pulls picture and sound from the same flow, so footsteps land on the right frame and a violin bow actually matches the note. It handles text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue, and chaining clips into multi-shot sequences up to 20 seconds each.
Then there is the part nobody saw coming: robots. BFL partnered with robot-learning startup mimic to build FLUX-mimic, a version that predicts dexterous manipulation actions — and they tested it on real production lines at Audi. The pitch is that a model which truly understands how the world moves is the same model that can act in it. Whether you buy that or not, it is a wild place for an “image lab” to land in two years.
Why You Should Care
For creative technologists, the interesting claim is not the robots — it is Self-Flow. BFL’s charts argue that unifying generation and understanding lifts both at once: lower error on video, image, and audio, plus a robot-control success rate that climbs to 47% versus 35% for plain flow matching, and learns twice as fast.
If that holds, it points at where AI creative tools are going. Not a stack of narrow models — an image gen here, a video gen there, a lip-sync tool taped on — but one backbone you prompt across formats. For a 3D and motion pipeline, a model that keeps sound, motion, and physical logic consistent is worth far more than one that just makes a prettier still.
And here is the line that matters most to this crowd: BFL says an open-weight FLUX 3 Dev is planned — the same multimodal backbone, released to the community, just like FLUX.1 Dev before it. If it ships, the entire ComfyUI ecosystem gets a native audio-video-image node instead of five stitched-together ones.
Try It / Follow Them
Temper the excitement: right now FLUX 3 is gated. FLUX 3 Video is in early access, FLUX 3 Image rolls out over the coming weeks, and FLUX 3 Action (the robotics variant) is in selected testing. There is no public API, no pricing, no disclosed parameter count, and — crucially — no open weights yet. Every benchmark above is BFL’s own; nobody outside has stress-tested the 20-second claim.
- Read the full announcement: bfl.ai/blog/flux-3
- Request early access and follow drops: @bfl_ai on X
- Watch for the FLUX 3 Dev open weights — that is the release that changes ComfyUI workflows.
IK3D Lab Take
This is a statement of intent, not a product you can build with today. But it is the right statement. The lab that made open image generation mainstream is now saying the future is one model for image, video, sound, and action — and it is putting that on a factory floor to prove it. The 20-second native-audio video is the flashy demo; Self-Flow is the real bet, and if BFL keeps its habit of shipping open weights, FLUX 3 Dev could be the single most important node the community gets all year. We will be first in line to test it — and first to check whether the charts survive contact with reality.



