
FLUX 3: Black Forest Labs Trains One Model to Generate Video, Audio and Robot Actions
The German lab's new multimodal flow model produces 20-second clips with natively synchronized audio from a single set of weights — and extends the same architecture to robotic action prediction, already tested on Audi production lines.
Black Forest Labs has released FLUX 3, and the headline is not any single capability but the architecture underneath all of them: a single multimodal foundation model, trained jointly on images, video and audio, that generates 20-second video clips with natively synchronized sound — and predicts robot actions from the same set of weights.
The Freiburg-based lab, best known for the FLUX image models that power much of the open image-generation ecosystem, is making an explicit bet with this release. As the company puts it in its announcement, a model "must learn a representation of the world" — not spatial structure, motion, or sound in isolation, but all of them together, the way physical events actually produce them.
One Flow, Every Modality
FLUX 3 is built on flow matching, extended through what Black Forest Labs calls its Self-Flow approach for aligning multimodal generation and understanding. Where most video systems bolt an audio model onto a video model — generating frames first, then scoring them with sound — FLUX 3 produces video and audio in a single inference pass from one flow-matching framework. The practical result is audio that is causally tied to what happens on screen: the model learned to associate sounds with the physical events that make them, rather than layering plausible ambience over finished footage.
The video system handles text-to-video, image-to-video, video-to-video, keyframe transitions and multilingual dialogue, with clips running up to 20 seconds at a time. For longer work, the lab demonstrates agentic chaining — stitching generated shots into multi-shot sequences lasting several minutes — along with typography generation strong enough for titles and on-screen text, a longstanding FLUX differentiator carried over from the image line.
From Pixels to Actuators
The most consequential extension may be the least flashy: native action prediction. Black Forest Labs argues that a model which has learned how the world moves and sounds has also learned much of what a robot needs to act in it. Working with robotics partner mimic, the lab built FLUX-mimic for dexterous manipulation tasks — and reports the system has been tested in production at Audi, moving the claim from demo reel to factory floor.
That framing places FLUX 3 in the same lineage as world-model efforts from DeepMind and NVIDIA, but with an unusual path: arriving at embodied prediction from a generative media model rather than from simulation infrastructure.
Early Numbers
The lab's preliminary human-preference evaluations put FLUX 3 ahead of the current video-generation field: preferred 77% of the time against Runway's Gen-4.5, 93% against Luma's Ray 3.2, and 69% against xAI's Grok Imagine Video. Those are vendor-run comparisons and should be read accordingly — no independent benchmark suite yet exists for joint audio-video generation — but the margins are wide enough to signal a genuine step, and Black Forest Labs says results will improve as the early-access models are refined.
Availability is staged. FLUX 3 Video is in early access now; FLUX 3 Image opens in the coming weeks; the action-prediction stack is limited to selected partners. Critically for the open ecosystem — including the Asian developer communities that built tooling economies around FLUX.1's open weights — the company has committed to an open-weight FLUX 3 Dev release later this year.
Why It Matters
FLUX 3 lands in a week when the frontier is consolidating around unified world representations rather than per-modality specialists. If a mid-sized European lab can ship competitive video, native audio and production-tested robot actions from one training run, the economics of maintaining separate video, audio and robotics models start to look questionable — for everyone from ByteDance's Seedance team to the robotics startups building on task-specific vision-language-action models. The open FLUX 3 Dev release, when it arrives, will be the real test: it would hand the first jointly trained video-audio foundation model to the open-weight community, and history suggests that community moves fast.
Newsletter
Get Lanceum in your inbox
Weekly insights on AI and technology in Asia.
More in Research

Kimi K3's Weights Are Out: 2.8T Parameters, a 1.4TB Download, and an Escalation Nobody Can Undo

MeetingToM: A Benchmark for the Social Skill AI Keeps Failing — Telling Real Agreement From Fake
