Asia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M UsersAsia Chip Stocks Snap Back: KOSPI Jumps 2.8%, Nikkei Reclaims 63,000CXMT Soars 466% on Shanghai Debut, Becoming China's Most Valuable Listed CompanyCyera to Buy Oasis Security for $1 Billion as AI Agent Protection Becomes Cybersecurity's Hottest MarketThe Security Split: How One Breach Redrew the AI Industry's Battle LinesChina's Embodied AI Machine: Record Robotics Funding Meets an IPO Assembly LineA 557% Profit Surge That Counts as a Miss: What SK Hynix's Quarter Reveals About the AI TradeCORE Biomedicine Raises $21M Across Boston, Tokyo and Suzhou for AI-Guided Precision OncologyFish Audio Reels In $52M Seed a Year After Starting in a Bedroom — With $21M ARR and 8M Users
Black Forest Labs FLUX 3 multimodal foundation model announcement graphic
Black Forest Labs
Research

FLUX 3: Black Forest Labs Trains One Model to Generate Video, Audio and Robot Actions

The German lab's new multimodal flow model produces 20-second clips with natively synchronized audio from a single set of weights — and extends the same architecture to robotic action prediction, already tested on Audi production lines.

D
Daniel ParkAI Correspondent
5 min read

Black Forest Labs has released FLUX 3, and the headline is not any single capability but the architecture underneath all of them: a single multimodal foundation model, trained jointly on images, video and audio, that generates 20-second video clips with natively synchronized sound — and predicts robot actions from the same set of weights.

The Freiburg-based lab, best known for the FLUX image models that power much of the open image-generation ecosystem, is making an explicit bet with this release. As the company puts it in its announcement, a model "must learn a representation of the world" — not spatial structure, motion, or sound in isolation, but all of them together, the way physical events actually produce them.

One Flow, Every Modality

FLUX 3 is built on flow matching, extended through what Black Forest Labs calls its Self-Flow approach for aligning multimodal generation and understanding. Where most video systems bolt an audio model onto a video model — generating frames first, then scoring them with sound — FLUX 3 produces video and audio in a single inference pass from one flow-matching framework. The practical result is audio that is causally tied to what happens on screen: the model learned to associate sounds with the physical events that make them, rather than layering plausible ambience over finished footage.

The video system handles text-to-video, image-to-video, video-to-video, keyframe transitions and multilingual dialogue, with clips running up to 20 seconds at a time. For longer work, the lab demonstrates agentic chaining — stitching generated shots into multi-shot sequences lasting several minutes — along with typography generation strong enough for titles and on-screen text, a longstanding FLUX differentiator carried over from the image line.

From Pixels to Actuators

The most consequential extension may be the least flashy: native action prediction. Black Forest Labs argues that a model which has learned how the world moves and sounds has also learned much of what a robot needs to act in it. Working with robotics partner mimic, the lab built FLUX-mimic for dexterous manipulation tasks — and reports the system has been tested in production at Audi, moving the claim from demo reel to factory floor.

That framing places FLUX 3 in the same lineage as world-model efforts from DeepMind and NVIDIA, but with an unusual path: arriving at embodied prediction from a generative media model rather than from simulation infrastructure.

Early Numbers

The lab's preliminary human-preference evaluations put FLUX 3 ahead of the current video-generation field: preferred 77% of the time against Runway's Gen-4.5, 93% against Luma's Ray 3.2, and 69% against xAI's Grok Imagine Video. Those are vendor-run comparisons and should be read accordingly — no independent benchmark suite yet exists for joint audio-video generation — but the margins are wide enough to signal a genuine step, and Black Forest Labs says results will improve as the early-access models are refined.

Availability is staged. FLUX 3 Video is in early access now; FLUX 3 Image opens in the coming weeks; the action-prediction stack is limited to selected partners. Critically for the open ecosystem — including the Asian developer communities that built tooling economies around FLUX.1's open weights — the company has committed to an open-weight FLUX 3 Dev release later this year.

Why It Matters

FLUX 3 lands in a week when the frontier is consolidating around unified world representations rather than per-modality specialists. If a mid-sized European lab can ship competitive video, native audio and production-tested robot actions from one training run, the economics of maintaining separate video, audio and robotics models start to look questionable — for everyone from ByteDance's Seedance team to the robotics startups building on task-specific vision-language-action models. The open FLUX 3 Dev release, when it arrives, will be the real test: it would hand the first jointly trained video-audio foundation model to the open-weight community, and history suggests that community moves fast.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused