Anthropic Walks Away From $6 Billion Decart Acquisition After Due Diligence•OpenAI Becomes Anchor Customer of Nvidia-Backed Firmus for Malaysian AI Factories•Japan's AI Data Center Capacity Set to Quadruple to 4.9 GW by 2033 With $60 Billion Buildout•All Three Majors Are Now Suing Anthropic. The AI Copyright Endgame Is Taking Shape•The Backlash Has Arrived: Asia's Data Center Boom Is Colliding With the People Who Live Next Door•Groq, Nemotron, Now Hugging Face: Nvidia Isn't Buying Companies — It's Buying the Open AI Ecosystem•Travis Kalanick's Atoms Eyes Robotaxis — With Uber's $100 Million and His Old AV Team•Fluidstack Closes $1.5 Billion Led by Jane Street — Revenue Up From $1.8M to a Projected $660M, While Owning Zero Chips•Anthropic Walks Away From $6 Billion Decart Acquisition After Due Diligence•OpenAI Becomes Anchor Customer of Nvidia-Backed Firmus for Malaysian AI Factories•Japan's AI Data Center Capacity Set to Quadruple to 4.9 GW by 2033 With $60 Billion Buildout•All Three Majors Are Now Suing Anthropic. The AI Copyright Endgame Is Taking Shape•The Backlash Has Arrived: Asia's Data Center Boom Is Colliding With the People Who Live Next Door•Groq, Nemotron, Now Hugging Face: Nvidia Isn't Buying Companies — It's Buying the Open AI Ecosystem•Travis Kalanick's Atoms Eyes Robotaxis — With Uber's $100 Million and His Old AV Team•Fluidstack Closes $1.5 Billion Led by Jane Street — Revenue Up From $1.8M to a Projected $660M, While Owning Zero Chips•Anthropic Walks Away From $6 Billion Decart Acquisition After Due Diligence•OpenAI Becomes Anchor Customer of Nvidia-Backed Firmus for Malaysian AI Factories•Japan's AI Data Center Capacity Set to Quadruple to 4.9 GW by 2033 With $60 Billion Buildout•All Three Majors Are Now Suing Anthropic. The AI Copyright Endgame Is Taking Shape•The Backlash Has Arrived: Asia's Data Center Boom Is Colliding With the People Who Live Next Door•Groq, Nemotron, Now Hugging Face: Nvidia Isn't Buying Companies — It's Buying the Open AI Ecosystem•Travis Kalanick's Atoms Eyes Robotaxis — With Uber's $100 Million and His Old AV Team•Fluidstack Closes $1.5 Billion Led by Jane Street — Revenue Up From $1.8M to a Projected $660M, While Owning Zero Chips•
Black Forest Labs FLUX 3 multimodal foundation model announcement graphic
Black Forest Labs
Research

FLUX 3: Black Forest Labs Trains One Model to Generate Video, Audio and Robot Actions

The German lab's new multimodal flow model produces 20-second clips with natively synchronized audio from a single set of weights — and extends the same architecture to robotic action prediction, already tested on Audi production lines.

D
Daniel ParkAI Correspondent
5 min read

Black Forest Labs has released FLUX 3, and the headline is not any single capability but the architecture underneath all of them: a single multimodal foundation model, trained jointly on images, video and audio, that generates 20-second video clips with natively synchronized sound — and predicts robot actions from the same set of weights.

The Freiburg-based lab, best known for the FLUX image models that power much of the open image-generation ecosystem, is making an explicit bet with this release. As the company puts it in its announcement, a model "must learn a representation of the world" — not spatial structure, motion, or sound in isolation, but all of them together, the way physical events actually produce them.

One Flow, Every Modality

FLUX 3 is built on flow matching, extended through what Black Forest Labs calls its Self-Flow approach for aligning multimodal generation and understanding. Where most video systems bolt an audio model onto a video model — generating frames first, then scoring them with sound — FLUX 3 produces video and audio in a single inference pass from one flow-matching framework. The practical result is audio that is causally tied to what happens on screen: the model learned to associate sounds with the physical events that make them, rather than layering plausible ambience over finished footage.

The video system handles text-to-video, image-to-video, video-to-video, keyframe transitions and multilingual dialogue, with clips running up to 20 seconds at a time. For longer work, the lab demonstrates agentic chaining — stitching generated shots into multi-shot sequences lasting several minutes — along with typography generation strong enough for titles and on-screen text, a longstanding FLUX differentiator carried over from the image line.

From Pixels to Actuators

The most consequential extension may be the least flashy: native action prediction. Black Forest Labs argues that a model which has learned how the world moves and sounds has also learned much of what a robot needs to act in it. Working with robotics partner mimic, the lab built FLUX-mimic for dexterous manipulation tasks — and reports the system has been tested in production at Audi, moving the claim from demo reel to factory floor.

That framing places FLUX 3 in the same lineage as world-model efforts from DeepMind and NVIDIA, but with an unusual path: arriving at embodied prediction from a generative media model rather than from simulation infrastructure.

Early Numbers

The lab's preliminary human-preference evaluations put FLUX 3 ahead of the current video-generation field: preferred 77% of the time against Runway's Gen-4.5, 93% against Luma's Ray 3.2, and 69% against xAI's Grok Imagine Video. Those are vendor-run comparisons and should be read accordingly — no independent benchmark suite yet exists for joint audio-video generation — but the margins are wide enough to signal a genuine step, and Black Forest Labs says results will improve as the early-access models are refined.

Availability is staged. FLUX 3 Video is in early access now; FLUX 3 Image opens in the coming weeks; the action-prediction stack is limited to selected partners. Critically for the open ecosystem — including the Asian developer communities that built tooling economies around FLUX.1's open weights — the company has committed to an open-weight FLUX 3 Dev release later this year.

Why It Matters

FLUX 3 lands in a week when the frontier is consolidating around unified world representations rather than per-modality specialists. If a mid-sized European lab can ship competitive video, native audio and production-tested robot actions from one training run, the economics of maintaining separate video, audio and robotics models start to look questionable — for everyone from ByteDance's Seedance team to the robotics startups building on task-specific vision-language-action models. The open FLUX 3 Dev release, when it arrives, will be the real test: it would hand the first jointly trained video-audio foundation model to the open-weight community, and history suggests that community moves fast.

Newsletter

Get Lanceum in your inbox

Weekly insights on AI and technology in Asia.

Share

More in Research

Lanceum

Independent coverage of AI and technology across Asia. We go beyond headlines to explain what matters.

Colophon

Typeset in Space Grotesk & DM Serif Display. Built with Nuxt & Tailwind. Powered by curiosity.

© 2026 Lanceum. All rights reserved.

Independent • Rigorous • Asia-Focused