FLUX 3: one model for 20-sec video+audio, images and robot actions
Black Forest Labs' unified multimodal model is real but gated: Video and Action are apply-only now, Image lands in weeks, open weights wait for late 2026.

Copy markdown
One model, four modalities
FLUX 3 jointly trains image, video, audio and robot-action prediction in one architecture — Black Forest Labs' first step past still images into video and physical AI. CEO Robin Rombach argues audio and language carry timing and intent that pixels alone can't.
20-second video with synced audio
The Video variant makes clips up to 20 seconds with synchronized audio from a single model, taking text, image or video in, with keyframe transitions, multilingual dialogue and multi-shot chaining. BFL says it already leads early evaluations against frontier video models.
The catch: apply-only, no public API
Only approved testers can touch Video and Action right now — you request access on BFL's site, there's no open endpoint. FLUX 3 Image joins the program in the coming weeks, and a FLUX 3 Dev build is aimed at developers.
Open weights wait for late 2026
No downloadable parameters yet: BFL says the open-weight FLUX 3 Dev edition lands "later in 2026," while pricing, service terms and full benchmarks stay undisclosed. Self-hosters and cost-planners are on hold for now.
What it unlocks once it opens
One call for a 20-second narrated product video, or an action model for a robot arm, collapses pipelines that today need three separate tools. Canva, Krea, Picsart, Magnific and Audi are already testing variants.