FLUX 3: one model for 20-sec video+audio, images and robot actions

Black Forest Labs' unified multimodal model is real but gated: Video and Action are apply-only now, Image lands in weeks, open weights wait for late 2026.

Nowline JUL 29 5:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • One model, four modalities

    FLUX 3 jointly trains image, video, audio and robot-action prediction in one architecture — Black Forest Labs' first step past still images into video and physical AI. CEO Robin Rombach argues audio and language carry timing and intent that pixels alone can't.

  • 20-second video with synced audio

    The Video variant makes clips up to 20 seconds with synchronized audio from a single model, taking text, image or video in, with keyframe transitions, multilingual dialogue and multi-shot chaining. BFL says it already leads early evaluations against frontier video models.

  • The catch: apply-only, no public API

    Only approved testers can touch Video and Action right now — you request access on BFL's site, there's no open endpoint. FLUX 3 Image joins the program in the coming weeks, and a FLUX 3 Dev build is aimed at developers.

  • Open weights wait for late 2026

    No downloadable parameters yet: BFL says the open-weight FLUX 3 Dev edition lands "later in 2026," while pricing, service terms and full benchmarks stay undisclosed. Self-hosters and cost-planners are on hold for now.

  • What it unlocks once it opens

    One call for a 20-second narrated product video, or an action model for a robot arm, collapses pipelines that today need three separate tools. Canva, Krea, Picsart, Magnific and Audi are already testing variants.