FLUX 3 generates synced video, audio and robot actions from one model

Black Forest Labs' debut video model is gated early-access for now; the real payoff is FLUX 3 Dev, an open-weight multimodal backbone promised later in 2026.

Nowline Jul 26 8:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • One model, three modalities

    FLUX 3 is Black Forest Labs' first video model: up to 20 seconds of 720p video with natively synchronized audio in a single pass, plus a FLUX-mimic head that outputs robot actions. It runs on one “Self-Flow” architecture, not stitched-together sub-models.

  • You can't run it yet — it's a waitlist

    Video and Action ship as gated early access only, by application at bfl.ai; there's no public API and no announced pricing. FLUX 3 Image is “coming weeks,” and the open-weight Dev tier is explicitly last in line — “later in 2026.”

  • The open-weight tier is the real builder story

    FLUX 3 Dev would be the first downloadable model that jointly generates video, audio and images from a single backbone — a plausible default for local multimodal video. Don't re-platform on the promise, and read the license when it lands: FLUX.1-era permissiveness isn't guaranteed to carry over.

  • Trust the tie, not the 93%

    BFL's own evals claim 77–93% preference over Runway Gen-4.5 and Luma Ray 3.2, but the methodology, sample size and prompt set are unpublished. The number that matters: a statistical tie (~52%) with Google's already-shippable Gemini Omni Flash and ByteDance's Seedance 2.0.

  • FLUX-mimic is already on a factory line

    The action model reports ~101ms end-to-end reaction time and is being tested on Audi production robots for soft-body manipulation, kitting and assembly. BFL is chasing physical AI, not just creative tooling — worth watching if you build robotics or simulation pipelines.