Seedance 2.5: 30s single-pass video, now with native audio
ByteDance's answer to Sora and Veo hits enterprise beta: 50 reference assets, in-clip edits and lip-synced audio, with general access set for Aug 7.

Copy markdown
30 seconds, one pass, with sound
Seedance 2.5 renders a single native 30-second clip, up from the old ~15s ceiling, with built-in audio in 10+ languages including audio-driven pacing, beat-matching and lip-sync. Text-, image- and reference-to-video all run through one model.
Up to 50 references, and real in-clip edits
It takes up to 50 multimodal reference assets (30 images, 10 videos, 10 audio; images up to 4K) to lock a character, product or style, and does localized edits, swap a product, change a background, extend a shot, without re-rolling the whole scene.
What you could build this weekend
One prompt plus a few brand assets now yields a 30-second spot, product demo or short-drama scene with synced voiceover, the kind of programmatic video pipeline that used to mean stitching several 5-8s clips. Run your own eval first: independent benchmarks aren't out.
The catch: enterprise beta now, general access Aug 7
As of the July 31 launch it's live only in enterprise beta with no free quota; ByteDance targets full availability August 7. Public per-second pricing and third-party endpoints (fal, Replicate) aren't confirmed, and aggregators like kie.ai still list it as coming soon.