MiniMax H3 opens its weights: first open model to top a video arena
The 33B omni-modal model does 768p video with native stereo audio and runs on a 24GB card — but the license bars the US, EU, UK and Korea.

Copy markdown
First open model to top a video arena
MiniMax declared H3 the SOTA open video generator on both Artificial Analysis and LMArena. It took #1 for audio-judged video editing (~1130 Elo) and top-three in text-to-video and image-to-video — the first open-weights model to lead a video board. Caveat: those scores measure MiniMax's hosted pipeline, not your local checkpoint render.
Video and native audio in one pass
H3 is a 33.1B omni-modal transformer: text, images, video and audio in, 4-15s clips at 768p/24fps with native 32kHz stereo audio out — no separate TTS or foley step. Its Omni-Reference mode pulls identity from a product image, motion from a reference clip and a soundscape from an audio file in a single generation.
Run it locally this weekend in ComfyUI
Native ComfyUI support merged Aug 3 — four new nodes and six official templates. Pruned INT8 checkpoints cut the download to ~19.5GB; a 24-32GB GPU runs it comfortably, and a 12GB card works with heavy RAM offload (just slowly). Local text-to-video with sound, zero API spend.
The license locks out the US, EU, UK and Korea
The MiniMax H3 Community License bars use — and even outputs generated — in the US, EU, UK and South Korea, under Hong Kong law. Elsewhere, commercial use needs a visible 'MiniMax H3' credit in your UI plus written sign-off above $20M revenue, and you can't distill it to train other models.
Or rent it if you're fenced out
In an excluded region the weights are off-limits, but hosted APIs aren't: MiniMax's own platform and fal.ai serve H3 at ~$0.13/sec for 2K (a 15s clip ≈ $1.95), and Atlas Cloud runs it globally at $0.14/sec. You can also apply for a separate license through the repo.