MiniMax H3 tops a video arena, but open weights lock out the West
The 33B omni-model makes 15-second clips with synced audio in one pass—yet the US, EU, UK and Korea can't run it locally, and its 2K stage stays closed.

Copy markdown
First downloadable model to top a video arena
MiniMax H3 ranks #1 in video editing on Artificial Analysis' arena—plus #2 in text-to-video and #3 in image-to-video—the first publicly downloadable model to lead any category. It's a 33.1B dense transformer that uses Qwen3-VL-32B as its text encoder.
Picture and sound in a single pass
H3 renders 4–15s clips at 768p/24fps with 32kHz stereo—dialogue, foley, score and room tone generated alongside the video, no separate audio model. You can condition on up to 9 images, 3 clips and 3 audio tracks, so lip-synced shorts and music-video edits become a weekend build.
The license bars the US, EU, UK and Korea
The community license prohibits local deployment in the United States, EU, UK and South Korea, and firms above $20M revenue need a separate commercial deal. Builders in blocked regions get only MiniMax's hosted API—not the 498GB weights on Hugging Face.
The 'open' release is deliberately capped
The download omits H3-Regenerate-2K (the 2K upscaler) and H3-Context-IR (prompt processing), so local output tops out at 768p versus the hosted service. It lands the same stretch Alibaba shipped Qwen Image 3.0 API-only with no weights at all—the open media wave is narrowing.