MiniMax Music 3.0 lands as open weights — full songs, run local
A state-of-the-art vocal model you can download today: five-minute tracks, running from 8GB of VRAM, ComfyUI on day one, and a license with no geo-blocks.

Copy markdown
A full song, sung, from lyrics and a vibe
Feed MiniMax Music 3.0 lyrics with [Verse]/[Chorus] tags plus a plain-text description and it composes, arranges, and sings a complete track up to five minutes — intro, chorus, bridge, outro — as 32kHz 16-bit stereo. The weights are open and live on Hugging Face now.
Runs on a 24GB card, or 8GB if you're patient
Standard inference wants 24GB+ of VRAM, but CPU offload and layer-streaming squeeze it onto an 8GB GPU (CUDA required). It ships with SGLang-Omni, Diffusers, and ComfyUI 0.33+ paths, so you can drop it into a local pipeline tonight.
The license finally dropped the geo-blocks
MiniMax's H3 video model carved out the US, EU, UK, and South Korea; Music 3.0's community license has no territorial exclusions at all. Commercial use is free under $20M in annual revenue — bigger companies need written sign-off — and any public output must be disclosed as AI-generated.
Under the hood, and why it sounds cleaner
The stack is an 8B 'global' LLM (initialized from Qwen3.5-8B), a 0.6B local LLM, a 2.4B flow-matching stage, and a 123M flow-VAE decoder — continuous audio, not token-only. MiniMax claims sharper reading of creative intent, fuller arrangements, and more natural sound than the 2.x line.
Build this weekend: your own soundtrack engine
With local weights and a permissive license, a game or app can spin up on-brand background music, jingles, or demo tracks on demand — no per-call meter, no royalty clearance. Just wire in the AI-disclosure label if the output ships publicly.