MiniMax Music 3 open weights: 5-minute songs with vocals, run local

An 8B model that fits a 24GB card — or 8GB with streaming, or $0.15/track hosted. Plus: Unsloth Desktop trains locally; Palmyra X6 cuts agent costs 52%.

Nowline AUG 14 2:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Full songs, vocals and all, up to five minutes

    Music 3 turns a prompt plus lyrics into a complete track — expressive vocals with controllable gender and timbre, 32kHz 16-bit stereo, up to five minutes long. That's a finished song, not a loop, and it handles multiple languages including Mandarin and English.

  • An 8B stack that fits a single consumer GPU

    The model is an 8B global LLM (initialized from Qwen3-8B) plus a 0.6B local LLM, a 2.4B flow-matching module and a 123M Flow-VAE. Full precision fits under 24GB (an RTX 4090); layer-streaming drops it onto an 8GB card, and one user generated on a 16GB AMD RX 7900 GRE in about 125 seconds.

  • Weights on Hugging Face, ComfyUI support day one

    The weights are live on Hugging Face with a community GGUF port already up, and ComfyUI 0.33.0 loads it out of the box (diffusers too) — so a local song-generation node is a weekend build. 'Open weights' is the pitch, but the license text isn't spelled out; read the HF LICENSE file before shipping anything commercial.

  • Or skip the GPU: $0.15 a track hosted, free tier included

    Don't want to run it locally? The hosted API is $0.15 per song at 120 RPM, and a Music-3.0-free tier costs nothing (throttled to 3 RPM) — enough to prototype a music feature into your app without buying hardware.

  • Elsewhere: Unsloth Desktop puts local training in one app

    Unsloth shipped a free, open-source desktop app (Apache-2.0 core) that both fine-tunes and runs 500+ models on Windows, macOS and Linux — LoRA training, GGUF export, and OpenAI-compatible serving with no notebook. It's an early beta, so expect rough edges.

  • Elsewhere: Writer's Palmyra X6 claims a 52% cut in agent costs

    Writer launched its Palmyra X6 flagship with an upgraded agent harness it says trims token spend on agentic work by about 52% — worth a look if you run long agent loops and watch the bill, though it leans enterprise.