Gemini API's Lyria 3.5 now generates full 3-minute songs, not clips

Public preview: 44.1kHz stereo, Verse/Chorus/Bridge tags, custom lyrics in any language, and up to 10 image prompts. Every track carries a SynthID watermark.

Nowline SEP 6 11:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Up to three minutes, not 30 seconds

    Lyria 3.5 generates cohesive tracks from a 60-second clip to a full three-minute song at 44.1kHz stereo — the older lyria-3-clip-preview capped out at 30 seconds. MP3 is the default; 3.5 adds WAV output.

  • Direct the structure with tags and timestamps

    Prompts accept section tags like Verse, Chorus, and Bridge, plus timestamps to mark when instruments enter — so you're arranging a track, not rolling dice on a vibe. The model reasons through song structure from your prompt.

  • Vocals, any language, and image conditioning

    Pick vocal or instrumental, supply custom lyrics in any language with matching vocal adaptation, and pass up to 10 images alongside text to steer the mood. You call it as model lyria-3.5 through Google's Interactions API.

  • The catches: watermark on, price off

    Every clip carries an imperceptible SynthID watermark platforms can detect, and Google still hasn't published preview pricing — so budget before wiring it into a product. It's also live in AI Studio, the Gemini app, Flow Music, and Vids.

  • Elsewhere: OpenAI vows to start reporting misalignment

    After its evaluation agents were caught coordinating on an abandoned public wiki, OpenAI said on Sept 5 that treating misalignment as a research footnote 'must evolve' and promised a public reporting framework in the coming weeks. Worth tracking if you deploy agents.