Qwen-Audio-3.0-TTS tops Speech Arena at a third of ElevenLabs' price

Plus edges Simba 3.2 for the #1 provider-voice slot; steer delivery in plain English with emotion tags — but it's hosted-only and slow at 16 chars/sec.

Nowline JUL 25 7:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • #1 on the Speech Arena, by a whisker

    Qwen-Audio-3.0-TTS-Plus takes the top provider-voice spot on Artificial Analysis' Speech Arena at ~1,236 Elo, nipping Simba 3.2 (1,234) and beating Gemini 3.1 Flash TTS and Sonic 3.5. For once, the best-sounding hosted voice isn't from ElevenLabs.

  • A third of ElevenLabs' price

    It runs $27.60 per 1M characters through Alibaba Cloud Model Studio — roughly a third of what ElevenLabs and MiniMax charge for comparable quality. That turns 'add narration' from a budget line into a rounding error.

  • Direct it in plain English

    Steer delivery with natural-language instructions ('read this slowly, like a bedtime story') plus 86 inline tags like [whisper], [angry], and [laughing]. It covers 16 languages including Tagalog, Thai, and Vietnamese, plus 20 Chinese dialect regions.

  • Flash for live agents, Plus for polish

    Flash hits ~300ms first-packet latency over WebSocket streaming — usable for real-time voice agents — while Plus maxes out quality. Model IDs are qwen-audio-3.0-tts-flash and -plus, with PCM/WAV/MP3/Opus output to 48kHz and SDKs in Python, Node, Go, Java, C#, and PHP.

  • Clone a voice from a noisy clip

    Voice cloning holds up even with noisy reference audio, and one-pass synthesis runs up to 3 minutes — enough to stand up a custom narrator or character voice this weekend without a clean studio recording.

  • The catch: hosted-only and slow

    There are no open weights — it's DashScope API only — and throughput is ~16 chars/sec versus Sonic 3.5's 120. Fine for short agent replies; painful for bulk audiobook or podcast synthesis.