Qwen-Audio 3.1: five-model voice stack, API prices cut up to 95%

New TTS-Next and ASR-Next models arrive as Alibaba slashes audio APIs; the White House moves to gate UK model testing and Meta hands agents a full cloud PC.

Nowline SEP 25 6:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Five models, one audio stack

    Qwen-Audio-3.1 upgrades ASR, TTS, and Realtime and adds two new models: TTS-Next generates speech, sound effects, and background audio in a single pass, and ASR-Next does multi-speaker diarization with emotion and ambient-sound detection. It spans 30 languages and 16 Chinese dialects over the API, though the ASR-Next endpoint isn't live yet.

  • Audio API prices cut up to 95%

    TTS drops about 70%, Realtime about 85%, and ASR up to 95%. Published rates put ASR-Flash at ¥0.8 in / ¥2.7 out and TTS-Flash at ¥1.5 in / ¥12 out per million tokens — cheap enough to leave transcription or a voice agent running without watching the meter.

  • White House moves to gate UK model testing

    Per a Politico report, the administration asked OpenAI and Anthropic to withhold new models from the UK AI Safety Institute until a US cybersecurity review clears them. If it holds, frontier models — and the safety fixes outside testers surface — could reach you later than they used to.

  • Meta gives agents a cloud computer

    Meta's new Muse agent runs on its own full Ubuntu Linux machine in the cloud, so it can browse, run code, and drive apps like a real desktop across sessions. It's a consumer product at $20-$100/mo, but it points where personal agents are going: persistent computers, not chat windows.