Alibaba ships Qwen-Audio 3.1 with five voice models and cuts ASR prices up to 95%
The Decoder (corroborated by Qwen on X and AlphaSignal)·medium signal
The stack covers ASR, TTS, Realtime and two new models. ASR-Next adds speaker, timestamp, emotion and background-sound analysis, and TTS-Next generates voices, effects and ambient audio together. Prices drop about 70% for TTS, 85% for Realtime and up to 95% for ASR. Standard ASR covers 30 languages plus 16 Chinese dialects at roughly 160ms to first character. The price cut makes Qwen a serious option for anyone building high-volume voice agents.