Change language to
0:00

Qwen Audio 3.1 brings five models to Alibaba’s voice stack, covering speech recognition, text-to-speech, real-time conversation and audio creation. Alibaba says API prices fall by about 70% for TTS, 85% for Realtime and up to 95% for ASR. The official Qwen announcement lists the same reductions, while The Decoder’s report outlines the new model roles.

Subscribe to our Newsletters for more Tech Stories

What Qwen Audio 3.1 adds

The Qwen Audio 3.1 family includes Qwen-Audio-3.1-ASR for speech recognition, Qwen-Audio-3.1-TTS for voice synthesis and Qwen-Audio-3.1-Realtime for live voice interaction. The two new models are Qwen-Audio-3.1-ASR-Next, which analyses speakers, timestamps, emotions and background sounds, and Qwen-Audio-3.1-TTS-Next, which can generate voices, sound effects and ambient audio together.

Qwen says the standard ASR model supports 30 languages and 16 Chinese dialects, with first-character response of about 160 milliseconds. Its built-in transcription polishing removes fillers and repetitions, which may save developers from adding a second clean-up pass to meeting and interview tools.

Qwen Audio 3.1 ASR chart comparing dialect recognition error rates

The technical report for Qwen Audio 3.1 Realtime says the model improves on the previous version in multilingual benchmarks, speech-conditioned tool use and selected full-duplex behaviours. Qwen describes Realtime as able to listen and speak at the same time, with interruptions supported during a conversation. It can also call APIs, knowledge bases and business systems.

Why the lower prices matter

The price cuts give developers more room to build voice features into customer service, transcription, accessibility and multilingual apps without treating every conversation as a small infrastructure event. The move also puts pressure on other voice API providers, following the voice AI systems that listen and talk at once already entering the market.

Alibaba has not published a UAE-specific price or availability detail in the material reviewed for this report. UAE developers can watch the growing market for developer voice models through Alibaba’s Qwen platform, but regional billing and data-handling terms still need checking before a production deployment.

Which developers are most likely to benefit from the Qwen Audio 3.1 cuts?

Teams building high-volume transcription, customer-service, accessibility or multilingual voice features are the clearest beneficiaries because lower per-call costs reduce the pressure to limit usage.

Does Qwen Audio 3.1 confirm UAE pricing?

No. The announcement material reviewed does not publish a UAE-specific price or regional billing detail, so developers should check Alibaba’s current Model Studio terms before deployment.

What is the difference between ASR-Next and TTS-Next?

ASR-Next focuses on understanding complete audio, including speakers and background sounds, while TTS-Next creates voice, effects and ambient audio together.