Alibaba's Qwen-Audio-3.0-TTS-Plus is the new leading model on the Artificial Analysis Speech Arena Leaderboard for Provider Voices, narrowly surpassing Simba 3.2 and ahead of Gemini 3.1 Flash TTS, and Sonic 3.5
Qwen-Audio-3.0-TTS-Plus is Alibaba's latest Text to Speech model, released with increased naturalness and contextually appropriate intonation. The release continues Alibaba's strong momentum across AI model releases, following recent leading launches in language, image, and video generation.
Key takeaways:
➤ Quality: Qwen-Audio-3.0-TTS-Plus has an Elo score of 1,236 (+17/-17) based on 1,305 arena appearances, narrowly ahead of Simba 3.2 at 1,234 (+17/-17), with overlapping confidence intervals, and ahead of Gemini 3.1 Flash TTS at 1,214 and Sonic 3.5 at 1,207.
➤ Throughput speed: The model generates 16 characters per second, below other leading models Simba 3.2 (30.2), Gemini 3.1 Flash TTS (27), and Sonic 3.5 (120)
➤ Pricing: Qwen-Audio-3.0-TTS-Plus is priced at $27.59 per 1M characters via Alibaba Cloud Model Studio, below Sonic 3.5 ($39.00/1M), above Gemini 3.1 Flash TTS ($18.31/1M) and Simba 3.2 ($10.00/1M).
See more details and listen to samples below ⬇️