# 阿里 Qwen-Audio-3.0-TTS-Plus 登顶文本转语音排行榜

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-07-21 19:31
- AIHOT 分数：44
- AIHOT 链接：https://aihot.virxact.com/items/cmrulfvjp0wz8bi07hjbe34nl
- 原文链接：https://the-decoder.com/alibabas-qwen-audio-3-0-tts-plus-tops-the-competition-in-the-text-to-speech-rankings

## AI 摘要

阿里云新模型 Qwen-Audio-3.0-TTS-Plus 以 1,236 的 Elo 评分登顶 Artificial Analysis 语音竞技场排行榜，略超 Simba 3.2（1,234）。

## 正文

Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, leads Artificial Analysis' Speech Arena leaderboard for provider voices. With an Elo score of 1,236, it sits just ahead of Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind.

Alibaba's Qwen-Audio-3.0-TTS-Plus takes the top spot on Artificial Analysis' Text to Speech Leaderboard for provider voices, edging out SpeechifyAI's Simba 3.2 by just two Elo points. | Image: Artificial Analysis

The model comes in two versions. Flash is built for real-time interaction with about 300 milliseconds of latency, while Plus targets high-quality speech output. It supports 16 languages, including less commonly covered ones like Tagalog, Malay, Thai, and Vietnamese, along with several Chinese dialects. Users can steer the speaking style with natural language or add nonverbal cues using tags like "[angry]" or "[giggles]." Alibaba also says the model handles noisy or echo-heavy reference recordings better than previous versions when cloning voices.

Speed is a weak spot: At 16 characters per second, it trails Sonic 3.5 (120) and Simba 3.2 (30.2) by a wide margin. Pricing lands at $27.60 per million characters through Alibaba Cloud Model Studio. A collection of audio samples is available here.
