Inworld's newly released Realtime TTS-2 is the new #1 on the Artificial Analysis Controlled Voice Arena, just ahead of Cartesia Sonic 3.6, and ranks #2 on our Provider Voice Arena behind Sonic 3.6
Realtime TTS-2 is a new Text to Speech model from @inworld_ai that supports over 100 languages, including English, Hindi, Spanish, French, German, Chinese, and Japanese. Language can be set explicitly or detected from the text, including multiple languages in a single request (e.g., "I'll grab a coffee. ¿Quieres uno? お疲れさま。"). It also supports delivery instructions written in plain text alongside the input (e.g., [speak tired but warm, like she just got home from a long day]).
Key takeaways: ➤ Controlled Voice: Realtime TTS-2 takes #1 on the Controlled Voice Arena with an Elo of 1,123 (+16/-16) across 1,292 appearances, 4 points ahead of Cartesia's Sonic 3.6 at 1,119. Sonic 3.5 follows at 1,096, then Inworld’s own Realtime TTS-2 Flash - Research Preview at 1,075, and ElevenLabs Eleven v3 at 1,062 ➤ Provider Voice: Realtime TTS-2 takes #2 on the Provider Voice Arena with an Elo score of 1,252 (+18/-18) based on 1,094 arena appearances, placing it ahead of Alibaba Qwen-Audio-3.0-TTS-Plus at 1,241 and Speechify Simba 3.2 at 1,240, but behind Sonic 3.6 at 1,282 ➤ Throughput: The model processes 106 characters per second of generation time, compared to 124 for Sonic 3.6, 97 for Simba 3.2, and 40 for Eleven v3 ➤ Pricing: Realtime TTS-2 is priced at $20.83 per 1M characters, lower priced than Sonic 3.6 ($49), Eleven v3 ($100), and Qwen-Audio-3.0-TTS-Plus ($27.59), but more expensive than Simba 3.2 ($10)
See more details and listen to samples in the thread below ⬇️