I wasn't expecting Soniox TTS v2 (a text-to-speech model) to sound this natural.
They just released this TTS v2
Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour)
while keeping the same model suitable for real-time agents, multilingual speech, expressive control, and cloning.
exceptional precision, high-fidelity voice cloning
more than 60 languages, natural language mixing, and low-latency streaming together in one model.
One model replaces a lot of voice-stack plumbing: expression control, voice cloning, 60+ languages, language mixing, pronunciation precision, and streaming all sit in the same system.
It is unusually well designed for live AI agents: low-latency streaming plus character-level timestamps let an agent start talking early, stop cleanly when interrupted, and resume without repeating itself.
Voice performance becomes programmable: developers can insert audio tags for whispering, excitement, laughter, pauses, and other delivery changes inside the generated text.