Soniox TTS v2 is one of those models you need to actually hear.
They just launched TTS v2, a text-to-speech model that speaks 60+ languages from one mode.
Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour).
One model replaces most of the voice-stack plumbing, since expression control, cloning, language mixing, pronunciation precision and streaming sit in the same system instead of 3 stitched-together vendors.
Voice performance becomes programmable. Because, audio tags allow developers to direct emotion, delivery, and vocal reactions throughout the text instead of selecting one fixed style for the entire passage. like [whispering], [excited] or [laughing] straight into the text,
Soniox TTS v2 is built for live agents, and voice agents require more than natural-sounding audio.
Speech must begin quickly, remain synchronized with the conversation, stop immediately when the user interrupts, and avoid repeating text that has already been spoken.
Interrupt a normal voice agent and it has no clue how much you actually heard, so it repeats itself or skips ahead.
TTS v2 timestamps every character down to the millisecond, so it knows the exact word you cut it off at and carries on from there.
Cloning from seconds of audio, keeping accent, rhythm and personality, with noise removed from the source clip first, so a phone recording still works. That voice holds its identity across all 60+ languages and switches language mid-sentence.
Precision where realistic models usually break, on codes like 7Q4M9B, prices, email addresses and medical terms.
Architecture, audio codec and inference engine are all in-house, which is what makes the price possible. Live as tts-rt-v2 in the US, Europe and Japan, backward compatible with v1.
🧵 1.