Inworld AI released Realtime TTS-2, a text-to-speech model that processes the full audio context of multi-turn exchanges before it speaks, adapting to the moment the way a person would.
One voice identity across 100+ languages.
Sub-200ms time-to-first-audio.
Natural-language voice direction, no emotion tag presets.
AI that hears how you sound, not only what you say, is now a real architecture decision.