Rohan Paul@rohanpaul_ai
43AI 编辑部评分,满分 100
2026-08-12 04:12· 13分钟前
AI 导读

Soniox 发布 TTS v2 文本转语音模型,以每小时 0.70 美元的价格提供高端音质,支持 60 多种语言、自然语言混合、低延迟流式传输和高保真语音克隆。该模型专为实时 AI 智能体设计,支持字符级时间戳和音频标签控制,便于打断与恢复。

I wasn't expecting Soniox TTS v2 (a text-to-speech model) to sound this natural.

They just released this TTS v2

Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour)

while keeping the same model suitable for real-time agents, multilingual speech, expressive control, and cloning.

exceptional precision, high-fidelity voice cloning

more than 60 languages, natural language mixing, and low-latency streaming together in one model.

One model replaces a lot of voice-stack plumbing: expression control, voice cloning, 60+ languages, language mixing, pronunciation precision, and streaming all sit in the same system.

It is unusually well designed for live AI agents: low-latency streaming plus character-level timestamps let an agent start talking early, stop cleanly when interrupted, and resume without repeating itself.

Voice performance becomes programmable: developers can insert audio tags for whispering, excitement, laughter, pauses, and other delivery changes inside the generated text.

SonioxIntroducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, ...

来源:Rohan Paul · x.com

Rohan Paul · @rohanpaul_ai · X·2026-08-12 04:12·13分钟前
AI 导读

Soniox 发布 TTS v2 文本转语音模型,以每小时 0.70 美元的价格提供高端音质,支持 60 多种语言、自然语言混合、低延迟流式传输和高保真语音克隆。该模型专为实时 AI 智能体设计,支持字符级时间戳和音频标签控制,便于打断与恢复。

I wasn't expecting Soniox TTS v2 (a text-to-speech model) to sound this natural.

They just released this TTS v2

Really premium voice quality at a dramatically lower price ($ 0.70-per-generated-hour)

while keeping the same model suitable for real-time agents, multilingual speech, expressive control, and cloning.

exceptional precision, high-fidelity voice cloning

more than 60 languages, natural language mixing, and low-latency streaming together in one model.

One model replaces a lot of voice-stack plumbing: expression control, voice cloning, 60+ languages, language mixing, pronunciation precision, and streaming all sit in the same system.

It is unusually well designed for live AI agents: low-latency streaming plus character-level timestamps let an agent start talking early, stop cleanly when interrupted, and resume without repeating itself.

Voice performance becomes programmable: developers can insert audio tags for whispering, excitement, laughter, pauses, and other delivery changes inside the generated text.

SonioxIntroducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, ...

来源:Rohan Paul· x.com