# There is a subtle architecture shift happening in voice AI. The voice stack is becoming part of the…

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-08 04:51
- AIHOT 分数：50
- AIHOT 链接：https://aihot.virxact.com/items/cmtrqiopr0cn3rotn4tju5wos
- 原文链接：https://x.com/rohanpaul_ai/status/2097065255839088698

## 正文

There is a subtle architecture shift happening in voice AI.

The voice stack is becoming part of the agent's execution loop.

@cartesia is combining the listening and speaking paths around that loop.
Sonic-3.6 turns text into speech (90ms latency) and Ink-2 turns speech into text (100ms transcript latency), faster than anything else streaming.

now holds the #1 spot for both speaking and listening models.

Becasue, voice agents are one of those systems where 100ms in the wrong place is very noticeable.

Cartesia is attacking both sides of that loop at once.

super interesting work by @krandiash and team.

### 引用推文

> Karan Goel：We released Sonic-3.5 and Ink-2, the #1 streaming models for text to speech and speech to text you can use in your voice agents today. New architectures enable ...
