# AssemblyAI 发布 Universal-3.5 Pro Realtime 流式语音转文本模型

- 来源：Artificial Analysis (@ArtificialAnlys)
- 发布时间：2026-07-06 23:54
- AIHOT 分数：57
- AIHOT 链接：https://aihot.virxact.com/items/cmr9fdamd03rmslsm5tv5nkh5
- 原文链接：https://x.com/ArtificialAnlys/status/2074160133702402314

## AI 摘要

AssemblyAI 推出流式 STT 模型 Universal-3.5 Pro Realtime，为 Universal-3 Pro Realtime 升级版。Max Accuracy 模式下 AA-WER Streaming 词错误率 4.1%，首次最终转录延迟 0.44 秒；Min Latency 模式 WER 4.3%，延迟 0.40 秒。准确率优于 Deepgram Flux（7.4%，0.02s）和 Nova-3 Realtime（6.6%，0.07s），略低于 Cartesia Ink-2（3.7%，0.09s）和 ElevenLabs Scribe v2 Realtime（3.6%，0.14s）。新增对话上下文动态更新功能（每轮 Agent 回复后可刷新，无需重连）。价格维持 $0.45/小时，语言支持从 6 种扩展至 18 种，支持句中代码切换。

## 正文

AssemblyAI has released Universal-3.5 Pro Realtime: a streaming Speech to Text model achieving 4.1% WER on AA-WER Streaming (~0.4s to first final), able to take in conversation context at the start of a call and after each agent turn, without reconnecting

Universal-3.5 Pro Realtime is AssemblyAI's latest streaming Speech to Text (STT) model, the successor to Universal-3 Pro Realtime. The model offers three default modes, Balanced (default), Max Accuracy, and Min Latency, each a combination of lower level streaming parameters. On AA-WER Streaming, the Max Accuracy variant is level with its predecessor, Universal-3 Pro Realtime, and the Min Latency variant is ~10% faster.

The model now takes conversation context that can be updated turn by turn. Agent's replies can be passed in at connection and refreshed mid-stream after each turn with no reconnect, giving more contextually relevant outputs. For example, priming it with "What's your email address?" yields "user@gmail.com" instead of "user at gmail dot com".

Key takeaways
➤ First Final Transcription: Universal-3.5 Pro Realtime achieves a 4.1% WER at 0.44s after end of speech in Max Accuracy mode, more accurate than the faster Deepgram Flux (7.4%, 0.02s) and Deepgram Nova-3 Realtime (6.6%, 0.07s), and behind the more accurate Cartesia Ink-2 external endpoints (3.7%, 0.09s) and ElevenLabs Scribe v2 Realtime (3.6%, 0.14s). Min Latency mode achieves a 4.3% WER, slightly faster at 0.40s.
➤ First Partial Transcription: WER for the Max Accuracy variant on First Partial is the same as First Final, 4.1% at 0.44s after end of speech, ahead of Cartesia Ink-2 external endpoints (4.3%, 0.07s) and behind only ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) on accuracy, though slower to emit than both. Min Latency mode trades accuracy for speed, returning a first partial at 6.1% WER at 0.39s
➤ Price: Universal-3.5 Pro Realtime costs $0.45/hr ($7.50 per 1,000 minutes), unchanged from Universal-3 Pro Realtime
➤ Language support: The model supports 18 languages, up from 6 in Universal-3 Pro Realtime, with mid-sentence code-switching.

See more details below ⬇️
