AssemblyAI 发布 Universal-3.5 Pro Realtime 流式语音转文本模型

Artificial Analysis · @ArtificialAnlys · X·2026-07-06 23:54·57天前
AI 导读

AssemblyAI 推出流式 STT 模型 Universal-3.5 Pro Realtime,为 Universal-3 Pro Realtime 升级版。Max Accuracy 模式下 AA-WER Streaming 词错误率 4.1%,首次最终转录延迟 0.44 秒;Min Latency 模式 WER 4.3%,延迟 0.40 秒。准确率优于 Deepgram Flux(7.4%,0.02s)和 Nova-3 Realtime(6.6%,0.07s),略低于 Cartesia Ink-2(3.7%,0.09s)和 ElevenLabs Scribe v2 Realtime(3.6%,0.14s)。新增对话上下文动态更新功能(每轮 Agent 回复后可刷新,无需重连)。价格维持 $0.45/小时,语言支持从 6 种扩展至 18 种,支持句中代码切换。

Artificial Analysis@ArtificialAnlys
57AI 编辑部评分,满分 100

AssemblyAI 发布 Universal-3.5 Pro Realtime 流式语音转文本模型

2026-07-06 23:54· 57天前
AI 导读

AssemblyAI 推出流式 STT 模型 Universal-3.5 Pro Realtime,为 Universal-3 Pro Realtime 升级版。Max Accuracy 模式下 AA-WER Streaming 词错误率 4.1%,首次最终转录延迟 0.44 秒;Min Latency 模式 WER 4.3%,延迟 0.40 秒。准确率优于 Deepgram Flux(7.4%,0.02s)和 Nova-3 Realtime(6.6%,0.07s),略低于 Cartesia Ink-2(3.7%,0.09s)和 ElevenLabs Scribe v2 Realtime(3.6%,0.14s)。新增对话上下文动态更新功能(每轮 Agent 回复后可刷新,无需重连)。价格维持 $0.45/小时,语言支持从 6 种扩展至 18 种,支持句中代码切换。

AssemblyAI has released Universal-3.5 Pro Realtime: a streaming Speech to Text model achieving 4.1% WER on AA-WER Streaming (~0.4s to first final), able to take in conversation context at the start of a call and after each agent turn, without reconnecting

Universal-3.5 Pro Realtime is AssemblyAI's latest streaming Speech to Text (STT) model, the successor to Universal-3 Pro Realtime. The model offers three default modes, Balanced (default), Max Accuracy, and Min Latency, each a combination of lower level streaming parameters. On AA-WER Streaming, the Max Accuracy variant is level with its predecessor, Universal-3 Pro Realtime, and the Min Latency variant is ~10% faster.

The model now takes conversation context that can be updated turn by turn. Agent's replies can be passed in at connection and refreshed mid-stream after each turn with no reconnect, giving more contextually relevant outputs. For example, priming it with "What's your email address?" yields "user@gmail.com" instead of "user at gmail dot com".

Key takeaways ➤ First Final Transcription: Universal-3.5 Pro Realtime achieves a 4.1% WER at 0.44s after end of speech in Max Accuracy mode, more accurate than the faster Deepgram Flux (7.4%, 0.02s) and Deepgram Nova-3 Realtime (6.6%, 0.07s), and behind the more accurate Cartesia Ink-2 external endpoints (3.7%, 0.09s) and ElevenLabs Scribe v2 Realtime (3.6%, 0.14s). Min Latency mode achieves a 4.3% WER, slightly faster at 0.40s. ➤ First Partial Transcription: WER for the Max Accuracy variant on First Partial is the same as First Final, 4.1% at 0.44s after end of speech, ahead of Cartesia Ink-2 external endpoints (4.3%, 0.07s) and behind only ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) on accuracy, though slower to emit than both. Min Latency mode trades accuracy for speed, returning a first partial at 6.1% WER at 0.39s ➤ Price: Universal-3.5 Pro Realtime costs $0.45/hr ($7.50 per 1,000 minutes), unchanged from Universal-3 Pro Realtime ➤ Language support: The model supports 18 languages, up from 6 in Universal-3 Pro Realtime, with mid-sentence code-switching.

See more details below ⬇️

来源:Artificial Analysis· x.com