# OpenAI 推出 GPT Transcribe 和 GPT Live Transcribe 语音识别模型

- 来源：The Decoder：AI News（RSS）
- 作者：Jonathan Kemper
- 发布时间：2026-07-29 20:45
- AIHOT 分数：54
- AIHOT 链接：https://aihot.virxact.com/items/cms638c9h16vxrobkc9wovjbh
- 原文链接：https://the-decoder.com/openais-new-transcription-models-reduce-error-rates-and-costs-but-still-lag-behind-elevenlabs-and-google

## AI 摘要

OpenAI 通过 API 推出 GPT Transcribe 和 GPT Live Transcribe 两款新语音识别模型。GPT Transcribe 处理预录音频的速度约为实时的 34 倍，词错误率为 3.31%，较前代降低 0.7 个百分点，定价降至每分钟 $0.0045。

## 正文

OpenAI has released GPT Transcribe and GPT Live Transcribe, two new speech recognition models available through its API. GPT Transcribe handles pre-recorded audio files, processing them about 34 times faster than real time. GPT Live Transcribe is built for real-time streaming with low latency.

According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe hits a word error rate of 3.31 percent. That's a 0.7 percentage point improvement over its year-old predecessor GPT-4o Transcribe. Pricing drops 25 percent at the same time, landing at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages.

视频 · 前往原文观看

In the AA-WER ranking, OpenAI still sits behind several competitors. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute.

Full details are in OpenAI's . The new transcription models complement OpenAI's recently announced Realtime model generation, which also includes the real-time transcription model GPT-Realtime-Whisper.

AI News Without the Hype – Curated by Humans
