# OpenAI 推出 GPT Transcribe 与 GPT Live Transcribe 语音识别模型

- 来源：The Decoder：AI News（RSS）
- 作者：Jonathan Kemper
- 发布时间：2026-07-29 20:45
- AIHOT 分数：44
- AIHOT 链接：https://aihot.virxact.com/items/cms64axgo184rrobkz58gdrwx
- 原文链接：https://the-decoder.com/gpt-transcribe-improves-on-its-predecessor-but-cant-catch-elevenlabs-google-or-mistral-on-error-rates

## AI 摘要

OpenAI 通过 API 推出 GPT Transcribe 和 GPT Live Transcribe 两款语音识别模型。GPT Transcribe 处理预录音频速度约为实时的 34 倍，词错误率为 3.31%，定价降至每分钟 $0.0045。

## 正文

OpenAI has released GPT Transcribe and GPT Live Transcribe, two new speech recognition models available through its API. GPT Transcribe handles pre-recorded audio files, processing them about 34 times faster than real time. GPT Live Transcribe is built for real-time streaming with low latency.

According to Artificial Analysis, which runs the AA-WER benchmark, GPT Transcribe hits a word error rate of 3.31 percent. That's a 0.7 percentage point improvement over its year-old predecessor GPT-4o Transcribe. Pricing drops 25 percent at the same time, landing at $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and multiple input languages.

视频 · 前往原文观看

In the AA-WER ranking, OpenAI still sits behind several competitors. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral recently undercut the market with Voxtral Transcribe V2, starting at just $0.003 per minute.

Full details are in OpenAI's . The new transcription models complement OpenAI's recently announced Realtime model generation, which also includes the real-time transcription model GPT-Realtime-Whisper.

AI News Without the Hype – Curated by Humans
