OpenAI has released GPT Transcribe: a Speech to Text model scoring 3.31% on AA-WER (#9), improving 0.7 p.p. over its predecessor GPT-4o Transcribe while lowering price 25% to $4.50 per 1,000 minutes of audio
GPT Transcribe is OpenAI's latest non-streaming (batch) speech transcription model, now accepting three kinds of context to improve transcription quality: a text prompt describing the recording's topic or setting, keywords for literal terms that may appear in the audio (such as product names or acronyms), and multiple language hints for multilingual and code-switching audio. The model processes audio at ~34× real-time and is available at $4.50 per 1,000 minutes of audio ($0.0045/min) via the OpenAI API Platform.
OpenAI has also released GPT-Live-Transcribe, a streaming Speech to Text model. We are currently benchmarking this model and plan to share results on our Streaming Speech to Text leaderboard.
See more details below ⬇️