我们最新的语音转文字模型,专为精准、智能的实时转写而设计。
Diego Melendo Casado
Gemini Audio 工程高级总监
Luke Leonhard
Gemini Audio 幕僚长,代表 Gemini Audio 团队

语音
今天,我们推出 Gemini 3.5 Transcribe——这是我们迄今为止最精准的语音转文字模型,专为智能语音交互而设计。与那些在背景噪音、复杂术语和语流不清清理方面表现不佳的传统语音识别模型不同,Gemini 3.5 Transcribe 能将原始音频直接转换为准确、精炼、格式规范的文本。
在我们的各类产品中(如 Gemini 应用及 Android 平台),我们已经看到消费者正从这一转写模型中受益,包括 Android 上的 Rambler 以及 macOS 上 Gemini 应用中的全新语音功能。现在,开发者可以通过 Google AI Studio 中的 Gemini API 以及 Gemini Enterprise Agent Platform,使用 Gemini 3.5 Transcribe 构建类似的功能。
我们将 3.5 Transcribe 设计为可无缝接入你的开发者工作流,无论你是在构建语音智能体、实时字幕工具,还是通话后分析管道。该模型可通过两个独立的 API 使用:
- 实时流式:通过 Live API 使用 gemini-3.5-transcribe-live****,提供亚秒级延迟的持续双向流式传输,适用于交互式语音应用。
- 预录音频处理:通过 Interactions API 使用 gemini-3.5-transcribe****,可转写录音音频、会议、通话记录等,并支持说话人归属和词级时间戳。
获得更精准、更智能的转写
Gemini 3.5 Transcribe 旨在捕捉你的自然说话风格,以更好地理解你的意图并识别自定义词汇,让你可以用语音完成任务。
- 智能转写:无缝处理自我纠正(如“我们周二见面——不,周三”),移除填充词(“嗯”和“啊”),并自动格式化你的文本。
- 函数调用:该模型可以通过函数调用将复杂任务(如图像生成和文件分析)委派给其他 Gemini 模型。目前已在 Gemini macOS 应用中提供。
- 更精准的转写:根据 Artificial Analysis 的评测,该模型在流式场景下平均词错误率(WER)为 4.0%,非流式场景下为 2.6%。在嘈杂的真实环境中表现出色,能够准确识别邮政编码、订单号等字母数字实体。
- 自定义词汇:通过将转写结果无缝适配到你提供的自定义词汇表,识别专业术语和特殊拼写。
- 全球语言支持:自动检测并转写超过 85 种语言,无缝处理地区口音和多样方言。
- 多说话人识别:在预录音频中准确区分说话人,并为最多三位说话人提供时间戳(3 人以上支持为实验性功能)。
Gemini 3.5 Transcribe 支持实时语言切换和无缝流式转写
观看 Gemini 3.5 Transcribe 如何通过智能转写能力清理语音中的不流畅表达。
3.5 Transcribe 提供带多说话人归因和词级时间戳的转写结果。
Gemini 3.5 Transcribe 的性能相比我们之前的转写模型 Chirp 3 实现了重大飞跃,带来了新能力、更优的词错误率以及显著更低的延迟。根据 Artificial Analysis 的评测,以最终转写完成时间为例,提升了 70%。在覆盖多种主要语言和地区的 FLEURS 基准上,该模型展现出精准的多语言表现,优于 Chirp 3,流式模式下 WER 为 5.50%,非流式场景下 WER 为 5.04%。


体验智能转写与高级听写
除了 Google AI Studio 和 Gemini Enterprise Agent Platform 中的 Gemini API 之外,3.5 Transcribe 还超越了标准语音转文本的能力,让 Google 生态内的操作更自然、更直观。通过将上下文感知理解直接带入 Gboard、Antigravity、Gemini 应用和 Chrome 等日常使用场景,它能够轻松捕捉细微差别、意图和行内编辑。
- 在 Android 版 Gboard 上,通过全新的 Rambler 功能,3.5 Transcribe 可将口述想法转换为格式规整的文本,并过滤掉口头禅。您还可以使用语音进行编辑、纠正拼写错误以及调整写作风格。
- 在 Google Antigravity 上,经您许可后,3.5 Transcribe 会结合屏幕上下文和聊天记录,确保对文件名、智能体思考过程以及当前活动文档实现精准的转写准确度。
- 在 Google AI Studio 中,您可以在 Build 模式下访问 3.5 Transcribe,随时用语音即时编写(vibe code)应用。
- 在 macOS 版 Gemini 应用中,3.5 Transcribe 不仅可将您自然流畅的语音转写为整洁的格式化文本,还支持语音指令,并能与屏幕上下文无缝结合,驱动复杂的工作流程。通过在后台调用其他 Gemini 模型来处理繁重任务,该模型让您仅凭语音即可轻松总结本地文件、跨应用复用文本,或直接在光标处生成图像。
- 即将在 Chrome 浏览器中推出,您将可以在任意网页输入框内通过语音输入文字——让您在 Chrome 中口述回复、撰写帖子或向 Gemini 提问变得更加自然、轻松。
Gemini 3.5 Transcribe 让您仅凭语音即可在 macOS 版 Gemini 应用中分析文件、生成图像和执行搜索。
了解 Gemini 3.5 Transcribe 如何在 Android 上利用 Rambler 自动去除口头禅并优化语音内容。
Gemini 3.5 Transcribe 利用 Google Antigravity 上的屏幕上下文来确保转写的准确性。
阅读早期评测
通过利用 Gemini Live API,Agora、Fishjam、LangChain、LiveKit、Pipecat、Vercel 和 Vision Agents 等开发者平台使开发者能够轻松构建和部署高性能的语音驱动界面。这些平台在后台管理复杂的实时媒体流基础设施,让开发者能够完全专注于打造用户体验。
Vivo、Intellitek Health 和 Lingopal 等公司也对 3.5 Transcribe 给予了积极反馈,称赞其出色的延迟表现、准确性以及广泛的语言支持。






立即开始使用 3.5 Transcribe
- 面向开发者:现已在 Gemini API 中通过 Google AI Studio 和 Google Antigravity 提供公开预览。
- 面向企业:现已在 Gemini Enterprise Agent Platform 中提供公开预览,并即将在 Gemini Enterprise for Customer Experience 中推出。
- 面向所有用户:现已在 macOS 版 Gemini 应用中提供英文版本,在部分国家和语言的 Android 版 Rambler 中提供,并即将在 Chrome 中推出。
将 Google 的最新资讯发送到您的收件箱
Our latest speech-to-text model designed for precise and intelligent real-time transcription.
Diego Melendo Casado
Senior Director, Engineering, Gemini Audio
Luke Leonhard
Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team

Voice
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:
- Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using
gemini-3.5-transcribe-live****. - Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using
gemini-3.5-transcribe****.
Get more precise and intelligent transcription
Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
- Smart transcription: Seamlessly handles self-corrections (like "let’s meet Tuesday—no, Wednesday"), removes filler words (“ums” and ‘“ahs"), auto-formats your text.
- Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app.
- More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.
- Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.
- Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.
- Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).
Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription
Watch Gemini 3.5 Transcribe clean up speech disfluencies with smart transcription capabilities.
3.5 Transcribe delivers transcription with multi-speaker attribution and word-level timestamps.
Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.


Experience smart transcription and advanced dictation
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.
- On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.
- On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.
- In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.
- In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.
- Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.
Gemini 3.5 Transcribe lets you analyze files, generate images, and search in the Gemini app on macOS using just your voice.
See how Gemini 3.5 Transcribe uses Rambler on Android to automatically remove filler words and clean up speech.
Gemini 3.5 Transcribe leverages screen context on Google Antigravity to ensure accurate transcription accuracy.
Read the early reviews
By leveraging the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.
Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.






Start using 3.5 Transcribe today
- For developers: In public preview in the Gemini API via Google AI Studio and Google Antigravity.
- For enterprises: In public preview via Gemini Enterprise Agent Platform and coming soon to Gemini Enterprise for Customer Experience.
- For everyone: In Gemini app on macOS in English, Rambler on Android in select countries and languages, and coming soon to Chrome.