Google 发布 Gemini 3.5 Transcribe 转录模型,可自动去除"um""ah"等填充词

The Verge:AI(RSS)·2026-08-27 01:00·1小时前·Jess Weatherbed
AI 导读

Google 推出 Gemini Audio 家族新成员 Gemini 3.5 Transcribe,可自动识别专业术语和超过 85 种语言,并自动格式化文本、去除“um”“uh”等填充词。该模型支持自定义词汇表、为最多三位说话人区分语音,并提供词级时间戳。今日起向 macOS Gemini 应用用户开放英语版本,开发者可通过 Gemini API 公开预览使用。

The Verge:AI(RSS)
63AI 编辑部评分,满分 100

Google 发布 Gemini 3.5 Transcribe 转录模型,可自动去除"um""ah"等填充词

2026-08-27 01:00· 1小时前· Jess Weatherbed
AI 导读

Google 推出 Gemini Audio 家族新成员 Gemini 3.5 Transcribe,可自动识别专业术语和超过 85 种语言,并自动格式化文本、去除“um”“uh”等填充词。该模型支持自定义词汇表、为最多三位说话人区分语音,并提供词级时间戳。今日起向 macOS Gemini 应用用户开放英语版本,开发者可通过 Gemini API 公开预览使用。

We got Gemini Audio models while we’re still waiting for the overdue Gemini 3.5 Pro launch.

We got Gemini Audio models while we’re still waiting for the overdue Gemini 3.5 Pro launch.

STK255_Google_Gemini_D STK255_Google_Gemini_D Jess Weatherbed

Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we’re still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June.

Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.”

Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps.

Gemini 3.5 Transcribe is rolling out starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. It’s also available for developers in public preview in the Gemini API via AI Studio and Antigravity. Google says that Chrome support is coming soon.

Correction, August 26th: Google previously mentioned two additional Gemini Audio models, 3.5 Live, and 3.5 Experimental, in information provided to The Verge prior to publication, but now says that only 3.5 Transcribe is being announced today.