新一代语音模型,专为自然的人机交互而生,现已为 ChatGPT 语音功能提供支持。
我们正在推出 GPT‑Live,这是一代全新的语音模型,能让与 AI 的对话感觉更像是在进行一场真实的交流。
GPT‑Live 基于全双工架构构建,这意味着它可以同时进行听和说。在对话过程中,GPT‑Live 可以通过“嗯哼”或“对”这样的词语来表明它在认真倾听,能够进行快速的你来我往,或者在你需要思考片刻时保持安静。最终带来的是一种令人耳目一新、易于交流的语音体验。
GPT‑Live 也是我们迄今为止最智能的语音模型。对于需要网络搜索、更深层次推理或更复杂工作的问题,它会在后台委托给我们最新的前沿模型,并在准备好后将结果带回对话中。在模型处理期间,GPT‑Live 可以继续与你交谈并保持对话的流畅性。上线之初,GPT‑Live 将在后台使用 GPT‑5.5。随着我们发布新的前沿模型,我们将持续更新 GPT‑Live 所使用的模型。
这些进步为全新的 ChatGPT 语音体验提供了动力,使其更加智能、使用起来也更加自然。随着时间的推移,我们相信这项研究还将解锁语音在日益复杂、耗时更长且更具智能体特性的工作中的应用能力。
从今天开始,我们开始向全球的 ChatGPT 用户逐步推出两个版本的 GPT‑Live——GPT‑Live‑1 和 GPT‑Live‑1 mini。我们还计划很快将其引入 API,开发者和企业可以通过此表格注册以接收通知。
进入人机交互的新时代
我们的愿景是实现真正自然的人机交互:一个与 AI 协作感觉就像与另一个人合作一样流畅和响应迅速的世界,同时推理和复杂任务的执行在后台无缝进行。
以往的方法
前几代的语音 AI 系统让我们更接近这一愿景,但伴随着重要的权衡取舍。
级联式语音系统
级联语音系统依赖一系列模型依次运作,以处理每一轮交互。最初的 ChatGPT 语音功能将三个模型串联在一起:一个语音转文本模型用于转录你的语音,一个大语言模型用于生成回复,以及一个文本转语音模型用于将回复转换回语音。这种方法使我们首次能够与前沿 AI 模型对话,但这种复杂性是有代价的:信息可能在模型间丢失,且回复缓慢而僵硬。
语音转文本
文本转语音
语音转文本
文本转语音
轮次式语音模型
像 ChatGPT 高级语音模式这样的轮次式语音模型,在单个模型内处理和生成音频,减少了延迟,使对话更流畅——但它们仍然通过离散的轮次运作。模型必须等待用户停止说话后才能回应,导致僵硬的来回交替。此外,由于轮次检测基于静默,即使是短暂的停顿或背景噪音也可能被误判为轮次结束——导致模型在不自然的时间点打断。
ChatGPT 高级语音模式对话示例
我们的新方法
GPT‑Live 通过两项架构变革解决了这些局限。
持续交互
首先,我们构建了 GPT‑Live 以实现持续交互,采用了全双工架构。GPT‑Live 并非处理一系列独立消息,而是在生成输出的同时持续处理输入。因此,模型可以每秒多次做出交互决策:是否说话、继续倾听、暂停、打断或调用工具。
这使得模型能够进行更自然的来回交流,保持更好的时间感,甚至执行实时翻译。
深度工作的委派
其次,我们将处理持续交互的 GPT‑Live 与深度工作解耦。当某个问题需要搜索、推理或更强的智能体能力时,GPT‑Live 可以将任务委派给另一个模型,如 GPT-5.5。这使得它能够在后台处理多个任务的同时,保持对话继续进行。
这一架构变化还让 GPT‑Live 能够持续使用最新模型和智能体,将前沿智能与自然交互融为一体。
评估
我们构建了全新的人工评估体系,用于衡量对话的愉悦度和流畅性。在这些一对一对比测试中,GPT‑Live‑1 和 GPT‑Live‑1 mini 在 5–10 分钟的匹配对话中,在整体偏好、话轮转换、打断、对话流畅度以及交互自然感等维度上,均显著优于高级语音模式。
GPQA:GPT‑Live‑1 在 GPQA 基准测试上大幅超越高级语音模式,该测试评估生物学、化学和物理学领域的专家级科学推理能力。
BrowseComp:GPT‑Live‑1 在 BrowseComp 基准测试上相比高级语音模式有显著提升,该测试评估智能体网络搜索以及查找难以定位信息的能力。
τ³-Voice Telecom(内部变体)**:GPT‑Live‑1 在 τ³-Voice Telecom 测试中优于高级语音模式,该测试评估语音智能体在真实的多轮电信支持任务中的表现。
GPT‑Live‑1(instant)和 GPT‑Live‑1 mini 在后台使用 GPT‑5.5 Instant 模型,而 GPT‑Live‑1 Medium 和 GPT‑Live‑1 High 则使用 GPT‑5.5 Thinking 模型,并分别采用中等和高推理努力程度。
** 本次评估中,我们使用了由最新推理模型驱动的定制化用户模型。
全新的 ChatGPT 语音体验
每周,超过 1.5 亿人使用语音和听写等功能与 ChatGPT 对话。他们用它来获取免提的日常帮助、练习语言、讲睡前故事,或是在通勤途中闲聊。
从今天起,当你点击语音按钮与 ChatGPT 对话时,将获得由 GPT‑Live 驱动的升级体验——对话更自然、回答更智能、倾听更精准,并支持视觉回应。
更自然的对话
与 ChatGPT 对话现在应该感觉更像真正的交谈了。你可以随时打断提问、暂停整理思绪,或请 ChatGPT 放慢语速。它会自然地用“嗯哼”或“明白了”这样的回应来确认你的话,让你知道它在认真倾听。我们还为 GPT‑Live 重新制作了 ChatGPT 中九种各具特色的语音。
更智能的回答
ChatGPT 语音现在可以调用我们最新的前沿模型,在你需要时给出更智能的回答。你还可以选择适合自己需求的推理层级:即时模式用于快速响应,中或高模式则在你希望 ChatGPT 花更多时间思考时使用。
更好的倾听
如果你需要片刻思考,ChatGPT 语音现在会耐心等待,而不会抢话打断。如果你要求它保持安静并倾听,它也会照做。当有背景噪音时,比如过往车辆或附近交谈,ChatGPT 能更好地聚焦你的声音,而不会被干扰分心。
一目了然的视觉答案
有些答案用眼睛看会更直观。在对话过程中,ChatGPT 现在可以针对天气、股票、体育等话题展示丰富的视觉卡片。语音功能也继续支持搜索、记忆、图片和文件上传。
最终呈现的 ChatGPT 语音体验,在日常使用中更加自然、更加强大、更加实用。
专为语音设计的安全保障
GPT‑Live 在设计之初就以安全为默认原则。它建立在我们最新模型的安全进展之上,同时针对关键风险领域增加了专门的安全训练,并加入了专为语音设计的新防护措施。
扩展的安全测试
为了更好地反映人们在真实场景中如何使用语音,我们首先扩展了安全测试,加入了全新的原生音频评估。我们还创建了合成评估,利用生成的音频更集中地针对关键安全领域进行测试,这些经验来自高级语音模式。这些领域包括自残、精神病与躁狂、对 AI 的情感依赖、暴力以及色情内容。内部专家还对模型进行了针对语音特有风险的红队测试。
在我们的测试中,GPT‑Live 在几乎所有评估维度上的表现都与高级语音模式相当或更优。您可以在 GPT‑Live 系统卡中了解更多关于我们的测试和安全保障措施的信息。
内置安全保障
由于语音对话是实时进行的,我们还构建了能够在模型说话时即时介入的安全保障措施。当系统检测到潜在的不安全输出时,它可以引导模型给出更安全的回应,显示额外的安全提示或资源,或在风险较高的情况下终止语音对话。对于涉及自残行为的对话,我们针对语音场景调整了 ChatGPT 的支持流程,包括提供经专家验证的危机热线支持。
我们设计了额外的保护措施来支持青少年用户,并直接在模型中训练了符合年龄特征的行为,以降低不当回应的风险。家长可以通过家长控制功能选择是否允许其青少年子女使用 ChatGPT 语音功能,并且在涉及潜在自残或自杀意图迹象的高风险情况下,关联的家长可能会收到通知。
从实际使用中学习
我们还在推出以情感依赖为重点的长期测量和发布后监控,以持续改进我们的理解并完善安全措施。基于我们先前在情感使用和情感健康方面的研究,这将帮助我们识别新出现的模式,并改进系统在情感敏感互动中的回应方式。
最后,GPT‑Live 是为对话而设计的,而非用于声音模仿。它使用 ChatGPT 中一组预定义的语音,并设有安全措施以防止其模仿真实人物的声音。
我们致力于支持安全与福祉,并将随着从实际使用中不断学习,持续加强这些保护措施。
可用性与限制
GPT‑Live 现已面向全球 ChatGPT 用户,在 iOS、Android 和 ChatGPT.com 平台上逐步推出。GPT‑Live‑1 将成为 Go、Plus 和 Pro 用户 ChatGPT 语音功能的默认模型,而 GPT‑Live‑1 mini 将成为免费用户的默认模型。更多可用性详情,请查阅我们的帮助中心。
我们已针对 ChatGPT 中最常用的几种语言对 GPT‑Live 进行了优化。对于某些语言,模型可能存在非母语口音或流利度不足的问题。我们正在积极努力,以提升跨语言的使用体验。
在发布之初,GPT‑Live 将不支持 ChatGPT 中的视频语音或屏幕共享功能,但我们正在努力尽快引入这些能力。您仍可访问 ChatGPT 语音的旧版本,包括标准语音模式和高级语音模式,这些功能在这些版本中是可用的。
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.
GPT‑Live is built on a full-duplex architecture, meaning it can listen and speak at the same time. During conversations, GPT‑Live can show it’s paying attention with phrases like “mhmm” or “yeah”, engage in quick back-and-forth, or just stay quiet when you need a moment to think. The result is a voice experience that is refreshingly easy to talk to.
GPT‑Live is also our smartest voice model yet. For questions that require web search, deeper reasoning, or more complex work, it delegates to our latest frontier model behind the scenes and brings the result back into the conversation when it’s ready. While it works, GPT‑Live can keep talking with you and maintain the flow of conversation. At launch, GPT‑Live will use GPT‑5.5 in the background. As we release new frontier models, we’ll continuously update the model used by GPT‑Live.
These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work.
We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today. We also plan to bring them to the API soon, and developers and enterprises can sign up to be notified using this form.
Entering a new era of human-AI interaction
Our vision is to enable truly natural human–AI interaction: a world where collaborating with AI feels as fluid and responsive as working with another person, while reasoning and complex task execution happen seamlessly in the background.
Previous approaches
Older generations of voice AI systems brought us closer to that vision, but with important tradeoffs.
Cascaded voice systems
Cascaded voice systems rely on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe your speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled us to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.
STT
GPT-5.5
TTS
STT
GPT-5.5
TTS
Example conversation with Standard Voice Mode, using GPT-5.5 Instant
Turn-based voice models
Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother — but they still operated through discrete turns. The model had to wait for the user to stop speaking before responding, resulting in rigid back-and-forth. In addition, because turn detection is based on silence, even a brief pause or background noise could be mistaken for the end of turn — causing the model to interrupt at unnatural times.
Example conversation with ChatGPT Advanced Voice Mode
Our new approach
GPT‑Live addresses these limitations through two architectural changes.
Continuous interaction
First, we built GPT‑Live for continuous interaction using a full-duplex architecture**.**Instead of processing a sequence of separate messages, GPT‑Live continuously processes input while generating output. The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool.
This allows the model to engage in more natural back-and-forth, maintain a better sense of time, and even perform live translation.
User
Example conversation with GPT-Live-1, using GPT-5.5 Instant
Delegation for deeper work
Second, we decoupled GPT‑Live — which handles continuous interaction — from deeper work. When a question requires search, reasoning, or more agentic capabilities, GPT‑Live can delegate the task to another model like GPT‑5.5. This allows it to keep the conversation going, even as it handles multiple tasks in the background.
This architectural change also allows GPT‑Live to continuously use the latest models and agents, combining frontier intelligence with natural interaction.
Search + reason
Example conversation with GPT-Live-1, using GPT-5.5 Instant
Evaluations
We built new human evaluations to measure pleasantness and the flow of conversation. In these head-to-head comparisons, GPT‑Live‑1 and GPT‑Live‑1 mini are strongly preferred over Advanced Voice Mode in matched 5–10 minute conversations that measure overall preference, turn-taking, interruptions, conversational flow, and how natural each interaction felt.
GPQA: GPT‑Live‑1 substantially outperforms Advanced Voice Mode on GPQA, which tests expert-level scientific reasoning across biology, chemistry, and physics.
BrowseComp: GPT‑Live‑1 shows strong gains over Advanced Voice Mode on BrowseComp, which tests agentic web search and the ability to find difficult-to-locate information.
τ³-Voice Telecom (internal variant)**: GPT‑Live‑1 outperforms Advanced Voice Mode on τ³-Voice Telecom, which tests voice agents on realistic, multi-turn telecom support tasks.
GPT‑Live‑1 (instant) and GPT‑Live‑1 mini use the GPT‑5.5 Instant model in the background, while GPT‑Live‑1 Medium and GPT‑Live‑1 High use the GPT‑5.5 Thinking model with medium and high reasoning effort.
** We used a customized user model, powered by our latest reasoning models, for this eval.
A new ChatGPT Voice experience
Each week, more than 150 million people talk to ChatGPT using features like Voice and Dictation. They use it to get hands-free everyday help, to practice languages, tell bedtime stories, or just chat during their commute.
Starting today, when you tap the Voice button to talk with ChatGPT, you’ll get an improved experience powered by GPT‑Live—with more natural conversations, smarter answers, better listening, and visual responses.
More natural conversations
Talking with ChatGPT should now feel much more like a real conversation. You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you’re saying with phrases like “mhmm” or “got it,” so you know it’s following along. We’ve also remastered the nine distinct voices in ChatGPT for GPT‑Live.
Smarter answers
ChatGPT Voice can now draw on our latest frontier models, giving you smarter answers when you need them. You can also choose the level of reasoning that fits your needs: Instant for fast responses, or Medium and High when you want ChatGPT to spend more time thinking.
Better listening
If you take a moment to think, ChatGPT Voice now waits instead of jumping in and interrupting. If you ask it to stay quiet and listen, it will. And when there’s background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on your voice instead of getting distracted.
Visual answers at a glance
Some answers are more useful when you can see them. While you’re talking, ChatGPT can now show rich visual cards for topics like weather, stocks, sports, and more. Voice also continues to support search, memory, images, and file uploads.
The result is a ChatGPT Voice experience that feels more natural, more capable, and more useful in everyday life.
Safety designed for voice
GPT‑Live was designed to be safe by default. It builds on the safety advances from our latest models while adding dedicated safety training across key risk areas and new safeguards designed specifically for voice.
Expanded safety testing
To better reflect how people use voice in real-life settings, we began by expanding our safety testing to include new audio-native evaluations. We also created synthetic evaluations that use generated audio to focus more intensively on key safety areas, drawing on what we learned from Advanced Voice Mode. Those areas include self-harm, psychosis and mania, emotional reliance on AI, violence, and sexual content. Internal experts also red-teamed the model for risks unique to voice.
In our testing, GPT‑Live performed comparably to or better than Advanced Voice Mode across nearly all of the areas we evaluated. You can read more about our testing and safeguards in the GPT‑Live system card .
Built-in safeguards
Because voice conversations unfold in real time, we also built safeguards that can act while the model is speaking. When the system detects potentially unsafe output, it can steer the model toward a safer response, surface additional safety messaging or resources, or end the voice conversation in higher-risk cases. For conversations involving self-harm, we adapted ChatGPT’s support flows for voice, including offering expert-vetted crisis helpline support.
We designed additional protections to support teen users, and trained age-appropriate behavior directly into the model to reduce the risk of inappropriate responses. Parents can choose whether their teen can use ChatGPT Voice through Parental Controls, and linked parents may be notified in higher-risk situations involving signs of potential self-harm or suicidal intent.
Learning from real-world use
We’re also rolling out longer-term measurement and post-launch monitoring focused on emotional reliance to continue improving our understanding and refining safeguards. Building on our previous research into affective use and emotional well-being, this will help us identify emerging patterns and improve how the system responds in emotionally sensitive interactions.
Finally, GPT‑Live is designed for conversation, not voice impersonation. It uses a set of predefined voices in ChatGPT, with safeguards to prevent it from imitating a real person’s voice.
We’re committed to supporting safety and well-being and will keep strengthening these protections as we learn from real-world use.
Availability & limitations
GPT‑Live is rolling out now to ChatGPT users globally across iOS, Android, and ChatGPT.com . GPT‑Live‑1 will become the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT‑Live‑1 mini will become the default for Free users. More availability details can be found in our Help Center .
We’ve optimized GPT‑Live for some of the most popular languages in ChatGPT. For certain languages, the model may have a non-native accent or gaps in fluency. We’re actively working to improve the experience across languages.
At launch, GPT‑Live will not support voice with video or screen sharing in ChatGPT, but we’re working to introduce these capabilities soon. You can still access legacy versions of ChatGPT Voice, including Standard and Advanced Voice Mode, where these features are available.