今天,我们在 ChatGPT(以 GPT‑5.4 思考模式)、API 和 Codex 中发布 GPT‑5.4。这是我们面向专业工作最强大、最高效的前沿模型。同时,我们还在 ChatGPT 和 API 中发布了 GPT‑5.4 Pro,面向希望在复杂任务上获得极致性能的用户。
GPT‑5.4 将我们近期在推理、编码和智能体工作流方面的最佳进展整合到一个前沿模型中。它融合了 GPT‑5.3‑Codex 业界领先的编码能力,同时改进了模型在工具、软件环境以及涉及电子表格、演示文稿和文档的专业任务中的工作方式。最终得到的模型能够准确、高效、有效地完成复杂的实际工作——用更少的来回交互交付你所要求的结果。
在 ChatGPT 中,GPT‑5.4 思考模式现在可以预先展示其思考计划,让你能在模型工作过程中中途调整方向,从而无需额外轮次即可获得更符合你需求的最终输出。GPT‑5.4 思考模式还改进了深度网络研究,尤其针对高度具体的查询,同时能更好地维持需要更长思考时间的问题的上下文。这些改进共同意味着更高质量的回答,响应更快且始终与当前任务相关。
在 Codex 和 API 中,GPT‑5.4 是我们发布的首个具备原生、最先进计算机使用能力的通用模型,使智能体能够操作计算机并在多个应用间执行复杂工作流。它支持高达 100 万 token 的上下文窗口,让智能体能够跨长周期规划、执行和验证任务。GPT‑5.4 还通过工具搜索改进了模型在大型工具和连接器生态系统中的工作方式,帮助智能体更高效地发现和使用合适的工具,同时不牺牲智能水平。最后,GPT‑5.4 是我们迄今为止 token 效率最高的推理模型,与 GPT‑5.2 相比,解决问题时使用的 token 显著减少——这意味着更低的 token 消耗和更快的速度。
结合在通用推理、编程和专业知识工作方面的进步,GPT‑5.4 能够支持更可靠的智能体、更快的开发者工作流程,并在 ChatGPT、API 和 Codex 中提供更高质量的输出。
GPT‑5.4GPT‑5.3‑CodexGPT‑5.2 GDPval(胜或平)83.0%70.9%70.9% SWE-Bench Pro(公开版)57.7%56.8%55.6% OSWorld-Verified 75.0%74.0%*47.3% Toolathlon 54.6%51.9%46.3% BrowseComp 82.7%77.3%65.8%
*此前报告为 64.7%。GPT‑5.3‑Codex 通过一个新增的 API 参数(用于保留原始图像分辨率)达到了 74.0%。
知识工作
在 GPT‑5.2 通用推理能力的基础上,GPT‑5.4 在专业人士关心的实际任务上,提供了更加稳定和精良的结果。
在测试智能体跨 44 个职业生成规范知识工作能力的 GDPval 上,GPT‑5.4 达到了新的最优水平,在 83.0% 的对比中达到或超越了行业专业人士,而 GPT‑5.2 的这一比例为 70.9%。
在 GDPval 中,模型尝试完成涵盖美国 GDP 贡献最大的前 9 个行业中 44 个职业的规范知识工作。任务要求生成真实的工作成果,例如销售演示文稿、会计电子表格、急诊排班表、制造图纸或短视频。GPT‑5.4 的推理努力设置为 xhigh,GPT‑5.2 设置为 heavy(在 ChatGPT 中略低一级)。
“GPT-5.4 是我们尝试过的最好的模型。它现在在我们衡量专业服务工作模型性能的 APEX-Agents 基准测试中位居榜首。它在创建长期交付物方面表现出色,例如幻灯片组、财务模型和法律分析,在提供顶级性能的同时,运行速度更快、成本低于其他前沿模型。”
—— Brendan Foody,Mercor 首席执行官
我们特别着重提升了 GPT‑5.4 在创建和编辑电子表格、演示文稿及文档方面的能力。在一项针对初级投行分析师可能完成的电子表格建模任务的内部基准测试中,GPT‑5.4 的平均得分为 87.3%,而 GPT‑5.2 为 **68.4%**。在一组演示文稿评估提示词中,人类评估员在 68.0% 的情况下更偏好 GPT‑5.4 生成的演示文稿,而非 GPT‑5.2 的,原因是其美学效果更佳、视觉多样性更强,并且更有效地运用了图像生成功能。
文档生成时推理力度设置为 xhigh。
您可以在 ChatGPT 中通过 GPT‑5.4 Thinking 或 Pro 模式体验这些功能。如果您是企业客户,我们建议使用我们今日同步发布的全新 ChatGPT for Excel 插件。我们还更新了 Codex 和 API 中可用的电子表格与演示文稿技能。
为了让 GPT‑5.4 在实际工作中表现更出色,我们在降低模型幻觉和错误方面持续取得进展。GPT‑5.4 是我们迄今为止事实准确性最高的模型:在一组用户标记了事实错误的脱敏提示词中,与 GPT‑5.2 相比,GPT‑5.4 的单个陈述错误概率降低了 33%,其完整回复包含任何错误的概率降低了 18%。
“GPT-5.4 为文档密集型法律工作树立了新标杆。在我们的 BigLaw Bench 评估中,它获得了 91% 的分数。与其他模型相比,GPT-5.4 目前在构建复杂交易分析、在长篇合同中保持准确性以及提供法律从业者所需的高水平细节方面表现更优。”
—— Niko Grupen,Harvey 应用研究主管
计算机使用与视觉能力
GPT‑5.4 是我们首款具备原生计算机使用能力的通用模型,这标志着开发者和智能体领域迈出了重要一步。对于构建能够在网站和软件系统中完成实际任务的智能体的开发者而言,它是目前可用的最佳模型。
我们设计了 GPT‑5.4,使其在广泛的计算机使用工作负载中均能表现出色。它擅长编写代码,通过 Playwright 等库来操作计算机,也能根据截图发出鼠标和键盘指令。其行为可通过开发者消息进行引导,这意味着开发者能够调整行为以适应特定使用场景。开发者甚至可以通过指定自定义确认策略,来配置模型的安全行为,以适应不同的风险承受水平。
该模型的性能与灵活性,体现在那些测试不同环境下计算机使用能力的各项基准中。在 OSWorld-Verified(该基准衡量模型通过截图及键盘/鼠标操作来导航桌面环境的能力)上,GPT‑5.4 取得了 75.0% 的先进成功率,远超 GPT‑5.2 的 47.3%,并且超越了人类 72.4% 的表现。
在测试浏览器使用的 WebArena-Verified 上,GPT‑5.4 在同时使用 DOM 和截图驱动的交互方式时,取得了领先的 67.3% 成功率,而 GPT‑5.2 为 65.4%。在同样测试浏览器使用的 Online-Mind2Web 上,GPT‑5.4 仅凭基于截图的观察就达到了 92.8% 的成功率,优于 ChatGPT Atlas 的智能体模式(其成功率为 70.9%)。
工具让步(tool yield)是指智能体暂停执行以等待工具响应的过程。如果 3 个工具被并行调用,接着又有 3 个工具被并行调用,那么让步次数就是 2。相比工具调用次数,工具让步次数能更好地反映延迟情况,因为它体现了并行化的优势。
GPT‑5.4 能够解读浏览器界面的截图,并通过基于坐标的点击与 UI 元素进行交互,从而发送电子邮件和安排日历事件。视频未经过加速处理。
GPT‑5.4 改进的计算机使用能力,建立在模型整体视觉感知能力提升的基础之上。在评估模型视觉理解与推理能力的 MMMU-Pro 测试中,GPT‑5.4 在不借助工具的情况下达到了 81.2% 的成功率,相比 GPT‑5.2 的 79.5% 有所提升。视觉感知能力的增强也带来了更好的文档解析表现。在 OmniDocBench 上,GPT‑5.4 在不启用推理努力的情况下,平均误差(以模型预测与真实结果之间的归一化编辑距离衡量)为 0.109,优于 GPT‑5.2 的 0.140。
MMMUPro 测试在推理努力设置为 xhigh 的条件下运行。OmniDocBench 测试在推理努力设置为 none 的条件下运行,以体现低成本、低延迟的性能表现。
我们也在改进针对高密度、高分辨率图像的视觉理解能力,这类场景下图像的全保真度至关重要。从 GPT‑5.4 开始,我们引入了原始图像输入细节级别,支持最高 10.24M 总像素或 6000 像素最大边长(取两者中较小值)的全保真感知;高图像输入细节级别现在支持最高 2.56M 总像素或 2048 像素最大边长。在与 API 用户的早期测试中,我们观察到在使用原始或高细节级别时,定位能力、图像理解能力和点击准确率均有显著提升。
“在我们对约 3 万个业主协会和房产税门户网站进行的计算机使用性能评估中,GPT-5.4 首次尝试的成功率达到 95%,三次尝试内成功率达到 100%,而此前 CUA 模型的成功率约为 73%–79%。同时,其会话完成速度提高了约 3 倍,token 消耗减少了约 70%,在大规模应用中显著提升了可靠性和成本效率。”
— Mainstay 首席执行官 Dod Fraser
在 API 中,开发者可以使用更新后的计算机工具来调用这些能力。请参阅我们更新的文档,了解推荐的最佳实践。
编程
GPT‑5.4 将 GPT‑5.3‑Codex 的编程优势与领先的知识工作和计算机使用能力相结合,这些能力在模型可以使用工具、进行迭代、以更少人工干预推进工作的较长运行任务中尤为重要。在 SWE-Bench Pro 上,GPT‑5.4 的表现与 GPT‑5.3‑Codex 持平或更优,同时在不同推理努力级别下延迟更低。
我们通过观察模型的线上行为并进行离线模拟来估算延迟。延迟估算考虑了工具调用时长(代码执行时间)、生成的 token 数量以及输入的 token 数量。实际延迟可能存在显著差异,并且取决于我们模拟中未涵盖的诸多因素。推理努力程度从无到极高进行了全面扫描。
开启后,Codex 中的 /fast 模式在 GPT‑5.4 上可实现高达 1.5 倍的 token 生成速度。它使用的是同一模型,具备相同的智能水平,只是速度更快。这意味着用户可以在编码任务、迭代和调试过程中保持流畅状态,无需中断。开发者也可以通过 API 使用优先处理功能,以同样快速的速度访问 GPT‑5.4。
在评估和内部测试中,我们发现 GPT‑5.4 在复杂前端任务上表现出色,其生成的结果在美观性和功能性上均显著优于我们此前发布的任何模型。
为了展示该模型在计算机使用与编码能力上的协同提升,我们还发布了一项实验性的 Codex 技能,名为“Playwright (Interactive)”。该技能使 Codex 能够对 Web 应用和 Electron 应用进行可视化调试;它甚至可以在构建应用的同时,对正在构建的应用进行测试。
使用 GPT‑5.4 从一条简单的提示词生成的游乐园模拟游戏,利用 Playwright Interactive 进行浏览器试玩测试,并借助图像生成了等距视角的素材集。该模拟包含基于瓦片的路径铺设、游乐设施与景观建造、游客寻路、排队以及游乐设施运行周期,同时,金钱、游客数量、满意度、清洁度和评分等公园指标会根据布局表现及游客反应而上下浮动。Playwright 被用于自动化浏览器试玩测试,通过建造和扩建公园、放置和移除路径与景点、检查摄像机导航,并验证游客、队列、游乐设施状态及 UI 指标在数轮游玩过程中是否正确更新。
提示词:使用 `$playwright-interactive` 和 `$imagegen`。创建一个交互式的等距视角主题公园模拟游戏,我可以在浏览器中建造并自由探索。利用 imagegen 确立整体视觉风格,并生成游戏资产,包括游乐设施、路径、地形、树木、水域、小吃摊、装饰物、建筑、图标和 UI 插图。游戏世界应具有统一感、精致感和丰富的视觉效果,采用高品质的美术方向,在等距视角下表现良好。让我能够放置和移除路径、添加景点、布置景观,并在公园中流畅移动,同时监控游客活动、游乐设施状态和公园发展情况。包含可信的游客移动逻辑,以及简单的公园管理系统,如金钱、清洁度、排队和满意度,让体验感觉有趣、清晰且完整,而非粗糙的原型。优先考虑魅力、可读性和强烈的游戏感,而非写实主义。
在试玩测试时,务必通过多轮游玩来建造和扩展公园,验证放置和导航功能是否流畅运行,确认游客会对公园布局和景点做出反应,并确保视觉效果、UI 和交互体验稳定且协调。
“GPT-5.4 目前在我们的内部基准测试中处于领先地位。我们的工程师发现它比之前的模型更自然、更果断。它能在处理模糊问题时不会自我怀疑,并且会主动并行化工作以保持进度推进。”
—— Lee Robinson,Cursor 开发者教育副总裁
工具使用
借助 GPT-5.4,我们显著改进了模型与外部工具的协作方式。AI 智能体现在能够在更大的工具生态系统中运作,更可靠地选择正确的工具,并以更低的成本和延迟完成多步骤工作流。
工具搜索
在 API 中,GPT-5.4 引入了工具搜索功能,这使得模型在拥有大量工具时也能高效工作。
此前,当模型被赋予工具时,所有工具定义都会预先包含在提示词中。对于拥有众多工具的系统,这可能会在每个请求中增加数千甚至数万个模型 token,从而提高成本、减慢响应速度,并用模型可能永远不会使用的信息挤占上下文窗口。
借助工具搜索功能,GPT‑5.4 会接收一个轻量级的可用工具列表,并附带工具搜索能力。当模型需要使用某个工具时,它可以查找该工具的定义,并立即将其追加到当前对话中。
这种方法大幅减少了工具密集型工作流所需的 token 数量,并保留了缓存,从而使请求更快、成本更低。同时,它也让智能体能够可靠地处理规模更大的工具生态系统。对于可能包含数万 token 工具定义的 MCP 服务器而言,效率提升尤为显著。
为了展示效率提升,我们在 Scale 的 MCP Atlas 基准测试中评估了 250 个任务,所有 36 个 MCP 服务器以两种模式启用:(1) 将每个 MCP 函数直接暴露在模型上下文中;(2) 将所有 MCP 服务器置于工具搜索之后。采用工具搜索配置后,总 token 用量减少了 47%,同时保持了相同的准确率。
示例 token 数量来自 MCP-Atlas 公开数据集中 250 个任务的平均值。
智能体工具调用
GPT‑5.4 还改进了工具调用,使其在推理过程中决定何时以及如何使用工具时更加准确和高效,尤其是在 API 中。与 GPT‑5.2 相比,它在 Toolathlon 基准测试中以更少的轮次实现了更高的准确率。该基准测试旨在评估 AI 智能体如何使用真实世界的工具和 API 完成多步骤任务。例如,智能体需要读取电子邮件、提取作业附件、上传文件、批改作业并将结果记录到电子表格中。
工具让步(tool yield)是指智能体让步以等待工具响应。如果 3 个工具被并行调用,接着又有 3 个工具被并行调用,那么让步次数为 2。与工具调用次数相比,工具让步次数能更好地反映延迟情况,因为它体现了并行化的优势。
对于偏好将推理努力(reasoning effort)设为“无”的延迟敏感型用例,GPT‑5.4 在其前代产品基础上进一步改进。
在 τ2-bench 基准测试中,模型必须使用工具来完成一项客服任务,其中可能存在一个模拟用户,该用户可以通信并对世界状态执行操作。推理努力被设置为“无”。
改进的网络搜索
GPT-5.4 在智能体网络搜索方面表现更优。在 BrowseComp(一项衡量 AI 智能体能否持续浏览网页以查找难以定位信息的评测)中,GPT-5.4 相比 GPT-5.2 实现了 17 个百分点的绝对提升,而 GPT-5.4 Pro 则以 89.3% 的成绩创下了新的最优水平。
在实际应用中,这意味着 GPT-5.4 Thinking 在回答需要整合网络上多个来源信息的问题时能力更强。它能够更持久地进行多轮搜索,以识别最相关的来源,尤其擅长处理“大海捞针”式的问题,并将这些信息综合成清晰、推理充分的答案。
在 BrowseComp 评测中,我们使用了一个搜索屏蔽列表,排除了包含评测基准答案的网站,以防止数据污染并确保性能评估的公平性。GPT-5.4 的评测时间晚于 GPT-5.2,因此分数反映了模型、我们的搜索系统以及互联网状态的变化。GPT-5.4 在测试时使用了更长、更新的屏蔽列表。模型使用的是 ChatGPT 搜索工具,该工具与 API 搜索可能存在细微差异。
“GPT-5.4 xhigh 是多步骤工具使用领域的新标杆。Zapier 运行着行业内最严格的工具使用基准测试之一,在数百个高级真实工作流中对模型进行测试。GPT-5.4 完成了此前模型放弃的任务——它是迄今为止最持久的模型。”
—— Wade,Zapier 首席执行官
可操控性
类似于 Codex 在开始工作时会概述其方法,ChatGPT 中的 GPT-5.4 Thinking 现在也会在处理更长、更复杂的查询时,通过一段前言来概述其工作步骤。您还可以在模型回答过程中添加指令或调整其方向。这使得引导模型获得您想要的确切结果变得更加容易,无需重新开始或进行多轮额外对话。该功能现已在 chatgpt.com 和 Android 应用上可用,即将登陆 iOS 应用。
该模型还能在困难任务上思考更长时间,同时保持对对话中先前步骤的更清晰认知。这使其能够处理更长的工作流和更复杂的提示词,同时在整个过程中保持答案的连贯性和相关性。
此视频已加速播放,仅供示意。
安全性
过去几个月里,我们在准备部署 GPT-5.4 的同时,持续改进了随 GPT-5.3-Codex 引入的安全防护措施。与 GPT-5.3-Codex 类似,我们根据《预备框架》将 GPT-5.4 归类为高网络安全能力模型,并按照系统卡中记录的相应保护措施进行部署。这些措施包括扩展的网络安全防护栈,涵盖监控系统、可信访问控制,以及对零数据留存(ZDR)场景下高风险客户的异步拦截机制,同时持续投入建设更广泛的安全生态系统。
由于网络安全能力天然具有双重用途,我们在持续校准策略和分类器的同时,对部署保持审慎态度。对于 ZDR 场景下的特定客户,请求级拦截仍是我们网络安全风险缓解措施的一部分;由于分类器仍在改进中,在我们持续完善这些防护手段的过程中,可能会出现一些误报。这些更新旨在改善安全防护在实际运行中的表现,包括减少不必要的拒绝和过度谨慎的回复,同时保持对滥用的强有力防范。
我们持续推进了关于思维链(CoT)可监控性的安全研究,以更好地理解模型的推理过程并帮助检测潜在的不当行为。作为这项工作的一部分,我们推出了一项新的开源评估——CoT 可控性,用于衡量模型是否能够故意混淆其推理过程以逃避监控。我们发现 GPT-5.4 Thinking 控制其思维链的能力较低,这对安全性而言是一个积极特性,表明该模型缺乏隐藏其推理过程的能力,也意味着思维链监控仍然是一种有效的安全工具。
可用性与定价
GPT-5.4 今日起在 ChatGPT 和 Codex 中逐步推出。在 API 中,GPT-5.4 现已作为 gpt-5.4 提供。GPT-5.4 Pro 也已在 API 中以 gpt-5.4-pro 的形式提供,供需要在最复杂任务上获得极致性能的开发者使用。
在 ChatGPT 中,GPT‑5.4 Thinking 即日起面向 ChatGPT Plus、Team 和 Pro 用户推出,取代 GPT‑5.2 Thinking。GPT‑5.2 Thinking 将在模型选择器的“旧版模型”区域为付费用户保留三个月,之后将于 2026 年 6 月 5 日退役。Enterprise 和 Edu 计划的用户可通过管理员设置启用早期访问。GPT‑5.4 Pro 面向 Pro 和 Enterprise 计划开放。ChatGPT 中 GPT‑5.4 Thinking 的上下文窗口与 GPT‑5.2 Thinking 保持一致。
GPT‑5.4 是我们的首款主流推理模型,它融合了 GPT‑5.3‑codex 的前沿编码能力,并正在 ChatGPT、API 和 Codex 中逐步推出。我们将其命名为 GPT‑5.4 以体现这一跃升,并简化在 Codex 中使用时的模型选择。随着时间的推移,您可以期待我们的 Instant 模型和 Thinking 模型以不同速度演进。
Codex 中的 GPT‑5.4 包含对 100 万 token 上下文窗口的实验性支持。开发者可以通过配置 `model_context_window` 和 `model_auto_compact_token_limit` 进行尝试。超出标准 272K 上下文窗口的请求将按正常速率的两倍计入使用限制。
在 API 中,GPT‑5.4 的每 token 定价高于 GPT‑5.2,以反映其增强的能力,同时其更高的 token 效率有助于减少许多任务所需的 token 总数。Batch 和 Flex 定价按标准 API 费率的一半提供,而 Priority 处理则按标准 API 费率的两倍提供。
API 模型 | 输入价格 | 缓存输入价格**** | 输出价格 gpt-5.2 | $1.75 / M tokens | $0.175 / M tokens | $14 / M tokens gpt-5.4 | $2.50 / M tokens | $0.25 / M tokens | $15 / M tokens gpt-5.2-pro | $21 / M tokens | - | $168 / M tokens gpt-5.4-pro | $30 / M tokens | - | $180 / M tokens
评测
专业
*评测 | GPT‑5.4* | GPT‑5.4
Pro | GPT‑5.3-Codex | GPT‑5.2**** | GPT‑5.2
Pro** | GDPval | 83.0% | 82.0% | 70.9% | 70.9% | 74.1% FinanceAgent v1.1 | 56.0% | 61.5% | 54.0% | 59.5% | — 投资银行建模任务(内部) | 87.3% | 83.6% | 79.3% | 68.4% | 71.7% OfficeQA | 68.1% | — | 65.1% | 63.1% | —
编码
*评测 | GPT‑5.4* | GPT‑5.4
Pro | GPT‑5.3-Codex | GPT‑5.2**** | GPT‑5.2
Pro** | SWE-Bench Pro(公开) | 57.7% | — | 56.8% | 55.6% | — Terminal-Bench 2.0 | 75.1% | — | 77.3% | 62.2% | —
计算机使用与视觉
*评测 | GPT‑5.4* | GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** OSWorld-Verified 75.0%—74.0%47.3%— MMMU Pro(无工具)81.2%——79.5%— MMMU Pro(有工具)82.1%——80.4%—
工具使用
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** BrowseComp 82.7%89.3%77.3%65.8%77.9% MCP Atlas 67.2%——60.6%— Toolathlon 54.6%—51.9%45.7%— Tau2-bench Telecom 98.9%——98.7%—
学术
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** 前沿科学研究 33.0%36.7%—25.2%— FrontierMath 第1–3级 47.6%50.0%—40.7%— FrontierMath 第4级 27.1%38.0%—18.8%31.3% GPQA Diamond 92.8%94.4%92.6%92.4%93.2% 人类最后的考试(无工具)39.8%42.7%—34.5%36.6% 人类最后的考试(有工具)52.1%58.7%—45.5%50.0%
长上下文
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** Graphwalks BFS 0K–128K 93.0%——94.0%— Graphwalks BFS 256K–1M 21.4%———— Graphwalks parents 0–128K(准确率)89.8%——89.0%— Graphwalks parents 256K–1M(准确率)32.4%———— OpenAI MRCR v2 8-needle 4K–8K 97.3%——98.2%— OpenAI MRCR v2 8-needle 8K–16K 91.4%——89.3%— OpenAI MRCR v2 8-needle 16K–32K 97.2%——95.3%— OpenAI MRCR v2 8-needle 32K–64K 90.5%——92.0%— OpenAI MRCR v2 8-needle 64K–128K 86.0%——85.6%— OpenAI MRCR v2 8-needle 128K–256K 79.3%——77.0%— OpenAI MRCR v2 8-needle 256K–512K 57.5%———— OpenAI MRCR v2 8-needle 512K–1M 36.6%————
抽象推理
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** ARC-AGI-1(已验证)93.7%94.5%—86.2%90.5% ARC-AGI-2(已验证)73.3%83.3%—52.9%54.2%(高)
无推理评测
*Eval***GPT‑5.4
(无)****GPT‑5.2
(无)**GPT‑4.1 OmniDocBench(归一化编辑距离)0.109 0.140— Tau2-bench Telecom 64.3%57.2%43.6%
除特别说明外,所有评测均将推理强度设置为 xhigh。基准测试在研究环境中进行,某些情况下其输出可能与生产环境中的 ChatGPT 略有不同。
Today, we’re releasing GPT‑5.4 in ChatGPT (as GPT‑5.4 Thinking), the API, and Codex. It’s our most capable and efficient frontier model for professional work. We’re also releasing GPT‑5.4 Pro in ChatGPT and the API, for people who want maximum performance on complex tasks.
GPT‑5.4 brings together the best of our recent advances in reasoning, coding, and agentic workflows into a single frontier model. It incorporates the industry-leading coding capabilities of GPT‑5.3‑Codex while improving how the model works across tools, software environments, and professional tasks involving spreadsheets, presentations, and documents. The result is a model that gets complex real work done accurately, effectively, and efficiently—delivering what you asked for with less back and forth.
In ChatGPT, GPT‑5.4 Thinking can now provide an upfront plan of its thinking, so you canadjust course mid-responsewhile it’s working**,** and arrive at a final output that’s more closely aligned with what you need without additional turns. GPT‑5.4 Thinking also improves deep web research, particularly for highly specific queries, while better maintaining context for questions that require longer thinking. Together, these improvements mean higher-quality answers that arrive faster and stay relevant to the task at hand.
In Codex and the API, GPT‑5.4 is the first general-purpose model we’ve released with native, state-of-the-art computer-use capabilities, enabling agents to operate computers and carry out complex workflows across applications. It supports up to 1M tokens of context, allowing agents to plan, execute, and verify tasks across long horizons. GPT‑5.4 also improves how models work across large ecosystems of tools and connectors with tool search, helping agents find and use the right tools more efficiently without sacrificing intelligence. Finally, GPT‑5.4 is ourmost token efficient reasoning modelyet, using significantly fewer tokens to solve problems when compared to GPT‑5.2—translating to reduced token usage and faster speeds.
Together with advances in general reasoning, coding, and professional knowledge work, GPT‑5.4 enables more reliable agents, faster developer workflows, and higher-quality outputs across ChatGPT, the API, and Codex.
GPT‑5.4GPT‑5.3‑CodexGPT‑5.2 GDPval (wins or ties)83.0%70.9%70.9% SWE-Bench Pro (Public)57.7%56.8%55.6% OSWorld-Verified 75.0%74.0%*47.3% Toolathlon 54.6%51.9%46.3% BrowseComp 82.7%77.3%65.8%
*Previously reported as 64.7%. GPT‑5.3‑Codex achieves 74.0% with a newly introduced API parameter that preserves the original image resolution.
Knowledge work
Building on GPT‑5.2’s general reasoning capabilities, GPT‑5.4 delivers even more consistent and polished results on real-world tasks that matter to professionals.
On GDPval, which tests agents’ abilities to produce well-specified knowledge work across 44 occupations, GPT‑5.4 achieves a new state of the art, matching or exceeding industry professionals in 83.0% of comparisons, compared to 70.9% for GPT‑5.2.
In GDPval, models attempt well-specified knowledge work spanning 44 occupations from the top 9 industries contributing to U.S. GDP. Tasks request real work products, such as sales presentations, accounting spreadsheets, urgent care schedules, manufacturing diagrams, or short videos. Reasoning effort was set to xhigh for GPT‑5.4 and heavy for GPT‑5.2 (a slightly lower level in ChatGPT).
“GPT-5.4 is the best model we’ve ever tried. It’s now top of the leaderboard on our APEX-Agents benchmark, which measures model performance for professional services work. It excels at creating long-horizon deliverables such as slide decks, financial models, and legal analysis, delivering top performance while running faster and at a lower cost than competitive frontier models.”
— Brendan Foody, CEO at Mercor
We put a particular focus on improving GPT‑5.4’s ability to create and edit spreadsheets, presentations, and documents. On an internal benchmark of spreadsheet modeling tasks that a junior investment banking analyst might do, GPT‑5.4 achieves a mean score of 87.3%, compared to **68.4%**for GPT‑5.2. On a set of presentation evaluation prompts, human raters preferred presentations from GPT‑5.4 68.0% of the time over those from GPT‑5.2 due to stronger aesthetics, greater visual variety, and more effective use of image generation.
Documents were generated with reasoning effort set to xhigh
You can try these capabilities in ChatGPT using GPT‑5.4 Thinking or Pro. If you’re an Enterprise customer, we recommend using our newly released ChatGPT for Excel add-in , which was also launched today. We've also updated our spreadsheet and presentation skills available in Codex and the API.
To make GPT‑5.4 better at real-world work, we continued our progress at driving down hallucinations and errors. GPT‑5.4 is our most factual model yet: on a set of de-identified prompts where users flagged factual errors, GPT‑5.4’s individual claims are 33% less likely to be false and its full responses are 18% less likely to contain any errors, relative to GPT‑5.2.
“GPT-5.4 sets a new bar for document-heavy legal work. On our BigLaw Bench eval, it scored 91%. Compared to other models, GPT-5.4 is currently better at structuring complex transactional analysis, maintaining accuracy across lengthy contracts, and delivering the high level of detail legal practitioners require.”
— Niko Grupen, Head of Applied Research at Harvey
Computer use and vision
GPT‑5.4 is our first general-purpose model with native computer-use capabilitiesand marks a major step forward for developers and agents alike. It’s the best model currently available for developers building agents that complete real tasks across websites and software systems.
We’ve designed GPT‑5.4 to be performant across a wide range of computer-use workloads. It is excellent at writing code to operate computers via libraries like Playwright, as well as issuing mouse and keyboard commands in response to screenshots. Its behavior is steerable via developer messages, meaning that developers can adjust behavior to suit particular use cases. Developers can even configure the model’s safety behavior to suit different levels of risk tolerance by specifying custom confirmation policies.
The model’s performance and flexibility are reflected across benchmarks that test computer use across different settings. On OSWorld-Verified, which measures a model’s ability to navigate a desktop environment through screenshots and keyboard/mouse actions, GPT‑5.4 achieves a state-of-the-art 75.0% success rate, far exceeding GPT‑5.2’s 47.3%, and surpassing human performance at **72.4%.**1
On WebArena-Verified, which tests browser use, GPT‑5.4 achieves a leading 67.3% success rate when using both DOM- and screenshot-driven interaction, compared to GPT‑5.2’s 65.4%. On Online-Mind2Web, which also tests browser use, GPT‑5.4 achieves a 92.8% success rate using screenshot-based observations alone, improving over ChatGPT Atlas’s Agent Mode, which achieves a success rate of 70.9%.
A tool yield is when an assistant yields to await tool responses. If 3 tools are called in parallel, followed by 3 more tools called in parallel, the number of yields would be 2. Tool yields are a better proxy of latency than tool calls because they reflect the benefits of parallelization.
GPT‑5.4 interprets screenshots of a browser interface and interacts with UI elements through coordinate-based clicking to send emails and schedule a calendar event. Video is not sped up.
GPT‑5.4’s improved computer use is built on the model’s improved general visual perception capabilities. On MMMU-Pro, a test of a model’s visual understanding and reasoning, GPT‑5.4 achieves an 81.2% success rate without tool use, an improvement over GPT‑5.2’s 79.5%. Improved visual perception also translates into better document parsing capabilities. On OmniDocBench, GPT‑5.4 without reasoning effort achieves an average error (measured by normalized edit distance between model prediction and ground truth) of 0.109, improved from GPT‑5.2’s 0.140.
MMMUPro was run with reasoning effort set to xhigh. OmniDocBench was run with reasoning effort set to none, to reflect low-cost, low-latency performance.
We’re also improving visual understanding for dense, high-resolution images where full fidelity matters. Starting with GPT‑5.4, we’re introducing an original image input detail level which supports full-fidelity perception up to 10.24M total pixels or 6000-pixel maximum dimension, whichever is lower; the high image input detail level now supports up to 2.56M total pixels or a 2048-pixel maximum dimension. In early testing with API users, we observed strong gains in localization ability, image understanding, and click accuracy when using original or high detail.
“In our evals measuring computer use performance across ~30K HOA and property tax portals, GPT-5.4 achieved a 95% success rate on the first attempt and 100% within three attempts, compared to ~73–79% with prior CUA models. It also completed sessions ~3x faster while using ~70% fewer tokens, materially improving reliability and cost efficiency at scale."
— Dod Fraser, CEO at Mainstay
In the API, developers can access these capabilities using the updated computer tool. Please see our updated documentation for recommended best practices.
Coding
GPT‑5.4 combines the coding strengths of GPT‑5.3‑Codex with leading knowledge work and computer-use capabilities, which matter most on longer-running tasks where the model can use tools, iterate, and push work further with less manual intervention. It matches or outperforms GPT‑5.3‑Codex on SWE-Bench Pro while being lower latency across reasoning efforts.
We estimate latency by looking at the production behavior of our models, and simulating this offline. The latency estimate accounts for tool call duration (code execution time), sampled tokens, and input tokens. Real-world latency may vary substantially, and depends on many factors not captured in our simulation. Reasoning efforts were swept from none to xhigh.
When toggled on, /fast mode in Codex delivers up to 1.5x faster token velocity with GPT‑5.4. It’s the same model and the same intelligence, just faster. That means users can move through coding tasks, iteration, and debugging while staying in flow. Developers can access GPT‑5.4 at the same fast speeds via the API by using priority processing .
In evaluation and internal testing we found that GPT‑5.4 excels at complex frontend tasks, with noticeably more aesthetic and more functional results than any models we’ve launched previously.
As a demonstration of the model’s improved computer-use and coding capabilities working in tandem, we’re also releasing an experimental Codex skill called “Playwright (Interactive) ”. This allows Codex to visually debug web and Electron apps; it can even be used to test an app it’s building, as it’s building it.
Theme park simulation game made with GPT‑5.4 from a single lightly specified prompt, using Playwright Interactive for browser playtesting and image generation for the isometric asset set. The simulation includes tile-based path placement, ride and scenery construction, guest pathfinding, queueing, and ride cycles, while park metrics like money, guest count, happiness, cleanliness, and rating rise or fall based on how the layout performs and how guests respond to it. Playwright was used to automate browser playtests by building and expanding the park, placing and removing paths and attractions, checking camera navigation, and verifying that guests, queues, ride states, and UI metrics updated correctly over several rounds of play.
Prompt:``Use $playwright-interactive and $imagegen. Create an interactive isometric theme park simulation game that I can build and navigate in the browser. Use imagegen to establish the overall visual vision and generate the game’s assets, including rides, paths, terrain, trees, water, food stalls, decorations, buildings, icons, and UI illustrations. The world should feel cohesive, polished, and visually rich, with a premium art direction that works well from an isometric perspective. Let me place and remove paths, add attractions, position scenery, and move around the park smoothly while monitoring guest activity, ride status, and park growth. Include believable guest movement, simple park management systems like money, cleanliness, queueing, and happiness, and make the experience feel playful, clear, and complete rather than like a rough prototype. Prioritize charm, readability, and strong game feel over realism.
When play testing, be sure to build and expand a park through several rounds of play, verify that placement and navigation work smoothly, confirm that guests react to the park layout and attractions, and ensure the visuals, UI, and interactions feel stable and cohesive.
“GPT-5.4 is currently the leader on our internal benchmarks. Our engineers find it to be more natural and assertive than previous models. It works through ambiguous problems without second-guessing itself, and it's proactive about parallelizing work to keep things moving.”
— Lee Robinson, VP of Developer Education at Cursor
Tool use
With GPT‑5.4, we’ve significantly improved how models work with external tools. Agents can now operate across larger tool ecosystems, choose the right tools more reliably, and complete multi-step workflows with lower cost and latency.
Tool search
In the API, GPT‑5.4 introduces tool search , which allows models to work efficiently when given many tools.
Previously, when a model was given tools, all tool definitions were included in the prompt upfront. For systems with many tools, this could add thousands—or even tens of thousands—of tokens to every request, increasing cost, slowing responses, and crowding the context with information the model might never use.
With tool search, GPT‑5.4 instead receives a lightweight list of available tools along with a tool search capability. When the model needs to use a tool, it can look up that tool’s definition and append it to the conversation at that moment.
This approach dramatically reduces the number of tokens required for tool-heavy workflows and preserves the cache, making requests faster and cheaper. It also enables agents to reliably work with much larger tool ecosystems. For MCP servers that may contain tens of thousands of tokens of tool definitions, the efficiency gains can be substantial.
To demonstrate the efficiency gains, we evaluated 250 tasks from Scale’s MCP Atlas benchmark with all 36 MCP servers enabled in two modes: (1) exposing every MCP function directly in the model context, and (2) placing all MCP servers behind tool search. The tool-search configuration reduced total token usage by 47% while achieving the same accuracy.
Example token counts come from averaging 250 tasks in the MCP-Atlas public dataset.
Agentic tool calling
GPT‑5.4 also improves tool calling, making it more accurate and efficient when deciding when and how to use tools during reasoning, particularly in the API. Compared to GPT‑5.2, it achieves higher accuracy in fewer turns on Toolathlon, a benchmark that tests how well AI agents can use real-world tools and APIs to complete multi-step tasks. For example, an agent needs to read emails, extract assignment attachments, upload them, grade them and record results in a spreadsheet.
A tool yield is when an assistant yields to await tool responses. If 3 tools are called in parallel, followed by 3 more tools called in parallel, the number of yields would be 2. Tool yields are a better proxy of latency than tool calls because they reflect the benefits of parallelization.
For latency-sensitive use cases where reasoning effort None is preferred, GPT‑5.4 further improves upon its predecessors.
In_τ2-bench_ , a model must use tools to accomplish a customer service task, where there may be a simulated user who can communicate and take actions on the world state. Reasoning effort was set to None.
Improved web search
GPT‑5.4 is better at agentic web search. On BrowseComp, a measurement of how well AI agents can persistently browse the web to find hard-to-locate information, GPT‑5.4 leaps 17%abs over GPT‑5.2, and GPT‑5.4 Pro sets a new state of the art of 89.3%.
In practice, this means GPT‑5.4 Thinking is stronger at answering questions that require pulling together information from many sources on the web. It can more persistently search across multiple rounds to identify the most relevant sources, particularly for “needle-in-a-haystack” questions, and synthesize them into a clear, well-reasoned answer.
In BrowseComp, we used a search blocklist excluding websites containing benchmark answers from evaluation to prevent contamination and ensure a fair measure of performance. GPT‑5.4 was measured on a later date than GPT‑5.2, so scores reflect changes in the model, our search system, and state of the internet. GPT‑5.4 was tested with a longer, updated blocklist. Models use the ChatGPT search tool, which can have small differences from API search.
“GPT-5.4 xhigh is the new state of the art for multi-step tool use. Zapier runs some of the most rigorous tool use benchmarks in the industry, testing models across hundreds of advanced real-world workflows. GPT-5.4 finished the job where previous models gave up - the most persistent model to date.”
— Wade, CEO at Zapier
Steerability
Similarly to how Codex outlines its approach when it starts working, GPT‑5.4 Thinking in ChatGPT will now outline its work with a preamble for longer, more complex queries. You can also add instructions or adjust its direction mid-response. This makes it easier to guide the model toward the exact outcome you want without starting over or requiring multiple additional turns. This feature is available now on chatgpt.com and the Android app, coming soon to the iOS app.
The model can also think longer on difficult tasks while maintaining stronger awareness of earlier steps in the conversation. This allows it to handle longer workflows and more complex prompts while keeping answers coherent and relevant throughout.
This video was sped up for illustrative purposes.
Safety
Over recent months, we’ve continued improving the safeguards we introduced with GPT‑5.3‑Codex while preparing GPT‑5.4 for deployment. Similar to GPT‑5.3‑Codex, we are treating GPT‑5.4 as High cyber capability under our Preparedness Framework, and we are deploying it with the corresponding protections as documented in the system card . These include an expanded cyber safety stack, including monitoring systems, trusted access controls, and asynchronous blocking for higher-risk requests for customers on Zero Data Retention (ZDR) surfaces, alongside ongoing investment in the broader security ecosystem.
Because cybersecurity capabilities are inherently dual-use, we maintain a precautionary approach to deployment while continuing to calibrate our policies and classifiers. For certain customers on ZDR surfaces, request-level blocking remains part of our cyber risk mitigation stack; because classifiers are still improving, some false positives may occur as we continue refining these safeguards. These updates aim to improve how safeguards operate in practice, including reducing unnecessary refusals and overly caveated responses, while preserving strong protections against misuse.
We’ve continued our safety research on Chain-of-Thought (CoT) monitorability to better understand how models reason and help detect potential misbehavior. As part of this work, we introduce a new open-source evaluation, CoT controllability, measuring whether models can deliberately obfuscate their reasoning to evade monitoring. We find that GPT‑5.4 Thinking’s ability to control its CoT is low, which is a positive property for safety, suggesting that the model lacks the ability to hide its reasoning and that CoT monitoring remains an effective safety tool.
Availability and pricing
GPT‑5.4 is rolling out gradually today across ChatGPT and Codex.In the API, GPT‑5.4 is available now as gpt-5.4. GPT‑5.4 Pro is also available in the API as gpt-5.4-pro for developers who need maximum performance on the most complex tasks.
In ChatGPT, GPT‑5.4 Thinking is available starting today to ChatGPT Plus, Team, and Pro users, replacing GPT‑5.2 Thinking. GPT‑5.2 Thinking will remain available for three months for paid users in the model picker under the Legacy Models section, after which it will be retired on June 5, 2026. Those on Enterprise and Edu plans can enable early access via admin settings. GPT‑5.4 Pro is available to Pro and Enterprise plans. Context windows in ChatGPT for GPT‑5.4 Thinking remain unchanged from GPT‑5.2 Thinking.
GPT‑5.4 is our first mainline reasoning model that incorporates the frontier coding capabilities of GPT‑5.3‑codex and that is rolling out across ChatGPT, the API and Codex. We're calling it GPT‑5.4 to reflect that jump, and to simplify the choice between models when using Codex. Over time, you can expect our Instant models and Thinking models to evolve at different speeds.
GPT‑5.4 in Codex includes experimental support for the 1M context window. Developers can try this by configuring model_context_window and model_auto_compact_token_limit. Requests that exceed the standard 272K context window count against usage limits at 2x the normal rate.
In the API, GPT‑5.4 is priced higher per token than GPT‑5.2 to reflect its improved capabilities, while its greater token efficiency helps reduce the total number of tokens required for many tasks. Batch and Flex pricing are available at half the standard API rate, while Priority processing is available at twice the standard API rate.
API modelInput priceCached input price****Output price gpt-5.2$1.75 / M tokens$0.175 / M tokens$14 / M tokens gpt-5.4$2.50 / M tokens$0.25 / M tokens$15 / M tokens gpt-5.2-pro$21 / M tokens-$168 / M tokens gpt-5.4-pro$30 / M tokens-$180 / M tokens
Evaluations
Professional
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** GDPval 83.0%82.0%70.9%70.9%74.1% FinanceAgent v1.1 56.0%61.5%54.0%59.5%— Investment Banking Modeling Tasks (Internal)87.3%83.6%79.3%68.4%71.7% OfficeQA 68.1%—65.1%63.1%—
Coding
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** SWE-Bench Pro (Public)57.7%—56.8%55.6%— Terminal-Bench 2.0 75.1%—77.3%62.2%—
Computer use and vision
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** OSWorld-Verified 75.0%—74.0%47.3%— MMMU Pro (no tools)81.2%——79.5%— MMMU Pro (with tools)82.1%——80.4%—
Tool use
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** BrowseComp 82.7%89.3%77.3%65.8%77.9% MCP Atlas 67.2%——60.6%— Toolathlon 54.6%—51.9%45.7%— Tau2-bench Telecom 98.9%——98.7%—
Academic
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** Frontier Science Research 33.0%36.7%—25.2%— FrontierMath Tier 1–3 47.6%50.0%—40.7%— FrontierMath Tier 4 27.1%38.0%—18.8%31.3% GPQA Diamond 92.8%94.4%92.6%92.4%93.2% Humanity's Last Exam (no tools)39.8%42.7%—34.5%36.6% Humanity's Last Exam (with tools)52.1%58.7%—45.5%50.0%
Long context
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** Graphwalks BFS 0K–128K 93.0%——94.0%— Graphwalks BFS 256K–1M 21.4%———— Graphwalks parents 0–128K (accuracy)89.8%——89.0%— Graphwalks parents 256K–1M (accuracy)32.4%———— OpenAI MRCR v2 8-needle 4K–8K 97.3%——98.2%— OpenAI MRCR v2 8-needle 8K–16K 91.4%——89.3%— OpenAI MRCR v2 8-needle 16K–32K 97.2%——95.3%— OpenAI MRCR v2 8-needle 32K–64K 90.5%——92.0%— OpenAI MRCR v2 8-needle 64K–128K 86.0%——85.6%— OpenAI MRCR v2 8-needle 128K–256K 79.3%——77.0%— OpenAI MRCR v2 8-needle 256K–512K 57.5%———— OpenAI MRCR v2 8-needle 512K–1M 36.6%————
Abstract reasoning
*EvalGPT‑5.4*GPT‑5.4
ProGPT‑5.3-CodexGPT‑5.2****GPT‑5.2
Pro** ARC-AGI-1 (Verified)93.7%94.5%—86.2%90.5% ARC-AGI-2 (Verified)73.3%83.3%—52.9%54.2% (high)
Evals without reasoning
*Eval***GPT‑5.4
(none)****GPT‑5.2
(none)**GPT‑4.1 OmniDocBench (normalized edit distance)0.109 0.140— Tau2-bench Telecom 64.3%57.2%43.6%
Evals were run with reasoning effort set to xhigh, except where specified otherwise. Benchmarks were conducted in a research environment, which may provide slightly different output from production ChatGPT in some cases.