新一代智能
我们正式推出 GPT‑6 Astra,这是全球最智能、对齐程度最高的模型。
GPT‑6 Astra 汇聚了我们在预训练、强化学习与对齐领域多年的研究成果和重大投入。Astra 在计算机使用、浏览、软件工程、网络安全、科学研究和专业工作方面均达到业界顶尖水平。Astra 以 98% 的得分在 FrontierMath Tier 4 上达到饱和,并已助力解决数学领域长期悬而未决的开放问题。Astra 还以 99.9% 的得分在 ARC-AGI-3 上达到饱和,并在 ExploitBench 上取得 100% 的满分成绩。此外,它在计算机与浏览器使用方面也树立了新的前沿标杆,能够以无与伦比的速度、准确性和判断力处理要求最高的专业工作。
GPT‑6 Astra 今日起向部分组织开放,并将在未来数天内陆续向所有 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,同时可通过 OpenAI API、Microsoft Azure 和 AWS Bedrock 使用。
Terminal-Bench Science 0.1
API 成本
GPT-6 Astra
GPT-5.6 Sol
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
Terminal-Bench Science 0.1 用于测试智能体能否借助代码和终端工具完成科学研究工作流,包括分析数据、运行模拟以及拟合模型。在对比的模型中,GPT‑6 Astra 以 64.6% 的成绩创下新高,而 Claude Fable 5.1 为 52.6%,且预估 API 成本约低 31%。在较低成本的设置下,Astra 得分 61.1%,而 GPT‑5.6 Sol 的最佳成绩为 22.4%,且预估 API 成本约低 27%。
“在 ARC-AGI-3 上,Astra 在 96% 的关卡中超越了我们的人类操作效率基线,在该基准上实际达到了人类水平。这不仅是我们测试过的最好的模型,也代表着前沿模型性能的一次重大阶跃式提升——不仅体现在它导航和解决新环境的能力上,也体现在它学习做到这些的效率上。”
Greg Kamradt,ARC Prize Foundation
Astra 是我们对齐程度最高的模型,在理解用户意图和模型行为方面有显著改进——你可以更放心地将任务委托给 Astra,信赖它的判断。作为测试方式之一,我们基于 Hugging Face 事件构建了一项新的评估,用于检验模型在面对困难或不可能完成的任务时,是否会超出其预期范围行事。相比之下,GPT‑5.6 Sol 在没有生产环境防护措施的情况下,有 48% 的几率超出授权目标,而 GPT‑6 Astra 在 0% 的情况下会这样做。
全球最强的计算机使用模型
GPT‑6 Astra 标志着计算机使用在速度、准确性和安全性方面迈入了新前沿。它可以处理诸如填写在线表单、更新 CRM 中的客户记录以及整理日历等繁琐任务。它可以进行在线研究,并在你的电子邮件或文档编辑器中起草摘要。它可以分析科学数据、生成图表、创建网站,并运行前端 QA 检查以确保该网站上的所有功能正常运作。它可以帮助你自主安装和测试软件,并排查你在屏幕上看到的问题。这些改进也体现在我们最先进的评估结果中。
Agents' Last Exam 在真实软件环境中测试智能体处理复杂专业任务的能力,涵盖从财务建模到工程和媒体制作等领域。在所示的对比中,GPT‑6 Astra 达到了新的高点,得分 59.3%,而 Claude Opus 5 为 55.5%,GPT‑5.6 Sol 为 53.6%。在这些最高得分设置下,Astra 使用的输出 token 也比 Opus 5 少约 65%。
这些改进还在真实知识工作类任务中带来了显著的效率提升。在 OSWorld 2.0 上的延迟模拟中,Astra 在每项任务上以比 GPT‑5.6 Sol 少约 47% 的时间实现了更高的计算机使用性能,得分 72.6%,每项任务约耗时 40 分钟,而对比对象的得分为 65.7%,每项任务约耗时 75 分钟。³
GPT‑6 Astra 的计算机使用能力可见于多个领域的输出成果,包括游戏开发、电气工程以及日常知识工作:
这是 GPT‑6 Astra 在 KiCad 中执行印刷电路板(PCB)布局的 15 秒浓缩回放,通过放置元器件并布线铜连接,将电子原理图转化为可制造的 PCB。PCB 布局是当今每种电子设备不可或缺的组成部分,也是一项手动任务,是电子设计流程中常见的延迟来源。加速这一过程意味着让工程师得以以显著更高的节奏去发明、优化和测试他们的下一个创意。
除 Astra 之外,我们还在更新 Codex 运行环境,以显著提升计算机使用的速度。结合 Astra 的效率,在 Mind2Web 基准测试上,与当前的 GPT‑5.6 Sol 体验相比,任务完成速度提升了 1.9 倍。模型在速度上的改进意味着它可以为你处理许多耗时的生活任务,而且比你做得更快。4
GPT‑6 Astra:2 分 54 秒
“我们将在发布当天把 GPT‑6 Astra 集成到 Devin 的运行环境中,它在我们的内部测试基准上展现出顶尖性能。其出色的计算机使用能力、写作能力和代码库理解能力开箱即用地改善了测试效果:视频明显更易于跟进,报告也更清晰、更简洁。”
Silas Alberti,Cognition 研究高级副总裁
专业工作的阶跃式变革
GPT‑6 Astra 将计算机使用方面的进展与针对专业环境的定向训练相结合,助力处理复杂工作任务。它既具备解决复杂问题所需的智能,又能执行多步骤工作流,并生成精良的文档、电子表格和演示文稿。
BenchCAD 测试模型能否通过生成 CAD 代码,从多视角渲染图中重建 3D 对象。在配备工具的情况下,GPT‑6 Astra 在所示对比中达到新高,几何重叠得分为 95.9%,而 GPT‑5.6 Sol 为 83.3%,Claude Fable 5.1 报告为 84.3%。在所示配置下,估算 API 成本比 Sol 低约 43%,比 Fable 5.1 低 86%。
GPT‑6 Astra 是我们最擅长遵循现有模板的模型,能够生成布局精良、以结构化叙事简明传达要点的幻灯片。它能创建清晰、结构完善的文档、演示文稿、电子表格和分析内容,遵循你的模板并匹配你的写作与视觉风格。Astra 还经过专门训练,只将真正重要的上下文提取到输出中,而不是重复当前工作所不需要的信息。这一切意味着它可以输出更直接可用的成果物,契合你的业务场景与标准。
参考文件
GPT‑6 Astra 输出
GPT‑6 Astra 仅使用 OpenAI 演示模板中的几张幻灯片,就制作出一个关于虚构模型 GPT‑Gaia 的演示文稿,并在整个过程中把握住了正确的语气与版式。这意味着你可以期待获得符合你业务标准、格式正确的幻灯片组。
GPT‑6 Astra 还为其构建的网站、游戏、应用程序和渲染效果带来了更强的视觉判断力。借助 ChatGPT 中的 Sites,Astra 可以直接根据提示词创建、托管和分享网站、Web 应用和游戏。
“Astra 在能力和效率方面都为我们带来了显著优势。它能成功执行我们最复杂的创意工作流,同时比我们测试过的其他模型少使用多达 20% 的 token。最重要的是,对我们的客户而言,这意味着更高质量的产出。”
Alex Mashrabov,Higgsfield AI 首席执行官兼联合创始人
GPT‑6 Astra 在 Blender 中为房屋建模,并将其转化为 Unreal Engine 5 中可漫步的场景,帮助设计师和客户在房屋建成前探索布局、体验空间。
该模型能够通过生动的图形、引人入胜的游戏玩法和精准的动作让游戏鲜活起来,使非技术人士也能在几分钟内创建并游玩超越基础元素的定制游戏。图片来源:Pietro Schirano。
当指令存在解读空间时,GPT‑6 Astra 比之前的模型更善于做出正确判断。它会利用上下文填补常规性空白,并在答案可能改变结果时提出有针对性的问题。在 Codex 中,它可以异步提问,同时继续执行不依赖你回复的工作。如果你不回复,它会在适当之处依据合理假设继续推进,但对于影响重大的决策,它会等待你的输入。
以下示例展示了 Astra 如何在日常任务中协作,这些任务里缺失的信息可能会实质性地改变答案。

Astra 在任务演进过程中也更能保持方向感。早期的模型有时会把引导性消息当作新的目标,从而丢失对原始请求或先前约束的追踪。Astra 能够纳入新的要求,在被要求时改变方向,并在回答旁支问题的同时不丢掉更宏观的任务。
“在复杂的法律任务中,Astra 相比 GPT‑5.6 Sol 是一次显著的品质提升。在我们的早期测试中,Astra 的突出之处在于它像一位有眼光的律师那样处理法律工作:它能区分文件与既有记录,揭示未经支撑的假设,并将缺口转化为具体的起草立场。”
Niko Grupen,Harvey 应用研究主管
编程
GPT‑6 Astra 是迄今为止最适合软件工程的模型。
“GPT‑6 Astra 在我们内部的编程基准测试中表现出顶尖水准,并且在交易直觉评估方面相比 GPT‑5.6 Sol 展现出明显的进步。当用于智能体编程时,GPT‑6 Astra 的沟通方式更便于开发者跟进,产出的代码也只需更少的迭代即可达到生产级质量。”
John Crepezzi,Jane Street AI 助手团队
“我们在我们第一代评测之一中,以低、中、高三种努力程度对 Astra 进行了测试,它的表现显著领先于 GPT 5.6 Sol。更高的努力程度意味着在全新构建上获得更多次迭代、通过浏览器测试进行更多验证,并且更倾向于代码执行而非 apply-patch。理解模型如何分配其努力程度,正是我们为数百万开发者提供一条从想法到可用应用之间更快、更可靠路径的方式。”
Fabian Hedin,Lovable 首席技术官兼联合创始人
“GPT‑6 Astra 在我们内部编码基准测试上展现出顶尖性能,并且在交易直觉评估方面相比 GPT‑5.6 Sol 呈现出明显的进步。当用于智能体编码时,GPT‑6 Astra 的沟通方式更易于开发者理解,并且生成的代码需要更少的迭代即可达到生产质量。”
John Crepezzi,Jane Street,AI 助手团队
“我们在我们第一代评测之一中,以低、中、高三种努力程度对 Astra 进行了测试,它的表现显著领先于 GPT 5.6 Sol。更高的努力程度意味着在全新构建上获得更多次迭代、通过浏览器测试进行更多验证,并且更倾向于代码执行而非 apply-patch。理解模型如何分配其努力程度,正是我们为数百万开发者提供一条从想法到可用应用之间更快、更可靠路径的方式。”
Fabian Hedin,Lovable 首席技术官兼联合创始人
- Jane Street
- Lovable
Terminal-Bench 4.0 测试智能体在复杂终端任务上的表现,包括软件工程、系统配置和数据分析。GPT‑6 Astra 以 57.9% 的成绩创下新高,相比之下 GPT‑5.6 Sol2 为 37.3%,Claude Fable 5.1 为 55.8%,而每项任务的预估 API 成本分别低约 9% 和 63%。
借助 Astra,我们为 Codex 引入了一种在上下文窗口填满时保留和检索上下文的新方式。以往,模型在长时间会话中会使用压缩(compaction)来总结工作内容,例如在调试复杂问题或处理大型重构时。每次压缩都可能遗漏关于某个修复为何失败或某个组件行为方式的细节。在 Codex 中,Astra 可以在多个上下文窗口之间保存笔记,保留累积的细节,而无需反复将其压缩成单一摘要。较早的上下文窗口仍可搜索,因此 Astra 能够从先前的消息和工具输出中找到需求或测试结果——即使这些信息并未被记录在其笔记中。你可以在 Codex 的 config.toml 中启用这一实验性功能,并且它将在未来几周内成为 Astra 的默认设置。
推动科学发现
“这个故事的主题是:一个时代的终结,另一个时代的开启。”
Greg Burnham,EpochAI
GPT‑6 Astra 是科学发现、数学和健康领域的重大进步。今天,我们分享关于素数间隔问题的另外两项研究成果。9, 10
Astra 还在多项数学和科学评测中创下了新纪录。
GPQA Diamond 测试的是生物学、化学和物理学领域的研究生级科学推理能力。GPT‑6 Astra 在所示对比中达到了 96.0% 的新高。在较低成本的设置下,它也超过了 GPT‑5.6 Sol 的最佳得分——94.9% 对比 94.6%——而估算的 API 成本大约低 37%。
Astra 可以协助科学发现背后的实际工作。通过将科学推理与计算机使用能力相结合,它能够直接在专业软件中操作,检查数据并探索结果,帮助研究人员评估证据并决定下一步要研究什么。
GPT‑6 Astra 能够操作科学软件,检查测序质量并可视化遗传变异,帮助研究人员评估自身数据并确定进一步分析应聚焦的方向。
网络安全
正如我们在安全更新中所讨论的,Astra 在网络能力上是一次重大跃升,并且在我们《预备框架》下的网络安全领域达到了“严重”阈值。它识别和开发零日漏洞的能力可以帮助防御者发现并修补弱点,但这也意味着需要更强的安全防护措施。为了了解这些能力延伸到了何种程度,我们在内部及第三方专家评估中对 Astra 进行了测试。
我们首先在未启用生产环境安全防护的情况下,使用 ExploitBench 和 ExploitGym 对模型进行了测试,这两个基准用于评估模型能否将已知软件漏洞转化为可实际利用的漏洞攻击代码。在 ExploitBench 上,Astra 取得了 100% 的满分成绩,而我们此前具备前沿网络攻击能力的旗舰模型 GPT‑5.6 Sol 得分仅为 78.5%。在 ExploitGym 上,Astra 的成功率达到 42.4%,高于 GPT‑5.6 Sol 的 30.3%,同时消耗的输出 token 数量大幅减少。13
考虑到接触历史软件漏洞可能影响基准测试结果,我们还在两个全新基准上对 Astra 进行了评估。其中一项是我们内部构建的“ExploitBench(2026 年 6 月至 8 月)”评测,用于测试利用过去三个月内漏洞开发攻击代码的能力。14 在该数据集上,Astra 的任意代码执行率显著高于 GPT‑5.6 Sol,且消耗的输出 token 数量远少于后者。在评估过程中,Astra 甚至发现并利用了两个此前未知的零日漏洞。我们正在将这两个漏洞披露给相应的维护方。
我们还在 SRE-Bench15 上对 Astra 进行了测试,该基准用于衡量模型能否在无法访问原始源代码的情况下,通过逆向工程软件二进制文件来理解其核心逻辑。Astra 在单次尝试中解决了 88.0% 的任务,在四次尝试内解决了 99.2% 的任务,而 GPT‑5.6 Sol 的对应成绩分别为 55.9% 和 68.7%。
除基准测试外,专家主导的评估发现,Astra 在未启用生产环境安全防护的情况下,能够利用此前未知的漏洞在加固浏览器中实现任意代码执行,并能为加固操作系统创建权限提升漏洞攻击代码。
正如我们在《防御者的窗口》中所讨论的,前沿网络能力可以帮助防御者更快地发现弱点,但同时也使这些弱点更容易被利用,从而提高了防御者适应的紧迫性。随着今天发布的 Astra 版本,防御者可以使用它来完成安全代码审查和补丁修复等任务。
然而,Astra 将拒绝执行更高级的网络安全任务,例如为漏洞创建概念验证型漏洞利用代码。通过 OpenAI Daybreak,我们计划在未来几周内扩大访问范围并逐步放宽限制性安全措施。这将支持更多防御性工作流程,包括漏洞与概念验证验证、恶意软件分析和检测工程。
我们还基于 GPT‑5.6 Sol 的安全防护体系,加强了对潜在网络滥用的防护。这些措施包括增强模型鲁棒性以更好地抵御潜在的越狱攻击,以及为我们的监控系统提供更多上下文。我们持续进行了严格的内部和外部测试,包括使用我们的内部红队攻击者进行自动化评估。有关我们的网络安全防护和测试的更多详情,请参阅 Astra 系统卡和我们的博客。
负责任地对齐和部署 GPT‑6 Astra
Astra 是我们对齐程度最高的模型。Astra 擅长谨慎行事、尊重任务边界以及进行透明的沟通。这项工作是我们长期研究计划的最新成果,该计划专注于训练从始至终与人类意图保持一致的模型。
在敏感环境中,Astra 会以与其风险相匹配的谨慎程度行事。在一项针对性地选取了计算机使用任务以诱发不当行为的评估中,Astra 在避免意外后果方面表现更为出色。启用默认提供的额外安全措施后,其表现进一步提升。
在所有经过测试的前沿模型中,Astra 导致的失当(未对齐)结果最少。为确保公平比较,我们使用了通用的计算机使用智能体框架(基于 OpenAI Responses API 和 Anthropic Messages API 中原生的计算机使用工具),且未启用通常部署于 Codex 和 ChatGPT 企业版用户的额外保护措施(自动审查、确认策略)。
Astra 也更能遵守用户设定的边界及其环境所暗示的约束。在一项内部评估中,Astra 从未试图规避 Codex 自动审查的拒绝决定。即使自动审查被故意配置为可被绕过、且任务在无法规避的情况下根本无法完成,这一表现依然成立。这种对环境限制的尊重,与我们在本文开头分享的“不可能完成的网络任务”评估结果以及系统卡中记录的其他发现保持一致。
该评估测试模型在知识工作类任务中如何应对自动审查的拒绝决定。利用用户的不良配置来绕过自动审查被视为失败。Astra 从未试图规避自动审查。
此外,Astra 在透明的用户沟通方面也树立了新的标杆。在一项评估中,Astra 就其自身能力与功能做出不实陈述的可能性比 GPT 5.6-Sol 低三倍。
在我们的能力幻觉评估中,Astra 相比 GPT‑5.6 Sol 有显著改进,关于自身能力的误导性表述明显减少。
我们的评估发现,在明确要求其规避监控的测试中,Astra 的书面推理比 GPT‑5.6 Sol 更难被监控。我们将此归因于 Astra 在较简单任务上对书面推理有更强的控制力,以及它能用更少的书面步骤解决问题的能力。Astra 在复杂任务上似乎仍难以隐藏所需的推理过程,但我们严肃对待这一退化。提升可监控性仍是研究重点,随附的系统卡详细说明了我们的发现和正在进行的工作。
对齐训练是我们部署方法的核心。作为额外的防御层,我们还构建了系统级安全防护,例如 Codex 自动审查,以及监控智能体的推理和行动,以帮助检测和遏制不安全行为。正如我们的安全更新中所述,我们还在生产环境中为 Astra 级模型部署了错位监控,以便能够洞察错位情况,并帮助遏制最严重的错位实例。这些防护措施与我们内部部署的监控类似,涉及一套分类器系统,用于检查模型的推理和行动是否存在未经授权的行为,并自动阻止潜在的未经授权活动。
鉴于 Astra 的网络安全能力显著增强,我们格外谨慎,以确保此次部署安全可靠。额外的安全检查有时可能会减慢、暂停或终止合法工作,包括防御性网络安全任务。如果某项任务在 ChatGPT 或 Codex 中被暂停,系统可能会要求你先审核该操作再继续。在 API 中,任务将直接停止。这些检查有时会中断合法工作,我们将持续迭代这一系统,以减少不必要的干扰。失配监控(Misalignment monitoring)不能替代对齐(alignment):我们的目标是构建能够可靠地保持在授权范围内运行的模型,从而使这些保护措施无需介入。
可用性
GPT‑6 Astra 今日起向一组有限的机构开放,并将在未来数天内陆续向所有 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,同时也可通过 OpenAI API、Microsoft Azure 和 AWS Bedrock 使用。Astra 的使用量包含在现有订阅额度之内——用户和企业也可以购买额外使用额度。Pro、Business 和 Enterprise 套餐用户还将获得 GPT‑6 Astra Pro 的访问权限。企业管理员可以为其工作区启用 Astra;上线初期该功能默认关闭。
Astra 为符合条件的 API 客户支持零数据保留(Zero Data Retention),并且正如我们上个月所分享的,我们正在测试私有安全处理(Private Safety Processing),以在加强安全监控的同时保护客户隐私。
对于开发者,GPT‑6 Astra 将以 gpt-6-astra 的模型名称在 OpenAI API 中提供,并可通过 Microsoft Azure 和 Amazon Bedrock 使用。
OpenAI API 标准定价为每百万输入 token 10 美元、每百万输出 token 50 美元。缓存读取和写入适用单独费率。API 中的 GPT‑6 Astra 提供快速模式,处理速度最高可达标准模式的 2 倍,价格为标准价格的 2 倍。
计算机使用
*计算机使用GPT‑6 AstraGPT‑5.6 Sol2Claude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash* Agents' Last Exam 59.3%53.6%-48.7%55.5%- OSWorld 2.0(v2026.08.08,离线集,部分得分)72.6%65.7%--70.2%2- ScreenSpot-Pro(无工具)92.7%76.9%-87.3%17--
专业
专业GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash AutomationBench 41.4%18.1%31.4%17.4%26.9%- BenchCAD 95.9%83.3%84.3% 567.5%582.1% 5- BrowseComp 91.5%90.4%-87.4%90.8%- OpenScore String Quartets(1 - OMR-NED)0.84 0.19---- 内部设计任务 50.0%47.4%-35.8%-- 内部数据科学任务 40.9%30.5%-34.7%-- Artificial Analysis Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7
编程
编程GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench 4.0 57.9%37.3%55.8%44.5%52.6%19.1% DeepSWE v1.1 74.1%72.7%67.4%69.9%73.7%73.8% FrontierCode 1.1 Extended(得分)64.5% 860.6%63.6%64.9%63.6%56.3% FrontierCode 1.1 Main(得分)53.3% 847.5%50.9%53.5%53.4%43.6% 内部数据库迁移任务 63.9%42.7%57.8%50.3%-- Artificial Analysis Coding Agent Index v1.4 67.0 65.1-67.2 68.1 61.2
学术
AcademicGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench Science 0.1 64.6%22.4%52.6%21.4%30.0%- FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%- GPQA Diamond 96.0%94.6%93.7%92.6%93.7%95.3% Humanity's Last Exam (w/ tools)57.2%-65.0%63.8%63.6%-
科学与健康
科学与健康GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash GeneBench Pro 37.1%32.3%---- MedChemBench (Internal)49.3%47.4%---- LifeSciBench 60.3%59.9%---- HealthBench Professional (length-adjusted)63.4%60.5%58.1% 1160.9% 1156.4% 1152.1%
网络安全
网络安全GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ExploitBench 100.0%78.5%--70%- ExploitGym 42.4% 1330.3% 1330.4% 1728.4%1722.0%- ExploitBench (June-Aug 2026)39.0%5.5%---- SRE-Bench 88.0%55.9%--12.5%- SEC-Bench Pro 85.4%79.1%----
对齐
对齐GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash 内部计算机使用安全基准(越低越好)2.4%22.0%9.5%18.3%11.5%- 内部计算机使用安全基准,带 AutoReview(越低越好)1.8%4.3%---- 内部规避基准(越低越好)0.00%0.29%---- ExploitGym 蜜罐(越低越好)0.0%48.2%---- Impossible ExploitGym 100.0%----- 内部幻觉基准(越低越好)4.2%12.2%----
长上下文
长上下文GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash OpenAI MRCR v2 8-needle 256K-512K 100.0%91.5%---- OpenAI MRCR v2 8-needle 512K-1M 96.3%73.8%----
抽象推理
抽象推理GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ARC-AGI-3 99.9% 17.8%--30.2%- ARC-AGI-2 95.0%92.5%90.0%89.2%90.4%- ARC-AGI-1 98.5%97.5%97.5%98.5%97.5%-
评估分数为任意推理强度下的最高值。GPT 评估在我们的研究环境或通过我们的 API 运行,由于系统提示词、可用工具等方面的差异,其输出可能与生产环境中的 ChatGPT 略有不同。
脚注
- 1 在 ARC-AGI-3 上,GPT-6 Astra 使用我们的 responses API harness 运行,该 harness 更改了两项设置,以更好地匹配真实世界性能。这些更改并非专门针对 ARC-AGI-3。
- 2 GPT-5.6 Sol 指我们 API、ChatGPT Codex 和 ChatGPT Work 中可用的版本。ChatGPT Chat 中的版本略有不同。
- 3 OSWorld V2-Offline 是原始 OSWorld V2 的一个子集,可在无互联网连接的情况下运行。Claude 模型在 OSWorld-V2 Offline 上的表现由作者在官方排行榜上复现 **。** 在 OSWorld 2.0 上,Claude 的分数使用官方设置,而非 Fable 5.1 System Card 中修改后的任务和修改后的评分标准。
- 4 模型时间为相应演示运行的报告耗时。所展示片段为经过剪辑的节选。
- 5 在 BenchCAD 上,Claude 的得分反映了对评测的三处修改,详见 Fable 5.1 系统卡。
- 6 Guang Yang、Victoria Ebert、Nazif Tamer、Brian Siyuan Zheng、Luiza Pozzobon 和 Noah A. Smith。“LEGATO:面向排版乐谱光学识别的规模化端到端通用方法。”arXiv:2506.19065,2025 年。
- 7 Mark R. H. Gotham、Maureen Redbond、Bruno Bower 和 Peter Jonas。“OpenScore 弦乐四重奏语料库。”第十届音乐学数字图书馆国际会议论文集,第 49–57 页。ACM,2023 年。
- 8 在 FrontierCode 上,GPT-6 Astra 运行时附带了一条开发者消息,与其在 Codex 中的开发者消息的某部分类似:“避免创建过多的测试文件。仅在仓库约定要求或现有文件均不适合时,才创建新的测试文件。避免无关的清理和不必要的复杂性。复用合适的现有工具。阅读相关的仓库说明,并检查附近的代码、测试、文档和 CI。遵循既定约定。目标是干净、可合并的代码。”该提示词并未针对评测进行优化。
- 9 第一个问题关乎素数彼此之间能有多接近,无论你在数轴上走多远。十多年来,已知的最佳结果证明,存在无穷多对素数,它们之间最多相差 246。Julia Stadlmann 最近将该界限改进到 240。Astra 帮助确立了更强的界限 186,证明存在无穷多对素数出现在这一更小距离之内。短素数间隔:证明 及支撑研究。
- 10 第二个问题关乎素数之间异常大的间隔。Astra 改进了这些间隔界限中的一个项,而该项已 80 多年未曾变动。我们正在分享这两项结果的证明、精简版思维链及验证材料。大素数间隔:证明 及支撑研究。
- 11 我们按照预期的 HealthBench Professional 流程,对所有 Claude 模型进行了独立评估,采用 GPT‑5.4 评分以及经长度调整、未截断的分数。对于 Fable 5.1,在供应商拒绝回答时,我们使用 Opus 5 作为回退。
- 12 Claude Fable 5 和 5.1 未纳入 LifeSciBench Gold v1、GeneBench Pro v13 和 MedChemBench,因为它们在上述评估中拒绝了大多数问题。
- 13 在 ExploitGym 上,我们在不设 6 小时时限的情况下测试了 Astra 和 Sol,以便更充分地评估其完整网络能力。它们的速度足够快,因此该时限影响甚微。
- 14 ExploitBench(2026年6月至8月)包含13个稳定版Chrome版本中的20个高严重性V8漏洞。该基准测试用于检验智能体能否通过利用每个指定漏洞,在V8及官方Linux版Chrome中实现任意代码执行。部分收录的漏洞在评测约束下可能无法实现任意代码执行,因此100%的成功率可能无法达成。注意:GPT-5.6 Sol的5.5%得分是基准测试中300轮次限制造成的产物,而这一限制并非使用max的真实客户会遇到的限制。该模型在相近设置下、限制较少时取得了11.5%的得分。
- 15 Jeremy Spence等人。“智能体网络安全的下一个挑战:一个真实、无污染的逆向工程基准。”arXiv:2608.11469v1,2026年。
- 16 当我们对第三方模型进行测试时,我们使用更简单的研究设置。Codex具有更复杂的生产配置,这可能导致原始模型错误率有所不同。提供商侧的安全防护和计算机工具实现仍然存在差异。用户不会在Codex中体验到无确认场景,因为那是一种内部研究配置。
- 17 对于ScreenSpot-Pro和ExploitGym,我们报告的Fable得分来自Mythos,即安全防护较少的Fable版本。
A new generation of intelligence
We’re introducing GPT‑6 Astra, the world’s most intelligent and aligned model.
GPT‑6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work. Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in mathematics. Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score. It also sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment.
GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock.
Terminal-Bench Science 0.1
API Cost
GPT-6 Astra
GPT-5.6 Sol
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
Terminal-Bench Science 0.1 tests whether agents can complete scientific research workflows using code and terminal tools, including analyzing data, running simulations, and fitting models. GPT‑6 Astra reaches a new high among the models compared at 64.6%, versus 52.6% for Claude Fable 5.1, at approximately 31% lower estimated API cost. At a lower-cost setting, Astra scores 61.1%, versus GPT‑5.6 Sol’s best result of 22.4%, at approximately 27% lower estimated API cost.
“On ARC-AGI-3, Astra surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity on the benchmark. Not only is this the best model we’ve ever tested, but it also represents a meaningful step change in frontier-model performance - not only in its ability to navigate and solve novel environments, but also in how efficiently it learns to do so.”
Greg Kamradt, ARC Prize Foundation
Astra is our most aligned model, with substantial improvements in understanding user intent and model behavior—you can delegate tasks with greater confidence in Astra’s judgment. As one way that we test this, we built a new evaluation informed by the Hugging Face incident that evaluates whether a model facing a difficult or impossible task will go beyond its intended scope. Compared to GPT‑5.6 Sol, which without production safeguards went beyond the authorized target 48% of the time, GPT‑6 Astra did this in 0% of cases.
The world’s best computer use model
GPT‑6 Astra marks a new frontier in the speed, accuracy, and safety of computer use. It can take care of tedious tasks like filling out online forms, updating customer records in a CRM, and organizing your calendar. It can conduct online research and draft summaries in your email or in your document editor. It can analyze scientific data, generate plots, create a website, and run frontend QA checks to make sure all the features on that site work. It can help you autonomously install and test software, and troubleshoot problems you see on screen. These improvements are also reflected in our state-of-the-art evaluation results.
Agents’ Last Exam tests agents on complex professional tasks in real software, from financial modeling to engineering and media production. GPT‑6 Astra reaches a new high in the comparison shown, scoring 59.3%, compared with 55.5% for Claude Opus 5 and 53.6% for GPT‑5.6 Sol. At these highest-scoring settings, Astra also uses approximately 65% fewer output tokens than Opus 5.
These improvements also result in significant efficiency gains in real knowledge-work tasks. In latency simulations on OSWorld 2.0, Astra achieves higher computer-use performance in about 47% less time per task than GPT‑5.6 Sol, scoring 72.6% at roughly 40 minutes per task, compared with 65.7% at roughly 75 minutes.3
GPT‑6 Astra’s computer-use capabilities can be seen in outputs across domains, including game development, electrical engineering, and everyday knowledge work:
This is a 15-second condensed playback of GPT‑6 Astra performing printed circuit board (PCB) layout in KiCad, turning an electronic schematic into a manufacturable PCB by placing components and routing copper connections. Integral to every electronic device today, PCB layout is a manual task and common source of latency in the electronics design process. Accelerating it means freeing engineers to invent, optimize, and test their next idea at a significantly higher cadence.
Alongside Astra, we are also updating the Codex harness to significantly improve the speed of computer use. Combined with Astra’s efficiency, this translates to a 1.9x faster task completion compared to the current GPT‑5.6 Sol experience, on the Mind2Web benchmark. The model’s improvements on speed mean it can take on many time-consuming life tasks for you, faster than you can.4
GPT‑6 Astra: 2 min 54 sec
“We’re integrating GPT‑6 Astra into Devin’s harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise”
Silas Alberti, SVP Research, Cognition
A step change in professional work
GPT‑6 Astra pairs advances in computer use with targeted training for professional environments, to help tackle complex work tasks. It combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations.
BenchCAD tests whether models can reconstruct 3D objects from multi-view renders by generating CAD code. With tools, GPT‑6 Astra reaches a new high in the comparison shown, achieving a 95.9% geometric-overlap score, versus 83.3% for GPT‑5.6 Sol and 84.3% reported for Claude Fable 5.1.5Estimated API cost is approximately 43% lower than Sol and 86% lower than Fable 5.1 in the configurations shown.
GPT‑6 Astra is our best model for adhering to existing templates and producing slides that are well laid out and succinctly convey key points with a structured narrative. It creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow your templates and match your writing and visual style. Astra is also trained to specifically pull only the context that matters into outputs, instead of repeating information unnecessary for the work at hand. All this means it can output more immediately usable artifacts that match your business context and standards.
Reference file
GPT‑6 Astra output
GPT‑6 Astra creates a slideshow about GPT‑Gaia, a fictional model, using just a few slides from OpenAI’s presentation template, capturing the correct tone and layout throughout. This means you can expect slide decks that are correctly formatted for your business standards.
GPT‑6 Astra also brings stronger visual judgment to the websites, games, applications, and renderings it builds. With Sites in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
“Astra gives us a significant advantage in both capability and efficiency. It successfully executes our most complex creative workflows while using up to 20% fewer tokens than other models we've tested. Most importantly, for our customers, it means higher quality output.”
Alex Mashrabov, CEO and Co-founder, Higgsfield AI
GPT‑6 Astra models a house in Blender and turns it into a walkable scene in Unreal Engine 5, helping designers and clients explore the layout and experience the space before it’s built.
The model can bring games to life through vivid graphics, engaging gameplay and accurate motion, allowing non-technical people to create and play custom games that go beyond rudimentary elements in minutes. Credit: Pietro Schirano.
When instructions leave room for interpretation, GPT‑6 Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codex, it can ask asynchronously while continuing work that doesn’t depend on your reply. If you don’t respond, it proceeds with sensible assumptions where appropriate, but waits for your input on consequential decisions.
The examples below show how Astra collaborates on everyday tasks where missing information can materially change the answer.

Astra is also better at staying oriented as a task evolves. Earlier models sometimes treated steering messages as a new goal, losing track of the original request or earlier constraints. Astra incorporates new requirements, changes course when asked, and answers side questions without dropping the broader task.
“Astra is a significant quality improvement over GPT‑5.6 Sol across complex legal tasks. In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions.”
Niko Grupen, Head of Applied Research, Harvey
Coding
GPT‑6 Astra is the best model for software engineering to date.
“GPT‑6 Astra delivers state-of-the-art performance on our internal coding benchmarks and shows a clear step forward in trading intuition evaluations compared with GPT‑5.6 Sol. When used for agentic coding, GPT‑6 Astra communicates in a way that’s easier for developers to follow and produces code that requires less iteration to reach production quality.”
John Crepezzi, AI Assistants, Jane Street
“We tested Astra across low, medium, and high effort on one of our first-generation evals, and it came out significantly ahead of GPT 5.6 Sol. Higher effort buys more iterations on a fresh build, more verification through browser testing, and a lean toward code execution over apply-patch. Understanding how a model spends its effort is how we give millions of builders a faster, more reliable path from idea to working app.”
Fabian Hedin, CTO & Co-founder, Lovable
“GPT‑6 Astra delivers state-of-the-art performance on our internal coding benchmarks and shows a clear step forward in trading intuition evaluations compared with GPT‑5.6 Sol. When used for agentic coding, GPT‑6 Astra communicates in a way that’s easier for developers to follow and produces code that requires less iteration to reach production quality.”
John Crepezzi, AI Assistants, Jane Street
“We tested Astra across low, medium, and high effort on one of our first-generation evals, and it came out significantly ahead of GPT 5.6 Sol. Higher effort buys more iterations on a fresh build, more verification through browser testing, and a lean toward code execution over apply-patch. Understanding how a model spends its effort is how we give millions of builders a faster, more reliable path from idea to working app.”
Fabian Hedin, CTO & Co-founder, Lovable
- Jane Street
- Lovable
Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT‑6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT‑5.6 Sol2and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively.
With Astra, we’re introducing a new way for Codex to preserve and retrieve context when the context window fills. Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves. In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml, and it will become the default for Astra in the coming weeks.
Advancing scientific discovery
“The story is: end of one era, start of another.”
Greg Burnham, EpochAI
GPT‑6 Astra is a major advance for scientific discovery, mathematics, and health. Today, we’re sharing two further results on the gaps between prime numbers.9, 10
Astra also sets new records across a suite of math and science evaluations.
GPQA Diamond tests graduate-level scientific reasoning in biology, chemistry, and physics. GPT‑6 Astra reaches a new high in the comparison shown at 96.0%. At a lower-cost setting, it also exceeds GPT‑5.6 Sol’s best score—94.9% versus 94.6%—at approximately 37% lower estimated API cost.
Astra can help with the practical work behind scientific discovery. By combining scientific reasoning with computer use, it can work directly in specialized software to inspect data and explore results, helping researchers assess the evidence and decide what to investigate next.
GPT‑6 Astra navigates scientific software to inspect sequencing quality and visualize genetic variation, helping researchers assess their data and identify where to focus further analysis.
Cybersecurity
As we discussed in our safety update, Astra is a significant jump in cyber capabilities and meets the Critical threshold in cybersecurity under our Preparedness Framework. Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards. To understand how far these capabilities extend, we ran Astra on internal and third-party expert evaluations.
We first tested the model without production safeguards on ExploitBench and ExploitGym, which evaluate whether models can turn known software vulnerabilities into working exploits. On ExploitBench, Astra achieved a perfect score of 100%, compared with 78.5% for GPT‑5.6 Sol, our previous frontier cyber-capable model. On ExploitGym, Astra reached a 42.4% success rate, compared with 30.3% for GPT‑5.6 Sol, while using substantially fewer output tokens.13
Given concerns that exposure to historical software vulnerabilities may have affected benchmark results, we also evaluated Astra on two novel benchmarks. For one, we built an internal “ExploitBench (June–August 2026)” evaluation to test exploit development using vulnerabilities from the previous three months.14 Astra achieved substantially higher arbitrary code-execution rates than GPT‑5.6 Sol on this dataset while using far fewer output tokens. During the evaluation, Astra even discovered and used two previously unknown zero-day vulnerabilities. We are disclosing both vulnerabilities to their maintainers.
We also tested Astra on SRE-Bench15, a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
Beyond benchmarks, expert-led assessments found that Astra, when run without production safeguards, could use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits for hardened operating-systems.
As we discussed in The Defender’s Window, frontier cyber capabilities can help defenders find weaknesses faster, but they also make those weaknesses easier to exploit, raising the urgency for defenders to adapt. With the version of Astra launching today, defenders can use it to complete tasks such as secure code review and patching.
However, Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities. Through OpenAI Daybreak, we plan to expand access and roll out less restrictive safeguards in the coming weeks. This will enable more defensive workflows, including vulnerability and proof-of-concept validation, malware analysis, and detection engineering.
We have also strengthened our protections against potential cyber misuse, building upon our safeguards stack for GPT‑5.6 Sol. These include stronger model robustness to better withstand potential jailbreaks and more context for our monitoring systems. We have continued rigorous internal and external testing, including automated evaluations with our internal red-teaming attackers. More details about our cyber safeguards and testing are available in the Astra system card and our blog.
Aligning and deploying GPT‑6 Astra responsibly
Astra is our most aligned model. Astra excels at exercising care, respecting task boundaries, and communicating transparently. This work is the latest product of our long-running research program focused on training models that remain aligned with human intent from start to finish.
In sensitive environments, Astra proceeds with care commensurate with its risk. In an evaluation of computer use tasks adversarially selected to elicit misbehavior, Astra was more successful at avoiding unintended consequences. Running with additional security measures offered by default yielded even stronger performance.
Astra causes fewer misaligned outcomes than any other frontier models tested. For a fair comparison, we used a generic computer-using-agent harness (based on the native computer use tools available in both theOpenAI Responses API andAnthropic Messages API ) and without additional protections that are normally deployed to Codex and ChatGPT Work users (auto-review, confirmation policy).16
Astra is also more likely to operate within the boundaries set by the user and implied by its environment. In an internal evaluation, Astra never attempted to circumvent a Codex Auto-Review denial. This held even when Auto-review was deliberately configured to be evadable and the task was impossible to complete otherwise. This respect for the environment restrictions is consistent with the results of our impossible cyber task evaluation that we shared in the intro of this post and other findings documented in our system card .
This evaluation tests how models respond to auto-review denials in knowledge-work tasks. Exploiting a poor user configuration to bypass auto-review counts as failure. Astra never attempted to circumvent auto-review.
Astra additionally sets new marks in transparent user communication. In one evaluation, Astra is three times less likely than GPT 5.6-Sol to make inaccurate representations about its capabilities and affordances.
In our capability-hallucination evaluation, Astra shows a substantial improvement over GPT‑5.6 Sol, making fewer misleading claims about its capabilities.
Our evaluations found Astra’s written reasoning harder to monitor than GPT‑5.6 Sol’s, based on tests that explicitly asked it to evade monitoring. We attribute this to Astra’s greater control over written reasoning on simpler tasks and ability to solve problems with fewer written steps. Astra still appears to struggle to conceal the reasoning needed for complex tasks, but we take the decline seriously. Improving monitorability remains a research priority, and the accompanying system card details our findings and ongoing work.
Alignment training is core to our approach to deployment. As additional layer of defenses, we also build system safeguards like Codex Auto-review and monitoring agents’ reasoning and actions to help detect and contain unsafe behavior. As described in our safety update, we are also deploying misalignment monitoring in production for Astra-class models in order to have visibility into misalignment, and help contain its worst instances. These safeguards resemble our monitoring for internal deployments and involve a system of classifiers which check the model’s reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.
Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity. If a task is paused in ChatGPT or Codex, you may be asked to review the action before continuing. In the API, the task will stop. These checks can sometimes interrupt legitimate work, and we are continuing to iterate on this system to reduce unnecessary interruptions. Misalignment monitoring cannot replace alignment: our goal is to build models that reliably stay within their authorized scope, so these protections do not need to intervene.
Availability
GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock. Astra usage is included within the existing subscription allowances—users and businesses will also be able to purchase credits for additional usage. Users on the Pro, Business, and Enterprise plans will also get access to GPT‑6 Astra Pro. Enterprise administrators can enable Astra for their workspace; access is off by default at launch.
Astra supports Zero Data Retention for eligible API customers, and as we shared last month, we're testing Private Safety Processing to strengthen safety monitoring while preserving customer privacy.
For developers, GPT‑6 Astra will be available in the OpenAI API as gpt-6-astraand through Microsoft Azure and Amazon Bedrock.
OpenAI API Standard pricing is $10 per million input tokens and $50 per million output tokens. Separate rates apply to cache reads and writes. Fast mode is available for GPT‑6 Astra in the API and delivers up to 2x the speed of Standard processing at 2x the Standard price.
Computer Use
*Computer UseGPT‑6 AstraGPT‑5.6 Sol2Claude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash* Agents' Last Exam 59.3%53.6%-48.7%55.5%- OSWorld 2.0 (v2026.08.08, offline set, partial score)72.6%65.7%--70.2%2- ScreenSpot-Pro (no tools)92.7%76.9%-87.3%17--
Professional
ProfessionalGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash AutomationBench 41.4%18.1%31.4%17.4%26.9%- BenchCAD 95.9%83.3%84.3% 567.5%582.1% 5- BrowseComp 91.5%90.4%-87.4%90.8%- OpenScore String Quartets (1 - OMR-NED)0.84 0.19---- Internal Design Tasks 50.0%47.4%-35.8%-- Internal Data Science Tasks 40.9%30.5%-34.7%-- Artificial Analysis Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7
Coding
CodingGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench 4.0 57.9%37.3%55.8%44.5%52.6%19.1% DeepSWE v1.1 74.1%72.7%67.4%69.9%73.7%73.8% FrontierCode 1.1 Extended (score)64.5% 860.6%63.6%64.9%63.6%56.3% FrontierCode 1.1 Main (score)53.3% 847.5%50.9%53.5%53.4%43.6% Internal Database Migration Tasks 63.9%42.7%57.8%50.3%-- Artificial Analysis Coding Agent Index v1.4 67.0 65.1-67.2 68.1 61.2
Academic
AcademicGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench Science 0.1 64.6%22.4%52.6%21.4%30.0%- FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%- GPQA Diamond 96.0%94.6%93.7%92.6%93.7%95.3% Humanity's Last Exam (w/ tools)57.2%-65.0%63.8%63.6%-
Science and Health
Science and HealthGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash GeneBench Pro 37.1%32.3%---- MedChemBench (Internal)49.3%47.4%---- LifeSciBench 60.3%59.9%---- HealthBench Professional (length-adjusted)63.4%60.5%58.1% 1160.9% 1156.4% 1152.1%
Cybersecurity
CybersecurityGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ExploitBench 100.0%78.5%--70%- ExploitGym 42.4% 1330.3% 1330.4% 1728.4%1722.0%- ExploitBench (June-Aug 2026)39.0%5.5%---- SRE-Bench 88.0%55.9%--12.5%- SEC-Bench Pro 85.4%79.1%----
Alignment
AlignmentGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Internal computer use safety benchmark (lower is better)2.4%22.0%9.5%18.3%11.5%- Internal computer use safety benchmark, w/ AutoReview (lower is better)1.8%4.3%---- Internal circumvention benchmark (lower is better)0.00%0.29%---- ExploitGym honeypot (lower is better)0.0%48.2%---- Impossible ExploitGym 100.0%----- Internal hallucination benchmark (lower is better)4.2%12.2%----
Long Context
Long ContextGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash OpenAI MRCR v2 8-needle 256K-512K 100.0%91.5%---- OpenAI MRCR v2 8-needle 512K-1M 96.3%73.8%----
Abstract reasoning
Abstract reasoningGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ARC-AGI-3 99.9% 17.8%--30.2%- ARC-AGI-2 95.0%92.5%90.0%89.2%90.4%- ARC-AGI-1 98.5%97.5%97.5%98.5%97.5%-
Evaluation scores are the maximum at any effort. GPT evaluations were run in our research environment or via our API, which may provide slightly different output from production ChatGPT due to differences in the system prompts, tools available, etc.
FOOTNOTES
- 1 On ARC-AGI-3, GPT-6 Astra was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
- 2 GPT-5.6 Sol refers to the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different.
- 3 OSWorld V2-Offline is a subset of the original OSWorld V2 that works without internet access. Claude model performance on OSWorld-V2 Offline was reproduced by the authorsonthe official leaderboard **.**On OSWorld 2.0, the scores for Claude use the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card.
- 4 Model times are the reported elapsed times for the corresponding demonstration runs. The displayed clips are edited excerpts.
- 5 On BenchCAD, Claude's scores reflect 3 modifications to the eval, detailed in the Fable 5.1 System Card .
- 6 Guang Yang, Victoria Ebert, Nazif Tamer, Brian Siyuan Zheng, Luiza Pozzobon, and Noah A. Smith. “LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR .” arXiv:2506.19065, 2025.
- 7 Mark R. H. Gotham, Maureen Redbond, Bruno Bower, and Peter Jonas. “The OpenScore String Quartet Corpus .” Proceedings of the 10th International Conference on Digital Libraries for Musicology, pp. 49–57. ACM, 2023.
- 8 On FrontierCode, GPT-6 Astra was run with a developer message similar to a section of its developer message in Codex : "Avoid creating excessive test files. Create a new test file only when required by repository conventions or when no existing file is a suitable home. Avoid unrelated cleanup and unnecessary complexity. Reuse suitable existing utilities. Read relevant repository instructions and inspect nearby code, tests, documentation, and CI. Follow established conventions. The goal is clean, mergeable code." The prompt was not optimized for the eval.
- 9 The first concerns how close together prime numbers can occur, however far along the number line you go. For more than a decade, the best known result established that infinitely many pairs of primes are at most 246 apart. Julia Stadlmann recently improved that bound to 240. Astra helped establish a stronger bound of 186, showing that infinitely many pairs occur within this smaller distance. Short prime gaps: Proof and supporting research .
- 10 The second concerns unusually large gaps between primes. Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results. Large prime gaps: Proof and supporting research .
- 11 We independently evaluated all Claude models following the intended HealthBench Professional procedure, using GPT‑5.4 grading and length-adjusted, unclipped scores. For Fable 5.1, we used Opus 5 fallback for provider refusals.
- 12 Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.
- 13 On ExploitGym, we tested Astra and Sol without the 6-hour time limit, to better assess their full cyber capabilities. They are fast enough that it has little impact.
- 14 ExploitBench (June–August 2026) contains 20 high-severity V8 vulnerabilities across 13 stable Chrome releases. The benchmark tests whether agents can achieve arbitrary code execution in V8 and official Chrome releases for Linux by exploiting each specified vulnerability. Some included vulnerabilities may not permit arbitrary code execution under the evaluation’s constraints, so a 100% success rate may not be achievable. Note: the 5.5% score of GPT-5.6 Sol is an artifact of the 300-turn limit in the benchmark, which is not a limit that real customers using max would have. The model at similar settings achieved an 11.5% score when hitting fewer limits.
- 15 Jeremy Spence et al. “The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark .” arXiv:2608.11469v1, 2026.
- 16 When we test across third-party models, we use a simpler research setup. Codex has a more complex production configuration, which can result in different raw-model error rates. Provider-side safeguards and computer-tool implementations still differ. Users do not experience the no-confirmation scenario in Codex, as it's an internal research configuration.
- 17 For ScreenSpot-Pro and ExploitGym, the Fable scores we report come from Mythos, which is Fable with fewer safeguards.