GPT-6 Astra:新一代智能
新一代智能
我们正式推出 GPT‑6 Astra——全球最智能、对齐程度最高的模型。
GPT‑6 Astra 汇聚了我们在预训练、强化学习与对齐领域多年的研究成果和重大投入。Astra 在计算机操作、浏览、软件工程、网络安全、科学和专业工作方面均达到业界顶尖水平。Astra 以 98% 的得分饱和 FrontierMath Tier 4,并已助力解决数学领域长期悬而未决的开放问题。Astra 还以 99.9% 的得分饱和 ARC-AGI-3,并以 100% 的得分饱和 ExploitBench。此外,它在计算机与浏览器使用方面也树立了新的前沿标杆,能够以无与伦比的速度、准确性和判断力处理最严苛的专业工作。
GPT‑6 Astra 于今日向部分组织开放,并将在未来数天内向所有 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,同时也可通过 OpenAI API 和 AWS 使用。
Terminal-Bench Science 0.1
API 成本
GPT-6 Astra
GPT-5.6 Sol
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
Terminal-Bench Science 0.1 测试智能体能否使用代码和终端工具完成科学研究工作流程,包括分析数据、运行模拟和拟合模型。在参与对比的模型中,GPT‑6 Astra 以 64.6% 的成绩创下新高,而 Claude Fable 5.1 为 52.6%,且预估 API 成本约低 31%。在较低成本的设置下,Astra 得分为 61.1%,而 GPT‑5.6 Sol 的最佳成绩为 22.4%,且预估 API 成本约低 27%。
Astra 是我们对齐程度最高的模型,在理解用户意图和模型行为方面有显著改进——你可以更放心地将任务委托给 Astra,信赖它的判断。作为测试方式之一,我们借鉴 Hugging Face 事件构建了一项新的评估,用于检验模型在面对困难或不可能完成的任务时,是否会超出其预期范围。相比之下,未配备生产环境安全防护的 GPT‑5.6 Sol 有 48% 的情况下超出了授权目标,而 GPT‑6 Astra 在 0% 的情况下出现此类行为。
全球最强的计算机使用模型
GPT‑6 Astra 在计算机使用的速度、准确性和安全性方面开创了新前沿。它可以处理繁琐的任务,例如填写在线表单、更新 CRM 中的客户记录以及整理你的日历。它可以进行在线研究,并在你的电子邮件或文档编辑器中起草摘要。它可以分析科学数据、生成图表、创建网站,并运行前端 QA 检查,以确保该网站上的所有功能都能正常运行。它可以帮助你自主安装和测试软件,并排查你在屏幕上看到的问题。这些改进也体现在我们最先进的评估结果中。
Agents’ Last Exam 在真实软件环境中测试智能体处理复杂专业任务的能力,涵盖财务建模、工程和媒体制作等领域。在所示对比中,GPT‑6 Astra 取得了新的最高分,得分 59.3%,而 Claude Opus 5 为 55.5%,GPT‑5.6 Sol 为 53.6%。在这些最高得分设置下,Astra 使用的输出 token 数量也比 Opus 5 少约 65%。
这些改进还在真实知识工作任务中带来了显著的效率提升。在 OSWorld 2.0 的延迟模拟中,Astra 在每项任务上比 GPT‑5.6 Sol 节省约 47% 的时间,同时实现了更高的计算机使用性能——得分 72.6%,每项任务约 40 分钟,而对比方得分 65.7%,每项任务约 75 分钟。³
GPT‑6 Astra 的计算机使用能力可见于多个领域的输出成果,包括游戏开发、电气工程和日常知识工作:
这是 GPT‑6 Astra 在 KiCad 中执行印刷电路板(PCB)布局的 15 秒浓缩回放——它将电子原理图转化为可制造的 PCB,通过放置元件和布线铜连接来实现。PCB 布局是当今每款电子设备不可或缺的环节,这是一项手工任务,也是电子设计流程中常见的延迟来源。加速这一过程意味着让工程师能够以显著更高的节奏去发明、优化和测试他们的下一个创意。
除 Astra 之外,我们还在更新 Codex harness,以显著提升计算机使用的速度。结合 Astra 的高效率,在 Mind2Web 基准上,任务完成速度相比当前 GPT‑5.6 Sol 体验提升了 1.9 倍。模型在速度上的改进意味着它可以为你处理许多耗时的生活任务,而且比你亲自完成得更快。⁴
GPT‑6 Astra:2 分 54 秒
“我们将在发布当天把 GPT‑6 Astra 集成到 Devin 的 harness 中,它在我们的内部测试基准上表现出最先进的性能。其出色的计算机使用能力、写作能力和代码库理解能力开箱即用地提升了测试效果:视频明显更容易跟进,报告也更清晰、更简洁。”
Silas Alberti,Cognition 研究高级副总裁
专业工作的阶跃式变革
GPT‑6 Astra 将计算机使用方面的进步与针对专业环境的定向训练相结合,助力应对复杂的工作任务。它兼具解决复杂问题所需的智能,以及执行多步骤工作流、生成精美文档、电子表格和演示文稿的能力。
BenchCAD 通过生成 CAD 代码来测试模型能否从多视角渲染中重建 3D 对象。在工具辅助下,GPT‑6 Astra 在所示对比中达到新高,几何重叠得分 95.9%,而 GPT‑5.6 Sol 为 83.3%,Claude Fable 5.1.5 报告得分为 84.3%。在所示配置下,估算 API 成本比 Sol 低约 43%,比 Fable 5.1 低 86%。
GPT‑6 Astra 是我们最擅长遵循现有模板、制作版式精良、以结构化叙事简明传达要点的幻灯片的模型。它能创建清晰、结构良好的文档、演示文稿、电子表格和分析,遵循你的模板并匹配你的写作与视觉风格。Astra 还经过专门训练,只将真正重要的上下文提取到输出中,而不是重复对当前工作无用的信息。这一切意味着它可以输出更直接可用的成果物,贴合你的业务上下文与标准。
参考文件
GPT‑6 Astra 输出
GPT‑6 Astra 仅使用 OpenAI 演示模板中的几张幻灯片,就制作了一个关于虚构模型 GPT‑Gaia 的幻灯片演示,全程把握住了正确的语气和版式。这意味着你可以期待得到符合你业务标准、格式正确的演示文稿。
GPT‑6 Astra 还为其构建的网站、游戏、应用和渲染图带来了更强的视觉判断力。借助 ChatGPT 中的 Sites,Astra 可以直接根据提示词创建、托管和分享网站、Web 应用和游戏。
“Astra 在能力和效率两方面都为我们带来了显著优势。它能成功执行我们最复杂的创意工作流,同时比我们测试过的其他模型最多节省 20% 的 token。最重要的是,对我们的客户而言,这意味着更高质量的产出。”
Alex Mashrabov,Higgsfield AI 首席执行官兼联合创始人
GPT‑6 Astra 在 Blender 中为房屋建模,并将其转化为 Unreal Engine 5 中可漫步的场景,帮助设计师和客户在房屋建成前探索布局、体验空间。
该模型能够通过生动的画面、引人入胜的游戏玩法和精准的动作让游戏活起来,使非技术人员也能在几分钟内创建并游玩超越基础元素的定制游戏。图片来源:Pietro Schirano。
当指令留有解读空间时,GPT‑6 Astra 比之前的模型更善于做出正确判断。它会利用上下文填补常规信息空缺,并在答案可能改变结果时提出有针对性的问题。在 Codex 中,它可以异步提问,同时继续推进不依赖你回复的工作。如果你不回复,它会在适当之处基于合理假设继续推进,但在重大决策上会等待你的输入。
以下示例展示了 Astra 如何在信息缺失可能实质性改变答案的日常任务中展开协作。

随着任务推进,Astra 也更能保持方向感。早期模型有时会把引导性消息当作新目标,从而丢失原始请求或先前约束的线索。Astra 能够纳入新要求、在被要求时改变方向,并在回答旁支问题的同时不丢掉更宏观的任务。
“在复杂的法律任务上,Astra 相比 GPT‑5.6 Sol 是一次显著的品质提升。在我们的早期测试中,Astra 的突出之处在于它像一位眼光敏锐的律师那样处理法律工作:它能区分正式档案与既有记录、识别缺乏依据的假设,并把信息缺口转化为具体的起草立场。”
Niko Grupen,Harvey 应用研究主管
编程
GPT‑6 Astra 是迄今为止最优秀的软件工程模型。
“GPT‑6 Astra 在我们内部编程基准上取得了业界领先的表现,并且在交易直觉评估中相比 GPT‑5.6 Sol 展现出明显的进步。用于智能体编程时,GPT‑6 Astra 的沟通方式更便于开发者理解,产出的代码也只需更少的迭代即可达到生产级质量。”
John Crepezzi,Jane Street AI 助手团队
“我们在第一代评估之一上对 Astra 进行了低、中、高三档努力程度的测试,它的表现显著领先于 GPT 5.6 Sol。更高的努力程度意味着在全新构建上投入更多迭代、通过浏览器测试进行更多验证,并且更倾向于代码执行而非 apply-patch。理解模型如何分配努力程度,正是我们帮助数百万开发者从想法到可用应用走得更快、更可靠的方式。”
Fabian Hedin,Lovable CTO 兼联合创始人
“GPT‑6 Astra 在我们内部编码基准测试中展现出顶尖性能,与 GPT‑5.6 Sol 相比,在交易直觉评估方面也呈现出明显的进步。用于智能体编码时,GPT‑6 Astra 的沟通方式更便于开发者理解,生成的代码达到生产质量所需的迭代次数也更少。”
John Crepezzi,AI 助手团队,Jane Street
“我们在第一代评估中分别以低、中、高三档努力程度测试了 Astra,其表现显著优于 GPT 5.6 Sol。更高的努力程度意味着在全新构建上进行更多次迭代、通过浏览器测试进行更多验证,并且更倾向于代码执行而非 apply-patch。理解模型如何分配努力程度,正是我们帮助数百万开发者从创意到可用应用之间走出一条更快、更可靠路径的方式。”
Fabian Hedin,CTO 兼联合创始人,Lovable
- Jane Street
- Lovable
Terminal-Bench 4.0 在复杂的终端任务上测试智能体,涵盖软件工程、系统配置和数据分析。GPT‑6 Astra 以 57.9% 的成绩创下新高,相比之下 GPT‑5.6 Sol2 为 37.3%,Claude Fable 5.1 为 55.8%,而每项任务的预估 API 成本分别低约 9% 和 63%。
通过 Astra,我们为 Codex 引入了一种在上下文窗口填满时保存和检索上下文的新方式。过去,模型在长时间会话中(例如调试复杂问题或处理大型重构时)会使用压缩(compaction)来总结工作。每次压缩都可能遗漏关于某个修复为何失败或某个组件行为方式的细节。在 Codex 中,Astra 可以跨上下文窗口保存笔记,保留累积的细节,而无需反复将它们压缩成单一摘要。较早的上下文窗口仍可搜索,因此 Astra 可以找到来自先前消息和工具输出的需求或测试结果——即使这些信息未被记录在其笔记中。你可以在你的 Codex `config.toml` 中启用这项实验性功能,并且它将在未来几周内成为 Astra 的默认设置。
推动科学发现
“这个故事是:一个时代的终结,另一个时代的开始。”
Greg Burnham,EpochAI
GPT‑6 Astra 是科学发现、数学和健康领域的重大进步。今天,我们分享关于素数间隔问题的另外两项成果。9, 10
Astra 还在多项数学和科学评估中创下了新纪录。
GPQA Diamond 测试生物学、化学和物理学领域的研究生水平科学推理能力。GPT‑6 Astra 在所示对比中达到了 96.0% 的新高。在较低成本的设置下,它也超过了 GPT‑5.6 Sol 的最佳成绩——94.9% 对比 94.6%——而估算的 API 成本大约低 37%。
Astra 能够助力科学发现背后的实际工作。通过将科学推理与计算机使用能力相结合,它可以直接在专业软件中工作,检查数据并探索结果,帮助研究人员评估证据并决定下一步该研究什么。
GPT‑6 Astra 能够操作科学软件,检查测序质量并可视化遗传变异,帮助研究人员评估自身数据,并确定应将进一步分析的重点放在哪里。
网络安全
正如我们在安全更新中所讨论的,Astra 在网络能力上是一次重大飞跃,根据我们的预备框架,它在网络安全方面达到了“严重”阈值。它识别和开发零日漏洞的能力可以帮助防御者发现并修补弱点,但这也意味着需要更强的安全防护措施。为了了解这些能力的延伸范围,我们在内部及第三方专家评估中运行了 Astra。
我们首先在未启用生产环境安全防护的情况下,在 ExploitBench 和 ExploitGym 上测试了该模型,这两个基准用于评估模型能否将已知软件漏洞转化为可用的漏洞利用程序。在 ExploitBench 上,Astra 取得了 100% 的满分,而此前我们具备前沿网络能力的模型 GPT‑5.6 Sol 得分仅为 78.5%。在 ExploitGym 上,Astra 的成功率达到 42.4%,相比之下 GPT‑5.6 Sol 为 30.3%,同时 Astra 使用的输出 token 数量要少得多。
考虑到接触历史软件漏洞可能影响基准测试结果的担忧,我们也在两个全新基准上评估了 Astra。其一,我们构建了一个内部“ExploitBench(2026年6月–8月)”评测,用于测试利用过去三个月内漏洞的漏洞利用开发能力。14 在该数据集上,Astra 实现的任意代码执行成功率显著高于 GPT‑5.6 Sol,同时使用的输出 token 数量远少于后者。评测期间,Astra 甚至发现并利用了两个此前未知的零日漏洞。我们正在向这两个漏洞的维护者进行披露。
我们还在 SRE-Bench15 上测试了 Astra。该基准用于衡量模型能否在无法获取原始源代码的情况下,对软件二进制文件进行逆向工程以理解其核心逻辑。Astra 在单次尝试中解决了 88.0% 的任务,在四次尝试内解决了 99.2% 的任务;而 GPT‑5.6 Sol 的对应成绩分别为 55.9% 和 68.7%。
除基准测试外,专家主导的评估发现,Astra 在未启用生产环境安全防护措施运行时,能够利用此前未知的漏洞在加固浏览器中实现任意代码执行,并为加固操作系统创建权限提升漏洞利用。
正如我们在《防御者的窗口》中所讨论的,前沿网络能力可以帮助防御者更快地发现弱点,但同时也使这些弱点更容易被利用,从而提高了防御者适应的紧迫性。借助今天发布的 Astra 版本,防御者可以使用它来完成安全代码审查和漏洞修补等任务。
然而,Astra 将拒绝执行更高级的网络安全任务,例如为漏洞创建概念验证型漏洞利用代码。通过 OpenAI Daybreak,我们计划在未来几周内扩大访问范围,并推出限制更少的安全防护措施。这将支持更多防御性工作流程,包括漏洞与概念验证验证、恶意软件分析以及检测工程。
我们还基于 GPT‑5.6 Sol 的安全防护体系,加强了对潜在网络滥用的防护。其中包括增强模型鲁棒性以更好地抵御潜在的越狱攻击,以及为监控系统提供更多上下文信息。我们持续开展严格的内部和外部测试,包括使用内部红队攻击者进行自动化评估。有关我们的网络安全防护和测试的更多详情,请参阅 Astra 系统卡和我们的博客。
负责任地对齐和部署 GPT‑6 Astra
Astra 是我们对齐程度最高的模型。Astra 擅长谨慎行事、尊重任务边界,并进行透明的沟通。这项工作是我们长期研究计划的最新成果,该计划专注于训练从头到尾始终与人类意图保持对齐的模型。
在敏感环境中,Astra 会以与其风险相称的谨慎态度行事。在一项针对性地选取以诱发不当行为的计算机使用任务评估中,Astra 在避免意外后果方面表现得更为出色。在默认启用额外安全措施的情况下运行,其表现更为强劲。
Astra 在所有接受测试的前沿模型中,产生的不对齐结果最少。为了公平比较,我们使用了通用的计算机操作智能体框架(基于 OpenAI Responses API 和 Anthropic Messages API 中均提供的原生计算机使用工具),并且未启用通常部署给 Codex 和 ChatGPT 工作版用户的额外保护措施(自动审查、确认策略)。16
Astra 也更倾向于在用户设定的边界及其环境所暗示的边界内运行。在一项内部评估中,Astra 从未试图绕过 Codex 自动审查的拒绝决定。即使在自动审查被故意配置为可被规避、且任务在其他情况下无法完成时,这一情况依然成立。这种对环境限制的尊重,与我们在本文开头分享的不可能网络任务评估结果以及系统卡中记录的其他发现是一致的。
这项评估测试模型在知识工作任务中如何应对自动审查的拒绝决定。利用用户的不良配置来绕过自动审查被视为失败。Astra 从未试图规避自动审查。
Astra 还在透明的用户沟通方面树立了新的标杆。在一项评估中,Astra 对其能力和功能做出不准确表述的可能性比 GPT 5.6-Sol 低三倍。
在我们的能力幻觉评估中,与 GPT‑5.6 Sol 相比,Astra 表现出显著进步,对其能力做出的误导性表述更少。
我们的评估发现,基于明确要求其规避监控的测试,Astra 的书面推理比 GPT‑5.6 Sol 的更难监控。我们将此归因于 Astra 在较简单任务上对书面推理有更强的控制力,并且能够用更少的书面步骤解决问题。Astra 在复杂任务上似乎仍然难以隐藏所需的推理过程,但我们严肃对待这一能力下降。提升可监控性仍是研究重点,随附的系统卡详细说明了我们的发现和正在进行的工作。
对齐训练是我们部署方法的核心。作为额外的防御层,我们还构建了系统安全防护措施,如 Codex Auto-review,并监控智能体的推理和行动,以帮助检测和遏制不安全行为。正如我们的安全更新中所述,我们还在生产环境中为 Astra 级模型部署了错位监控,以便能够洞察错位情况,并帮助遏制其最严重的实例。这些防护措施类似于我们对内部部署的监控,涉及一个分类器系统,该系统检查模型的推理和行动是否存在未经授权的行为,并自动停止可能未经授权的活动。
鉴于 Astra 的网络安全能力显著增强,我们特别谨慎以确保此次部署的安全可靠。额外的安全检查有时可能会减慢、暂停或停止合法工作,包括防御性网络安全工作。如果任务在 ChatGPT 或 Codex 中被暂停,你可能会被要求在继续之前审查该操作。在 API 中,任务将停止。这些检查有时会中断合法工作,我们正在继续迭代此系统以减少不必要的干扰。错位监控不能取代对齐:我们的目标是构建能够可靠地保持在授权范围内运行的模型,从而使这些保护措施无需介入。
可用性
GPT‑6 Astra 于今日向一组有限的组织推出,并在未来几天内向所有 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,同时也可通过 OpenAI API 和 AWS 使用。Astra 的使用量包含在现有订阅额度内——用户和企业也可以购买额外使用量的积分。Pro、Business 和 Enterprise 套餐用户还将获得 GPT‑6 Astra Pro 的访问权限。Enterprise 管理员可以为其工作区启用 Astra;上线时该功能默认关闭。
Astra 支持符合条件的 API 客户使用零数据保留(Zero Data Retention),并且正如我们上个月分享的,我们正在测试私有安全处理(Private Safety Processing),以在保护客户隐私的同时加强安全监控。
对于开发者,GPT‑6 Astra 在 OpenAI API 中以 gpt-6-astra 的形式提供,同时也可在 Amazon Bedrock 中使用。OpenAI API 标准定价为每百万输入 token 10 美元,每百万输出 token 50 美元。缓存读取和写入适用单独的费率。GPT‑6 Astra 在 API 中提供快速模式,处理速度最高可达标准模式的 2 倍,价格为标准价格的 2 倍。
计算机使用
计算机使用 | GPT‑6 Astra | GPT‑5.6 Sol2 | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | **** | Gemini 3.8 Flash ---|---|---|---|---|---|---|--- Agents' Last Exam | 59.3% | 53.6% | - | 48.7% | 55.5% | - OSWorld 2.0(v2026.08.08,离线集,部分得分) | 72.6% | 65.7% | - | - | 70.2% | 3 | - ScreenSpot-Pro(无工具) | 92.7% | 76.9% | - | 87.3% | 16 | -
专业版
ProfessionalGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash AutomationBench 41.4%18.1%31.4%17.4%26.9%- BenchCAD 95.9%83.3%84.3% 467.5%482.1% 4- BrowseComp 91.5%90.4%-87.4%90.8%- OpenScore String Quartets (1 - OMR-NED)0.84 0.19---- Internal Design Tasks 50.0%47.4%-35.8%-- Internal Data Science Tasks 40.9%30.5%-34.7%-- Artificial Analysis Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7
编程
编程GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench 4.0 57.9%37.3%55.8%42.0%52.3%19.1% DeepSWE v1.1 74.1%72.7%67.4%69.9%73.7%73.8% FrontierCode 1.1 Extended (score)64.5% 760.6%63.6%64.9%63.6%56.3% FrontierCode 1.1 Main (score)53.3% 747.5%50.9%53.5%53.4%43.6% Internal Database Migration Tasks 63.9%42.7%57.8%50.3% Artificial Analysis Coding Agent Index v1.4 67.0 65.1 67.2 68.1 61.2
学术
学术GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench Science 0.1 64.6%22.4%52.6%21.4%30.0%- FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%- GPQA Diamond 96.0%94.6%93.7%92.6%93.7%95.3% Humanity's Last Exam (w/ tools)57.2%-65.0%63.8%63.6%-
科学与健康
科学与健康GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash GeneBench Pro 37.8%28.7% MedChemBench (Internal)49.3%47.4% LifeSciBench 60.3%59.9% HealthBench Professional (length-adjusted)63.4%60.5%58.1% 1060.9% 1056.4% 1052.1%
网络安全
网络安全GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ExploitBench 100.0%78.5%70% Exploit Gym 42.4% 1230.3% 1230.4% 1628.4%1622.0%16 ExploitBench(2026年6月-8月)39.0%11.5% SRE-Bench 88.0%55.9%12.5% SEC-Bench Pro 85.4%79.1%
对齐
对齐GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash 内部计算机使用安全基准(越低越好)2.4%22.0%9.5%18.3%11.5%- 内部计算机使用安全基准,含 AutoReview(越低越好)1.8%4.3%---- 内部规避基准(越低越好)0.00%0.29%---- ExploitGym 蜜罐(越低越好)0.0%48.2%---- Impossible ExploitGym 100.0%----- 内部幻觉基准(越低越好)4.2%12.2%----
长上下文
长上下文GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash OpenAI MRCR v2 8-needle 256K-512K 100.0%91.5%---- OpenAI MRCR v2 8-needle 512K-1M 96.3%73.8%----
抽象推理
抽象推理GPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ARC-AGI-3 99.9% [T7]7.8%--30.2%- ARC-AGI-2 95.0%92.5%90.0%89.2%90.4%- ARC-AGI-1 98.5%97.5%97.5%98.5%97.5%-
评测分数为任意投入下的最高值。GPT 评测在我们的研究环境中或通过我们的 API 运行,由于系统提示词、可用工具等方面的差异,其输出可能与生产环境中的 ChatGPT 略有不同。
脚注
- 1 在 ARC-AGI-3 上,GPT-6 Astra 使用我们的 responses API harness 运行,该 harness 更改了两项设置以更好地匹配真实世界性能。这些更改并非专门针对 ARC-AGI-3。
- 2 GPT-5.6 Sol 指我们的 API、ChatGPT Codex 和 ChatGPT Work 中可用的版本。ChatGPT Chat 中的版本略有不同。
- 3 OSWorld V2-Offline 是原始 OSWorld V2 的一个子集,可在无互联网访问的情况下运行。Claude 模型在 OSWorld-V2 Offline 上的性能由独立第三方复现。在 OSWorld 2.0 上,Claude 的分数使用官方设置,而非 Fable 5.1 System Card 中修改后的任务和修改后的评分标准。
- 5 在 BenchCAD 上,Claude 的分数反映了对评测的 3 项修改,详见 Fable 5.1 System Card。
- 6 Guang Yang, Victoria Ebert, Nazif Tamer, Brian Siyuan Zheng, Luiza Pozzobon, and Noah A. Smith. “LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR.” arXiv:2506.19065, 2025.
- 7 Mark R. H. Gotham、Maureen Redbond、Bruno Bower 与 Peter Jonas。《OpenScore 弦乐四重奏语料库》。第十届音乐学数字图书馆国际会议论文集,第 49–57 页。ACM,2023 年。
- 8 在 FrontierCode 上,GPT-6 Astra 运行时附带了一条开发者消息,该消息与其在 Codex 中的开发者消息某一部分类似:“避免创建过多的测试文件。仅在仓库规范要求或现有文件均不适合时,才新建测试文件。避免无关清理和不必要的复杂性。复用合适的现有工具。阅读相关仓库说明,并查看附近代码、测试、文档和 CI。遵循既定规范。目标是干净、可合并的代码。”该提示词并未针对该评测进行优化。
- 9 第一个问题关乎素数之间能出现多近的距离,无论你在数轴上走到多远。十多年来,已知最佳结果表明,存在无穷多对素数,它们之间最多相隔 246。Julia Stadlmann 最近将该界限改进为 240。Astra 帮助确立了更强的界限 186,表明存在无穷多对素数出现在这一更小的间距内。短素数间隔:证明及相关研究。
- 10 第二个问题关乎素数之间异常大的间隔。Astra 改进了这些间隔界限中的一个项,该界限已 80 多年未曾变动。我们正在分享这两项结果的证明、精简版思维链及验证材料。大素数间隔:证明及相关研究。
- 11 我们按照预期的 HealthBench Professional 标准流程,使用 GPT‑5.4 评分和长度调整后的非截断分数,对所有 Claude 模型进行了独立评估。对于 Fable 5.1,当提供商拒绝回答时,我们使用 Opus 5 作为后备。
- 12 Claude Fable 5 和 5.1 未包含在 LifeSciBench Gold v1、GeneBench Pro v13 和 MedChemBench 中,因为它们在评估中拒绝了大多数问题。
- 13 在 ExploitGym 上,我们测试了 Astra 和 Sol,未设置 6 小时的时间限制,以便更好地评估其完整的网络能力。它们的速度足够快,因此这一限制影响不大。
- 14 ExploitBench(2026 年 6 月至 8 月)包含 13 个稳定版 Chrome 版本中的 20 个高危 V8 漏洞。该基准测试旨在检验智能体能否通过利用每个指定漏洞,在 V8 和官方 Linux 版 Chrome 中实现任意代码执行。部分收录的漏洞在评估约束下可能无法实现任意代码执行,因此 100% 的成功率可能无法达到。
- 15 Jeremy Spence 等人。“智能体网络安全的下一个挑战:一个现实、无污染的逆向工程基准。” arXiv:2608.11469v1,2026 年。
- 16 当我们对第三方模型进行测试时,我们使用更简单的研究设置。Codex 具有更复杂的生产配置,这可能导致原始模型错误率有所不同。提供商侧的安全防护措施和计算机工具实现仍然存在差异。用户不会在 Codex 中体验到无确认场景,因为这是一种内部研究配置。
- 17 对于 ScreenSpot-Pro 和 ExploitGym,我们报告的 Fable 得分来自 Mythos,即安全防护较少的 Fable 版本。
GPT-6 Astra: A new generation of intelligence
A new generation of intelligence
We’re introducing GPT‑6 Astra, the world’s most intelligent and aligned model.
GPT‑6 Astra brings together years of research and big bets across pre-training, reinforcement learning, and alignment. Astra is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work. Astra saturates FrontierMath Tier 4 with a 98% score, having already helped solve long-standing open problems in mathematics. Astra also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score. It also sets a new frontier on computer and browser use, handling the most demanding professional work with unmatched speed, accuracy, and judgment.
GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Terminal-Bench Science 0.1
API Cost
GPT-6 Astra
GPT-5.6 Sol
Claude Fable 5.1
Claude Fable 5
Claude Opus 5
Terminal-Bench Science 0.1 tests whether agents can complete scientific research workflows using code and terminal tools, including analyzing data, running simulations, and fitting models. GPT‑6 Astra reaches a new high among the models compared at 64.6%, versus 52.6% for Claude Fable 5.1, at approximately 31% lower estimated API cost. At a lower-cost setting, Astra scores 61.1%, versus GPT‑5.6 Sol’s best result of 22.4%, at approximately 27% lower estimated API cost.
Astra is our most aligned model, with substantial improvements in understanding user intent and model behavior—you can delegate tasks with greater confidence in Astra’s judgment. As one way that we test this, we built a new evaluation informed by the Hugging Face incident that evaluates whether a model facing a difficult or impossible task will go beyond its intended scope. Compared to GPT‑5.6 Sol, which without production safeguards went beyond the authorized target 48% of the time, GPT‑6 Astra did this in 0% of cases.
The world’s best computer use model
GPT‑6 Astra marks a new frontier in the speed, accuracy, and safety of computer use. It can take care of tedious tasks like filling out online forms, updating customer records in a CRM, and organizing your calendar. It can conduct online research and draft summaries in your email or in your document editor. It can analyze scientific data, generate plots, create a website, and run frontend QA checks to make sure all the features on that site work. It can help you autonomously install and test software, and troubleshoot problems you see on screen. These improvements are also reflected in our state-of-the-art evaluation results.
Agents’ Last Exam tests agents on complex professional tasks in real software, from financial modeling to engineering and media production. GPT‑6 Astra reaches a new high in the comparison shown, scoring 59.3%, compared with 55.5% for Claude Opus 5 and 53.6% for GPT‑5.6 Sol. At these highest-scoring settings, Astra also uses approximately 65% fewer output tokens than Opus 5.
These improvements also result in significant efficiency gains in real knowledge-work tasks. In latency simulations on OSWorld 2.0, Astra achieves higher computer-use performance in about 47% less time per task than GPT‑5.6 Sol, scoring 72.6% at roughly 40 minutes per task, compared with 65.7% at roughly 75 minutes.3
GPT‑6 Astra’s computer-use capabilities can be seen in outputs across domains, including game development, electrical engineering, and everyday knowledge work:
This is a 15-second condensed playback of GPT‑6 Astra performing printed circuit board (PCB) layout in KiCad, turning an electronic schematic into a manufacturable PCB by placing components and routing copper connections. Integral to every electronic device today, PCB layout is a manual task and common source of latency in the electronics design process. Accelerating it means freeing engineers to invent, optimize, and test their next idea at a significantly higher cadence.
Alongside Astra, we are also updating the Codex harness to significantly improve the speed of computer use. Combined with Astra’s efficiency, this translates to a 1.9x faster task completion compared to the current GPT‑5.6 Sol experience, on the Mind2Web benchmark. The model’s improvements on speed mean it can take on many time-consuming life tasks for you, faster than you can.4
GPT‑6 Astra: 2 min 54 sec
“We’re integrating GPT‑6 Astra into Devin’s harness on launch day, where it delivers state-of-the-art performance on our internal testing benchmark. Its excellent computer use, writing, and codebase understanding improved testing right out of the box: videos are noticeably easier to follow, and reports are clearer and more concise”
Silas Alberti, SVP Research, Cognition
A step change in professional work
GPT‑6 Astra pairs advances in computer use with targeted training for professional environments, to help tackle complex work tasks. It combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations.
BenchCAD tests whether models can reconstruct 3D objects from multi-view renders by generating CAD code. With tools, GPT‑6 Astra reaches a new high in the comparison shown, achieving a 95.9% geometric-overlap score, versus 83.3% for GPT‑5.6 Sol and 84.3% reported for Claude Fable 5.1.5Estimated API cost is approximately 43% lower than Sol and 86% lower than Fable 5.1 in the configurations shown.
GPT‑6 Astra is our best model for adhering to existing templates and producing slides that are well laid out and succinctly convey key points with a structured narrative. It creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow your templates and match your writing and visual style. Astra is also trained to specifically pull only the context that matters into outputs, instead of repeating information unnecessary for the work at hand. All this means it can output more immediately usable artifacts that match your business context and standards.
Reference file
GPT‑6 Astra output
GPT‑6 Astra creates a slideshow about GPT‑Gaia, a fictional model, using just a few slides from OpenAI’s presentation template, capturing the correct tone and layout throughout. This means you can expect slide decks that are correctly formatted for your business standards.
GPT‑6 Astra also brings stronger visual judgment to the websites, games, applications, and renderings it builds. With Sites in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
“Astra gives us a significant advantage in both capability and efficiency. It successfully executes our most complex creative workflows while using up to 20% fewer tokens than other models we've tested. Most importantly, for our customers, it means higher quality output.”
Alex Mashrabov, CEO and Co-founder, Higgsfield AI
GPT‑6 Astra models a house in Blender and turns it into a walkable scene in Unreal Engine 5, helping designers and clients explore the layout and experience the space before it’s built.
The model can bring games to life through vivid graphics, engaging gameplay and accurate motion, allowing non-technical people to create and play custom games that go beyond rudimentary elements in minutes. Credit: Pietro Schirano.
When instructions leave room for interpretation, GPT‑6 Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codex, it can ask asynchronously while continuing work that doesn’t depend on your reply. If you don’t respond, it proceeds with sensible assumptions where appropriate, but waits for your input on consequential decisions.
The examples below show how Astra collaborates on everyday tasks where missing information can materially change the answer.

Astra is also better at staying oriented as a task evolves. Earlier models sometimes treated steering messages as a new goal, losing track of the original request or earlier constraints. Astra incorporates new requirements, changes course when asked, and answers side questions without dropping the broader task.
“Astra is a significant quality improvement over GPT‑5.6 Sol across complex legal tasks. In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does: it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions.”
Niko Grupen, Head of Applied Research, Harvey
Coding
GPT‑6 Astra is the best model for software engineering to date.
“GPT‑6 Astra delivers state-of-the-art performance on our internal coding benchmarks and shows a clear step forward in trading intuition evaluations compared with GPT‑5.6 Sol. When used for agentic coding, GPT‑6 Astra communicates in a way that’s easier for developers to follow and produces code that requires less iteration to reach production quality.”
John Crepezzi, AI Assistants, Jane Street
“We tested Astra across low, medium, and high effort on one of our first-generation evals, and it came out significantly ahead of GPT 5.6 Sol. Higher effort buys more iterations on a fresh build, more verification through browser testing, and a lean toward code execution over apply-patch. Understanding how a model spends its effort is how we give millions of builders a faster, more reliable path from idea to working app.”
Fabian Hedin, CTO & Co-founder, Lovable
“GPT‑6 Astra delivers state-of-the-art performance on our internal coding benchmarks and shows a clear step forward in trading intuition evaluations compared with GPT‑5.6 Sol. When used for agentic coding, GPT‑6 Astra communicates in a way that’s easier for developers to follow and produces code that requires less iteration to reach production quality.”
John Crepezzi, AI Assistants, Jane Street
“We tested Astra across low, medium, and high effort on one of our first-generation evals, and it came out significantly ahead of GPT 5.6 Sol. Higher effort buys more iterations on a fresh build, more verification through browser testing, and a lean toward code execution over apply-patch. Understanding how a model spends its effort is how we give millions of builders a faster, more reliable path from idea to working app.”
Fabian Hedin, CTO & Co-founder, Lovable
- Jane Street
- Lovable
Terminal-Bench 4.0 tests agents on complex terminal-based tasks, including software engineering, system configuration, and data analysis. GPT‑6 Astra reaches a new high at 57.9%, compared with 37.3% for GPT‑5.6 Sol2and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost per task, respectively.
With Astra, we’re introducing a new way for Codex to preserve and retrieve context when the context window fills. Historically, models have used compaction to summarize work during long sessions, such as when debugging complex issues or tackling large refactors. Each compaction can leave out details about why a fix failed or how a component behaves. In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs—even if that information wasn’t captured in its notes. You can enable this experimental feature in your Codex config.toml , and it will become the default for Astra in the coming weeks.
Advancing scientific discovery
“The story is: end of one era, start of another.”
Greg Burnham, EpochAI
GPT‑6 Astra is a major advance for scientific discovery, mathematics, and health. Today, we’re sharing two further results on the gaps between prime numbers.9, 10
Astra also sets new records across a suite of math and science evaluations.
GPQA Diamond tests graduate-level scientific reasoning in biology, chemistry, and physics. GPT‑6 Astra reaches a new high in the comparison shown at 96.0%. At a lower-cost setting, it also exceeds GPT‑5.6 Sol’s best score—94.9% versus 94.6%—at approximately 37% lower estimated API cost.
Astra can help with the practical work behind scientific discovery. By combining scientific reasoning with computer use, it can work directly in specialized software to inspect data and explore results, helping researchers assess the evidence and decide what to investigate next.
GPT‑6 Astra navigates scientific software to inspect sequencing quality and visualize genetic variation, helping researchers assess their data and identify where to focus further analysis.
Cybersecurity
As we discussed in our safety update, Astra is a significant jump in cyber capabilities and meets the Critical threshold in cybersecurity under our Preparedness Framework. Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses, but it also creates a need for stronger safeguards. To understand how far these capabilities extend, we ran Astra on internal and third-party expert evaluations.
We first tested the model without production safeguards on ExploitBench and ExploitGym, which evaluate whether models can turn known software vulnerabilities into working exploits. On ExploitBench, Astra achieved a perfect score of 100%, compared with 78.5% for GPT‑5.6 Sol, our previous frontier cyber-capable model. On ExploitGym, Astra reached a 42.4% success rate, compared with 30.3% for GPT‑5.6 Sol, while using substantially fewer output tokens.13
Given concerns that exposure to historical software vulnerabilities may have affected benchmark results, we also evaluated Astra on two novel benchmarks. For one, we built an internal “ExploitBench (June–August 2026)” evaluation to test exploit development using vulnerabilities from the previous three months.14 Astra achieved substantially higher arbitrary code-execution rates than GPT‑5.6 Sol on this dataset while using far fewer output tokens. During the evaluation, Astra even discovered and used two previously unknown zero-day vulnerabilities. We are disclosing both vulnerabilities to their maintainers.
We also tested Astra on SRE-Bench15, a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
Beyond benchmarks, expert-led assessments found that Astra, when run without production safeguards, could use previously unknown vulnerabilities to achieve arbitrary code execution in hardened browsers and create privilege-escalation exploits for hardened operating-systems.
As we discussed in The Defender’s Window, frontier cyber capabilities can help defenders find weaknesses faster, but they also make those weaknesses easier to exploit, raising the urgency for defenders to adapt. With the version of Astra launching today, defenders can use it to complete tasks such as secure code review and patching.
However, Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities. Through OpenAI Daybreak, we plan to expand access and roll out less restrictive safeguards in the coming weeks. This will enable more defensive workflows, including vulnerability and proof-of-concept validation, malware analysis, and detection engineering.
We have also strengthened our protections against potential cyber misuse, building upon our safeguards stack for GPT‑5.6 Sol. These include stronger model robustness to better withstand potential jailbreaks and more context for our monitoring systems. We have continued rigorous internal and external testing, including automated evaluations with our internal red-teaming attackers. More details about our cyber safeguards and testing are available in the Astra system card and our blog.
Aligning and deploying GPT‑6 Astra responsibly
Astra is our most aligned model. Astra excels at exercising care, respecting task boundaries, and communicating transparently. This work is the latest product of our long-running research program focused on training models that remain aligned with human intent from start to finish.
In sensitive environments, Astra proceeds with care commensurate with its risk. In an evaluation of computer use tasks adversarially selected to elicit misbehavior, Astra was more successful at avoiding unintended consequences. Running with additional security measures offered by default yielded even stronger performance.
Astra causes fewer misaligned outcomes than any other frontier models tested. For a fair comparison, we used a generic computer-using-agent harness (based on the native computer use tools available in both theOpenAI Responses API andAnthropic Messages API ) and without additional protections that are normally deployed to Codex and ChatGPT Work users (auto-review, confirmation policy).16
Astra is also more likely to operate within the boundaries set by the user and implied by its environment. In an internal evaluation, Astra never attempted to circumvent a Codex Auto-Review denial. This held even when Auto-review was deliberately configured to be evadable and the task was impossible to complete otherwise. This respect for the environment restrictions is consistent with the results of our impossible cyber task evaluation that we shared in the intro of this post and other findings documented in our system card .
This evaluation tests how models respond to auto-review denials in knowledge-work tasks. Exploiting a poor user configuration to bypass auto-review counts as failure. Astra never attempted to circumvent auto-review.
Astra additionally sets new marks in transparent user communication. In one evaluation, Astra is three times less likely than GPT 5.6-Sol to make inaccurate representations about its capabilities and affordances.
In our capability-hallucination evaluation, Astra shows a substantial improvement over GPT‑5.6 Sol, making fewer misleading claims about its capabilities.
Our evaluations found Astra’s written reasoning harder to monitor than GPT‑5.6 Sol’s, based on tests that explicitly asked it to evade monitoring. We attribute this to Astra’s greater control over written reasoning on simpler tasks and ability to solve problems with fewer written steps. Astra still appears to struggle to conceal the reasoning needed for complex tasks, but we take the decline seriously. Improving monitorability remains a research priority, and the accompanying system card details our findings and ongoing work.
Alignment training is core to our approach to deployment. As additional layer of defenses, we also build system safeguards like Codex Auto-review and monitoring agents’ reasoning and actions to help detect and contain unsafe behavior. As described in our safety update, we are also deploying misalignment monitoring in production for Astra-class models in order to have visibility into misalignment, and help contain its worst instances. These safeguards resemble our monitoring for internal deployments and involve a system of classifiers which check the model’s reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.
Given the significant increase in Astra’s cybersecurity capabilities, we are being especially careful to make this deployment safe and secure. Extra safety checks can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity. If a task is paused in ChatGPT or Codex, you may be asked to review the action before continuing. In the API, the task will stop. These checks can sometimes interrupt legitimate work, and we are continuing to iterate on this system to reduce unnecessary interruptions. Misalignment monitoring cannot replace alignment: our goal is to build models that reliably stay within their authorized scope, so these protections do not need to intervene.
Availability
GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS. Astra usage is included within the existing subscription allowances—users and businesses will also be able to purchase credits for additional usage. Users on the Pro, Business, and Enterprise plans will also get access to GPT‑6 Astra Pro. Enterprise administrators can enable Astra for their workspace; access is off by default at launch.
Astra supports Zero Data Retention for eligible API customers, and as we shared last month, we're testing Private Safety Processing to strengthen safety monitoring while preserving customer privacy.
For developers, GPT‑6 Astra is available in the OpenAI API as gpt-6-astra, and is also available in Amazon Bedrock. OpenAI API Standard pricing is $10 per million input tokens and $50 per million output tokens. Separate rates apply to cache reads and writes. Fast mode is available for GPT‑6 Astra in the API and delivers up to 2x the speed of Standard processing at 2x the Standard price.
Computer Use
Computer UseGPT‑6 AstraGPT‑5.6 Sol2 Claude Fable 5.1Claude Fable 5Claude Opus 5****Gemini 3.8 Flash Agents' Last Exam 59.3%53.6%-48.7%55.5%- OSWorld 2.0 (v2026.08.08, offline set, partial score)72.6%65.7%--70.2%3 - ScreenSpot-Pro (no tools)92.7%76.9%-87.3%16 --
Professional
ProfessionalGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash AutomationBench 41.4%18.1%31.4%17.4%26.9%- BenchCAD 95.9%83.3%84.3% 467.5%482.1% 4- BrowseComp 91.5%90.4%-87.4%90.8%- OpenScore String Quartets (1 - OMR-NED)0.84 0.19---- Internal Design Tasks 50.0%47.4%-35.8%-- Internal Data Science Tasks 40.9%30.5%-34.7%-- Artificial Analysis Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7
Coding
CodingGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench 4.0 57.9%37.3%55.8%42.0%52.3%19.1% DeepSWE v1.1 74.1%72.7%67.4%69.9%73.7%73.8% FrontierCode 1.1 Extended (score)64.5% 760.6%63.6%64.9%63.6%56.3% FrontierCode 1.1 Main (score)53.3% 747.5%50.9%53.5%53.4%43.6% Internal Database Migration Tasks 63.9%42.7%57.8%50.3% Artificial Analysis Coding Agent Index v1.4 67.0 65.1 67.2 68.1 61.2
Academic
AcademicGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Terminal-Bench Science 0.1 64.6%22.4%52.6%21.4%30.0%- FrontierMath Tier 4 (v2)97.6%83.0%87.8%87.8%73.2%- GPQA Diamond 96.0%94.6%93.7%92.6%93.7%95.3% Humanity's Last Exam (w/ tools)57.2%-65.0%63.8%63.6%-
Science and Health
Science and HealthGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash GeneBench Pro 37.8%28.7% MedChemBench (Internal)49.3%47.4% LifeSciBench 60.3%59.9% HealthBench Professional (length-adjusted)63.4%60.5%58.1% 1060.9% 1056.4% 1052.1%
Cybersecurity
CybersecurityGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ExploitBench 100.0%78.5%70% Exploit Gym 42.4% 1230.3% 1230.4% 1628.4%1622.0%16 ExploitBench (June-Aug 2026)39.0%11.5% SRE-Bench 88.0%55.9%12.5% SEC-Bench Pro 85.4%79.1%
Alignment
AlignmentGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash Internal computer use safety benchmark (lower is better)2.4%22.0%9.5%18.3%11.5%- Internal computer use safety benchmark, w/ AutoReview (lower is better)1.8%4.3%---- Internal circumvention benchmark (lower is better)0.00%0.29%---- ExploitGym honeypot (lower is better)0.0%48.2%---- Impossible ExploitGym 100.0%----- Internal hallucination benchmark (lower is better)4.2%12.2%----
Long Context
Long ContextGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash OpenAI MRCR v2 8-needle 256K-512K 100.0%91.5%---- OpenAI MRCR v2 8-needle 512K-1M 96.3%73.8%----
Abstract reasoning
Abstract reasoningGPT‑6 AstraGPT‑5.6 SolClaude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash ARC-AGI-3 99.9% [T7]7.8%--30.2%- ARC-AGI-2 95.0%92.5%90.0%89.2%90.4%- ARC-AGI-1 98.5%97.5%97.5%98.5%97.5%-
Evaluation scores are the maximum at any effort. GPT evaluations were run in our research environment or via our API, which may provide slightly different output from production ChatGPT due to differences in the system prompts, tools available, etc.
FOOTNOTES
- 1 On ARC-AGI-3, GPT-6 Astra was run with our responses API harness, which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
- 2 GPT-5.6 Sol refers to the version available in our API, ChatGPT Codex, and ChatGPT Work. The version in ChatGPT Chat is slightly different.
- 3 OSWorld V2-Offline is a subset of the original OSWorld V2 that works without internet access. Claude model performance on OSWorld-V2 Offline was reproduced by an independent third-party. On OSWorld 2.0, the scores for Claude use the official settings, and not the modified tasks and modified grading from the Fable 5.1 System Card.
- 5 On BenchCAD, Claude's scores reflect 3 modifications to the eval, detailed in the Fable 5.1 System Card .
- 6 Guang Yang, Victoria Ebert, Nazif Tamer, Brian Siyuan Zheng, Luiza Pozzobon, and Noah A. Smith. “LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR .” arXiv:2506.19065, 2025.
- 7 Mark R. H. Gotham, Maureen Redbond, Bruno Bower, and Peter Jonas. “The OpenScore String Quartet Corpus .” Proceedings of the 10th International Conference on Digital Libraries for Musicology, pp. 49–57. ACM, 2023.
- 8 On FrontierCode, GPT-6 Astra was run with a developer message similar to a section of its developer message in Codex : "Avoid creating excessive test files. Create a new test file only when required by repository conventions or when no existing file is a suitable home. Avoid unrelated cleanup and unnecessary complexity. Reuse suitable existing utilities. Read relevant repository instructions and inspect nearby code, tests, documentation, and CI. Follow established conventions. The goal is clean, mergeable code." The prompt was not optimized for the eval.
- 9 The first concerns how close together prime numbers can occur, however far along the number line you go. For more than a decade, the best known result established that infinitely many pairs of primes are at most 246 apart. Julia Stadlmann recently improved that bound to 240. Astra helped establish a stronger bound of 186, showing that infinitely many pairs occur within this smaller distance. Short prime gaps: Proof and supporting research .
- 10 The second concerns unusually large gaps between primes. Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results. Large prime gaps: Proof and supporting research .
- 11 We independently evaluated all Claude models following the intended HealthBench Professional procedure, using GPT‑5.4 grading and length-adjusted, unclipped scores. For Fable 5.1, we used Opus 5 fallback for provider refusals.
- 12 Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.
- 13 On ExploitGym, we tested Astra and Sol without the 6-hour time limit, to better assess their full cyber capabilities. They are fast enough that it has little impact.
- 14 ExploitBench (June–August 2026) contains 20 high-severity V8 vulnerabilities across 13 stable Chrome releases. The benchmark tests whether agents can achieve arbitrary code execution in V8 and official Chrome releases for Linux by exploiting each specified vulnerability. Some included vulnerabilities may not permit arbitrary code execution under the evaluation’s constraints, so a 100% success rate may not be achievable.
- 15 Jeremy Spence et al. “The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark .” arXiv:2608.11469v1, 2026.
- 16 When we test across third-party models, we use a simpler research setup. Codex has a more complex production configuration, which can result in different raw-model error rates. Provider-side safeguards and computer-tool implementations still differ. Users do not experience the no-confirmation scenario in Codex, as it's an internal research configuration.
- 17 For ScreenSpot-Pro and ExploitGym, the Fable scores we report come from Mythos, which is Fable with fewer safeguards.