要点
- OpenAI 正在发布其迄今最强大的模型 GPT-6 Astra。付费 ChatGPT 用户和云平台应在未来几天内获得访问权限。
- 在基准测试中,Astra 在逻辑、数学和软件工程等关键学科上显著优于其前代 GPT-5.6 Sol 以及 Anthropic 的 Fable 5 系列模型。
- Token 价格高出 2.5 倍,与 Anthropic 的 Fable 5.1 持平,但 OpenAI 认为,根据使用场景不同,每个已完成任务的实际成本实际上更低。
OpenAI 已发布其迄今最强大的模型 GPT-6 Astra。总裁 Greg Brockman 表示,按照 OpenAI 自己的定义,它或许已经够格称为“AGI”,或者至少已触手可及——即在大多数具有经济价值的工作上超越人类表现的 AI 系统。
GPT-6 Astra 首先通过 OpenAI 的 Daybreak 计划向部分组织推出。在未来几天内,它将向 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,同时也可通过 API 以及 AWS Bedrock 和 Microsoft Azure 等云平台获取。
Astra 在位于得克萨斯州的 Stargate 设施中,使用超过 100,000 块 GPU 完成了预训练。OpenAI 研究员 Aidan Clark 称这是公司有史以来规模最大的一次训练。Clark 表示,从 Sol 到 Astra 的能力跃升幅度大于从早期模型到 Sol 的跃升,部分原因在于之前的 AI 模型在监控训练过程中发挥了作用。
在 OpenAI 发布的基准测试中,Astra 的得分远高于其前代 GPT-5.6 Sol 以及 Anthropic 的 Fable 系列模型。Astra 在多个领域均取得顶尖成绩:逻辑推理(ARC-AGI-3 上达到 98.6%,不过是在其自身测试条件下)、数学(FrontierMath Tier 4 v2 上达到 97.6%)、软件工程(DeepSWE v1.1 上达到 74.1%)、专家知识(GPQA Diamond 上达到 96%)、工程(BenchCAD 上达到 95.9%)以及网络安全(ExploitBench 上达到 100%)。
计算机使用
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Agents' Final Exam | 59.3% | 53.6% | 48.7% | 55.5% | ||
| OSWorld 2.0(离线,部分) | 72.6% | 65.7% | 70.2% ³ | |||
| ScreenSpot-Pro(无工具) | 92.7% | 76.9% | 87.3% ¹⁷ |
专业
| 基准 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | |
| BenchCAD | 95.9% | 83.3% | 84.3% ⁵ | 67.5% ⁵ | 82.1% ⁵ | |
| BrowseComp | 91.5% | 90.4% | 87.4% | 90.8% | ||
| OpenScore String Quartets | 0.84 | 0.19 | ||||
| 内部设计任务 | 50.0% | 47.4% | 35.8% | |||
| 内部数据科学任务 | 40.9% | 30.5% | 34.7% | |||
| AA Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
BenchCAD 成本:比 Sol 低约 43%,比 Fable 5.1 低约 86%。
编码
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Terminal Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | 19.1% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended | 64.5% ⁸ | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% |
| FrontierCode 1.1 Main | 53.3% ⁸ | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% |
| 内部数据库迁移 | 63.9% | 42.7% | 57.8% | 50.3% | ||
| AA 编程智能体指数 v1.4 | 67.0 | 65.1 | 67.2 | 68.1 | 61.2 |
Terminal-Bench 4.0 成本:比 Sol 低约 9%,比 Fable 5.1 低约 63%。
学术
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| 人类最后的考试(工具版) | 57.2% | 65.0% | 63.8% | 63.6% |
低成本设置:Terminal-Bench Science 为 61.1%,成本降低约 27%;GPQA Diamond 为 94.9%,成本降低约 37%。素数间隔从 240 改进到 186,并且一个大间隔边界项在 80 多年来首次得到改进。
科学与健康
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| GeneBench Pro | 37.8% | 28.7% | ||||
| MedChemBench(内部) | 49.3% | 47.4% | ||||
| LifeSciBench | 60.3% | 59.9% | ||||
| HealthBench Professional | 63.4% | 60.5% | 56.6% ¹¹ | 60.9% ¹¹ | 54.5% ¹¹ | 52.1% |
Fable 5 和 5.1 未纳入 LifeSciBench、GeneBench Pro 和 MedChemBench,因为它们会拒绝回答大多数问题。
网络安全
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| ExploitBench | 100.0% | 78.5% | 70% | ||
| ExploitGym | 42.4% ¹³ | 30.3% ¹³ | 30.4% ¹⁷ | 28.4% ¹⁷ | 22.0% ¹⁷ |
| ExploitBench(2026年6月–8月) | 39.0% | 5.5% | |||
| SRE-Bench | 88.0% | 55.9% | 12.5% | ||
| SEC-Bench Pro | 85.4% | 79.1% |
SRE-Bench 四次尝试内:99.2%,而 Sol 为 68.7%。Astra 在评估过程中发现了两个此前未知的零日漏洞。
对齐(除 Impossible ExploitGym 外,数值越低越好)
| 指标 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| 计算机安全 | 2.4% | 22.0% | 9.5% | 18.3% | 11.5% |
| 相同,带 AutoReview | 1.8% | 4.5% | |||
| 规避 | 0.00% | 0.29% | |||
| ExploitGym 蜜罐 | 0.0% | 48.2% | |||
| 不可能完成的 ExploitGym 任务 | 100.0% | ||||
| 模型幻觉 | 4.2% | 12.2% |
不可能任务范围测试:Sol 有 48% 的时间超出了其授权目标,Astra 为 0%。Astra 错误陈述自身能力的可能性比 Sol 低 3 倍。
长上下文
| 基准测试 | Astra | Sol |
|---|---|---|
| MRCR v2 8 针 256K–512K | 100.0% | 91.5% |
| MRCR v2 8 针 512K–1M | 96.3% | 73.8% |
抽象推理
| 基准测试 | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| ARC-AGI-3 | 99.9% ¹ | 7.8% | – | – | 30.2% |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% |
据报道,在科学工作方面,该模型改进了一项关于素数间隔的数学结果,并在生物学、化学、医学和物理学评测中创下新纪录。OpenAI 还将 Astra 定位为一款能够像人类一样可靠操作计算机的模型,在衡量该能力的 OSWorld 2.0 上,Astra 以每个任务约 40 分钟的速度取得了 72.6% 的成绩,而 Sol 的得分是 65.7%,每个任务大约耗时 75 分钟。
比 Sol 更贵,但 OpenAI 表示每个任务的成本更低
GPT-6 Astra 通过 API 使用标准模式时,输入 token 价格为每百万个 10 美元,输出 token 价格为每百万个 50 美元。快速模式承诺提供 2.5 倍速度,但价格翻倍,这使得 Astra 比 GPT-5.6 Sol 贵 2.5 倍,与 Anthropic 的 Fable 5.1 处于同一价格区间。
布罗克曼认为,按 token 价格来比较模型正变得越来越不靠谱,因为 OpenAI 的 token 与竞争对手的并不相同,甚至在其自身的不同模型系列之间也不具可比性。他表示,真正重要的是每个完成任务的价格,而 OpenAI 已经在尝试这种定价模式。据该公司称,在 DeepSWE v1.1 上,Astra 的最高配置相比 Sol,可将每项任务的预估 API 成本降低约 57%。
在新闻发布会上,布罗克曼承认并不存在一个明确定义的 AGI 时刻,他表示团队原本以为 OpenAI 成立时会有一个人人都能识别的明显门槛。但实际情况并非如此,这一转变比预期更为渐进,这也是 OpenAI 去年春天阐述过的立场。布罗克曼在发布会结束时仍表示:“欢迎来到 AGI 时代。”OpenAI 首席执行官 Sam Altman 此前曾表示,他预计今年年底前会出现一个他称之为 AGI 的模型。
首个达到 OpenAI 关键网络安全阈值的模型
Astra 是 OpenAI 根据其《预备框架》归类为“关键”级别的首个模型。这意味着,在获得适当工具和访问权限的情况下,该模型能够发现此前未知的漏洞,并在防御严密的系统中构建漏洞利用链,全程无需人类逐步指导。OpenAI 也承认,Astra 的书面推理过程比 GPT-5.6 Sol 的更难监控。
在 ExploitBench 这一衡量模型发现并利用真实软件漏洞能力的基准测试中,Astra 取得了 100% 的满分成绩。OpenAI 表示,该模型在评估过程中发现了两个此前未知的漏洞,公司随后已将其报告给相关厂商。人类专家确认,Astra 能够识别涵盖浏览器和操作系统等多个软件类别的新型零日漏洞。
目前,最先进的网络安全能力仅限于 Daybreak Blue 项目中的可信防御方使用。OpenAI 此前曾推迟 Astra 的发布,以进行更多的安全测试。这些双重用途的能力是一把双刃剑:一个能自主发现漏洞的智能体,帮助防御者修补漏洞的便利程度,与帮助攻击者利用漏洞的便利程度不相上下。
CNET
Axios
Key Points
- OpenAI is launching GPT-6 Astra, its most capable model to date. Paying ChatGPT customers and cloud platforms should get access in the coming days.
- In benchmarks, Astra significantly outperforms its predecessor GPT-5.6 Sol and Anthropic's Fable 5 models across key disciplines including logic, math, and software engineering.
- Token prices are 2.5x higher and on par with Anthropic's Fable 5.1, but OpenAI argues that the cost per completed task is actually lower depending on the use case.
OpenAI has shipped GPT-6 Astra, its most capable model to date. President Greg Brockman says it might already qualify as "AGI" or is at least within reach, meaning an AI system that outperforms humans at most economically valuable work by OpenAI's own definition.
GPT-6 Astra is rolling out first to select organizations through OpenAI's Daybreak program. Over the coming days, it will become available to ChatGPT Plus, Pro, Business, and Enterprise customers, as well as through the API and cloud platforms like AWS Bedrock and Microsoft Azure.
Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark called it the company's largest training run ever. The jump from Sol to Astra represents a bigger capability gain than the jump to Sol from earlier models, Clark said, in part because previous AI models played a role in monitoring training.
In benchmarks OpenAI published, Astra scores well above its predecessor GPT-5.6 Sol and Anthropic's Fable models. Astra hits top marks across a range of disciplines: logical reasoning (98.6 percent on ARC-AGI-3, though under its own test conditions), math (97.6 percent on FrontierMath Tier 4 v2), software engineering (74.1 percent on DeepSWE v1.1), expert knowledge (96 percent on GPQA Diamond), engineering (95.9 percent on BenchCAD), and cybersecurity (100 percent on ExploitBench).
Computer Use
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Agents' Final Exam | 59.3% | 53.6% | 48.7% | 55.5% | ||
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% | 70.2% ³ | |||
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | 87.3% ¹⁷ |
Professional
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | |
| BenchCAD | 95.9% | 83.3% | 84.3% ⁵ | 67.5% ⁵ | 82.1% ⁵ | |
| BrowseComp | 91.5% | 90.4% | 87.4% | 90.8% | ||
| OpenScore String Quartets | 0.84 | 0.19 | ||||
| Internal Design Tasks | 50.0% | 47.4% | 35.8% | |||
| Internal Data Science Tasks | 40.9% | 30.5% | 34.7% | |||
| AA Intelligence Index v4.1.1 | 61.2 | 60.9 | 65.7 | 62.1 | 63.1 | 58.7 |
BenchCAD cost: ~43% below Sol, ~86% below Fable 5.1.
Coding
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Terminal Bench 4.0 | 57.7% | 37.3% | 55.8% | 42.0% | 52.3% | 19.1% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| FrontierCode 1.1 Extended | 64.5% ⁸ | 60.6% | 63.6% | 64.9% | 63.6% | 56.3% |
| FrontierCode 1.1 Main | 53.3% ⁸ | 47.5% | 50.9% | 53.5% | 53.4% | 43.6% |
| Internal Database Migration | 63.9% | 42.7% | 57.8% | 50.3% | ||
| AA Coding Agent Index v1.4 | 67.0 | 65.1 | 67.2 | 68.1 | 61.2 |
Terminal-Bench 4.0 cost: ~9% below Sol, ~63% below Fable 5.1.
Academic
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 87.8% | 87.8% | 73.2% | |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| Humanity's Last Exam (tools) | 57.2% | 65.0% | 63.8% | 63.6% |
Lower-cost settings: Terminal-Bench Science 61.1% at ~27% lower cost; GPQA Diamond 94.9% at ~37% lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years.
Science and Health
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 | Gem 3.8 F |
|---|---|---|---|---|---|---|
| GeneBench Pro | 37.8% | 28.7% | ||||
| MedChemBench (internal) | 49.3% | 47.4% | ||||
| LifeSciBench | 60.3% | 59.9% | ||||
| HealthBench Professional | 63.4% | 60.5% | 56.6% ¹¹ | 60.9% ¹¹ | 54.5% ¹¹ | 52.1% |
Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions.
Cybersecurity
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| ExploitBench | 100.0% | 78.5% | 70% | ||
| ExploitGym | 42.4% ¹³ | 30.3% ¹³ | 30.4% ¹⁷ | 28.4% ¹⁷ | 22.0% ¹⁷ |
| ExploitBench (Jun–Aug 2026) | 39.0% | 5.5% | |||
| SRE-Bench | 88.0% | 55.9% | 12.5% | ||
| SEC-Bench Pro | 85.4% | 79.1% |
SRE-Bench within four attempts: 99.2% versus 68.7% for Sol. Astra found two previously unknown zero-days during evaluation.
Alignment (lower is better except for Impossible ExploitGym)
| Metric | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| Computer Safety | 2.4% | 22.0% | 9.5% | 18.3% | 11.5% |
| Same, with AutoReview | 1.8% | 4.5% | |||
| Circumvention | 0.00% | 0.29% | |||
| ExploitGym honeypot | 0.0% | 48.2% | |||
| Impossible ExploitGym | 100.0% | ||||
| Hallucination | 4.2% | 12.2% |
Impossible-task scope test: Sol exceeded its authorized target 48% of the time, Astra 0%. Astra is 3x less likely to misstate its own capabilities.
Long Context
| Benchmark | Astra | Sol |
|---|---|---|
| MRCR v2 8-needle 256K–512K | 100.0% | 91.5% |
| MRCR v2 8-needle 512K–1M | 96.3% | 73.8% |
Abstract Reasoning
| Benchmark | Astra | Sol | Fable 5.1 | Fable 5 | Opus 5 |
|---|---|---|---|---|---|
| ARC-AGI-3 | 99.9% ¹ | 7.8% | – | – | 30.2% |
| ARC-AGI-2 | 95.0% | 92.5% | 90.0% | 89.2% | 90.4% |
| ARC-AGI-1 | 98.5% | 97.5% | 97.5% | 98.5% | 97.5% |
In scientific work, the model reportedly improved a mathematical result on prime gaps and set new records in biology, chemistry, medicine, and physics evaluations. OpenAI also positions Astra as a model that can reliably operate a computer the way a human would, and on OSWorld 2.0, which measures that ability, Astra scored 72.6 percent at about 40 minutes per task compared to Sol's 65.7 percent at roughly 75 minutes.
More expensive than Sol, but OpenAI says cheaper per task
GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens in standard mode through the API. Fast mode, which promises 2.5x speed, doubles the price, making Astra 2.5x more expensive than GPT-5.6 Sol and putting it in the same price range as Anthropic's Fable 5.1.
Brockman argued that token prices are becoming a poor way to compare models, since OpenAI's tokens aren't the same as a competitor's and aren't even comparable across its own model families. What matters, he said, is the price per completed task, and OpenAI is already experimenting with that pricing model. On DeepSWE v1.1, Astra's top configuration cuts estimated API costs per task by about 57 percent compared to Sol, according to the company.
During the press briefing, Brockman acknowledged there's no clearly defined AGI moment, saying the team originally thought there would be an obvious threshold everyone would recognize when OpenAI was founded. That's not how it played out, and the transition has been more gradual than expected, a position OpenAI laid out last spring. Brockman still closed the briefing by saying, "Welcome to the AGI era." OpenAI CEO Sam Altman had previously said he expects a model he'd call AGI by the end of the year.
First model to hit OpenAI's critical cybersecurity threshold
Astra is the first model that OpenAI classifies as "critical" under its Preparedness Framework. That means the model can find previously unknown vulnerabilities and build exploit chains across well-defended systems when given the right tools and access, all without a human guiding it step by step. OpenAI also admits that Astra's written reasoning is harder to monitor than GPT-5.6 Sol's.
On ExploitBench, a benchmark that measures a model's ability to discover and exploit real software vulnerabilities, Astra scores a perfect 100 percent. OpenAI says the model discovered two previously unknown vulnerabilities during evaluation, which the company then reported to the affected vendors. Human experts confirmed that Astra can identify novel zero-day vulnerabilities across several software categories, including browsers and operating systems.
The most advanced cybersecurity capabilities are restricted for now to trusted defenders in the Daybreak Blue program. OpenAI had previously delayed Astra's release to run more safety testing. These dual-use capabilities cut both ways: an agent that can autonomously find a vulnerability helps a defender patch it just as easily as it helps an attacker exploit it.
CNET
Axios