GPT-6 Astra“今日起向一小部分组织推出,未来几天将向所有 ChatGPT Plus、Pro、Business 和 Enterprise 用户开放,也可通过 OpenAI API 和 AWS 使用”——我自己还没试过,所以暂时没有太多可说的。
其 API 定价将与 Claude Fable 5 和 5.1 相同:输入 $10/百万 token,输出 $50/百万 token。这显然是 OpenAI 对标 Fable 的产品,而且在 OpenAI 自行报告的多数基准测试中,其得分似乎都高于 Fable。
最令人印象深刻的是,Astra 在近期(3 月发布)的 ARC-AGI 3 基准测试中取得了 99.9% 的得分——不过值得注意的是,Fable 5 尚未公布测试结果,而且 ARC-AGI 博客指出,99.9% 的得分是在花费 $19K、使用 OpenAI 定制的“Provider Adapter harness”的情况下取得的,而默认 ARC-AGI harness 的得分为 62.7%,花费 $26K。
Provider Adapter harness 会在请求之间保留不透明的推理状态,并对较长的对话使用压缩处理,使模型能够复用先前的工作成果。
考虑到近期 Hugging Face 事件,Astra 在安全任务上表现出色并不令人意外。它在 ExploitBench 上得分 100%(GPT-5.6 Sol 为 78.5%),在 ExploitGym 上得分 42.4%(Sol 为 30.3%),在 SRE-Bench 二进制逆向工程任务中四次尝试内得分 99.2%,而 Sol 为 68.7%。
它在长上下文方面也更出色:在 OpenAI 的八针基准测试中,它在 256K–512K token 范围内得分 100%,在 512K–1M token 范围内得分 96.3%。OpenAI 或许已经攻克了长上下文处理中长期存在的一个挑战。
不过它并非在所有方面都胜出。Artificial Analysis 指出,Astra 在其 Intelligence Index(智能指数)上仍被 Fable 击败:
在智能水平上与 GPT-5.6 Sol 比肩:GPT-6 Astra 在该指数上的得分为 61,与 GPT-5.6 Sol 持平。这比 Claude Fable 5.1(含回退的最高分)低 5 分。该模型也落后于 Meta 新发布的 Muse Spark 1.3(最高分)。
它在 Coding Agent Index(编程智能体指数)上表现更好:
领跑 Coding Agent Index 成本效益前沿:在最大推理强度下,GPT-6 Astra 的成本与 GPT-5.6 Sol(最大强度)大致相当,而指数得分高出 2 分。在相同得分下,该模型每项任务的成本不到 Claude Fable 5 的一半。
等我获得 Astra 的访问权限后,会再写更多相关内容。该 API 模型一旦上线,其标签将是 gpt-6-astra。
GPT-6 Astra is "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS" - I've not tried it yet myself, so I don't have a great deal to say about it yet.
It's going to be API priced at the same rate as Claude Fable 5 and 5.1: $10/million input and $50/million output. This is clearly OpenAI's Fable competitor, and appears to score higher than Fable on most of OpenAI's self-reported benchmarks.
Most impressively, Astra scores 99.9% on the recent (released in March) ARC-AGI 3 benchmark - though notably Fable 5 does not yet have a published result, and the ARC-AGI blog notes that the 99.9% score was achieved for $19K using OpenAI's custom "Provider Adapter harness", while the default ARC-AGI harness scored 62.7% for $26K.
The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.
Unsurprisingly, given the recent Hugging Face incident, Astra is a beast at security tasks. It scores 100% on ExploitBench (GPT-5.6 Sol got 78.5%), 42.4% on ExploitGym (Sol got 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering compared to Sol's 68.7%.
It's also better at long context: on OpenAI's eight-needle benchmark it got 100% at 256K–512K tokens and 96.3% at 512K–1M tokens. OpenAI may have vanquished one of the ongoing challenges with long context processing.
It doesn't win at everything though. Artificial Analysis note that Astra is still beaten by Fable on their Intelligence Index:
Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).
It did better on their Coding Agent Index:
Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.
I'll write more about Astra once I get access to it. The API model label once it rolls out will be gpt-6-astra.