Cognition:GPT-6 Astra 在 FrontierCode 1.1 上以低 64% 成本逼近 Fable 5 编码水平

Rohan Paul · @rohanpaul_ai · X·2026-09-04 07:02·30分钟前
AI 导读

Cognition 宣布 GPT-6 Astra 将接入 Devin,在 FrontierCode 1.1 上得分与 Fable 5 相差不到 0.4 分,成本降低 64%,并在其内部测试基准上创造 SOTA,生成更全面的测试、更清晰的报告和更好的视频证据。

Rohan Paul@rohanpaul_ai
51AI 编辑部评分,满分 100

Cognition:GPT-6 Astra 在 FrontierCode 1.1 上以低 64% 成本逼近 Fable 5 编码水平

2026-09-04 07:02· 30分钟前
AI 导读

Cognition 宣布 GPT-6 Astra 将接入 Devin,在 FrontierCode 1.1 上得分与 Fable 5 相差不到 0.4 分,成本降低 64%,并在其内部测试基准上创造 SOTA,生成更全面的测试、更清晰的报告和更好的视频证据。

Cognition says GPT-6 Astra reached near-Fable 5 coding quality (within 0.4 points) at 64% lower rollout cost on FrontierCode 1.1

Now, FrontierCode benchmark is quite unusual because it asks whether an AI coding agent can produce a pull request a maintainer would actually merge, rather than stopping at functional correctness.

Cognition built 150 tasks with maintainers from 36 open-source repositories, grading correctness, regression safety, tests, scope, style, and adherence to each codebase's conventions.

The score is a weighted rubric aggregate rather than a solve rate, and any run that misses a blocking requirement receives zero.

METR separately found that roughly half of earlier SWE-bench Verified patches that passed automated tests still would not have been merged by maintainers.

i.e. FrontierCode really measures review-quality code.

CognitionGPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal te...