Artificial Analysis 发布 Intelligence Index v4.2,GPT-6 Astra 评分争议后调整基准

The Decoder:AI News(RSS)·2026-09-06 02:21·31分钟前·Matthias Bastian
AI 导读

Artificial Analysis 发布 Intelligence Index v4.2,此前其对 GPT-6 Astra 的评分与 Epoch AI、ARC-AGI-3 等评估结果相左而受到质疑。

The Decoder:AI News(RSS)
58AI 编辑部评分,满分 100

Artificial Analysis 发布 Intelligence Index v4.2,GPT-6 Astra 评分争议后调整基准

2026-09-06 02:21· 31分钟前· Matthias Bastian
AI 导读

Artificial Analysis 发布 Intelligence Index v4.2,此前其对 GPT-6 Astra 的评分与 Epoch AI、ARC-AGI-3 等评估结果相左而受到质疑。

Image description

Artificial Analysis has released version 4.2 of its Intelligence Index, likely in response to criticism that its benchmarks failed to capture GPT-6 Astra's actual progress.

Earlier, several evaluations, including OpenAI's own, had placed Astra well ahead of the field. Epoch AI ranked it first out of 267 models with 169 points across more than 50 benchmarks, and ARC-AGI-3 showed a large to very large jump depending on the harness used. Artificial Analysis scored Astra just on par with its predecessor.

With the updated index, GPT-6 Astra now shows a four-point gain over its predecessor. Anthropic's Claude Fable 5.1 still leads the ranking, followed by Astra in second and Meta in third. Astra also uses fewer tokens per task than every other frontier model, according to Artificial Analysis. On cost-to-performance ratio, Anthropic, OpenAI, Meta, and Zhipu AI all share the lead.

Artificial Analysis Intelligence Index v4.2: Claude Fable 5.1 leads the ranking, with GPT-6 Astra in second, now four points ahead of its predecessor Sol. Both OpenAI models were previously tied. | Image: Artificial Analysis

The index adds two new benchmarks: AA-Briefcase for real-world knowledge work and GDP.pdf from Surge AI for PDF document analysis. GPQA-Diamond has been dropped because models have solved it. Private test data now makes up 40 percent of the weighting to make gaming harder. Artificial Analysis also fixed scoring errors across several benchmarks and tweaked its grading systems for more stable results.

Artificial Analysis says it held off on updates to keep scores stable during major model launches, but the top of the leaderboard moved so fast that an interim update was necessary. Version 5 has been in the works for eight months and will roll out in stages, it says.

Artificial Analysis

来源:The Decoder:AI News(RSS)· the-decoder.com