The Decoder:AI News(RSS)
52AI 编辑部评分,满分 100

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 上超越 Opus 5

2026-07-30 17:03· 37分钟前· Matthias Bastian
跳到正文
AI 摘要

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 逻辑基准测试中达到 38.3% 的分数,超过 Claude Opus 5 的 30.2%。但该结果使用了自有 API 的“Retained Reasoning”和“Compaction”设置,在官方测试中仅得 7.8%。ARC Prize 联合创始人回应称通用 API 设置允许使用,但存在“潜在的公平性问题”。

Image description

Update – Jul 30, 2026

  • Added ARC Prize statements

ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.

He noted that ARC Prize has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer." Different providers using different settings does create "a potential parity issue," Chollet said, but he considers that acceptable "as long as the settings and the cost are clearly reported."

OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmarkOpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

GPT-5.6 Sol's ARC-AGI-3 scores jump dramatically when using OpenAI's custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.

视频 · 前往原文观看

OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's true, and ARC-AGI-3 is designed to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons, ARC Prize said in response to OpenAI's results. The sticking point is whether ARC Prize used an older "OpenAI-style completions API" that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI.

AI News Without the Hype – Curated by Humans

via X

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 上超越 Opus 5

The Decoder:AI News(RSS)·2026-07-30 17:03·37分钟前·Matthias Bastian
阅读原文· the-decoder.com
AI 摘要

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 逻辑基准测试中达到 38.3% 的分数,超过 Claude Opus 5 的 30.2%。但该结果使用了自有 API 的“Retained Reasoning”和“Compaction”设置,在官方测试中仅得 7.8%。ARC Prize 联合创始人回应称通用 API 设置允许使用,但存在“潜在的公平性问题”。

原文 · 保持原样,未翻译
Image description

Update – Jul 30, 2026

  • Added ARC Prize statements

ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.

He noted that ARC Prize has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer." Different providers using different settings does create "a potential parity issue," Chollet said, but he considers that acceptable "as long as the settings and the cost are clearly reported."

OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmarkOpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

GPT-5.6 Sol's ARC-AGI-3 scores jump dramatically when using OpenAI's custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.

视频 · 前往原文观看

OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's true, and ARC-AGI-3 is designed to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons, ARC Prize said in response to OpenAI's results. The sticking point is whether ARC Prize used an older "OpenAI-style completions API" that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI.

AI News Without the Hype – Curated by Humans

via X

阅读原文the-decoder.com