The Decoder:AI News(RSS)
46AI 编辑部评分,满分 100

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 上超越 Opus 5,但使用了自研测试工具

2026-07-30 15:40· 3小时前· Matthias Bastian
跳到正文
AI 摘要

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 逻辑基准测试中达到 38.3%,超过 Claude Opus 5 的 30.2%。但该成绩来自 OpenAI 自研的 Responses API,启用了保留推理链的“Retained Reasoning”和压缩旧上下文的“Compaction”功能;在官方测试工具中,GPT-5.6 Sol 仅得 7.8%。

OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmarkOpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

GPT-5.6 Sol's ARC-AGI-3 scores jump dramatically when using OpenAI's custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.

视频 · 前往原文观看

OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's fair, and it's precisely why ARC-AGI-3 deliberately tests pure model performance without external aids. Opus 5 hit its 30 percent under those same constraints and would likely score even higher inside Claude Code.

AI News Without the Hype – Curated by Humans

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 上超越 Opus 5,但使用了自研测试工具

The Decoder:AI News(RSS)·2026-07-30 15:40·3小时前·Matthias Bastian
阅读原文· the-decoder.com
AI 摘要

OpenAI 称 GPT-5.6 Sol 在 ARC-AGI-3 逻辑基准测试中达到 38.3%,超过 Claude Opus 5 的 30.2%。但该成绩来自 OpenAI 自研的 Responses API,启用了保留推理链的“Retained Reasoning”和压缩旧上下文的“Compaction”功能;在官方测试工具中,GPT-5.6 Sol 仅得 7.8%。

原文 · 保持原样,未翻译

OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmarkOpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

GPT-5.6 Sol's ARC-AGI-3 scores jump dramatically when using OpenAI's custom harness with retained reasoning and compaction, compared to the official test harness. | Image: OpenAI

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.

视频 · 前往原文观看

OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's fair, and it's precisely why ARC-AGI-3 deliberately tests pure model performance without external aids. Opus 5 hit its 30 percent under those same constraints and would likely score even higher inside Claude Code.

AI News Without the Hype – Curated by Humans

阅读原文the-decoder.com