# Anthropic Claude Opus 5 在 ARC-AGI-3 基准上以 30.2% 得分大幅领先 GPT-5.6 Sol，展现全新推理行为

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-07-26 17:43
- AIHOT 分数：55
- AIHOT 链接：https://aihot.virxact.com/items/cms1mf9bq00qxro05gwqhi1m4
- 原文链接：https://the-decoder.com/anthropics-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-measure-real-intelligence

## AI 摘要

Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准上取得 30.2% 的得分，是此前 OpenAI GPT-5.6 Sol (Max) 创下的 7.8% 纪录的近四倍。

## 正文

Key Points

Anthropic's Claude Opus 5 scored 30.2 percent on the ARC-AGI-3 benchmark, nearly four times the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max).

The ARC Prize team attributes the lead to genuinely stronger logical reasoning that enables more autonomous exploration and planning in unfamiliar environments.

During testing, Opus 5 displayed behavior not previously seen from an AI model, including translating tasks into algebraic notation and independently formulating reflection equations, while also solving five previously unsolved environments.

The creators of the ARC-AGI benchmark say Anthropic's Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning.

The model scored 30.2 percent on ARC-AGI-3, making it the new leader. The previous record was 7.8 percent, set by OpenAI's GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level. That also puts it ahead of Anthropic's "Fable-class" models, which hit around 20 percent according to ARC Prize.

ARC Prize's analysis credits the lead to stronger logical reasoning, "which enables more autonomous exploration, planning, and execution across unfamiliar environments." During testing, Opus 5 also showed behavior that researchers hadn't seen from a model before. It translated tasks into algebraic notation and independently formulated reflection equations for the first time.

Claude Opus 5 leads the ARC-AGI-3 leaderboard by a wide margin, far ahead of GPT-5.6 Sol (Max) and other competitors. | Image: ARC Prize

Six of the 25 public demo environments have now been solved. The full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Both results match previous top scores, though at slightly higher costs, according to ARC Prize.

ARC-AGI-3 measures how well AI models solve new tasks they didn't encounter during training, including ones humans can usually handle with ease. The current version works like a game. The model must infer the rules of an interactive environment, plan its actions, and carry them out step by step. This tests general reasoning rather than stored knowledge.

Some AI systems may have already passed the benchmark, but they rely on extra software known as a harness. Official scores count only the language model's own performance. ARC Prize argues that future AGI systems shouldn't need outside help to solve new tasks. Opus 5 would likely score even higher if used within Claude Code.

Independent tests suggest narrower gains

Anthropic hasn't explained the gain, but targeted data labeling and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after ARC-AGI-3 and its format became public. That may have let Anthropic target the benchmark's skills and puzzle formats, though it doesn't show the company trained on the exact tasks. Annotators could have labeled reasoning traces, useful actions, failed attempts, and recovery steps from similar puzzles. Reinforcement learning could then reward exploration, planning, rule discovery, and self-correction.

Tests on Witness, Guanghan Ning's private benchmark for interactive puzzle games, point to narrower gains. Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5 while improving far less over Opus 4.8 than it did on ARC-AGI-3. It identified a conventional puzzle's hidden rules before taking any action but trailed Opus 4.8 on a game with less familiar mechanics. Ning says that pattern fits training on genre-specific data, though Witness can't identify what data Anthropic used.

Greg Kamradt, one of the researchers behind ARC-AGI-3, said the results don't rule out broader reasoning gains. A game based on familiar mechanics doesn't test adaptation to novelty, while one weak result doesn't outweigh the model's overall improvement without detailed scores for that task. Witness was also designed around ARC-AGI-3-style puzzles, so better performance could reflect real transfer rather than memorization.

Ning later clarified that Opus 5 did generalize to Witness, just far less than on ARC-AGI-3. He compared the process to the evolution of coding benchmarks. As a major target for interactive reasoning, ARC-AGI-3 will likely attract the most training effort first. Covering more edge cases could then help models generalize to a wider range of abstract reasoning tasks. Ning said coding followed a similar path, moving from saturated benchmarks such as HumanEval to frequently updated competitions and today's coding agents.

AI News Without the Hype – Curated by Humans

via X
