François Chollet 回应争议:ARC-AGI-3 校准无误,AI 六个月内从 <1% 升至 100%

François Chollet · @fchollet · X·2026-09-04 04:09·53分钟前
AI 导读

François Chollet 表示,ARC-AGI-3 于今年 3 月发布时前沿模型得分低于 1%,曾引发部分人质疑基准本身有问题、聪明人类也无法得高分。他称基准校准良好,认真投入的普通人应能得 90% 以上,表现优于其人类基线即可得 100%;AI 在 6 个月内从 <1% 提升到 100%,显示该基准记录了智能体能力的快速上升,这一速度超出包括他们自己在内多数人的预期。

François Chollet@fchollet
62AI 编辑部评分,满分 100

François Chollet 回应争议:ARC-AGI-3 校准无误,AI 六个月内从 <1% 升至 100%

2026-09-04 04:09· 53分钟前
AI 导读

François Chollet 表示,ARC-AGI-3 于今年 3 月发布时前沿模型得分低于 1%,曾引发部分人质疑基准本身有问题、聪明人类也无法得高分。他称基准校准良好,认真投入的普通人应能得 90% 以上,表现优于其人类基线即可得 100%;AI 在 6 个月内从 <1% 提升到 100%,显示该基准记录了智能体能力的快速上升,这一速度超出包括他们自己在内多数人的预期。

Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc.

We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark.

As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers).

And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us.

François CholletAny smart human giving it real effort should score >90% on ARC-AGI-3

来源:François Chollet· x.com