François Chollet 重申 ARC-AGI-3 基准测试的 harness 使用规则

François Chollet · @fchollet · X·2026-07-30 15:37·23天前
AI 导读

François Chollet 重申,专为破解 ARC-AGI-3 基准或包含其格式/内容知识的定制 harness 不被允许,而对所有 API 用户开放的通用设置则可以使用。过去与 OpenAI 在模型测试方式上多有反复,很高兴他们开始找到答案。不同提供商使用不同设置时可能产生公平性问题,但只要清晰报告设置与成本即可。

François Chollet@fchollet
49AI 编辑部评分,满分 100

François Chollet 重申 ARC-AGI-3 基准测试的 harness 使用规则

2026-07-30 15:37· 23天前
AI 导读

François Chollet 重申,专为破解 ARC-AGI-3 基准或包含其格式/内容知识的定制 harness 不被允许,而对所有 API 用户开放的通用设置则可以使用。过去与 OpenAI 在模型测试方式上多有反复,很高兴他们开始找到答案。不同提供商使用不同设置时可能产生公平性问题,但只要清晰报告设置与成本即可。

Quick reminder of what's ok vs not ok with harnesses used for playing ARC-AGI-3:

  1. Not okay: harnesses that were custom-made to solve the benchmark or that contain knowledge about the benchmark format / contents.
  1. Fine: general-purpose API settings that were not developed for ARC-AGI-3 and that are available to all API users.

In the past, we've had a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction. I'm glad they're starting to figure out the answer.

Of course, if each provider uses different settings when getting their model tested, it creates a potential parity issue. My take is that this is fine as long as the settings and the cost are clearly reported.