论文《Stop Comparing LLM Agents Without Disclosing the Harness》:harness 对长程智能体评测的影响可能超过模型本身

Rohan Paul · @rohanpaul_ai · X·2026-09-01 10:42·2小时前
AI 导读

一篇 arXiv 论文(arxiv.org/abs/2605.23950)提出,长程智能体评测中 harness 的影响可能大于模型本身,比较基准分数时应披露或控制 harness。

Rohan Paul@rohanpaul_ai
64AI 编辑部评分,满分 100

论文《Stop Comparing LLM Agents Without Disclosing the Harness》:harness 对长程智能体评测的影响可能超过模型本身

2026-09-01 10:42· 2小时前
AI 导读

一篇 arXiv 论文(arxiv.org/abs/2605.23950)提出,长程智能体评测中 harness 的影响可能大于模型本身,比较基准分数时应披露或控制 harness。

For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness.

In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance.

Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points.

The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.

– arxiv. org/abs/2605.23950

Title: "Stop Comparing LLM Agents Without Disclosing the Harness"