# 论文《Stop Comparing LLM Agents Without Disclosing the Harness》：harness 对长程智能体评测的影响可能超过模型本身

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-09-01 10:42
- AIHOT 分数：64
- AIHOT 链接：https://aihot.virxact.com/items/cmti2ia5c08iyrofqlgwzhjmt
- 原文链接：https://x.com/rohanpaul_ai/status/2094616734435336462

## AI 摘要

一篇 arXiv 论文（arxiv.org/abs/2605.23950）提出，长程智能体评测中 harness 的影响可能大于模型本身，比较基准分数时应披露或控制 harness。

## 正文

For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness.

In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance.

Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points.

The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.

– arxiv. org/abs/2605.23950

Title: "Stop Comparing LLM Agents Without Disclosing the Harness"
