For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness.
In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance.
Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points.
The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.
– arxiv. org/abs/2605.23950
Title: "Stop Comparing LLM Agents Without Disclosing the Harness"