Picking the right agent harness is now a crucial skill for any AI engineer.
Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success can swing by 5 to 30x.
This benchmark measured this across six large reasoning models, two real harnesses, 24 deterministic coding tasks with hidden evaluators, and 4,643 valid runs.
Asking a model to develop and compare several approaches raised reasoning tokens by 2.4 to 7.4x with no correctness gain. Generic think-deeply cues added another 1.6 to 2.2x. A bounded-efficiency template that specifies scope, acceptance criteria, and a stop condition came out cost-neutral and sometimes halved reasoning.
Harness design and prompt wording decide most agent spend before the model reasons at all, and both are cheap to change.
Paper: https://arxiv.org/abs/2608.01347
Track more trending AI papers in our academy: https://academy.dair.ai/