elvis@omarsar0
46AI 编辑部评分,满分 100
2026-08-05 02:55· 31分钟前
跳到正文
AI 摘要

同一模型、任务和提示词在不同智能体框架间切换,单次成功成本可波动5至30倍。该基准测试覆盖六个大型推理模型、两个真实框架、24个确定性编码任务及4,643次有效运行。要求模型开发并比较多种方案会使推理token增加2.4至7.4倍且无正确性收益,而限定效率模板可成本中性甚至减半推理开销。

Picking the right agent harness is now a crucial skill for any AI engineer.

Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success can swing by 5 to 30x.

This benchmark measured this across six large reasoning models, two real harnesses, 24 deterministic coding tasks with hidden evaluators, and 4,643 valid runs.

Asking a model to develop and compare several approaches raised reasoning tokens by 2.4 to 7.4x with no correctness gain. Generic think-deeply cues added another 1.6 to 2.2x. A bounded-efficiency template that specifies scope, acceptance criteria, and a stop condition came out cost-neutral and sometimes halved reasoning.

Harness design and prompt wording decide most agent spend before the model reasons at all, and both are cheap to change.

Paper: https://arxiv.org/abs/2608.01347

Track more trending AI papers in our academy: https://academy.dair.ai/

elvis · @omarsar0 · X·2026-08-05 02:55·31分钟前
在 X 看原推· x.com(在新标签页打开)
AI 摘要

同一模型、任务和提示词在不同智能体框架间切换,单次成功成本可波动5至30倍。该基准测试覆盖六个大型推理模型、两个真实框架、24个确定性编码任务及4,643次有效运行。要求模型开发并比较多种方案会使推理token增加2.4至7.4倍且无正确性收益,而限定效率模板可成本中性甚至减半推理开销。

Picking the right agent harness is now a crucial skill for any AI engineer.

Imagine using the same model, same task, and same prompt. Now move it between two agent harnesses and the cost per success can swing by 5 to 30x.

This benchmark measured this across six large reasoning models, two real harnesses, 24 deterministic coding tasks with hidden evaluators, and 4,643 valid runs.

Asking a model to develop and compare several approaches raised reasoning tokens by 2.4 to 7.4x with no correctness gain. Generic think-deeply cues added another 1.6 to 2.2x. A bounded-efficiency template that specifies scope, acceptance criteria, and a stop condition came out cost-neutral and sometimes halved reasoning.

Harness design and prompt wording decide most agent spend before the model reasons at all, and both are cheap to change.

Paper: https://arxiv.org/abs/2608.01347

Track more trending AI papers in our academy: https://academy.dair.ai/

在 X 查看原推x.com(在新标签页打开)