智能体排行榜信号不足,方差分解揭示真相

DAIR.AI · @dair_ai · X·2026-08-14 05:00·12天前
AI 导读

DAIR.AI 分享的一项研究对 TheAgentCompany、tau-squared-bench 和 AppWorld 进行四方面方差分解,发现智能体主效应仅占总方差不到 3%,智能体-任务交互效应占 7%-23%,排行榜实际衡量的是特化程度。

DAIR.AI@dair_ai
47AI 编辑部评分,满分 100

智能体排行榜信号不足,方差分解揭示真相

2026-08-14 05:00· 12天前
AI 导读

DAIR.AI 分享的一项研究对 TheAgentCompany、tau-squared-bench 和 AppWorld 进行四方面方差分解,发现智能体主效应仅占总方差不到 3%,智能体-任务交互效应占 7%-23%,排行榜实际衡量的是特化程度。

There is much less signal in agent leaderboards than the rankings imply.

A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main effect accounts for under 3% of total variance in every dataset and check type. The agent-by-task interaction accounts for 7 to 23%. Leaderboards are ranking specialization.

Three estimators agree to three decimal places, Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM.

Aggregate reliability collapses where deployment actually hurts. On the hardest task quartile, reliability on tau-squared action checks falls from 0.752 to 0.000. Training-cell reliability also correlates negatively with held-out reliability at minus 0.90, so the designs that look most reliable replicate worst.

Population-level diagnostics do transfer, with the capability-gap ratio stable at 0.35 to 0.40 across enterprise benchmarks. Per-family rankings invert.

Paper: https://arxiv.org/abs/2608.11323

Track more trending AI papers in our academy: https://academy.dair.ai/