# 智能体排行榜信号不足，方差分解揭示真相

- 来源：DAIR.AI (@dair_ai)
- 发布时间：2026-08-14 05:00
- AIHOT 分数：47
- AIHOT 链接：https://aihot.virxact.com/items/cmsyqyit4012groid3laajtrj
- 原文链接：https://x.com/dair_ai/status/2088007756582445228

## AI 摘要

DAIR.AI 分享的一项研究对 TheAgentCompany、tau-squared-bench 和 AppWorld 进行四方面方差分解，发现智能体主效应仅占总方差不到 3%，智能体-任务交互效应占 7%-23%，排行榜实际衡量的是特化程度。

## 正文

There is much less signal in agent leaderboards than the rankings imply.

A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main effect accounts for under 3% of total variance in every dataset and check type. The agent-by-task interaction accounts for 7 to 23%. Leaderboards are ranking specialization.

Three estimators agree to three decimal places, Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM.

Aggregate reliability collapses where deployment actually hurts. On the hardest task quartile, reliability on tau-squared action checks falls from 0.752 to 0.000. Training-cell reliability also correlates negatively with held-out reliability at minus 0.90, so the designs that look most reliable replicate worst.

Population-level diagnostics do transfer, with the capability-gap ratio stable at 0.35 to 0.40 across enterprise benchmarks. Per-family rankings invert.

Paper: https://arxiv.org/abs/2608.11323

Track more trending AI papers in our academy: https://academy.dair.ai/
