There is much less signal in agent leaderboards than the rankings imply.
A four-facet Generalizability Theory decomposition across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent main effect accounts for under 3% of total variance in every dataset and check type. The agent-by-task interaction accounts for 7 to 23%. Leaderboards are ranking specialization.
Three estimators agree to three decimal places, Henderson Method-I, REML via lme4, and a Bayesian binomial GLMM.
Aggregate reliability collapses where deployment actually hurts. On the hardest task quartile, reliability on tau-squared action checks falls from 0.752 to 0.000. Training-cell reliability also correlates negatively with held-out reliability at minus 0.90, so the designs that look most reliable replicate worst.
Population-level diagnostics do transfer, with the capability-gap ratio stable at 0.35 to 0.40 across enterprise benchmarks. Per-family rankings invert.
Paper: https://arxiv.org/abs/2608.11323
Track more trending AI papers in our academy: https://academy.dair.ai/