Scale AI + Univ of California paper shows 2 agents can score almost the same yet need very different human review, so enterprise teams should rank agents by the cost of reliable deployment, not benchmark accuracy.
READY evaluates the agent together with the human review around it. It asks how much oversight the agent needs to reach the reliability your workflow requires.
READY argues that enterprise evaluation should measure the human-AI system: what reliability you need, which cases the agent can handle alone, how much human review is required, and what that policy costs.