I think "time horizon" is becoming one of the more useful ways to talk about agent capability.
At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common.
A model can be extremely capable locally and still be unreliable over a long execution trajectory.
Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene?
maybe we are moving towards a benchmark, something like
human minutes consumed per completed task.