Rohan Paul 谈用时间跨度衡量智能体能力,并引用 OpenAI 自动化研究实习生里程碑

Rohan Paul · @rohanpaul_ai · X·2026-09-07 06:00·11小时前
AI 导读

Rohan Paul 认为「时间跨度」正成为谈论智能体能力更有用的指标,任务越长无干预成功率越低,OpenAI 数据从 15 分钟内任务的 86% 降到 64-128 小时档的约 16%。

Rohan Paul@rohanpaul_ai
59AI 编辑部评分,满分 100

Rohan Paul 谈用时间跨度衡量智能体能力,并引用 OpenAI 自动化研究实习生里程碑

2026-09-07 06:00· 11小时前
AI 导读

Rohan Paul 认为「时间跨度」正成为谈论智能体能力更有用的指标,任务越长无干预成功率越低,OpenAI 数据从 15 分钟内任务的 86% 降到 64-128 小时档的约 16%。

I think "time horizon" is becoming one of the more useful ways to talk about agent capability.

At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common.

A model can be extremely capable locally and still be unreliable over a long execution trajectory.

Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene?

maybe we are moving towards a benchmark, something like

human minutes consumed per completed task.

Rohan PaulOpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that wo...