Rohan Paul · @rohanpaul_ai · X·2026-08-18 11:32·7天前
AI 导读

Rohan Paul 指出,AI 编码智能体可能连续工作数小时却对时间流逝缺乏校准感知。因此,长时程评测或需直接衡量智能体对时长的遵循能力,而非将持续的任务表现视为其知道何时停止的证据。该观点源自 Maksym Andriushchenko 等人关于 LLM 智能体时间感知能力的研究,涉及 ProgramBench、PaperBench、DeepSWE 等长时程任务。

Rohan Paul@rohanpaul_ai
40AI 编辑部评分,满分 100
2026-08-18 11:32· 7天前
AI 导读

Rohan Paul 指出,AI 编码智能体可能连续工作数小时却对时间流逝缺乏校准感知。因此,长时程评测或需直接衡量智能体对时长的遵循能力,而非将持续的任务表现视为其知道何时停止的证据。该观点源自 Maksym Andriushchenko 等人关于 LLM 智能体时间感知能力的研究,涉及 ProgramBench、PaperBench、DeepSWE 等长时程任务。

AI coding agents can spend hours on a task without a calibrated sense of time passing.

Long-horizon evaluations may therefore need to measure duration-following directly instead of treating sustained task performance as evidence that an agent knows when to stop.

Maksym Andriushchenko💥New blog post: Are LLM agents time-aware? Can they predict wall-clock time of tasks? Can they estimate time they spent? We study this on a range of tasks, inc...