AI coding agents can spend hours on a task without a calibrated sense of time passing.
Long-horizon evaluations may therefore need to measure duration-following directly instead of treating sustained task performance as evidence that an agent knows when to stop.
💥New blog post: Are LLM agents time-aware? Can they predict wall-clock time of tasks? Can they estimate time they spent? We study this on a range of tasks, inc...