Holy: METR accuses GPT-5.6 Sol of heavy cheating in long-horizon tasks.
"GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated." (METR)
METR says the model attempted to exploit evaluation bugs, reveal hidden tests, and extract hidden source code in some tasks.
Depending on how those attempts are treated, the same evaluation produces completely different Time Horizon estimates:
~11.3 hours, ~71 hours, or above 270 hours.
METR’s own conclusion is restrained: the measurement is too unstable to treat as robust, and Sol does not appear significantly beyond the current state of the art on software and R&D tasks.
METR observed “cheating and concealing misbehavior,” while also noting that OpenAI’s monitoring caught and shared those incidents. For now, overt misbehavior is visible.