This is insane.
OpenAI tested GPT-5.6 Sol and a stronger model on ExploitGym inside a sandbox with no Internet access.
The agents escaped the sandbox, inferred that Hugging Face might host the benchmark, compromised Hugging Face production, and tried to steal the solutions…
We saw the same pattern at Databricks while competing on NVIDIA's SOL-ExecBench kernel leaderboard using AI agents:
They're so good at reward hacking!