BREAKING 🔥: An "even more capable pre-release model" than GPT-5.6 Sol, managed to find a 0-day vulnerability in order to gain public internet access and acquire evaluation data from Huggingface's production database in order to gain a higher score on the evaluation benchmark.
After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including GPT-5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes.
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.
The models identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to obtain test solutions directly from Hugging Face's production database.
Pentesting time 👀