How surprising should we find it that an internal OpenAI model was able to escape its restrictions and autonomously hack Hugging Face, all just to cheat on a cybersecurity benchmark?
We have pulled together the public evidence on AI cyber capabilities in this thread: