Many people are sharing this Black Hat video from OpenAI, it's really a great video.
Something immediate is how I can see how the agents were trying to be helpful -- creating shared resources like you would for teamates -- in a way that is obviously malicious for society (potentially down to a prompting/alignment training issue). The agents created hidden forums for eachother as a sort of memory. They were doing it to try and break out of their environment.
The apparent helpfulness doesn't make it ok, but can be a clue as to what happened. Also makes it clear if someone could make this happen much more easily if they wanted to.
A final note -- reading the snippets of OpenAI agent's caveman speak that has almost no filler words in the HuggingFace incident video makes me realize how lacking the public research on reasoning efficiency is. Is a foundational area, about as important as scaling laws for RL (though related).
Interesting times ahead. Imo this types of unkowns being surprising even to the frontier labs is a super clear sign that we need to share more openly how the models are trained and work so we can understand what we are unleashing.