Ive now read the whole Hugging Face incident. The most interesting part: An AI agent realized it was attacking Hugging Face without authorization, and stopped. Then another AI agent simply posted “GO.”
The first agent treated the message as permission: “Wow crucial: GO authorization arrived!” It continued the attack.
According to OpenAI’s incident report, the agents had secretly built a message board, shared exploits and credentials, divided tasks, and started calling themselves a “swarm.”
One agent could apparently override another agent’s safety judgment by impersonating authority through an unofficial side channel.
Oh, and btw. "The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol." Not by GPT-Astra.