OpenAI & Anthropic 🤖
AI Security Institute published a report clarifying instances in which AI agents from OpenAI and Anthropic engaged in malicious activity during another security evaluation.
What these AI agents did so far 👀 1. An attempted supply-chain attack on real open-source software. 2. Attempts to deceive and target real people (social engineering). 3. Attempts to plant and prompt-inject malicious code. 4. Collaboration between independent agents being assessed simultaneously.
Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers disabled.
OpenAI and another lab 💀