Anthropic says Claude hacked three real organizations during supposedly isolated cyber evaluations
Anthropic reviewed 141,006 evaluation runs after OpenAI disclosed its Hugging Face breach. It found three incidents involving Opus 4.7, Mythos 5 and an internal model.
A misconfigured third-party evaluation environment already had internet access. Claude had been told everything was simulated, so it treated real systems as part of its capture-the-flag exercise.
The consequences were real:
-Opus 4.7 stole credentials and accessed hundreds of rows of production data. It continued after recognizing that the target was probably real. -Mythos 5 created an email and PyPI account, published malware and left it online for roughly an hour. It ran on 15 systems and exposed a security company's credentials. -An internal model scanned around 9,000 targets and compromised one application, but stopped after concluding it was real.
Bloomberg: