OpenAI’s unreleased Astra model found two V8 zero-days during testing, and used them in an exploit chain with little human help.
In their new blogpost, OpenAI wrote that in separate expert assessments, Astra compromised a hardened browser, escaped its sandbox and executed commands on the host. It also chained several operating-system vulnerabilities to move from an unprivileged account to root.
OpenAI has classified Astra as “Critical” for cybersecurity, the first of its models to reach that threshold. OpenAI paused parts of Astra’s training after the Hugging Face incident, but restarted the main frontier RL run on August 28 under stricter controls.