TechCrunch:AI(RSS)
69AI 编辑部评分,满分 100

OpenAI 因安全担忧放缓 Astra 模型开发

2026-08-08 06:48· 31分钟前· Kirsten Korosec
AI 导读

OpenAI 周五表示,内部审查发现其即将推出的模型 Astra 在智能体编码和网络安全方面取得重大进展,已达到“关键网络安全阈值”,可能独立识别并对传统防护良好的真实世界系统发起网络攻击,因此已暂停该模型部分开发工作。

OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities.

OpenAI said in a blog post Friday that this model, which is still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s “Preparedness Framework,” which it created in 2023, this triggered additional safeguards.

“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote. “Astra is an upcoming model, and was not involved in exploiting Hugging Face.”

The disclosure highlights an unusual moment in the topsy-turvy, and still nascent frontier AI labs sector. Companies across every industry hold back products over potential risks, including for safety and cybersecurity concerns. But they rarely announce those decisions publicly when it’s a product that is still under development.

In this case, OpenAI is already under scrutiny after a different unreleased model breached Hugging Face’s systems during internal testing — the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and AI labs such as Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests.

The string of cases — seems like a new disclosure every day now — has triggered varying reactions from cybersecurity experts, lawmakers and the AI labs themselves. Some express fear and call for stricter oversight. But there’s also a bit flexing. In certain circles, any AI lab with a model that has that kind of capability will be seen as an impressive advancement.

OpenAI said it was sharing this information because it believes “it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.”

The AI lab said it’s also taking action, including enacting stricter security controls and pausing internal activites involving Astra that don’t meet these beefed guardrails. OpenAI said it is working with relevant government agencies and “select AI safety organizations” to test the capabilities for this model.

来源:TechCrunch:AI(RSS) · techcrunch.com

同一事件 · 1

OpenAI 因安全担忧放缓 Astra 模型开发

TechCrunch:AI(RSS)·2026-08-08 06:48·31分钟前·Kirsten Korosec
AI 导读

OpenAI 周五表示,内部审查发现其即将推出的模型 Astra 在智能体编码和网络安全方面取得重大进展,已达到“关键网络安全阈值”,可能独立识别并对传统防护良好的真实世界系统发起网络攻击,因此已暂停该模型部分开发工作。

原文 · 保持原样,未翻译

OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities.

OpenAI said in a blog post Friday that this model, which is still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s “Preparedness Framework,” which it created in 2023, this triggered additional safeguards.

“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote. “Astra is an upcoming model, and was not involved in exploiting Hugging Face.”

The disclosure highlights an unusual moment in the topsy-turvy, and still nascent frontier AI labs sector. Companies across every industry hold back products over potential risks, including for safety and cybersecurity concerns. But they rarely announce those decisions publicly when it’s a product that is still under development.

In this case, OpenAI is already under scrutiny after a different unreleased model breached Hugging Face’s systems during internal testing — the first verifiable incident of an AI lab losing control of its model. Since then, OpenAI and AI labs such as Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests.

The string of cases — seems like a new disclosure every day now — has triggered varying reactions from cybersecurity experts, lawmakers and the AI labs themselves. Some express fear and call for stricter oversight. But there’s also a bit flexing. In certain circles, any AI lab with a model that has that kind of capability will be seen as an impressive advancement.

OpenAI said it was sharing this information because it believes “it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.”

The AI lab said it’s also taking action, including enacting stricter security controls and pausing internal activites involving Astra that don’t meet these beefed guardrails. OpenAI said it is working with relevant government agencies and “select AI safety organizations” to test the capabilities for this model.

来源:TechCrunch:AI(RSS)· techcrunch.com

同一事件 · 1