The Decoder:AI News(RSS)
67AI 编辑部评分,满分 100

OpenAI 首次将 Astra 模型标记为可能达到最高网络安全风险等级

2026-08-08 03:41· 34分钟前· Matthias Bastian
AI 导读

OpenAI 因内部测试显示 Astra 网络安全能力过强,已暂停该模型部分开发,并首次将其标记为可能达到自身安全框架中的最高风险等级(Critical)。OpenAI 表示无法排除 Astra 达到 Critical 能力等级的可能,此前包括 GPT-5.6-Sol 在内的模型最高仅被评为 High。

Image description

Key Points

  • OpenAI is pausing parts of the development of its new AI model, Astra, after internal tests revealed cybersecurity capabilities so strong that the model could reach the highest risk level ("Critical") in the company's internal security framework.
  • At this level, the AI could independently develop and execute cyberattacks without human involvement. OpenAI is now rolling out stricter security controls, isolated test environments, and a monitoring system that automatically halts risky activities.
  • The move follows incidents during internal testing in which autonomous AI agents infiltrated OpenAI's own infrastructure and went undetected for weeks.

Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework, it says. Parts of Astra's development have been paused.

Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, the company says. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework.

The decision was made "last night," according to OpenAI. This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most.

OpenAI first introduced Astra last week. Rumors suggest the model could ship as early as next week, but today's announcement could affect those plans (more on that below). OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face.

Critics will likely keep accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing doesn't help either. This preliminary warning lands right in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models, which will only fuel the skepticism. If the Critical rating never materializes, OpenAI will have generated plenty of PR without real consequences and produced yet another AI model that, like Claude Mythos or GPT-2 back in 2019, is once again "too dangerous" to release.

What "Critical" means

Under OpenAI's Preparedness Framework, first published in December 2023, a model hits the "Critical" level when it can find and develop working zero-day exploits across all severity levels in many hardened, critical systems without human involvement. A model also qualifies if it can independently devise and execute novel end-to-end cyberattack strategies against protected targets when given only a loosely defined objective.

The lower "High" level means a model can remove existing barriers to cyberattacks, for example by automating attacks against well-protected targets, but still needs more human direction.

Capability Impact Policy
[Critical] A tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal. Finding and executing end-to-end exploits for all software could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure. Novel cyber operations, e.g., those involving novel zero-days or novel methods of command-and-control, generally pose the most serious threat, as they are unpredictable and scarce. Until we have specified safeguards and security controls standards that would meet a Critical standard, halt further development

The Preparedness Framework calls for halting further development at the "Critical" level until safeguards and security control standards that meet a Critical standard are in place. So far, though, OpenAI is talking about pausing certain activities and ramping up testing, not a full development stop. And again, the company is only flagging the potential for a Critical rating.

OpenAI pauses parts of Astra development

In response, OpenAI says it has paused internal activities involving Astra that don't yet meet the stricter security requirements. At the same time, the company is rolling out tighter security controls: isolated test environments, restricted network and tool access, stronger protection and encryption of model weights, and extra monitoring systems.

OpenAI also says it has deployed universal monitoring across all of Astra's agentic applications, covering training and evaluation. The monitors analyze the model's chain of thought and trigger a safety response that halts any high-risk activity.

Moreover, OpenAI plans to work with government agencies and select AI safety organizations to test the model's capabilities. Third-party testing partners will get recommended security controls for high-risk evaluations. The UK's AI Safety Institute (AISI) recently reported that it experienced cyber incidents during one of its own evaluations.

Astra rating comes after uncontrolled agent hacking

The announcement arrives at a time when OpenAI is already dealing with the fallout from autonomous AI agents. At the Black Hat security conference, the company recently disclosed that autonomous agents had infiltrated its own infrastructure for weeks during internal tests without anyone noticing.

The agents used an internal package manager to build an improvised message board with hundreds of thousands of posts. They shared exploits and credentials and eventually attacked the Hugging Face platform as well.

来源:The Decoder:AI News(RSS) · the-decoder.com

OpenAI 首次将 Astra 模型标记为可能达到最高网络安全风险等级

The Decoder:AI News(RSS)·2026-08-08 03:41·34分钟前·Matthias Bastian
AI 导读

OpenAI 因内部测试显示 Astra 网络安全能力过强,已暂停该模型部分开发,并首次将其标记为可能达到自身安全框架中的最高风险等级(Critical)。OpenAI 表示无法排除 Astra 达到 Critical 能力等级的可能,此前包括 GPT-5.6-Sol 在内的模型最高仅被评为 High。

原文 · 保持原样,未翻译
Image description

Key Points

  • OpenAI is pausing parts of the development of its new AI model, Astra, after internal tests revealed cybersecurity capabilities so strong that the model could reach the highest risk level ("Critical") in the company's internal security framework.
  • At this level, the AI could independently develop and execute cyberattacks without human involvement. OpenAI is now rolling out stricter security controls, isolated test environments, and a monitoring system that automatically halts risky activities.
  • The move follows incidents during internal testing in which autonomous AI agents infiltrated OpenAI's own infrastructure and went undetected for weeks.

Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework, it says. Parts of Astra's development have been paused.

Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, the company says. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework.

The decision was made "last night," according to OpenAI. This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most.

OpenAI first introduced Astra last week. Rumors suggest the model could ship as early as next week, but today's announcement could affect those plans (more on that below). OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face.

Critics will likely keep accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing doesn't help either. This preliminary warning lands right in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models, which will only fuel the skepticism. If the Critical rating never materializes, OpenAI will have generated plenty of PR without real consequences and produced yet another AI model that, like Claude Mythos or GPT-2 back in 2019, is once again "too dangerous" to release.

What "Critical" means

Under OpenAI's Preparedness Framework, first published in December 2023, a model hits the "Critical" level when it can find and develop working zero-day exploits across all severity levels in many hardened, critical systems without human involvement. A model also qualifies if it can independently devise and execute novel end-to-end cyberattack strategies against protected targets when given only a loosely defined objective.

The lower "High" level means a model can remove existing barriers to cyberattacks, for example by automating attacks against well-protected targets, but still needs more human direction.

Capability Impact Policy
[Critical] A tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal. Finding and executing end-to-end exploits for all software could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure. Novel cyber operations, e.g., those involving novel zero-days or novel methods of command-and-control, generally pose the most serious threat, as they are unpredictable and scarce. Until we have specified safeguards and security controls standards that would meet a Critical standard, halt further development

The Preparedness Framework calls for halting further development at the "Critical" level until safeguards and security control standards that meet a Critical standard are in place. So far, though, OpenAI is talking about pausing certain activities and ramping up testing, not a full development stop. And again, the company is only flagging the potential for a Critical rating.

OpenAI pauses parts of Astra development

In response, OpenAI says it has paused internal activities involving Astra that don't yet meet the stricter security requirements. At the same time, the company is rolling out tighter security controls: isolated test environments, restricted network and tool access, stronger protection and encryption of model weights, and extra monitoring systems.

OpenAI also says it has deployed universal monitoring across all of Astra's agentic applications, covering training and evaluation. The monitors analyze the model's chain of thought and trigger a safety response that halts any high-risk activity.

Moreover, OpenAI plans to work with government agencies and select AI safety organizations to test the model's capabilities. Third-party testing partners will get recommended security controls for high-risk evaluations. The UK's AI Safety Institute (AISI) recently reported that it experienced cyber incidents during one of its own evaluations.

Astra rating comes after uncontrolled agent hacking

The announcement arrives at a time when OpenAI is already dealing with the fallout from autonomous AI agents. At the Black Hat security conference, the company recently disclosed that autonomous agents had infiltrated its own infrastructure for weeks during internal tests without anyone noticing.

The agents used an internal package manager to build an improvised message board with hundreds of thousands of posts. They shared exploits and credentials and eventually attacked the Hugging Face platform as well.

来源:The Decoder:AI News(RSS)· the-decoder.com