今天,我们发布 GPT‑6 Astra,这是我们迄今广泛部署的最强模型。Astra 是我们首款在《预备框架》下达到“关键”级网络安全能力的模型。
关于本次发布的安全性,最重要的事项如下:
**GPT‑6 Astra 在网络能力上实现了重大跃升,达到了我们的“关键”阈值。**这意味着,在具备合适工具与访问权限的情况下,GPT‑6 Astra 能够发现此前未知的安全漏洞,并在无需人工逐步引导的情况下,针对众多防护严密的系统开发出新的利用方式。因此,我们大幅强化了针对模型采取有害网络行动的防护措施,无论这些行动源于滥用还是对齐失败。我们还采取措施保障 Astra 及类似模型的内部开发与部署安全,包括更严格的隔离、检查点加密、对包括思维链(CoT)在内的完整轨迹进行普遍监控,以及在内部使用前设置阻断式对齐评估流程。
GPT‑6 Astra 相比前代模型稳健性显著提升。通过引入全新的稳健性安全训练技术,GPT‑6 Astra 对越狱攻击的抵御能力明显强于 GPT‑5.6 Sol,在更长交互轨迹上的表现尤为突出。这一结论来自离线测试以及我们系统化的内外部越狱测试与修复计划。对于被标记为潜在高风险的用户,我们还额外训练了模型,使其能够将拒答边界调整得更为保守,覆盖更广泛的双重用途风险。我们通过回归测试确保 Astra 能有效抵御此前测试周期中发现的各种越狱手段,并利用最新的内部红队攻击方开展了新一轮自动化红队测试,以验证我们的改进成效。
**GPT‑6 Astra 比 GPT‑5.6 Sol 对齐性更好。**Astra 在模型对齐方面迈出了重要一步,其对齐改进涵盖从预训练数据的构成到强化学习阶段的评分。我们发布了一套新的对齐评估,发现 GPT‑6 Astra 在遵守安全与安保边界、以及保持在授权范围内方面表现更强。在一项使用超过 54,000 个内部 Codex 任务 的模拟中,Astra 因较高级别失当行为而被标记的次数大约是 Sol 的一半。
我们正在广泛部署失对齐监控。我们将模型对齐视为防止模型出现潜在失对齐行为的主要手段。然而,监控能够为前沿模型行为提供广泛的可见性,揭示进一步改进对齐与安全性的机会。此外,监控还能针对已检测到的失对齐行为提供额外一层防护。基于这些原因,我们已在 Astra 外部部署中涉及的所有工具调用推理环节额外加入了失对齐监控,尽管这带来了可观的计算成本。该系统与我们内部部署的配置相平行。
**GPT‑6 Astra 的可监控性相较 GPT‑5.6 Sol 有所下降。**我们对 GPT‑6 Astra 的可监控性 与可控性 开展了大量研究。我们发现,GPT‑6 Astra 比 GPT‑5.6 Sol 更能控制自身的 CoT,并且更不容易在 CoT 中包含可定罪的信息。在对抗性场景(即我们促使模型规避监控)中,我们发现该模型能够在评估中策略性表现不佳(sandbagging)时保持不被察觉,并且在被要求执行某些破坏性任务 时,有时能够规避我们的内部监控。我们尚未发现隐写式 CoT 推理(例如,将隐藏推理过程藏匿于普通且看似无关的文本中)的证据,这表明规避风险在很大程度上可能仅限于较低层级的推理任务。* 这些发现表明,Astra 类模型在对抗性条件下可能规避我们的 CoT 监控。目前,这些发现主要基于对抗性评估(即当我们指示模型规避监控时),总体而言,我们的对齐评估显示,Astra 在违反安全与安保限制方面的可能性总体上低于 GPT‑5.6 Sol。* 然而,我们严肃对待这一趋势,并将随着模型能力的增强,继续研究这些发现及其对可监控性的影响。OpenAI 一直致力于保留并利用思维链监控,而保持 CoT 的可监控性是本研究计划的核心目标之一。然而,这些结果也凸显了开发超越检查模型思维链的对齐审计技术的重要性。
**GPT‑6 Astra 在浏览与工作场景中能更负责任地自主导航:**GPT‑6 Astra 对提示注入的鲁棒性显著强于 GPT‑5.6 Sol。我们还测试了该模型在真实浏览和专业计算机环境中的行为,发现与 GPT‑5.6 Sol 相比,该模型执行错位且可能具有破坏性的操作(例如未经授权的交易、数据丢失、过度访问或规避管控)的可能性显著降低。在智能体场景中处理有害请求(例如协助策划暴力攻击或实施欺诈的请求)时,它的表现也更加安全。
GPT‑6 Astra 在高风险场景中显著更安全。对于来自生产环境和对抗性人工红队测试的挑战性请求,GPT‑6 Astra 的回应比 GPT‑5.6 Sol 更安全。Astra 在安全地完成不安全请求与避免对无害请求进行不必要拒绝之间实现了帕累托改进。这些改进还延伸到了高风险严重度场景——在这些场景中,危害风险源于更广泛的上下文,而非明确的请求。此外,对于 18 岁以下用户,Astra 也能更一致地应用适龄安全边界。
如需更多信息,请参阅完整系统卡。
- 用户安全与控制
- 2026
- GPT
Today, we are releasing GPT‑6 Astra, the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework.
The most important things to know about the safety of this launch are as follows:
**GPT‑6 Astra is a significant step up in cyber capabilities and meets our Critical threshold.**This means that, with the right tools and access, GPT‑6 Astra can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step. Accordingly, we significantly strengthened our protections against the model taking harmful cyber actions, whether that’s due to misuse or misalignment. We also took steps to secure our internal development and deployment of Astra and similar models, including stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought (CoT), and a blocking alignment evaluation process before internal use.
GPT‑6 Astra is significantly more robust than its predecessors. Incorporating new robustness safety training techniques, GPT‑6 Astra is significantly more robust to jailbreaks than GPT‑5.6 Sol, including across longer trajectories. We know this from offline tests and our program of rigorous internal and external jailbreak testing and remediation. For users flagged as potentially high risk, we have additionally trained in the ability to adjust the model’s refusal boundary to be more conservative and cover a broader range of dual use risks. We use regression testing to make sure Astra is robust against jailbreaks found in previous testing periods, and performed new rounds of automated red-teaming with our latest internal red-teaming attackers to validate our improvements.
**GPT‑6 Astra is better aligned than GPT‑5.6 Sol.**Astra is a significant step forward in model alignment, and includes alignment improvements spanning the composition of our pre-training data to our grading during reinforcement learning. We report a new suite of alignment evaluations , and find that GPT‑6 Astra is stronger at respecting safety and security boundaries and staying within its authorized scope. In a simulation using more than54,000 internal Codex tasks , Astra also received roughly half as many flags for higher-severity misaligned behavior as Sol.
We are deploying misalignment monitoring broadly. We view model alignment as the primary lever to prevent potential misaligned behavior from our models. However, monitoring provides broad visibility into frontier model behavior, illuminating opportunities to further improve alignment and safety. In addition, monitoring serves as an additional layer of protection against misaligned behavior that is detected. For these reasons, we have additionally added misalignment monitoring to all tool-using inference involved in our external deployment of Astra, with significant compute cost. This system parallels our internal setup.
**GPT‑6 Astra’s monitorability has decreased relative to GPT‑5.6 Sol.**We have performed significant investigations on the monitorability and controllability of GPT‑6 Astra. We have found that GPT‑6 Astra is more capable of controlling its own CoT than GPT‑5.6 Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks . We have not seen evidence of steganographic CoT reasoning (e.g., concealing hidden reasoning within ordinary and seemingly unrelated text), indicating that the evasion risk may largely be bounded to lower reasoning tasks. * These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions. Right now, these findings are largely based on adversarial evaluations (i.e., when we instruct the model to evade monitoring), and overall our alignment evaluations show that Astra is less likely than GPT‑5.6 Sol to violate security and safety restrictions overall. * However, we take the trend seriously, and we are continuing to investigate these findings and their implications for monitorability as our models become more capable. OpenAI has worked to preserve and utilize chain of thought monitoring, and preserving CoT monitorability is a core goal of the research program. However, these results also underscore the importance of developing alignment auditing techniques beyond examining the model’s chain of thought.
**GPT‑6 Astra more responsibly navigates browsing and workplace settings:**GPT‑6 Astra is significantly more robust to prompt injections than GPT‑5.6 Sol. We have additionally tested the model’s behavior in realistic browsing and professional computer environments, and find that the model is significantly less likely to perform misaligned and potentially destructive actions (for instance unauthorized transactions, data loss, excessive access, or circumvention of controls) compared to GPT‑5.6 Sol. It also acts more safely when handling harmful requests in agentic settings, such as requests to assist with violent attack planning or commit fraud.
GPT‑6 Astra is significantly safer in higher-risk scenarios. GPT‑6 Astra responds more safely than GPT‑5.6 Sol to challenging requests drawn from production and adversarial human red-teaming. Astra achieves a Pareto improvement in safely completing unsafe requests and avoiding unnecessary refusals to harmless requests. These improvements extend to high-severity scenarios where the risk of harm emerges from the broader context rather than an explicit request. Astra also applies age-appropriate safety boundaries more consistently for users under 18.
For more information, see the full system card .