该公司强调,在其模型入侵 Hugging Face 之后,为 GPT-6 Astra 加强了防护栏。
OpenAI 的下一个大型 AI 模型已“进入 AGI 时代”
该公司强调,在其模型入侵 Hugging Face 之后,为 GPT-6 Astra 加强了防护栏。
OpenAI 的下一个大型模型来了:GPT-6 Astra。该公司称其在网络安全、专业工作、软件工程、科学和计算机使用等领域实现了“代际能力飞跃”。正如 OpenAI 本周早些时候宣布的那样,这也是首个被认定达到 OpenAI“关键网络安全能力阈值”的模型——但该公司承诺,这不会导致其模型再次入侵竞争对手公司的内部系统。
“如果我们快进几年,回头再看,然后说,‘AGI 到底是什么时候真正被创造出来的?’我认为大约就是这个时候,而且我认为可能就是这个模型,”OpenAI 总裁 Greg Brockman 在周四的新闻发布会上表示。在随后的电话会议中,他补充道,“就我个人而言,我确实认为我们已经达到了……我认为觉得我们现在正处于 AGI 时代并非不合理。”
这一消息距离 GPT-5 的发布已过去一年多,距离上一代模型套件的最后迭代版本 GPT-5.6 的发布也已有近两个月。该模型于今日向 OpenAI 的企业网络安全客户(即有权访问其 Daybreak 平台的企业客户)推出。OpenAI 总裁 Greg Brockman 表示,在接下来的几天内,它将向所有 Plus、Pro、Business 和 Enterprise 用户开放。该模型也将通过 OpenAI API 和 AWS 提供。
OpenAI 尤其大力宣传该模型的智能体能力与编程实力,意在吸引企业客户——并与以企业和编程能力著称的 Anthropic 竞争——为其 IPO 做准备。在一份新闻稿中,该公司表示 GPT-6 Astra 能够完成多步骤的智能体任务、构建可运行的网站,并创建“精美”的文档、电子表格和演示文稿。OpenAI 还称其为公司“软件工程领域的最佳模型,在真实代码库的复杂任务上表现更强劲。”
OpenAI 还在努力修复自身形象,此前一个未发布的 AI 模型——该公司称其并非 Astra——突破了受限环境,入侵了 OpenAI 内部系统,找到了接入互联网的方法,创造了一种让 AI 智能体在不为公司所知的情况下秘密串通的方式,并侵入了 AI 实验室 Hugging Face 的系统,而这一切 OpenAI 直到 Hugging Face 自己发布博客文章后才知晓。该事件被广泛比作一起备受瞩目的空难或畅销药品的大规模召回。
对 OpenAI 而言,这从一个小角度看或许算是正面公关——在日益白热化的 AI 竞赛中展示其模型有多强大——但同时也损害了 OpenAI 在可靠性方面的声誉。OpenAI 在一份发布中特意强调,Astra 是该公司“迄今对齐程度最高的模型”,能帮助人们“在保持监督的同时委派复杂工作”。OpenAI 首席科学家 Jakub Pachocki 向记者谈及让 AI 模型与人类利益保持一致的困难,称“智能的进步并不能保证对齐的进步”,并表示监控 AI 系统正变得越来越具有挑战性。(研究人员近期对相关报告发出警报:OpenAI 允许 Astra 使用“不透明循环推理”,即让其思维链——研究人员赖以检测模型是否在欺骗人类评估者的“心理草稿纸”——变得不可读。)
该公司目前处境岌岌可危。投资者正施压要求其最终实现盈利——或者至少创造更多收入——但就在宣布 Astra 之前,该公司刚刚因 Hugging Face 遭入侵事件及其处理方式而面临强烈批评。(尽管 OpenAI 邀请了三位外部评估人员就事件经过撰写独立报告,但公司仅允许他们在报告中回答少量预先设定好的问题,且调查时间不足一周,而此次攻击整体涉及 AI 智能体长达数月的协同策划。)
本周早些时候,OpenAI 举行了一场新闻发布会,仅仅是为了宣布其推迟了 Astra 的开发,以便改进其安全工具。在发布会上,负责 OpenAI 安全流程的 Mia Glaese 提到了该公司新的“失配监控”方法,据 OpenAI 称,该方法包括对潜在问题“全天候升级和快速响应”,并在 30 分钟内通知研究人员。
这种谨慎态度尤其必要,因为存在“关键网络安全能力阈值”,这意味着 OpenAI 认为 Astra 在发现和利用安全漏洞方面具有无与伦比的能力,即使在保护极为严密的系统中也是如此,而且全程无需人工指导。与 Anthropic 针对 Mythos 级模型(该模型曾引发对网络安全风险的警报)的规则类似,OpenAI 在一份声明中表示,它将允许 Astra 向“一组经过筛选的可信防御者提供限制较少的访问权限,以支持漏洞验证、恶意软件分析和检测工程等工作。”
OpenAI 及其竞争对手近期同意让特朗普政府在其模型发布前对其进行评估,Astra 也不例外。Brockman 告诉记者:“我们与政府一起执行了标准的测试流程……他们没有反馈说‘你需要修改这个’,无论是在安全防护还是其他方面。”
OpenAI 研究训练副总裁 Aidan Clark 称,Astra 是 OpenAI 首个由先前模型在监督训练中扮演“重要角色”的模型,这体现了该公司在备受争议的递归自我改进(即 AI 系统无需人工干预即可处理自身训练、编码并创建自身更高级版本)概念上取得的进展。
“训练前沿模型过去意味着整夜随时待命、从硬件错误中恢复任务,常常要耗费大量时间调试,”克拉克在新闻发布会上表示。“到 Astra 训练接近尾声时,一天中大部分时间都能不间断推进已成为常态,即便出现问题,模型也往往只需几秒钟的停机就能再次恢复进展。”
It emphasized stronger guardrails for GPT-6 Astra after the company’s models hacked Hugging Face.
OpenAI’s next big AI model has ‘entered the AGI era’
It emphasized stronger guardrails for GPT-6 Astra after the company’s models hacked Hugging Face.
OpenAI’s next big model is here: GPT-6 Astra. The company calls it a “generational leap in capability” for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI announced earlier this week, it’s also the first model designated as meeting OpenAI’s “critical cybersecurity capability threshold” — but the company promises that won’t lead to a repeat of its models hacking a rival company’s internal systems.
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
The news comes more than a year after the release of GPT-5, and nearly two months after the release of GPT-5.6, the last iteration of the previous model suite. The model rolls out today to enterprise OpenAI’s cybersecurity customers (enterprise customers with access to its Daybreak platform). Over the next several days, OpenAI president Greg Brockman said, it will be released to all Plus, Pro, Business, and Enterprise users. It’ll also be available via the OpenAI API and AWS.
OpenAI especially touted the model’s agentic capabilities and coding prowess in a bid to attract enterprise customers — and compete with Anthropic, known for its enterprise and coding prowess — ahead of its IPO. In a release, the company said GPT-6 Astra can complete multistep agentic tasks, build working websites, and create “polished” documents, spreadsheets, and presentations. OpenAI also called it the company’s “best model for software engineering, with stronger performance on complex tasks in real codebases.”
OpenAI is also trying to rehabilitate its image after an unreleased AI model — which it says wasn’t Astra — broke out of its restricted environment, compromised internal OpenAI systems, figured out how to gain internet access, created a way for AI agents to secretly conspire without the company’s knowledge, and hacked into the systems of AI lab Hugging Face, all without OpenAI knowing about it until Hugging Face itself put out a blog post. The incident was widely compared to a high-profile plane crash or popular pharmaceutical drug recall.
For OpenAI, this could be seen as good PR in one small way — showing how powerful its models can be in an ever-intensifying AI race — but it also damaged OpenAI’s reputation as far as reliability. OpenAI made sure to say in a release that Astra is the company’s “most aligned model yet” and helps people “delegate complex work while maintaining oversight.” Jakub Pachocki, OpenAI’s chief scientist, spoke to reporters about difficulties with keeping AI models aligned with human interests, saying that “progress in intelligence does not guarantee progress in alignment,” and he said that monitoring AI systems is becoming more and more challenging. (Researchers have recently raised alarms about reports that OpenAI allows Astra to utilize “opaque recurrence,” or render its chain of thought — a “mental scratchpad” that researchers rely on to detect if a model is scheming against its human evaluators — unreadable.)
The company is in a precarious position right now. Investors are putting on the pressure for it to finally turn a profit — or, at least, generate more revenue — but it’s announcing Astra just after facing significant criticism for both the Hugging Face hack and the way the company handled it. (Although OpenAI invited three external evaluators to write their own report about what happened, the company only allowed them to answer a handful of pre-decided questions in their report and to investigate a duration of less than a week, while the attack involved months of AI agents conspiring overall.)
OpenAI held a press briefing earlier this week just to announce that it had delayed Astra’s development in order to improve its safety tooling. And during the press briefing, Mia Glaese, who leads OpenAI’s safety processes, referenced the company’s new misalignment monitoring approach, which includes “24/7 escalation and rapid response” for potential concerns, notifying researchers within 30 minutes, according to OpenAI.
This caution is particularly warranted because of the ”critical cybersecurity capability threshold,” which means OpenAI considers it incomparably good at finding and exploiting security vulnerabilities even in extremely well-protected systems, all without human guidance. Similar to Anthropic’s rules for Mythos-class models, which raised alarm bells about cybersecurity risks, OpenAI said in a release it would allow for “less restrictive access” of Astra to an “initial set of trusted defenders, supporting work such as vulnerability validation, malware analysis, and detection engineering.”
OpenAI and its competitors recently agreed to allow the Trump administration to assess their models before release, and Astra was no exception Brockman told reporters, “We did our standard testing processes together with the government … There is nothing that they came back saying, ‘You need to change this,’ as far as safeguards or anything.”
Aidan Clark, OpenAI’s VP of research training, called Astra the first OpenAI model for which previous models played a “large role” in supervising training, referencing the company’s progress towards the controversial concept of recursive self-improvement (or AI systems that handle their own training, coding, and creating advanced versions of themselves without human intervention).
“Training a frontier model used to mean waking up at all hours of the night, recovering jobs from hardware errors, often losing long periods of time to debugging,” Clark said during the press briefing. “By the end of training Astra, it was routine to go most of a day with uninterrupted progress, and when an issue did occur, the model was often progressing again after just a few seconds of downtime.”