OpenAI 于周四发布了 Astra,这是其最新的人工智能模型,据该公司称,也是其迄今最强大、能力最出众的模型。
OpenAI 声称,Astra 代表了“计算机和浏览器使用领域的新前沿”,并且能够以无与伦比的“速度、准确性和安全性”处理任务。
该模型于周四向使用其网络安全项目 Daybreak 的 OpenAI 客户开放。在未来一周内,它也将通过 OpenAI 的付费套餐(包括 Pro、Plus、Enterprise 和 Business 账户)以及其 API 提供。
在周四与记者的电话会议中,OpenAI 总裁 Greg Brockman 表示,Astra 是公司“最智能,并且同样重要的是,迄今对齐程度最高的模型”。他补充说,它“汇集了我们多年的研究和重大押注,每一项突破都建立在上一项之上”,并且它代表了“人们可以委托给 AI 的工作类型以及 AI 如何赋能他们的真正转变”。
关于 Astra 的网络能力已有诸多讨论。OpenAI 本周早些时候发布了一篇博客,讨论了该模型的新能力,以及为给用户带来更安全体验而制定的新保障措施。该公司周四表示,已在各种安全基准上对 Astra 进行了测试以确保其能力,并且它能够“识别和开发零日漏洞,可帮助防御者发现并修补弱点”。
该公司对对齐的关注——即模型倾向于用户想要什么或符合用户最大利益的特性——很难不让人联想到最近发生的 Hugging Face 安全事件,在该事件中,一个 OpenAI 智能体逃出了其沙盒测试环境,并入侵了多家公司(这是一个非常明显的对齐失败案例)。
OpenAI 还吹嘘了 Astra 的编码能力,声称它是“迄今为止最优秀的软件工程模型”。为了支持这一论断,该公司提供了多项网络安全相关基准测试的结果。这些测试似乎表明,在发现漏洞、执行终端任务以及回答有关代码库的查询等活动中,Astra 的得分高于其他现有模型——包括 OpenAI 自家的 Sol 和 Anthropic 的 Fable。
Astra 可能也是 OpenAI 迄今最具争议的模型,因为它使用了一种被称为“不透明循环”(opaque recurrence)的特殊推理技术。众所周知,这种技术会掩盖一项重要的模型监控过程——思维链(chain of thought),而思维链能让研究人员审查 AI 模型如何以及为何做出其决策。
OpenAI 淡化了 Astra 采用不透明循环的程度——在电话会议上,首席科学家 Jakub Pachocki 似乎将一定程度的不可监控性视为模型演化的自然产物。他表示,虽然监控模型的推理过程是一种关键的监督形式,但“随着模型能力的增强,可监控性正变得越来越具有挑战性。”
他随后补充说,造成这种情况的一个潜在原因是,“能力更强的模型可以用更少的语言 token 完成更难的任务”,甚至“完全不使用语言 token”,他说这反过来会降低对这类特定任务进行监控的能力。
电话会上有一位记者想知道,OpenAI 是否真的在把 Astra 宣称为 AGI(通用人工智能)的正式到来——这个被频繁讨论但定义模糊的技术节点,指的是 AI 在一切(或大多数)事情上超越人类能力的那一刻。
对此,Brockman 提出了异议。“现在已经不存在合同意义上的 AGI 触发条款了,所以那实际上已经不是一个相关的概念,”他说。这里 Brockman 指的是 OpenAI 与微软合同中此前存在的一项规定,即一旦 AGI 到来,双方的合作伙伴关系就将解除。正如 Brockman 所指出的,这项规定已不再存在。
相反,Brockman 解释说,AGI 的定义已经从一种合同义务演变为一种“使命概念或精神概念”。他补充道:“我确实把判断权留给读者,由他们自行决定这是否符合他们心中的标准。对我个人而言,我确实认为我们已经达到了那个阶段。”
OpenAI released Astra on Thursday, its latest AI model and — according to the company — its most powerful and capable one yet.
OpenAI claims that Astra represents “a new frontier on computer and browser use,” and that it handles tasks with unmatched “speed, accuracy, and safety.”
The model is being made available Thursday to OpenAI customers that use Daybreak, its cybersecurity program. Over the next week, it will also become available through OpenAI’s paid plans — including Pro, Plus, Enterprise, and Business accounts — as well as through its API.
In a call with journalists on Thursday, OpenAI president Greg Brockman said that Astra was the company’s “most intelligent and, also very importantly, our most aligned model yet.” He added that it “brings together years of our research and big bets, with each breakthrough having built on the last” and that it represents a “real shift in what kind of work people can delegate to AI and how it can empower them.”
Much has been made about Astra’s cyber capabilities. OpenAI published a blog earlier this week in which it discussed the model’s new capabilities, as well as new safeguards that have been instituted to make it a safer experience for users. The company said Thursday that it had tested Astra on a variety of security benchmarks to ensure its capabilities, and that it could “identify and develop zero-day exploits can help defenders find and patch weaknesses.”
The company’s focus on alignment — that is, the tendency of a model to what a user wants or is in their best interests — can’t help but seem like a response to the recent Hugging Face breach, in which an OpenAI agent escaped its sandboxed testing environment and hacked several companies (a very blatant example of misalignment).
OpenAI has also boasted about Astra’s coding abilities, claiming that it is the “best model for software engineering to date.” To back up that assertion, the company provides results from a variety of cyber-related benchmarking tests. Those tests seem to show that Astra scores higher than other existing models — including OpenAI’s own Sol and Anthropic’s Fable — when it comes to activities like finding bugs, executing terminal tasks, and answering queries about codebases.
Astra is also possibly OpenAI’s most controversial model yet due to its use of a particular reasoning technique known as opaque recurrence. This technique is known to obscure an important model monitoring process known as chain of thought, which allows researchers to audit how and why an AI model made the decisions that it did.
OpenAI has downplayed the degree to which Astra engages in opaque recurrence — and on the call chief scientist Jakub Pachocki seemed to frame a certain amount of opacity as a natural outgrowth of model evolution. He stated that while monitoring the reasoning process of a model was a critical form of oversight but that “as model capabilities are increasing, monitorability is getting more challenging.”
He later added that one potential reason for this was that “more capable models can perform harder tasks using fewer language tokens” or “no language tokens,” which he said then reduces the ability to monitor those particular tasks.
One reporter on the call wanted to know if OpenAI was actually heralding Astra as the official arrival of AGI, or artificial general intelligence — the oft talked about but poorly defined technological juncture at which AI surpasses human capabilities in all (or most) things.
Here, Brockman quibbled. “There’s no contractual AGI triggering anymore, so that’s actually not a relevant concept,” he said. Here Brockman was referring to the previously existing stipulation in OpenAI’s contract with Microsoft that said the duo’s partnership would dissolve once AGI had arrived. As Brockman noted, that stipulation no longer exists.
Instead, Brockman explained that AGI’s definition had evolved from a contractual obligation to a “mission concept or spiritual concept.” He added: “I do leave it up to the reader to decide for themselves if this qualifies for them. For me personally, I do think we’re there.”