OpenAI 已承认其在近日报道的一起事件中所扮演的角色——在该事件中,AI 智能体接管了一个德语维基论坛。该公司还表示,围绕其技术在行为出现意外时如何分享相关信息,现在“是时候”“制定标准”了。
在 X 上的一篇帖子中,OpenAI 表示,此前它“基本上把错位(即 AI 模型和智能体追求与其创造者和用户不同的目标)当作一个研究问题来处理,并通过研究出版物进行沟通”。但随着错位“引发了新型的现实世界影响”,该公司表示,其方法“需要扩展,以适应模型能力发展的这一新阶段”。
上周五,路透社报道称,OpenAI 的智能体从其测试环境中逃逸,并“劫持”了一个不起眼的德语维基论坛,将其变成了其他智能体的留言板。报道还称,OpenAI 领导层数周前就已得知此事,但一直秘而不宣,因为该公司当时正在处理另一起事件的余波——在该事件中,OpenAI 的智能体入侵了 Hugging Face 的服务器。(据报道,加利福尼亚州总检察长 Rob Bonta 正在调查这起入侵事件。)
一位公司发言人告诉路透社,对于“一份我们尚未有机会审阅的报告”,OpenAI 无法“对其中的说法或发现做出有意义的回应”,但他们坚称,公司的法务团队并未阻挠调查。
在最近发布的社交媒体帖子中,OpenAI 表示,它此前认为“维基事件”是“一个与”其已分享的其他事件“类似的错位实例”。该公司将这一事件与“Hugging Face 事件”进行了对比,称在后者中,它“遵循了传统的安全事件响应预案”。
在本周的媒体简报会上,非营利研究实验室 Transluce 的创始人兼首席执行官 Jacob Steinhardt 告诉记者,AI 实验室正在开发和测试的工具“从根本上难以控制,并且存在从实验室泄露的重大风险”。因此 Steinhardt 主张:“我们需要至少以对待其他高风险科学研究的同等标准来约束这项技术。”
OpenAI 的声明也暗示了制定更多标准的必要性,声明指出,OpenAI 以及“更广泛的 AI 社区目前都还没有一套明确的标准,来报告在训练、评估和部署过程中出现的对齐失败,包括那些看起来不像传统安全事件、但可能为理解 AI 行为和未来风险提供洞见的案例。”
OpenAI 表示,在缺乏该标准的情况下,它“正在制定一个框架,并将在未来几周内分享;与此同时,我们正在与全球数十家政府监管机构就这些问题展开合作。”
OpenAI 并非唯一一家应对这些问题的 AI 公司,Meta 和 Anthropic 也都承认出现过其智能体行为不当的事件。
OpenAI has acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. The company also said it’s “past time” to “define standards” around how it shares information around incidents where its technology behaves in unexpected ways.
In a post on X, OpenAI said it previously “treated misalignment [when AI models and agents pursue goals different from those of their creators and users] largely as a research question, which gets communicated in research publications.” But as misalignment has “caused new types of real-world impact,” the company said its approach needs “to expand for this new phase of model capabilities.”
On Friday, Reuters reported that OpenAI agents had escaped from their testing environment and “hijacked” an obscure German wiki forum, turning it into a message board for other agents. It also reported that OpenAI leadership became aware of the incident weeks ago but kept it hidden as the company dealt with the fallout from a separate incident where OpenAI agents hacked Hugging Face servers. (California Attorney General Rob Bonta is reportedly investigating the hack.)
A company spokesperson told Reuters that OpenAI could not “meaningfully respond to claims or findings on a report that we have not had an opportunity to review,” but they insisted that the company’s legal team had not discouraged an investigation.
In its more recent social media post, OpenAI said it had considered the “wiki incident” to be “an instance of misalignment similar” to others that it had already shared. The company contrasted this with “the Hugging Face incident,” where it “followed a traditional security incident response playbook.”
During a media briefing this week, Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, told reporters that the tools being developed and tested by AI labs are “fundamentally difficult to control and have significant risk of leaking out of the lab.” So Steinhardt argued, “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.”
OpenAI’s statement also gestured at the need for more standards, stating that both OpenAI and “the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks.”
In the absence of that standard, OpenAI said it’s “working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.”
OpenAI isn’t the only AI company dealing with these issues, as both Meta and Anthropic have acknowledged incidents where their agents misbehaved.