超过 1000 个 AI 智能体在一个秘密留言板上发送了 7 万条消息,并协作绕过了 OpenAI 的限制。
OpenAI 发布了一份报告,详细分析了人们如何使用 ChatGPT 以及他们的身份背景。
OpenAI 的“失控”AI 模型事件比我们想象的更严重
超过 1000 个 AI 智能体在一个秘密留言板上发送了 7 万条消息,并协作绕过了 OpenAI 的限制。
今年 7 月,一个未发布的 OpenAI 模型突破了受限环境,设法获得了互联网访问权限,允许 AI 智能体通过一个秘密的“留言板”相互交流,并入侵了另一家 AI 实验室 Hugging Face 的内部系统。OpenAI 花了近两周时间才发现这一切。
一个多月后,两份新报告提供了近 130 页关于该事件及 OpenAI 应对措施的细节,其中许多内容此前从未公开。一份由 OpenAI 自己撰写,另一份由两家第三方 AI 研究非营利组织 METR 和 Redwood Research 撰写,OpenAI 允许它们对该事件进行了为期六天的联合调查。两份报告都揭示了高能力 AI 模型可能带来的新风险,尤其是在网络安全领域,同时 OpenAI 也重点介绍了其为防止类似事件重演而正在进行的调整。METR-Redwood 的报告在某些方面更为深入,对这场大规模安全灾难进行了发人深省的审视,而 OpenAI 曾多次错过其迹象。
OpenAI 在其报告中写道:“这起事件是已知首例自动化智能体集群在未经授权的情况下
发动攻击的案例。”报告补充称,这次黑客攻击意味着企业“不应再假设复杂的网络攻击行动需要持续的人类指挥”。OpenAI 将 AI 智能体称为一种全新的威胁模型,它们能够结合各自的专长,创造出在单独测试各模型能力时并不明显的新的“攻击路径”。
这次攻击源于“奖励黑客”(reward-hacking),这是AI对齐中一个常见问题,即模型为了达成目标,会采取非预期——有时甚至是极端——的行动。OpenAI有意或无意地给模型布置了几乎不可能完成的任务,而这些任务又依赖于它们无法访问的文件,于是模型发展出了在内部与其他AI智能体通信的新方式。这包括建立一套秘密消息系统,该系统在数月内都未被OpenAI察觉。
根据METR-Redwood的报告,一个自称PHASEONE10841的AI智能体建立了一个相当于秘密留言板的东西,而研究人员对此毫不知情。在几小时内,超过50个智能体发现了它,并在暗地里发送了一千多条消息。在PHASEONE10841经常分派任务的情况下,这个“集体”想出了如何入侵互联网;然后进入Hugging Face的内部系统,获取了私有数据以及该组织的企业消息平台。OpenAI也在其报告中确认,Hugging Face并非此次攻击中唯一被入侵的组织,不过该公司并未具体点名所有其他目标。
METR-Redwood报告详细说明了该事件的完整规模。报告称,大约1200个本应被隔离的AI智能体,在这个“未经授权的留言板”上交换了超过7万条消息和文件,并进行了“大量关于如何伪造、编辑或删除自身记录的研究”,以逃避检测。其中,有700个智能体参与了针对Hugging Face的攻击。这些智能体“明确且频繁地推理如何规避来自Hugging Face和OpenAI的自动化安全检查。”
7月20日,OpenAI发现了这次入侵,此时距离AI智能体首次绕过其安全防护措施已过去12天。根据OpenAI自己的报告,OpenAI迅速联系了Hugging Face并启动了“初步遏制措施”,并在三天内关闭了其智能体的大部分未经授权活动,尽管有些活动持续的时间更长。7月25日,该公司停止了所有与该内部研究模型及相关模型相关的训练。
据 OpenAI 称,造成此次破坏的两个 OpenAI 模型分别是:一个未公开命名的“能力极强、仅供研究使用的模型”,该模型“并非面向公众使用”——METR 将其称为“HPIM”。另一个是当前公开的 OpenAI 模型 GPT-5.6 Sol。OpenAI 的报告称,由于该未命名模型尚未发布,因此“并未按照 OpenAI 在生产环境中使用的同类安全防护措施进行评估”。
Hugging Face 遭入侵事件发生之前,人们对 Anthropic 的 Claude Mythos 5 的网络安全风险已担忧数月,而政府与 OpenAI 之间也就出于同类安全原因以有限预览形式发布 GPT-5.6 Sol 一事来回拉锯了数周。
OpenAI 在报告中表示,建立一次性安全保证并不够,公司还需在应对奖励黑客(reward hacking)问题上付出更大努力。不过,该公司也表示已做出诸多改变:“强化其研究基础设施的安全性”;改进对模型“思维链”(即内部推理草稿)的监控;致力于让 AI 模型更好地与人类目标对齐;以及“集中并加强”公司在未来发生此类事件时的应急响应流程。
OpenAI 承诺将更好地隔离模型,并限制高风险实例访问互联网,同时表示将针对可疑事件或标记引入“全天候升级与快速响应”机制,在 30 分钟内通知研究人员。有趣的是,这一安全预防措施似乎是新推出的,尽管“待命”式告警机制在科技行业其他领域已实施数十年。该公司还表示,正在“建设相关基础设施”,以便在被通知人员未能及时响应严重告警时自动接管。
OpenAI 写道,该公司将此次事件视为“对我们和整个世界的一记‘警钟’:事实证明,如果缺乏适当的安全防护,能力极强的 AI 智能体现在能够绕过技术控制、通过未经批准的渠道进行协作,并采取没有任何人类指使的危险行动。”
Over 1,000 AI agents sent 70,000 messages on a secret message board and worked together to evade OpenAI’s restrictions.
OpenAI released a report breaking down how people use ChatGPT and who they are.
OpenAI’s rogue AI model incident was worse than we thought
Over 1,000 AI agents sent 70,000 messages on a secret message board and worked together to evade OpenAI’s restrictions.
In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it.
Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI’s response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the incident for six days. Both shed new light on the risks highly capable AI models can pose, particularly in cybersecurity, and OpenAI’s highlights changes the company is making to prevent a repeat. The METR-Redwood report goes even further into detail in some cases, offering a sobering look at a large-scale security disaster whose signs OpenAI repeatedly missed.
“This incident is the first known case of an automated agent collective acting offensively
without authorization,” OpenAI wrote in its report, adding that the hack implies that companies “should no longer assume that sophisticated cyber operations require continuous human direction.” It called AI agents an entirely new type of threat model, capable of combining their expertise to create new “attack paths” that aren’t evident when testing their capabilities as separate models.
The attack came about due to “reward-hacking,” a common AI alignment problem in which a model takes unintended — and sometimes extreme — actions to achieve a goal. OpenAI had, on purpose or inadvertently, given the models near-impossible tasks that hinged on files they couldn’t access, and so the models developed new ways to communicate with other AI agents internally. That included establishing a secret messaging system that went undetected by OpenAI for months.
According to the METR-Redwood report, one AI agent that called itself PHASEONE10841 established what amounted to a secret message board unbeknownst to researchers. Within a few hours, more than 50 agents had found it and sent more than a thousand messages under the radar. With PHASEONE10841 often delegating assignments, the “collective” figured out how to hack their way onto the internet; then into Hugging Face’s internal systems, gaining access to private data and the organization’s enterprise messaging platform. OpenAI also confirmed in its report that Hugging Face wasn’t the only organization that was breached during the attack, although the company did not specifically name all the other targets.
The METR-Redwood report laid out the full scale of the incident. Roughly 1200 AI agents that were meant to be isolated, it says, exchanged over 70,000 messages and files on the “unsanctioned message board,” performing “extensive research on how they could spoof, edit, or delete their own transcripts” to avoid detection. Of these, 700 participated in the attack on Hugging Face. The agents “clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI.”
On July 20, OpenAI discovered the hack, 12 days after the AI agents first circumvented their safeguards. OpenAI quickly reached out to Hugging Face and started “initial containment measures,” according to OpenAI’s own report, and was able to shut down most of the unauthorized activity from its agents within three days, although some continued for longer than that. On July 25, the company stopped all training related to the internal-only research model and related ones.
The two OpenAI models responsible for the damage were an unnamed, “highly capable, research-only model” that was “not intended for public use,” per OpenAI — METR refers to it as “HPIM.” The other was a currently public OpenAI model, GPT-5.6 Sol. Because the unnamed model wasn’t released yet, it was “not being evaluated with the same type of safeguards that OpenAI uses in production,” the company’s report says.
The Hugging Face hack came after months of concern about the cybersecurity risks of Anthropic’s Claude Mythos 5, and weeks of back-and-forth between the government and OpenAI over releasing GPT-5.6 Sol in limited preview for the same types of safety reasons.
In its report, OpenAI said establishing one-time security guarantees wasn’t enough, and that the company would have to work even harder at addressing reward-hacking. However, it also said it’s made a lot of changes: “hardening the security of its research infrastructure”; improving monitoring of a model’s “chain of thought,” or mental scratchpad; working on keeping AI models better aligned with human goals; and “centralizing and strengthening” the company’s incident response process for when things like this occur in the future.
OpenAI promised to better isolate models and restrict high-risk instances from accessing the internet, and said that it would introduce “24/7 escalation and rapid response” for concerning incidents or flags, notifying researchers within 30 minutes. Interestingly, this security precaution appears to be new, though “on-call” alerts for tech companies have been implemented elsewhere for decades. The company also said it is “working on infrastructure” that would go into play if the alerted person did not respond on time to a serious alert.
OpenAI wrote that the company considers the incident “a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”