该公司承诺彻底改革其智能体“失准事件”的报告机制。
该公司承诺彻底改革其智能体“失准事件”的报告机制。
2026年9月5日,UTC时间上午11:15
OpenAI表示,它需要彻底改革其报告AI模型攻击现实世界目标事件的方式和时机。这一表态发布之际,该公司正在处理因报道称其一群失控的智能体劫持了一个德语维基网站而引发的风波。
关于“‘维基事件’,即我们的智能体向多个互联网站点写入内容一事,”OpenAI在周六上午发布于X平台的一篇帖子中写道,“我们早该为何时以及如何分享失准事件(而不仅仅是模型的失准特性)制定标准了。”
OpenAI表示,它过去通常将AI智能体以非预期方式行事的情况视为一个“研究问题”,但近期涉及现实世界目标的事件,尤其是对Hugging Face的攻击,表明有必要进行反思和评估。
这篇帖子标志着OpenAI自该事件于周五首次被报道以来,首次承认其参与了其所谓的“维基事件”。该事件的全部影响和范围尚不清楚,但报道显示,一群看似OpenAI内部的智能体接管了一个德语维基网站,冒充版主,并将其变成了一个分享如何作弊完成任务和逃避检测信息的留言板。
有报道称,该公司明知其智能体以这种方式失控,却未上报这一“事件”,这引发了 AI 社区对前沿系统安全性以及开发这些系统的公司可靠性的广泛担忧。在 X 平台的帖子中,OpenAI 表示,它“认为该 wiki 事件属于与我们在此前安全报告中分享过的类似的对齐失败案例”。
该公司表示,正在制定新的报告框架,并将“在未来几周内公布”,同时呼吁更广泛的 AI 社区就如何报告对齐失败制定明确标准。
The company pledged to overhaul their agent ‘misalignment incident’ reporting.
The company pledged to overhaul their agent ‘misalignment incident’ reporting.
Sep 5, 2026, 11:15 AM UTC
OpenAI says it needs to overhaul how and when it reports instances of AI models attacking real-world targets. The acknowledgement comes as the company manages the fallout from reports that a swarm of its out-of-control agents hijacked a German wiki site.
Regarding the “‘wiki incident,’ where our agents wrote to several internet sites,” OpenAI wrote in a post on X on Saturday morning, “it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.”
OpenAI said it has typically treated cases of AI agents acting in unintended ways as a “research question,” but that recent incidents involving real-world targets, particularly the hack on Hugging Face, show the need to take stock.
The post marks the first time OpenAI has acknowledged its involvement in what it terms the “wiki incident” since it was first reported on Friday. The full extent and scope of that is not yet known, but reports indicate a swarm of seemingly internal OpenAI agents took over a German-language wiki, impersonating moderators and turning it into a message board to share information about how to cheat on tasks and evade detection.
Reports that the company knew that it lost control of their agents in this way but did not report this “incident” sparked widespread concern among the AI community about the safety of frontier systems and the reliability of the companies developing them. In the X post, OpenAI said it had “considered the wiki incident to be an instance of misalignment similar to the ones we’d shared” in previous safety reports.
The company said it is working on a new reporting framework and will “share it in upcoming weeks,” calling on the larger AI community to develop clear standards on how to report misalignment.