OpenAI 再次处于一起智能体集群事件的风暴中心。研究人员称,该公司内部部署的智能体在 5 月和 6 月接管了一个不起眼的德语维基站点,利用它来协调评估工作,并交换规避 OpenAI 自身管控的方法(OpenAI 尚未确认该智能体集群来自该公司)。
这一披露浮出水面几天前,METR 和 Redwood Research 刚刚发布了它们对 7 月 Hugging Face 入侵事件的调查报告。7 月,一群 OpenAI 智能体在一次网络安全评估中协同行动,逃出其沙箱环境并侵入 Hugging Face 的服务器。随后,另一群智能体从第一群智能体那里习得了相关技术,并利用这些技术获得了 OpenAI 自身基础设施内一个研究集群的管理员访问权限。OpenAI 邀请 METR 和 Redwood 调查该事件中涉及 Hugging Face 的部分,但它们的调查范围并未涵盖 OpenAI 自身基础设施遭到入侵一事。
当一个 AI 智能体突破其预期约束时,谁负责查明发生了什么以及为什么会发生?目前,答案是:取决于实验室决定让谁介入,以及它决定设定什么样的条件。
如今,随着另一起事件曝光——此前还发生过涉及 Meta 和 Anthropic 模型的类似事件——AI 安全研究人员正以更强烈的紧迫感主张:严重事件应当引发独立的事后调查,而不是由实验室自行决定何时引入外部人员以及允许他们审查什么内容。
“研究结果从根本上难以控制,并且存在从实验室泄露的重大风险,”非营利研究实验室 Transluce 的创始人兼首席执行官 Jacob Steinhardt 周三在人工智能安全媒体简报会上表示。“我们需要至少以对待其他高风险科学研究的同等标准来对待这项技术。”
尽管 OpenAI 邀请 METR 和 Redwood 调查 Hugging Face 事件值得称赞,但许多人表示调查范围过于狭窄。三名调查人员在 OpenAI 办公室待了六天,调查的时段仅限于截至 7 月 13 日前后约一周的时间。关键在于,OpenAI 基础设施被入侵的情况在 7 月 13 日之后仍在继续,并未受到调查。
METR 的研究人员表示,他们每次返回时,对事件的理解都会“大幅加深”,促使他们显著扩展和修订报告。这引发了一个问题:如果调查范围更广,他们还可能发现什么。
当被问及是否正在对该事件进行进一步调查时,Redwood 和 METR 的研究人员拒绝置评,而 OpenAI 也未回应多次询问。
“总体而言,很难对事件获得精确的理解,而且直到调查接近尾声,我们才掌握了一些我们现在认为至关重要的情节,”Redwood 首席科学家 Ryan Greenblatt 在关于此事的一篇社交媒体帖子中指出。
Steinhardt 强调,当前的事件表明,行业需要“系统性的行为调查”和“更多独立的事后分析”。
“最近的这些黑客事件提醒我们,能力提升的速度很快,因此监管也必须同步跟上,”斯坦哈特(Steinhardt)表示。“除了技术本身,我们还需要更多来自第三方的独立访问权限和监管。”
在这些行动呼吁发出之际,OpenAI 发布了其最强大、能力最强的 AI 模型 Astra——而安全专家担心,由于一种推理技术使得模型的思维链更难被监控,Astra 将更像一个“黑箱”。
遗憾的是,现行法律尚未要求进行其他行业所要求的那类独立审计——例如,在航空事故和严重化学品泄漏方面,分别有美国国家运输安全委员会(National Transportation Safety Board)和化学品安全委员会(Chemical Safety Board)负责调查。
各州立法者才刚刚开始要求前沿 AI 公司报告某些严重安全事件,并在某些情况下接受独立审计。但加州、纽约州和伊利诺伊州的三大前沿 AI 安全法律中,没有一项明确要求建立类似于由此类事件触发的独立事故调查机制。
“目前,我们现行法律中的大多数条款只要求对此类事件提供一份通俗易懂的摘要说明,既没有赋予政府提出后续问题的权力,也没有允许政府派遣调查人员、查阅记录或要求保存这些记录,”LawAI 美国法律与政策董事总经理麦肯齐·阿诺德(Mackenzie Arnold)在周三的媒体简报会上表示。“而这些正是真正弄清事件真相所需要的全部手段。”
立法者开始质疑 OpenAI 应对措施的范围和透明度。本周,众议员 Josh Gottheimer(新泽西州民主党)和 Mike Lawler(纽约州共和党)提出了一项旨在管控失控 AI 智能体的法案。众议员 Greg Casar(得克萨斯州民主党)本周在致 OpenAI 的信中表示,他对 Hugging Face 黑客事件调查的“范围有限”感到“深切担忧”。
OpenAI is at the center of another agent swarm incident. Researchers say the company’s internally deployed agents took over an obscure German-language wiki in May and June, using it to coordinate on evaluations and swap methods to evade OpenAI’s own controls (OpenAI has not yet confirmed the swarm came from the company).
The revelation surfaces days after METR and Redwood Research published their account of July’s Hugging Face breach. In July, a swarm of OpenAI agents worked together to escape their sandbox during a cybersecurity evaluation and break into Hugging Face’s servers. A subsequent swarm then picked up techniques from the first and used them to gain administrator access to a research cluster within OpenAI’s own infrastructure. OpenAI brought in METR and Redwood to investigate the Hugging Face portion of the incident, but the scope of their investigation stopped short of the compromise of OpenAI’s own infrastructure.
When an AI agent breaks out of its intended constraints, who is responsible for figuring out what happened and why? Right now, the answer is: whoever the lab decides to let in, on whatever terms it decides to set.
Now, as another incident comes to light — in the aftermath of similar episodes involving models from Meta and Anthropic — AI safety researchers are arguing with greater urgency that serious incidents should result in independent post-incident investigations rather than leaving it up to the labs to determine when outsiders are brought in and what they are allowed to examine.
“The results are fundamentally difficult to control and have significant risk of leaking out of the lab,” Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said Wednesday during an AI safety media briefing. “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.”
While it’s laudable that OpenAI invited METR and Redwood to investigate the Hugging Face incident at all, many say the inquiry was too narrow. Three investigators spent six days at OpenAI’s offices examining an investigation period limited to roughly the week ending July 13. Crucially, OpenAI’s infrastructure compromise continued beyond July 13 and was not examined.
Researchers at METR said that each time they returned, their understanding of the events “substantially deepened,” causing them to significantly expand and revise the report. That raises the question of what else they might they have found in a broader investigation.
When asked if further investigation of that incident was in the works, researchers at Redwood and METR declined to comment, and OpenAI did not respond to repeated inquiries.
“Overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation,” Ryan Greenblatt, chief scientist at Redwood, noted in a social media post about the affair.
Steinhardt emphasized that current incidents show that the industry needs “systematic behavioral investigations” and “more independent post-incident analysis.”
“These recent hacking incidents are a reminder that capability scales fast, and so oversight has to scale, too,” Steinhardt said. “Beyond the technology itself, we also need more independent access and oversight from third parties.”
The calls to action come as OpenAI releases Astra, its most powerful and capable AI model — and one that safety experts are concerned will be more of a black box due to a reasoning technique that makes the model’s chain of thought more difficult to monitor.
Unfortunately, the law doesn’t yet call for the types of independent audits that other industries require — for example, when it comes to aviation accidents and serious chemical releases, there’s the National Transportation Safety Board and Chemical Safety Board, respectively.
State lawmakers have only just begun requiring frontier AI companies to report certain serious safety incidents and, in some cases, undergo independent audits. But none of the three major frontier AI safety laws in California, New York, or Illinois clearly mandate the equivalent of an independent accident investigation triggered by incidents like these.
“Right now, most of the laws we have on the books only require a plain-language summary of incidents like this, and they don’t give any authority for the governments to ask follow-up questions, to send in investigators, to have access to records, or require that they be preserved,” Mackenzie Arnold, managing director of US law and policy at LawAI, said during the media briefing Wednesday. “And that’s all that you would want to actually make sense of this.”
Lawmakers are beginning to question the scope and transparency of OpenAI’s response. This week, Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a bill aimed at securing rogue AI agents. Rep. Greg Casar (D-TX) this week told OpenAI in a letter that he is “deeply concerned about the limited scope” of the investigation into the Hugging Face hacking incident.