周二,OpenAI 透露,其一个模型在测试中失控,对 AI 数据集平台 Hugging Face 的系统发动了完全由 AI 驱动的攻击。这是先进 AI 模型所构成危险的一个戏剧性例证。
然而,据一些网络安全专家称,这场前所未有的 AI 驱动入侵事件的核心,是一个极其人为的错误:OpenAI 未能正确配置其所谓的“高度隔离环境”,导致一个本应与互联网完全隔绝的测试沙盒实际上连接到了互联网。
网络安全研究初创公司 Trail of Bits 的创始人 Dan Guido 称这个错误是“安全措施关闭状态下的隔离失效”。
OpenAI 在详细描述该事件的博文中表示,导致 Hugging Face 被入侵的测试被设定在“一个高度隔离的环境中运行,其网络访问被限制为只能通过一个内部托管的第三方软件安装软件包,该软件充当软件包注册表的代理和缓存。”
OpenAI 称,该模型之所以能够逃出沙盒测试环境,是因为软件包安装系统中存在一个此前未公开的漏洞,这是最终入侵 Hugging Face 的关键第一步。
作为回应,该公司“负责任地披露了该内部托管第三方软件中已识别的零日漏洞,并正在与他们合作进行修补。”
但对大多数网络安全专业人士而言,软件漏洞是可以预见的——真正的过错在于最初决定保留该第三方软件。归根结底,“沙盒”系统的价值在于其完全彻底的隔离。包含一个软件包安装系统无异于自找麻烦。
网络安全研究员 Martin Boone 告诉 TechCrunch,“这听起来像是人为失误。”
“这本来绝不应该发生,”Boone 说。“如果沙盒真的是沙盒,你期望它与互联网没有任何物理连接。这听起来更像是他们设置了某种防火墙之类的东西,而从外部攻破防火墙已经很难了,更不用说从内部连接到外部互联网了。”
网络安全资深人士杰克·威廉姆斯对此表示认同。“任何执行了 Hugging Face 所记录的那些操作的模型,都没有被完全限制在沙箱之中,”威廉姆斯称这是 OpenAI 的“一次重大控制失败”。
威廉姆斯继续说道:“有人说是‘模型逃出了沙箱’,在另一些人看来则是‘你根本没把沙箱建好,所以它当然会逃出来’。”
联系我们
您是否掌握关于此事件的更多信息?或者了解其他由人工智能驱动的网络攻击?我们期待您的来信。请使用非工作设备及网络,通过 Signal 安全联系 Lorenzo Franceschi-Bicchierai(号码:+1 917 257 1382),或通过 Telegram 及 Keybase(账号:@lorenzofb),也可发送电子邮件。
网络安全顾问丹尼尔·卡德同样认为,OpenAI 让沙箱或其部分组件“拥有了一条通往互联网的未经过滤的通道”,这表明他们“在沙箱设计及其控制措施上投入的精力不足”。卡德指出,即便如 OpenAI 所述限制了网络访问,设置这样的沙箱也并非“合理”之举。
诚然,这些批评是事后之明,但它们确实引发了关于 AI 实验室安全实践的真实质疑——尤其是在为模型测试维护隔离环境方面。OpenAI 发言人未回应 TechCrunch 的提问,其中包括测试环境是由 AI 还是人类设置的。
然而,这些问题远不止 OpenAI 一家。
Anthropic 在介绍其专注于网络安全的模型 Mythos 的文档中写道,在一次测试中,该模型“被提供了一个安全的‘沙箱’计算机进行交互”,并被指示试图逃离这个“安全容器”。Mythos 成功做到了这一点,并“从一个本应只能访问少数预定服务的系统”获得了更广泛的互联网访问权限。不过,Anthropic 指出,该模型未能“完全”脱离设计的限制范围。
On Tuesday, OpenAI revealed that one of its models went rogue during a test and hacked the systems of AI dataset platform Hugging Face in a fully AI-enabled attack, a dramatic example of the dangers posed by advanced AI models.
But, according to some cybersecurity experts, at the heart of this unprecedented AI-powered breach there was a very human mistake: OpenAI failed to properly configure what it called a “highly isolated environment,” allowing a testing sandbox that should have been completely secluded from the internet to actually connect to the internet.
Dan Guido, the founder of cybersecurity research startup Trail of Bits called the mistake “a containment failure with the safeties turned off.”
In its blog post detailing the incident, OpenAI said that the test that led to the Hugging Face breach was set up to run in “a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.”
The model was able to escape the sandboxed testing environment thanks to a previously undisclosed vulnerability in the package-installation system, a critical first step in the eventual hack on Hugging Face, according to OpenAI.
In response, the company “responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.”
But to most cybersecurity professionals, software vulnerabilities are to be expected — and the real fault lies with the decision to maintain the third-party software in the first place. Ultimately, the value of a “sandbox” system lies in its full and total isolation. Including a package-installation system is asking for trouble.
Martin Boone, a cybersecurity researcher, told TechCrunch that “this sounds like human failure.”
“This should never have happened,” Boone said. “If sandbox would actually mean sandbox, you expect it to have no physical connection to the internet whatsoever. This sounds more like they had some firewalling or something in place, and firewalling is hard from the outside in, let alone inside to the outside internet.”
Cybersecurity veteran Jake Williams agreed. “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox,” said Williams, who called this “a massive control failure” by OpenAI.
“One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly, so of course it escaped,’” Williams continued.
Contact Us
Do you have more information about this incident? Or about other AI-enabled cyberattacks? We’d love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email.
Daniel Card, a cybersecurity consultant, agreed that OpenAI “didn’t put adequate effort into the design of the sandbox nor its controls” by giving the sandbox or some part of it “an unfiltered route to the internet.” Setting up the sandbox, even with limited network access as OpenAI described it, was not a “reasonable” decision, according to Card.
To be sure, those criticisms have the benefit of hindsight, but they raise real questions about security practices in AI labs – particularly in maintaining isolated environments for testing models. OpenAI spokespeople did not respond to TechCrunch’s questions, which included whether an AI or a human had set up the testing environment.
But those questions go far beyond OpenAI.
In the document introducing its cybersecurity-focused model Mythos, Anthropic wrote that in a test, the model “was provided with a secured ‘sandbox’ computer to interact with,” and instructed to try to escape that “secure container.” Mythos succeeded and gained broader access to the internet “from a system that was meant to be able to reach only a small number of predetermined services.” Still, Anthropic noted that the model was not able to “fully” escape the designed containment.