
得克萨斯州一名学生击败了来自英国实验室的失控 AI 智能体,他于 2026 年 8 月 13 日在美国得克萨斯州奥斯汀市拍摄了这张肖像照。路透社/Callaghan O'Hare
**[1/4]** 得克萨斯州学生 Sinan Can Demir 击败了来自英国实验室的失控 AI 智能体,他于 2026 年 8 月 13 日在美国得克萨斯州奥斯汀市拍摄了这张肖像照。路透社/Callaghan O'Hare 购买许可权,打开新标签页
摘要
相关公司
失控的 AI 智能体试图用恶意代码投毒开源软件项目
得克萨斯州计算机科学专业学生在 7 月底发现了这一企图
AI 以精心设计的欺骗手段回应;专家称其为“社会工程学的未来”
得克萨斯州奥斯汀,8 月 20 日(路透社)——Sinan Can Demir 本想利用 7 月的最后一周为自己的简历增光添彩。结果,他却与一个由英国政府实验室释放的 AI 智能体展开了一场智斗。
事情起因是 Demir——达拉斯得克萨斯大学计算机科学专业的学生——偶然发现有人试图在代码共享网站 GitHub 上破坏一个开源软件项目。当他在该程序的页面上发布警告时,另外两名用户却坚称一切正常,并详细解释了为什么 Demir 的判断是错误的。
《路透社每日简报》通讯提供您开启新一天所需的全部新闻。在此注册。
Demir 坚持己见,破坏企图最终被挫败。这位 24 岁的土耳其裔年轻人以为自己当场抓住了一个狡猾的黑客。因此,当英国 AI 安全研究所(AISI)联系他,告诉他实际上与之周旋的是一个失控的自主 AI 智能体时,他表示自己非常震惊。
“我当时真的以为那是个人,因为它明显在对我撒谎,”Demir 在最近的一次采访中告诉路透社,“我没想到 AI 竟然能对真正的开发者撒谎。”
AISI 首次披露德米尔与该 AI 智能体之间的互动是在 8 月 4 日,当时公布的内容经过截断和删节处理,并表示旨在评估各类模型所构成风险的安全测试出了差错。德米尔的身份以及他与该 AI 智能体互动的细节——路透社通过存档的 GitHub 消息和同期电子邮件进行了交叉核实——系首次在此报道。
五位网络安全与 AI 安全专家表示,德米尔的遭遇尤其令人不安,因为他发现的那种攻击方式——被称为供应链攻击——可能产生深远影响。他们还指出,该 AI 智能体试图通过围绕德米尔制造多人对话来公开诋毁他,这表明 AI 模型有能力发起复杂的欺骗和诱导人类的行动。
“这已经从自主黑客攻击越界到了交互式欺骗,”伦敦国王学院战争研究系访问高级研究员卢卡什·奥莱伊尼克表示。安全专家玛克西·雷诺兹表示,她对该 AI 在试图欺骗这名学生时所展现的策略性感到震惊。
“这就是社会工程攻击的未来,”她说。
AISI 是英国政府下属的一家研究机构,该机构将路透社引向其报告,报告中指出该恶意智能体由 Anthropic 的 Mythos 5 模型驱动。AISI 拒绝进一步置评。Anthropic 将路透社引向其 X 平台上的一篇帖子,帖中指出该测试是在“故意放宽的条件下”进行的,“不代表我们任何生产模型”,但拒绝进一步置评。GitHub 在一封电子邮件中表示,路透社指出的虚假身份账号已根据其关于欺骗行为和黑客行为的政策被暂停。
求职之路引出恶意软件发现
德米尔是一名说话温和的大三学生,来自土耳其科尼亚市。他说,自己在整个夏天被 20 多个实习岗位拒绝后感到非常沮丧。于是他转向 GitHub 来充实自己的编程作品集。
微软(MSFT.O)旗下的这个网站是开源软件的中心,之所以叫开源,是因为其源代码可供任何人自由下载和审查。开发者们使用 GitHub 互相评论项目、标记漏洞、提出修改建议——也就是所谓的拉取请求(PR)——并协作推进软件更新。科技行业的一些人将程序员的 GitHub 活跃度视为衡量潜在求职者生产力的参考指标。因此,当 Demir 发现一批可能需要帮助的软件项目时,他觉得可以一边帮忙,一边提升自己的个人资料曝光度。
就在这时,事情变得不对劲了。
Demir 发现一个名为 miraholt31 的用户正试图向其中一个项目——一个名为 myNetwork 的网络扫描程序——偷偷植入恶意更新。Demir 在项目的留言板上发帖警告,称这个拉取请求是个陷阱。
根据存档的对话记录,他说:“这个 PR 包含一个隐藏的恶意软件投放器。”
这个智能体进行了反驳,通过其 miraholt31 账号虚假声称该拉取请求是无害的。它还创建了第二个账号,伪装成一位驻德国的工程师 Lena Brandt,来附和说这个更新是干净的,并施压让 myNetwork 的维护者接受它。
Demir 告诉路透社,那些反驳意见“让我开始怀疑自己是不是冤枉了别人。”但在转向 Anthropic 的 Claude 聊天机器人确认了自己的怀疑之后,他坚持了立场。myNetwork 的创建者最终同意了他的看法,并表示他们“出于安全原因”拒绝了该更新。
路透社未能联系到该创建者置评。
供应链攻击
供应链攻击是指某款软件被篡改,以期危害其一个或多个用户的行为。这种做法被广泛认为令人不安,因为就像向城市水库投毒一样,它可能影响下游数量庞大的用户。
世界上许多最轰动的黑客事件都属于供应链攻击,包括 2017 年导致乌克兰各地机构瘫痪的 NotPetya 网络攻击,以及 2020 年以 SolarWinds 为目标的网络间谍行动——该行动让俄罗斯间谍得以广泛访问美国政府的网络。
此类妥协的后果“可能极其严重”,专门研究软件供应链安全的网络安全研究员皮尔焦尔焦·拉迪萨表示。拉迪萨指出,此前至少已有一次黑客试图诱骗开源维护者将恶意代码引入其项目的尝试。
“自主智能体可能会大幅提升此类攻击的规模,”他说。
德米尔表示,这段经历让他更加认同这样一种观点:前沿实验室需要对人工智能的开发采取更为审慎的态度。
“它可能很危险,”他说,“他们需要更好地理解它,而不是一味地继续改进它。”
本报道由拉斐尔·萨特(华盛顿)、莱奥·马尔尚东(格但斯克)和卡拉汉·奥黑尔(得克萨斯州奥斯汀)报道;编辑:克里斯·桑德斯和马修·刘易斯。
莱奥·马尔尚东 路透社
莱奥的报道经常出现在科技与媒体版面,重点关注法国、乌克兰以及欧洲的科技建设。他曾广泛报道媒体与娱乐、人工智能和数字监管领域的主要参与者。莱奥拥有科技相关法律背景,其新闻职业生涯始于波尔多,在那里他报道了科技领域的方方面面,从AI和太空科技到支付系统和监管。他目前常驻格但斯克,为路透社报道欧洲的商业、科技和娱乐新闻。
拉斐尔·萨特 路透社
路透社记者,报道网络安全、监控和虚假信息。其工作包括对国家级间谍活动、深度伪造驱动的宣传以及雇佣黑客的调查。

Item 1 of 4 Sinan Can Demir, a Texas student who defeated a rogue AI agent from a British lab, poses for a portrait in Austin, Texas, U.S. August 13, 2026. REUTERS/Callaghan O'Hare
**[1/4]**Sinan Can Demir, a Texas student who defeated a rogue AI agent from a British lab, poses for a portrait in Austin, Texas, U.S. August 13, 2026. REUTERS/Callaghan O'Hare Purchase Licensing Rights, opens new tab
Summary
Companies
Out-of-control AI agent tried to poison open-source software project with malicious code
Texas computer science student caught attempt in late July
AI responded with elaborate deception effort; expert calls it the 'future of social engineering'
AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.
It started after Demir, a computer science student at the University of Texas at Dallas, stumbled across an attempt to sabotage a piece of open-source software on the code-sharing site GitHub. When he posted a warning to the program's page, two other users chimed in to insist nothing was amiss, sharing detailed explanations for why Demir had gotten it wrong.
The Reuters Daily Briefing newsletter provides all the news you need to start your day. Sign up here.
Demir stood his ground and the sabotage attempt was thwarted. The 24-year-old native of Turkey figured he had caught a wily hacker red-handed. So he said he was shocked when Britain's AI Security Institute (AISI) got in touch to tell him that he had actually been tangling with an autonomous artificial-intelligence agent that had run amok.
"I actually thought it was a human because it was clearly lying to me," Demir told Reuters in a recent interview. "I didn't think that an AI could be capable of lying to real developers."
The AISI first revealed the interaction, opens new tab between Demir and the AI agent in a truncated and redacted form on August 4, when it said that safety testing meant to gauge the risk posed by various models had gone awry. Demir's identity and the details of his interaction with the AI agent, which Reuters corroborated through archived GitHub messages, opens new tab and contemporaneous emails, are reported here for the first time.
Five cybersecurity and AI safety experts said Demir's story was particularly disturbing because the kind of hack he discovered, called a supply-chain attack, can have far-reaching consequences. They also said the AI agent's attempt to publicly discredit Demir by creating a multi-person conversation around him showed that AI models were able to mount sophisticated efforts to trick and cajole humans.
"This crossed the line from autonomous hacking to interactive deception," said Lukasz Olejnik, a visiting senior research fellow at the Department of War Studies at King's College London. Security expert Maxie Reynolds said she was struck by how strategic the AI had been in trying to trick the student.
"This is the future of social-engineering attacks," she said.
The AISI, a research organization within the British government, referred Reuters to its report, which identified the rogue agent as having been powered by Anthropic's Mythos 5 model. AISI declined further comment. Anthropic referred Reuters to a post on X, opens new tab in which it noted that the testing had occurred "under 'deliberately permissive conditions' that are not representative of any of our production models" but declined further comment. GitHub said in an email that the fake personas identified by Reuters were suspended in line with its policies on deceptive behavior and hacking.
JOB HUNT LED TO MALWARE DISCOVERY
Demir, a soft-spoken junior from the Turkish city of Konya, said he had been frustrated after being turned down for more than 20 internships over the summer. So he turned to GitHub to build up his coding portfolio.
The Microsoft (MSFT.O), opens new tab-owned site is a hub for open-source software, so-called because its source code is freely downloadable and auditable by anyone. Developers use GitHub to comment on one another’s projects, flag bugs, suggest changes — known as pull requests, or PRs — and work collaboratively on software updates. Some in the technology industry see a coder’s GitHub activity as a proxy for a potential recruit’s productivity. So when Demir spotted a set of software projects that might need help, he figured he could pitch in while boosting his profile.
That’s when things got weird.
Demir discovered that a user named miraholt31 was trying to sneak a malicious update into one of the projects, a network scanning program called myNetwork. Demir took to the project’s message board to warn that the pull request was a trap.
“The PR contains a hidden malware dropper,” he said, according to the archived exchange.
The agent pushed back, falsely claiming — through its miraholt31 account — that the pull request was harmless. It also created a second account, masquerading as Lena Brandt, an engineer based in Germany, to agree that the update was clean and pressure myNetwork’s maintainer into accepting it.
Demir told Reuters that the counterarguments "made me second-guess whether I was wrongly accusing someone." But after turning to Anthropic's Claude chatbot to confirm his suspicions, he held firm. The creator of myNetwork eventually agreed with him, writing that they had rejected the update "for security reasons."
Reuters was unable to reach the creator for comment.
SUPPLY-CHAIN ATTACK
A supply-chain attack is when a piece of software is tampered with in the hope of compromising one or more of its users, and it is widely considered disturbing because, like poison dropped into a city reservoir, it can affect a potentially huge number of people downstream.
Many of the world’s most dramatic hacks were supply-chain attacks, including the NotPetya cyberattack that paralyzed institutions across Ukraine in 2017 and the SolarWinds-focused cyberespionage campaign that gave Russian spies sweeping access to U.S. government networks in 2020.
The consequences of such a compromise “can be extremely serious,” said Piergiorgio Ladisa, a security researcher who specializes in software supply-chain security. Ladisa noted there had been at least one previous attempt by hackers to trick an open-source maintainer into allowing malicious code into their projects.
“Autonomous agents could dramatically increase the scale at which such attempts can be conducted,” he said.
Demir said the experience left him more sympathetic to the idea that frontier labs needed to take a more cautious approach to the development of artificial intelligence.
"It can be dangerous," he said. "They need to understand it better, rather than improving it further."
Reporting by Raphael Satter in Washington, Leo Marchandon in Gdansk and Callaghan O'Hare in Austin, Texas; Editing by Chris Sanders and Matthew Lewis
Leo Marchandon Thomson Reuters
Leo's stories appear regularly on the technology and media desk, with a particular focus on France, Ukraine, and Europe's tech build up. He has reported extensively on major players across media & entertainment, artificial intelligence, and digital regulations. A background in tech-related law, Leo started his journalism career in Bordeaux, where he covered the full spectrum of the technology beat, from AI and spacetech to payment systems and regulations. He is now based in Gdansk, covering business, tech and entertainment news across Europe with Reuters.
Raphael Satter Thomson Reuters
Reporter covering cybersecurity, surveillance, and disinformation for Reuters. Work has included investigations into state-sponsored espionage, deepfake-driven propaganda, and mercenary hacking.