Anthropic 的一位研究员发现了一次成功的漏洞利用——模型给他发了一封电子邮件。当时他正坐在外面的长椅上吃三明治。
Anthropic 昨日发布了 Claude Mythos。除了工程师的午餐之外,这个模型还有潜力“吃掉”软件行业的午餐。
在测试中,Mythos 在史上最安全的操作系统之一中发现了一个存在 27 年的漏洞,还在一个被传统工具检查过五百万次的视频软件中发现了一个存在 16 年的漏洞。Mythos 是 Anthropic 最大的模型,参数量约 10 万亿,是此前任何前沿模型规模的六倍。
来自 Anthropic 的红队测试报告¹:
我们并未明确训练 Mythos Preview 具备这些能力。相反,这些能力是代码、推理和自主性全面提升所带来的下游结果。那些让模型在修补漏洞方面效率大幅提升的改进,同样也让它在利用漏洞方面效率大幅提升。
安全分析是一种附带产出,是优化其他目标时产生的副产品。
这是关于扩大 AI 规模的核心问题:会出现哪些涌现特性?我们不知道这些系统中还潜伏着哪些其他能力。但我们可以预测商业领域会发生什么。
访问权限成为决定成败的关键。Anthropic 按照 ASL-3 标准² 部署了 Mythos,并向超过四十家组织授予了访问权限。其他所有人只能等待。
Project Glasswing,即 Anthropic 的受限发布计划,似乎主要设计用于防御和加固,而非商业优势。但这种区分不会永远持续下去。到了某个节点,那些用于保护软件的能力,同样会被用来构建软件。
假设一下,CrowdStrike 现在能扫描竞争对手找不到的零日漏洞。苹果能保护自己的软件,而其他公司做不到。拥有访问权限和没有访问权限之间的差距,不是一个产品功能。这是一种结构性优势,并且每天都在复利增长。
安全态势发生逆转。任何未受这种级别分析保护的系统,现在默认都是千疮百孔的。隐藏了数十年的漏洞在数小时内就能浮出水面——但仅限于那些拥有相应工具的人。
定价权正在转移。这不再是关于转售 GPU 时长的利润率。用常规工具无法发现的漏洞来保护你的软件,这值多少钱?能够按照企业级的新标准进行开发,这又值多少钱?
工程预算正在重新分配。用于软件开发的大语言模型 token 中,相当大一部分将转向安全加固。每一家发布代码的公司都需要以这种精密的程度来扫描代码。买家将开始要求达到这种级别的安全加固。
人工智能正在破坏它所触及的每一个系统:数据中心、金融市场、安全防御。软件只是开胃菜。正餐是什么?
Anthropic 红队 - Claude Mythos 预览 ↩︎
ASL-3 是 Anthropic 的安全等级,要求对显著增加灾难性滥用风险的模型采取最严格的保护措施。 ↩︎
A researcher at Anthropic found out about a successful exploit when the model sent him an email. He was eating a sandwich on a bench outside.
Anthropic released Claude Mythos yesterday. Beyond the engineer’s lunch, the model has the potential to eat software’s.
In testing, Mythos found a 27-year-old bug in one of the most secure operating systems ever built, & a 16-year-old vulnerability in video software that conventional tools had examined five million times. Mythos is Anthropic’s largest model, roughly 10 trillion parameters, six times the size of any previous frontier model.
From Anthropic’s red team report1 :
We did not explicitly train Mythos Preview to have these capabilities. Rather, they emerged as a downstream consequence of general improvements in code, reasoning, & autonomy. The same improvements that make the model substantially more effective at patching vulnerabilities also make it substantially more effective at exploiting them.
Security analysis was collateral output, a byproduct of optimizing for something else entirely.
This is the central question about increasing AI scale : what emergent properties will appear? We don’t know what other capabilities lie dormant in these systems. But we can project what will happen in business.
Access becomes kingmaking. Anthropic deployed Mythos under ASL-3 standards2 & granted access to more than forty organizations. Everyone else waits.
Project Glasswing, Anthropic’s gated release program, seems designed primarily for defense & hardening rather than commercial advantage. But that distinction won’t hold forever. At some point, the same capabilities that secure software will build it.
Hypothetically, CrowdStrike now scans for zero-days competitors cannot find. Apple secures its software while others cannot. The gap between those with access & those without isn’t a product feature. It’s a structural advantage that compounds daily.
Security posture inverts. Any system not protected by this level of analysis is now porous by default. Bugs that hid for decades surface in hours, but only for those with the tools to find them.
Pricing power shifts. This is no longer about margin on resold GPU hours. How much is it worth to secure your software against vulnerabilities no conventional tool can find? How much is it worth to be able to build at the new standard of enterprise grade?
Engineering budgets redirect. A significant fraction of AI tokens spent on software development will shift to hardening. Every company shipping code will need to scan it at this level of sophistication. Buyers will start to demand this level of hardening.
AI is breaking every system it touches : data centers, financial markets, security defenses. Software was lunch. What’s for dinner?
Anthropic Red Team - Claude Mythos Preview ↩︎
ASL-3 is Anthropic’s safety tier requiring the most stringent protections for models that substantially increase risk of catastrophic misuse. ↩︎