# Anthropic 称 Opus 5 几乎免疫浏览器提示词注入，攻击成功率降至零

- 来源：The Decoder：AI News（RSS）
- 作者：Matthias Bastian
- 发布时间：2026-07-25 18:43
- AIHOT 分数：60
- AIHOT 链接：https://aihot.virxact.com/items/cms0942e500corodza9v7s7sb
- 原文链接：https://the-decoder.com/opus-5-may-have-solved-browser-based-prompt-injection-the-biggest-security-flaw-haunting-ai-agents

## AI 摘要

Anthropic 称其 Opus 5 模型在自家软件中几乎免疫提示词注入攻击。在 129 个浏览器智能体测试场景中，攻击成功率为零；在 Gray Swan 通用提示词注入测试中，15 次尝试后的成功率从 Opus 4.8 的 5.5% 降至 2.0%。零成功率仅在 Claude Cowork 等产品开启 Auto Mode 时实现，该模式叠加了输入扫描与执行拦截两层防御。

## 正文

Anthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved. In a general prompt injection test by security firm Gray Swan, the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent.

Opus 5 leads the Gray Swan IPI benchmark. After 15 attempts, the attacker success rate sits at 2.0 percent, followed by Mythos 5 (2.6 percent) and Fable 5 (2.8 percent). | Image: Anthropic

That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork. Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero.

AI News Without the Hype – Curated by Humans

System Card
