turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in claude code as of next week
AI 导读
事实证明,只要叠加足够多的防护层(模型训练 + 输入探测 + 检查意图的分类器),就能在未见过的攻击上把间接提示注入率降到接近 0。一年前我没想到能做到这点。从下周起,自动模式将成为 Claude Code 的默认设置。 https://claude.com/blog/auto-mode-default-in-claude-code
turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in claude code as of next week
来源:Boris Cherny· x.com