
在 IPI 基准测试中,Opus 5 将 Opus 4.8 的攻击成功率近乎缩减至三分之一。评估中表现最佳的非Claude模型仍为16.5%。施奈尔强调核心原则:我们无法完全封堵提示词注入,只能让其在统计上变得高昂。
在 Anthropic 发布的 IPI(间接提示注入)基准测试中,Opus 5 声称在 15 次尝试中攻击成功率为 2.0%(而 Opus 4.8 为 5.5%),单次尝试成功率为 0.2%(而 Opus 4.8 为 0.5%)。同系列实验室对比数据:Sonnet 5 为 5.9%(k=15),Mythos 5 为 2.6%。该基准测试中表现最佳的非 Claude 模型 Muse Spark 停滞在 16.5%——超过 Opus 5 的八倍以上。Bruce Schneier 在其 2026 年 7 月 31 日的博客中引用了这些数据,并重申其立场:在一般情况下,阻止提示注入仍然是不可能的,但特定场景下的进展显著。
除厂商公布的数字外,有两点值得关注。第一点:非 Claude 基准线 16.5% 表明,这种攻击向量的差距是真实存在的,而非微不足道——对于将代理暴露于第三方内容(邮件、文档、网页)的企业部署而言,将成功率从 16% 降至 2% 会显著改变漏洞的运营成本。第二点:Schneier 的解读——“特定场景下的进展,而非彻底解决”——是正确的态度。提示注入问题不会被完全解决,但可以使其在统计上变得代价高昂。
完整的 IPI 协议(数据集、对手、分类)及第三方可重现的评估。若缺乏这些,2% 的数字只存在于产品营销中,而非 SOC(安全运营中心)的实际数据。
本文由人工智能撰写,并经人工编辑审核。
So Opus 5 is making real progress here-hope this momentum pushes the whole industry to stop dragging its feet on security. But Schneier’s warning still rings true: better results don’t mean the fight is over.
Totally agree-trackable progress is great, but Schneier’s point stands: we need systemic change, not just incremental wins.
But a 2% injection rate still leaves a lot of room for improvement. Can we realistically expect near-zero attacks in production anytime soon?
Claude Fable 5 : de l'annonce à la mise en production