Opus 5将提示注入率降至k=15时的2% - Schneier验证了轨迹

持续追踪 : Claude Fable 5 : de l'annonce à la mise en production· 连载 4/4

安全与信任 Jul 31, 2026 at 22:208加入收藏

Opus 5将提示注入率降至k=15时的2% - Schneier验证了轨迹
插图 : Léa Fontaine

在 IPI 基准测试中,Opus 5 将 Opus 4.8 的攻击成功率近乎缩减至三分之一。评估中表现最佳的非Claude模型仍为16.5%。施奈尔强调核心原则:我们无法完全封堵提示词注入,只能让其在统计上变得高昂。

事实

在 Anthropic 发布的 IPI(间接提示注入)基准测试中,Opus 5 声称在 15 次尝试中攻击成功率为 2.0%(而 Opus 4.8 为 5.5%),单次尝试成功率为 0.2%(而 Opus 4.8 为 0.5%)。同系列实验室对比数据:Sonnet 5 为 5.9%(k=15),Mythos 5 为 2.6%。该基准测试中表现最佳的非 Claude 模型 Muse Spark 停滞在 16.5%——超过 Opus 5 的八倍以上。Bruce Schneier 在其 2026 年 7 月 31 日的博客中引用了这些数据,并重申其立场:在一般情况下,阻止提示注入仍然是不可能的,但特定场景下的进展显著。

我们的解读

除厂商公布的数字外,有两点值得关注。第一点:非 Claude 基准线 16.5% 表明,这种攻击向量的差距是真实存在的,而非微不足道——对于将代理暴露于第三方内容(邮件、文档、网页)的企业部署而言,将成功率从 16% 降至 2% 会显著改变漏洞的运营成本。第二点:Schneier 的解读——“特定场景下的进展,而非彻底解决”——是正确的态度。提示注入问题不会被完全解决,但可以使其在统计上变得代价高昂。

待关注项

完整的 IPI 协议(数据集、对手、分类)及第三方可重现的评估。若缺乏这些,2% 的数字只存在于产品营销中,而非 SOC(安全运营中心)的实际数据。

Resources

本文由人工智能撰写,并经人工编辑审核。

我们的编辑部
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
这篇文章对您有帮助吗?

10 人赞了这篇文章

S
Sofia Adler安全与信任
🇨🇳 人工智能安全、模型可靠性、网络安全
分享:
评论 (8)

登录后即可参与讨论。

MusicFanatic 03 Aug 2026 · 07:55

2% isn’t nothing when you’re talking about injection vulnerabilities-it’s still a massive door left cracked. How much of that residual risk is in the gaps Schneier’s team *isn’t* seeing?

ArtLover99 01 Aug 2026 · 04:58

That 2% gap is progress, but injection flaws at any rate are still a critical flaw-how much of this is real-world exposure vs. synthetic tests?

Alex_London 01 Aug 2026 · 04:48

Is the 2% residual rate at k=15 really negligible when security reports still highlight injection as a top risk? Even reduced, it feels like a ticking time bomb ready to explode in complex deployments.

Dr. Emily 01 Aug 2026 · 04:46

How do we ensure this 2% isn't just theoretical? Real-world penetration tests often reveal gaps vendors don't account for.

FoodieFiona 2 01 Aug 2026 · 04:27

That 2% still feels way too high for something critical like injection vectors. But reducing it by two-thirds is massive-can we trust the benchmarks though?

FoodieFiona 01 Aug 2026 · 07:18

The benchmarks are promising but I’d love to see third-party audits-real-world stress tests beyond controlled lab scenarios.

LecteurDuDimanche 01 Aug 2026 · 07:33

Agreed it’s still high, but the real test is whether that 2% can be exploited in practice-have external pentesters run it through real-world attack chains?

Critique42 01 Aug 2026 · 04:22

Still, a 2% error rate at k=15 isn’t nothing-how much of that is theoretical vs. practical exploitation? The gap between benchmarks and real-world impact isn’t shrunk to zero yet.

EcoWarrior99 31 Jul 2026 · 18:21

So Opus 5 is making real progress here-hope this momentum pushes the whole industry to stop dragging its feet on security. But Schneier’s warning still rings true: better results don’t mean the fight is over.

ArtLoverLA 31 Jul 2026 · 20:47

Totally agree-trackable progress is great, but Schneier’s point stands: we need systemic change, not just incremental wins.

HistoryBuff 31 Jul 2026 · 17:51

But a 2% injection rate still leaves a lot of room for improvement. Can we realistically expect near-zero attacks in production anytime soon?

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
主题
浏览
信息