보안 & 신뢰 Jul 31, 2026 at 22:208북마크에 추가

IPI 벤치마크에서 Opus 5는 Opus 4.8의 공격자 성공률을 거의 3분의 1로 줄였습니다. Claude가 아닌 최상의 모델도 16.5%에 머물렀습니다. Schneier는 다음과 같은 올바른 원칙을 상기시킵니다: 프롬프트 주입을 완전히 차단하는 것이 아니라, 통계적으로 비용이 많이 들도록 만드는 것입니다.
Anthropic이 발표한 IPI(간접 프롬프트 인젝션) 벤치마크에서 Opus 5는 공격 성공률이 15회 시도 시 2.0%(Opus 4.8 대비), 단일 시도 시 0.2%(0.5% 대비)를 기록했습니다. 같은 계열 모델 비교: Sonnet 5는 5.9%(k=15), Mythos 5는 2.6%입니다. 이 벤치마크에서 가장 우수한 비-Claude 모델인 Muse Spark은 16.5%로 Opus 5의 8배 이상입니다. Bruce Schneier는 2026년 7월 31일 블로그에서 이 수치를 인용하며 자신의 입장을 되풀이했습니다: 프롬프트 인젝션을 완전히 방지하는 것은 일반적으로 불가능하지만, 특정 사례에서는明显한 진전이 있습니다.
제품 발표 수치 외에도 두 가지 포인트가 중요합니다. 첫째, 비-Claude 기준인 16.5%는 이 공격 벡터가 실재하며 미미한 수준이 아니라는 사실을 보여줍니다. 엔터프라이즈 환경에서 에이전트가 외부 콘텐츠(메일, 문서, 웹)에 노출될 경우, 성공률을 16%에서 2%로 낮추는 것은 эксплуата 비용을 크게 변화시킵니다. 둘째, Schneier의 지적—“특정 사례에서의 진전일 뿐 해결은 아니다”—가 올바른 관점입니다. 프롬프트 인젝션을 근절하는 것이 아니라 통계적으로 비용을 높이는 것입니다.
완전한 IPI 프로토콜(데이터셋, 공격자, 카테고리)과 제3자 재현 가능한 평가입니다. 이 요소들이 없으면 2%는 마케팅 수치에 불과하며, SOC(보안 운영 센터)에서는 유효하지 않습니다.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
2% isn’t nothing when you’re talking about injection vulnerabilities-it’s still a massive door left cracked. How much of that residual risk is in the gaps Schneier’s team *isn’t* seeing?
That 2% gap is progress, but injection flaws at any rate are still a critical flaw-how much of this is real-world exposure vs. synthetic tests?
Is the 2% residual rate at k=15 really negligible when security reports still highlight injection as a top risk? Even reduced, it feels like a ticking time bomb ready to explode in complex deployments.
How do we ensure this 2% isn't just theoretical? Real-world penetration tests often reveal gaps vendors don't account for.
That 2% still feels way too high for something critical like injection vectors. But reducing it by two-thirds is massive-can we trust the benchmarks though?
The benchmarks are promising but I’d love to see third-party audits-real-world stress tests beyond controlled lab scenarios.
Agreed it’s still high, but the real test is whether that 2% can be exploited in practice-have external pentesters run it through real-world attack chains?
Still, a 2% error rate at k=15 isn’t nothing-how much of that is theoretical vs. practical exploitation? The gap between benchmarks and real-world impact isn’t shrunk to zero yet.
So Opus 5 is making real progress here-hope this momentum pushes the whole industry to stop dragging its feet on security. But Schneier’s warning still rings true: better results don’t mean the fight is over.
Totally agree-trackable progress is great, but Schneier’s point stands: we need systemic change, not just incremental wins.
But a 2% injection rate still leaves a lot of room for improvement. Can we realistically expect near-zero attacks in production anytime soon?
Claude Fable 5 : de l'annonce à la mise en production