메뉴
BL
The Decoder 3일 전

오퍼스 5, AI 에이전트 최대 취약점인 브라우저 프롬프트 인젝션 해결

IMP
9/10
핵심 요약

Anthropic의 신모델 Opus 5가 자체 보안 소프트웨어와 결합했을 때, AI 에이전트의 가장 심각한 보안 취약점으로 꼽히는 '프롬프트 인젝션(Prompt Injection)' 공격을 0%의 성공률로 사실상 완벽하게 차단했습니다. 모델 자체의 방어력도 크게 향상되었지만, 완벽한 차단은 데이터를 사전에 검사하고 위험 명령을 차단하는 이중 보안 레이어(Auto Mode)가 활성화되었을 때 달성됩니다. 이는 프롬프트 인젝션을 완전히 막는 것은 불가능할 수 있다고 언급했던 기존 업계의 우려를 뒤집는 중요한 보안적 성과입니다.

번역된 본문

오퍼스(Opus) 5, 브라우저 기반 프롬프트 인젝션(Prompt Injection) 해결... AI 에이전트를 괴롭히는 최대 보안 결함 작성자: 마티아스 바스티안 (Matthias Bastian) 2026년 7월 25일

Anthropic은 Opus 5가 자사 소프트웨어 내에서 프롬프트 인젝션에 거의 면역이라고 밝혔습니다. 공격자가 웹페이지의 숨겨진 텍스트와 같이 조작된 입력을 통해 AI 모델의 지시를 우회하는 방식인 프롬프트 인젝션은 Opus 5를 상대로 거의 모든 경우에서 실패했습니다. 시스템 카드(System Card)에 따르면, 브라우저 에이전트의 경우 129개의 테스트 시나리오 전체에서 공격 성공률이 0%에 도달했습니다. OpenAI가 지난 12월에 프롬프트 인젝션은 결코 완전히 해결되지 않을 수 있다고 인정했던 점을 고려하면, 이는 매우 중요한 의미를 갖습니다.

보안 업체인 그레이 스완(Gray Swan)의 일반적인 프롬프트 인젝션 테스트에서는 15번의 시도 후 공격 성공률이 5.5%(Opus 4.8 기준)에서 2.0%로 크게 떨어졌습니다. 다만, 0%의 성공률은 Claude Cowork와 같은 제품에서 '오토 모드(Auto Mode)'를 켰을 때만 유지됩니다. 오토 모드는 두 가지 방어 계층을 겹쳐서 사용합니다. 첫 번째 계층은 모델이 처리하기 전에 들어오는 데이터를 스캔하여 숨겨진 지시를 찾아냅니다. 두 번째 계층은 실행 전에 위험한 동작을 차단합니다. 공격자가 이 두 가지 방어벽을 독립적으로 모두 통과해야만 공격이 성립됩니다.

이러한 보호 장치가 없다면 Opus 5의 공격 성공률은 3.7%에 머물며, Sonnet 5는 0.93%로 실제로 더 나은 성능을 보여줍니다. 결국 모델 자체의 성능과 보호 소프트웨어의 결합이 결합되어야만 공격 성공률을 0%로 낮출 수 있습니다.

원문 보기
원문 보기 (영어)
Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 25, 2026 Anthropic says Opus 5 is nearly immune to prompt injections in its own software. Prompt injection , where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card . That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved . In a general prompt injection test by security firm Gray Swan , the success rate after 15 attempts dropped from 5.5 percent (Opus 4.8) to 2.0 percent. That zero percent rate only holds with Auto Mode turned on in products like Claude Cowork . Auto Mode stacks two defense layers. One scans incoming data for hidden instructions before the model processes them. The other blocks dangerous actions before execution. An attacker has to beat both independently. Without them, Opus 5 sits at 3.7 percent, and Sonnet 5 actually does better at 0.93 percent. Only the combination of model and protective software pushes the rate to zero. Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: System Card Ask about this article… Search