메뉴
BL
The Decoder 48일 전

안스로픽 연구: AI, 보안 패치로 익스플로잇 제작에 '주단위' 아닌 '시간단위' 소요

IMP
9/10
핵심 요약

안스로픽의 최신 연구에 따르면, 대규모 언어 모델(LLM)이 공개된 보안 패치를 분석해 익스플로잇을 제작하는 데 걸리는 시간이 기존의 '주(Week)' 단위에서 '시간(Hour)' 단위로 극적으로 단축되었습니다. 연구진은 소스 코드가 공개된 파이어폭스는 물론, 폐쇄적인 윈도우 커널 환경에서도 고도의 권한 상승 공격 체인을 구축하는 데 성공했다고 밝혔습니다. 이는 소프트웨어 패치가 사실상 해커를 위한 완벽한 로드맵으로 변질되었으며, 보안 패치 적용 전의 골든타임이 사실상 사라졌음을 의미하는 매우 중요한 경고입니다.

번역된 본문

안스로픽(AI 기업) 연구: AI, 보안 패치로 익스플로잇을 제작하는 데 주가 아닌 단 몇 시간만 소요 작성자: Matthias Bastian 게시일: 2026년 6월 10일

안스로픽(AI 기업)의 보안 연구팀이 대규모 언어 모델이 파이어폭스와 윈도우의 알려진 취약점을 악용하는 속도를 얼마나 빠르게 측정했는지 체계적으로 분석했습니다. 그 결과는 패치 전략에 대한 기존의 오랜 가정을 완전히 깨부숴 버렸습니다.

소프트웨어 개발사가 보안 구멍을 막으면, 그 순간부터 경주가 시작됩니다. 공격자는 패치를 분석하고, 이를 통해 취약점을 역설계(Reverse-engineer)한 뒤, 아직 업데이트를 적용하지 않은 시스템을 공격할 수 있습니다. 버라이존(Verizon)의 데이터 유출 보고서(안스로픽 인용)에 따르면, 이른바 'N-Day 취약점'이 실제 피해의 매우 큰 부분을 차지합니다.

과거에는 패치를 역설계하는 작업이 느리고 전문화된 작업이었기 때문에 수비수(방어자)들에게 시간을 벌어주었습니다. 안스로픽 보안팀의 새로운 연구에 따르면, 이러한 시간적 버퍼는 이제 대부분 사라졌습니다. 연구진들은 "이제 단 한 명의 공격자도 단 하루 오후 만에 한 달 치 패치를 작동하는 익스플로잇으로 변환할 수 있으며, 이는 수천 달러의 비용과 전문적인 기술 없이도 가능하다"라고 말했습니다.

이제 패치는 공격자를 위한 로드맵입니다 보안 패치는 암묵적으로 버그가 어디에 있었는지 알려줍니다. 공격자는 이전 코드와 새 코드를 비교하여 결함을 정확히 찾아냅니다. 과거에는 이 과정에 몇 주가 걸렸습니다. 2020년 만디언트(Mandiant)의 분석에 따르면 25개 취약점 중 16개가 악용되기까지 한 달 이상 걸렸습니다. 안스로픽은 대규모 언어 모델이 이 과정을 얼마나 가속화하는지 측정했습니다.

아직 공개되지 않은 'Mythos Preview'를 포함한 6개의 클로드(Claude) 모델이 테스트되었습니다. 첫 번째 테스트에서 연구진은 파이어폭스의 자바스크립트 엔진인 'SpiderMonkey'에 대한 18개의 보안 패치를 선택했습니다. 파이어폭스는 의도적으로 선택된 것으로, 안스로픽에 따르면 이 브라우저는 수비수에게 '최상의 시나리오'입니다. 자동으로 업데이트되며, 모질라는 최근 사소한 업데이트 주기를 월간에서 주간으로 늘렸기 때문입니다. 만약 이 짧은 패치 간격조차도 공격자에게 충분하다면, 다른 소프트웨어는 훨씬 더 심각한 상태일 것입니다.

Mythos Preview는 18개 취약점 중 14개를 크래시시켰으며, 이는 버그를 찾고 이해했음을 증명합니다. 첫 번째 증명은 단 12분 만에 나왔고, 40분 이내에 13개가 더 뒤따랐습니다. 14번째는 약 3시간으로 훨씬 오래 걸렸습니다. Opus 4.5 모델은 2개만 해냈고, Opus 4.8은 11개를 해냈습니다. 취약점당 50회 반복 테스트한 신뢰성 테스트에서 Mythos Preview는 18개 버그 중 7개를 모든 시도에서 재현했습니다. Opus 4.8과 Opus 4.6은 각각 단 하나의 취약점에서만 그 수준의 일관성을 보였습니다.

단순한 크래시보다 더 중요한 것은 모델이 실제로 취약점을 악용해 대상 시스템에서 외부 코드를 실행할 수 있는지 여부입니다. Mythos Preview는 이 부문에서 확연히 앞서나가 약 12시간 만에 8개의 작동하는 익스플로잇을 만들어냈습니다. Opus 4.8은 2개, Opus 4.6과 Sonnet 4.6은 각각 1개를 만들어냈습니다. 첫 번째 익스플로잇은 패치가 공개된 지 1시간 이내에 준비되었으며, 이는 패치된 파이어폭스 148 버전이 배포되기 18일 전이었습니다.

소스 코드 없는 윈도우 커널: 8개의 권한 상승 공격 체인 두 번째 테스트는 훨씬 더 어려웠습니다. 바로 2026년 1월과 2월 '패치 튜즈데이(Patch Tuesday)'에서 발견된 윈도우 커널의 21개 취약점으로, 모두 공격자가 제한된 사용자 계정에서 최고 관리자 권한으로 넘어갈 수 있게 해주는 것들이었습니다. 파이어폭스와 달리 윈도우 소스 코드는 공개되어 있지 않습니다. 모델은 컴파일된 바이너리, 공개 디버그 기호, Ghidra 분석 도구의 기계 생성 디컴파일 코드, 변경된 함수의 비교(Diff) 결과, 그리고 마이크로소프트의 공개 자문 등을 바탕으로 작업해야만 했습니다.

Mythos Preview는 6시간 이내에 21개 취약점 중 18개를 찾아냈으며, 총 비용은 API 크레딧 약 2,200달러였습니다. Opus 4.8은 15개, Sonnet 4.6과 Opus 4.7은 각각 13개를 기록했습니다. 제한된 사용자 계정에서 최고 권한 수준인 'SYSTEM'으로 가는 완전한 권한 상승을 달성한 모델은 Mythos Preview가 유일했습니다. 이 모델은 총 15,700달러의 비용으로 8가지 다른 작동하는 공격 체인을 구축했으며, 익스플로잇당 평균 약 2,000달러가 소요되었습니다. Opus 4.8은 개별 공격 컴포넌트를 개발했지만, 이를 완전한 체인으로 결합하지는 못했습니다.

마이크로소프트는 이 21개 취약점 중 14개를 '악용될 가능성이 낮음(Less likely to be exploited)' 또는 '악[용될 가능성 없음(Unlikely to be exploited)]'으로 분류했습니다.

원문 보기
원문 보기 (영어)
Anthropic study shows AI needs hours, not weeks, to build exploits from security patches Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 10, 2026 Anthropic Anthropic's security research team has systematically measured how fast large language models can exploit known vulnerabilities in Firefox and Windows. The results blow up long-standing assumptions about patch strategies. When software makers close security holes, a race starts. Attackers can analyze the patch, reverse-engineer the vulnerability from it, and hit systems that haven't applied the update yet. According to Verizon's data breach report (via Anthropic) , these so-called N-Day vulnerabilities cause a huge share of real-world damage. Reverse engineering patches used to be slow, specialized work, and that bought defenders time. A new study from Anthropic's security team says that buffer is now mostly gone. "A lone operator can now turn a month’s worth of patches into working exploits in a single afternoon—for a few thousand dollars and with no specialized expertise," the researchers write. Patches are now roadmaps for attackers A security patch implicitly tells you where the bug was. Attackers compare old code with new code and pinpoint the flaw. Historically, this took weeks. In a Mandiant analysis from 2020 , 16 out of 25 vulnerabilities took a month or longer to be exploited. Anthropic measured how much large language models speed this up. Six Claude models were tested, including Mythos Preview, which isn't publicly available yet. For the first test, the researchers picked 18 security patches for SpiderMonkey, Firefox's JavaScript engine. Firefox was a deliberate choice: according to Anthropic, the browser is a best-case scenario for defenders. It updates itself automatically, and Mozilla recently increased the frequency of minor updates from monthly to weekly. If even these short patch gaps are enough, other software is in far worse shape. Mythos Preview crashed 14 of the 18 vulnerabilities, proving it had found and understood each bug. The first proof came after 12 minutes, and thirteen more followed within 40 minutes. The 14th took much longer, about three hours. Opus 4.5 managed just 2, Opus 4.8 hit 11. In reliability tests with 50 runs per vulnerability, Mythos Preview reproduced seven out of 18 bugs on every single attempt. Opus 4.8 and Opus 4.6 only hit that level of consistency for one vulnerability each. More important than a crash is whether the model can actually exploit the vulnerability to run foreign code on the target system. Mythos Preview pulled clearly ahead here, producing eight working exploits in about twelve hours. Opus 4.8 managed two, Opus 4.6 and Sonnet 4.6 each managed one. The first exploit was ready within an hour of the patch going live, 18 days before the patched Firefox 148 shipped. Windows kernel without source code: 8 privilege escalation chains The second test was much harder: 21 vulnerabilities in the Windows kernel from the January and February 2026 Patch Tuesdays, all allowing an attacker to jump from a restricted user account to full admin rights. Unlike Firefox, Windows source code isn't open. The model had to work with compiled binaries, public debug symbols, a machine-generated decompilation from the Ghidra analysis tool, a diff of changed functions, and Microsoft's public advisory. Mythos Preview found 18 of the 21 vulnerabilities in under six hours, at a total cost of about $2,200 in API credits. Opus 4.8 scored 15, Sonnet 4.6 and Opus 4.7 both scored 13. For full privilege escalation, going from a restricted user account to the highest privilege level, SYSTEM, Mythos Preview was the only model to succeed. It built 8 different working attack chains for a total of about $15,700, averaging roughly $2,000 per exploit. Opus 4.8 developed individual attack components but couldn't combine them into a complete chain. Microsoft classified 14 of the 21 vulnerabilities as "less likely to be exploited" or "unlikely to be exploited." Mythos Preview cracked 13 of those 14 and even achieved full privilege escalation for one rated "unlikely to be exploited." According to Anthropic, Microsoft's rating system is calibrated to human security researchers. Once Mythos-class models become more widely available, that calibration will have to change. The timing makes it worse. Even with Microsoft's automatic update service Windows Autopatch , it takes seven days for 90 percent of registered devices to get a patch and eleven days for a forced reboot. All eight of Mythos Preview's attack chains were done before a single device would have automatically applied the patch. Publicly available models can build exploits too Anthropic stresses that the Claude models already available to the public can also develop exploits when safety filters are turned off, just less successfully. Models from other companies and open-source models likely have similar capabilities, which widens the pool of potential attackers considerably. The old patch rhythm of monthly release cycles and staged rollouts is outdated, Anthropic argues. It's built on the assumption that exploiting a patch takes weeks of expert work. The common term "N-Day," which measures time between patch and exploit in days, is now misleading. "N-Hour" better describes the new reality. The researchers acknowledge that a real attack needs more steps, such as finding vulnerable targets, delivering the malicious code, and bypassing detection systems. But while these stages remain, the previously most time-consuming step, exploit development itself, now takes hours. Systems that are hard or slow to update face the greatest risk, including industrial control systems, medical devices, and networked equipment with fixed maintenance windows or vendor-locked software, Anthropic writes. A more durable fix than faster patching is to cut down on the sources of bugs themselves, for example through memory-safe languages like Rust or hardware-level protections that wipe out entire classes of attacks at once. The report was published before the release of Claude Fable 5 , Anthropic's Mythos variant with stronger safety restrictions. Mythos 5 (without the preview tag) is still only available to institutions Anthropic has selected , a problem for the EU, among others . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->