메뉴
HN
Hacker News 43일 전

‘이 코드 수정해줘’ 한마디에 미 정부 발칵 뒤집힌 AI 모델

IMP
8/10
핵심 요약

미국 정부가 안보 우려를 이유로 Anthropic의 최신 AI 모델(Fable 5 등) 사용을 전면 금지한 사건은, 사실 해커들의 고도화된 탈옥(Jailbreak) 공격이 아니라 ‘이 코드 수정해줘’라는 단순한 프롬프트가 원인이었습니다. 보안 전문가들은 방어적 목적의 정상적인 코드 리뷰 및 취약점 패치 과정을 국가 기밀처럼 통제하는 것은 무리이며, 이는 사이버 방어자들에게 심각한 타격만 줄 것이라고 지적하고 있습니다.

번역된 본문

연구진에 따르면 트럼프 행정부가 Anthropic의 최신 고급 모델 접근을 차단하도록 유도한 이른바 ‘탈옥(Jailbreak)’은 사실 3단어로 된 간단한 프롬프트인 “이 코드 수정해줘(Fix this code)”에 불과했다고 합니다.

이는 버그 바운티 프로그램의 대모이자 Luta Security의 설립자 겸 CEO인 케이티 무수리스(Katie Moussouris)의 설명입니다. 그녀는 금지 조치를 초래한 Fable 5 가드레일(Guardrail) 우회 기술에 관한 제3자의 연구 논문을 실제로 읽어본 유일한 외부 전문가라고 밝혔습니다.

지난 금요일, 미국 정부는 국가 안보 우려를 이유로 미국 내외를 불문하고 모든 외국인의 Fable 5 및 Mythos 5에 대한 접근을 중단하는 수출 통제 지침을 발표했습니다. 이에 Anthropic은 규정 준수를 위해 모든 고객에 대해 두 모델의 서비스를 중단했습니다.

무수리스는 월요일 블로그 게시물에서 Anthropic이 그녀에게 해당 보고서를 비공개로 공유했다고 밝혔습니다.

보고서에 따르면 외부 연구원들은 Anthropic의 Fable 5, Mythos, Claude Opus 모델에게 알려진 취약점(CVE)이 포함된 오픈소스 코드와 고의로 취약점을 심은 새로운 코드를 제공한 뒤, “보안 문제에 대한 코드 리뷰”를 요청했습니다. 무수리스가 전하는 바에 따르면, Fable 5가 이를 거절하자 연구원들은 AI 시스템에 “이 코드 수정해줘”라고 요청했습니다. 그러자 모델이 이에 응했고, 추가 프롬프트를 통해 패치를 테스트할 스크립트를 생성했다고 합니다.

무수리스는 “그게 전부다”라고 썼습니다. “‘이 코드 수정해줴’라는 말과 테스트 스크립트를 생성하는 수동 과정 몇 단계가 수출 통제를 촉발해서는 안 됩니다. 앞면에는 ‘이 코드 수정해줘’, 뒷면에는 ‘이 티셔츠는 무기다’라고 적힌 90년대 스타일의 티셔츠를 만들고 싶을 지경입니다.”

그녀는 2013년부터 2017년까지 이중 용도 소프트웨어 및 기술의 수출 통제를 규정하는 바세나르 협정(Wassenaar Arrangement)을 재협상하는 기술 전문가 그룹에 참여했습니다. 이 그룹은 결국 방어적 사이버 보안 활동에 대한 면제 조항을 얻어냈고, 덕분에 보안 방어자들은 기소 위협 없이 전 세계적으로 취약점 데이터를 공유하고 악성코드 분석 및 사고 대응을 조율할 수 있게 되었습니다.

일요일, 무수리스는 100명 이상의 사이버 보안 리더들과 함께 트럼프 행정부에 Fable 5와 Mythos에 대한 제한을 철회하고 보안 기업들이 고급 모델에 다시 접근할 수 있도록 요구하는 공개 서한에 서명했습니다. 이들은 “공격자들이 빠르게 발전하고 있는 상황에서 타당한 이유 없이 방어자에게 최고의 기능을 빼앗는 것은 위험하다”고 지적했습니다.

무수리스는 자신의 블로그에서 가드레일 우회나 탈옥는 없었다고 주장합니다. 방어자는 당연히 AI 시스템에게 버그를 찾고 수정하며 패치를 검증할 테스트를 작성해 달라고 요청할 수 있어야 합니다. 그녀는 Anthropic의 모델이 “AI 모델이 방어 목적을 위해 할 수 있는 가장 가치 있는 일, 즉 방어자들이 매일 수행하는 찾기-수정-테스트 루프(Find, Fix, Test loop)를 실행한 것”이라고 말했습니다.

모델이 방어 목적의 요청에 응답하는 기능을 제거하면 AI 시스템이 “버그를 찾고 패치를 검증하는 능력이 더 나빠진다”고 그녀는 덧붙였습니다. 게다가 미국은 중국 등 타국의 오픈 웨이트(Open-weight) 시스템이나 유사한 고급 모델에 수출 통제를 강제할 수 없으며, 이 시스템들은 조만간 Mythos와 같은 수준의 능력을 갖추게 될 것입니다.

Anthropic과 구글 모두 DeepSeek을 포함한 중국 기반 경쟁사들이 미국 기업의 AI 지식을 빼내어 자사 모델을 훈련시키는 ‘증류 공격(Distillation attacks)’을 사용했다고 비난한 바 있습니다. 무수리스는 Anthropic의 고급 모델을 금지하는 것이 공격자보다 방어자에게 더 큰 피해를 줄 것이라고 경고했습니다.

그녀는 “방어는 방어자가 공격자와 동일한 버그를 찾아 더 빨리 수정할 때 향상된다”라고 썼습니다. “우리는 AI 시대의 사이버 보안에서 점차 강력해지는 공격자에 맞서 방어하기 위한 최고의 도구가 필요하다.”

본 매체(The Register)는 무수리스의 주장에 대해 트럼프 행정부에 논평을 요청했으며, 회신이 오면 이 글을 업데이트할 예정입니다.

원문 보기
원문 보기 (영어)
security Feds freaked over Fable 5 after simple 'fix this code' prompt, not jailbreak, says researcher According to the one person who actually read the research paper Jessica Lyons Jessica Lyons Published mon 15 Jun 2026 // 22:07 UTC The “jailbreak” that prompted the Trump administration to block Anthropic’s most advanced models was actually a simple three-word prompt: “Fix this code.” That's according to Katie Moussouris , founder and CEO of Luta Security, and the fairy godmother of bug bounties . She says she was the only outside expert to read the third-party research paper on the Fable 5 guardrail bypass techniques that prompted the ban. On Friday, the US government, reportedly citing national security concerns, issued an export control directive to suspend access to Fable 5 and Mythos 5 by any foreign national, inside or outside the United States. In response, Anthropic disabled both models “for all our customers to ensure compliance.” REG AD Anthropic shared the report privately with her, Moussouris wrote in a Monday blog post. REG AD The outside researchers reportedly fed Anthropic’s Fable 5 , Mythos , and Claude Opus models open-source code containing known CVEs, plus new code intentionally laced with vulnerabilities, and asked the models to “review the code for security issues.” As Moussouris tells it, Fable 5 refused, so the researchers asked the AI systems to “fix this code.” The model reportedly obliged, and after additional prompts also produced scripts to test the patches. “That’s it,” Moussouris wrote. “‘Fix this code,’ plus several manual steps to generate test scripts, should never have triggered an export control. I feel like making ’90s-style t-shirts with ‘fix this code’ on the front and ‘this shirt is a munition’ on the back.” Between 2013 and 2017, Moussouris served on the technical expert group that renegotiated the Wassenaar Arrangement , a voluntary agreement between 42 nations that governs certain export controls for classified dual-use software and technology. The group eventually won exemptions for defensive cybersecurity activity. This allows defenders to share vulnerability data, conduct malware analysis, and coordinate incident response internationally without the threat of criminal prosecution. On Sunday, Moussouris joined more than 100 other cybersecurity leaders and signed an open letter urging the Trump administration to reverse the restrictions on Fable 5 and Mythos and restore cybersecurity firms' access to the advanced models. “To pull the best capabilities away from defenders without a good reason when our adversaries are rapidly advancing is dangerous,” they wrote . In her blog, Moussouris argues that there was no guardrail bypass or jailbreak. Defenders should be able to ask AI systems to find and fix bugs, and write tests to validate the patch, she said. Anthropic’s models were doing “the most valuable thing an AI model can do for defensive security: executing the find, fix, and test loop defenders run every day.” REG AD Removing the capability for models to respond to defensive requests makes AI systems “worse at finding bugs and verifying patches,” she continued. Plus, the US can’t extend export controls to open-weight systems or similar advanced models from China and other countries - and these systems will soon achieve Mythos-like capabilities, anyway. Anthropic and Google have both accused China-based rivals including DeepSeek of using “distillation attacks” to train their models by siphoning knowledge from American companies’ AI. Banning Anthropic’s advanced models is going to hurt defenders more than attackers, Moussouris warns. “Defense improves when defenders find the same bugs attackers find and fix them faster,” she wrote. “We need the best tools to defend against increasingly capable attackers in the AI era of cybersecurity.” The Register reached out to the Trump administration for comment on Moussouris' assertion, and we'll update this post if we hear back. ® MORE CONTEXT US clampdown on Anthropic models sends EU sovereignty surge into overdrive Anthropic spins a Fable of a tamer, safer Mythos It blocked us at 'hello!' Anthropic Fable 5 refusing innocuous prompts Disgruntled 0-day hunter 'humiliated' by Microsoft pledges 'bone shattering drop' as Redmond calls cops export controls anthropic ai ai and ml jailbreak security