메뉴
BL
The Decoder 28일 전

탈옥(Jailbreak) 논란 2주 만에 앤스로픽 '페이블 5' 전 세계 재출시

IMP
8/10
핵심 요약

앤스로픽의 차세대 AI 모델 '페이블 5(Fable 5)'가 안전장치 우회(탈옥) 취약점 논란으로 인한 2주간의 미국 정부 수출 통제를 마치고 전 세계적으로 다시 서비스를 재개했습니다. 회사는 문제가 된 해킹 요청을 차단하는 새로운 보안 필터를 도입했으나, 이로 인해 일반적인 프로그래밍 요청까지 과도하게 거절되는 딜레마에 직면했습니다. 앤스로픽은 AI 모델의 완벽한 탈옥 방어는 불가능하다고 인정하며, 업계 표준 보안 프레임워크 구축과 선제적 정부 규제를 적극적으로 촉구하고 있습니다.

번역된 본문

2주간의 사용 중단 조치 이후, 미국 정부는 앤스로픽의 두 번째로 강력한 AI 모델인 '페이블 5(Fable 5)'에 대한 수출 통제를 해제하여 전 세계적으로 다시 사용할 수 있도록 허용했습니다. 상대적으로 제한이 적은 버전인 '미토스 5(Mythos 5)'는 여전히 승인을 받은 일부 미국 기관으로만 사용이 제한됩니다. 앤스로픽은 해결책으로 이러한 요청을 차단하는 새로운 필터를 학습시켰지만, 이 필터는 무해한 프로그래밍 작업까지 더 자주 거절하는 부작용을 낳고 있습니다.

2주간의 규제 끝에 미국 정부는 앤스로픽이 자사의 가장 강력한 AI 모델을 다시 전 세계적으로 출시하도록 허용했습니다. 오늘부터 페이블 5는 클로드 플랫폼(Claude Platform), Claude.ai, Claude Code, 그리고 Claude Cowork를 통해 전 세계에서 다시 사용할 수 있습니다. Pro, Max, Team 및 일부 엔터프라이즈 요금제 사용자는 7월 7일까지 주간 사용량 한도의 최대 50% 범위 내에서 이 모델을 이용할 수 있습니다. 이 기간 이후에는 사용량 기반 크레딧으로 과금될 예정입니다. AWS, 구글 클라우드(Google Cloud), 마이크로소프트 파운드리(Microsoft Foundry)를 통한 접근은 "가능한 한 빨리" 복구될 예정입니다.

동일한 기본 모델의 상대적으로 제한이 적은 버전인 미토스 5는 6월 26일 정부 승인을 받은 일부 미국 기관 그룹으로만 여전히 제한됩니다. 앤스로픽은 이른바 '글래스윙(Glasswing)' 프로그램에서 더 많은 파트너에게 접근 권한을 확대하기 위해 정부와 계속 협의 중이라고 밝혔습니다. 유럽연합(EU)이 이 프로그램에 참여할지는 아직 불확실합니다.

이번 규제 조치는 아마존 연구원들의 보안 발견에서 비롯된 것이라고 앤스로픽이 확인했습니다. 연구원들은 페이블 5의 안전 가드레일(보호망)을 우회하는 방법을 발견했습니다. 이후 이 모델은 여러 소프트웨어 취약점을 식별했으며, 특정 사례에서는 그중 하나를 악용하는 방법을 보여주는 코드를 생성해 냈습니다.

미국 정부와 앤스로픽은 이 취약점을 조사하는 데 2주를 보냈습니다. 보고서에서 페이블 5가 발견한 것과 동일한 결함은 클로드 오퍼스 4.8(Claude Opus 4.8), GPT-5.5, 미티(Mimi) K2.7을 포함해 능력이 떨어지는 다른 많은 모델에서도 발견할 수 있었습니다. 특정 익스플로잇(취약점 공격) 데모의 경우, 클로드 하이쿠 4.5(Claude Haiku 4.5) 같은 훨씬 작은 규모의 모델을 포함하여 테스트된 모든 모델이 동일한 결과를 생성했습니다. 앤스로픽은 이를 일상적인 방어적 사이버 보안 작업에만 해당하는 극단적인 사례(Edge case)라고 불렀습니다.

이에 대응하여 회사는 아마존 보고서에 나온 기술을 99% 이상의 사례에서 차단하는 개선된 안전 분류기(classifier)를 학습시켰습니다. 요청이 차단되면 사용자에게 알림이 표시되며, 해당 요청은 이전 버전인 오퍼스 4.8 모델로 라우팅(전달)됩니다.

하지만 이 새로운 분류기는 트레이드오프(단점)를 동반합니다. 일상적인 코딩 및 디버깅 중에 무해한 요청을 더 자주 플래그(차단 대상)로 지정한다는 것입니다. 사용자들은 첫 번째 페이블 릴리스 당시에도 모델이 너무 제한적이라고 불만을 제기했었습니다. 출시 당시에는 보편적인 탈옥(Universal jailbreak)이 발견되지 않았습니다. 하지만 회사는 "탈옥으로부터 완전히 견고한(즉, 영향을 받지 않는) AI 모델을 만드는 것은 아마도 불가능할 것"이라고 인정합니다. 이는 페이블 5가 출시되기 전부터 잘 알려진 사실이었습니다.

앤스로픽은 AI 업계에 탈옥을 평가하고 대응 조치를 촉발하기 위한 공유 표준이 필요하다고 주장합니다. 회사는 아마존, 마이크로소프트, 구글 및 기타 글래스윙 파트너들과 함께 이러한 프레임워크를 구축하고 있다고 밝혔습니다. 또한 앤스로픽은 탈옥 제보 채널을 24시간 모니터링하는 전담 팀을 구성했으며, 보안 연구원들이 페이블 5의 잠재적인 사이버 탈옥을 신고할 수 있는 새로운 해커원(HackerOne) 프로그램을 출시했습니다.

앤스로픽은 프론티어 모델(최첨단 AI 모델)에 대한 정부의 더 긴밀한 감독을 원합니다. 앤스로픽은 행정명령과 관련된 공동 노력을 바탕으로 미국 정부와의 협력을 확대하고 있습니다. 회사는 여러 가지를 약속했습니다. 정부 파트너는 보안에 민감한 영역에서 역량을 강화하는 모델에 대한 사전 출시 액세스 권한을 얻게 됩니다. 발견된 탈옥 또는 남용 패턴은 신속하게 공유될 것입니다. 앤스로픽는 공동 연구를 위해 전담 자원과 상당한 컴퓨팅 파워를 제공할 것입니다. 그리고 회사는 프론티어 모델 제공업체를 위한 공유 업계 표준 구축을 지원할 것입니다. 앤스로픽은 이 모든 것을 '강력한 규제'에 명시하기를 원합니다.

원문 보기
원문 보기 (영어)
Anthropic's Fable 5 is back worldwide after a two-week government ban over a jailbreak Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 1, 2026 Key Points After a two-week suspension, the U.S. government has lifted export controls on Anthropic's second most powerful AI model, Fable 5, making it available worldwide again. The less restricted version, Mythos 5, remains limited to a select group of U.S. organizations. Anthropic has trained a new filter that blocks such requests as a fix, but the filter also rejects harmless programming tasks more often. Ask about this article… Search After a two-week ban, the US government is letting Anthropic ship its most powerful AI model globally again. Fable 5 is back worldwide starting today through the Claude Platform, Claude.ai , Claude Code, and Claude Cowork. Pro, Max, Team, and select Enterprise plans include the model through July 7 at up to 50 percent of weekly usage limits. After that, it'll be billed through usage credits. Access on AWS, Google Cloud, and Microsoft Foundry will be restored "as quickly as possible." Mythos 5, the less restricted version of the same base model, remains limited to a group of US organizations that got government approval on June 26. Anthropic says it's still working with the government to expand access to more partners in the so-called Glasswing program. Whether the EU will join remains unclear. Ad Anthropic confirms that the ban stemmed from a security finding by Amazon researchers. They had found a way to bypass Fable 5's safety guardrails. The model then identified several software vulnerabilities and, in one case, produced code showing how to exploit one of them. Ad DEC_D_Incontent-1 Anthropic says it's "probably impossible" to make an AI model that can't be jailbroken The US government and Anthropic spent two weeks investigating the vulnerability. Many less capable models could spot the same flaws Fable 5 found in the report, including Claude Opus 4.8, GPT-5.5, and Kimi K2.7. For the specific exploit demo, every model tested produced the same result, even much smaller ones like Claude Haiku 4.5. Anthropic calls it an edge case that only involved routine defensive cybersecurity work. In response, the company trained an improved safety classifier that blocks the technique from the Amazon report in more than 99 percent of cases. When a request gets blocked, users see a notification, and the request gets routed to the older Opus 4.8 model. Ad The new classifier comes with a tradeoff, though. It flags harmless requests more often during everyday coding and debugging. Users had already complained the model was too restrictive during the first Fable release . No universal jailbreak was found at the time of release. But the company admits it's "probably impossible to make any AI model fully robust (that is, impervious) to jailbreaks." That was well known before Fable 5 shipped. Ad DEC_D_Incontent-2 The AI industry needs a shared standard for rating jailbreaks and triggering countermeasures, Anthropic argues. The company says it's building such a framework with Amazon, Microsoft, Google, and other Glasswing partners. Anthropic is also standing up a team for 24/7 monitoring of jailbreak submission channels and launched a new HackerOne program where security researchers can report potential cyber jailbreaks for Fable 5. Ad Anthropic wants closer government oversight of frontier models Anthropic is expanding its work with the US government, building on their joint efforts tied to the executive order . The company is making several commitments. Government partners will get pre-release access to models that advance capabilities in security-sensitive areas. Discovered jailbreaks or abuse patterns will be shared quickly. Anthropic will put up dedicated resources and significant compute for joint research. And the company will help build a shared industry standard for frontier model providers. Anthropic wants all of this written into "strong regulation" and applied equally to every frontier model developer. "Government involvement in AI releases requires a durable, transparent process that gives cyber defenders and others the certainty they need about access to powerful models," the company writes. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Anthropic