메뉴
BL
TechCrunch AI 48일 전

안스로픽 'Fable' 보안 가드레일, 전문가들 불호

IMP
6/10
핵심 요약

안스로픽이 강력한 사이버보안 모델인 'Mythos'의 퍼블릭 버전인 'Fable'을 공개했습니다. 그러나 Fable에 적용된 과도한 안전 가드레일이 정상적인 코드 리뷰나 블로그 포스트 분석 같은 무해한 요청까지 무차별적으로 차단하여 사이버보안 전문가들의 강한 불만을 사고 있습니다. 현재의 키워드 기반 제한 방식은 실무자들의 업무 효율을 떨어뜨린다는 지적이 나오며, 향후 가드레일의 정교한 개선이 필요해 보입니다.

번역된 본문

안스로픽은 화요일 최신 모델인 '페이블(Fable)'을 출시하며, 이를 강력하고 큰 화제를 모은 사이버보안 모델인 '미토스(Mythos)'의 퍼블릭 제한판이라고 설명했습니다. 하지만 모든 사람이 이러한 제한에 만족하는 것은 아니며, 수많은 사이버보안 연구원과 전문가들이 온라인에서 불만을 터뜨렸습니다.

IBM X-Force에서 근무하는 유명 보안 연구원인 발렌티나 "촘피(Chompie)" 팔미오티(Valentina Palmiotti)는 "[페이블은] 사이버와 간접적으로만 관련될 수 있는 모든 요청을 거부합니다. 블로그 게시물을 읽는 것과 같은 무해한 작업도 포함해서요"라고 말했습니다. 프롬프트가 가드레일을 건드리면, 페이블은 채팅을 일시 정지하고 자신의 "안전 조치가 이 메시지를 사이버보안 또는 생물학 주제로 플래그했습니다"라고 알려줍니다.

이러한 가드레일은 페이블이 멀웨어를 개발하거나 소프트웨어를 손상시키는 데 악용될 위험을 줄이기 위해 마련되었으며, 이는 안스로픽 내부의 오랜 우려 사항이었습니다. 생물학에 대한 제한은 생물학 무기 개발과 관련된 유사한 우려에서 비롯되었습니다. 지난 4월 AI 거대 기업이 미토스를 출시했을 때, 중요 소프트웨어 및 인프라를 보호하기 위해 모델을 배포하는 노력인 '프로젝트 글래스윙(Project Glasswing)'이라는 이름으로 모델 액세스를 제한된 수의 기업 및 조직으로 제한했습니다. 지난주 안스로픽은 15개국 수백 개 조직으로 미토스에 대한 액세스를 확대했습니다.

하지만 좋은 의도에도 불구하고, 많은 사이버보안 전문가들은 제한의 무작위적인 성격에 여전히 불만을 품고 있습니다. 사이버보안 베테랑인 맷 수이체(Matt Suiche)는 테크크런치(TechCrunch)에 "보안 코드를 작성해달라고 요청하면, 소프트웨어 엔지니어링 모범 사례가 아닌 사이버보안 관련 작업으로 간주하여 요청이 거부됩니다"라고 말했습니다. 페이블은 가드레일에 걸릴 경우 기본적으로 Claude Opus 4.8 모델로 전환(Fall back)되도록 프로그래밍되어 있습니다. "키워드 기반인 것 같습니다. 그래서 '사이버보안'의 어휘적 영역에 속하는 모든 것이 가드레일을 작동시킵니다."

AI 사이버보안 스타트업인 톨모(Tolmo)의 기술 직원인 수이체는 "하지만 아직 초기 단계이고 그들이 가드레일을 계속 조정하고 있기 때문에 이해할 수 있습니다. 안스로픽과 다른 프론티어 모델 기업들이 차세대 사이버보안 기업들과 더 많이 협력함에 따라 시간이 지나면 분명 진화할 것입니다"라고 말했습니다. "이런 출시를 할 때는 충분한 제한을 두는 것이 부족한 것보다 낫고, 시간이 지나면서 가드레일을 완화하는 것이 낫습니다."

또 다른 연구원은 X(옛 트위터)에서 "코드 리뷰를 요청하는 것만으로도 페이블의 가드레일이 작동한다"고 불평했습니다. 안스로픽은 코멘트 요청에 즉시 응답하지 않았습니다.

모델 내부의 가드레일 외에도 안스로픽은 사이버보안 전문가들에게 '사이버 검증 프로그램(Cyber Verification Program)'에 신청할 것을 요구합니다. 신청자가 승인을 받으면 Claude를 사용하여 사이버보안 작업을 수행할 때 제한이 줄어듭니다. 오픈AI(OpenAI)에도 '사이버를 위한 신뢰할 수 있는 액세스(Trusted Access for Cyber)'라는 유사한 프로그램이 있습니다.

원문 보기
원문 보기 (영어)
Anthropic released its latest model Fable on Tuesday, billing it as a public and limited version of its powerful and much-hyped cybersecurity model Mythos. But not everyone is happy with the restrictions, and a number of cybersecurity researchers and professionals have aired complaints online. “[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post,” said Valentina “Chompie” Palmiotti, a well-known security researcher who works at IBM X-Force. When a prompt triggers its guardrails, Fable pauses the chat and says that its “safety measures flagged this message for cybersecurity or biology topics.” The guardrails were put in place to limit the risk that Fable could be used to develop malware or compromise software — a longstanding concern within Anthropic. The restrictions on biology come from a similar concern around developing biological weapons . When the AI giant released Mythos in April, it restricted the model to a limited number of companies and organizations in what it called Project Glasswing , an effort to deploy the model to secure critical software and infrastructure. Last week, Anthropic expanded access to Mythos to hundreds of organizations in 15 countries. But despite the good intentions, many cybersecurity experts are still put off by the haphazard nature of the restrictions. Matt Suiche, a cybersecurity veteran, told TechCrunch that "if you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded.” Fable is programmed to fall back to Claude Opus 4.8 if it hits a guardrail. “It seems to be keyword based, so anything in the lexical field of ‘cybersecurity' triggers the guardrails.” Contact Us Do you have more information about how hackers are using AI? Or how cybersecuity companies are using AI? We'd love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email . “But it is understandable as we are still in the early days and they are still adapting their guardrails. I am sure they are going to evolve over time as Anthropic and other frontier model companies will collaborate more with the current new generation of cybersecurity companies,” said Suiche, who is a member of the technical staff at Tolmo, an AI cybersecurity startup. “It's better to catch more people than not enough when you do such a release and to relax the guardrails over time.” Another researcher griped on X that “even asking for a code review” triggers Fable’s guardrails. Anthropic did not immediately respond to a request for comment. Apart from guardrails inside its models, Anthropic requires cybersecurity professionals to apply to the Cyber Verification Program . If they get approved, the applicants have fewer limitations on using Claude for cybersecurity work. OpenAI has a similar program called Trusted Access for Cyber . Topics AI , ai safety , Anthropic , cybersecurity , fable , Mythos , Security When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Lorenzo Franceschi-Bicchierai Senior Reporter, Cybersecurity Lorenzo Franceschi-Bicchierai is a Senior Writer at TechCrunch, where he covers hacking, cybersecurity, surveillance, and privacy. You can contact or verify outreach from Lorenzo by emailing lorenzo@techcrunch.com , via encrypted message at +1 917 257 1382 on Signal, and @lorenzofb on Keybase/Telegram. View Bio June 18 Los Angeles Get an inside look at what it takes to scale and succeed from leaders at Mach Industries, Founders Fund, and Shinkei Systems. Through candid fireside chats and high-impact networking, you'll walk away with valuable insights and new connections. REGISTER NOW Most Popular Google just fired a warning shot in the AI subscription price wars Lucas Ropek Connie Loizos WWDC 2026: Everything announced on Siri AI, iOS 27, Apple Intelligence, and more Morgan Little Aisha Malik Anthropic's Claude Fable 5 is a version of Mythos the public can access today Rebecca Bellan It's not FAANG anymore. It's MANGOS. Julie Bort Microsoft's open source tools were hacked to steal passwords of AI developers Zack Whittaker Google will pay SpaceX $920M per month for compute Sean O'Kane Mira Murati steps back into the spotlight, carefully Connie Loizos
관련 소식