메뉴
HN
Hacker News 48일 전

Anthropic, 숨겨진 Claude Fable 안전장치 사과

IMP
9/10
핵심 요약

Anthropic이 새로운 AI 모델인 Claude Fable에 모델 증류(Model Distillation)를 막기 위한 보이지 않는 안전장치를 몰래 적용한 논란에 대해 공식 사과했습니다. 사용자 및 연구원들의 불만이 제기되자, 이제는 응답을 조용히 저하시키는 대신 다른 안전 규정처럼 제한 사실을 명확히 안내하는 방식으로 투명하게 개선하겠다고 밝혔습니다. 이는 AI 모델의 평가 투명성과 경쟁사 간의 데이터 활용을 둘러싼 업계의 중요한 정책적 기준을 제시하는 사건입니다.

번역된 본문

AI 뉴스 - Anthropic

Anthropic, 보이지 않는 Claude Fable 안전장치에 대해 사과하다

이 회사는 모델 증류(model distillation)를 방지하기 위해 은밀하게 설정했던 안전장치를 다른 안전 조치들과 같이 사용자에게 공개적으로 보이도록 만들겠다고 밝혔다.

By Robert Hart / 2026년 6월 11일

Anthropic은 새로운 AI 모델인 Claude Fable 5에 숨겨진 안전장치를 은밀히 장착해 성능을 제한한 것에 대해 사과했다. 이 조치는 이 모델을 사용해 경쟁 시스템을 개발하려는 연구원과 경쟁사 모두에게 방해가 되었다. 이 회사는 방침을 바꾸어, Fable이 더 많은 질문을 거절하게 되더라도 제한이 발동하는 시기를 더 투명하게 알리겠다고 말했다.

Fable은 Anthropic의 AI 시스템 등급인 'Mythos' 클래스에서 대중에게 널리 제공되는 첫 번째 모델이다. 이 회사는 수개월 동안 이 모델들이 대중에게 공개하기에는 너무 위험하다고 경고해 왔다. Anthropic은 Fable이 특정 '고위험' 질문에 응답하지 못하도록 막는 안전장치를 적용해 출시함으로써 이러한 위험 중 일부를 해결했다고 밝혔다.

Anthropic이 Fable의 응답을 제한할 것이라고 밝힌 영역 중 하나는 '증류(distillation)'이다. 이는 더 큰 모델의 출력을 사용해 더 작은 AI 모델을 훈련시키는 기술이다. 시스템 작동 방식을 설명하기 위해 AI 개발자들이 공개하는 문서인 '시스템 카드(system card)'에서 Anthropic은 증류 시도로 의심되는 질문을 처리할 때 모델의 답변을 직접 수정하고 성능을 저하시키겠다고 밝혔다. 이때 사용자는 안전장치가 작동했음을 알리는 알림을 받거나 응답이 변경되었다는 사실을 전혀 통보받지 못했다.

Anthropic은 이제 증류에 대한 접근 방식을 변경하고 있다. X(옛 트위터)에 올린 게시물에서 회사 측은 이제 해당 질문들이 이전 플래그십 모델인 Claude Opus 4.8로 돌아가도록(Fallback) 처리할 것이라고 밝혔다. 또한 Anthropic은 "이 일이 발생할 때마다 여러분은 이를 볼 수 있을 것"이라며 사용자에게 명확하게 알릴 것이라고 강조했다.

이는 Fable이 다른 고위험 영역의 질문을 처리하는 방식과 유사하다. 생물학, 화학, 사이버 보안과 같은 영역에서 안전 기능이 작동하면, 마약, 무기 또는 기타 금지된 콘텐츠를 다루는 회사의 광범위한 안전 규칙에 따라 완전히 차단되지 않는 한 질문은 Opus 4.8을 통해 라우팅된다. 특히 생물학의 경우 안전장치가 너무 광범위하게 설정되어 있어 기본적인 질문에도 Fable을 사실상 사용할 수 없는 경우도 있었으며, Anthropic은 The Verge와의 논평에서 이를 인정했다.

Anthropic은 "눈에 보이는 안전장치는 테스트와 우회가 가능하므로 튼튼하게 설계되어야 하며, 이를 제대로 완성하는 데는 시간이 걸립니다. 보이지 않는 안전장치는 더 좁은 범위를 타겟팅할 수 있어 오탐지(False positive)를 최소화하며 빠르게 출시할 수 있습니다. 저희는 이런 이유로 보이지 않는 안전장치를 선택했지만, 이는 잘못된 트레이드오프였습니다. 여러분은 저희가 마련한 안전장치와 그 이유에 대해 알 권리가 있으며, 균형을 맞추지 못해 죄송합니다"라고 밝혔다.

이번 변화는 경쟁 모델로 증류하려는 사용자를 조용히 제한하려는 Anthropic의 결정에 AI 연구 커뮤니티의 강력한 반발이 일어난 이후에 나왔다. 비평가들은 이 안전장치가 최첨단 모델을 평가하려는 제3자에게도 영향을 미칠 수 있다고 경고했다. 시스템 카드에서 Anthropic은 최신 모델의 AI 개발 가속화 능력이 이러한 요청을 차단할 정당한 이유가 된다고 지적하며, "Claude을 사용해 경쟁 모델을 개발하는 것은 이미 당사의 서비스 약관을 위반하는 것"이라고 덧붙였다. Anthropic은 이전에도 중국 경쟁사인 DeepSeek가 '산업적' 규모로 자사 모델을 부당하게 증류했다고 비난한 바 있다.

원문 보기
원문 보기 (영어)
AI News Anthropic Anthropic apologizes for invisible Claude Fable guardrails The company says it will make the covert safeguard preventing model distillation as visible as other safety measures. The company says it will make the covert safeguard preventing model distillation as visible as other safety measures. by Robert Hart Jun 11, 2026, 11:40 AM UTC Link Share Gift Image: The Verge Robert Hart is a London-based reporter at The Verge covering all things AI and a Senior Tarbell Fellow. Previously, he wrote about health, science and tech for Forbes . Anthropic has apologized for stealthily throttling its new AI model, Claude Fable 5 , with hidden guardrails that undermine both researchers and rivals using it to develop competing systems. The company says it is reversing course and will be more transparent about when the restrictions kick in, even if that means Fable refuses more queries. Fable is the first widely available model in Anthropic’s Mythos class of AI systems, a group the company has spent months warning are too dangerous for public release . Anthropic says it has addressed some of those risks by launching Fable with safeguards that prevent it from responding to certain “high-risk” queries. One of the areas Anthropic said it would restrict Fable’s responses is distillation, a technique for training smaller AI models using the outputs of larger ones. In Fable’s system card — a public document AI developers release to explain how a system works — Anthropic said it would handle queries it believed were distillation attempts by altering and degrading the model’s answers directly. Users would not be notified that they had triggered the safety measure or informed that the responses had been changed. Anthropic said it is now changing its approach to distillation: Queries will now fall back to Claude Opus 4.8, Anthropic’s previous flagship model , the company said in a post on X. Anthropic will prominently tell users too: “You will see this every time it happens.” This is similar to how Fable handles queries in other high-risk areas. When safety features are triggered in areas like biology, chemistry, and cybersecurity, queries are routed through Opus 4.8 unless they are blocked outright under the company’s broader safety rules, such as those covering drugs, weapons, or other prohibited content. In some cases, notably biology, the safeguards have been calibrated so broadly that Fable is practically unusable for even basic queries , something Anthropic acknowledged in a comment to The Verge . “Visible safeguards can be probed, so they have to be robust, which takes time to get right,” Anthropic wrote. “Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason—and that was the wrong tradeoff. You should have visibility into the safeguards we have in place, and why. We’re sorry for not getting the balance right.” The change follows intense backlash from the AI research community over Anthropic’s decision to silently limit users suspected of trying to distill Fable into competing models — a safeguard critics warned could also affect third parties trying to evaluate the frontier model. In the system card, Anthropic said newer models’ ability to accelerate AI development justified targeting those requests, noting that “using Claude to develop competing models already violates our Terms of Service.” Anthropic has previously accused Chinese rivals like DeepSeek of unfairly distilling its models on an “industrial” scale. Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates. Robert Hart AI Anthropic News Most Popular Most Popular Xbox warns of a ‘reset’ as it prepares for layoffs iFixit Trump phone teardown confirms it’s an HTC dupe Microsoft restricts Claude Fable for employees over data retention concerns Nearly a million passports and photo IDs were left unprotected on the public internet Claude Fable won’t answer basic biology questions The Verge Daily A free daily digest of the news that matters most. Email (required) Sign Up By submitting your email, you agree to our Terms and Privacy Notice . This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply. Advertiser Content From This is the title for the native ad