메뉴
HN
Hacker News • 24일 전

클로드 페이블 5.1 및 클로드 미토스 5.1 출시

IMP
9/10
핵심 요약

Anthropic이 코딩·지식 작업용 최첨단 모델 '클로드 페이블 5.1'과 사이버보안·생명과학 연구용 안전장치가 적용된 '클로드 미토스 5.1'을 공개했습니다. 페이블 5.1은 캐시 읽기 요금 인하로 일반 워크로드 기준 약 25%, 에이전트형 작업에서는 최대 약 45% 비용이 절감되며, 기업 고객을 위한 완전한 프라이버시 보장 방식(Enterprise Frontier Safeguards)도 단계적으로 도입됩니다. 특히 밀레니엄 투자사 내부 시스템의 희귀 크래시 원인을 수년간의 인간 엔지니어와 기존 모델 모두 찾지 못했던 것을 해결하는 등 장기 과제 해결 능력에서 새로운 기준을 제시했습니다.

번역된 본문

저희는 클로드 페이블(Claude Fable) 5.1과 클로드 미토스(Claude Mythos) 5.1을 소개합니다. 이들은 코딩과 지식 작업 분야에서 세계 최고 수준의 모델이며, 그 연구 기능은 AI 모델이 과학적 발전에 어떻게 기여할지를 미리 보여줍니다.

클로드 페이블 5.1과 클로드 미토스 5.1은 동일한 모델이지만 안전장치 수준이 다릅니다. 페이블 5.1은 일반적으로 이용 가능하며, 미토스 5.1은 신뢰 기반 접근 프로그램을 통해서만 제공되고, 그 안전장치는 사이버보안 및 생명과학 분야의 작업을 지원하도록 특별히 설계되었습니다.

성능 향상과 함께, 페이블 5.1은 가격, 데이터 보존, 안전장치에 대해 고객으로부터 받은 피드백을 해결하기 위한 중요한 조치들을 취했습니다.

가격. 페이블 5.1은 토큰 기반으로 과금되는 모든 곳에서 일반적인 워크로드 기준 페이블 5보다 약 25% 저렴해질 것으로 예상됩니다. 이는 캐시 읽기(모델이 이미 처리·저장된 입력을 읽는 것) 요금을 인하했기 때문입니다. 에이전트형 작업이 많은 경우 절감 효과는 훨씬 커서 최대 약 45%에 달합니다.

데이터 보존. 새로운 엔터프라이즈 프론티어 세이프가드(Enterprise Frontier Safeguards, EFS) 시스템은 제로 데이터 보존 정책과 동일한 완전한 프라이버시를 고객에게 제공하면서도, 악의적 이용 방지 면에서는 여전히 최고 수준을 유지합니다. EFS는 데이터를 Anthropic이 아니라 고객이 완전히 통제하는 클라우드 인프라에 저장하는 방식으로 작동합니다. 이번 가을 말부터 단계적으로 기업 고객에게 제공될 예정이며, EFS가 출시되기 전까지는符合条件的 고객이 페이블 5.1을 제로 데이터 보존 방식으로 사용할 수 있습니다.

안전장치. 저희는 오탐(시스템이 무해한 콘텐츠를 플래그하는 경우)을 줄이도록 안전장치를 개선했습니다. 사이버보안 분야에서 새로운 안전장치는 이전보다 60% 적은 오탐을 기록합니다. 이는 부분적으로는 페이블 5.1이 이제 소프트웨어 취약점을 발견하는 데 사용될 수 있기 때문입니다(다만 해당 취약점을 이용한 익스플로잇 개발은 불가). 생물학 분야에서는 미국 정부와 협력하여 개발한 접근 프로그램을 통해 클로드 미토스 5.1의 고급 생물학 기능에 대한 접근을 허용하고 있으며, 과학자들의 등록을 곧 오픈할 예정입니다.

새로운 성능 기준

클로드 페이블 5.1은 코딩, 지식 작업, 장기 실행 문제 해결 과제에서 새로운 기준을 제시합니다. 아래 차트는 페이블 5.1이 이전 모델인 페이블 5보다 훨씬 높은 성능을 달성할 수 있음을 보여줍니다. 그리고 낮음(Low) 또는 중간(Medium) 노력 수준으로 설정하면 페이블 5.1은 페이블 5와 동등하거나 더 나은 결과를 훨씬 낮은 비용으로 달성합니다. (페이블 5.1은 Claude Code에서는 기본적으로 높음(High) 노력 수준, Claude Cowork와 Claude.ai에서는 중간(Medium)으로 설정되어 있습니다.)

  • 에이전트형 과학 연구
  • 에이전트형 터미널 코딩
  • 학제간 추론
  • 에이전트형 코딩

페이블 5.1은 낮은 품질의 결과로 이어지는 지름길을 피하며, 소프트웨트 문제의 근본 원인을 해결할 만큼 똑똑합니다. 예를 들어 투자사 밀레니엄(Millennium)의 테스트에서 페이블 5.1은 수년간 시도했음에도 그 어느 엔지니어도(또는 다른 모델도) 설명하지 못했던 내부 시스템의 희귀 크래시 원인을 찾아냈습니다.

여기에서 다양한 벤치마크에서 페이블 5.1의 성적을 확인할 수 있습니다.

얼리 액세스 파트너들은 이러한 성능 향상을 체감했으며, 모델 출력의 정성적 개선 또한 알아차렸습니다. 파트너들의 평가는 다음과 같습니다:

"내부 벤치마크에서 클로드 페이블 5.1은 페이블 5나 오퍼스 5보다 더 많은 코딩 문제를 해결하며, 트레이딩 직관 분야에서 최고 수준(SOTA)을 달성했습니다. 이전 모델들은 작업 시간이 길어지면 따라가기 어려워졌지만, 페이블 5.1은 길고 다단계인 작업에서도 읽기 쉬운 결과를 유지합니다."

— Jane Street Capital, 크레이그 폴스(Craig Falls), 퀀트 리서치 총괄

원문 보기
원문 보기 (영어)
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress. Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences. Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards. Price . Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%. Data retention . Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention. Safeguards . We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon. A new performance frontier Claude Fable 5.1 sets a new standard on coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves similar or better results than Fable 5 at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.) Agentic scientific research Agentic terminal coding Multidisciplinary reasoning Agentic coding Agentic scientific research Agentic terminal coding Multidisciplinary reasoning Agentic coding Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying. Here, you can see how Fable 5.1 compares across various benchmarks: Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us: Previous 1 of 22 Next Quote “In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.” Company Jane Street Capital Author Craig Falls, Head of Quantitative Research