메뉴
HN
Hacker News 49일 전

앤스로픽, 최고 성능 모델 클로드 페이블 5 및 미토스 5 출시

IMP
9/10
핵심 요약

앤스로픽이 AI 성능 벤치마크를 대부분 석권한 최신 모델 '클로드 페이블 5(Claude Fable 5)'와 사이버 보안 특화 모델 '미토스 5(Mythos 5)'를 공개했습니다. 특히 복잡하고 장기적인 소프트웨어 엔지니어링 작업에서 압도적인 성능을 보여주며, 가격은 기존 프리뷰 모델 대비 절반 이하로 책정되었습니다. 보안상 위험을 방지하기 위해 일부 민감한 주제는 필터링되며, 사이버 방어를 위한 미토스 5는 미국 정부와의 협력을 통해 제한적으로 배포됩니다.

번역된 본문

발표: 클로드 페이블 5 및 클로드 미토스 5 2026년 6월 9일

오늘 우리는 일반 사용이 안전하도록 조치한 미토스(Mythos)급 1 모델, 클로드 페이블 5(Claude Fable 5)를 출시합니다. 페이블 5의 성능은 우리가 지금까지 일반에 공개한 그 어떤 모델보다 뛰어납니다. 소프트웨어 엔지니어링, 지식 노동(Knowledge work), 비전(Vision), 과학 연구 등 AI 성능을 평가하는 거의 모든 테스트 벤치마크에서 최고 수준(State-of-the-art)의 성과를 보여줍니다. 작업이 길고 복잡할수록 페이블 5가 다른 모델들에 비해 압도적인 우위를 차지합니다.

이토록 뛰어난 성능의 모델을 출시하는 것은 리스크가 따릅니다. 안전장치가 없다면 사이버 보안과 같은 분야에서 페이블 5의 능력이 악용되어 심각한 피해를 초래할 수 있습니다. 따라서 우리는 특정 주제에 대한 질문은 차순위 성능 모델인 클로드 오퍼스 4.8(Claude Opus 4.8)의 응답으로 대체되도록 안전장치를 마련하여 모델을 출시했습니다. 안전성과 신속한 출시를 동시에 달성하기 위해 이러한 안전장치를 보수적으로 조정했습니다. 가끔 무해한 요청이 걸리기도 하지만, 평균적으로 5% 미만의 세션에서만 작동합니다. 향후 몇 달 내에 더 강력한 모델이 출시될 예정이며, 우리는 안전장치를 개선하고 오탐(False positive)을 최대한 신속하게 줄이기 위해 노력하고 있습니다.

소수의 사이버 방어자 및 인프라 제공업체를 위해 클로드 미토스 5(Claude Mythos 5)도 함께 출시합니다. 이는 페이블 5와 동일한 기반 모델이지만, 일부 분야에서 안전장치가 해제된 버전입니다. 미토스 5는 미국 정부와의 협력을 통해 프로젝트 글래스윙(Project Glasswing)의 일환으로 클로드 미토스 프리뷰(Mythos Preview)를 대체하는 업그레이드 형태로 우선 배포될 예정입니다. 이 모델은 전 세계 그 어떤 모델보다 강력한 사이버 보안 역량을 갖추고 있습니다. 조만간 더 폭넓은 신뢰 기반 접근 프로그램을 통해 미토스 5에 대한 접근성을 확대할 계획입니다.

페이블 5와 미토스 5 같은 모델의 능력은 세상에 깊은 선한 영향을 미칠 잠재력을 가지고 있습니다. 우리는 프로젝트 글래스윙에서 이러한 초기 징후를 확인했습니다. 여기서 모델들은 사이버 방어자들이 매우 중요한 소프트웨어를 안전하게 보호하도록 도왔습니다. 또한 생명 과학 연구에서도 그 가능성을 보았는데, 모델이 새로운 가설을 제시하고 새로운 치료제 개발 속도를 높이고 있습니다.

페이블 5와 미토스 5는 백만 입력 토큰당 10달러, 백만 출력 토큰당 50달러에 제공되며, 이는 클로드 미토스 프리뷰 가격의 절반 이하입니다. 오늘의 공동 출시는 최대한 많은 사용자에게 최대한 빠르고 안전하게 고급 AI 기능을 제공하려는 우리의 목표를 향한 또 하나의 발걸음입니다.

클로드 페이블 5 및 클로드 미토스 5 평가

아래 표는 페이블 5와 미토스 5의 성능을 다른 최고 수준 모델들과 비교한 것입니다. 페이블 5와 미토스 5는 기존의 그 어떤 클로드 모델보다 더 오랜 시간 자율적으로 작업할 수 있습니다. 아래에서는 이러한 기술이 소프트웨어 엔지니어링에 어떻게 적용되는지 논의하고, 지식 노동, 비전, 메모리 및 생명 과학 연구 분야에서 향상된 모델의 성능을 다룹니다.

소프트웨어 엔지니어링(Software engineering). 초기 테스트 동안 스트라이프(Stripe)는 페이블 5가 수개월의 엔지니어링 작업을 단 며칠로 압축했다고 보고했습니다. 5천만 줄짜리 루비(Ruby) 코드베이스에서 이 모델은 수작업으로는 전체 팀이 두 달 이상 걸렸을 전체 코드 마이그레이션 작업을 단 하루 만에 수행했습니다. 또한 페이블 5는 이전 클로드 모델들보다 토큰 효율성이 더 뛰어납니다. 고품질의 유지보수가 가능한 에이전트 코딩을 평가하는 코그니션(Cognition)의 프론티어코드(FrontierCode) 평가에서 페이블 5는 중간 수준의 노력만 기울여도 최고 수준 모델들 중 가장 높은 점수를 기록했습니다.

지식 노동(Knowledge work). 페이블 5는 복잡한 분석 작업에서 강력한 성능을 보여줍니다. 시니어급 추론을 평가하는 헤비아(Hebbia)의 금융 벤치마크에서 페이블 5는 모델 중 가장 높은 점수를 받았으며, 문서 기반 추론, 차트 및 표 해석, 문제 해결 측면에서 상당한 성능 향상을 보여주었습니다. IMC는 페이블 5가 실제 정보 조회, 개념적 추론, 근본 원인 분석, 기댓값 분석 등 모든 트레이딩 분석 평가에서 거의 만점에 가까운 결과를 냈다고 언급했습니다.

비전(Vision). 페이블 5는 비전과 관련된 작업을 수행하는 데 있어 새로운 최고 수준(State-of-the-art)의 모델입니다. 상세한 과학 도표에서 정확한 숫자를 추출할 수 있으며, 웹 애플리케이션의 소스 코드를 재구축하는 것과 같은 복잡한 시각 기반 작업도 수행할 수 있습니다.

원문 보기
원문 보기 (영어)
Announcements Claude Fable 5 and Claude Mythos 5 Jun 9, 2026 Today we’re launching Claude Fable 5 : a Mythos-class 1 model that we’ve made safe for general use. Fable 5’s capabilities exceed those of any model we’ve ever made generally available. It is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas. The longer and more complex the task, the larger Fable 5’s lead over our other models. Releasing a model this capable comes with risks. Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage. We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8. To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we’re working to improve our safeguards and reduce false positives as quickly as we can. For a small group of cyberdefenders and infrastructure providers, we’re also launching Claude Mythos 5 . It’s the same underlying model as Fable 5, but with the safeguards lifted in some areas. 2 Mythos 5 will initially be deployed through Project Glasswing , in collaboration with the US Government, as an upgrade to Claude Mythos Preview. It has the strongest cybersecurity capabilities of any model in the world. Soon, we intend to expand access to Mythos 5 through a broader trusted access program. The capabilities of models like Fable 5 and Mythos 5 have the potential to do profound good for the world. We’ve seen the beginnings of this in Project Glasswing, where the models have helped cyber defenders secure critically important software. We’ve also seen it in life sciences research, where the models are positing novel hypotheses and speeding up the development of new therapeutics. Fable 5 and Mythos 5 are being offered at $10 per million input tokens and $50 per million output tokens—less than half the price of Claude Mythos Preview. Today’s joint launch is another step towards our goal of bringing advanced AI capabilities to as many users as possible, as quickly and as safely as we can. Evaluating Claude Fable 5 and Claude Mythos 5 The table below compares the capabilities of Fable 5 and Mythos 5 to other leading models. Fable 5 and Mythos 5 can work autonomously for longer than any previous Claude models. Below we discuss how these skills apply to software engineering, and cover the model’s improved capabilities in knowledge work, vision, memory, and life sciences research. Software engineering. During early testing, Stripe reported that Fable 5 compressed months of engineering into days. In a 50-million-line Ruby codebase, the model performed a codebase-wide migration in a day that would otherwise have taken a whole team over two months by hand. Fable 5 is also more token-efficient than past Claude models: on Cognition’s FrontierCode evaluation, which tests high-quality, maintainable agentic coding, Fable 5 scores highest among frontier models, even at medium effort. Knowledge work . Fable 5 shows strong performance on complex analytical tasks. On Hebbia ’s Finance Benchmark for senior-level reasoning, Fable 5 has the highest score of any model, with substantial gains in document-based reasoning, chart and table interpretation, and problem solving. IMC noted that Fable 5 aced their trading-analysis evaluations nearly across the board, including factual lookup, conceptual reasoning, root-cause analysis, and expected-value analysis. Vision. Fable 5 is the new state-of-the-art model for tasks involving vision. It can extract precise numbers from detailed scientific figures and can perform complex vision-based tasks like rebuilding a web app’s source code from screenshots alone. It also needs less scaffolding: for example, previous Claude models struggled to play Pokémon Fire even with harnesses that gave them additional helpful tools, but Fable 5 beat FireRed with a minimal, vision-only harness. Memory and long-context. Fable 5 stays focused across millions of tokens in long-running tasks and improves its outputs using its own notes. When we had the model play the deck-building game Slay the Spire , giving it access to persistent file-based memory improved its performance three times more than for Opus 4.8; Fable also reached the game’s final act three times more often. Solar eclipses Factorio VibeCAD Fluid with Classical EDM Drug design: Using Mythos 5, our internal protein design experts accelerated aspects of the drug design process by around ten times. In one example, they found that Mythos 5, with protein design and bioinformatics tools but no human assistance, matches or beats skilled human operators. In doing so, the model executes all of the tasks that are normally completed by a scientist: choosing binding sites, selecting and running protein design tools, and recovering from failures along the way. Nine of the 14 protein targets from this study (shown below) yielded strong candidates for drug design that we’re currently investigating. Novel hypotheses in molecular biology. Mythos 5 is our first model to consistently produce novel, compelling scientific hypotheses. In blinded head-to-head comparisons against Opus-class models, our scientists preferred Mythos’s molecular biology hypotheses ~80% of the time, and have advanced several to experimental evaluation. In the meantime, one Mythos hypothesis—a novel mechanism for an E. coli protein—was corroborated in a study from a lab independently working on the same problem. Novel research in genomics. Mythos 5 conducted novel genomics research in over a week of largely autonomous work. It assembled single-cell data for millions of cells spanning 138 animal species and designed and trained a custom machine learning model to identify cells performing the same role in even distantly-related organisms. With only high-level human input, Mythos 5’s trained model outperformed a recent model published in the journal Science —despite being 100 times smaller. We intend to publish these results in the coming months. Alignment . In our automated alignment assessment we found that Mythos 5’s level of misaligned behavior (including misaligned actions taken by the model such as deception, and cooperation with misuse of the model by a user) was low, and similar to that of Opus 4.8. Given they are the same underlying model, Fable 5’s level of alignment will be similar. The assessment is described in full, along with a detailed suite of other safety and capabilities tests, in the model’s system card . Early feedback for Claude Fable 5 Customers with early access ran their own tests on Fable 5. Below, in their words, is a selection of what they’re seeing: Claude Fable 5 is the state of the art model on CursorBench. It's opened up a class of long-horizon problems that were out of reach for earlier models. Claude Fable 5 is a real step forward for the developers GitHub serves. In our early testing, it took on complex, long-horizon coding tasks with a level of autonomy and reliability that exceeded previous benchmarks. But what excites us most is the direction it points: a future where developers can hand increasingly ambitious work to agents and trust the results across the software lifecycle. These are the strongest results of any Claude model we've had the opportunity to test. Claude Fable 5 is a clear step forward on agentic coding and prototyping. Claude Fable 5's reasoning is a clear step beyond Opus 4.8. It works at senior research scientist grade — picking directions, allocating resources, killing its incorrect beliefs, and producing novel first-principles outputs. Claude Fable 5 understands what builders mean, not j