메뉴
BL
The Decoder 49일 전

앤스로픽, 코딩·과학 대폭 향상된 5세대 모델 발표

IMP
9/10
핵심 요약

앤스로픽이 5세대 클로드 모델인 클로드 페이블 5(Claude Fable 5)와 클로드 미토스 5(Claude Mythos 5)를 공개했습니다. 범용 모델인 페이블 5는 코딩, 데이터 분석, 시각 처리 등 대부분의 벤치마크에서 기존 최고 성능 모델들을 뛰어넘는 압도적인 성능을 입증했으며, 보안 및 바이오 분야에 특화된 미토스 5는 신약 설계 및 유전체 연구에서 사람 수준 이상의 자율적 성과를 보여주며 AI의 실질적 업무 자동화 가능성을 입증했습니다.

번역된 본문

앤스로픽, 코딩 및 과학 분야에서 대폭적인 성능 향상을 이룬 클로드 페이블 5와 미토스 5 발표

작성자: Matthias Bastian | 2026년 6월 9일

핵심 요약 앤스로픽(Anthropic)은 5세대 AI 모델 두 가지를 새롭게 출시했습니다. 범용으로 사용 가능한 '클로드 페이블 5(Claude Fable 5)'와 초기에는 사이버 보안과 같은 전문 분야를 위해 선별된 파트너에게만 제공되는 '클로드 미토스 5(Claude Mythos 5)'입니다.

페이블 5는 프로그래밍, 이미지 처리, 복잡한 데이터 분석 벤치마크에서 최고 점수를 기록하며 앤스로픽의 이전 모든 모델을 능가하는 성능을 보여주었습니다. 반면 미토스 5는 신약 설계에서 강력한 성능을 발휘하며 유전체 연구에서는 대부분의 과정을 자율적으로 수행합니다.

새로운 모델의 가격은 백만 입력 토큰당 10달러로 책정되어, 기존 클로드 오퍼스 4.8(Claude Opus 4.8) 모델보다 거의 두 배 비싼 가격에 제공됩니다. 다만 토큰 효율성 측면에서의 최종적인 판단은 아직 남겨진 상태입니다.


앤스로픽은 5세대 클로드 모델 두 가지를 새롭게 출시했습니다. 클로드 페이블 5는 거의 모든 벤치마크에서 최고 자리를 차지했으며, 클로드 미토스 5(이제 프리뷰 버전이 아님)는 여전히 선별된 파트너들에게만 제공됩니다. 두 모델은 동일한 기반 모델(Base model)을 공유합니다.

페이블 5는 일반적인 사용자를 위한 보수적인 안전 가드레일(Safety guardrails)을 갖추고 출시되었습니다. 반면 미토스 5는 사이버 보안과 같은 특정 분야에서는 이러한 제한을 완화하며, 극소수의 파트너 그룹에게만 예약적으로 제공됩니다.

앤스로픽은 "페이블 5가 회사가 출시한 모든 범용 모델을 뛰어넘으며, 테스트된 거의 모든 벤치마크에서 최고 수준(State-of-the-art)의 결과를 달성했다"고 밝혔습니다. 또한 길고 복잡한 작업일수록 기존 모델들과의 성능 격차가 더욱 벌어진다고 덧붙였습니다.

소프트웨어 엔지니어링 실무 작업 해결 능력을 평가하는 벤치마크인 SWE-Bench Pro(공개된 GitHub 저장소를 기반으로 도움 없이 수행)에서 페이블 5는 80.3%를 기록했습니다. 같은 테스트에서 클로드 오퍼스 4.8은 69.2%, GPT 5.5는 58.6%, 제미나이 3.1 프로(Gemini 3.1 Pro)는 54.2%를 기록했습니다.

프로덕션 표준에 따라 까다로운 코딩 작업을 테스트하는 Cognition의 FrontierCode 벤치마크에서 페이블 5는 29.3%의 점수를 받았습니다. 클로드 오퍼스 4.8은 13.4%, GPT 5.5는 겨우 5.7%를 기록했습니다.

앤스로픽은 페이블 5가 이전 클로드 모델들보다 토큰 효율성도 더 뛰어나다고 주장합니다. 중간 수준의 노력(medium effort)을 기울였을 때, FrontierCode 벤치마크에서 모든 최신 프론티어 모델(Frontier models) 중 최고 점수를 기록했습니다.

결제 처리 업체인 스트라이프(Stripe)는 페이블이 5개월에 걸친 엔지니어링 작업을 단 며칠 만에 압축해 완료했다고 밝혔습니다. 5천만 줄 규모의 Ruby 코드베이스 환경에서, 이 모델은 온전한 팀이 두 달 이상 걸렸을 마이그레이션 작업을 단 하루 만에 끝냈습니다.

지식 노동, 시각, 장기 기억의 비약적 도약 앤스로픽에 따르면 페이블 5는 복잡한 분석 작업에서도 최고 수준의 성능을 보여줍니다. 노련한 금융 애널리스트 수준의 AI 추론 능력을 테스트하는 Hebbia의 Finance Benchmark에서 문서 기반 추론, 차트 및 표 해석 능력이 크게 향상되며 모델 중 최고 점수를 기록했습니다. 트레이딩 그룹인 IMC는 페이블 5가 자사의 트레이딩 분석 평가를 거의 전 분야에서 통과했다고 전했습니다.

시각(비전) 작업에서도 페이블 5는 새로운 최고 수준의 모델입니다. 자세한 과학 삽화에서 정확한 수치를 추출하거나, 스크린샷만으로 웹 앱의 소스 코드를 재구성할 수 있습니다. 한 데모 시연에서는 게임 스크린샷만을 이용해 '포켓몬스터 파이어레드(Pokemon FireRed)'를 플레이해 보였습니다. 기존 모델들이 지도와 같은 추가 게임 데이터 접근 및 복잡한 헬퍼 프레임워크가 필요했던 것과는 대조적인 성과입니다.

앤스로픽은 페이블 5가 수백만 토큰에 걸쳐 집중력을 유지하며, 스스로 메모(Note-taking)를 작성해 최종 결과물의 품질을 높인다고 설명했습니다. 단, 이와 관련된 구체적인 벤치마크 수치는 공유하지 않았습니다.

신약 설계 및 자율적인 유전체 연구 앤스로픽의 내부 단백질 설계 전문가들은 미토스 5가 신약 설계 과정의 일부를 10배나 가속화했다고 밝혔습니다. 한 테스트에서 단백질 설계 및 생물정보학(Bioinformatics) 도구만 부여된 채 사람의 도움 없이 작동한 이 모델은 숙련된 인간 작업자와 동등하거나 그보다 뛰어난 성과를 보였습니다.

과학자가 수행하는 모든 단계를 직접 처리했습니다. 결합 부위(Binding sites)를 선정하고, 단백질 설계 도구를 실행 및 관리하며, 발생한 오류를 스스로 수정했습니다. 14개의 단백질 타겟 중 9개에서 현재 연구 중인 강력한 신약 설계 후보군을 도출해 냈습니다.

앤스로픽의 가장 큰 주장은 미토스 5가 최초의... (원문 누락)

원문 보기
원문 보기 (영어)
Anthropic releases Claude Fable 5 and Mythos 5 with major gains in coding and science Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 9, 2026 Anthropic Key Points Anthropic has released two new fifth-generation AI models: Claude Fable 5 for general use and Claude Mythos 5, which is initially available only to selected partners for specialized areas such as cyber security. Fable 5 outperforms all of Anthropic's previous models, achieving top scores in benchmarks for programming, image processing, and complex data analysis, while Mythos 5 shows strong performance in drug design and operates largely autonomously in genomics research. The new models come at a price of 10 US dollars per million input tokens, making them nearly twice as expensive as the Claude Opus 4.8 model, with token efficiency still to be determined. Ask about this article… Search Anthropic releases two new models in the fifth Claude generation. Claude Fable 5 claims the top spot in nearly all benchmarks, while Claude Mythos 5 (no longer in preview) is still only available to select partners. Both models share the same base model. Fable 5 ships with conservative safety guardrails for general use. Mythos 5 drops those restrictions in areas like cybersecurity and is reserved for a small group of partners. Anthropic says Fable 5 beats every generally available model the company has ever shipped and claims state-of-the-art results in nearly all benchmarks tested. The gap widens on long, complex tasks, the company states. Ad On SWE-Bench Pro, a benchmark for solving real software engineering tasks from public GitHub repos without help, Fable 5 hits 80.3 percent. Claude Opus 4.8 lands at 69.2 percent, GPT 5.5 at 58.6 percent, and Gemini 3.1 Pro at 54.2 percent. Ad DEC_D_Incontent-1 On Cognition's FrontierCode benchmark, which tests demanding coding tasks under production standards, Fable 5 scores 29.3 percent. Claude Opus 4.8 manages 13.4 percent. GPT 5.5 gets just 5.7 percent. Fable 5 is also more token-efficient than earlier Claude models, Anthropic claims. At medium effort, it posts the top score among all frontier models on FrontierCode. Payment processor Stripe says Fable compressed five months of engineering work into days. In a Ruby codebase with 50 million lines, the model finished a migration in one day that would have taken a full team over two months. Ad Knowledge work, vision, and long-term memory jump ahead Fable 5 also tops the charts on complex analytical tasks, according to Anthropic. On Hebbia's Finance Benchmark, which tests AI reasoning at the level of seasoned financial analysts, it posted the highest score of any model, with gains in document-based reasoning and chart and table interpretation. Trading group IMC says Fable 5 passed their trading analysis evaluations almost across the board. On vision tasks, Fable 5 is the new state-of-the-art model, Anthropic says. It can pull precise figures from detailed scientific illustrations and rebuild a web app's source code from screenshots alone. As a demo, Fable 5 played through Pokemon FireRed using only game screenshots. Earlier models needed a complex helper framework with extra tools and access to additional game data like maps. Ad DEC_D_Incontent-2 Anthropic says Fable 5 stays focused across millions of tokens and boosts its own results by taking notes. The company didn't share specific benchmarks here. Ad Drug design and autonomous genomics research Anthropic's internal protein design experts say Mythos 5 sped up parts of the drug design process by 10x. In one test, the model, equipped with protein design and bioinformatics tools but no human help, matched or beat experienced human operators. It handled every step a scientist would: it picked binding sites, launched and ran protein design tools, and fixed errors on its own. Nine out of 14 protein targets yielded strong drug design candidates, which are now being studied. Anthropic's biggest claim is that Mythos 5 is the first model to consistently produce novel and convincing scientific hypotheses, something that's highly debated when it comes to current LLMs . In blinded comparisons, Anthropic's scientists preferred Mythos' molecular biology hypotheses over those from Opus-class models about 80 percent of the time. One hypothesis, a novel mechanism for an E. coli protein, was backed by an independent study , the company says. In genomics, Mythos 5 worked largely on its own for over a week, Anthropic claims. The model compiled single-cell data for millions of cells from 138 animal species, then designed and trained its own machine learning model to identify cells with the same function across distantly related organisms. The result reportedly outperformed a model recently published in Science, despite being 100 times smaller. Anthropic plans to publish these results in the coming months. Mythos 5 stays locked to cyber defenders for now Claude Mythos 5 will continue to be offered through Project Glasswing in partnership with the US government. It replaces the earlier Claude Myth Preview. Anthropic calls it the world's strongest cybersecurity model. It scored 78 percent on the ExploitBench benchmark, up from 69 percent for Mythos Preview and 40 percent for Opus 4.8. All current Mythos Preview users can upgrade to Mythos 5. Access will expand gradually in coordination with the US government. Anthropic is also planning a Trusted Access Program for biology. Select researchers will get access to Fable 5 without biology and chemistry safeguards, though cyber safeguards will stay in place. Price almost doubles compared to Opus, and Fable gets kicked out of subscription plans Both models cost $10 per million input tokens and $50 per million output tokens. Anthropic says that's less than half the price of Claude Mythos Preview, but it's much more than current Claude Opus models . How expensive Fable and Mythos actually are in practice depends on token consumption per task relative to cost per million tokens . On Claude.ai plans, the new models count as 2x usage, though it's not clear whether usage here maps directly to token burn. Model Base Input Tokens 5m Cache Writes 1h Cache Writes Cache Hits & Refreshes Output Tokens Claude Fable 5 $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok Claude Mythos 5 (limited availability) $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok Claude Opus 4.8 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok Fable 5 is available now through the Claude API and usage-based Enterprise plans. Subscription plans (Pro, Max, Team, seat-based Enterprise) follow a staggered rollout. Until June 22, Fable 5 is included at no extra cost. Starting June 23, access requires usage credits. Down the line, Anthropic plans to fold the model back into regular subscription plans once it has enough capacity. Opus fallback is supposed to keep dangerous prompts in check Anthropic says Mythos-class models are powerful enough to pose real risks, in cyberattacks or bioweapons research, for example. Fable 5 uses new AI classifiers that flag dangerous requests and automatically route them to the weaker Claude Opus 4.8 model. Over 95 percent of sessions aren't affected. The classifiers cover three areas: cybersecurity, biology and chemistry, and distillation, where third parties try to extract model capabilities, something all major Western AI companies suspect Chinese labs of doing . In the web interface and apps, users get a notification when the system falls back to Opus 4.8. In the Messages API, the request is blocked by default, though developers can turn on a server-side fallback. In cyber tests, Fable 5 scored a zero percent success rate on offensive tasks. External testers couldn't find a universal jailbreak in over 1,000 hours, Anthropic says. The company also built in another layer of protection that's invisible to users. Prompts aimed at building frontier LLMs, like setting up pretraining pipelines, distributed training infrastr