메뉴
BL
The Decoder • 24일 전

안스로픽, 클로드 페이블 5.1 출시…코딩·연구 성능 향상에 최대 45% 저렴

IMP
8/10
핵심 요약

안스로픽이 클로드 페이블 5.1과 미토스 5.1을 출시했습니다. 에이전틱 코딩 및 연구 벤치마크에서 큰 성능 향상을 보이며, 캐시 읽기 비용을 4분의 1로 낮춰 일반 작업 약 25%, 복잡한 에이전틱 워크플로우에서는 최대 45% 비용을 절감했습니다. 또한 워터마크 내장과 사이버보안 등 특수 접근 프로그램 확대도 주요 변화입니다.

번역된 본문

안스로픽이 클로드 페이블 5.1과 미토스 5.1이라는 최강의 AI 모델을 출시했습니다. 에이전틱 코딩과 텍스트 품질이 향상된 것과 함께, 비용도 최대 45% 절감했습니다.

이전 세대와 마찬가지로 두 5.1 모델은 동일한 기반 모델을 공유하지만 안전 가드레일에서 차이가 있습니다. 페이블 5.1은 일반적으로 이용 가능하며, 미토스 5.1은 사이버보안과 생명과학 분야의 특별 접근 프로그램으로 제한됩니다. 또한 최초로 워터마크를 내장한 클로드 모델이기도 합니다. 안스로픽은 비공개 프리뷰로 탐지 API를 출시해 규제기관, 언론, 팩트체커, 연구기관이 텍스트에 워터마크가 포함되어 있는지 검증할 수 있게 했습니다. 회사는 점진적으로 접근을 확대할 계획이며, 관심 있는 기관은 별도로 신청할 수 있습니다.

페이블 5.1은 일반적인 워크로드에서 페이블 5보다 약 25% 저렴하며, 많은 도구 호출이 포함된 긴 자율 실행 같은 고도의 에이전틱 작업에서는 절감 폭이 약 45%까지 늘어납니다. 이는 캐시 읽기 비용을 백만 토큰당 1달러에서 0.25달러로 대폭 인하했기 때문입니다. 나머지 API 가격은 변경 없이 백만 입력 토큰당 10달러, 백만 출력 토큰당 50달러입니다. 참고로 오퍼스 5는 백만 입력 토큰당 5달러, 출력 토큰당 25달러로 절반 가격입니다.

높은 비용은 페이블 5에 대한 가장 큰 불만이었으며, 기업 고객들 사이에서 채택률이 낮은 원인으로 추정됩니다. 오퍼스 5가 7월 말 출시되면서 더 낮은 가격으로 대부분의 벤치마크에서 페이블 5와 맞먹거나 능가하면서 가격 압박은 계속 커져 왔습니다.

기존 클로드 모델처럼 새 버전에도 연산량을 조절하는 노력 수준(effort level) 시스템이 있습니다. 낮거나 중간 노력 수준에서는 페이블 5.1이 더 낮은 비용으로 페이블 5와 동등한 결과를 낼 것으로 기대됩니다.

코딩·연구 벤치마크에서 큰 도약

페이블 5.1은 에이전틱 벤치마크에서 큰 향상을 보였습니다. Terminal-Bench-Science 0.1에서는 52.6%를 기록해 페이블 5의 24.7%를 두 배 이상 뛰어넘었고, 22.4%의 GPT-5.6 Sol도 크게 앞섰습니다. 에이전틱 코딩 벤치마크인 Terminal-Bench 4.0에서는 페이블 5.1이 55.8%, 미토스 5.1이 60.9%를 기록해 페이블 5의 42.0%와 GPT-5.6 Sol의 37.3%를 웃돌았습니다. 이러한 성과가 실제 사용 환경에서도 같은 규모로 이어질지는 앞으로 몇 주간 지켜봐야 알 수 있습니다.

주요 벤치마크 결과:

  • 에이전틱 과학 연구 Terminal-Bench-Science 0.1: 페이블 5.1 52.6%, 페이블 5 24.7%, 오퍼스 5 29.0%, GPT-5.6 Sol 22.4%
  • 에이전틱 코딩 Terminal-Bench 4.0: 페이블 5.1 55.6%/미토스 5.1 60.9%, 페이블 5 42.0%, 오퍼스 5 52.3%, GPT-5.6 Sol 37.3%
  • 지식 작업 GDPval-AA v2: 페이블 5.1 1853점, 페이블 5 1723점, 오퍼스 5 1824점, GPT-5.6 Sol 1711점
  • 컴퓨터 사용 OSWorld 2.0(부분): 페이블 5.1 77.9%, 페이블 5 72.9%, 오퍼스 5 75.4%
  • 컴퓨터 사용 OSWorld 2.0(엄격): 페이블 5.1 41.7%, 페이블 5 36.1%, 오퍼스 5 39.6%
  • 다학제 추론 Humanity's Last Exam(도구 없음): 페이블 5.1 60.9%, 페이블 5 57.8%, 오퍼스 5 56.6%
  • 다학제 추론 Humanity's Last Exam(도구 사용): 페이블 5.1 65.0%, 페이블 5 63.8%, 오퍼스 5 63.6%
  • 비즈니스 워크플로우 AutomationBench: 페이블 5.1 31.4%, 페이블 5 17.1%, 오퍼스 5 26.9%, GPT-5.6 Sol 19.6%
  • 에이전틱 코딩 CursorBench 3.2.0: 페이블 5.1 73.4%, 페이블 5 70.5%, 오퍼스 5 70.0%, GPT-5.6 Sol 67.2%

안스로픽 연구원 펠릭스 리제베르크(Felix Rieseberg)에 따르면 페이블 5.1은 글쓰기 스타일도 개선되었습니다. 이전 모델들은 채팅에서 글머리 기호와 굵은 글씨에 지나치게 의존했는데, 페이블 5.1은 이를 줄이고 스타일 지시를 더 충실히 따르며 전반적으로 더 자연스럽게 들립니다.

페이블 5.1, 사이버보안 작업 개방

사이버보안, 생물학, 의학 질문에 대한 안전 필터가 이전 버전보다 덜 공격적입니다. 5.0 모델에서는 필터가 너무 민감해서 잘 알려진(원문 여기서 잘림)

원문 보기
원문 보기 (영어)
Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent less Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 1, 2026 Key Points Anthropic has released two new models, Claude Fable 5.1 and Mythos 5.1, delivering stronger coding performance and better text quality. Cheaper cache reads cut costs across the board, saving about 25 percent on typical tasks and enabling savings of up to 45 percent on complex agentic workflows. Fable 5.1 outperforms its predecessor Fable 5 and rival GPT-5.6 Sol on agentic benchmarks. It's available now. Ask about this article… Search Anthropic launches Claude Fable 5.1 and Mythos 5.1, its most capable AI models yet. Along with gains in agentic coding and text quality, the company cuts costs by up to 45 percent. Like their predecessors, both 5.1 models share the same base model but differ in safety guardrails. Fable 5.1 is broadly available, while Mythos 5.1 is restricted to special access programs for cybersecurity and life sciences. They're also the first Claude models to ship with built-in watermarks . Anthropic is launching a detection API in private preview that lets regulators, media outlets, fact-checkers, and research institutions verify whether a text contains the watermark. The company plans to expand access over time, and interested parties can sign up here . Ad Fable 5.1 costs about 25 percent less than Fable 5 for typical workloads, with savings climbing to roughly 45 percent for heavily agentic tasks involving long, autonomous runs with many tool calls. Anthropic made that possible by slashing cache reads from $1 to $0.25 per million tokens. All other API prices remain unchanged at $10 per million input tokens and $50 per million output tokens. For comparison, Opus 5 runs at half that price with $5 input and $25 output per million tokens. Ad High cost was the biggest complaint about Fable 5, and it likely contributed to the model seeing low adoption among enterprise customers . The price pressure has been building since Opus 5 launched in late July, already matching or beating Fable 5 on most benchmarks at a lower price. Like earlier Claude models, the new versions come with an effort-level system that controls compute usage. At low or medium effort, Fable 5.1 should match Fable 5's results at lower cost. Big jumps in coding and research benchmarks Fable 5.1 posts major gains on agentic benchmarks. On Terminal-Bench-Science 0.1, the model hits 52.6 percent, more than double Fable 5's 24.7 percent and far ahead of GPT-5.6 Sol at 22.4 percent. On Terminal-Bench 4.0 for agentic coding, Fable 5.1 scores 55.8 percent while Mythos 5.1 reaches 60.9 percent, compared to 42.0 percent for Fable 5 and 37.3 percent for GPT-5.6 Sol. Whether these gains translate to real-world use at the same scale will become clear over the coming weeks. Ad Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol Agentic Scientific Research Terminal-Bench-Science 0.1 52.6% 24.7% 29.0% 22.4% Agentic Coding Terminal-Bench 4.0 55.8% / 60.9% (Mythos 5.1) 42.0% 52.3% 37.3% Knowledge Work GDPval-AA v2 1853 1723 1824 1711 Computer Use OSWorld 2.0 (partial) 77.9% 72.9% 75.4% — Computer use OSWorld 2.0 (strict) 41.7% 36.1% 39.6% — Multidisciplinary reasoning : Humanity's Last Exam (no tools) 60.9% 57.8% 56.6% — Multidisciplinary Reasoning : Humanity's Last Exam (with tools) 65.0% 63.8% 63.6% — Business Workflows AutomationBench 31.4% 17.1% 26.9% 19.6% Agentic Coding CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2% Anthropic researcher Felix Rieseberg says Fable 5.1 also improves its writing style. Earlier models leaned too heavily on bullet points and bold text in chat, and Fable 5.1 dials that back while following style instructions more closely and sounding more natural overall. Fable 5.1 opens up for cybersecurity work The safety filters for cybersecurity, biology, and medical questions are less aggressive than in earlier versions. In the 5.0 models, they were so sensitive that they triggered on well-intentioned requests too, producing false positives at a frustrating rate. Ad The cybersecurity filters in 5.1 generate 60 percent fewer false positives, and for biology-related queries, the filters fire 85 percent less often on harmless questions about basic biology and medicine. Ad Fable 5.1 can now identify software vulnerabilities for the first time, though not develop exploits. Penetration testing and exploit generation still get routed to the Opus models. Mythos is also part of Claude Security for defensive purposes . How to access Fable 5.1 and Mythos 5.1 Claude Fable 5.1 is available immediately on all platforms, including AWS, Google Cloud, and Microsoft Azure. Developers can access it via the API using claude-fable-5-1 . For enterprise customers, Anthropic is rolling out Enterprise Frontier Safeguards (EFS), which store customer data solely on the customer's own cloud infrastructure. Claude Mythos 5.1 is currently limited to US organizations through two programs: the Cyber Verification Program for defensive security work and the Life Sciences Verification Program, developed with the US government. Anthropic plans to expand access to international partners. Anthropic is also cracking down on distillation attacks, where a model's capabilities are systematically extracted through thousands of fake accounts. New API accounts can no longer edit Claude's prior context in multi-turn conversations while keeping the thinking transcript, which according to Anthropic closes a documented distillation technique. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Anthropic