메뉴
BL
The Decoder 46일 전

성능은 5.7% 오르고 가격은 2배, 앤스로픽 '클로드 페이블 5'

IMP
7/10
핵심 요약

앤스로픽의 최신 AI 모델인 클로드 페이블 5(Claude Fable 5)가 여러 벤치마크를 석권하며 인공지능 성능 1위를 차지했습니다. 하지만 전작 대비 고작 5.7%의 성능 향상을 위해 토큰당 가격이 2배로 인상되어 가성비 논란이 예상됩니다. 이제 기업들은 실제 업무에 이 비싼 모델을 도입할지 신중한 경제성 평가에 나서야 할 시점입니다.

번역된 본문

앤스로픽의 최신 AI 모델인 '클로드 페이블 5(Claude Fable 5)'가 GPT-5.5 등 경쟁 모델들을 제치고 Artificial Analysis 인텔리전스 지수 1위를 차지했습니다. 전작인 오퍼스 4.8(Opus 4.8)과 비교했을 때 전반적인 성능 향상은 여러 벤치마크에서 평균 5.7%에 그쳤습니다.

하지만 이 미미한 성능 향상에는 뼈아픈 비용이 따릅니다. 백만 토큰당 입력 10달러, 출력 50달러로 가격이 2배로 인상되어 전체 벤치마크 실행 비용이 오퍼스 4.8(최대 추론 시 4,970달러)의 두 배인 약 9,940달러에 육박합니다. 오퍼스 4.8과 4.7 역시 4.6에 비해 적은 성능 향상에도 가격을 크게 올린 바 있으며, 앤스로픽 스스로도 4.8의 개선을 '겸손하지만 실질적인 수준'이라고 평분한 바 있습니다.

기업들은 약 5%의 성능 향상을 위해 두 배의 비용을 지불할 만한 사용 사례가 무엇인지 신중하게 저울질해야 합니다. 벤치마크 회의론자들은 어떤 테스트 스위트도 실제 현실 세계의 능력을 완벽하게 반영할 수 없다고 지적하지만, 최소한 Artificial Analysis(AA) 지수는 10개의 평가를 종합하여 단일 벤치마크보다는 폭넓은 기반을 제공합니다. 기업의 과도한 사용량에 따른 월간 청구액은 숙련된 개발자를 고용하는 비용과 맞먹을 수 있으며, '토크노믹스(Tokeneconomics)'라는 측면에서 경제성은 점차 핵심 요소가 되고 있습니다.

그럼에도 불구하고 원시 벤치마크 수치는 주목할 만합니다. 페이블 5는 10개의 인텔리전스 지수 벤치마크 중 5개에서 새로운 기록을 세웠습니다. 지식 및 환각(Hallucination) 벤치마크인 AA-Omniscience에서 이전 최고 기록인 제미나이 3.1 프로 프리뷰(Gemini 3.1 Pro Preview)보다 7점 높은 40점을 기록했습니다. 이 우위는 주로 낮은 환각륭이 아니라 더 높은 정확도에서 비롯되며, 환각륭 자체는 중간 수준에 머물렀습니다. Artificial Analysis는 오픈 웨이트(Open-weight) 모델 간에 AA-Omniscience 정확도와 모델 크기 사이에 강력한 상관관계가 있다고 지적했는데, 이는 페이블 5가 기존의 어떤 공개된 앤스로픽 모델보다 클 수 있음을 시사합니다.

에이전트(Agentic) 작업에서도 페이블 5는 앤스로픽의 우위를 더욱 확고히 했습니다. 실제 지식 작업 벤치마크인 GDPval-AA에서 오퍼스 4.8(1,890점) 대비 2.2% 상승한 1,932 Elo를 기록했으며, 에이전트 코딩을 위한 Terminal-Bench Hard와 도구 사용을 위한 Tau2-bench Telecom에서도 1위를 차지했습니다.

'인류의 마지막 시험(Humanity's Last Exam, HLE)'에서는 53%를 기록하며 오퍼스 4.8을 7점 이상 앞섰습니다. 단 한 번의 HLE 실행(대체 라우팅 포함) 비용은 약 2,200달러로, Artificial Analysis가 테스트한 모델 중 가장 비싸며 기존 오퍼스 모델의 최대 비용(1,974달러)을 크게 웃돕니다.

한편, 페이블 5의 비용을 더욱 높이는 요인으로 안전 필터가 작동합니다. 앤스로픽에 따르면 페이블 5는 클로드 미토스 5(Claude Mythos 5)와 동일한 기본 모델을 사용하지만, 사이버 보안, 생물학, 화학, 모델 증류(Model distillation)와 관련된 쿼리에 대해서는 추가적인 안전 장치가 적용됩니다. 필터가 작동하면 대체(Fallback) 메커니즘이 요청을 오퍼스 4.8으로 우회하는데, 이렇게 우회된 요청 역시 요금 청구에 포함되어 전체 비용을 상승시킵니다.

앤스로픽은 세션의 5% 미만만이 이 영향을 받는다고 밝혔습니다. 그러나 Artificial Analysis가 인텔리전스 지수를 평가하는 동안 측정된 우회 비율은 작업의 약 8%였으며, 주로 GPQA, AA-Omniscience, 그리고 인류의 마지막 시험(HLE)의 과학 질문에서 발생했습니다. 특히 HLE 테스트만 놓고 보면 우회율은 무려 9%에 달했습니다.

원문 보기
원문 보기 (영어)
Anthropic's Claude Fable 5 costs twice as much for 5.7 percent more performance Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 12, 2026 Key Points Anthropic's latest AI model, Claude Fable 5, has claimed the top spot in the Artificial Analysis Intelligence Index, surpassing rivals including GPT-5.5. The overall measured performance leap compared to its predecessor, Opus 4.8, is 5.7 percent across many benchmarks. That performance gain comes at a steep cost: token prices have doubled, with a full benchmark run now approaching $10,000, twice the price of Opus 4.8. Ask about this article… Search Claude Fable 5 tops the Artificial Analysis Intelligence Index and sets new highs across several benchmarks. But the gain over its predecessor might be slim, while costs more than double. Anthropic's new flagship model, Claude Fable 5 , scores 64.9 points in the Artificial Analysis Intelligence Index, claiming first place. The gap to the best non-Anthropic model, GPT-5.5, is about five points. Anthropic now holds the top two spots on the leaderboard. That crown comes at a cost. Fable 5 runs $10 and $50 per million input and output tokens, double Opus 4.8's $5 and $25. A full index run hits $9,940, versus $4,970 for Opus 4.8 at max reasoning. That premium buys a 5.7 percent performance gain. Opus 4.8 and 4.7 already followed the same pattern versus Opus 4.6, with steep price bumps for small gains . Anthropic itself called 4.8's improvement over 4.7 "modest but tangible." Ad Companies need to weigh carefully which use cases actually justify paying double for about five percent more performance. Benchmark skeptics will note that no test suite fully captures real-world ability. The AA Index at least aggregates ten evaluations, giving it a broader base than any single benchmark. Ad DEC_D_Incontent-1 Depending on the region, the monthly bill for heavy enterprise use could cover the cost for experienced developers. The Artificial Analysis data makes clear that economics is becoming a key factor, as our subscriber special on "Tokeneconomics" explores in depth. Top scores across most benchmarks The raw benchmark numbers are nonetheless notable. Fable 5 sets records in five of the ten Intelligence Index benchmarks. On AA-Omniscience , the knowledge and hallucination benchmark, the model hits 40 points, seven more than the previous leader Gemini 3.1 Pro Preview. That lead comes mainly from higher accuracy, not a lower hallucination rate. On hallucinations, the model lands squarely in the middle of the pack. Ad Artificial Analysis notes a strong link between AA-Omniscience accuracy and model size among open-weight models. That hints Fable 5 may be larger than any previous public Anthropic model. On agentic tasks, Fable 5 widens Anthropic's lead. On GDPval-AA, a real-world knowledge work benchmark, it reaches an Elo of 1,932, up 2.2 percent from Opus 4.8 at 1,890. It also tops Terminal-Bench Hard for agentic coding and Tau2-bench Telecom for tool use. Ad DEC_D_Incontent-2 On Humanity's Last Exam, the model scores 53 percent, over seven points ahead of Opus 4.8. A single HLE run with fallback costs about $2,200, the most expensive of any model Artificial Analysis has tested. Previous Opus models topped out at $1,974. Ad Safety filters drive costs up even more Fable 5 uses the same base model as Claude Mythos 5 according to Anthropic, plus extra safeguards for queries touching cybersecurity, biology, chemistry, and model distillation. When a filter trips, a fallback mechanism reroutes the request to Opus 4.8. Those rerouted requests still count toward billing, pushing total costs higher. Anthropic says fewer than five percent of sessions are affected. But Artificial Analysis measured fallback routing in about eight percent of tasks during its Intelligence Index evaluation, mainly on science questions from GPQA, AA-Omniscience, and Humanity's Last Exam. On the HLE test alone, the fallback rate hit nine percent. Access comes with an expiration date Fable 5 keeps the same one-million-token context window as Opus 4.8. Pro, Max, Team, and Enterprise subscribers can use it through June 22, with usage counting at double the Opus rate. After that, it moves to credit-based billing. That makes it even pricier than token rates suggest. Anthropic says it'll bring back subscription access once capacity allows. My colleague Max recently analyzed Fable 5's strengths and weaknesses and found the safety filters blocking large numbers of harmless requests, from medical physics questions to basic security reviews. Anthropic's system card also revealed invisible throttling that degrades Fable's performance when users try to build competing frontier models, though Anthropic has since walked that back . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Artificial Analysis