메뉴
BL
The Decoder 20일 전

앤스로픽 '클로드 페이블 5', 업계 벤치마크 석권... 단, 고비용

IMP
7/10
핵심 요약

AI 벤치마크 플랫폼 인공 분석(Artificial Analysis)이 금융, 법률 등 산업별 성능 평가 지수를 발표했으며, 앤스로픽의 '클로드 페이블 5'가 전 부문 1위를 차지했습니다. 하지만 클로드는 타 모델 대비 압도적인 성능 우위를 점친 것은 아니며, 작업당 비용이 최대 100배 이상 높게 책정되어 기업들의 비용 효율성 고민이 커지고 있습니다.

번역된 본문

앤스로픽의 '클로드 페이블 5(Claude Fable 5)', 비싼 가격 속 새로운 업계 벤치마크 석권 작성자: 마티아스 바스티안 (Matthias Bastian) 2026년 7월 8일

AI 벤치마크 플랫폼 인공 분석(Artificial Analysis)이 AI 모델의 산업별 성능을 평가하는 6개의 새로운 지수를 발표했습니다. 앤스로픽의 '클로드 페이블 5'가 모든 부문에서 선두를 차지했지만, 더 저렴한 대안 모델들은 극히 일부의 비용으로도 유사한 작업을 수행할 수 있는 것으로 나타났습니다.

현재 AI 모델들은 금융, 법률, 의학과 같은 특정 산업 분야의 실무 작업에서 얼마나 우수한 성능을 발휘할까요? 인공 분석은 재무·회계, 법률, 헬스케어·의료, 전략·운영, 엔지니어링, 경제 6개 분야의 AI 모델을 비교하는 새로운 역량 지수(Capability Indices)를 도입했습니다. 이 새로운 지수는 해당 플랫폼의 기존 에이전트(Agentic) 및 코딩(Coding) 지수를 확장한 것입니다.

인공 분석에 따르면, 이 평가 방법론은 미국 O*NET 고용 정보 시스템의 직업 분류를 기반으로 합니다. 재무 모델링, 법률 연구 및 계약서 검토, 임상 진료 지원 등을 포함하는 특정 분야의 핵심 기술은 해당 직무에서 요구되는 업무에서 도출됩니다. 각 분야별로 벤치마크 테스트는 새롭게 구성되며, 특정 기술이 해당 산업에서 얼마나 자주 쓰이는지에 따라 가중치가 부여됩니다. 인공 분석은 모든 벤치마크가 독립적으로 실행된다고 밝혔습니다.

클로드 페이블 5, 8개 지수 전체 석권 앤스로픽의 '클로드 페이블 5(Opus 4.8 대체 모드 포함)'는 8개 지수 모두에서 1위를 차지했습니다. 인공 분석에 따르면, '클로드 오퍼스 4.8(max)'는 8개 중 6개 부문에서 2위를 기록했으며, 나머지 2개 부문에서는 오픈AI의 'GPT-5.5(xhigh)'가 2위를 차지했습니다. 내일 출시될 예정인 'GPT-5.6'이 앤스로픽의 현재 모델과 GPT-5.5 사이의 격차를 좁힐 수 있을 것으로 보입니다.

최상위권 아래의 순위는 분야별로 크게 달라집니다. 구글의 '제미나이 3.5 플래시(Gemini 3.5 Flash)', '제미나이 3.1 프로 프리뷰', 오픈AI의 'GPT-5.5(xhigh)', 앤스로픽의 '클로드 소네트 5(max)', 중국의 'GLM-5.2(max)' 등은 작업의 성격에 따라 순위가 엇갈리고 있습니다. 오픈웨이트(Open-weights) 모델 중에서는 'GLM-5.2(max)'가 6개 산업 지수 중 5개에서 선두를 달리고 있습니다. 53점을 기록한 'GLM-5.2(max)'는 엔지니어링 벤치마크 전체 5위에 올랐으며, 공동 55점을 받은 '클로드 소네트 5(max)' 및 'GPT-5.5(xhigh)'와 불과 2점 차이입니다. 전략 및 운영 지수에서는 '딥시크 V4 프로(DeepSeek V4 Pro, max)'가 38점으로 오픈웨이트 모델 중 선두를 차지했습니다.

인공 분석의 결과는 수백만 명의 실제 사용자가 블라인드 테스트를 통해 모델을 평가하는 독립 플랫폼인 엘엠 아레나(LMArena)의 최신 데이터와 일치합니다. 7월 7일 기준 리더보드에 따르면, '클로드 페이블 5'는 텍스트 아레나, 코드 아레나, 에이전트 아레나에서 모두 1위를 차지했습니다. 앤스로픽은 이 세 가지 주요 부문을 모두 석권한 유일한 연구소입니다. 에이전트 아레나에서 '페이블 5'는 모델 평균보다 16.58% 높은 점수를 기록하며 8.66%를 기록한 오픈AI의 'GPT-5.5 xHigh'와 6.62%의 Z.ai 'GLM 5.2'를 크게 앞섰습니다. 오픈AI는 텍스트 아레나에서 10위, 구글은 7위에 그쳤으며, 딥시크는 38위에 머물렀습니다.

약간의 성능 우위, 따라오는 압도적인 가격표 최고 수준의 성능은 그에 상응하는 높은 비용을 동반하며, 이는 앤스로픽의 최첨단 모델에서 특히 두드러집니다. 비용 분석에 따르면, '딥시크 V4 플래시(DeepSeek V4 Flash, max)'는 6개 지수에 걸친 작업을 작업당 0.04달러 미만의 비용으로 처리하며 중간 수준의 점수를 기록했습니다. 'GLM-5.2(max)'는 작업당 0.26달러에서 0.58달러 사이의 비용으로 오픈웨이트 모델 중 최고의 성능을 제공했습니다.

반면, '클로드 페이블 5'는 전략 및 운영 지수에서 작업당 3.48달러의 비용이 발생합니다. 이는 작업당 0.03달러가 드는 '딥시크 V4 프로(max)'보다 100배 이상 비싼 가격이지만, 얻을 수 있는 점수 우위는 겨우 12점에 불과합니다. 이전의 벤치마크에서도 이미 미미한 성능 향상에 비해 '페이블 5'의 가격이 지나치게 높다는 평가가 있었습니다. 이 성능 격차가 비용을 정당화할 수 있는지에 대해서는 기업들 사이에서도 뜨거운 논쟁이 진행 중입니다. 한 가지 대안으로는 여러 모델을 혼합하여 사용하는 방법입니다. 성능이 뛰어난 오케스트레이터(orchestrator)를 활용해 가성비가 좋은 저렴한 작업용 모델에 태스크를 분배하는 것입니다. 최첨단 모델은 특정 언어 모델이 해당 작업을 해결할 수 있는지 확인하는 '최초 검증(First check)' 용도로만 활용할 수도 있습니다. 해결 가능성이 확인되면, 다음 단계는 가장 저렴한

원문 보기
원문 보기 (영어)
Anthropic's Claude Fable 5 dominates new industry benchmarks at a steep premium Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 8, 2026 Artificial Analysis Ask about this article… Search Benchmarking platform Artificial Analysis has released six new industry-specific performance indices for AI models. Anthropic's Claude Fable 5 leads every category, but cheaper alternatives handle tasks at a fraction of the cost. How well do current AI models perform on industry-specific tasks in fields like finance, law, or medicine? Artificial Analysis has introduced six new Capability Indices that compare AI models across Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, and Economics. The new indices expand on the platform's existing Agentic and Coding indices. According to Artificial Analysis , the methodology is based on occupational classifications from the US O*NET system . Domain-specific skills are derived from the job tasks defined there, covering things like financial modeling, legal research and contract analysis, or clinical decision support. For each domain, the benchmark suite is assembled fresh and weighted by how often a given skill shows up in that industry. Artificial Analysis says all benchmarks are run independently. Ad Claude Fable 5 leads all eight indices Anthropic's Claude Fable 5 (with Opus 4.8 fallback) takes first place across all eight indices. Claude Opus 4.8 (max) comes in second in six of eight categories, according to Artificial Analysis, while OpenAI's GPT-5.5 (xhigh) grabs second in the remaining two. GPT-5.6, set to launch tomorrow , could close the gap between Anthropic's current models and GPT-5.5. Ad DEC_D_Incontent-1 Below the top tier, rankings shift considerably by domain. Google's Gemini 3.5 Flash, Gemini 3.1 Pro Preview, OpenAI's GPT-5.5 (xhigh), Anthropic's Claude Sonnet 5 (max), and the Chinese GLM-5.2 (max) trade places depending on the task. Among open-weights models, GLM-5.2 (max) leads in five of the six industry indices. With 53 points, GLM-5.2 (max) reaches fifth place overall in the Engineering benchmark, just two points behind Claude Sonnet 5 (max) and GPT-5.5 (xhigh), which both score 55. In the Strategy & Ops Index , Deepseek V4 Pro (max) takes the open-weights lead with 38 points. Ad The Artificial Analysis results track with current data from LMArena , an independent platform where millions of real users rank models through blind comparisons. As of the July 7 leaderboard snapshot, Claude Fable 5 holds first place in the Text Arena, Code Arena, and Agent Arena. Anthropic is the only lab leading all three main categories. In the Agent Arena, Fable 5 scores 16.58 percent above the model average, well ahead of OpenAI's GPT-5.5 xHigh at 8.66 percent and Z.ai's GLM 5.2 at 6.62 percent. OpenAI sits at 10th in the Text Arena, Google at 7th. DeepSeek lands at 38th. Ad DEC_D_Incontent-2 A small quality edge comes with a much bigger price tag Top performance comes at a premium, and that's especially true for Anthropic's frontier models. DeepSeek V4 Flash (max) handles tasks across all six indices for less than $0.04 per task, according to the cost analysis, while scoring in the mid-range. GLM-5.2 (max) offers the best open-weights performance at costs between $0.26 and $0.58 per task. Ad Claude Fable 5, by contrast, costs $3.48 per task in the Strategy & Ops Index. That's over 100 times more than DeepSeek V4 Pro (max) at $0.03, for a lead of just 12 points. A previous benchmark already showed the steep premium for Fable 5 relative to its modest performance gains . Whether that gap justifies the price is already a live debate among enterprises . One option is to pair models together, using a capable orchestrator to hand off tasks to cheaper worker models that deliver better value per dollar. Frontier models can also serve as a first check on whether a language model can solve a task at all. Once that's confirmed, the next step is finding the cheapest model that still gets the job done. Some voices, like US economist and LLM enthusiast Ethan Mollick , argue that the performance leap of a model like Fable 5 simply can't be captured by current benchmarks yet. There's no evidence to back that claim so far. All indices, including full weightings and benchmark components, are available on the Artificial Analysis website , along with the documented methodology. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Artificial Analysis