메뉴
BL
The Decoder 4일 전

앤스로픽 클로드 오푸스 5, 대부분 벤치마크서 페이블 5 맞먹거나 능가

IMP
8/10
핵심 요약

앤스로픽의 '클로드 오푸스 5(Claude Opus 5)'가 여러 평가에서 최고 수준의 성능을 기록하며 경쟁 모델인 '페이블 5(Fable 5)'를 압도하는 가성비를 보여주었습니다. 특히 코딩 및 소프트웨어 엔지니어링 부문에서 강세를 보이지만, 환각 현상(거짓 정보 생성) 비율이 50%에 달해 고위험 작업 적용 시 신뢰성에 대한 우려도 함께 제기됩니다. 또한 가장 높은 추론 단계보다 'high' 수준의 설정에서 비용 대비 최고의 효율과 코딩 결과물을 제공하는 것으로 확인되었습니다.

번역된 본문

앤스로픽의 클로드 오푸스 5(Claude Opus 5)는 더 낮은 비용으로 페이블 5(Fable 5)를 능가하며, 여러 벤치마크 평가 기준 오늘날 사용 가능한 가장 뛰어난 AI 모델로 평가받고 있습니다.

주요 요점:

  • 인공지능 분석 업체 Artificial Analysis에 따르면, 앤스로픽의 클로드 오푸스 5는 현재 가장 성능이 뛰어난 모델로, 분석적 품질 및 지식 기반 작업에서 특히 강점을 보이며 지능 지수(Intelligence Index) 61점을 기록했습니다.
  • 오푸스 5는 확신이 없는 경우에도 더 자주 답변을 생성하려는 경향이 있어 환각(Hallucination) 비율이 50%까지 치솟았으며, 이는 고위험 애플리케이션에서의 신뢰성에 대한 의문을 제기합니다.
  • 'high' 및 'xhigh' 성능 수준에서 오푸스 5는 더 낮은 비용을 유지하면서 오푸스 4.8과 소넷 5(Sonnet 5)를 모두 능가합니다.

여러 벤치마크에 따르면 앤스로픽의 클로드 오푸스 5는 오늘날 사용할 수 있는 가장 뛰어난 AI 모델이며, 더 낮은 비용으로 페이블 5의 성능을 뛰어넘습니다. 오푸스 5는 지식 작업, 코딩, 과학적 추론 및 사실적 정확성을 다루는 9가지 테스트를 결합한 Artificial Analysis 지능 지수에서 61점을 받았습니다. 이는 클로드 페이블 5(60점), GPT-5.6 Sol(59점), 키미 K3(57점), 클로드 오푸스 4.8(56점)을 약간 앞서는 수치입니다. Artificial Analysis는 모델이 공식 출시되기 전에 앤스로픽과 협력하여 이를 테스트했습니다.

코딩 부문에서 'xhigh' 수준의 클로드 오푸스 5는 클로드 코드(Claude Code)와 페어링되어 AI 모델이 버그를 찾고 수정하는 것을 포함해 프로그래밍 작업을 얼마나 잘 독립적으로 처리하는지 측정하는 Artificial Analysis 코딩 지수(Coding Index)에서 공동 1위를 차지했습니다. AI 에이전트가 실제 터미널 환경에서 자율 엔지니어로서 어떻게 수행하는지 테스트하는 Terminal-Bench v2.1에서는 'max' 설정에서 89%의 점수를 기록하며 이전 선두 주자인 GPT-5.6 Sol과 타이 기록을 세웠습니다.

과학적 추론의 경우, 오푸스 5는 여러 학문 분야를 아우르는 매우 어려운 지식 테스트인 '인류의 마지막 시험(Humanity's Last Exam)'에서 53%의 점수를 받았으며, 이는 페이블 5와 동일한 수치입니다. 미국 아곤 국립연구소(Argonne National Laboratory)와 일리노이 대학교(UIUC) 연구진이 만든 물리학 벤치마크인 CritPt에서도 다시 페이블 5와 타이를 이루었지만, GPT-5.6 Sol, GPT-5.5 Pro, GPT-5.6 Terra에는 뒤졌습니다.

팩트 정확도는 여전히 약점입니다. 모델의 지식 주장에 대한 정확도를 테스트하는 AA-Omniscience에서 오푸스 5는 오푸스 4.8 대비 7점이 개선되었으나 여전히 페이블 5에 뒤집니다. 또한, 확실하지 않을 때도 답변을 생성하는 빈도가 늘어 환각 발생률을 14점 끌어올려 50%에 달했습니다.

에포크 AI(Epoch AI) 확인, 최첨단 모델 간의 판세는 초접전 에포크 AI 또한 클로드 오푸스 5를 테스트했습니다. 이 연구 기관은 종합 에포크 역량 지수(Epoch Capability Index)에서 159점을 부여했으며, 이는 161점을 받은 페이블 5에 불과하게 뒤지는 수치입니다. 하지만 소프트웨어 엔지니어링 벤치마크(SWE-ECI)만 놓고 보면 오푸스 5는 161점으로 페이블 5와 타이를 이루며 GPT-5.6 Terra와 클로드 오푸스 4.8을 압도했습니다. GPT-5.6 Sol은 두 부문 모두에서 선두를 달리고 있습니다. 전반적으로 이는 최첨단 모델 간의 경쟁이 매우 치열하다는 것을 확인시켜 줍니다. 어느 단일 모델도 턱걸이하여 명확한 우위를 차지할 수 없습니다. 이는 결국 AI 모델이 대중화(Commoditized)될 것이라는 주장에 무게를 실어줍니다.

낮은 추론 단계가 더 나은 가성비와 코딩 결과를 제공 오푸스 5의 평균 지능 지수 작업 비용은 2.03달러로, 폴백(fallback)이 적용된 클로드 페이블 5의 2.75달러보다 저렴합니다. 물론 오푸스 4.8의 1.80달러 및 소넷 5의 1.53달러보다는 비쌉니다. 그러나 'high' 및 'xhigh' 단계에서 오푸스 5는 더 낮은 비용을 유지하면서 오푸스 4.8과 소넷 5를 모두 능가합니다.

Vals.ai는 프로그래밍 작업을 위한 벤치마크인 Vibe Code Bench를 사용하여 5가지 추론 단계에 걸쳐 클로드 오푸스 5를 테스트했습니다. 점수는 'low' 단계에서 76.7%로 시작해 'medium'에서 82%, 'high'에서 89.8%로 상승했습니다. 하지만 훨씬 높은 비용에도 불구하고 가장 높은 두 단계에서는 성능이 떨어져 'xhigh'는 88.3%, 'max'는 88.4%를 기록했습니다. Vals.ai는 가장 높은 단계일수록 오류를 포함할 확률이 더 높은 복잡한 솔루션을 생성하는 경향이 있다는 것을 발견했습니다. 반면 'high' 단계는 요구 사항을 더욱 안정적으로 충족하는 더 단순한 솔루션을 만들어냅니다. Terminal-Bench 2.1 역시 비슷한 패턴을 보여줍니다. 최상위 단계일수록 각 시도에 더 많은 시간을 소비하여 주어진 시간 내에 시도할 수 있는 횟수가 줄어들기 때문에, 'high' 단계가 'max' 단계를 능가하는 결과를 보였습니다.

원문 보기
원문 보기 (영어)
Anthropic's Claude Opus 5 costs well below Fable 5 while matching or beating it across most benchmarks Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 25, 2026 Key Points According to Artificial Analysis, Anthropic's Claude Opus 5 is currently the most capable model, achieving an Intelligence Index score of 61, with particular strengths in analytical quality and knowledge-based tasks. Because Opus 5 tends to respond more frequently even when uncertain, its hallucination rate climbs to 50 percent, raising questions about reliability in high-stakes applications. At the "high" and "xhigh" performance levels, Opus 5 outperforms both Opus 4.8 and Sonnet 5 while maintaining lower costs. Ask about this article… Search Anthropic's Claude Opus 5 is the most capable AI model available today, according to several benchmarks, outperforming Fable 5 while costing less. Opus 5 scored 61 on the Artificial Analysis Intelligence Index , which combines nine tests covering knowledge work, coding, scientific reasoning, and factual accuracy. That puts it just ahead of Claude Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and Claude Opus 4.8 (56). Artificial Analysis worked with Anthropic to test the model before its public release. In coding, Claude Opus 5 at "xhigh" paired with Claude Code shares first place on the Artificial Analysis Coding Index, which measures how well AI models handle programming tasks on their own, including finding and fixing bugs. On Terminal-Bench v2.1, which tests agents as autonomous engineers in real terminal environments, Opus 5 scored 89 percent at "max," matching the previous leader GPT-5.6 Sol. Ad For scientific reasoning, Opus 5 scored 53 percent on Humanity's Last Exam, a very difficult knowledge test covering many academic fields. That ties it with Fable 5. On CritPt, a physics benchmark from researchers at Argonne National Laboratory and UIUC, it again matches Fable 5 but trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra. Ad DEC_D_Incontent-1 Factual accuracy remains a weak spot: On AA-Omniscience, which tests the accuracy of a model's knowledge claims, Opus 5 improved by 7 points over Opus 4.8 but still trails Fable 5. Opus 5 also answers more often when it's uncertain, pushing its hallucination rate up 14 points to 50 percent. Epoch AI confirms tight race among frontier models Epoch AI has also tested Claude Opus 5. The research institute gave it an overall Epoch Capability Index score of 159, just below Fable 5 at 161. When looking solely at the software engineering benchmarks (SWE-ECI), though, Opus 5 ties with Fable 5 at 161, outperforming GPT-5.6 Terra and Claude Opus 4.8. GPT-5.6 Sol leads in both categories. Ad Overall, this confirms that the race among frontier models is tight. No single model can pull away or claim a clear advantage. That lends weight to the argument that AI models will eventually become commoditized . Lower reasoning tiers deliver better value and coding results The average Intelligence Index task costs $2.03 with Opus 5, less than Claude Fable 5 with fallback at $2.75. It costs more than Opus 4.8 at $1.80 and Sonnet 5 at $1.53. At the "high" and "xhigh" tiers, however, Opus 5 beats both Opus 4.8 and Sonnet 5 while costing less. Ad DEC_D_Incontent-2 Vals.ai tested Claude Opus 5 across all five reasoning tiers using Vibe Code Bench, a benchmark for programming tasks. Scores climb from 76.7 percent at "low" to 82 percent at "medium" and 89.8 percent at "high." Performance dips at the two highest tiers, with "xhigh" scoring 88.3 percent and "max" scoring 88.4 percent despite much higher costs. Ad Vals.ai found that the highest tiers tend to produce more complex solutions that contain errors more often. The "high" tier produces simpler solutions that meet the requirements more reliably. Terminal-Bench 2.1 shows a similar pattern. The "high" tier beats "max" because the model spends more time on each attempt at the top tier, leaving fewer attempts within the time limit. That matches Anthropic's guidance, with "high" set as the default tier in both the API and Claude Code. Token pricing stays at $5 per million input tokens and $25 per million output tokens. Cache writes cost $6.25 per million tokens with a five-minute lifetime, while cache hits run just $0.50 per million tokens. Opus 5 pulls ahead in knowledge work Opus 5 performs especially well on the AA-Briefcase benchmark, which measures how well AI models handle typical office tasks like writing research reports, building presentations, and analyzing spreadsheets based on thousands of input files. Performance is scored across correctness, analytical quality, and presentation quality, then rolled into an Elo rating similar to chess rankings. At max reasoning, Opus 5 reaches an Elo of 1720, a full 146 points ahead of Claude Fable 5 (1574). Its three highest tiers (max, xhigh, high) sweep the top three spots. Combined with Fable 5, Sonnet 5, and Opus 4.8, Anthropic now holds the vast majority of top-10 positions. Cost per task dropped 20 percent to $17.79, down from $22.30 for Fable 5. The "xhigh" variant costs $14.26, and "high" comes in at just $10.41, less than half of Fable 5. Both still beat Fable 5 in the Elo ranking. At medium performance, Opus 5 hits an Elo of 1470, just behind GPT-5.6 Sol (max, 1505). At low performance, it lands at 1223 Elo, slightly below GLM-5.2 (max, 1254). Opus 5's biggest gains show up in analytical quality. At "max," it reaches an Analytical Quality Elo of 2016, almost 300 points ahead of Fable 5. Its Rubric Pass Rate, which tracks how often the model meets predefined quality standards, sits at 58 percent ("max"), 57.2 percent ("xhigh"), and 56 percent ("high"). Presentation quality is a different story. Opus 5 scores a Presentation Elo of 1628, about 40 points behind GPT-5.6 Sol at "max" (1666). More performance takes more time: At "max," Opus 5 needs over 36 minutes per task and averages 103 passes, about 50 percent longer than Opus 4.8 at 24 minutes and 55 passes. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: AA 1 | AA 2 | Epoch AI