Anthropic이 2026년 9월 출시한 독점 모델 Claude Opus 5.5(적응형 추론, 최대 노력)이 Artificial Analysis 지능 지수 58점으로 동급 모델 중 1위를 차지했습니다. 다만 가격은 입력 100만 토큰당 4달러, 출력 100만 토큰당 20달러로 중앙값보다 비싸고, 평가 중 2억 6천만 토큰을 생성해 상당히 장황한 편입니다. 텍스트와 이미지 입력을 지원하며 컨텍스트 창은 100만 토큰입니다.
번역된 본문
Claude Opus 5.5(적응형 추론, 최대 노력, 기본 폴백) 지능, 성능 및 가격 분석 (Artificial Analysis, Anthropic, 2026년 9월 출시, 독점 모델)
모델 요약
지능: 206개 모델 중 1위, Artificial Analysis 지능 지수 58점 (지능 항목 만점 4단위 중 4단위)
속도: 초당 출력 토큰 수 미공개 (속도 항목 알 수 없음)
비용: 206개 모델 중 87위, 입력 $4.00/100만 토큰, 출력 $20.00/100만 토큰, 캐시 할인 95%, 지능 지수 과제당 비용 $5.98 (비용 항목 4단위 중 4단위)
장황함: 206개 중 89위, 지능 지수 평가에서 2억 6천만 출력 토큰 생성 (장황함 항목 4단위 중 4단위)
비교 요약
Claude Opus 5.5(적응형 추론, 최대 노력, 기본 폴백)은 지능 면에서 최상위 모델군에 속하지만, 비슷한 가격대의 다른 모델과 비교하면 다소 비싼 편입니다. 이 모델은 텍스트와 이미지 입력을 지원하고 텍스트를 출력하며, 100만 토큰의 컨텍스트 창을 갖추고 있습니다. 지능 지수에서 58점을 기록해 동급 모델의 중앙값(25점)보다 훨씬 높은 수준입니다. 지능 지수 평가 과정에서 2억 6천만 토큰을 생성했는데, 이는 중앙값인 9,200만 토큰에 비해 매우 장황한 수치입니다. 가격은 입력 100만 토큰당 $4.00(중앙값 $2.00 대비 다소 비쌈), 출력 100만 토큰당 $20.00(중앙값 $10.00 대비 다소 비쌈)입니다. 이 모델의 지능 지수 평가에 총 $8,708.20의 비용이 들었습니다.
기술 사양
추론: 지원 (이 페이지는 추론 버전 기준이며, 비추론 버전도 별도로 존재할 수 있음)
입력 모달리티: 텍스트 및 이미지 지원
출력 모달리티: 텍스트 지원
컨텍스트 창: 100만 토큰 (12pt Arial 폰트 기준 약 1,500페이지 분량)
이 클래스의 모델 수: 206개
지표는 같은 클래스의 모델과 비교됩니다:
비추론 모델 → 다른 비추론 모델과만 비교
추론 모델 → 추론 및 비추론 모델 전체와 비교
오픈 웨이트 모델 → 같은 크기 등급의 오픈 웨이트 모델과만 비교 (초소형: 4B 파라미터 이하 / 소형: 4B–40B / 중형: 40B–150B / 대형: 150B 초과)
독점 모델 → 같은 가격대의 독점 및 오픈 웨이트 모델과 비교 (입력:출력 3:1 가격 비율 사용, 100만 토큰당 $0.15 미만 / $0.15–$1 / $1 초과)
하이라이트 (업데이트됨)
지능: Artificial Analysis 지능 지수, 높을수록 좋음
속도: 초당 출력 토큰 수, 높을수록 좋음
과제당 비용: 지능 지수 과제당 가중 평균 비용(USD), 낮을수록 좋음
Artificial Analysis 지능 지수 v4.3.2는 10개 평가를 포함합니다: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1. 661개 모델 중 29개 표시됨. 각 평가의 세부 내용과 실행 방법은 지능 지수 방법론 문서를 참고하세요.
Artificial Analysis K Anthropic • Claude Opus 5.5 max • Proprietary model • Released September 2026 Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) Intelligence, Performance & Price Analysis Compare Try it out API Provider Benchmarks Model summary Intelligence Updated # 1 / 206 58 Artificial Analysis Intelligence Index 4 out of 4 units for Intelligence. Speed N/A Output tokens per second Unknown out of 4 units for Speed. Cost # 87 / 206 In $4.00 Out $20.00 Cache Discount 95% $5.98 Cost per Intelligence Index task 4 out of 4 units for Cost. Verbosity # 89 / 206 260M Output tokens from Intelligence Index 4 out of 4 units for Verbosity. Comparison Summary Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is amongst the leading models in intelligence, but somewhat expensive when comparing to other models of similar price. The model supports text and image input, outputs text, and has a 1M tokens context window. Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) scores 58 on the Artificial Analysis Intelligence Index, placing it well above average among comparable models (median: 25). When evaluating the Intelligence Index, it generated 260M tokens, which is very verbose in comparison to the median of 92M. Pricing for Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) is $4.00 per 1M input tokens (somewhat expensive, median: $2.00) and $20.00 per 1M output tokens (somewhat expensive, median: $10.00). In total, it cost $8708.20 to evaluate Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) on the Intelligence Index. Technical specifications Reasoning Yes This page shows the reasoning version of this model. A non-reasoning variant may also exist. Input modality Supports: text and image Output modality Supports: text Context window 1M ~1500 A4 pages of size 12 Arial font 206 models in this class Metrics are compared against models of the same class: Non-reasoning models → compared only with other non-reasoning models Reasoning models → compared across both reasoning and non-reasoning Open weights models → compared only with other open weights models of the same size class: Tiny: ≤4B parameters Small: 4B–40B parameters Medium: 40B–150B parameters Large: >150B parameters Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio: <$0.15 per 1M tokens $0.15–$1 per 1M tokens >$1 per 1M tokens Highlights Updated Intelligence Artificial Analysis Intelligence Index · Higher is better Speed Output tokens per second · Higher is better Cost per Task Weighted average cost (USD) per Intelligence Index task · Lower is better Prompt Options Intelligence Updated Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 29 of 661 models Add model from specific provider Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights / Proprietary Reasoning / Non-Reasoning Text Only / Multimodal Inputs Artificial Analysis Intelligence Index by Open Weights / Proprietary Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 29 of 661 models Add model from specific provider Proprietary Open Weights Open Weights (Commercial Use Restricted) Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if commercial use is limited by conditions, and as 'Non-commercial' if the license prohibits commercial use. Capability Indexes Measures the performance of models on specific capabilities and industries Finance & Accounting Strategy & Ops Legal Engineering Economics Artificial Analysis Finance & Accounting Index Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity's Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better 29 of 158 models Add model from specific provider Benchmarks Intelligence Evaluations Intelligence evaluations measured independently by Artificial Analysis · Higher is better Coding Agentic Tool Use Private Dataset User Interaction Finance Medical Legal Intelligence Index Long Context Multimodal Instruction Following Faithfulness Writing Business See more 18 of 26 evaluations 29 of 661 models Add model from specific provider AA-Briefcase v1.1 Updated Agentic knowledge work, (Elo-500)/2000 GDPval-AA v2.1 Updated Agentic real-world work tasks, (Elo-500)/2000 AutomationBench-AA Updated Agentic SaaS workflows Terminal-Bench 4.0 New Agentic coding & terminal use SciCode Coding Humanity's Last Exam Reasoning & knowledge GDP.pdf New Professional document reasoning, All-pass CritPt Physics reasoning AA-Omniscience Accuracy Knowledge AA-Omniscience Non-Hallucination Rate 1 - hallucination rate AA-LCR v1.1 Long context reasoning Harvey LAB-AA Legal agentic work, criterion pass rate EnterpriseOps-Gym-AA Agentic business operations AA-AnalystAgent Quantitative analysis on spreadsheets & documents 𝜏³-Banking Agentic tool use ITBench-AA Kubernetes incident root-cause analysis MMMU-Pro Visual reasoning MLCR-AA New Medical long context reasoning Intelligence Evaluation Relevance While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases. Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. AA-Briefcase v1.1 Updated AA-Briefcase Elo AA-Briefcase Rubric Score (%) Analytical Quality & Presentation Elo AA-Briefcase Elo AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better 29 of 179 models Add model from specific provider AA-Briefcase Elo AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0. AA-Omniscience AA-Omniscience Index AA-Omniscience Accuracy AA-Omniscience Hallucination Rate AA-Omniscience Index AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct. 29 of 536 models Add model from specific provider AA-Omniscience Index AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It r