메뉴
BL
The Decoder 27일 전

클로드 소넷 5, 동결된 토큰 단가 뒤에 숨긴 '눈에 띄는 가격 인상'

IMP
8/10
핵심 요약

Anthropic의 새로운 모델인 Claude Sonnet 5는 토큰당 공식 단가는 동일하게 유지하면서도, 복잡한 작업을 수행하기 위해 토큰 소모량이 크게 증가해 결과적으로 실제 사용자 부담 비용은 전작 대비 약 2배 가까이 상승했습니다. 복잡한 추론 작업에서는 여전히 대형 모델들에 뒤처지는 등 기술적 한계가 존재하는 상황에서, 이러한 '숨겨진 비용 상승' 전략은 저렴한 경쟁 모델들과의 가격 경쟁력 측면에서 개발자들에게 중요한 고려 사항이 됩니다.

번역된 본문

클로드 소넷 5, 변하지 않는 토큰 단가 뒤에 숨겨진 가격 인상 패턴을 지속하다

독립 테스트 결과, Claude Sonnet 5는 5위를 차지했으며 일부 에이전트 기반 작업에서는 더 비싼 Opus 4.8을 성능으로 제쳤습니다. 하지만 이 모델은 토큰 소비량이 기하급수적으로 늘어나 작업당 실제 비용은 Anthropic의 기존 최고 수준 모델보다도 더 비싸졌습니다.

평가 기관인 Artificial Analysis는 Claude Sonnet 5의 출시 전 테스트를 진행하고 이를 지능 지수(Intelligence Index)에 반영했습니다. Sonnet 5는 최고 성능 설정에서 53점을 받아 GPT-5.5(고성능 모드)와 공동 5위를 기록했습니다. 더 높은 순위의 4개 모델은 각각 55점의 GPT-5.5(최고 성능 모드), 54점의 Opus 4.7, 56점의 Opus 4.8, 그리고 오늘부터 다시 일반 이용이 가능해진 60점의 Claude Fable 5입니다.

이는 이전 모델인 Sonnet 4.6(47점)에 비해 6점 상승한 수치지만, Sonnet 5가 이 점수에 도달하기 위해서는 훨씬 더 많은 토큰을 소모해야 합니다.

동일한 토큰 가격, 두 배가 되는 실제 비용

표면적으로 Sonnet 5는 이전 모델과 동일한 토큰 가격을 유지하고 있습니다. 백만 입력 토큰당 3달러, 백만 출력 토큰당 15달러이며, 상위 모델인 Opus 4.8은 각각 5달러와 25달러입니다.

그러나 Artificial Analysis에 따르면, 지능 지수 내의 평균적인 작업을 수행할 때 Sonnet 5는 약 2.29달러가 드는 반면 Opus 4.8은 약 1.97달러가 소요됩니다. 최고 성능 설정('max')에서 Sonnet 5는 작업당 이전 모델보다 약 40% 더 많은 출력 토큰을 소모합니다.

특히 AA-Briefcase 및 GDPval-AA와 같은 에이전트 기반 지식 작업 벤치마크에서는 이전 모델보다 약 3배나 더 많은 에이전트 루프를 실행합니다. Sonnet 4.6은 작업당 약 1.20달러가 들었습니다. Sonnet 5가 이러한 작업들 중 일부에서 Opus 4.8보다 뛰어난 성능을 보여주긴 했지만, 비용은 거의 두 배로 뛰어버린 셈입니다.

Anthropic은 9월 1일까지 백만 토큰당 2달러 및 10달러의 프로모션 가격을 적용하고 있지만, Artificial Analysis의 이번 결과는 정가를 기준으로 산출되었습니다.

여전히 복잡한 추론에서 한계를 드러내는 Sonnet 5

Sonnet 5는 추론 및 지식 집약적 벤치마크에서는 여전히 대형 모델들에 뒤처집니다. 미국 아곤 국립연구소(Argonne National Labs)와 일리노이 대학교의 최첨단 물리학 추론 테스트인 CritPt에서 17%의 점수를 기록했습니다. 이는 전작보다 14점이나 오른 점수지만, 고성능 설정의 GLM-5.2, Claude Opus, Fable, GPT-5.5보다는 여전히 낮은 수치입니다.

다른 평가에서는 Sonnet 4.6 대비 확실한 성능 향상을 보여줬습니다. Terminal-Bench v2.1에서 9점, '인류의 마지막 시험(Humanity's Last Exam)'에서 10점, SciCode에서 7점이 각각 상승했습니다. 나머지 평가 지표의 점수는 대체로 비슷하게 유지되었습니다.

숨기고 있는 Anthropic의 가격 인상 전략

Anthropic은 이런 방식을 이미 썼습니다. Opus 4.7이 출시되었을 때도 토큰 가격은 표면상 동일하게 유지되었지만, 새로운 토크나이저(tokenizer)가 동일한 텍스트를 약 '30%' 더 많은 토큰으로 분해하여 실제 청구액을 부풀렸습니다.

개발자인 Abhishek Ray는 1.325배에서 1.47배의 비용 증가를 측정했으며, 483개 이상의 제출물을 분석한 커뮤니티 보고서는 요청당 토큰 사용량이 37.4%나 급증했다는 사실을 발견했습니다.

Sonnet 5에서는 토크나이저 문제에 더해 작업당 훨씬 더 많은 토큰을 소모하는 모델의 '에이전트적 성향'이 겹쳐 문제가 더욱 복잡해졌습니다.

Anthropic의 모델들은 세대를 거듴할 때마다, 때로는 극적으로 비싸지고 있지만 공식 가격표에는 이것이 반영되지 않습니다. Sonnet이 속한 중급 모델 시장에서 Deepseek V4 Pro 및 GLM-5.2 같은 중국 경쟁사들이 저렴한 비용으로 경쟁력 있는 성능을 제공하고 있기 때문에, 이러한 숨겨진 비용 증가는 소비자를 설득하기 어렵습니다.

의미를 잃어가는 단순한 원시 토큰 가격보다는, 표준화된 작업당 비용이나 실제 지식 노동 기반의 비용 체계 등 AI 제공업체들의 보다 투명한 가격 책정이 필요한 시점입니다.

원문 보기
원문 보기 (영어)
Claude Sonnet 5 continues Anthropic's pattern of hiding price increases behind unchanged token rates Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 1, 2026 In an independent test, Claude Sonnet 5 placed fifth and beat the pricier Opus 4.8 on some agent-based tasks. But its massive jump in token consumption makes the model more expensive per task than Anthropic's previous top model. Artificial Analysis evaluated Claude Sonnet 5 before its release and added it to its Intelligence Index. Sonnet 5 scored 53 points at peak performance, tying with GPT-5.5 (high) for fifth place. Four models rank higher: GPT-5.5 (xhigh) at 55, Opus 4.7 at 54, Opus 4.8 at 56, and Claude Fable 5, once again generally available as of today , at 60 points. That's a six-point jump over Sonnet 4.6 (47 points), but Sonnet 5 chews through far more tokens to get there. Same token prices, double the real cost On paper, Sonnet 5 keeps the same token prices as its predecessor: $3 per million input tokens and $15 per million output tokens, while Opus 4.8 sits at $5 and $25. Yet according to Artificial Analysis, an average task in the Intelligence Index costs $2.29 with Sonnet 5, versus about $1.97 with Opus 4.8. At the maximum performance setting ("max"), Sonnet 5 burns through about 40 percent more output tokens per task than Sonnet 4.6. In agent-based knowledge work benchmarks like AA-Briefcase and GDPval-AA , it runs about three times as many agent loops as its predecessor. Sonnet 4.6 cost about $1.20 per task. That's nearly doubled, even though Sonnet 5 beats Opus 4.8 on some of these tasks. Anthropic is running a promotional rate of $2 or $10 per million tokens through September 1, but Artificial Analysis based its results on regular prices. Complex reasoning still exposes Sonnet 5's limits Sonnet 5 still falls short of larger models on reasoning- and knowledge-heavy benchmarks. On CritPt, a frontier physics reasoning test from Argonne National Labs and the University of Illinois, it scored 17 percent. That's 14 points above its predecessor but below GLM-5.2, Claude Opus, Fable, and GPT-5.5 in their higher configurations. Elsewhere, Sonnet 5 shows solid gains over Sonnet 4.6: a 9-point jump on Terminal-Bench v2.1, 10 points on Humanity's Last Exam, and 7 points on SciCode. Scores on the remaining evaluations stayed roughly flat. Anthropic keeps raising prices without saying so Anthropic has done this before. When Opus 4.7 launched , token prices stayed flat on paper, but a new tokenizer chopped the same text into "approximately 30%" more tokens , inflating the real bill. Developer Abhishek Ray measured a 1.325x to 1.47x increase, and a community analysis of over 483 submissions found a 37.4 percent jump in tokens per request. With Sonnet 5, the tokenizer issue is compounded by the model's more agentic behavior, which eats through far more tokens per task. Anthropic's models keep getting pricier with each generation, sometimes dramatically so, yet the official price lists don't reflect it. That kind of hidden cost creep is a hard sell when Chinese competitors like Deepseek V4 Pro and GLM-5.2 offer competitive performance at a fraction of the cost in the mid-range segment where Sonnet sits. AI providers need more transparent pricing, like cost per standardized task or real-world knowledge work job, rather than raw token prices that lose meaning . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->