메뉴
BL
The Decoder 4일 전

앤스로픽, 클로드 오퍼스 5 공개... 페이블 5 대비 절반 가격

IMP
8/10
핵심 요약

앤스로픽이 새로운 플래그십 모델인 클로드 오퍼스 5를 공개했습니다. 이 모델은 자체 코드 작성 및 지식 작업 벤치마크에서 최고 수준의 성능을 기록하며, 기존 최상위 모델인 페이블 5에 필적하는 성능을 토큰당 절반 가격에 제공합니다. GPT-5.6 Sol 및 중국 경쟁사들에 대응하여 가성비를 대폭 높린 것이 특징이며, 특히 새로운 문제 해결 능력을 측정하는 ARC-AGI-3 테스트에서 경쟁 모델을 압도하는 성과를 보여주었습니다.

번역된 본문

Anthropic(앤스로픽)은 새로운 플래그십 모델인 Claude Opus 5가 코딩 및 지식 작업에서 최고 점수를 기록하는 동시에, 토큰당 단가는 절반 수준으로 유지하며 Claude Fable 5와 맞먹는 성능을 발휘한다고 밝혔습니다.

앤스로픽은 GPT-5.6 Sol 및 중국 경쟁사들로 인한 가격 경쟁 압력에 대응하고 있습니다. 새로운 Opus 5 모델은 훨씬 비싼 Fable 5와의 가격 대비 성능 격차를 줄이기 위해 설계되었습니다.

Opus 5는 Claude Max의 기본 모델이자 Claude Pro에서 사용할 수 있는 가장 성능이 뛰어난 모델이 됩니다. 100만 토큰의 컨텍스트 창(context window)과 토큰당 단가는 기존과 동일하게 유지됩니다. 이전 모델인 Opus 4.8과 마찬가지로 Opus 5는 100만 입력 토큰당 5달러, 100만 출력 토큰당 25달러의 요금이 책정되었습니다.

또한 새로운 패스트 모드(Fast Mode)는 속도를 2.5배 높이지만 가격은 두 배로 증가합니다.

하지만 실제 비용을 논하려면 토큰 효율성을 고려해야 합니다. 기존 Opus 4.7은 동일한 기본 요금을 적용받았음에도 불구하고 작업당 Opus 4.6보다 30~40% 더 많은 비용이 발생했습니다. 최근 Claude Sonnet 5에서도 이와 유사한 패턴이 나타났습니다.

투입 노력(effort)이 높을수록 비용이 증가하지만 성능은 오히려 저하될 수 있습니다.

사용자는 'low(낮음)', 'medium(중간)', 'high(높음)', 'xhigh(최고)', 'max(최대)'라는 5가지 노력 설정을 통해 성능과 토큰 사용량 간의 균형을 맞출 수 있습니다. 앤스로픽은 Opus 5가 모든 노력 수준에서 이전 모델보다 더 나은 가성비를 제공한다고 말합니다.

프롬프트 가이드에서 앤스로픽은 'low' 및 'medium' 설정을 폭넓게 활용할 것을 권장합니다. 회사 측은 이 설정들이 극히 일부의 토큰 사용량과 지연 시간으로도 훌륭한 결과를 내며, 기존 Opus 모델의 동일한 설정을 능가한다고 설명합니다. 다만 코딩 및 에이전트 작업의 경우 여전히 'xhigh' 설정에서 시작할 것을 권장합니다.

실제로 Opus 5는 두 가지 벤치마크에서 두 번째로 높은 설정인 'xhigh'보다 가장 높은 'max' 설정에서 비용이 더 많이 들었음에도 불구하고 점수가 약간 더 낮게 나타났습니다. 이러한 성능 저하는 Frontier-Bench v0.1과 Artificial Analysis Coding Agent Index에서 확인되었습니다.

에이전트 코딩 부문에서 Opus 5는 Fable 5를 능가합니다.

앤스로픽 자체 벤치마크에 따르면, Opus 5는 여러 평가에서 신기록을 세웠습니다. Frontier-Bench v0.1 테스트에서 이 모델은 에이전트 터미널 코딩(Agentic terminal coding) 부문에서 43.3%의 기록을 세워 Fable 5(33.7%), GPT-5.6 Sol(34.4%), 이전 모델인 Opus 4.8(21.1%)을 큰 차이로 제쳤습니다.

지식 작업 벤치마크인 GDPval-AA v2에서도 Opus 5는 1,861의 Elo 점수로 Fable 5(1,747)와 GPT-5.6 Sol(1,736)을 앞섰습니다.

물론 Opus 5가 모든 영역에서 1위를 차지한 것은 아닙니다. DeepSWE v1.1을 통한 에이전트 코딩 평가에서는 GPT-5.6 Sol이 72.7%로 1위를 차지했고, Fable 5(69.7%)와 Opus 5(68.8%)가 그 뒤를 따랐습니다. 의료 작업 및 법률 벤치마크에서는 각각 Fable 5와 Mythos 5가 Opus 5보다 더 나은 성능을 보였습니다.

특히 ARC-AGI-3 결과는 앤스로픽 벤치마크에서 가장 큰 서프라이즈이자 이례적인 기록입니다. 이 벤치마크는 암기된 패턴 없이 새로운 문제 해결 능력을 측정하는데, Opus 5는 30.2%의 점수를 기록했습니다. 이전 모델인 Opus 4.8은 1.5%에 그쳤고 GPT-5.6 Sol은 7.8%를 기록했기 때문에, 2위 모델과의 격차는 거의 4배에 달합니다. 이 벤치마크에 대한 Fable 5의 결과는 존재하지 않으며, 테스트에서 이렇게 큰 격차를 보인 것이 실제 사용 환경에서도 나타날지는 불분명합니다.

한편 Opus 5는 사이버 보안 작업에서는 Mythos 5에 뒤처지는 모습을 보였습니다.

원문 보기
원문 보기 (영어)
Anthropic claims its new Claude Opus 5 delivers near-Fable 5 performance at half the token price Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 24, 2026 Anthropic Key Points Anthropic is positioning Claude Opus 5 as a cheaper alternative to Fable 5 that beats it in several benchmarks. Opus 5 leads tests in agentic coding and knowledge work. On ARC-AGI-3, which measures novel problem-solving, the model scores 30.2 percent, nearly four times higher than GPT-5.6 Sol. Anthropic says Opus 5 can check and improve its own work through iteration, and build its own tools through code when it needs them. Ask about this article… Search Anthropic's new flagship model Claude Opus 5 posts top scores in coding and knowledge work while approaching Claude Fable 5's performance at half its token rates. Anthropic is responding to pricing pressure from GPT-5.6 Sol and Chinese competitors . Its new Opus 5 model is designed to close the price-performance gap with the much pricier Fable 5 . Opus 5 becomes the default model on Claude Max and the most capable model available on Claude Pro. The 1 million-token context window and token rates remain unchanged . Like its predecessor Opus 4.8, Opus 5 costs $5 per million input tokens and $25 per million output tokens. A new Fast Mode increases speed by 2.5x but doubles the price. Ad Model Input Tokens Cache Writes (5 min.) Cache Writes (1 hr) Cache Hits & Updates Output tokens Claude Fable 5 $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok Claude Mythos 5 (limited availability) $10 / MTok $12.50 / MTok $20 / MTok $1 / MTok $50 / MTok Claude Opus 5 $5 / MTok $6.25 / MTok $10 / MTok $0.50 / MTok $25 / MTok But token rates don't tell the full story without factoring in token efficiency. Opus 4.7 ended up costing 30 to 40 percent more per task than Opus 4.6 , even though both models had the same base rates. A similar pattern showed up recently with Claude Sonnet 5 . Ad DEC_D_Incontent-1 Higher effort can cost more and perform worse Users can trade off performance against token use through five effort settings called low, medium, high, xhigh, and max. Anthropic says Opus 5 offers better value than its predecessor at every effort level. In its prompting guide , Anthropic recommends making broad use of the "low" and "medium" settings. The company says they deliver good results with a fraction of the token use and latency while beating the same settings on earlier Opus models. For coding and agentic tasks, Anthropic still recommends starting with "xhigh." Ad Opus 5 scores slightly worse at the max effort setting than at the second-highest setting on two benchmarks, despite costing more. The drop appears on Frontier-Bench v0.1 and the Artificial Analysis Coding Agent Index. Opus 5 beats Fable 5 at agentic coding According to Anthropic's own benchmarks, Opus 5 sets records across several evaluations. On Frontier-Bench v0.1, the model hits 43.3 percent on agentic terminal coding, beating Fable 5 (33.7 percent), GPT-5.6 Sol (34.4 percent), and its predecessor Opus 4.8 (21.1 percent) by wide margins. On the knowledge work benchmark GDPval-AA v2, Opus 5 leads with an Elo score of 1,861, ahead of Fable 5 (1,747) and GPT-5.6 Sol (1,736). Ad DEC_D_Incontent-2 Opus 5 doesn't win everywhere. On agentic coding via DeepSWE v1.1, GPT-5.6 Sol leads with 72.7 percent, followed by Fable 5 (69.7 percent) and Opus 5 (68.8 percent). On health tasks and legal benchmarks, Fable 5 and Mythos 5 outperform Opus 5, respectively. Ad The ARC-AGI-3 result is likely the biggest surprise and outlier in Anthropic's benchmarks. The benchmark measures novel problem-solving without memorized patterns , and Opus 5 scores 30.2 percent. Opus 4.8 managed just 1.5 percent, GPT-5.6 Sol hit 7.8 percent. That's nearly a 4x gap over the next-best model. There's no Fable 5 result for this benchmark, and it's unclear whether such a large lead on a test will show up in actual use. Opus 5 falls behind Mythos 5 on cybersecurity tasks. Anthropic says it deliberately didn't train Opus 5 on cyber tasks, as was also the case with its predecessor. The model comes close to Mythos 5 at finding vulnerabilities but performs much worse when asked to exploit them. Anthropic also says Opus 5 has improved at generating visual outputs and analyzing visual content such as charts and diagrams. Opus 5 builds its own tools when it needs them Anthropic describes Opus 5 as much better at checking its own work and improving it through iteration. In a Frontier-Bench task, Opus 5 received a drawing of a machine part and had to create a 3D model in FreeCAD. The catch was that the model intentionally had no way to view the drawing directly. Opus 5 responded by writing its own computer vision pipeline to extract the geometry from raw pixels, then reconstructed the complete machine part. No other model solved this task after five attempts, Anthropic says. Opus 5 also worked on a real bug in a popular open-source package manager. According to Anthropic, it found the root cause and fixed an edge case that the community patch had missed. A competing model fixed only the surface symptom before reporting the bug as resolved. An engineer at a trading firm reportedly used Opus 5 to build a market data feed for a new exchange in one session. Previous models couldn't complete the task, even with detailed plans. Opus 5's cyber filters trigger less often than Fable 5's Opus 5's safety setup allows source code vulnerability research but blocks binary-based vulnerability scanning, penetration testing, and exploit generation, according to Anthropic. The cyber classifiers trigger about 85 percent less often than on Fable 5. Fable 5's frequent interventions drew heavy criticism . Blocked requests in Claude.ai, Claude Code, and Claude Cowork default to Opus 4.8 as a fallback, the same approach used with Fable 5. Anthropic calls Opus 5 the most capable generally available model for scientific research. It shows gains over Opus 4.8 across all life sciences evaluations, with standout improvements in organic chemistry (plus 10.2 percentage points on deriving molecular structures from spectroscopy data) and protein-related tasks (plus 7.7 percentage points). Alongside Opus 5, Anthropic is releasing two beta features. Mid-Conversation Tool Changes on the Claude Platform let developers swap available tools during a conversation without invalidating the prompt cache. Automatic Fallbacks on the API route blocked requests to a different model automatically. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Anthropic