메뉴
HN
Hacker News 4일 전

클로드 오푸스 5 (Claude Opus 5) 발표

IMP
9/10
핵심 요약

앤스로픽(Anthropic)이 최신 AI 모델인 클로드 오푸스 5(Claude Opus 5)를 공개했습니다. 이 모델은 이전 세대인 오푸스 4.8 대비 동일한 비용으로 압도적인 성능 향상을 보여주며, 코딩 및 지식 작업 벤치마크에서 새로운 SOTA(State-of-the-Art)를 달성했습니다. 특히 모델의 노력도(effort) 설정을 통해 토큰 비용과 지능도를 최적화할 수 있어 실무자의 일일 업무 효율성을 극대화하는 데 중요한 의미가 있습니다.

번역된 본문

제품 발표: 클로드 오푸스 5 (Claude Opus 5) 소개 (2026년 7월 24일)

오늘 클로드 오푸스 5(Claude Opus 5)를 출시합니다. 이 모델은 사려 깊고 주도적인 모델로, 클로드 페이블 5(Claude Fable 5)에 필적하는 최고 수준의 지능을 절반의 가격에 제공합니다. Frontier-Bench 및 GDPval-AA와 같은 코딩 및 지식 작업 평가에서 오푸스 5는 새로운 최고 수준(State-of-the-Art)의 모델입니다. 단, 사이버 보안 작업에서는 여전히 Mythos 5 뒤처집니다. 오푸스 5는 매일 사용할 수 있도록 설계되었으며, 다른 모델들보다 더 효율적으로 작동합니다. 이 모델은 Claude Max의 새로운 기본 모델이자, Claude Pro에서 가장 강력한 모델입니다.

성능 및 비용 효율성 클로드 오푸스 5는 이전 세대인 오푸스 4.8과 동일한 비용으로 대폭 향상된 성능을 제공합니다. 이 섹션의 차트는 모델의 노력도(effort) 설정에 따라 성능이 어떻게 변하는지 보여주며, 고객은 이를 활용해 지능을 극대화하거나 토큰을 절약해 더 빠르고 저렴한 결과를 얻을 수 있습니다.

오푸스 5는 가치 있는 소프트웨어 엔지니어링 작업에서 탁월합니다. 예를 들어, Frontier-Bench v0.1에서 오푸스 5는 다른 모든 모델을 능가하며, 작업당 더 낮은 비용으로 오푸스 4.8의 성능을 두 배 이상 달성했습니다. CursorBench 3.2에서 최대 노력(max effort)으로 설정 시, 이 모델은 페이블 5의 최고 점수와 0.5% 이내의 성능을 보여주지만 작업당 비용은 절반에 불과합니다. 또한 high, xhigh, max 노력 설정에서 주어진 비용 대비 다른 모든 모델보다 더 높은 성능을 달성합니다.

지식 작업 및 문제 해결 작업에서도 유사한 결과를 확인했습니다. 예를 들어:

  • 모델이 새로운 문제를 해결해야 하는 평가인 ARC-AGI 3에서 오푸스 5의 점수는 차상위 모델보다 3배 높습니다.
  • 모델이 비즈니스 작업을 처음부터 끝까지 완료할 수 있는지 측정하는 Zapier AutomationBench에서 오푸스 5의 통과율은 작업당 동일한 비용 대비 차상위 모델의 약 1.5배입니다. 최저 노력 설정에서도 오푸스 5는 다른 모든 모델보다 더 많은 작업을 통과합니다.
  • 컴퓨터 사용 벤치마크인 OSWorld 2.0에서 오푸스 5는 주어진 모든 비용 조건에서 다른 모든 모델을 능가하며, 단 3분의 1 약간 넘는 비용으로 페이블 5의 최고 결과를 뛰어넘습니다.

또한 여러 관련 평가에서 우리의 가장 좋고 비용 효율적인 모델입니다: ARC-AGI 3, GDPval-AA v2, OSWorld 2.0, HLE, AutomationBench, DeepSearchQA.

오푸스 5는 과학 연구 분야에서도 오푸스 4.8 대비 의미 있는 발전을 이루었습니다. 구조생물학, 유기화학, 생물정보학 등을 다루는 모든 생명 과학 평가에서 오푸스 4.8보다 더 나은 성능을 보여줍니다. 유기화학 작업(내부 벤치마크에서 분광 데이터로부터 분자 구조를 추론하여 오푸스 4.8보다 10.2% 높은 점수 기록)과 단백질 관련 작업(단백질 서열 변이가 기능에 미치는 영향 예측에서 7.7% 높은 점수 기록)에서 향상이 두드러집니다. 마지막으로 오푸스 5는 훨씬 더 강력한 시각적 출력물을 생성할 수 있습니다: 풍동(Wind tunnel), 셀 아티팩트(Cell artifact).

클로드 오푸스 5와 함께 일하기 클로드 오푸스 5는 자신의 작업을 검증하고 성공할 때까지 꼼꼼하게 반복하는 능력이 훨씬 강화되었습니다. 평가 및 얼리 액세스 테스트에서 우리와 사용자는 오푸스 5의 주도성과 철저함에 대한 많은 사례를 확인했습니다:

  • 한 Frontier-Bench 작업에서 오푸스 5에게 기계 부품 도면이 주어졌고, 이를 3D FreeCAD 모델로 재구성하는 코드를 작성하도록 요청받았습니다. 그러나 이 작업에서 모델은 도면을 직접 볼 수 없도록 의도적으로 제한되었습니다. 이에 오푸스 5는 원시 픽셀에서 기하학적 형상을 추출하기 위해 자체 컴퓨터 비전 파이프라인을 작성한 다음, 전체 기계 부품을 재구성했습니다. 이 과정에서 반복적으로 성공했으며, 동일한 설정을 사용한 경쟁 모델은 5번의 시도 끝에도 이를 해결하지 못했습니다.
  • 인기 있는 오픈 소스 패키지 관리자의 실제 버그가 주어졌을 때, 오푸스 5는 근본 원인을 찾아내고 커뮤니티 패치가 놓친 엣지 케이스를 수정했습니다. 반면 경쟁 모델은 표면적인 증상만 수정하고 근본 원인을 파악하지 못한 채 버그가 해결되었다고 보고했습니다.
  • 한 트레이딩 회사의 엔지니어는 오푸스 5를 사용해 단 한 번의 세션으로 새로운 거래소를 위한 시장 데이터 피드를 구축했습니다. 이전 모델들은 이 작업을 완수할 수 없었습니다.
원문 보기
원문 보기 (영어)
Product Announcements Introducing Claude Opus 5 Jul 24, 2026 Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA , Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro. Performance and cost-effectiveness Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results. Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2 , at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort. Frontier-Bench v0.1 CursorBench AA Coding Agent Index We see similar results on knowledge work and problem-solving tasks. For example: On ARC-AGI 3 , an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model. On Zapier AutomationBench , which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model. On OSWorld 2.0 , a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost. It’s also our best and most cost-efficient model on several related evaluations: ARC-AGI 3 GDPval-AA v2 OSWorld 2.0 HLE AutomationBench DeepSearchQA Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher). Finally, Opus 5 is capable of producing much stronger visual outputs: Wind tunnel Cell artifact Working with Claude Opus 5 Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness: On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts. Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved. An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly. Below are further reports from our early-access customers on their experience of working with Opus 5: On FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost. Within Devin, it also shows particular strength on difficult debugging and root-cause analysis tasks. Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. On CursorBench it’s just under Fable 5 and has many of the same behaviors. We are excited to see how developers use it in Cursor. Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn’t pass; Opus 5 hit 100%. On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run. It reaches for the right statistical tests to rule out confounders, cross-checks its own results by independent methods, and stays on track through long multi-step analyses. Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn’t just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build. Claude Opus 5 is the biggest leap in the Opus family since 4.5. On the same full-stack app builds, the front end shows it first: the best animations, games, and 3D work we have seen from an Opus model. We’re loving Claude Opus 5. For the kind of open-ended analytical work our agent handles, it’s a strict upgrade over Opus 4.8, and the gains are biggest exactly where it matters: the harder, vaguer tasks. Responses are clearer and more concise, and we see improved efficiency at higher effort levels too. Claude Opus 5 is a striking improvement over Opus 4.8 for the financial research workflows our analysts run every day. It stands out on numerical reasoning, table work, and sharper critical thinking where precision matters. Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily. Claude Opus 5 is a clear generational step up from Opus 4.8. Over one weekend I gave it a chief-of-staff role over my dev environments: it built its own monitor, drove each box, and pulled me in only for the judgment calls. Claude Opus 5 made large scale changes across our Fundamental Research Assistant codebase, adapting to feedback throughout an agentic workflow and explaining its reasoning more clearly than any model we’ve used. It handled work we would normally have broken into much smaller pieces. On some of our hardest financial-modeling tasks, Claude Opus 5 is a clear step up from Opus 4.8 in both accuracy and efficiency. Its performance floor is materially higher, especially on deep finance domain logic. Across effort levels it averaged 9 percentage points higher accuracy with a third fewer turns and tool calls and 60% less time. Claude Opus 5 checks its own work the way a real frontend developer would. On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off-screen checkout button, and fixed both before handing the work back. Claude Opus 5 is a clear step up in performance on legal agent work compared to prior Opus models, and
관련 소식