메뉴
HN
Hacker News 28일 전

클로드 소넷 5 출시: 최고 수준의 자율 에이전트 모델

IMP
9/10
핵심 요약

앤스로픽(Anthropic)이 대규모 모델에 맞먹는 추론 및 도구 활용 능력을 갖춘 '클로드 소넷 5'를 공개했습니다. 이 모델은 기존 소넷 모델 대비 자율성이 크게 향상되었으며, 복잡한 코딩과 소프트웨어 엔지니어링 작업을 독립적으로 수행하면서도 합리적인 가격을 유지하여 실무 개발자들에게 효율적인 옵션을 제공합니다.

번역된 본문

제품 소개: 클로드 소넷 5 (Claude Sonnet 5) 2026년 6월 30일

클로드 소넷 5는 역대 가장 강력한 에이전트(Agent) 중심 소넷 모델로 설계되었습니다. 스스로 계획을 수립하고, 브라우저 및 터미널과 같은 도구를 사용하며, 불과 몇 달 전만 해도 더 크고 비싼 대형 모델이 필요했던 수준으로 자율적으로 실행할 수 있습니다. 많은 개발자들에게 에이전트 AI 시대는 소넷(Sonnet)급 모델과 함께 시작되었습니다. 클로드 소넷 3.5, 3.6, 3.7은 코딩과 도구 사용에서 놀라운 기술력을 보여준 첫 번째 모델들이었습니다. 하지만 최근 에이전트 역량의 가장 큰 발전은 오푸스(Opus)급 모델에서 이루어졌습니다. 소넷 5는 이러한 격차를 좁혔습니다. 오푸스 4.8에 근접하는 성능을 발휘하지만 가격은 더 저렴합니다. 추론, 도구 사용, 코딩 및 지식 작업과 같은 에이전트 성능의 핵심 측면에서 기존 소넷 4.6 대비 실질적인 향상을 이루었습니다.

당사의 안전성 평가 결과, 소넷 5는 소넷 4.6보다 전반적으로 바람직하지 않은 행동 발생률이 낮으며, 에이전트 맥락에서 사용할 때 일반적으로 더 안전한 것으로 나타났습니다. 또한, 평가 결과에 따르면 현재의 오푸스 모델들에 비해 사이버 보안(Cybersecurity) 공격 작업을 수행할 수 있는 능력은 훨씬 낮습니다.

오늘부터 클로드 소넷 5는 모든 요금제에서 사용할 수 있습니다. 프리(Free) 및 프로(Pro) 플랜의 기본 모델이며, 맥스(Max), 팀(Team), 엔터프라이즈(Enterprise) 사용자에게도 제공됩니다. 또한 클로드 코드(Claude Code)와 클로드 플랫폼(Claude Platform)에서도 이용 가능합니다. 이 플랫폼에서는 2026년 8월 31일까지 백만 입력 토큰당 2달러, 백만 출력 토큰당 10달러의 도입 기간 할인 가격이 적용되며, 이후에는 백만 입력 토큰당 3달러, 백만 출력 토큰당 15달러로 책정됩니다. 개발자들은 클로드 API를 통해 claude-sonnet-5를 사용할 수 있습니다.

클로드 소넷 5와 함께 작업하기 아래 차트는 다양한 노력(Effort) 수준에서 소넷 5의 성능을 소넷 4.6 및 오푸스 4.8과 비교한 것입니다. 여기에는 에이전트 검색 평가인 BrowseComp와 컴퓨터 사용 평가인 OSWorld-Verified가 포함됩니다. 소넷 5(주황색 선)는 소넷 4.6(회색 선) 대비 확실한 성능 향상을 보여줍니다. 이러한 작업에서 더 높은 정확도가 필요한 경우 여전히 오푸스 4.8(노란색 선)이 최적의 선택이지만, 소넷 5는 개발자들에게 기존보다 훨씬 높은 품질을 제공하면서도 가격은 낮은 옵션을 제공합니다. 소넷 5와 오푸스 4.8 사이에서 사용자는 노력 수준을 조절하여 비용과 성능의 최적의 균형을 찾을 수 있습니다.

[에이전트 검색 / 에이전트 컴퓨터 사용 차트 이미지 설명]

早期 액세스 파트너들의 피드백은 일관적입니다. 소넷 5는 이전 모델들보다 훨씬 뛰어난 자율성을 보여줍니다. 테스터들은 기존 소넷 모델이 진행을 멈추던 복잡한 작업을 완성하는 방식, 명시적인 요청 없이도 자체 출력을 확인하는 방식, 그리고 이 모든 에이전트 작업을 매력적인 가격에 수행하는 방식에 대해 다음과 같이 설명했습니다:

  • "클로드 소넷 5는 여러 단계로 이루어진 소프트웨어 엔지니어링 작업에서 우리 에이전트에게 강력한 실행 계층을 제공합니다. 복잡하고 어려운 기술적 환경에서도 지속적인 코딩, 도구 사용 및 디버깅을 잘 처리하며, 끝까지 밀고 나가는 실행력과 기술적 기반이 중요한 워크플로우에서 특히 유용했습니다."

  • "우리는 클로드 소넷 5에게 두 가지 임무를 주었습니다. 세일즈포스(Salesforce) 계정 등급을 업데이트하고 엔터프라이즈 담당자들에게 출시 알림을 보내는 것이었는데, 이를 처음부터 끝까지 완수했습니다. 과거에는 중간에 멈추곤 했었죠. 일상적인 자동화를 위해서는 망설일 필요 없는 선택입니다. 클로드 소넷 5는 적은 리소스로 더 많은 성과를 냅니다. 동일한 출력 품질을 유지하면서도 거기에 도달하는 단계는 더 적습니다. 또한 안전하지 않은 요청도 깔끔하고 일관되게 거부합니다. 러버블(Lovable)에서 우리는 수백만 명의 빌더들에게 강력한 도구를 쥐여주고 있습니다. 구축하는 방법을 아는 모델만큼이나, 언제 거절해야 하는지 아는 모델이 중요합니다."

  • "우리는 클로드 소넷 5를 우리가 가진 가장 도전적인 수십 개의 실제 풀 리퀘스트(Pull Request)에 투입해 보았습니다. 그结果是 각 작업을 테스트 및 검증된 결과물로 끝까지 완성하여, 엔지니어들이 판단과 의사결정, 최종 승인에 집중할 수 있게 해주었습니다."

  • "클로드 소넷 5에게 버그를 조사해달라고 요청했습니다. 별도의 지시 없이도 테스트 코드를 작성하고, 수정 사항을 구현한 뒤, 해당 변경 사항이 없으면 버그가 다시 나타나는지 확인하기 위해 이를 임시 저장(stash)했습니다. 이 모든 과정이 단 한 번의 실행으로 이루어졌습니다. 클로드 소넷 5를 사용하면 에이전트가 계획을 유지하고 당사의 규칙을 따르며, 복잡한 다중 단계의 변경 사항을 클린하게 배포(Ship)합니다. 게다가 매우 효율적입니다."

원문 보기
원문 보기 (영어)
Product Introducing Claude Sonnet 5 Jun 30, 2026 Claude Sonnet 5 is built to be the most agentic Sonnet model yet. It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models. For many developers, the agentic AI era began with Sonnet-class models: Claude Sonnet 3.5, 3.6, and 3.7 were the first models that showed impressive skills in coding and tool use. More recently, though, the clearest gains in agentic capabilities have been in our Opus-class models. Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. It’s a substantial improvement over its predecessor, Sonnet 4.6, on important aspects of agentic performance like reasoning, tool use, coding, and knowledge work: Our safety assessments found that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6, and is generally safer to use in agentic contexts. Evaluations also show that it has a much lower ability to perform cybersecurity tasks than our current Opus models. From today, Claude Sonnet 5 is available across all plans: it is the default model for Free and Pro plans, and is available to Max, Team, and Enterprise users. It’s also available in Claude Code and on the Claude Platform, where it launches with introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026, after which it will be priced at $3 per million input tokens and $15 per million output tokens. Developers can use claude-sonnet-5 via the Claude API . Working with Claude Sonnet 5 The charts below compare the performance of Sonnet 5 with Sonnet 4.6 and Opus 4.8 at different effort levels on the agentic search evaluation BrowseComp and the computer use evaluation OSWorld-Verified . Sonnet 5 (orange line) is a strict improvement over Sonnet 4.6 (gray line). Opus 4.8 (yellow line) is still the model of choice for higher accuracy on these tasks, but Sonnet 5 provides developers with lower-priced options that are of much higher quality than what was previously available. Between Sonnet 5 and Opus 4.8, users can adjust the effort level to find the right balance of cost and performance. Agentic search Agentic computer use Feedback from our early access partners has been consistent: Sonnet 5 is much more agentic than its predecessors. Testers described how it finishes complex tasks where previous Sonnet models would stop short, how it checks its own output without explicitly being asked, and how it does all this agentic work at an attractive price point: Claude Sonnet 5 gives our agents a strong execution layer for multi-step software engineering work. It handles sustained coding, tool use, and debugging well across messy technical contexts, and has been especially useful for workflows where follow-through and technical grounding matter. We handed Claude Sonnet 5 a two-part job—update Salesforce account tiers, send a launch announcement to enterprise contacts—and it finished end to end. That used to stall halfway. For day-to-day automation, it’s a no-brainer Claude Sonnet 5 gets more done with less. Same output quality, fewer steps to get there. It refuses unsafe requests cleanly and consistently, too. At Lovable, we’re putting powerful tools in the hands of millions of builders. A model that knows when to say no is just as important as one that knows how to build. We ran Claude Sonnet 5 against dozens of our most challenging real pull requests, and it carried each one through to a tested, verified result on its own — freeing our engineers to focus on the judgment, the decision, and the final sign-off. I asked Claude Sonnet 5 to investigate a bug. Unprompted, it wrote a reproducing test, implemented the fix, then stashed it to confirm the bug came back without the change. All in a single pass. With Claude Sonnet 5, agents stay on plan, follow our conventions, and ship clean multi-step changes, all at an efficient cost. Claude Sonnet 5 is at its best on brownfield code—race conditions, hidden tests, the parts nobody wants to touch. It traces a failure to its actual root cause and ships a durable fix instead of patching the symptom. Claude Sonnet 5 sits on the Pareto frontier for Eve’s plaintiff-law tasks. We see the clearest gains in legal research and analysis, at a price-to-performance ratio that made the choice to migrate easy. ClickHouse agents explore live data and produce insights on the fly, so time-to-insight matters when testing new models. Claude Sonnet 5 reasons in tighter steps and gets our users to answers noticeably faster. That speed is a difference our customers feel. At Pace, our computer-use agents run insurance workflows—submission intake, FNOL, loss runs—on the systems our operations teams already use. Claude Sonnet 5 consistently takes the right action and does it quickly, which is what real insurance work demands. 01 / 10 Safety evaluations Our pre-deployment safety evaluations found that Sonnet 5 was overall an improvement on Sonnet 4.6. On agentic safety, the model is better at refusing malicious requests and resisting hijack attempts in prompt injection attacks. The model shows lower rates of hallucination and sycophancy than Sonnet 4.6. On our automated behavioral audit, which tests a wide range of misaligned behaviors such as cooperation with misuse and deception, Sonnet 5 scored lower (that is, safer) overall. However, it did show somewhat higher rates of misaligned behavior on this assessment compared to the more capable Opus 4.8 and Claude Mythos Preview. We did not deliberately train Sonnet 5 on cybersecurity tasks. It can perform some routine, non-harmful cyber tasks, but on evaluations testing potentially dangerous cyber skills, such as developing software exploits, it shows substantially poorer performance than models such as Opus 4.8 and Mythos 5. Scores from one evaluation, which tested models’ ability to develop exploits for vulnerabilities in the Firefox browser, are shown in the chart below. Sonnet 5 was never able to develop a full working exploit, but it does show a slightly higher rate of partial success than Sonnet 4.6. This latter change is likely due to improvements in general intelligence rather than specific training. Since Sonnet 5 is somewhat stronger than its predecessor on these tasks, we’ve launched it with cyber safeguards enabled by default. These safeguards —which detect and block dangerous cyber usage in real time—are the same as those present in Claude Opus 4.7 and 4.8 (because we judged that the overall level of cybersecurity risk from Sonnet 5 was low, the safeguards are less strict than those launched with Fable 5, which block a much wider range of cybersecurity tasks). 1 Our full assessment of Sonnet 5 across many safety and capability evaluations is reported in the Claude Sonnet 5 System Card . Availability and pricing Claude Sonnet 5 is available everywhere today at an introductory price of $2 per million input tokens and $10 per million output tokens through August 31, 2026. It then moves to standard pricing at $3 per million input tokens and $15 per million output tokens. 2 We’ve increased rate limits across Chat, Cowork, Claude Code, and the Claude Platform 3 to accommodate the higher token usage of higher effort levels; users can select whichever level makes sense for their particular project.