메뉴
BL
The Decoder 28일 전

앤스로픽, '클로드 소네트 5' 공개... 비싼 오퍼스 모델과의 격차를 줄이다

IMP
8/10
핵심 요약

앤스로픽이 자율적 에이전트 능력을 대폭 강화한 '클로드 소네트 5'를 출시했습니다. 이 모델은 이전 모델을 전 분야에서 압도할 뿐만 아니라, 실제 지식 작업(Task)에서는 기존 최상위 모델인 오퍼스 4.8을 근소하게 뛰어넘는 성능을 보여줍니다. 사이버 보안 위험을 낮추면서도 100만 토큰의 컨텍스트 윈도우를 지원하며, 8월 말까지는 할인된 도입가로 제공됩니다.

번역된 본문

앤스로픽, 비싼 오퍼스 모델과의 격차를 줄인 '클로드 소네트 5' 공개

앤스로픽은 클로드 소네트 5(Claude Sonnet 5)를 출시했습니다. 벤치마크 결과에 따르면 이 모델은 더 큰 규모의 오퍼스 4.8(Opus 4.8)에 근접했으며, 일부 분야에서는 이를 능가하기도 했습니다. 현재 도입 할인가로 모델을 사용할 수 있습니다.

앤스로픽은 이 모델을 '지금까지 가장 에이전트(Agent)화된 소네트'라고 부릅니다. 이 모델은 스스로 계획을 세우고, 브라우저나 터미널 같은 도구를 잡아(활용해) 작업을 수행할 수 있습니다. 불과 몇 달 전만 해도 더 크고 비싼 모델만이 소화할 수 있었던 수준으로 독립적으로 작동하며, 소네트 5는 이러한 격차를 줄이는 것을 목표로 합니다.

소네트 4.6 대비 확실한 도약을 보여주는 벤치마크

앤스로픽이 공개한 벤치마크에 따르면, 소네트 5는 모든 테스트 카테고리에서 이전 버전인 소네트 4.6을 뛰어넘었으며, 더 비싼 오퍼스 4.8의 성능에 근접했습니다. 에이전트 코딩(Agentic coding)을 측정하는 SWE-bench Pro에서 소네트 5는 63.2%를 기록하며 소네트 4.6(58.1%)을 앞섰습니다. 오퍼스 4.8은 69.2%를 기록했습니다. Terminal-Bench 2.1에서는 소네트 5가 80.4%를 기록한 반면, 소네트 4.6은 67.0%에 그쳤습니다. 다학제간 추론(Humanity's Last Exam)에서는 도구 사용 시 57.4%에 도달하며 57.9%를 기록한 오퍼스 4.8과 거의 비슷한 수준을 보여주었습니다. 컴퓨터 사용(OSWorld-Verified) 분야에서는 소네트 5가 81.2%를 기록해 이전 모델(78.5%)을 능가했습니다.

특히 실제 지식 작업(Real-world knowledge tasks)에서 AI를 테스트하는 GDPval-AA v2 벤치마크에서는 소네트 5가 1,618점을 기록하며 1,615점을 받은 더 큰 규모의 오퍼스 4.8을 실제로 이겼습니다. 앤스로픽은 얼리 액세스 파트너들의 피드백 역시 같은 결과를 보여주었다고 밝혔습니다. 소네트 5는 검색 작업을 처리하는 방식 등에서 이전 버전들보다 훨씬 더 능동적으로 에이전트처럼 행동합니다.

이번에는 사이버 보안이 우려사항이 아니다

최근 앤스로픽은 출시하지 못하는 모델들로 뉴스거리가 되었습니다. 미국 정부는 사이버 보안 우려를 이유로 앤스로픽의 가장 성능이 좋은 두 모델인 Mythos 5와 Fable 5의 출시를 막고 있습니다. 이러한 상황은 소네트 5 출시에도 영향을 미쳤고, 앤스로픽은 유사한 우려가 생기는 것을 방지하고자 명확히 선제 대응했습니다.

앤스로픽에 따르면 이 모델은 사이버 보안 작업으로 훈련되지 않았으며, 소프트웨어 취약점 공격 코드 작성과 같은 위험한 능력을 테스트할 때 오퍼스 4.8과 Mythos 5보다 훨씬 낮은 점수를 받습니다. 다만 이러한 작업에서 이전 모델보다는 약간 높은 점수를 기록했기 때문에, 앤스로픽은 기본적으로 사이버 보안 장치를 활성화했습니다. 이는 위험한 사이버 사용을 실시간으로 표시하고 차단하며, 클로드 오퍼스 4.7 및 4.8에 이미 적용된 보호 장치와 동등한 수준입니다. 사용자들이 불만을 제기했던 Fable 5의 제한적인 가드레일과 비교하면 다소 완화된 수준입니다. 앤스로픽은 소네트 5의 전반적인 사이버 보안 위험이 낮다고 평가하고 있습니다.

안전성 측면에서 앤스로픽에 따르면, 이 모델은 소네트 4.6에 비해 악의적인 요청을 거부하고 프롬프트 삽입 공격(Prompt Injection)을 방어하는 능력이 향상되었습니다. 또한 환각(Hallucination) 현상과 사용자가 하는 말이라면 무조건 동조하는 아첨(Sycophantic) 행동도 줄어들었습니다. 앤스로픽의 전체 안전성 평가는 '클로드 소네트 5 시스템 카드(System Card)'에 자세히 나와 있습니다.

2026년 8월까지 적용되는 도입 할인가

클로드 소네트 5는 현재 모든 요금제에서 사용 가능합니다. 이 모델은 무료(Free) 및 프로(Pro) 사용자의 새로운 기본 모델이며, 맥스(Max), 팀(Team), 엔터프라이즈(Enterprise) 구독자도 접근할 수 있습니다. 개발자들은 클로드 코드(Claude Code)와 클로드 플랫폼(Claude Platform)에 이 모델을 연동할 수 있습니다. API 측면에서는 'claude-sonnet-5'라는 이름으로 제공됩니다. 훈련 데이터 기준일은 2026년 1월이며, 100만 토큰의 컨텍스트 윈도우(Context Window)를 지원합니다.

2026년 8월 31일까지 앤스로픽은 백만 입력 토큰당 2달러, 백만 출력 토큰당 10달러의 요금을 부과합니다.

원문 보기
원문 보기 (영어)
Anthropic's new Claude Sonnet 5 closes the gap to the pricier Opus model series Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 30, 2026 Anthropic Key Points Anthropic released Claude Sonnet 5, which the company calls its most agentic Sonnet yet. It can build plans on its own and use tools like browsers and terminals. In benchmarks, Sonnet 5 beats its predecessor, Sonnet 4.6, across the board and closes in on the larger Opus 4.8. On real-world knowledge work tasks, it even edges past Opus 4.8. The model is available now on all Anthropic platforms at an introductory discount, with pricing rising to standard Sonnet rates after August 2026. Ask about this article… Search Anthropic released Claude Sonnet 5. In benchmarks, it closes in on the larger Opus 4.8 and even beats it in some areas. The model is available now at an introductory price. Anthropic calls it the most agentic Sonnet yet: it can build plans, grab tools like browsers and terminals, and work on its own at a level that just months ago only bigger, pricier models could pull off, according to the company. Sonnet 5 is meant to close that gap. Benchmarks show a clear jump over Sonnet 4.6 Anthropic's published benchmarks show Sonnet 5 beating its predecessor Sonnet 4.6 in every tested category while gaining ground on the pricier Opus 4.8. On agentic coding, Sonnet 5 hits 63.2 percent on SWE-bench Pro, up from 58.1 percent for Sonnet 4.6. Opus 4.8 sits at 69.2 percent. On Terminal-Bench 2.1, Sonnet 5 pulls 80.4 percent versus Sonnet 4.6's 67.0 percent. For multidisciplinary reasoning (Humanity's Last Exam), the model reaches 57.4 percent with tools, nearly matching Opus 4.8 at 57.9 percent. On computer use (OSWorld-Verified), Sonnet 5 posts 81.2 percent compared to 78.5 percent for its predecessor. Ad On the knowledge work benchmark GDPval-AA v2, which tests AI on real-world knowledge tasks , Sonnet 5 actually beats the larger Opus 4.8, scoring 1,618 to Opus's 1,615. Anthropic says feedback from early-access partners told the same story. Sonnet 5 acts far more agentically than previous versions, showing up in things like how it handles search tasks. Ad DEC_D_Incontent-1 Cybersecurity isn't a concern this time Lately, Anthropic has been making news for models it can't ship. The US government is blocking the company's two most capable models, Mythos 5 and Fable 5 , over cybersecurity concerns. That context hangs over the Sonnet 5 launch. Anthropic is clearly eager to get ahead of any similar worries. The model wasn't trained on cybersecurity tasks, the company says, and in tests for risky capabilities like writing software exploits, it scores far below both Opus 4.8 and Mythos 5. Sonnet 5 does score a bit higher than its predecessor on these tasks, though. So Anthropic has switched on cyber safeguards by default. They flag and block risky cyber usage in real time, on par with the protections already in place for Claude Opus 4.7 and 4.8. They're dialed back compared to Fable 5's guardrails, which users complained about almost immediately . Anthropic says it views the overall cybersecurity risk from Sonnet 5 as low. Ad On the safety front, the model does a better job turning down malicious requests and fending off prompt injection attacks than Sonnet 4.6, according to Anthropic. Hallucinations and sycophantic behavior , the tendency to just agree with whatever the user says, are down as well. Anthropic's full safety evaluation is in the Claude Sonnet 5 System Card . Introductory pricing runs through August 2026 Claude Sonnet 5 is live now on all plans. It's the new default for Free and Pro users, and Max, Team, and Enterprise subscribers can access it too. Developers can plug it into Claude Code and the Claude Platform. On the API side, it goes by "claude-sonnet-5" . The training cutoff is January 2026, with a one-million-token context window. Ad DEC_D_Incontent-2 Until August 31, 2026, Anthropic is charging $2 per million input tokens and $10 per million output tokens. After that , prices jump to $3 and $15, which is what previous Sonnet models cost. Ad Real-world costs might tell a different story: Because the model works more agentically, it's likely to chew through more tokens per task . So even at the same per-token rate, running Sonnet 5 could end up costing more than its predecessors. The same thing happened when Opus went from 4.6 to 4.7. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Anthropic