메뉴
HN
Hacker News • 3일 전

앤스로픽, 클로드 오퍼스 5.5 발표

IMP
9/10
핵심 요약

앤스로픽이 새로운 Claude 5.5 시리즈의 첫 모델인 클로드 오퍼스 5.5를 공개했습니다. 대부분의 작업에서 Claude Fable 5.1 수준의 성능을 보이며 운영 비용은 Opus 5보다 40% 저렴하고 출력 속도도 30% 이상 빠릅니다. 안전성 면에서도 자동 행동 감사에서 역대 최고 점수를 기록했으며, 생명과학·사이버보안 등 고위험 분야는 검증된 조직에 한해 접근이 허용됩니다.

번역된 본문

새로운 Claude 5.5 시리즈의 첫 모델인 Claude Opus 5.5를 소개합니다. 이 모델은 대부분의 작업에서 Claude Fable 5.1과 동등한 수준의 성능을 발휘하며, 운영 비용은 Opus 5보다 40% 저렴합니다.

Claude Opus 5.5는 우리가 '프론티어 개발 속도 조절(frontier pacing)'을 촉구한 이후 첫 출시입니다. 출시 전에 Frontier Design과 METR을 포함한 외부 평가기관의 테스트를 거쳤습니다. 우리가 수행하는 가장 포괄적인 정렬(alignment) 테스트인 자동 행동 감사에서 Opus 5.5는 지금까지 테스트한 모델 중 가장 뛰어난 성적을 보였습니다. 또한 가장 강력한 모델을 위해 개발한 안전장치를 함께 갖추고 있습니다.

Opus 5.5에서 기대할 수 있는 개선 사항은 다음과 같습니다.

성능. Opus 5.5는 Opus 5 대비 큰 도약입니다. 새로운 선도 모델로서, 초기 테스터들은 가장 복잡한 작업에서 성능이 크게 향상된 것을 확인했습니다. 한 테스터는 68만 줄짜리 코드 마이그레이션을 하루도 걸리지 않아 완료했는데, 이는 엔지니어링 팀이라면 몇 주가 걸릴 작업이었습니다. 소프트웨어의 비효율을 찾아 수정하는 데에도 뛰어납니다. 웹 앱의 모든 페이지 로딩 시간을 단축하라고 요청했을 때 Opus 5.5는 40회 중 39회 성공했으며, Opus 5는 더 작은 개선을 보였을 뿐 아니라 앱의 동작도 바꿔버렸습니다. 다른 테스터는 여러 Claude 모델에게 단일 프롬프트로 게임을 만들게 했는데, Opus 5.5가 그래픽과 완성도 면에서 가장 높은 점수를 받았습니다.

안전성. Opus 5.5는 수천 개의 시뮬레이션 시나리오에서 Claude를 테스트하는 정렬 테스트 모음인 자동 행동 감사에서 역대 최고 점수를 달성했습니다. 최근 모델들에 비해 되돌리기 어려운 행동을 하거나 주어진 경계를 벗어나 행동할 가능성이 훨씬 낮고, 프롬프트 인젝션(prompt injection)에 대한 저항성도 Opus 5보다 높습니다. 또한 정렬 테스트를 더 긴 작업, 불가능한 작업, 실제 사고를 모델링한 시나리오로 확장했지만 여전히 한계는 존재합니다. 평가에 대한 전체 내용은 Opus 5.5 시스템 카드에서 확인할 수 있습니다.

Opus 5.5는 생물학과 사이버보안 분야에서 Claude Mythos 5.1과 동등한 수준이기 때문에, Claude Fable 5.1과 유사한 안전장치를 갖추고 배포합니다. 검증된 조직은 오늘부터 생명과학 검증 프로그램(Life Sciences Verification Program)에 신청하여 생물학 연구에 Opus 5.5를 사용할 수 있습니다. 향후 몇 주 내로 사이버 검증 프로그램(Cyber Verification Program) 접근 범위도 확대하여, 검증된 사이버보안 실무자들이 업무에 Opus 5.5를 사용할 수 있게 됩니다.

비용과 속도. Opus 5.5는 Opus 5보다 적은 컴퓨팅으로 서비스할 수 있으며, 가격도 이를 반영합니다. 테스트 결과 기본 설정에서 일반적인 워크로드 기준 Opus 5보다 40% 저렴합니다. 입력·출력 토큰 가격은 백만 토큰당 각각 4달러와 20달러로, Opus 5보다 20% 저렴합니다. 캐시 읽기(에이전트·코딩 작업 비용의 대부분을 차지)는 백만 토큰당 0.20달러로, Opus 5보다 60% 저렴합니다. 출력 생성 속도도 Opus 5보다 30% 이상 빠릅니다. 가격 인하 외에도 Pro, Max, Team 플랜의 5시간 사용 한도를 늘립니다. 또한 구독 사용자에게 요율 제한 리셋(rate limit reset)을 제공하며, 이를 저장해 원하는 때에 사용할 수 있습니다.

소통. Opus 5.5는 이전 모델보다 더 자연스럽게 소통합니다. 초기 테스터들은 글이 더 명확하고 따라가기 쉬워졌다고 평가했는데, 이는 Opus 5에 대한 일반적인 피드백을 해소한 것입니다. 가장 중요한 정보를 앞에 배치하고, 그 스타일 덕분에 긴 세션에서 더 나은 작업 파트너가 됩니다. 한 초기 테스터는 말하길, "내가 쓰는 방식으로 글을 씁니다"라고 했습니다. 우리 자신의 사용 경험에서도 Opus 5.5의 작업을 더 따라가고 검증하기 쉬워졌는데, 이는 안전성 측면에서도 실용적 측면에서도 이점입니다.

Claude Sonnet 5.5와 Claude Haiku 5.5도 향후 몇 주 내에 출시되며, 성능·효율성·안전성에서 동일한 개선 사항을 대부분 포함합니다.

성능과 비용 효율성 벤치마크에서 Claude Opus 5.5는 에이전트 코딩, 컴퓨터 사용, 지식 작업 분야를 선도합니다. 다만 이 수준의 성능에서는 벤치마크 격차가 실제 사용 환경의 차이를 보여주는 신뢰할 만한 지표가 되지 못한다는 점을 확인했습니다. 우리의 실제 사용에서 Opus 5.5와 Claude Fable 5.1의 격차는... (원문 누락)

원문 보기
원문 보기 (영어)
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. Claude Opus 5.5 is our first release since we called for pacing the frontier . It was tested before release by external evaluators, including Frontier Design and METR . On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models. Here are some of the improvements you can expect from Opus 5.5: Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish. Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card . Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program , and verified cybersecurity practitioners will be able to use Opus 5.5 for their work. Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5. In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, and Team plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose. Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one. Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety. Performance and cost-effectiveness On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest. Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs. Pricing Prices per 1M tokens Claude Opus 5.5 Claude Opus 5 Cache reads $0.20 $0.50 Input tokens $4 $5 Output tokens $20 $25 Cache writes $5 $6.25 Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens. Coding Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less. Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost. Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench Our early testers reported similar efficiency and intelligence gains: GitHub Clio Lovable Quantium Spotify Optiver Column Kiro GitHub Clio Lovable Quantium Spotify Optiver Column Kiro Quote “Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.” Company GitHub Author Mario Rodriguez, Chief Product Officer