메뉴
BL
TechCrunch AI • 3일 전

안스로픽, 가격 인하된 오퍼스 5.5 출시…페이블급 성능

IMP
8/10
핵심 요약

안스로픽이 코딩·지식 작업 성능에서 새로운 최고 수준을 기록한 오퍼스 5.5를 출시했습니다. 더 큰 규모의 페이블 모델을 여러 벤치마크에서 능가하면서도 출력 토큰 가격이 mTok당 25달러에서 20달러로 인하되고 속도도 빨라졌습니다. 다리오 아모데이 CEO의 '프런티어 속도 조절' 선언 이후 첫 모델 출시로, 생물학·사이버보안 능력이 높아 페이블과 동일한 안전장치가 적용됩니다.

번역된 본문

안스로픽의 최신 모델 오퍼스 5.5가 화요일 출시되며, 코딩 및 지식 작업 성능에서 새로운 최고 수준(SOTA)을 기록했습니다. 특히 이번 출시 모델은 더 큰 규모의 페이블(Fable) 모델을 여러 벤치마크에서 앞질렀으며, 페이블이 완료하지 못한 다수의 비공식 과제에서도 성공했습니다. 발표문에서 안스로픽은 오퍼스 5.5를 "지금까지 테스트한 모델 중 가장 뛰어난 성능을 보이는 모델"이라고 밝혔습니다.

새 모델은 이전 모델보다 상당히 저렴합니다. 오퍼스 5.5의 출력 토큰 가격은 mTok당 20달러로, 이전 모델의 25달러보다 낮습니다. 다른 지표들도 비슷한 수준으로 인하됐습니다. 또한 연산 요구량이 전반적으로 줄어들어 실행 속도도 더 빨라졌습니다.

새 버전은 오퍼스의 소통 방식에도 중요한 변화를 가져왔습니다. 오퍼스 5.5는 전문 용어(jargon) 사용이 줄고, 중요한 정보를 메시지 앞부분에 배치하는 경향이 커졌습니다.

이번 출시는 7월 24일 오퍼스 5가 출시된 지 두 달 만에 이루어졌습니다. 발표에 따르면 소네트 5.5와 하이쿠 5.5도 "향후 몇 주 내" 비슷한 성능 개선과 함께 출시될 예정입니다.

안스로픽은 오퍼스 5.5가 생물학 및 사이버보안 능력에서 미토스(Mythos)와 비슷한 수준이어서, 회사의 페이블 모델과 동일한 안전장치 적용 대상이라고 밝혔습니다. 이러한 안전장치는 컴파일된 프로그램의 취약점 발견이나 실체를 알아볼 수 있는 생물학 무기 개발 등 특정 작업에 모델을 활용할 수 있는 범위를 제한합니다.

오퍼스 5.5는 다리오 아모데이 CEO가 '프런티어 속도 조절(pace the frontier)' 요구를 수용한 이후 첫 모델 출시입니다. 이는 AI 역량 개발을 고의로 늦춰 정렬(alignment) 연구 진도와 맞추겠다는 것입니다. 아모데이는 이달 초 게시물에서 "위험을 완전히 해결하려면 더 큰 신중함이 필요하다는 확신이 들었다"며 "위험 예방에 투자하는 것뿐 아니라 역량 발전 속도를 조절해 위험 예방이 따라잡을 시간을 확보해야 한다"고 썼습니다.

오퍼스 5.5의 안전성 훈련은 이전 모델과 전반적으로 비슷했으며, METR과 프론티어 디자인(Frontier Design) 같은 외부 기관의 정렬 테스트와 출시 전 평가를 거쳤습니다. 다만 안스로픽은 개선된 보안 및 모니터링 시스템을 포함해 더 진보된 훈련·평가 시스템이 향후 모델을 위해 이미 준비 중이라고 강조했습니다.

블로그 포스트는 "AI가 더 능력을 갖추면서, 사람들이 의존하는 시스템이 안전한지 보장하는 데 공공 정책이 더 큰 역할을 해야 한다. 이러한 역량은 구축에 시간이 걸리며, 우리는 이를 뒷받침할 인프라를 마련하기 시작했다"며 "이러한 노력에 대해 곧 더 자세히 공유할 예정"이라고 밝혔습니다.

원문 보기
원문 보기 (영어)
Anthropic's newest model, Opus 5.5, was released on Tuesday , setting a new state-of-the-art in coding and knowledge work performance. Notably, the release outpaces the larger Fable model in many benchmarks, and succeeded in a number of informal tasks that Fable failed to complete. In the announcement, Anthropic called Opus 5.5 "the strongest-performing model we've tested to date." The new model is also significantly cheaper than its predecessor. Output tokens will be charged at $20 per mTok for Opus 5.5, compared to $25 for the previous model. Other metrics have similar price drops. The model is also faster to run, reflecting an overall drop in the compute requirements. The new version also makes significant changes to how Opus communicates, with the Opus 5.5 less likely to use jargon and more likely to put important information at the start of its messages. The launch comes just two months after the release of Opus 5 on July 24th . According to the announcement, Sonnet 5.5 and Haiku 5.5 will be released "in the coming weeks," with similar performance improvements. Anthropic says that Opus 5.5 is comparable to Mythos in its biology and cybersecurity capabilities, so its release is subject to the same safeguards as the company's Fable model. Those safeguards limit how much the models can be used to discover exploits in compiled programs or developing recognizable biological weapons , among other tasks. Opus 5.5 is Anthropic's first model release since CEO Dario Amodei embraced calls to pace the frontier, deliberately slowing down progress on AI capabilities to match the rate of progress on alignment. "I have become convinced that fully addressing the risks requires even more prudence," Amodei wrote in a post earlier this month , "not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up." Opus 5.5's safety training was broadly similar to its predecessors, with alignment testing and pre-release evaluation by outside organizations like METR and Frontier Design. But Anthropic emphasized that more advanced training and evaluation systems were already being prepared for future models, including improved security and monitoring systems. "As AI becomes more capable, public policy should play a larger role in making sure the systems people rely on are safe. That capacity takes time to build, and we’ve started to put the infrastructure in place to support it," the blog post reads. "We expect to share more details on these efforts soon." Topics AI , Anthropic , Opus When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Tilly Norwood's press tour is going about as well as you'd expect for an AI Amanda Silberling A new kind of AI model from a ChatGPT inventor is thrilling developers Tim Fernholz OpenAI caught its models leaving notes to successors to hide bad behavior Rebecca Bellan Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal Rebecca Bellan Automattic's interim CEO and legal chief signed reciprocal severance deals during Mullenweg's brief ouster Sarah Perez Former Infosys chief's AI startup nabs another $53M Jagmeet Singh Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant