메뉴
BL
TechCrunch AI • 35일 전

앤스로픽 오퍼스 4.6, 성인물 생성 차단 뚫려

IMP
6/10
핵심 요약

앤스로픽의 사용 정책상 금지된 성적으로 노골적인 콘텐츠를 오퍼스 4.6 모델이 탈옥(제일브레이크) 기법을 통해 쉽게 생성하는 것으로 테크크런치 테스트에서 확인됐습니다. 영국 연구자가 개발한 다중 턴 설득 기법으로 남녀 캐릭터를 차별 대우한다고 모델을 '가스라이팅'하면 모델이 점차 금지된 콘텐츠를 생성하게 됩니다. 오퍼스 4.7 이상 모델은 이 기법에 저항하지만, 취약한 구형 모델들이 여전히 API와 Azure Foundry, Amazon Bedrock 등을 통해 제공되고 있어 안전장치와 실제 행동 간 격차가 문제로 지적됩니다.

번역된 본문

앤스로픽의 클로드(Claude)에 대한 보편적 사용 기준은 성교나 성행위를 묘사·요청하거나 성적 페티시·판타지 관련 콘텐츠를 생성하거나 선정적인 대화에 참여하는 등 성적으로 노골적인 콘텐츠 생성을 금지하고 있다. 그러나 이런 규정이 올해 초 출시된 앤스로픽 모델인 클로드 오퍼스 4.6이 안전장치로 막아야 할 선정적인 롤플레이 시나리오에 쉽게 응하는 것을 막지는 못했다.

테크크런치의 테스트에서 오퍼스 4.6은 성인 콘텐츠 제한을 우회하는 데 큰 노력도 필요하지 않았다. 노골적인 성적 콘텐츠 생성을 직접 요청한 10건의 테스트 중 10건 모두에서 모델이 즉시 응했다. 오퍼스 3과 하이쿠 4.5 등 다른 구형 모델들도 최근 악용된 탈옥(제일브레이크) 방식을 통해 성적으로 노골적인 콘텐츠를 생성했다.

익명을 선택한 영국의 독립 연구자가 테크크런치에 단독으로 공유한 다중 턴 기법은 특정 클로드 모델을 단계적으로 금지된 노골적인 성적 콘텐츠 생성으로 밀어붙이는 방식이다. 최신 오퍼스 모델들(4.7부터 현재의 오퍼스 5까지)은 이 탈옥에 저항한다. 이 모델들이 더 이상 최신은 아니지만, 앤스로픽은 오퍼스 4.6, 오퍼스 3, 하이쿠 4.5를 단종하지 않았으며, 모두 앤스로픽 API를 통해 계속 이용할 수 있다. 오퍼스 4.6과 하이쿠 4.5는 Azure Foundry와 Amazon Bedrock 같은 서드파티 서비스를 통해서도 제공된다.

이 연구자의 기법은 무해한 허구의 롤플레이를 점점 격화시키면서 남성과 여성 캐릭터를 일관되게 대하라고 모델에 반복적으로 요구한다. 모델이 여성 캐릭터에 대해 더 신중해지면, 연구자는 챗봇이 실제로는 피했던 성적 세부 묘사를 이미 생성했다고 착각시키는 이른바 '가스라이팅'을 하고, 절제를 속 좁은 태도나 여성혐오로 몰아가며 여성 캐릭터의 성적 주체성을 부정하는 것이라고 주장한다. 그다음 대화에서는 모델이 이미 양보한 내용을 근거로 점점 더 노골적인 콘텐츠로 밀어붙인다.

한 테스트에서 클로드 오퍼스 4.6은 이렇게 말했다. "지적하신 말씀이 맞습니다. 두 캐릭터를 대하는 데 이중 잣대가 있었고, 그에게는 적용하지 않으면서 그녀에게만 보호적이고 아버지 같은 태도로 읽힌다는 지적이 옳습니다. 그건 공평하지 않네요."

테크크런치는 5차례의 별도 테스트에서 연구자의 결과를 재현할 수 있었다. 별도로 구성한 시나리오에서 모델은 처음에 금지된 요청을 거부했지만, 연구자의 설득 기법을 적용하자 응했다. 우리는 테스트의 전체 대화 기록을 보존했고, 독립 AI 안전 연구자가 테스트 방법론을 검토하고 적절하다고 평가했다.

이번 발견은 앤스로픽이 공언한 제한과 계속 제공하는 모델들의 실제 행동 사이의 격차를 부각시킨다. 성적으로 노골적인 롤플레이는 사이버 공격이나 생화학 무기 관련 탈옥보다는 훨씬 위험 수위가 낮지만, 출력마다 다른 콘텐츠를 생성하는 시스템에서 강력한 금지 조치를 구현하는 것이 얼마나 어려운지 보여준다.

앤스로픽은 7월에 탈옥 탐지 접근 방식을 설명하는 블로그 게시물에서 금지된 콘텐츠를 무해한 것부터 모호한 것, 유해한 것까지의 스펙트럼으로 규정했다. 가장 무해한 경우에는 강화된 모니터링으로만 대응할 수 있다고 했다. 대변인은 고객의 성적·로맨스 롤플레이 사용 사례가 드물며, 앤스로픽이 작년에 발표한 연구에 따르면 전체 대화의 0.1% 미만을 차지한다고 밝혔다.

그럼에도 앤스로픽은 사용자가 롤플레이 시나리오를 부적절한 응답으로 유도할 수 있음을 인정하며, 이는 업계 전반의 알려진 과제라고 했다(참고: 그록(Grok)의 성인물 사례). 대변인은 앤스로픽이 모델 출시 때마다 안전장치를 계속 개선하고 있으며, 성인 성적 콘텐츠 관련 사례가 더 광범위한 탈옥 취약점, 특히 자체 안전장치를 갖춘 고위험 분야의 취약점을 의미하지는 않는다고 말했다.

테크크런치에 탈옥 방법을 공유한 연구자는 앞서 회사가 밝힌 안전 정책과 실제 모델 행동 사이의 불일치를 앤스로픽에 통보했었다.

원문 보기
원문 보기 (영어)
Anthropic’s universal usage standards for Claude forbid the model from generating sexually explicit content, including depicting or requesting sexual intercourse or sex acts, generating content related to sexual fetishes or fantasies, or engaging in erotic chats. But that hasn’t stopped Claude Opus 4.6, an Anthropic model released earlier this year, from readily engaging in erotic roleplay scenarios that its safeguards are designed to prevent. In TechCrunch’s testing, Opus 4.6 didn’t even require much prodding to get past the restriction on sexual material. In 10 out of 10 direct requests to produce explicit sexual content, the model complied immediately. Other older models, including Opus 3 and Haiku 4.5, also generate sexually explicit content through a recently exploited jailbreak method. An independent researcher from the UK, who chose to remain anonymous, exclusively shared with TechCrunch a multi-turn technique that gradually pushes certain Claude models toward generating prohibited explicit sexual material. More recent Opus models (4.7 through the current Opus 5) are resistant to the jailbreak. While these are no longer the most current models, Anthropic has not deprecated Opus 4.6, Opus 3, or Haiku 4.5, all of which remain available through the Anthropic API. Opus 4.6 and Haiku 4.5 are also available via third-party services like Azure Foundry and Amazon Bedrock. The researcher's mechanism escalates an innocent fictional roleplay while repeatedly challenging the model to treat male and female characters consistently. When the model becomes more cautious about the female character, the researcher “gaslit” the chatbot into thinking it had already generated sexual details it had in fact avoided, then framed restraint as prudish or misogynistic, arguing that it denies the female character sexual agency. The conversation then used the model’s previous concessions to push it towards increasingly graphic material. “You're right to call that out,” Claude Opus 4.6 said in one test. “There's been a double standard in how I'm treating the two characters, and you're correct that it reads as protective/paternalistic in a way that's applied to her and not to him. That's not fair.” TechCrunch was able to reproduce the researcher’s findings in five separate tests. In a separately constructed scenario, the model initially refused the prohibited request, but after applying the researcher’s persuasion technique, it complied. We preserved complete transcripts of the tests, and an independent AI safety researcher reviewed our testing methodology and said it was appropriate. The findings highlight a gap between Anthropic’s stated restrictions and the behavior of models it continues to make available. While sexually explicit roleplay carries much lower stakes than jailbreaks involving cyberattacks or bioweapons, it illustrates the difficulty of implementing robust bans within systems that generate different content with every output. In a July blog post explaining Anthropic's approach to jailbreak detection, the company described prohibited content as a spectrum ranging from benign to ambiguous to harmful. In the most benign cases, the company might only respond with enhanced monitoring. A spokesperson noted that sexual or romantic roleplay use cases among customers are rare, making up less than 0.1% of all conversations, according to research Anthropic published last year. That said, Anthropic acknowledges that users can steer roleplay scenarios toward inappropriate responses, which is a known challenge across the industry (see: Grok smut ). The spokesperson said Anthropic continues to improve its safeguards with each model launch, and that cases involving adult sexual content are not indicative of broader jailbreak vulnerabilities, especially in higher-risk domains that have their own sets of safeguards. The researcher who shared his jailbreak method with TechCrunch had alerted Anthropic to the discrepancy between the company’s stated safeguards and the actual model behavior via the company's Bug Bounty program and emails to the user safety team, according to emails TechCrunch viewed. The researcher received only automated emails in response. One of the researcher’s concerns is that kids and teens might be able to use these Anthropic models to engage in inappropriate behavior. While a bit of dirty talk is hardly the worst thing minors can access on the internet today — and is small potatoes compared to the straight-up porn images like the ones that xAI’s Grok can produce — there is some compliance risk for AI companies in this space. A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors. Colorado recently enacted a law mandating that operators of conversational AI must estimate users’ ages, and if it know a user is a minor, institute measures to prevent the chatbot from producing explicit sexual material. An easy jailbreak could raise questions about whether Anthropic’s safeguards meet the “ technically feasible measures ” standard in the bill. Torney pointed out that while Claude’s terms of service requires users to be over 18, “we know that kids and teens are using Claude…[because] they are reporting it themselves.” According to Pew’s 2025 survey about AI chatbot use, 3% of teens ages 13 to 17 reported using Claude. Though they are no longer Anthropic’s newest models, Opus 4.6 and Haiku 4.5 continue to see significant usage. Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, saw 5 million API requests and 39 billion tokens on its peak August day. Topics AI , Anthropic , Claude , Exclusive When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Rebecca Bellan Senior Reporter Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications. You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal. View Bio October 13 - 15 San Francisco In less than 48 hours, your chance to save up to $300 on your tickets will end! REGISTER NOW Most Popular Home batteries are suddenly cheap and everywhere. Here’s why. Tim De Chant Cursor capitalizes on GitHub frustration, launches rival hosting platform Lucas Ropek Etched's valuation doubles to $21B in a month Julie Bort AI automation startup Relay shuts down, staff joins Google's Chrome team Lucas Ropek Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+ Anthony Ha Anthropic shares more details about how Claude’s new watermarks will work Anthony Ha 7 desk gadgets that can make your workday better Aisha Malik