메뉴
BL
Wired AI • 37일 전

개발자들, 클로드 워터마크 우회법 이미 찾아내

IMP
8/10
핵심 요약

Anthropic이 EU AI법 준수를 위해 Claude 생성 텍스트에 보이지 않는 워터마크(SynthID)를 삽입하겠다고 발표한 지 불과 4시간 만에 개발자 기욤 메이어가 워터마크 제거 코드를 공개해 큰 인기를 얻고 있습니다. 이 사건은 AI 콘텐츠 투명성 규제와 기술적 우회 사이의 공방가 현실화되었음을 보여줍니다.

번역된 본문

Anthropic이 Claude 모델이 AI 생성 콘텐츠에 보이지 않는 기계 판독 가능 워터마크를 전 세계적으로 삽입할 것을 확인한 지 불과 네 시간 만에, 개발자 기욤 메이어(Guillaume Meyer)는 자신의 우회 코드를 공개했다. Claude 생성 텍스트에서 워터마크를 제거하는 그의 코드는 GitHub에서 큰 인기를 얻었고, X에서 2만 회 이상 북마크되었으며, 100명 이상의 기여자가 참여했고 더 많은 사람들이 자신의 프로젝트에 이 기술을 통합하고 있다.

한 AI 전문가는 "Anthropic이 Claude 텍스트에 워터마크를 삽입하고 있지만... 하루 만에 이 문제는 사실상 끝났다"고 쓰며, 쇠사슬을 끊고 구겨진 EU와 Anthropic 깃발 위에 서 있는 메이어의 이미지를 함께 게시했다.

메이어와 다른 사람들은 Anthropic이 지난주 EU AI법을 준수하기 위해 Claude가 워터마크를 도입할 것이라고 발표한 후 워터마킹의 작동 방식을 조사하기 시작했다. 메이어는 WIRED에 모든 AI 생성 콘텐츠에 라벨이 붙어야 한다는 생각에 동의하지 않아 워터마킹을 회피하려는 사람들도 있지만, 자신을 포함한 다른 사람들은 단순히 기술적 도전을 즐긴다고 말했다. 그는 프리랜서 콘텐츠 작가와 소셜 미디어 크리에이터들도 이 코드 사용을 도와달라고 연락해 왔다고 밝혔다.

이번 달 초 시행된 새 규정에 따르면 Anthropic과 OpenAI 같은 모델 제공자는 합성 음성, 이미지, 비디오, 텍스트에 라벨을 붙여 기계가 AI 생성 콘텐츠로 탐지할 수 있도록 해야 하며, 위반 시 연매출의 최대 3%에 해당하는 벌금을 물 수 있다. 규정은 제공자가 우회 도구를 판매하는 것을 금지하지만, 독립적인 도구에 대한 법적 제한은 없다.

메이어는 "나는 투명성에 반대하지 않고 콘텐츠 귀속 attribution을 지지한다. 다만 워터마킹 자체가 심각한 단점과 위험을 가진 매우 나쁜 해결책이라고 생각한다"고 말했다. 그는 오탐 위험과, 워터마크가 AI 사용 정도(가벼운 사용인지 많은 사용인지)를 구별하지 못할 가능성을 우려한다. 특히 그는 프랑스어 원어민으로서 글쓰기를 편집할 때 Claude나 Grammarly 같은 AI 도구를 자주 사용한다. Anthropic조차 텍스트가 Claude와 접촉했을 확률만 생성할 수 있다고 인정하는 상황에서 워터마크를 증거로 사용하면, 고용주가 지원자를 부당하게 거부하거나 탐지기에 걸렸다는 이유만으로 연구자들에 대한 과장된 AI 사용 고발로 이어질 수 있다고 그는 말한다.

Anthropic은 Claude가 단어와 구문을 선택하는 방식에 패턴을 남겨 워터마크를 삽입한다. 이 패턴은 인간 독자에게는 알아차릴 수 없지만 찾는 법을 아는 기계는 탐지할 수 있다. 이 방식이 Claude의 출력에 영향을 미치기 때문에 일부 사용자는 응답 품질 저하를 우려하지만, Anthropic은 그렇지 않을 것이라고 주장한다. SynthID라 불리는 이 기술은 Google이 개발한 것으로, Google은 2023년부터 AI 생성 콘텐츠에 워터마크를 삽입하는 데 사용해 왔다. 컴퓨터 과학자 스콧 애런슨(Scott Aaronson)은 OpenAI에서 근무할 때 비슷한 방법을 제안했지만, 워터마크가 고객을 제품에서 멀어지게 할까 우려한 회사가 이를 배포하지 않았다고 말한다.

메이어의 제거 방법은 워터마크를 삽입하지 않는 대규모 언어 모델을 사용해 동의어를 교체하고 콘텐츠를 약간 재구성하는 여러 버전의 재작성 텍스트를 생성하는 방식이다. 물론 이는 워터마크를 삽입하지 않는 다른 대규모 언어 모델을 사용한다는 전제에 의존하는데, OpenAI, Microsoft, Meta 등 제공자 190개 기업이 EU 투명성 실무 강령에 서명했기 때문에 안전한 선택이 아닐 수 있다. 이들 연구소 중 얼마나 많은 곳이 워터마크를 구현할지는 아직 지켜봐야 한다. 워터마크는 8월부터 출시되는 모든 신규 모델에 포함되어야 하며, 기존 모델에도 12월까지 통합되어야 한다. Anthropic이 워터마크 탐지 소프트웨어를 공개하기 전까지는 이 도구가 작동한다는 확신은 없지만, Claude 워터마킹의 기반이 되는 SynthID-텍스트 접근 방식을 이해하면 상당히 확신할 수 있다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Within four hours of Anthropic confirming that Claude models would globally embed invisible, machine-readable watermarks into any AI-generated content, developer Guillaume Meyer had published his override. His code to remove watermarks from Claude-generated text has since gone viral on GitHub, has been bookmarked more than 20,000 times on X, and has drawn more than 100 contributors, with many more incorporating the technology into their own projects. “Anthropic is embedding watermarks in its Claude texts … the issue is practically history just one day later,” wrote one AI specialist, accompanied by an image of Meyer breaking out of chains and standing on crumpled EU and Anthropic flags. Meyer and others started investigating how watermarking works after Anthropic announced last week that Claude would adopt it in order to comply with the European Union’s AI Act . Some are trying to evade the watermarking because they disagree with the idea that all AI-generated content should be labeled as such, Meyer told WIRED, while others, including himself, say they simply relish the technical challenge. Freelance content writers and social media creators have also contacted Meyer asking for assistance using the code, he says. The new rules, which came in earlier this month, stipulate that model providers like Anthropic and OpenAI must label synthetic audio, image, video, or text so that this material can be detected by a machine as AI-generated—or face fines of up to 3 percent of annual turnover. While the rules say providers cannot market circumvention tools, there is no legal restriction on independent tools. “I'm not against transparency, and I'm all for content attribution,” says Meyer. “I just think watermarking in itself is a really bad solution, because it has major drawbacks and risks.” He is concerned about the risk of false positives and that the watermarking might not distinguish between light or heavy AI use, especially since, as a native French speaker, he often uses Claude and other AI tools like Grammarly to edit his writing. Using the watermark as evidence–when even Anthropic admits it can only generate a probability that the text has been touched by Claude–could lead to employers unfairly rejecting candidates or overblown accusations of researchers using artificial intelligence just because the detector flags it, he says. Anthropic watermarks text invisibly by leaving a pattern in Claude’s choice of words and phrases that is indiscernible to a human reader but would be detectable by a machine that knows how to look for it. Because this influences Claude’s output, some users are concerned this will degrade the quality of Claude’s responses, though Anthtropic insists this won’t be the case. The technique, called SynthID , was developed by Google, which has been using it to watermark its AI-generated content since 2023. Computer scientist Scott Aaronson proposed a similar method when working at OpenAI but says the firm never deployed it because the company was worried that watermarks would put customers off its product. Meyer’s removal method uses a non-watermarking large language model to generate multiple rewrites, swapping in synonyms and slightly reorganizing content. Of course, this relies on using other large language models which do not insert watermarks—possibly not a safe bet since 190 organizations—providers OpenAI, Microsoft, and Meta among them—have signed the EU’s transparency code of practice. It remains to be seen how many of these laboratories are going to implement their watermarks, which must be included in all new models released from August and must be integrated into existing models by December. While there’s no certainty this tool works until Anthropic releases the software it uses to detect a watermark, understanding the basic SynthID-text approach underpinning Claude’s watermarking makes them fairly sure the method works, says Wayne Pan, chief technology and cofounder at Silicon Valley–based sovereign AI startup Haimaker. He incorporated Meyer’s open-source tool into his platform because he similarly disliked the idea of Claude watermarking content even when it’s only been lightly edited and disagreed with the watermark being invisible to the user. Other coders have developed their own removal tools: Software engineer Erik Hughes took 15 minutes to knock up a tool with Claude that removes invisible and look-alike characters, reorders sentences within paragraphs, and swaps several words for synonyms. Leon Chlon, a Visiting Fellow at the University of Oxford, says the watermarks can be removed by condensing Claude’s response, translating it into a dialect like Arabic, which has very different semantics compared to English, and then translating it back. Anthropic itself acknowledged that heavily edited, paraphrased, or translated content might not carry a watermark. In a statement to WIRED, a spokesperson for Anthropic said: "We're adding marking to Claude's output to comply with the EU AI Act, and other labs are taking similar steps. It’s hard to identify AI-generated text, and this gives people better tools for identification. Text from supported Claude models, including output from Claude Code, will carry an invisible watermark, and it doesn't change the meaning, quality, or readability of Claude's responses. We also plan to ship a text-detection API so users can do more of this themselves.” Anthropic says it’s working out how to implement watermark detection for text and plans to release a tool to do so soon—at which point developers will finally be able to see whether their methods are foolproof. It’s also continuing to work on improving the watermarking system. “I think they wanted to show that they're in good faith doing it,” says Pan, “but I don't think you can ever have a watermark that will withstand everything.”