메뉴
BL
The Decoder • 46일 전

안스로픽, 전 세계 클로드 출력물에 워터마크 도입

IMP
8/10
핵심 요약

안스로픽(Anthropic)은 EU AI 법안을 준수하기 위해 2026년 8월부터 모든 신규 클로드 모델의 텍스트 및 이미지 출력물에 기계 판독이 가능한 워터마크와 출처 메타데이터를 전 세계적으로 적용할 예정입니다. 이는 AI가 생성한 콘텐츠를 투명하게 식별하고 위변조를 방지하기 위한 조치이며, 향후 개발자와 일반 사용자가 이를 검증할 수 있는 도구도 제공될 계획입니다.

번역된 본문

안스로픽, 전 세계 클로드 출력물에 '일부 편집에서도 유지되는' 워터마크 도입 Matthias Bastian / 2026년 8월 11일

핵심 요약:

  • 안스로픽은 EU AI 법안 실무 강령(Code of Practice)에 서명했으며, 2026년 8월부터 모든 신규 클로드(Claude) 모델에 기계 판독이 가능한 라벨을 부착할 예정입니다. 이 요건은 전 세계적으로 적용됩니다.
  • 텍스트에는 복사해도 유지되는 보이지 않는 워터마크가 적용됩니다.
  • 이미지와 같은 파일에는 C2PA 표준을 따르는 서명된 출처 메타데이터가 첨부됩니다.
  • 안스로픽은 한계를 인정합니다. 사람들이 주로 AI를 자신의 글을 편집하거나 번역하는 데 사용하기 때문에, 워터마크가 있다고 해서 클로드가 해당 내용을 직접 작성했다고 증명할 수는 없습니다. 과도한 편집이나 형식 변환은 표시를 완전히 제거할 수도 있지만, "일부 편집에서는 유지될 수 있습니다."

안스로픽은 AI 생성 콘텐츠의 투명성에 관한 EU AI 법안 실무 강령에 서명했습니다. 2026년 8월부터 새로운 클로드 모델은 텍스트에 워터마크를 삽입하고 파일에는 서명된 출처 메타데이터를 첨부할 것입니다. 2026년 8월 2일 이후 EU에서 출시되는 클로드 모델에는 이러한 라벨링 기능이 기본적으로 탑재되어 제공됩니다.

이 요구 사항은 EU 국경에서만 멈추지 않습니다. 이는 API, 클로드, 클로드 코드(Claude Code), 클로드 코워크(Claude Cowork) 및 클로드 태그(Claude Tag)를 포함한 모든 클로드 제품에 전 세계적으로 적용될 것입니다. 생성된 텍스트는 삽입된 워터마크를 갖게 되며, 생성된 파일에는 디지털 서명된 출처 메타데이터가 부여됩니다. 기존 모델들은 법률에 따른 유예 기간이 주어지지만, 안스로픽은 이미 이를 보완하여 적용하기 위한 작업을 진행 중이라고 밝혔습니다.

또한 사용자와 제3자가 라벨을 확인할 수 있도록 검증 도구를 출시할 계획이지만, 구체적인 시기는 밝히지 않았습니다. 안스로픽에 따르면, 자체 서비스에 클로드를 통합하는 개발자들은 자신의 서비스에 EU AI 법안 제50조의 어떤 요구 사항이 적용되는지 스스로 판단해야 합니다.

AI 콘텐츠를 식별 가능하게 만드는 두 가지 방법

안스로픽은 두 가지 유형의 라벨을 사용할 계획입니다. 회사에 따르면, 클로드가 생성한 텍스트는 의미, 품질, 가독성에 영향을 주지 않는 보이지 않는 워터마크를 포함하게 됩니다. 이 워터마크는 복사 및 붙여넣기를 해도 살아남으며 "일부 편집 과정을 거쳐도 유지될 수 있습니다." 이는 모델 수준에서 적용되므로 어떤 클로드 제품을 사용하든 상관이 없습니다.

.svg, .png, .jpg 이미지를 포함한 지원 파일들은 콘텐츠 출처 및 신뢰성 연합(Coalition for Content Provenance and Authenticity)이 개발한 개방형 C2PA 표준을 기반으로 서명된 출처 메타데이터를 갖게 됩니다. 이 서명은 해당 파일이 클로드에 의해 처리되었음을 나타내며, 향후 변조 여부를 밝혀낼 수 있습니다. 텍스트 워터마크는 AWS, 구글 클라우드(Google Cloud), 마이크로소프트 파운드리(Microsoft Foundry)와 같은 클라우드 파트너를 통해서도 작동해야 하지만, 이러한 플랫폼 자체에서는 서명된 메타데이터를 지원하지 않을 수도 있습니다.

탐지의 한계

안스로픽은 이러한 한계점에 대해 솔직하게 공개하고 있습니다. 워터마크가 탐지되었다고 해서 클로드가 실제로 해당 콘텐츠를 작성했다는 의미는 아닙니다. 사람들은 교정, 번역 또는 요약을 위해 클로드를 지속적으로 사용하므로, 아이디어는 인간에게서 나왔더라도 출력물에는 워터마크가 포함될 수 있습니다.

반대로 워터마크가 없다고 해서 명확해지는 것도 아닙니다. 모델이 워터마크 도입 이전에 출시되었을 수도 있고, 텍스트가 심하게 편집되거나 번역되었을 수도 있으며, 텍스트가 너무 짧아 신뢰할 수 있는 탐지가 불가능할 수도 있고, 형식 변환이나 스크린샷을 통해 메타데이터가 제거되었을 수도 있기 때문입니다.

진짜 시험대는 이 워터마크들이 편집, 형식 변환, 그리고 번역 과정에서 얼마나 잘 버텨내느냐는 것입니다. 만약 제 역할을 한다면, 알려진 워터마크를 확인하는 것이 결과를 유발한 요인이 무엇인지 밝히지 않는 독점적 탐지 방식을 사용하는 팽그램(Pangram)과 같은 도구들보다 훨씬 더 신뢰할 수 있을 것입니다. 서드파티 탐지기들이 안스로픽의 워터마크를 지원하게 되면, 훨씬 더 신뢰할 수 있는 신호를 제공받을 수 있을 것입니다.

사회적으로 여전히 민감한 AI 텍스트 탐지 문제

이 문제에 있어 안스로픽만 혼자인 것은 아닙니다. 구글 딥마인드(Google Deepmind)는 SynthID 워터마킹 시스템을 오픈소스화하고 이를 제미니(Gemini) 모델에 통합했습니다. SynthID는 텍스트 품질을 저하시키지 않고 워터마크를 생성하기 위해 토큰 예측 중에 확률 값을 약간 변경하는 방식을 사용합니다. 이 기술은 여러 환경에서 작동합니다.

원문 보기
원문 보기 (영어)
Anthropic watermarks all Claude outputs globally with marks that "may persist through some editing" Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 11, 2026 GPT-Image-2 prompted by THE DECODER Key Points Anthropic signed the EU AI Act Code of Practice and will equip all new Claude models with machine-readable labels starting in August 2026. The requirement applies worldwide. Text gets invisible watermarks that persist when copied. Files like images get signed provenance metadata following the C2PA standard. Anthropic acknowledges limits: a watermark doesn't prove Claude wrote the content, since people often use the AI to edit or translate their own text. Heavy editing or format changes can also strip the markings entirely; however, it "may persist through some editing." Ask about this article… Search Anthropic has signed the EU AI Act Code of Practice on transparency for AI-generated content. Starting in August 2026, new Claude models will embed watermarks in text and attach signed provenance metadata to files. Claude models that launch in the EU on or after August 2, 2026, will ship with this labeling baked in. The requirement won't stop at EU borders either. It'll apply globally across all Claude products, including the API, Claude, Claude Code, Claude Cowork, and Claude Tag. Generated text will carry embedded watermarks , while generated files will get digitally signed provenance metadata. Existing models get a transition period under the law, but Anthropic says it's already working on retrofitting them. Ad The company also plans to release verification tools so users and third parties can check the labels, though it hasn't said when. Developers that integrate Claude into their products must determine which Article 50 requirements apply to their services, according to Anthropic. Ad DEC_D_Incontent-1 Two methods for making AI content identifiable Anthropic plans to use two types of labels. Text generated by Claude will carry an invisible watermark that doesn't affect its meaning, quality, or readability, according to the company. The watermark survives copying and pasting and "may persist through some editing." It gets applied at the model level, so it doesn't matter which Claude product you're using. Supported files, including .svg, .png, and .jpg images, will carry signed provenance metadata based on the open C2PA standard , developed by the Coalition for Content Provenance and Authenticity. The signature indicates that Claude processed the file and can reveal later tampering. Text watermarks should also work through cloud partners such as AWS, Google Cloud, and Microsoft Foundry, though those platforms may not support signed metadata. Ad Detection has some limits Anthropic is upfront about the limits. A detected watermark doesn't mean Claude actually wrote the content. People use Claude all the time for proofreading, translating, or summarizing, so the output might carry a watermark even though the ideas came from a human. No watermark doesn't clear things up either. The model might have shipped before watermarking rolled out, the text could have been heavily edited or translated, the passage might be too short for reliable detection, or the metadata got stripped through format conversion or a screenshot. Ad DEC_D_Incontent-2 The real test is how well these watermarks survive editing, reformatting, and translation. If they hold up, checking for a known watermark should be more reliable than tools like Pangram, whose proprietary detection methods don't reveal what triggered a result . Third-party detectors could add support for Anthropic's watermark, giving them a more reliable signal. Ad AI text detection remains a sensitive issue across society Anthropic isn't alone here. Google Deepmind open-sourced its SynthID watermarking system , building it into the Gemini models. SynthID slightly tweaks probability values during token prediction to create a watermark without degrading text quality. It works across languages but struggles with text that's been edited after generation. OpenAI has been sitting on a text detector with 99.9 percent accuracy for about two years and still hasn't released it. The reasons include how easily users can beat it through translation or rewriting, the risk of stigmatizing certain groups, and likely worries that a public detector could hurt OpenAI's own business. That risk is especially serious in education, where unreliable detectors can lead to false cheating allegations . At the same time, there are good reasons to want to know when and how much AI was used. Studies show that heavy reliance on AI tools can weaken critical thinking and writing skills , particularly among students who treat them as a shortcut rather than a learning aid . The problem goes beyond academics too, with scammers now enrolling fake students at US colleges and using AI to breeze through coursework and collect financial aid . Anthropic's decision could also affect its business. Claude is popular for knowledge work, especially among school and college students, because even older models produce fairly natural prose. With schools and universities already fighting over AI use in academic work , more reliable detection could make Claude less appealing to those users. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Claude