AI가 생성한 텍스트에 숨겨진 워터마크는 글자의 픽셀이나 메타데이터가 아닌, 단어를 선택하는 확률적 '주사위 굴리기' 과정에 비밀 키를 적용하여 작동합니다. 이를 통해 구글 제미나이(Gemini)나 앤스로픽(Anthropic)의 클로드(Claude) 모델은 외형상 문장을 자연스럽게 유지하면서도, 오직 비밀 키를 소유한 자만이 패턴을 통계적으로 감지해낼 수 있는 보이지 않는 표식을 남깁니다. 이는 AI 창작물을 식별하고 표절 및 오용을 방지하는 데 핵심적인 기술적 기반으로 평가받습니다.
번역된 본문
← declaude AI 텍스트 워터마크의 작동 원리. 일반 텍스트에 워터마크를 넣는다는 것은 불가능해 보입니다. 텍스트에는 데이터를 숨길 픽셀이 없고, 복사해서 붙여넣기를 해도 살아남는 메타데이터도 없습니다. 모든 글자는 눈앞에 그대로 드러나 있습니다. 도대체 표식이 어디에 들어갈 수 있을까요? 하지만 이 표식은 분명히 존재합니다. 구글은 2024년부터 제미나이(Gemini) 앱과 웹 환경의 텍스트에 워터마크를 적용해 왔습니다(이 글을 작성하는 시점 기준, API는 공식적인 예외 사항입니다). 그리고 2026년 8월부터 새로운 클로드(Claude) 모델들은 모델 수준에서 텍스트에 마크를 새기고 있으며, 이후 이전 모델들에도 순차적으로 적용될 예정입니다. 이 표식은 보이지 않고, 복사해도 유지되며, 글자 자체에 존재하지 않기 때문에 작동합니다. 그들은 '단어 사이의 선택지' 속에 존재합니다.
글쓰기는 수많은 작은 선택의 연속입니다
이 단계의 핵심 아이디어: 모델은 각자 적절해 보이는 여러 단어들 사이에서 가중치가 적용된 주사위를 굴려 글을 씁니다.
모델이 문장을 작성하는 도중일 때, '다음 단어'가 무엇인지 정확히 알지 못합니다. 자동완성처럼 선호도가 반영된 후보 단어 목록을 가지고 있습니다. 문장의 끝에서 한 단어 앞선 실제 상황의 예시를 살펴봅시다.
작성 중인 문장: The results of the study were quite
🎲 주사위 굴리기 🎲 x20 굴리기
주사위를 굴릴 때마다 후보 목록을 훑고, 확률(막대그래프와 일치함)에 따라 한 단어에 떨어져 문장 뒤에 추가됩니다. 점들은 주사위가 어디에 떨어졌는지 집계합니다. x20을 시도하고 확률의 형태를 띠는 것을 지켜보세요.
결코 변하지 않는 것을 주목하세요: 주사위가 어디에 떨어지든 완벽하고 자연스러운 문장이 만들어집니다.
한 페이지의 텍스트에는 단어마다 수백 개의 이러한 작은 분기점이 포함되어 있으며, 많은 경우 여러 옵션이 동등하게 괜찮습니다. 이러한 유연함이 바로 워터마크의 원재료입니다. 주사위가 어떻게 떨어질지 결정하는 주체는 문장의 의미를 바꾸지 않고도 텍스트 내에 패턴을 숨길 수 있습니다.
비밀 키가 선택에 개입합니다
이 단계의 핵심 아이디어: 비밀 키는 후보 목록을 은밀하게 색칠하고, 특정 색상에 가벼운 가중치를 더해줍니다. 텍스트는 여전히 자연스럽게 읽힙니다.
다음은 고전적인 방식입니다(Kirchenbauer et al. 2023; 구글의 SynthID는 더 정교한 토너먼트 방식을 통해 같은 결과에 도달합니다). 각 분기점에서 비밀 키가 적용된 수학 연산은 후보 단어들을 녹색(green)과 적색(red)으로 나눕니다. 이는 오직 키 보유자만 재현할 수 있는 임의의 색칠입니다. 그런 다음 주사위가 녹색 쪽으로 약간 기울어집니다.
작성 중인 문장: The results of the study were quite
비밀 키 적용 🎲 주사위 굴리기 🎲 x20 굴리기
키 미적용: 이는 모델 자체의 선호도입니다. 점선 윤곽선은 키가 적용되었을 때 이전 확률이 어떻게 변하는지 보여줍니다.
이 과정이 은밀하게 작동하는 데에는 두 가지 이유가 있습니다. 첫째, 가중치는 아주 미세하게 주어집니다. 적색 단어도 여전히 선택될 수 있으며, 단지 확률이 약간 낮아질 뿐입니다. 둘째, 색상은 단어의 고정된 속성이 아닙니다. 키는 직전 단어들의 짧은 흐름을 바탕으로 이를 계산하므로, 특정 앞부분이 오면 녹색이었던 동일한 후보 단어도 다른 앞부분이 오면 적색이 될 수 있습니다.
동일한 네 개의 후보 단어가 여섯 개의 다른 앞부분 이후 키에 의해 다르게 색칠됩니다. 키는 직전의 단어만 볼 뿐이며, 전체 텍스트에서 그 단어의 위치는 보지 못합니다. 오직 녹색으로의 전반적인 기울어짐만 누적되며, 어느 단어가 녹색이었는지는 키 보유자만이 알고 있습니다.
(같은 원리의 다른 형제 기술들: 실제 서비스에 적용 중인 구글의 SynthID는 단순한 가중치 조정 대신 작고 비밀스러운 '토너먼트'로 이를 대체합니다. 모델 자체의 확률에서 몇 개의 후보를 추출하고, 키를 이용해 점수를 매긴 뒤, 키 추첨의 평균이 모델이 의도한 원래의 확률과 정확히 일치하도록 짝을 맞춥니다. OpenAI에서 구축한 아론슨(Aaronson)의 방식은 여기서 더 나아가 아예 주사위 굴리기 결과 자체를 키에서 도출합니다. 수학적 방법은 다르지만 원리는 같습니다. 표식은 '선택지' 속에 존재합니다.)
키를 가진 사람만이 집계할 수 있습니다
이 단계의 핵심 아이디어: 키가 있다면 어떤 텍스트든 다시 색칠하여 단순히 셀 수 있습니다. 워터마크가 있는 텍스트는 우연이라고 보기 힘들 정도로 녹색이 자주 등장합니다.
감지 과정은 텍스트의 문맥을 읽거나 문체를 판단하지 않습니다. 감지기는 단어 위에 키 보유자의 색칠을 다시 재생하고, 녹색이 몇 번이나 나왔는지 횟수를 셉니다. 워터마크가 없거나(또는 올바른 키가 없다면) 녹색은 약 절반의 확률로 나와야 합니다. 동전 던지기처럼 말이죠.
다음은 평범해 보이는 단락입니다. 두 가지 키를 모두 적용해 보세요.
올바른 키로 집계 / 잘못된 키로 집계
📖 계속 읽기: 동일한 마크, 4배의 텍스트 / 녹색 수: 55회 중 –회
← declaude How AI text watermarking works . A watermark in plain text sounds impossible. Text has no pixels to hide data in, and no metadata survives copy-and-paste; every character is right there in front of you. Where could a mark possibly go? And yet the marks are real. Google has watermarked text from the Gemini app and web experience since 2024 (its API is, at the time of writing, a documented exception ), and as of August 2026, new Claude models mark text at the model level, with earlier models to follow. They're invisible, they survive copying, and they work because they don't live in the characters at all. They live in the choices between words . 1. Writing is a series of small choices The one idea in this step: a model writes by rolling weighted dice between several words that would each be fine. When a model is mid-sentence, it doesn't know "the next word." It has a shortlist, like autocomplete, with preferences. Here's a real kind of moment, one word from the end of a sentence: the sentence being written The results of the study were quite 🎲 roll the dice 🎲 roll ×20 Each roll sweeps the shortlist, lands on one word (odds matching the bars) and drops it into the sentence above. The dots tally where the rolls land: try ×20 and watch the pile take the shape of the odds. Notice what never changes: every landing makes a perfectly good sentence. A page of text contains hundreds of these little forks, one per word, and at many of them several options are equally fine. That slack is the raw material. Whoever gets to lean on how the dice land can hide a pattern in the text without changing what it says. 2. A secret key leans on those choices The one idea in this step: the key secretly colours the shortlist and gives one colour a gentle nudge. The text still reads normally. Here is the classic recipe (Kirchenbauer et al. 2023; Google's SynthID reaches the same end by a subtler, tournament-style route). At each fork, secret-keyed maths splits the candidate words into green and red , an arbitrary colouring only the key-holder can reproduce. Then the dice get tilted a little toward green. the sentence being written The results of the study were quite apply the secret key 🎲 roll the dice 🎲 roll ×20 No key applied: these are the model's own preferences. Dashed outlines will show the old odds once the key is on. Two things make this sneaky. The nudge is mild: a red word can still win — it's just a little less likely. And the colouring is not a fixed property of the word: the key computes it from a short run of the words just before, so the same candidate is green after one prefix and red after another: The same four candidate words, coloured by the key after six different prefixes. The key sees the words just before it; the position in the wider text is invisible to it. Only the overall lean toward green accumulates, and only the key-holder knows which words were green where. (Two siblings, same principle. Google's SynthID — the one in production — replaces the nudge with a tiny secret tournament : a few candidates are drawn from the model's own odds, the key scores them, and the bracket is arranged so that, averaged over the key's draws, every word's odds stay exactly what the model intended. Aaronson's scheme, built at OpenAI, skips even that and derives the dice-rolls themselves from the key. Different maths, same principle: the mark lives in the choices.) 3. Whoever holds the key can count The one idea in this step: with the key, you can re-colour any text and simply count. Marked text lands green too often to be luck. Detection doesn't read the text or judge its style. The detector replays the key-holder's colouring over the words and counts how many came up green. Without a mark (or without the right key), green should win about half the time. A coin flip. Here's an ordinary-looking paragraph; try both keys on it: count with the right key count with the wrong key 📖 keep reading: same mark, ×4 the text greens: – of 55 coin flip flag bar (this length) Filled-and-underlined chips are green, dashed outlines are red. The words read identically either way; the colouring exists only in the key-holder's maths. With the wrong key the split is meaningless, and the count sits at chance. (This demo's tilt is drawn strong so you can see it; a production mark leans far more gently and needs correspondingly more text. In this demo's 50/50 model, a 1,500-word document would flag at only ~55% green: small leans become persuasive only through length, which is why short texts are genuinely hard to call.) 4. What editing does to the mark The one idea in this step: the mark lives in runs of untouched wording. Editing erases it exactly where the runs break, and nowhere else. Each word's colouring is derived from a short run of the words just before it (one to a handful, depending on the scheme). So a position only counts as evidence if a short window of the original wording (the word plus its neighbours) survives intact. Here is the same paragraph from step 3, at five edit depths. Drag the slider and watch the highlighted runs shrink. A highlight means that run of wording still matches the original exactly, so the detector can count there. Everything faded is new wording, where there is nothing but coin-flip noise left to count. fix typos light touch tighten sentences heavy edit full rewrite fix typos · surviving windows: – % The verdict reads the surviving fraction measured from the highlights above , projected to a 1,500-word document. Two things to notice: how much a "heavy edit" leaves standing, and how far toward a full rewrite you have to drag before the evidence actually dies. On real implementations (MarkLLM's KGW and EXP schemes on an open model, washed by declaude's full-rewrite route): about 0.5% of windows survive, and detector accuracy falls from essentially certain to a coin flip. The published literature agrees on the shape of this. Light or one-pass paraphrase dilutes the mark rather than deleting it; in Kirchenbauer et al.'s experiments, the detector recovers given enough text, with even human paraphrase becoming detectable again after roughly 800 tokens (about 600 words). What removes the mark is re-composition that shares no runs of wording with the original. That is why a tool that rewrites from the meaning (like declaude 's full-rewrite route) is what actually erases this family of marks, and why a light pass that keeps most of the phrasing does not. One boundary stated plainly: those numbers come from open implementations we can measure. Anthropic's production scheme is undisclosed, so no one outside Anthropic can yet run this test against Claude's own mark. 5. What this means in practice The one idea in this step: detection is private, probabilistic, and about processing , not authorship. Only the key-holder can check. Your teacher, editor, or favourite "AI detector" website cannot run this test; a genuine check needs the provider's secret key, or a checking service the provider runs. Google runs an early-access detector portal for SynthID; Anthropic says detection tooling is forthcoming. A watermark check is not an "AI detector." Tools like GPTZero guess from style and are famously unreliable. A watermark is the opposite: a deliberate, key-gated statistical test. Don't let the two blur. A found mark means "processed by", not "written by". Anthropic's own documentation notes that human text merely proofread or translated by Claude picks up the mark. And absence proves even less: old models or heavy editing yield clean results on genuine AI text. Short and low-choice text carries little mark. Evidence grows with length, and text with only one right continuation (code, quotations, lists of facts) offers the dice too little slack to hide anything in. Certain marks outlive a rewrite. Schemes keyed on the word itself rather than its neighbours hold up far better: a same-meaning rewrite keeps enough of the words that much of the mark survives. (Th