메뉴
BL
404 Media • 38일 전

안스로픽 텍스트 워터마킹, 글쓰기에 대한 무관심 드러내

IMP
7/10
핵심 요약

Anthropic이 향후 Claude가 생성하는 텍스트에 워터마크를 삽입하겠다고 발표했습니다. 방식은 텍스트에 아무것도 추가하지 않고, 단어 선택의 무작위성 원천을 자사 알고리즘으로 바꿔 패턴을 남기는 것입니다. 그러나 'grey'와 'overcast' 같은 동의어를 완전히 상호 교환 가능하다고 보는 접근은 작가의 단어 선택을 무의미하게 만든다는 비판을 받고 있습니다.

번역된 본문

이달 초 Anthropic은 향후 Claude 버전이 AI가 생성했음을 보여주는 워터마크가 포함된 텍스트를 생성할 것이라고 발표했다. 당시 Anthropic은 이것이 어떻게 작동할지 설명하지 않아 팟캐스트에서 추측할 수밖에 없었다. 텍스트에 인코딩하는 것인가? 보이지 않는 문자를 포함하는 것인가? 메타데이터로 뭔가 하는 것인가? 이제 주말에 게시된 블로그 포스트 덕분에 Anthropic이 AI의 글쓰기 방식 자체를 바꾸는 방법으로 이를 수행할 것임을 알게 되었다. "텍스트에 아무것도 추가되지 않고 숨겨진 문자도 없습니다"라고 Anthropic은 해당 블로그 포스트에서 밝혔다. "워터마크가 있는 텍스트와 없는 텍스트의 차이는 독자가 구별할 수 없습니다."

이 포스트의 설명에 따르면, Anthropic은 자사와 자사 알고리즘만 아는 방식으로 AI 생성 텍스트의 단어 선택을 미묘하게 변경할 것이다. Anthropic은 "단어를 선택하는 데 사용되는 무작위성의 원천"을 바꾸는 워터마킹 알고리즘을 알고 있으므로, 특정 텍스트가 AI 생성 여부를 탐지하는 도구를 만들 수 있다.

이 연구와 접근 방식은 데이터 과학적인 측면에서 흥미롭지만, Anthropic의 일반인을 위한 설명은 이 회사가 글쓰기 기술이나 인간 작가가 자신의 생각을 전달하기 위해 사용하려는 단어들 간의 미묘한 차이를 얼마나 거의 고려하지 않는지 보여준다. Anthropic은 독자에게 두 문장의 차이를 생각해보라고 한다. "'오늘 날씨가 춥고…'라는 문장을 보자. 다음 단어가 '달콤한(sugary)'일 가능성은 매우 낮다. 하지만 '흐린(overcast)' 또는 '잿빛의(grey)'일 가능성은 꽤 높다. 대부분의 상황에서 모델이 후자의 두 단어 중 무엇을 선택하든 독자에게는 크게 중요하지 않다. 어느 쪽이든 문장의 의미는 거의 같기 때문이다. 이런 경우 선택은 난수로 결정된다"고 Anthropic은 쓰고 있다. "워터마킹은 이처럼 생성된 텍스트 전체에 걸쳐 수없이 발생하는 낮은 위험도의 선택들을 이용해 Claude의 응답에 패턴을 남깁니다. 이 패턴은 독자에게는 탐지할 수 없지만, 이를 인코딩하는 키를 가진 사람은 누구나 탐지할 수 있습니다. 워터마킹이 사용될 때도 여전히 무작위로 선택이 이루어지지만, 무작위성의 원천이 다릅니다."

무언가를 써본 사람이라면 누구나, "오늘 날씨가 춥고 잿빛이다"와 "오늘 날씨가 춥고 흐리다"라는 문장의 차이가 Anthropic의 표현대로 때로는 "위험도가 낮을" 수 있지만 항상 그런 것은 아니라는 점을 이해하리라 희망한다. "잿빛(grey)"과 "흐린(overcast)"은 다른 단어이며, 인간 작가가 특정 맥락에서 하나를 다른 것보다 선택할 수 있는 이유는 무수히 많다. 그러나 이 예시에서 Anthropic의 알고리즘은 이 단어들을 완전히 상호 교환 가능한 것으로 보며, 따라서 워터마킹 알고리즘은 워터마킹을 위해 단어 선택을 이쪽이나 저쪽으로 "조정(nudge)"할 수 있다고 결정했다.

Anthropic은 계속한다. "임의의 난수 생성기를 사용해 다음 단어를 선택하는 대신, 워터마킹은 키와 그 앞에 나온 몇 개의 단어를 사용해 모델이 선택해야 할 단어를 결정합니다. 즉, Claude가 선택하는 단어는 여전히 무작위이지만, 이제 단어의 연속을 확인하고 해당 키를 사용할 때 Claude가 내릴 선택과 일치하는지 볼 수 있습니다. 일치한다면 해당 텍스트가 Claude에 의해 생성되었을 확률을 부여할 수 있습니다."

Anthropic은 "워터마킹은 Claude 출력의 품질에 영향을 주지 않습니다. 독자에게 워터마크된 응답은 워터마크되지 않은 응답과 구별할 수 없으며", "내부 테스트에서 워터마킹이 Claude 텍스트의 내용, 창의성 수준, 가독성에 아무런 영향을 미치지 않는 것을 확인했다"고 주장한다.

사람들은 Anthropic의 워터마킹 시스템에 대해 상당히 화가 나 있으며, 그럴 만하다. 동의어는 때때로 상호 교환 가능하지만 항상 그런 것은 아니라는 점은 Daring Fireball의 John Gruber의 훌륭한 에세이와 저널리즘 학자 Jeff Jarvis의 글에서 지적된 바 있으며, Jarvis는 Anthropic이 "글쓰기를 저평가한다"고 주장했다. 이러한 선택을 통해 "Anthropic은 단어를 대체 가능한 것으로, 언어를 무작위적인 것으로, 선택을 무의미한 것으로 선언하고 있다."

원문 보기
원문 보기 (영어)
Earlier this month, Anthropic announced that future versions of Claude will generate text that includes watermarks showing it was AI-generated. At the time, Anthropic did not explain how this would work, leaving us to speculate on the podcast: Would it somehow encode this into the text? Include invisible characters? Do something with the metadata? We now know, thanks to a blog post over the weekend, that Anthropic will do this by changing how its AI writes altogether. “Nothing is added to the text and there are no hidden characters,” Anthropic wrote in that company blog post . “The difference between watermarked and un-watermarked text will not be distinguishable to readers.” The way it will work, the post explained, is that Anthropic will subtly alter the word choices in AI-generated text in a way that is only known to Anthropic and its algorithms. Anthropic will know the watermarking algorithm, which will change “the source of the randomness used to pick among words” and thus can write a tool to detect whether something has been AI-generated. This research and approach is interesting in a data science kind of way, but Anthropic’s layperson explanation for how this will work shows how little the company thinks about the craft of writing or the subtle differences between words a human author might want to use to convey their thoughts. Anthropic asks us to consider the difference between two sentences: “Take the sentence ‘The weather today was cold and…’. The next word is very unlikely to be ‘sugary.’ But it is quite likely to be ‘overcast’ or ‘grey.’ Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number,” Anthropic writes. “Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it. When watermarking is used, choices are still made at random, but the source of the randomness is different.” Anyone who has written anything would, I hope, understand that the difference between the sentences “The weather today was cold and grey” and “The weather today was cold and overcast” are sometimes “low stakes,” as Anthropic describes, but not always. “Grey,” and “overcast” are different words, and there are any number of reasons why a human author might pick one over the other in a given context. In this example, however, Anthropic’s algorithm sees these words as totally interchangeable and thus its watermarking algorithm has decided that it can “nudge” the word choice one way or the other for the purposes of watermarking. Anthropic continues: “Instead of using an arbitrary random number generator to pick the next word, watermarking uses the key and a few words that come before to settle what word the model should pick. That is, the words that Claude picks are still random, but now, one can check the sequence of words and see if it’s consistent with the choices Claude would make if it was using the key. If it is, one can assign a probability that the text was generated by Claude.” Anthropic claims “Watermarking does not impact the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one,” and that “in internal testing, we’ve seen no impact of watermarking on the content, level of creativity, or readability of Claude’s text.” People are quite mad about Anthropic’s watermarking system, and understandably so. Synonyms are sometimes interchangeable, but not always, as is pointed out in this excellent essay by John Gruber of Daring Fireball , and by journalism academic Jeff Jarvis, in which he claims Anthropic “devalues writing .” In making this choice, “Anthropic declares words fungible, language random, choice meaningless,” Jarvis writes. When I sat down to write this post, I was mad because it seems like Anthropic is putting its thumb on the scale, messing with the outputs of its machine and saying that the resulting text is qualitatively just the same as the other AI text it was probably going to output. But as I began writing this, I realized that my problem is not necessarily with text watermarking but with AI-generated text altogether. It does not matter to me, necessarily, whether the output of Claude’s garbage AI text is one way or is a slightly different way. But it does matter to me that AI data scientists at huge tech companies think that word choice doesn’t matter, or that it is possible to statistically use synonyms wherever without fucking with the meaning of a sentence. Throughout the blog post, Anthropic describes the act of writing as being akin to a probabilistic game of chance. In Anthropic’s own words, its writing is sometimes the result of an “arbitrary random number generator,” and “random” whenever its systems encounter a situation where its tool believes, based on pattern recognition, that the choice between several possible next words isn’t all that important. That may be true for LLM garbage, but is not true for the human experience of writing, which is why human writing almost always feels different than AI writing. This watermarking approach, and Anthropic’s blog post about it, highlights something that should already be clear about a company that famously scanned and destroyed huge numbers of printed books and has trained its LLMs on stolen content: Anthropic does not care about the craft or effort of writing, and sees words as fungible and unimportant. Anthropic says it is making this change as part of the European Union’s new AI regulations, which are well-intentioned but problematic. While it can definitely be useful to have additional ways of detecting AI-generated content, the carelessness with which Anthropic has announced this decision highlights the broader problem with using LLMs to write: They are, as Anthropic notes, probabilistic tools that do not “write” in the way that humans do, rather, they mimic their training data which is, by definition, things that have already happened and been ingested. Contrast this with how Anthropic sees code, something where it says an “exact output is required.” In writing, meanwhile, Anthropic suggests different words are often “equally good.” Over and over again, Anthropic and the researchers who work on this type of watermarking claim that text can be “nudged” in this way without being noticeable to humans or without impacting “quality.” But it is worth noting that the people judging the “quality” of the AI-generated outputs are either data scientists or people asking AI tools to do their writing for them, not, say, people who care about reading or writing. The scientific paper that Anthropic cites was done by Google researchers on a Google watermarking tool called “SynthID,” which Anthropic’s watermarking is based on. In the SynthID study, quality was assessed by randomly putting watermarking on some Gemini outputs, then asking Gemini users to either thumbs-up or thumbs-down the response: “A random fraction of queries were routed to a watermarked model and an equivalent number to the unwatermarked counterpart. The Gemini user interface allows users to provide feedback on model responses via a thumbs-up (good response) and a thumbs-down (bad response). We analysed approximately 20 million watermarked and unwatermarked responses and computed the thumbs-up and thumbs-down rates (both as a fraction of the total number of thumbs-up and thumbs-down feedback received). We found that the thumbs-up rate for the two models differed by 0.01%.” I hope it is clear to anyone who has clicked on this article that asking someone who asked a chatbot something to thumbs up or thumbs down a response is not a very good way of assessing the “quality” of “writing.” The other human assessme