메뉴
BL
404 Media • 8일 전

'둠 루프': 오픈AI·마이크로소프트, LLM이 웹을 파괴하고 도용으로 만들어졌음을 인정

IMP
9/10
핵심 요약

뉴욕타임스 vs 오픈AI 저작권 소송에서 공개된 수정되지 않은 법원 서류에 따르면, 마이크로소프트와 오픈AI 임원들이 내부 문서와 증언에서 LLM이 도용된 콘텐츠로 학습되었으며 웹 생태계를 파괴하는 '둠 루프'를 만들고 있다고 인정했습니다. 마이크로소프트 내부 문서는 이를 '전례 없는 규모의 놀라운 도난'이자 '인류 역사상 최대의 노동 도난'으로 표현했으며, 빙에서 뉴스 사이트 클릭률이 90% 이상 급감했다는 증언도 포함되어 있습니다.

번역된 본문

마이크로소프트와 오픈AI에서 AI 업무를 맡고 있는 임원들이 비평가들이 줄곧 주장해온 내용을 인정했다. 대규모 언어 모델(LLM)은 마이크로소프트 임원이 '전례 없는 규모의 놀라운 도난'이자 '인류 역사상 최대의 노동 도난'이라고 부른 것 위에 세워진 약탈적인 기술이라는 것이다. 마이크로소프트 내부 문서는 생성형 AI 제품이 '웹 전체'를 죽이는 '둠 루프(doom loop)'를 만들어냈다고 밝혔다. 이러한 진술과 그 밖의 여러 '가면 벗기' 순간들은 수년간 법정을 오가며 진행되어온 거대한 뉴욕타임스 vs 오픈AI 저작권 소송에서 목요일 공개된 수정되지 않은(unredacted) 법원 제출 서류에 핵심적으로 담겨 있다.

뉴욕타임스 측 변호사들이 약식 판결(summary judgment, 즉 법원에 판결을 내려달라고 요청하는 신청)을 요청하며 제출한 이 서류에는 그동안 마이크로소프트와 오픈AI의 요청으로 봉인되거나 편집(검열)되어 왔던 문서와 증언에서 두 회사 임원들이 한 일련의 인정 사실들이 정리되어 있다. AI 기업들이 이를 대중에게 숨기고 싶어했던 이유는 쉽게 짐작할 수 있다. 이 진술들을 종합해보면, LLM이 어떻게 학습되었고 어떻게 작동하며 인간 노동에 어떤 즉각적인 위협이 되는지에 대한 가장 뼈아픈 고발 중 하나이기 때문이다. 이는 AI가 더 강력해지고 기업들이 서사를 '초지능' AI의 가상 실존적 위험으로 옮기려 하더라도, 이미 만들어진 도구들은 인간의 창의성과 노동을 훔쳐 만들어졌으며 정의상 인간 노동 시장에 실존적 위협이라는 사실을 일깨워준다.

소송에 인용된 마이크로소프트 내부 문서에는 "전 세계 수백만 명이 곧 대규모 모델이 자신들의 작업을 모두 '빨아들이는' 것을 전례 없는 규모의 놀라운 도난으로 여기게 될 것"이라고 적혀 있으며, "콘텐츠를 만든 사람들 중 거의 아무도 이런 방식으로 사용되기를 의도하지 않았고, 사용에 대한 보상도 받지 못한다"고 덧붙였다.

이 서류는 뉴욕타임스 변호사들이 작성했지만, 대부분 빅테크 임원들의 진술과 인터뷰로 구성되어 있으며, 이들은 LLM이 대부분 도용된 콘텐츠로 학습되었다는 점, 훔친 대상인 인간 작가·예술가·미디어 기업에 실존적 위험이 된다는 점, 그리고 자사 제품이 웹을 잠식하고 훔친 대상 기업들의 사업을 파괴하는 '둠 루프'를 시작했다는 점을 모두 인정했다.

수정되지 않은 서류는 디지털 미디어 기업들을 대표하는 업계 단체인 Digital Content Next의 CEO 제이슨 킨트(Jason Kint)가 발견했다. 킨트는 거대 AI 저작권 소송들을 면밀히 추적하며 게시해오고 있다.

법원 제출 서류는 마이크로소프트 내부 문서를 인용하고 있는데, 이 문서에 따르면 AI 제품은 인간 콘텐츠 제작자에게서 훔친 뒤, 훔친 사람들과 웹사이트의 클릭을 잠식함으로써 그들의 비즈니스 모델을 파괴한다. 해당 문서는 이렇게 말한다. "우리의 AI 콘텐츠 전략은 우리 모델의 성능과 웹 전체를 동시에 저해할 '둠 루프'를 시작했다. 최종 제품이 필수 공급자들의 경제적 기반을 위협하는 것은 매우 이례적인 일이지만, 우리는 LLM 비즈니스의 '콘텐츠 공급망'에 대해 바로 그런 상황을 만들어냈다."

사티아 나델라(Satya Nadella) CEO를 포함한 마이크로소프트 임원들은 선서 하에, 뉴욕타임스와 다른 뉴스 사이트의 콘텐츠를 긁어온 후 해당 뉴스 사이트의 클릭수가 완전히 폭락하여 빙(Bing)에서 90% 이상 감소했다고 증언했다. 법정 절차에서 확보된 문서에 따르면 오픈AI는 "뉴욕타임스 유료벽(paywall)을 우회하는 핵(hack)"을 만들었고, 이에 대해 오픈AI 공동창립자 그렉 브록먼(Greg Brockman)은 "아, 멋지네"라고 말했다. 마이크로소프트 임원 브렌트 헤흐트(Brent Hecht)는 LLM이 "경제적 가치를 공급망 아래로 분배할 방법 없이" 콘텐츠를 훔치며, 이는 필연적으로 콘텐츠 제작자들의 경제적 안정을 위협한다고 작성했다. 오픈AI 정책 디렉터 잭 클라크(Jack Clark)는 회사가 "사회의 '문화'를 정의하는 사람들의 노동을 대체하는 시스템"을 만들고 있다고 썼으며, 마이크로소프트도 정책 문서에서 이를 인정했다.

원문 보기
원문 보기 (영어)
Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.” Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI. It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market. “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” an internal Microsoft document cited in the case read, adding “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.” The filing was written by lawyers for the New York Times but is largely comprised of statements and interviews with big tech executives that admit both that LLMs are largely trained on stolen content, that they represent an existential risk for the human writers, artists, and media companies that they stole from, and that their products have started a “doom loop” that is eating the web and destroying the businesses that these companies stole from. The unredacted filing was found by Jason Kint, the CEO of Digital Content Next, a trade organization that represents digital media companies. Kint has been closely following and posting about massive AI copyright lawsuits . The court filing cites an internal Microsoft document that found that AI products steal from human content creators, then cannibalize clicks from the people and websites they’ve stolen from, thereby destroying their business models. “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said. Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing. Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.” OpenAI’s policy director Jack Clark wrote that the company was “creating systems that substitute for the labor of the people that define the ‘culture’ of society” and Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.” OpenAI called itself an “existential threat” to news publishers, and an OpenAI software engineer testified that “no matter how prominently we show the links, users won’t click.” None of this is at all surprising to anyone who has been paying attention to the development of generative artificial intelligence, but the document, taken in whole, is a real they-admit-it situation. OpenAI’s and Microsoft’s lawyers have been trying to argue that their model training is fair use and transformative under copyright law and that they are building something that is fundamentally different from the human labor that it was trained on. But internally, these executives know that what they have built has been built on stolen content and that the products they’ve made are cannibalizing the sources they’ve stolen from and destroying the internet as we know it. About the author Jason is a cofounder of 404 Media. He was previously the editor-in-chief of Motherboard. He loves the Freedom of Information Act and surfing. More from Jason Koebler