메뉴
BL
TechCrunch AI • 8일 전

마이크로소프트 임원, AI 스크래핑을 '인류 역사상 최대 노동 절도'로 규정

IMP
9/10
핵심 요약

뉴욕타임스가 오픈AI와 마이크로소프트를 상대로 제기한 저작권 소송의 새로운 공개 자료에서, 마이크로소프트 최고 경영진이 AI 학습용 데이터 수집을 '절도'로 내부적으로 인정한 사실이 드러났습니다. 내부 문서는 AI 챗봇이 뉴스 매체에 '실존적 위협'이며 유료 콘텐츠를 무단으로 긁어갔고 저작권 표시를 의도적으로 제거했음을 시사합니다. 이는 오픈AI의 공정 이용(fair use) 방어 논리를 정면으로 반박하는 증거로, AI 저작권 소송의 향방에 큰 영향을 줄 수 있습니다.

번역된 본문

3년 전 뉴욕타임스가 오픈AI와 마이크로소프트를 상대로 제기한 저작권 소송의 새로운 수정 공개 자료에 따르면, AI 데이터 수집(스크래핑)이 절도에 해당한다는 인정과 AI 제품이 언론 매체에重大한 위협이 된다는 사실이 드러났습니다. 소송 자료에 따르면 마이크로소프트의 최고위 임원은 사내에서 이 회사들의 AI 학습 방식을 '절도'라고 표현했으며, 오픈AI 경영진 역시 자사 AI 모델이 학습 데이터를 제공한 언론인과 출판사에 '실존적 위협'이 된다고 말했습니다. 공개된 자료에는 또한 이 회사들이 유료 회원제 결제벽(paywall)을 몰래 우회해 콘텐츠를 확보하고, 대량 스크래핑으로 학습 데이터셋을 구축했으며, 학습 데이터에서 저작권 고지를 의도적으로 제거한 방식이 상세히 담겨 있습니다. 다만 새로운 정보의 상당 부분은 여전히 봉인된 원본 증거 자료가 아닌 뉴욕타임스 측의 소송 요약서에서 나온 것이며, 아래 인용문들은 원래 맥락 없이 제시된 것이라는 점에 유의해야 합니다.

이번 수정 공개는 3년째 이어지고 있는 소송의 최신 국면으로, 뉴욕타임스는 처음에 이 회사들이 자사 콘텐츠로 생성형 AI 모델을 학습시키며 저작권법을 위반했다고 주장했습니다. AI 기업이 저작권 자료를 AI 학습에 합법적으로 사용할 수 있는지에 대한 명확한 답은 없지만, 법원들은 대체로 학습이 '공정 이용(fair use)'에 해당한다는 AI 기업 측 주장에 우호적인 모습을 보여왔습니다. 공정 이용은 패러디, 뉴스 보도, 비평 등 특정 경우에 허가 없이 저작물을 사용할 수 있게 하는 법 원칙입니다. 이달 초 트럼프 행정부는 오픈AI의 무허가 저작물 사용을 옹호하는 의견서를 소송에 제출하기도 했습니다.

하지만 새로 드러난 여러 인정 사항은 특히 원저작물의 시장을 대체하거나 해쳐서는 안 된다는 공정 이용 요건 측면에서 오픈AI의 방어 논리와 배치됩니다. 예를 들어 마이크로소프트 자체 데이터에 따르면, Copilot '답변 엔진'은 전통적인 빙(Bing) 검색 대비 뉴욕타임스 도메인의 클릭률을 최대 93%까지 떨어뜨린 것으로 나타났습니다. 마이크로소프트 응용과학 책임자 브렌트 헥트(Brent Hecht)가 2024년 1월에 작성한 내부 발표자료는 이러한 하락을 '우리 모델과 웹 전체의 성능을 동시에 해칠' '파멸의 고리(doom loop)'라고 묘사했습니다. 소송 자료에 인용된 마이크로소프트 문서는 “최종 제품이 필수 공급자의 경제적 기반을 위협하는 것은 매우 이례적인 일이지만, 이것이 바로 우리가 LLM 사업의 '콘텐츠 공급망'과 관련하여 만들어낸 상황이다”라고 적고 있습니다.

사티아 나델라(Satya Nadella) 마이크로소프트 CEO도 올해 초 증언에서 “유료 결제벽 뒤에 있는 모든 콘텐츠는 이를 사용하려는 누구든—근거 제공(grounding)이든 학습이든—라이선스를 받아야 한다”고 밝혔으며, '오픈AI가 결제벽 뒤의 정보를 긁어가 학습에 사용했음을 알았다면' '오픈AI에 모델을 재학습하도록 요구했을 것'이라고 말했습니다.

다른 인정 사항들은 공정 이용 판단 기준의 다른 축과 배치됩니다. ChatGPT 총괄인 닉 털리(Nick Turley)는 내부 커뮤니케이션에서 출판사들이 챗봇 같은 제품으로부터 '실존적 위협'에 직면해 있으며, 이러한 제품은 '대체로 대체적(substitutive)'이고 '더 좋아질수록 점점 더 대체적이 될 것'이라고 썼습니다. 그레그 브록만(Greg Brockman) 오픈AI 사장은 이 모델들을 '뉴스에 탁월하다'고 표현했습니다. 나델라 역시 올해 선서 하에 챗봇과의 대화가 “원 출처를 방문할 필요 없이 AI 플랫폼에서 바로 정보를 제공하는 방식으로 대체했다”는 점을 인정했습니다. 이런 표현들은 이 기술이 원저작물을 변혁하기보다 직접 경쟁할 수 있음을 보여줍니다. 마이크로소프트 문서는 생성형 AI가 '기반 모델을 학습시킨 데이터를 만든 바로 그 사람들의 고용을 크게 위협할' '실질적 위험'이 있다고 명시합니다.

복제의 규모 또한 충격적입니다. 문서에는 오픈AI의 중간 학습(mid-training) 데이터셋에만 뉴욕타임스가 출판한 저작물의 복사본이 91,692건 이상 포함되어 있다는 사실이 처음으로 드러났습니다.

원문 보기
원문 보기 (영어)
New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals an admission that AI scraping was tantamount to theft, and that AI products pose a major threat to publications. Per the lawsuit, a top Microsoft executive privately described the companies’ AI training practices as “theft,” and OpenAI’s own leadership said its AI models posed an “existential threat” to the publishers and journalists whose work trained them. The unsealed material also details how the companies allegedly obtained and used that content by bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data. It's worth noting that much of the new information comes from The Times' own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context. The unredacted filing is the latest escalation in the three-year-old lawsuit, in which The New York Times initially alleged the firms violated copyright law by training generative AI models on its content. The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. Several of the new admissions, however, run counter to OpenAI’s fair use defense, particularly the rule's requirement that use doesn’t substitute or harm the market for the original work. For example, Microsoft’s own data shows its Copilot “answer engine” caused click-through rates for The New York Times' domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft’s director of Applied Science, Brent Hecht, in January 2024 describes the decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.” “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” reads the Microsoft document, as quoted in the filing. Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and made clear that, if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.” Other admissions cut against different pillars of the fair-use test: OpenAI’s head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.” OpenAI President Greg Brockman described the models as “excellent at news.” Nadella agreed under oath earlier this year that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.” That kind of language speaks to how the technology could directly compete with, rather than transform, the original work. A Microsoft document states that there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.” The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. In a January 2023 internal memo, Hecht called it “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs' content, including scraping it from the Bing Index. “OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its own commercial products,” the filing reads. “Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.” The companies allegedly assembled the Project Mango data into a training dataset that contains copies of at least 160,903 unique works from the news publishers. In order to get the most out of their scraping, OpenAI employees allegedly came up with a plan to circumvent paywalls without detection. The filings show that when OpenAI researcher Nick Ryder told Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice.” OpenAI employees also allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data. The findings also describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers "wouldn't want model outputting” “copyright notices” to users. OpenAI and Microsoft did not return requests for comment. Topics AI , copyright , Government & Policy , Microsoft , new york times , OpenAI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Rebecca Bellan Senior Reporter Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications. You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Clean tech startup Fluxnium found a way to tap 50,000 years' worth of nuclear fuel Tim De Chant Salesforce and Nvidia's new reasoning model is everything the AI labs should fear Julie Bort Jensen Huang took a call from Trump, and showed off something else, too Connie Loizos The 9 buzziest startups from Y Combinator’s latest Demo Day, according to VCs Marina Temkin Dominic-Madori Davis Tesla says it will finally unveil the second-generation Roadster on October 1 Anthony Ha Revolut confirms customer data breach through fake government requests Jagmeet Singh Matt Mullenweg tells (trolls?) Automattic staff, saying he's back in control after CEO ouster Sarah Perez