메뉴
BL
The Decoder • 7일 전

AI 학습 '공정 이용' 논리, 내부 임원들조차 '경악할 도난'이라 불러 흔들린다

IMP
9/10
핵심 요약

뉴욕타임스 등 미국 억론사들이 OpenAI와 마이크로소프트를 상대로 수십억 달러의 손해배상을 청구한 소송에서, AI 기업 내부 문서와 증언이 공정 이용(fair use) 방어 논리를 정면으로 반박하고 있습니다. 마이크로소프트 임원은 AI 학습 데이터 수집을 '인류 역사상 최대 규모의 도난'이라 불렀고, 내부 자료는 챗봇이 원 저널리즘을 대체하며 미디어 산업에 '둠 루프'를 만들고 있음을 보여줍니다.

번역된 본문

AI 학습의 '공정 이용' 근거, 회사 내부 인사들이 '경악할 도난'이라 부르며 흔들려

뉴욕타임스를 포함한 여러 미국 언론사들이 AI 학습 과정의 저작권 침해를 이유로 OpenAI와 마이크로소프트에 수십억 달러의 손해배상을 요구하고 있다.

원고 측은 경영진들이 자사의 공정 이용 방어 논리에 의문을 제기하고 챗봇을 원 저널리즘의 대체재로 묘사한 내부 메시지와 선서 증언을 증거로 제시했다.

소장은 또한 OpenAI가 유료회원 전용 기사(페이월)를 체계적으로 우회하고, 학습 데이터 라이선스 조건을 위반했으며, 소송 제기 후 증거를 은폐하기 위한 필터를 도입했다고 주장한다.

뉴욕타임스와 다른 원고들의 새 법원 제출 서류는 OpenAI와 마이크로소프트 경영진의 이전에 공개되지 않았던 내부 이메일과 선서 증언을 인용하고 있다.

뉴욕타임스와 여러 언론사들이 뉴욕 연방지방법원에 통합 준판정(요약판결) 서면을 jointly 제출했다. 시카고트리뷴과 덴버포스트를 거느린 데일리뉴스 그룹, CNET·IGN·PCMag를 보유한 지프 데이비스(Ziff Davis)가 타임스와 함께 원고로 나섰다. 매더존스가 속한 조사보도센터(Center for Investigative Reporting)와 더 인터셉트(The Intercept)도 원고에 포함된다. 파이낸셜타임스에 따르면 원고 측은 수십억 달러의 손해배상을 구하고 있다.

92쪽에 달하는 이 서면은 2023년 12월 뉴욕타임스 소송을 시작으로 여러 소송을 묶은 다구역 통합 소송(MDL)의 일환이다.

경영진의 내부 발언이 공정 이용 방어 논리를 무너뜨린다

이 서면은 증거개시 과정에서 공개된 내부 이메일, 슬랙 메시지, 선서 증언을 인용하며, 이중 일부는 AI 기업들의 공정 이용 방어 논리를 정면으로 반박한다.

마이크로소프트 응용과학 국장 브렌트 헤크트(Brent Hecht)는 이 관행을 '전례 없는 규모의 경악할 도난'이자 어쩌면 '인류 역사상 최대 규모의 노동 도난'이라고 불렀다. 그는 또 공정 이용 방어가 성공한다면 '공정 이용이라는 개념을 완전히 조롱하는 것'이 될 것이라고 적었다. 마이크로소프트는 FT에 이 발언들은 '한 직원의 개인적 견해'이며 '법적 분석이 아니다'라고 밝혔다.

미국 저작권청도 2025년 5월, AI 기업들이 데이터를 복사하는 규모를 고려하면 공정 이용이 광범위하게 적용될 수 없다고 결론지었다. 이 보고서를 총괄한 관리는 이후 트럼프 행정부에 의해 해고되었다.

서면에 따르면 OpenAI의 챗GPT 총책임자 닉 털리(Nick Turley)는 언론사들이 '실존적 위협'에 직면했다고 적었다. 그는 이 제품들이 '명백히 대부분 대체적'이며 '성능이 좋아질수록' 언론사의 콘텐츠를 점점 더 대체할 것이라고 말했다. 한 OpenAI 엔지니어는 '아무리 링크를 눈에 띄게 표시해도 사용자는 클릭하지 않을 것'이라고 덧붙였다.

마이크로소프트의 사티아 나델라(Satya Nadella) CEO는 선서 하에 챗봇 대화가 원 출처 방문을 대체했다는 사실을 확인했다. 마이크로소프트는 FT에 대한 성명에서 그의 증언은 '사람들이 정보를 찾고 소비하는 방식의 변화에 대한 광범위한 원칙과 진행 중인 변화'에 관한 것이며, 소송의 쟁점인 '저작권 문제에 대한 결론'이 아니라고 주장했다.

마이크로소프트의 내부 문서는 '둠 루프(doom loop)'라 불리는 자기강화 악순환을 묘사한다. '우리의 AI 콘텐츠 전략이 우리 모델의 성능과 웹 전체의 성능을 동시에 저해할 둠 루프를 시작했다. 최종 제품이 필수 공급자의 경제적 기반을 위협하는 것은 매우 이례적이지만, 그것이 우리가 LLM 사업의 콘텐츠 공급망에 관하여 만들어낸 상황이다.'

마이크로소프트 자체 데이터에 따르면 Copilot의 클릭률은 전통적인 빙(Bing) 검색보다 훨씬 낮았다. 뉴욕타임스는 8793%, 데일리뉴스 그룹은 8391%, 지프 데이비스는 51~94% 감소했다.

OpenAI는 내부에서 데일리뉴스 그룹의 핵심 사업인 지역 뉴스를 '예쁜'(prett…) 것으로 묘사한 것으로 나타났다.

원문 보기
원문 보기 (영어)
AI training built on fair use looks shaky when the companies' own people call it "astonishing theft" Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 18, 2026 Nano Banana Pro prompted by THE DECODER Key Points Several US media companies, including The New York Times, are seeking billions of dollars in damages from OpenAI and Microsoft over alleged copyright infringement in AI training. The plaintiffs cite internal messages and sworn testimony in which executives questioned their own fair use defense and described chatbots as substitutes for original journalism. The filing also accuses OpenAI of systematically bypassing paywalls, violating license terms for training data, and deploying filters to suppress evidence after the lawsuits were filed. Ask about this article… Search A new court filing by The New York Times and other plaintiffs cites previously undisclosed internal emails and sworn testimony from OpenAI and Microsoft executives. The New York Times and several other media companies have filed a joint summary judgment brief in US District Court in New York. Joining the Times are the Daily News group, which includes the Chicago Tribune and Denver Post, and Ziff Davis , which owns CNET, IGN, and PCMag. The plaintiffs also include the Center for Investigative Reporting, home to Mother Jones, and The Intercept . The plaintiffs are seeking billions of dollars in damages, according to the Financial Times . The 92-page brief is part of consolidated multidistrict litigation that brings together several lawsuits, beginning with the New York Times suit filed in December 2023. Ad Executives' internal statements undercut the fair use defense The brief draws on internal emails, Slack messages, and sworn testimony disclosed during discovery, including statements that undercut the AI companies' fair use defense. Ad Microsoft's director of applied science, Brent Hecht, called the practice "an astonishing theft of unprecedented proportions" and possibly the "largest theft of labor in human history." He also wrote that a successful fair use defense would arguably "make a complete mockery of the idea of 'fair use.'" Microsoft told the Financial Times these comments "reflect one employee's individual perspective" and "are not a legal analysis." The US Copyright Office also concluded in May 2025 that fair use cannot apply broadly given the sheer scale at which AI companies copy data. The official who oversaw the report was later fired by the Trump administration . Ad Nick Turley, OpenAI's head of ChatGPT, wrote that publishers face an "existential threat," according to the brief. He said the products "are largely substitutive, period" and would increasingly replace publishers' offerings "as they get better." An OpenAI engineer added that "no matter how prominently we show the links, users won't click." Microsoft CEO Satya Nadella confirmed under oath that chatbot conversations had replaced visits to original sources. In a statement to the FT, Microsoft said his testimony concerned "broad principles and changes under way in how people find and consume information." It was not a "conclusion about copyright questions" at the center of the case, the company said. Ad An internal Microsoft document describes a self-reinforcing cycle it calls a "doom loop." Ad Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain. Microsoft's own data shows that click-through rates on Copilot were much lower than on traditional Bing search. Rates were down 87 to 93 percent for the New York Times, 83 to 91 percent for the Daily News group, and 51 to 94 percent for Ziff Davis. OpenAI internally described local news, a core business for the Daily News group, as a "pretty common quer[y]" in ChatGPT. Other studies have also found that chatbot answers sharply reduce traffic to the open web . OpenAI co-founder Greg Brockman discussed the models' ability to handle news in an internal message. we are excellent at news btw. every time i do any generative stuff on NYT it seems to predict the next sentence pretty well. Paywall workarounds and license restrictions complicate OpenAI's defense The plaintiffs describe how OpenAI systematically bypassed paywalls and ignored terms of service. When an employee told Brockman about "a hack to get around nytimes paywall," he replied "ah nice." Around 2017, Brockman also wrote that he was "deeply motivated by the gazillions" he hoped to earn by commercializing OpenAI's technology. OpenAI's corporate representative testified that he knew of no method for detecting paywalled content in the training data and no effort to remove it. The company's standard crawling process "did not include reviewing websites['] ... Terms of Use or Service." Nadella testified under oath that "anything that is paywalled should be licensed by anyone who wants to use it." He said he would have forced OpenAI to retrain its models had he known the company had scraped paywalled content and used it for training. OpenAI also acquired the "New York Times Annotated Corpus," a collection of 1.8 million articles, through a third party. Its license restricted use to "non-commercial linguistic education, research and technology development." OpenAI employees knew using the corpus to train models "would not be appropriate," but did so anyway. Plaintiffs accuse OpenAI of suppressing evidence and challenge all four fair use factors Immediately after the lawsuits were filed, OpenAI built a filter to suppress output most likely drawn from the plaintiffs' publications, according to the brief. Content from companies that hadn't sued remained unaffected. The plaintiffs argue that the filter was designed not to protect copyrights but to prevent them from gathering evidence. Microsoft's Hecht had raised concerns that OpenAI might deploy such a filter, calling it an "accidental cover up." He warned it would result in "people who have a right over the content having less visibility into what was used for training." Even so, the brief includes numerous examples of problematic outputs, with ChatGPT reproducing exact copies and summaries of NYT articles on request, including articles behind a paywall. The plaintiffs also argue that the fair use defense fails on all four statutory factors. They say the use is substitutive and commercial rather than transformative, while the articles are expressive works at the core of copyright protection. The defendants also copied entire works, even though their own experts admitted that no single work was necessary. On market harm, the plaintiffs point to existing licensing markets and those that could realistically develop. OpenAI and Microsoft have signed licensing deals with other publishers, as have Amazon, Google, Meta, and Perplexity. But rather than pay for the plaintiffs' content, the defendants took it for free and traded it among themselves, the brief argues. A functioning licensing market makes the fair use argument harder to sustain. The defendants can't credibly claim licensing wasn't possible when they're already paying for comparable content from other publishers. AI can generate a million "pink slime" articles for $6,800 The plaintiffs also challenge the argument that AI models don't compete with publishers, further undercutting the fair use defense. They argue that OpenAI's products let users flood the market with "pink slime," meaning low-quality, often plagiarized pseudo-news. At OpenAI's API prices, generating one million news-style articles of 500 words each costs about $6,800, without involving a single journalist. The brief cites Prism News, which used just four employees to run 200 AI-generated publications posing as local newsrooms and hobby sites. Ope