메뉴
HN
Hacker News • 6일 전

마이크로소프트 디렉터: AI 스크래핑은 '인류 역사상 최대 규모의 노동 도난'

IMP
8/10
핵심 요약

뉴욕타임스가 OpenAI와 마이크로소프트를 상대로 제기한 저작권 침해 소송에서, 양사 내부 문서를 인용한 법정 서류가 공개되었습니다. 마이크로소프트 응용과학 디렉터 브렌트 헥트는 AI가 콘텐츠를 긁어모으는 것을 '인류 역사상 최대의 노동 도난'이라 표현했으며, OpenAI 임원들도 퍼블리셔에게 실존적 위협이 된다고 인정한 것으로 드러났습니다. 이는 AI 학습 데이터의 '공정 사용(fair use)' 주장을 약화시킬 수 있는 중요한 증거입니다.

번역된 본문

링크 복사 페이스북 X 왓츠앱 레딧 핀터레스트 플립보드 이메일 이 기사 공유하기 대화 참여하기 뉴스레터 구독하기

뉴욕타임스는 2023년 말 OpenAI와 마이크로소프트를 저작권 침해로 고소했으며, 이 소송은 거의 3년이 지난 지금도 진행 중인 것으로 보인다. 이제 해당 언론사의 법률팀이 피고들의 진술과 문서를 바탕으로 작성된 충격적인 법정 의견서를 제출하며 법원에 즉결 판결(summary judgment)을 요청했다.

404 Media에 따르면, 이 문서들은 양사의 요청으로 여전히 봉인되거나 편집된 상태이며, 공개된 내용은 경영진의 잠재적으로 치명적인 발언들을 보여준다. 여기에는 AI 스크래핑이 인류 역사상 최대 규모의 노동 도난이자 퍼블리셔에게 실존적 위협이라는 주장도 포함되어 있다.

TH Premium으로 더 깊이 들어가기: AI와 데이터센터

  • 데이터센터 냉각 기술의 현재 상황
  • 커스텀 AI ASIC의 현주소
  • 미국의 AI 칩 규정은 계속 바뀌고 있다 — 그리고 나머지 세계가 그 대가를 치르고 있다
  • GTC 2026: 이언 벅 프레스 Q&A 전문 — 하이퍼스케일 및 HPC 부문 부사장이 CPX 보류와 올해 LPU 디코드 출시에 대해 발언하다
  • 데이터센터 CPU 수요가 급증했으며, AI 에이전트가 그 원인이다

법정 의견서는 마이크로소프트 응용과학(Applied Science) 디렉터인 브렌트 헥트(Brent Hecht)의 2023년 1월자 내부 메모를 인용했다. 그는 이 메모에서 "전 세계 수백만 명이 곧 대형 모델이 자신들의 작업을 모두 '빨아들이는' 것을 전례 없는 규모의 놀라운 도난으로 간주하게 될 것"이라고 말했으며, 이를 "인류 역사상 최대 규모의 노동 도난"이라고도 불렀다.

또 다른 마이크로소프트 문서에서는 "콘텐츠를 만든 사람들 중 이런 방식으로 사용되기를 의도한 사람은 거의 없으며, 사용에 대한 보상도 받지 못한다"고 언급했다.

ChatGPT가 2023년 내내 인기를 끌면서, 이 소프트웨어 거인의 자체 데이터에 따르면 Copilot은 빙(Bing) 검색에 비해 뉴욕타임스의 클릭률을 최대 93%까지 떨어뜨린 것으로 나타났다.

응용과학 디렉터의 또 다른 메모는 이를 "파멸의 고리(doom loop)"라고 부르며 "우리 모델의 성능과 웹 전체를 동시에 저해할 것"이라고 말했다. 뉴욕타임스의 의견서는 해당 문서에서 헥트의 발언을 인용했다: "최종 제품이 필수 공급자들의 경제적 기반을 위협하는 것은 매우 이례적인 일이지만, 이것이 우리가 LLM 비즈니스의 '콘텐츠 공급망'에 관해 만들어낸 상황이다."

OpenAI의 ChatGPT 책임자 닉 털리(Nick Turley)는 내부 커뮤니케이션에서 AI 챗봇이 퍼블리셔에게 "실존적 위협"이라고 말했다. 그들은 "대부분 대체재적 성격"을 가지며 "더 좋아질수록 점점 더 대체재가 될 것"이라는 것이다. 또 다른 OpenAI 엔지니어는 "아무리 링크를 눈에 띄게 표시해도 사용자는 클릭하지 않을 것"이라고 증언했다.

OpenAI의 또 다른 연구원인 닉 라이더(Nick Ryder)는 회사 사장 그렉 브록맨(Greg Brockman)에게 "뉴욕타임스 유료벽을 우회하는 핵(hack)"에 대해 알렸고, 브록맨은 "아, 좋네"라고 답했다.

AI 기업들은 모델에 공급할 데이터를 위해 인터넷을 긁어모으는 것이 '공정 사용(fair use)'에 해당한다고 주장하며, 한 법원은 Anthropic의 출판물 사용이 이 범주에 속한다고 판결한 바 있다. 법은 공정 사용을 "비평, 논평, 뉴스 보도, 교육(교실용 다수 복사물 포함), 학문 또는 연구"로 정의한다. 특정 사용이 '공정 사용'에 해당하는지를 판단하는 요소로는 "(1) 사용의 목적과 성격(해당 사용이 상업적 성격인지 또는 비영리 교육 목적인지 포함), (2) 저작물의 성격, (3) 저작물 전체와 관련하여 사용된 부분의 양과 실질성, (4) 저작물의 잠재적 시장이나 가치에 대한 사용의 영향" 등이 있다.

하지만 뉴욕타임스 의견서에서 드러난 이 모든 내용은 이러한 공정 사용 주장에 심각한 의문을 제기할 수 있다.

원문 보기
원문 보기 (영어)
Copy link Facebook X Whatsapp Reddit Pinterest Flipboard Email Share this article 2 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter The New York Times sued OpenAI and Microsoft for copyright infringement in late 2023, with the case apparently still ongoing almost three years later. Now, the publication’s legal team has asked the court for a summary judgment after it filed a revealing legal brief based on statements and documents from the defendants. According to 404 Media , these documents remain sealed or redacted at the request of both companies, with the revelations showing potentially damaging statements from their leadership, including claims AI scraping is the biggest theft of labor in human history and an existential threat to publishers. Go deeper with TH Premium: AI and data centers The data center cooling state of play The custom AI ASIC state of play America’s AI chip rules keep changing — and the rest of the world is paying the price GTC 2026: Ian Buck press Q&A transcript — VP of Hyperscale and HPC speaks out on shelving CPX and shipping LPU decode this year Demand for data center CPUs has surged, and AI agents are responsible The brief cited an internal memo dated January 2023 by Microsoft director of Applied Science Brent Hecht, where he allegedly said, “Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions” and also called it “the largest theft of labor in human history.” Another Microsoft document was cited saying, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.” As ChatGPT surged in popularity throughout 2023, the software giant’s own data revealed that Copilot dropped click-through rates for The New York Times by as much as 93% compared to Bing search. Another memo by the Applied Science director called it a “doom loop” and said it would “hurt the performance of our models and the entire web at the same time.” The NYT brief quoted Hecht from the document, saying, “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’” Latest Videos From Tom's Hardware Watch full video here: OpenAI Head of ChatGPT Nick Turley said in internal communications that the AI chatbot is an “existential threat” to publishers as they are “largely substitutive” and “will get more and more substitutive as they get better,” while another OpenAI engineer testified that “no matter how prominently we show the links, users won’t click.” Nick Ryder, another OpenAI researcher, told company president Greg Brockman about a “hack to get around nytimes paywall,” to which he replied, “ah nice.” AI companies argue that scraping the internet for data to feed to their models is “fair use,” with one court agreeing that Anthropic’s use of published material falls under this category . The law defines this as “criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research.” Some of the factors that determine whether a particular use falls under “fair use” include “(1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work.” However, all these revelations in NYT’s brief could complicate OpenAI’s fair use defense, especially as it shows that the leadership of both companies are aware of the possible market repercussions of AI scraping. Microsoft CEO Satya Nadella said in a deposition from earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training” and that if he “had been made aware that OpenAI has scraped and trained on information that was behind a paywall,” the company would have required OpenAI “to retrain its models.” Follow Tom's Hardware on Google News , or add us as a preferred source , to get our latest news, analysis, & reviews in your feeds. Stay On the Cutting Edge: Get the Tom's Hardware Newsletter Get Tom's Hardware's best news and in-depth reviews, straight to your inbox. Contact me with news and offers from other Future brands Receive email from us on behalf of our trusted partners or sponsors TOPICS See all comments (2) Jowi Morales Contributing Writer Jowi Morales is a tech enthusiast with years of experience working in the industry. He’s been writing with several tech publications since 2021, where he’s been interested in tech hardware and consumer electronics. 2 Comments Comment from the forums This is still going on. I was just reading an article in a British newspaper (The Guardian), and then did a Google search with some of the key concepts discussed in that article. What Google's AI Overview gave me immediately was partly a word-by-word copy of parts of that article. Reply ttquantia said: This is still going on. I was just reading an article in a British newspaper (The Guardian), and then did a Google search with some of the key concepts discussed in that article. What Google's AI Overview gave me immediately was partly a word-by-word copy of parts of that article. Search AI should be putting the descriptions that come up into the context, not baking them into the model. If that's an infringement, so is Web search. Reply View All 2 Comments