메뉴
BL
TechCrunch AI • 33일 전

AI 학습에 저작권 도서 사용, 합법일까? 복잡한 문제다

IMP
8/10
핵심 요약

ChatGPT, Claude 등 AI 모델이 수억 권의 책과 문서를 저자 동의 없이 학습해 왔지만, 저작권법은 1976년 이후 개정되지 않아 법적 판단이 공정 이용(fair use) 여부를 놓고 법원마다 엇갈리고 있다. 작년 앤스로픽(Anthropic)에 15억 달러 배상을 명령한 판결도 실제로는 AI 학습 자체를 합법으로 인정하고 불법 다운로드만 처벌한 것으로, AI 기업에 유리한 선례로 해석된다.

번역된 본문

이제쯤이면 알고 계실 겁니다. ChatGPT, Gemini, Claude 등 챗봇을 구동하는 AI 모델은 수억 권의 책, 온라인 기사, 학술 논문 등 인터넷에서 찾을 수 있는 거의 모든 것을 담은, 사실상 무한한 데이터베이스로 학습됩니다. 대부분의 출판 작가들은 자신도 모르는 사이, 동의 없이 자신의 생계를 위협하는 바로 그 AI 도구의 개발에 기여해 왔습니다. 불법처럼 보이죠? 하지만 현실은 그렇게 단순하지 않습니다.

"이 법률 분야와 기술 분야 전반의 문제는 정말 많은 일이 벌어지고 있다는 점입니다"라고 지식재산권, 저작권, 기술 분야 전문 변호사인 캐시 겔리스(Cathy Gellis)가 테크크런치에 말했습니다. "매우 복잡하고, 찬반 양측 모두 지금 벌어지는 일에 대해 강한 감정이 얽혀 있습니다."

작년, 이런 유형의 첫 판결 중 하나로 윌리엄 알서프(William Alsup) 판사는 자사 AI 모델 학습에 작품이 사용된 작가 그룹에게 앤스로픽(Anthropic)이 거액의 15억 달러 저작권 배상금을 지급하라고 명령했습니다. 겉보기에는 작가들에게 유리한 도덕적 승리처럼 보였지만, 알서프 판사는 실제로 앤스로픽의 AI 학습 자체는 합법이라고 판결했습니다. 판사가 앤스로픽에 제재를 가한 것은 불법 온라인 섀도우 라이브러리(불법 공유 사이트)에서 이 책들을 해적판으로 다운로드한 행위 때문이었습니다.

"작가를 지망하는 독자처럼, 앤스로픽의 LLM은 작품을 학습한 것은 그것을 복제하거나 대체하기 위해서가 아니라, 새로운 것을 창조하기 위해 방향을 틀기 위해서였습니다"라고 판사는 쓰면서, LLM이 수조 개의 단어를 흡수하는 방식을 작가의 문학 연구에 비유했습니다.

겔리스는 이 판결이 AI 기업에 더 유리하다고 봅니다. 2028년까지 연간 약 2,000억 달러 매출을 전망하는 기업에게 15억 달러 벌금이 무슨 대수일까요? "판사가 상황을 살펴보고 저작권 작품을 복사하는 것이 아니라 읽는 것과 유사하다고 판단한 것은 AI 학습에 전반적으로 좋은 소식이라고 생각합니다"라고 겔리스는 말했습니다. "저작권법의 핵심은 복제(copying)에 있지, 작품을 사용하거나 경험하거나 소비하거나 읽는 것에 있지 않습니다."

저작권법은 1976년 이후 개정되지 않았습니다. 즉, 판사들은 AI 산업의 미래를 좌우할 수 있는 법적 질문에 직면했을 때 50년 전의 기준을 어떻게 해석할지 스스로 파악해야 합니다.

"지금 모두가 매우 걱정하고 있습니다. 법이 제각각이기 때문이고, 바로 이 질문 때문입니다"라고 JWL International의 지식재산·미디어 부문 선임 변호사이자 창립자인 제이슨 헨더슨(Jason Henderson)이 테크크런치에 말했습니다. "AI 모델이 방대한 자료로 학습되었다는 것을 모두 알고 있지만, 법은 그 질문을 따라잡지 못하고 있습니다."

이러한 질문은 대체로 공정 이용(fair use) 법리에 달려 있습니다. 즉, 저작권 작품의 사용이 법적으로 허용될 만큼 '변형적(transformative)'인가 하는 문제입니다. 공정 이용은 명시적 허락 없이 저작권 자료를 사용할 수 있게 하는 저작권법의 예외 조항으로, 비평, 패러디, 교육 등을 통해 저작권 작품에 대해 논평하고 반복적으로 활용할 수 있는 능력을 보호합니다. 판사들은 무언가가 공정 이용인지 판단할 때 작품의 목적과 성격, 사용량, 시장에 미치는 영향 등 구체적인 요소를 고려합니다.

"저작권은 언제나 시장을 보호하고 키우는 것에 관한 것입니다"라고 헨더슨은 지적했습니다. "법원들은 [AI 사건에서] 판단 논리가 제각각입니다. 경향적으로 승소하는 경우는, 직접 경쟁할 목적으로 남의 재산으로 학습하는 것이라면 법원이 이를 눈살을 찌푸리는 반면, 경쟁하지 않는 일이라면 법원이 괜찮다고 인정할 방법을 찾는 경향이 있습니다."

헨더슨은 미디어·기술 기업 톰슨로이터(Thomson Reuters)가 경쟁하는 AI 기반 법률 플랫폼을 만들기 위해 자사 콘텐츠를 복제한 리서치 기업 로스 인텔리전스(Ross Intelligence)를 고소한 사건을 언급한 것입니다. "로스의 사용은 톰슨로이터와 '추가적 목적이나 다른 성격'을 가지지 못하므로 변형적이지 않습니다."

원문 보기
원문 보기 (영어)
You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right? The reality isn’t that simple. “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.” Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a mammoth $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models. At face value, this seemed like a moral victory favoring authors, but Judge Alsup actually ruled that Anthropic’s AI training was lawful. What Alsup penalized Anthropic for was pirating these books from illegal online shadow libraries. “Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” the judge wrote, comparing the way an LLM ingests trillions of words to a writer’s study of literature. Gellis thinks the ruling is more advantageous for AI companies. What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028? “I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work,” Gellis said. “Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work.” Copyright law hasn’t been updated since 1976, which means that judges have to figure out how to interpret guidelines from 50 years ago when confronting legal questions that have the potential to shape the future of the AI industry. “Everybody is very worried right now because the law is all over the place, and it’s because of this question,” Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.” These questions often hinge on fair use law — namely, whether use of a copyrighted work is "transformative" enough to be considered legally permissible. Fair use is a carve out of copyright law that allows for the use of copyrighted materials without explicit permission, protecting the ability to comment and iterate on copyrighted works through criticism, parody, education, and other means. Judges consider specific factors when deciding if something is fair use, including the purpose and nature of the work, the amount used, and its impact on the market. “Copyright is always about protecting and growing the market,” Henderson noted. “The courts are kind of all over the place in their reasoning [in AI cases]. What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.” Henderson is referencing a case in which the media and technology company Thomson Reuters sued the research firm Ross Intelligence for copying its content in order to build a competing, AI-based legal platform. “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s,” Judge Stephanos Bibas wrote last year. In that case, Judge Bibas decided that it was not fair use to train on Reuters’ content to make a new platform that would directly compete with it. While authors could potentially argue that chatbots are competing with them by using their works to generate new, synthetic books, that argument has not yet prevailed in court. When it comes to the relationship between AI and copyright, Gellis finds it helpful to narrow down what we’re actually talking about – the way we think about copyright in terms of AI training is quite different from how we think about copyrighting AI-generated content. In one case, Thaler v. Perlmutter, the court ruled that if a work is 100% AI-generated, it’s not copyrightable, which opens a whole new can of worms – how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI? “If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” Gellis said. “[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.” Most AI companies are still lodged in pending litigation over these issues, which means that we won't have a definitive solution to these problems any time soon. "What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it'll take later states of litigation to figure out which one will prevail," Gellis said. "But in the meantime, all these decisions are shaping everything that's happening. It would be kind of foolish for the AI companies to ignore them." Topics AI , Exclusive When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Amanda Silberling Senior Writer Amanda Silberling is a senior writer at TechCrunch covering the intersection of technology and culture. She has also written for publications like Polygon, MTV, the Kenyon Review, NPR, and Business Insider. She is the co-host of Wow If True, a podcast about internet culture, with science fiction author Isabel J. Kim. Prior to joining TechCrunch, she worked as a grassroots organizer, museum educator, and film festival coordinator. She holds a B.A. in English from the University of Pennsylvania and served as a Princeton in Asia Fellow in Laos. You can contact or verify outreach from Amanda by emailing amanda@techcrunch.com or via encrypted message at @amanda.100 on Signal. View Bio October 13 - 15 San Francisco In less than 48 hours, your chance to save up to $300 on your tickets will end! REGISTER NOW Most Popular Inherent, founded by DeepMind alumni, says its AI ‘teammate' just outperformed Anthropic and OpenAI at replicating research Anna Heim Michael Polansky is training an AI model on skin that’s still alive Connie Loizos How AI accounting startup Rillet raised $100M and became a unicorn in 48 hours Dominic-Madori Davis Tesla’s solar roof is dead — here’s what went wrong Tim De Chant Oura faces lawsuit accusing it of misleading consumers about sleep-tracking accuracy Aisha Malik Home batteries are suddenly cheap and everywhere. Here’s why. Tim De Chant Cursor capitalizes on GitHub frustration, launches rival hosting platform Lucas Ropek