메뉴
BL
404 Media • 56일 전

AI 학습용 종이책 대량 매입 논란… ISBNdb 서비스 중단

IMP
7/10
핵심 요약

책 데이터베이스 기업 ISBNdb가 대량의 종이책을 스캔해 AI 학습 데이터로 판매하려 한다는 단독 보도가 나오자 관련 서비스를 전면 중단했습니다. 이 과정에서 책이 파괴되는 문제와 저작권 침해 우려가 제기되며 윤리적 논란이 불거졌습니다.

번역된 본문

책 데이터베이스 기업인 ISBNdb가 종이책을 조달해 AI 모델 학습용으로 판매한다는 404 Media(404 미디어)의 보도가 나온 후, 이 회사는 관련 웹사이트 페이지를 삭제하고 AI 모델을 학습시킨다는 주장을 철회했습니다. 대신 이를 단지 '시장 수요 테스트'라고 해명했습니다.

404 Media의 보도로부터 9일 뒤인 7월 30일, ISBNdb는 홈페이지에 공지를 추가하고 변경 사항에 대한 업데이트를 게재했습니다. ISBNdb는 다음과 같이 밝혔습니다. "당사 웹사이트의 마케팅 랜딩 페이지와 관련된 최근 보도를 확인했으며, 이로 인해 제기된 우려를 충분히 이해하고 있습니다. 팩트는 다음과 같습니다. ISBNdb는 AI 학습이나 그 외 어떤 목적으로도 책을 구매, 스캔, 판매한 적이 없습니다. 당사는 AI 모델을 훈련시키지 않으며, 그런 적도 없습니다. 해당 페이지는 시장 수요를 타진하기 위한 테스트였으며, 실제로 그러한 서비스가 제공된 적은 없습니다. 해당 페이지는 내려놓은 상태입니다. 당사의 본질은 사람들이 책을 찾을 수 있도록 돕는 것입니다. 20년 이상 ISBNdb는 책 세계의 목차 역할을 해왔습니다. 즉, 서점, 도서관, 독서 앱이 독자와 책을 연결해 주는 기저의 데이터입니다. 책 자체가 아닌 책에 대한 데이터를 다루며, 이 점은 변하지 않았습니다."

ISBNdb는 7월 28일에 "AI 대형 언어 모델(LLM) 데이터셋 요구를 위한 종이책 조달"이라는 랜딩 페이지를 삭제했습니다. 사이트 측은 "이는 수요를 탐색하기 위한 과정의 일부였으며, 해당 방향에서 벗어나기로 결정했습니다. 당사의 주력 서비스인 ISBNdb(책 메타데이터 API) 서비스는 영향을 받지 않고 정상적으로 운영되고 있다"고 덧붙였습니다. ISBNdb는 404 Media의 코멘트 요청에 일절 응답하지 않았습니다.

한편, 현재는 삭제된 ISBNdb 사이트 내의 한 기사에서는 2022년 이전에 출판된 종이책에는 AI가 생성한 텍스트가 포함되어 있지 않아 AI 학습 데이터로 이상적이며, 이를 학습에 활용하면 모델 붕괴(Model Collapse)를 방지할 수 있다고 주장했습니다. 해당 기사는 또한 자신의 작품이 이런 식으로 활용되는 것을 반대할 수도 있는 작가들에게, 그저 AI 모델을 조작하고 방해할 의도를 가지고 글을 쓰면 된다고 제안하기도 했습니다.

404 Media는 중고 서점들과 대화를 나눴고, 최근 매출이 갑자기 급증했다는 온라인 제보들을 확인했습니다. 한 서점주는 404 Media에 다음과 같이 말했습니다. "단순히 주문량만 늘어난 게 아니라 주문이 기묘합니다. 구매하는 책의 종류를 보면 (...) 전혀 이유나 규칙을 알 수 없습니다. 또한 책의 가격에 완전히 무관심합니다. 이런 방식으로 팔린 적이 있는 책들 중에는 [...] 가격이 터무니없이 비싼 것들도 있었습니다. 그들은 돈이 워낙 많기 때문에 가격을 신경 쓰지 않는다는 건데, AI 측의 소행이라는 것을 알려주는 단서입니다."

ISBNdb의 현재 삭제된 페이지 중 하나에는 다음과 같은 내용이 적혀 있었습니다. "중고 시장에서 종이책을 대량으로 구매하는 것은 창작자들이 기존에 얻었을 수익을 박탈하는 것이 아닙니다. 이 책들은 이미 창작자들에 대한 재정적 의무를 다 충족한 것들입니다."

지난 1월, 책 작가들은 Anthropic(앤스로픽)을 상대로 저작권 소송을 제기했으며, 앤스로픽이 수백만 권의 종이책을 구매 및 스캔하면서 원본을 파괴할 계획이라는 내부 문서가 공개되었습니다. 워싱턴 포스트의 조사에 따르면 앤스로픽은 도서관, 소매업체, 개인이 책을 판매하는 여러 마켓플레이스 중 하나인 Better World Books라는 회사에서 책을 구매하고 있는 것으로 드러났습니다.

ISBNdb의 현재 삭제된 마케팅 자료에는 스캔 과정에서 종이책을 파괴하다 적발되는 것이 좋지 않은 인상을 줄 수 있다는 점을 인정하는 내용이 담겨 있었습니다. ISBNdb 사이트는 이렇게 적고 있었습니다. "인식의 문제는 현실입니다. 'AI 기업, 책 200만 권을 파괴하다'라는 기사 제목은 사람들의 동정심을 얻을 수 없습니다."

Emanuel Maiberg(에마누엘 마이베르크)가 이 기사의 취재에 기여했습니다.

저자 소개 Sam Cole(샘 콜)는 인터넷의 외곽 지대에서 성, 성인 산업, 온라인 문화, 인공지능에 대해 글을 씁니다. 그녀는 저서로 《How Sex Changed the Internet and the Internet Changed Sex》가 있습니다. Samantha Cole의 더 많은 기사 보기

원문 보기
원문 보기 (영어)
Following 404 Media’s reporting that book database company ISBNdb claimed to source printed books to then sell to AI companies for AI training, the company deleted the part of its website offering the service and walked back claims that it would train AI models, and instead called it “a test of market interest.” On July 30, nine days after 404 Media’s reporting, ISBNdb added a note to its homepage and an update on its news page about the change. “We've seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised. The facts: ISBNdb has never purchased, scanned, or sold a book — for AI training or anything else,” ISBNdb wrote. “We don't train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We've taken the page down. Our job is helping people find books. For more than two decades, ISBNdb has been the card catalog of the book world — the data behind how bookstores, libraries, and reading apps connect readers with titles. Data about books, not the books themselves. That hasn't changed.” ISBNdb removed the landing page for “Printed Books Sourcing for Your AI LLMs Dataset Needs” on July 28. “It was part of exploring demand, and we've chosen to pivot away from that direction. Our main ISBNdb (book metadata API) services are unaffected and running as usual,” the site says. ISBNdb has not responded to any requests for comment from 404 Media. In one now-removed article on its site , ISBNdb said that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text, and ingesting these books could prevent model collapse. The article also suggested that book authors who might be opposed to this use of their work could just write with manipulating and sabotaging AI models in mind. 404 Media spoke to booksellers and saw reports online detailing a sudden uptick in sales recently. “It's not just the quantity, but the weirdness of the orders,” one bookseller told 404 Media. “The type of books [...] there's no rhyme or reason to it. Also, there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.” “Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received,” one of ISBNdb’s now-removed pages said. “These are books that have already fully discharged their financial obligation to their creators.” In January, book authors filed a copyright lawsuit against Anthropic, revealing internal documents that showed Anthropic planned to obtain and scan millions of printed books and destroy them in the process. An investigation by the Washington Post found that Anthropic was buying books from a company called Better World Books, one of several marketplaces where libraries, retailers, and individuals sell books. In ISBNdb’s now-removed marketing materials, it admitted that getting caught destroying printed books during the scanning process would not be a good look. “The optics problem is real,” ISBNdb’s site said . “‘AI company destroys two million books’ is not a headline that generates sympathy.” Emanuel Maiberg contributed reporting to this story. About the author Sam Cole is writing from the far reaches of the internet, about sexuality, the adult industry, online culture, and AI. She's the author of How Sex Changed the Internet and the Internet Changed Sex. More from Samantha Cole