메뉴
BL
404 Media 8일 전

AI '모델 붕괴' 막는다…AI 기업들, 옛날 책 사들이기 나서

IMP
8/10
핵심 요약

AI 기업들이 모델 성능 저하(모델 붕괴)를 막기 위해 2022년 이전에 출판된 아날로그 인쇄 책들을 대량으로 매입하고 있습니다. 온라인상에는 이미 AI가 생성한 저품질 데이터(AI 슬롭)가 넘쳐나고 작가들의 방해 공작이 포함될 위험이 크기 때문입니다. 이 과정에서 유통업체들은 철저한 비밀유지 계약(NDA)을 통해 구매자의 신원을 숨기며, 스캔을 위해 수백만 권의 실물 책을 파기하는 과정이 대중에게 알려지는 것을 방어하고 있습니다.

번역된 본문

AI 기업들이 자사 모델을 개선하기 위해 더 많은 학습 데이터를 찾으면서, 한 회사는 AI 기업들이 만들어내는 바로 그 'AI 쓰레기(AI slop)'가 전혀 없다는 점에서 구식 인쇄 책을 이상적인 데이터 소스로 제공하고 있다.

“세계에서 가장 우수한 AI 학습 데이터가 서가 위에 방치되어 있습니다.”라고 세계 최대 도서 데이터베이스를 구축하고 AI 기업을 위한 대규모 도서 매입 서비스를 제공하는 회사인 ISBNdb가 자사 웹사이트에서 밝혔다. “책은 웹 크롤링으로는 결코 복제할 수 없는 방식으로 구조화된, 엄선되고 동료 평가를 거친 특정 분야의 인간 지식을 나타냅니다. 정보 밀도가 높고, 편집 과정을 거쳤으며, 권위가 있습니다.”

ISBNdb는 자사 웹사이트의 한 글에서 2022년 이전에 출판된 인쇄 책에는 AI가 생성한 텍스트가 포함되어 있지 않기 때문에 AI 학습 데이터로 이상적이라고 설명했다. 이 글에서 지적하듯, 오늘날 AI 기업들이 인터넷에서 수집할 수 있는 대부분의 데이터에는 AI가 생성한 텍스트가 포함될 가능성이 높다. 이는 AI가 생성한 데이터로 학습된 모델이 오류에 더 취약한 성능 저하 모델을 초래하는 '모델 붕괴(model collapse)' 현상을 유발할 수 있다.

이 글은 또한 자신의 글이 학습 목적으로 무단 수집되는 것에 반대하는 도서 작가들이 이제 결과물인 AI 모델을 교란하고 방해하기 위해 고안된 글을 생성하여 AI 모델을 쉽게 오염(poison)시킬 수 있다고 지적했다. “LLM(대형 언어 모델) 이전 시대의 인쇄 책은 구조적으로 이러한 오염이 없음이 보장됩니다. 그 자체만으로도 큰 이점입니다 [...] [2022년 이전에] 출판된 실물 책은 현대적 오염 도구가 전혀 없는 구조적으로 깨끗한 상태입니다.”

ISBN은 국제 표준 도서 번호(International Standard Book Number)를 의미하며, 대부분의 책 뒷면에 있는 숫자 상업 식별자이자 바코드이다. 수년간 ISBNdb는 도서 판매자, 도서관 및 유통업체가 재고를 관리하고 책을 찾고 판매하도록 도왔지만, 생성형 AI 붐으로 인해 AI 기업들에게도 귀중한 존재가 되었다. ISBNdb는 이제 도서 메타데이터에 대한 접근 권한을 파는 것 외에도, AI 연구소가 주문당 1,000권에서 최대 100만 권의 인쇄 책을 대량으로 매입할 수 있도록 지원한다. ISBNdb의 데이터를 통해 AI 기업들은 중복을 피하면서 인쇄된 책을 체계적으로 매입, 스캔 및 학습 데이터로 변환할 수 있다.

AI 기업들이 학습 데이터를 위해 인쇄 책을 빨아들이려는 시도는 도서 작가들이 Anthropic을 상대로 제기한 저작권 소송에서 수백만 권의 인쇄 책을 구해 스캔하고 그 과정에서 책을 파기하려는 계획을 상세히 담은 내부 문서가 공개되면서 올해 1월 큰 관심을 받았다. 워싱턴 포스트(Washington Post)의 기사에 따르면 Anthropic은 도서관, 소매업체 및 개인이 책을 판매할 수 있는 여러 마켓플레이스 중 하나인 Better World Books라는 회사에서 책을 구매하고 있었다. 최근 구글 역시 저작권이 있는 책으로 Google Gemini를 학습시켰다는 이유로 도서 출판사들로부터 소송을 당했다.

ISBNdb는 AI 기업의 신원을 비밀로 유지할 수 있다고 광고한다. “모든 거래에 대해 엄격한 NDA(비밀유지계약)를 체결합니다.”라고 ISBNdb의 웹사이트는 명시한다. “모든 프로젝트는 법적 구속력이 있는 비밀유지계약으로 시작됩니다. 고객의 신원, 전략 및 인수 대상은 절대 공개되지 않습니다.”

ISBNdb는 AI 기업들이 스캔 과정에서 인쇄된 책을 파괴하는 모습이 적발되는 것을 원하지 않을 수 있다고 언급했다. “인식 문제는 현실입니다.”라며 ISBNdb 웹사이트는 전했다. “'AI 기업, 200만 권의 책을 파기하다'라는 헤드라인은 대중의 동정심을 얻지 못합니다.”

이러한 마켓플레이스에서 외국 도서 판매를 전문으로 하는 한 전문 서적상은 필자에게 올해 4월부터 자신과 다른 서적상들이 역사적인 매출 급증을 목격했다고 전했다. 이 서적상은 플랫폼에서 계속 비즈니스를 하기 위해 익명을 요구했다.

AI 기업에 학습 데이터용으로 수백 권의 책을 판매했을 것으로 의심되는 이 서적상은 필자에게 이렇게 말했다. “저는 개인적으로 이 모든 상황에 복잡한 감정을 느낍니다. 그렇지 않으면 판매될 가능성이 희박한 오래된 재고를 처리해준다는 점 등 재정적으로 저에게 이익이 됩니다. 저는 해외 및 외국어 도서 재고를 보유하고 있어 이러한 판매에 적합했습니다. 반면에 저는 최종 용도가 마음에 들지 않으며, 그것이... (이하 생략)”

원문 보기
원문 보기 (영어)
As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. “The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site . “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.” In one article on its site , ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models. “Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.” ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication. AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books. ISBNdb advertises that it can keep the identity of AI companies secret. “Strict NDA [non-disclosure agreement] on every engagement,” ISBNdb’s site says. “Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed.” ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. “The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.” One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms. “I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.” This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain. The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were. First of all, he said, the kind of books he sells are specialized and are usually bought by schools and libraries. Purchases from these organizations have been trending downward because of reduced funding, he said. Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. I have not seen any evidence that this bookseller’s recent sales were facilitated by ISBNdb or that the client was an Anthropic or another AI company. “It's not just the quantity, but the weirdness of the orders,” the bookseller told me. “I've had library orders before, and usually they're mostly confined to a single subject or maybe a slightly broader range of subjects. But basically, almost every library in the world has lost their budget. I know all the U.S. college libraries don't buy much anymore. The Australian libraries don't buy much anymore. The type of books [...] there's no rhyme or reason to it. Also, there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.” “Is it just me, or has there been an uptick in the number of AutoBuy orders since the tail end of last year?” one bookseller wrote on the forums for Alibris , another marketplace for selling books, in February. The AutoBuy function allows a customer to flag books they want to automatically purchase once they become available for sale on Alibris. “Any comment on what is happening? Is an AI going to read every single book? Any insight into how the selections are made? They seem to vary quite a bit in condition, format (hardcover and softcover), price and so on.” “We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders,” Mike Feldman, director of client services at Alibris, responded. One bookseller told me that similarly large orders of books were coming through another marketplace called Biblio. Customers can provide Biblio with a spreadsheet of ISBNs they want to purchase and the company takes it from there. In June, a publication in the Netherlands talked to several rare booksellers who reported similar large bulk purchases they assumed were coming from AI companies. It’s hard to say for a fact that the books are being bought for training data and possibly being destroyed by AI companies because ISBNdb and book marketplaces like Biblio and Alibris keep the identity of the buyer hidden. Large bulk purchases of books are first sent to distribution centers where, for example, Alibris checks the quality of the books before sending them off to the client. Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both “high volume destructive and non-destructive book scanning” services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster a