메뉴
HN
Hacker News • 56일 전

AI 학습을 위해 희귀 도서들이 대량 파괴되고 있다

IMP
8/10
핵심 요약

AI 기업들이 고품질 학습 데이터 확보를 위해 희귀 및 절판 도서를 대량으로 매입해 파괴 스캐닝(destructive scanning)한 뒤 폐기하고 있습니다. 2022년 이전 출간된 책들은 'AI 오염'이 없어 프리미엄 데이터로 취급되며, 특히 Anthropic은 이 작업을 비밀리에 산업 규모로 진행한 것으로 확인되었습니다.

번역된 본문

알렉산드리아 도서관이 불타다: AI 기업들이 책을 대량으로 파괴하고 있습니다.

AI 기업들은 수백만 권의 책을 스캔하고 파괴하기 위해 대가를 지불하고 있으며, 이 사실을 대중이 알지 않기를 원합니다. (이미지: Anonymous/Derek McArthur) 이 기사는 자매지인 USA Today와의 독점 구독자 파트너십을 통해 제공되었으며, 미국 동료들이 작성했습니다. 이 글이 The Herald의 견해를 반드시 반영하는 것은 아닙니다.

이미 공개된 웹의 표면을 모두 긁어모으고 불법으로 확보된 라이브러리를 남김없이 탐닉한 인공지능 산업이 새로운 식욕을 드러냈습니다. 현재 수백만 권의 실물 책이 대량으로 구매되어 업계가 불안해하며 부르는 '파괴적 스캐닝(destructive scanning)' 과정에 투입되고 있습니다.

데이터 확보 과정은 점차 수백만 권의 책 등에서 본드(철사)를 풀어내고, 스캔한 뒤 종이 절단기로 세절하여 폐기하는 방식으로 이루어지고 있습니다. 이 서비스는 ISBNdb라는 회사를 통해 제공됩니다. 이들은 "주문당 최대 100만 권의 도서를 확보"해 줄 것을 약속합니다. 특히 2022년 이전에 출간된 책들은 AI가 생성한 텍스트에 오염되지 않았다는 '구조적 보장'이 있기 때문에, AI 회사들은 "오래되고 희귀하며 전문적인 서적"을 프리미엄 학습 데이터로 크게 홍보하고 있습니다. AI가 자기 자신을 학습하는 데리리어스(delirious, 혼란)한 루프에 빠지면서, 기존 세계의 정보는 대형 언어 모델(LLM)의 발전과 유지에 더욱 귀중한 자원이 되었습니다.

(필자 Derek McArthur의 AI와 예술에 대한 추가 기사 목록)

  • 저명한 문학상이 AI가 작성한 작품에 상을 수여한 혐의를 받고 있다.
  • 모든 예술가를 대체할 것인가? AI 예술 옹호자들의 기묘한 논리.
  • 한 래퍼가 자신의 새로운 바이럴 히트곡이 노골적인 AI 창작물이 아니라고 거짓말을 하고 있다.
  • OpenAI의 소라(Sora)가 할리우드를 파괴할 것이었다... 갑자기 추락하기 전까지는.
  • 하이랜드 소(Highland cow) AI 아트 판매를 금지하는 법을 통과시킬 수 있을까?
  • 넌센스한 AI 생성 벽화는 글래스고(Glasgow)가 원하는 마지막 것이다.
  • 이 베테랑 영화인은 AI에게 아이디어를 달라고 했을 때 충격을 받았다.

이 회사의 마케팅 자료는 책이 겪게 될 현실에 대해 매우 직설적입니다. 웹사이트에는 "'AI 기업이 2백만 권의 책을 파괴했다'는 헤드라인은 동정심을 유발하지 못한다"고 명시되어 있습니다. 이들은 구매자들에게 비밀유지계약(NDA)을 제공하며, 이 작업을 '디지털 보존(digital preservation)'으로 포장하라고 조언합니다.

우리가 말하는 책들은 역사의 수많은 재앙 속에서도 살아남았고, 정보화 시대 이전의 어둠 속에서 희귀해진 책들입니다. 유럽과 미국 전역의 소규모 서점상들은 희귀 및 절판 도서에 대한 대량 주문이 갑자기 급증한 것을 눈치챘습니다. 한 서점상은 404 미디어(404 Media)에 주당 약 20권의 책을 팔던 것에서 구하기 힘든 외국어 서적을 포함해 수백 권을 판매하는 것으로 변했다고 전했습니다.

그는 "이것은 나에게 재정적 이익을 줄 뿐만 아니라, 그렇지 않으면 판매될 가능성이 희박한 오래된 재고를 처리해 준다"고 말했습니다. "반면에, 나는 그것이 어떻게 최종 사용되는지 마음에 들지 않으며, 희귀한 책들이 종이 펄프로 만들어지는 것을 좋아하지 않습니다."

이 책들은 대체할 수 없는 것들입니다. 디지털화되지 않은 실물 책의 마지막 사본이 세절되어 파괴되면, 그것은 영원히 사라지는 것입니다. 그리고 디지털화된 버전은 저작물의 정확하고 직접적인 표현이라기보다는 AI 해석에 휘둘리는 신세가 됩니다. (이미지: 옥스퍼드 대학교의 희귀 도서 및 원고 컬렉션 / 게티이미지)

챗봇 클로드(Claude)를 개발한 기업인 앤스로픽(Anthropic)은 이러한 관행을 견인해 온 주역이었습니다. 저작권 소송 과정에서 공개된 내부 문서에 따르면 2024년 '프로젝트 파나마(Project Panama)'라는 프로그램이 시작된 것으로 드러났습니다. 법원 제출 서류는 '유압식 절단 기계'를 사용해 수백만 권의 책 등을 잘라낸 뒤, 이를 '고속, 고품질, 산업용 수준의 스캐너'에 통과시키는 산업 규모의 작전을 묘사했습니다. 내부 메모는 이를 "전 세계 모든 책을 파괴적으로 스캔(disruptively scan)하려는 우리의 노력"이라고 불렀으며, "우리가 이 프로젝트를 추구하고 있다는 사실이 알려지는 것을 원하지 않는다"고 덧붙였습니다. 앤스로픽은 이를 위해 구글 북스(Google Books)의 전 총괄 책임자를 고용했습니

원문 보기
원문 보기 (영어)
Derek McArthur Library of Alexandria burns: AI companies are destroying books in bulk AI (Artificial Intelligence) Artificial intelligence Entertainment Literature By Derek McArthur Share 0 Comments AI companies are paying for millions of books to be scanned and destroyed - and they'd rather you didn't know about it (Image: Anonymous/Derek McArthur) This article is brought to you by our exclusive subscriber partnership with our sister title USA Today, and has been written by our American colleagues. It does not necessarily reflect the view of The Herald. The artificial intelligence industry, having already stripped the open web of its bark and devoured every morsel of its illegally acquired libraries, has developed a new appetite. Millions of physical books are now being bought in bulk and fed into a process that the industry nervously calls “destructive scanning”. Data acquisition increasingly involves cutting the spines from millions of books, scanning them and then shredding and discarding them afterward. The service comes courtesy of a company called ISBNdb. They promise to obtain “up to one million titles per order.” They specifically tout “older, rare and specialist volumes” as premium training data because books from the pre-2022 era are “structurally guaranteed” to be free of AI contamination. As AI starts to feed on itself into a delirious loop, information from the old world becomes more valuable to the momentum and continuation of large language models. Read more about AI and the arts from Derek McArthur: A prestigious literary competition is being accused of rewarding AI-written works Replace all artists? The strange logic of the AI art defenders A rapper is pretending his new viral hit isn't a blatant AI creation OpenAI’s Sora was going to destroy Hollywood... until it came abruptly crashing down Can we pass a law banning the sale of Highland cow AI art? A nonsensical AI-generated mural is the last thing Glasgow ever needs This film veteran was stunned when he asked AI to give him ideas The company's marketing material is blunt about the reality of what happens to the books. "'AI company destroys two million books' is not a headline that generates sympathy," its website states. It offers non-disclosure agreements to buyers and advises them to frame the work as "digital preservation". We are talking about books that have survived through the many scourges of history and have become rare due to their life as an obscurity before the information age. Small booksellers across Europe and the United States have noticed a sudden surge in bulk orders for rare and out-of-print titles. One bookseller told 404 Media that he went from selling about 20 books a week to hundreds, including hard-to-find foreign language works. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell," he said. "On the other hand, I don't like the end-use, and I don't like that uncommon books are being pulped." Okay, and on the other hand, these things cannot be replaced. Once a final non-digitised copy of a book is shredded, it’s gone forever. And the version that is digitised becomes at the mercy of AI interpretation rather than a straight, accurate representation of the work. A collection of rare books and manuscripts at Oxford University (Image: Getty) Anthropic, the company behind the chatbot Claude, has been a major driver of this practice. Internal documents made public during a copyright lawsuit revealed a program called “Project Panama” launched in 2024. Court filings described an industrial-scale operation using a "hydraulic powered cutting machine" to slice the spines off millions of books, which were then fed through "high speed, high quality, production level scanners". An internal memo called it "our effort to disruptively scan all the world's books" and noted, "we do not want it to be known that we are pursuing this project." Anthropic hired the former head of Google Books specifically to help acquire what the company described as "all the books in the world". The legal framework enabling this comes from a federal judge's ruling in the Anthropic copyright case. The court determined that by making a digital copy and destroying the physical original in a one-for-one transfer, the process qualifies as "transformative" and falls under fair use protections. The reasoning is that since only one copy exists at a time, it's not traditional piracy. This ruling sets a precedent that allows the practice to wildly accelerate. Read more from Derek McArthur: Does the Mack need to be faithfully rebuilt or are we just scared of our own future? Movie studios are circling like sharks to buy the biggest online review platform This year’s Cannes came and went. Is cinema in danger of becoming a niche hobby? Swapping the boss won't put back together a broken Creative Scotland Rare booksellers across multiple countries are now receiving similar suspicious purchase requests. In the Netherlands, an antiquarian bookseller received an email from a Singapore-based company claiming to be working on "a new project focused on collecting books in multiple languages". The attached list contained 3,000 English-language titles organized by ISBN number, ranging from folklore and fairytales to technical manuals and science texts. Similar reports have emerged from Switzerland, Spain, and Germany. The booksellers involved concluded the books were destined for AI training because no collector or reseller would need such a random assortment of obscure titles. So who will be the knight in shining armour and stop such a senseless and wasteful practice? Elon Musk stated on his social media platform X that he instructed his own AI team "to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning". Thanks for that, Elon, but it’s hard to take the moral right of someone seriously when their own AI model generated child sexual abuse material. Destroying books is very efficient for AI companies and, according to the courts, legal. But the cost is a permanent loss that carries historical and cultural significance beyond words printed on the page. The books that survive this process will live the rest of days as mindless data points in a model, their physical forms long turned to pulp. Unless the legal framework changes or public pressure forces companies to preserve the originals, the Library of Alexandria will burn yet again. Derek McArthur is an artist and arts writer specialising in cinema and culture. He writes a weekly arts column for The Herald. AI (Artificial Intelligence) Artificial intelligence Entertainment Literature Share 0 Comments Get involved with the news Send your news & photos