메뉴
HN
Hacker News • 39일 전

에어태그가 밝혀낸 아마존의 희귀서적 파괴 — AI 학습용 스캔 실태

IMP
7/10
핵심 요약

희귀서적에 숨긴 AirTag 추적 결과, 아마존이 AI 모델 학습을 위해 대량으로 구매한 희귀서적을 스캔한 뒤 파괴하는 것으로 확인됐습니다. 라스베이거스 시설(VGT3)에서 책등을 뜯어 페이지를 스캔하는 전담팀이 운영되고 있으며, ISBN/바코드를 먼저 스캔하도록 훈련받은 점에서 AI 기업들이 ISBN 목록을 따라 전 세계 출간 도서를 체계적으로 스캔하려 한다는 서점 주인들의 추측이 사실로 확인되었습니다.

번역된 본문

최근 1년여 동안 책 판매업자들은 AI 기업들이 희귀서적을 대량으로 구매한 뒤 AI 학습용으로 스캔하고 파기한다고 의심해 왔습니다. 하지만 이를 입증하기는 어려웠는데, 404 Media가 희귀서적 안에 숨긴 AirTag를 통해 최소한 한 기술 대기업, 즉 아마존이 이러한 대량 주문의 배후에 있음을 밝혀냈습니다. 404 Media는 월요일, 대량 주문된 희귀서적에 AirTag를 심는 데 동의한 서점업자와 접촉했다고 밝혔습니다. 이 AirTag는 라스베이거스에 있는 아마존 AI 학습 시설까지 추적되었으며, 이곳에는 책을 책등에서 뜯어내 페이지를 스캔하는 전담팀이 있었습니다. 파괴적 스캔에 대한 비판이 커지는 가운데도 무신경하게도, 해당 팀 창고 VGT3의 문에는 공룡 티라노사우루스가 책을 삼키려는 로고가 걸려 있었습니다. 아마존의 해명: 아마존은 404 Media의 조사 결과에 대한 코멘트를 거부하며, AI 학습을 구체적으로 언급하지 않은 성명만을 제공했습니다. "아마존은 고객이 사용하는 제품과 서비스를 개발·개선하기 위해 상업적 경로를 통해 도서를 구매한다"는 것이었습니다. 그러나 아마존은 구글, OpenAI, Anthropic 같은 선도 기업들과 경쟁하기 위해 방대한 독자적 학습 데이터가 필요한 '프론티어 AI 모델'을 개발 중입니다. 현재 기업들은 경쟁 우위를 잃지 않으려고 학습 데이터를 철저히 보호하고 있습니다. 구하기 어려운 희귀서적의 텍스트로 모델을 학습시키는 것은 아마존에 분명 이점이 될 것입니다. 특히 Anthropic과 xAI 같은 경쟁사들은 희귀·고서를 학습에 사용하지 않는다고 공개적으로 밝힌 바 있습니다. 또한 아마존은 새로운 원문 텍스트 원천이 필요했던 것으로 보입니다. 404 Media는 온라인 포럼에서 VGT3 노동자들이 올해 초 아마존이 스캔할 책이 부족해졌다고 언급한 논의를 발견했습니다. 부족이 심각해져 더 이상 스캔할 책을 구하지 못하면 창고가 문을 닫을 수 있다는 우려까지 나왔으며, 한때는 공급이 완전히 끊기기도 했다고 노동자들은 말했습니다. 하지만 이 시설은 여전히 가동 중이며, 그곳으로 배달된 희귀서적 대량 주문이 이 팀에 의해 체계적으로 파괴되는 것이 확인되었습니다. 많은 애서가들은 (대안이 있는데도) 파괴적 책 스캔에 경악하고 있습니다. 대표적 비판자인 마이클 버리는 이 관행을 '악 그 자체'라고까지 불렀다고 알려졌습니다. 반면 아마존 노동자들은 포럼에서 유연한 근무시간의 단조로운 일자리를 원하는 이들에게는 '좋은' 기회라고 언급했습니다. 희귀서적을 둘러싼 논쟁: 서점업자들이 소중히 여기는 희귀 작품들이 아마존의 AI 기계에 의해 찢겨 삼켜지고 있다는 사실 외에도, 404 Media의 조사는 AI 기업들이 특정 희귀서적만 주문하는 이유에 대한 또 다른 추측을 뒷받침했습니다. 역사적 판매 실적을 기록한 뒤, 서점업자들은 AI 기업들이 고유한 저작물을 최대한 많이 학습 데이터에 담기 위해 ISBN 번호가 있는 책들을 표적으로 삼는다고 의심해 왔습니다. 404 Media가 아마존 노동자들의 온라인 논의를 검토한 결과, 그들은 책을 스캔하기 전에 바코드나 ISBN을 먼저 스캔하도록 훈련받은 것으로 나타났습니다. 이러한 관행은 "AI 기업들이 ISBN 목록을 따라 전 세계 모든 인쇄된 책을 체계적으로 스캔하려 한다"는 서점업자들의 이론에 "추가적인 신빙성을 부여"한다고 404 Media는 보도했습니다. 서점업자들에게는 돈벌이가 될지 몰라도, 정성껏 수집한 컬렉션이 아마존 같은 파괴적 책 스캔行태로 향할지 모른다는 위험은 윤리적 딜레마를 제기합니다. 서점업자들은 다양한 희귀서적의 가치를 평가하는 방법을 알고 있지만, AI 기업들은 저렴하고 고유한 텍스트를 찾는 과정에서 그 단계를 생략하고 있는 듯합니다.

원문 보기
원문 보기 (영어)
Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon. On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that was part of a bulk order. That Airtag was then tracked to an Amazon AI training facility in Las Vegas that housed a team focused on tearing books from their spines and scanning pages, 404 Media reported. Apparently tone-deaf to the escalating backlash over destructive book scanning, a logo on the door of that team’s warehouse, VGT3, showed a Tyrannosaurus rex preparing to devour a book, 404 Media documented. Amazon deflects Amazon declined to comment on 404 Media’s findings, only providing Ars with the same statement it gave to 404 Media, which does not mention AI training specifically. “Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” Amazon’s statement said. However, Amazon is developing what it considers frontier AI models, which require a massive amount of unique training data to stay competitive with leading firms like Google, OpenAI, or Anthropic. Right now, firms carefully guard their training data to avoid losing an edge. And training models on text from rare books that are difficult to find would seemingly offer an advantage for Amazon, especially since rivals like Anthropic and xAI have publicly stated that they are not training on rare or antique books. Further, it seems that Amazon needed a new source of original text. 404 Media flagged discussions in online forums where VGT3 workers suggested that earlier this year Amazon had run low on books to scan. The shortage was so alarming that they worried the warehouse might shut down if Amazon couldn’t find more books to scan. At one point, the supply completely ran out, workers said. But the facility is still operational, 404 Media reported. And it’s now confirmed that bulk orders of rare books delivered there are systematically destroyed by this crew. Many book lovers are horrified by destructive book-scanning (since there is an alternative) , with one staunch critic, Michael Burry, even reportedly labeling the practice to be “evil incarnate.” However, Amazon workers reported in forums that they consider the gig to be a “nice” opportunity for those drawn to a humdrum job with flexible hours. Debate rages over rare books On top of revealing that rare works that booksellers value are getting chewed up and swallowed by Amazon’s AI machine, 404 Media suggested that its investigation helped firm up another bookseller theory about why AI firms might be ordering certain rare books and not others. After a reportedly historic year of sales, booksellers had suspected that AI firms were targeting books with ISBN numbers in order to ensure that the highest volume of unique works were present in training data sets. And 404 Media’s review of Amazon workers’ online discussions indicated that they were trained to scan barcodes or ISBNs before scanning books. That practice, 404 Media reported, “gives further credence” to booksellers’ theory that “AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs.” For booksellers, the money may be good, but the risk that their carefully sourced collections will be destined for destructive book scanning like Amazon’s raises an ethical dilemma. They know how to assess a wide range of rare books to determine their value, and AI firms seem to be skipping that step in hunting low-cost, unique ISBNs to complete their checklists. Right now, the books that AI firms are apparently buying up aren’t necessarily the kind of prized first editions of celebrated works that are typically valued quite highly. Instead, AI firms often target older books with lower monetary value, such as books that were never translated from a foreign language that’s not widely used today or books that were never popular enough to be widely distributed. However, these works may still have “historical value, intellectual value, sentimental value” that AI firms overlooked, the bookseller who planted the Airtag told 404 Media. A rare book’s value can be derived from “all sorts of things” that “the AI companies don’t care about. They just want the content as a bunch of words strung together.” Redditors weigh in On Reddit, some book fans debated whether it was that problematic that companies are destroying rare books to train AI, especially since, as the BBC reported , some of these books have been sitting on booksellers’ shelves for decades gathering dust. “You were not going to buy that old paper book,” one Redditor commented in response to a post lamenting that “an obscure book from 1700 is now a museum piece and may reveal day-to-day stuff that we didn’t know.” In that thread, the original poster said that the real problem was that tiny details and even major historical insights that can be gleaned from reviewing rare books will be lost to AI greed. AI firms will “never share the contents” of books they scan, the poster said, “as they don’t want anyone else to be able to train their AI” on the same works. “This sounds like propaganda from the AI haters,” another Redditor pushed back, but a subsequent commenter shared similar fears. Although people might clash with someone who argues that training AI on these works will make all the knowledge that a work contains more accessible online, the commenter suggested instead that AI models would only make available “warped, censored, and paywalled fragments of ideas” from forgotten works. “Even if you love the technology, you can admit that the concept of an AI literally eating books to become more powerful is pretty dystopian,” someone else on the thread said. Some booksellers agree that not every text needs to be saved, the BBC reported. However, Scottish bookseller Derek Walker told the BBC that AI firms should be striving to distinguish between works that won’t be missed much—such as little-known academic texts that may technically be rare—and lesser-known antique works that may be “the only known surviving example of an edition.” “It would be a much more significant problem if one like that were to be bought for destruction, having survived this long,” Walker said. For AI firms guarding their training data and using services that mask their identities as buyers, there’s likely little desire to discuss their bulk buying directly with booksellers. Allowing booksellers to weigh in on the works they plan to feed into their models would require a level of transparency that may seem riskier than a possible reputation hit if it’s ever proven that a treasured first edition was destroyed in the name of advancing AI. Ashley Belanger Senior Policy Reporter Ashley Belanger Senior Policy Reporter Ashley is a senior policy reporter for Ars Technica, dedicated to tracking social impacts of emerging policies and new technologies. She is a Chicago-based journalist with 20 years of experience. 109 Comments