메뉴
HN
Hacker News • 36일 전

아마존은 '공정 이용'이라 부리, 인터넷 아카이브는 소송당하는 시대

IMP
7/10
핵심 요약

404 Media의 조사에 따르면 아마존이 희귀 도서를 대량 구매해 제본을 잘라내 파괴적으로 스캔한 뒤 AI 학습 데이터로 활용하고 있습니다. 반면 비파괴 방식으로 책을 디지털화하는 인터넷 아카이브는 대기업들의 소송으로 존폐 위기에 처해 있어, 거대 기업과 공공 이익 단체에 대한 상반된 대우가 도마에 올랐습니다.

번역된 본문

2026년 8월 20일 목요일. 아마존이 이런 행위를 '공정 이용(fair use)'이라고 부를 수 있는 반면, 거대 기업들은 인터넷 아카이브를 고소해 파산으로 몰아가려 하는 것이야말로 요즘 시대의 단면이다. 이 404 Media 보도는 당연히 큰 주목을 받고 있다.

아마존은 책을 대량으로 구매해 AI 학습 데이터로 쓰기 위해 스캔하고, 그 과정에서 책을 파괴하고 있다. 404 Media의 조사는 AI 기업이 학습 데이터용으로 구매할 것으로 예상되는 희귀 도서에 추적 장치를 싣고, 전국을 따라가 최종 목적지에 도달하기까지의 경로를 추적함으로써 지금까지 알려지지 않았던 아마존의 도서 구매 작전을 폭로했다. 그 최종 목적지는 네바다주 라스베이거스에 있는 아마존 창고였다.

이 시설에서 일하는 아마존 직원들에 따르면, 그들이 하는 일이라곤 대량으로 들어오는 인쇄 도서를 받아 더 빨리 스캔하기 위해 제본을 잘라내는 것뿐이다. 이 과정에서 인쇄된 책은 파괴된다. VGT3라 불리는 이 창고에서 일하는 아마존 팀의 로고는 이빨을 드러내고 손에 책을 든 공룡이다.

… 우리는 추적한 화물에 포함된 책의 제목은 공개하지 않지만, 그것들은 희귀본, 즉 유통되는 복사본이 많지 않은 책들이다. 그 이유는 처음부터 많이 인쇄되지 않았거나, 많은 사람이 사용하지 않는 외국어 책이기 때문이기도 하다. 이 책을 판 서점주는 이렇게 말했다. 《올리버 트위스트》 초판처럼 세상 사람들이 관심을 가질 만한 책은 아니지만, 그렇다고 가치가 없는 것은 아니라고.

서점주는 말했다. "가치에는 여러 종류가 있습니다. 당연히 금전적 가치도 있지만, 그 외에도 많은 가치가 있죠. 역사적 가치, 지적 가치, 감상적 가치 같은 것들이요. 그런 것들은 전부 AI 기업들은 신경 쓰지 않습니다. 그들에게는 그저 줄지어 나열된 단어 덩어리로서의 콘텐츠만이 중요할 뿐이죠."

다른 곳에서 언급했듯이, 이것은 대규모 지식재산권 도난이기도 하다. 하지만 아직 제대로 조명받지 못한 측면이 하나 있다. 진정으로 윤리적이고 공공의 이익을 추구하는 기업은 같은 문제를 어떻게 다루는지가 바로 그것이다.

"모든 책을 스캔하다: 인터넷 아카이브의 스크라이브 작업" — 안 로르 프레앙(Anne-Laure Freant)

1996년, 컴퓨터 엔지니어 브루스터 케일(Brewster Kahle)은 '모든 지식에 대한 보편적 접근'이라는 사명을 가지고 인터넷 아카이브를 설립했다. 오늘날 이 비전은 스캔 센터에서 이루어지는 꼼꼼한 작업을 이끈다. 작업자들은 수백 년 된 책의 내용과 물리적 온전성을 모두 보존하며 한 페이지씩 조심스럽게 디지털화한다.

인터넷 아카이브의 접근 방식은 2000년대 초반에 등장한 디지털화 방식에 대한 근본적인 반발에서 비롯되었다. 구글이 2004년 도서 프로젝트를 시작했을 때, 디지털 도서관의 규모에 혁명을 일으켰지만 '속도 대 보존'이라는 문제 있는 트레이드오프를 가져왔다. 구글의 산업적 방식은 종종 파괴적 스캔을 수반했다. 빠른 자동화 처리를 위해 책 등줄기를 자르고 제본을 해체하는 것이다.

인터넷 아카이브는 다른 길을 선택했다. "인터넷 아카이브에서는 절대 제본을 잘라 책을 파괴하지 않습니다. 대신 어려운 길을 택해 한 페이지씩 디지털화합니다." 이로 인해 인터넷 아카이브의 매우 특수한 목적에 맞춰 기계와 소프트웨어를 개조하게 되었고, '북 스캐너' 혹은 '스크라이브(scribe) 운영자'라는 직무가 탄생했다.

분명히 말해두자면, 파괴적 방식이 더 빠르고 저렴하다. 하지만 아마존이 보유한 막대한 자원과, AI 버블을 부풀리는 데 쏟아부어지고 많은 경우 낭비된 것으로 입증된 엄청난 자금을 생각하면 그다지 변명이 되지 못한다. 이 모든 것이 아마존의 형편없는 Nova 시리즈의 최신 모델을 학습시키기 위한 것이라는 점을 떠올리면 더더욱 그렇다. 이는 그들이 내린 선택이었다. 마치 원자력과 재생에너지 발전 용량을 키울 시간을 들이는 대신 지금 당장 대규모 화석연료 발전소를 짓기로 한 것과 같은 선택이고, 무수한 작가들의 지식재산을 훔치기로 한 것과 같은 선택이었다.

원문 보기
원문 보기 (영어)
Thursday, August 20, 2026 It is a sign of the times that Amazon gets to call this fair use while huge corporations try to sue the Internet Archive out of business. This 404 report justly been getting considerable coverage. Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process. A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination. That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands. ... We're not revealing the titles of the books included in the shipment we tracked, but they are rare, meaning there are not many copies of them in circulation. Sometimes that's because not many copies of them were ever printed, and sometimes because they are in a foreign language not many people speak. As the bookseller who sold them told me, there are not many people in the world who would care about them in the same way people might care about the first edition of Oliver Twist , but that doesn't mean they're not valuable. "There are different types of value," the bookseller said. "There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together." As mentioned elsewhere, this is also IP theft on a massive scale. There is, however, one aspect which hasn't gotten to play it deserves, namely how a genuinely ethical and public spirited company handles the same problem . Scanning all the Books: The Work of Scribes for the Internet Archive Anne-Laure Freant In 1996, computer engineer Brewster Kahle founded the Internet Archive with the mission to provide "universal access to all knowledge." Today, that vision drives the methodical work happening in scanning centers where operators carefully digitize books one page at a time, preserving both the content and the physical integrity of centuries-old volumes. The Internet Archive's approach stems from a fundamental disagreement with the digitization methods that emerged in the early 2000s. When Google launched its Books project in 2004, it revolutionized the scale of digital libraries but introduced a troubling trade-off: speed versus preservation . Google's industrial approach often involved destructive scanning—cutting book spines and dismantling bindings to facilitate rapid automated processing. The Internet Archive chose a different path. "At the Internet Archive, we never destroy a book by cutting off its binding. Instead, we digitize it the hard way, one page at a time". That led to the adaptation of machines and software to fit the very specific purpose of the Internet Archive, and to a job: book scanner, or scribe operator . Just to be clear, the destructive method is faster and cheaper but given the tremendous resources of Amazon and the spectacular amount of money that has been spent and in many cases demonstrably wasted pumping up the AI bubble, that's not much of an excuse. Arguably even worse when you remember this is all going to train the latest of Amazon's crappy Nova series. This was a choice they made, just like setting up massive fossil fuel power plants now rather than taking the time to increase nuclear and renewable capacity was a choice, just like stealing the intellectual property of countless writers and artists was a choice, just like rolling out products that weren't ready for prime time was a choice, just like setting up ridiculously at complex and deceptive financing schemes rather than growing the industry in a sustainable way was a choice. Posted by Mark at 7:30 AM No comments: Post a Comment Older Post Home Subscribe to: Post Comments (Atom)