메뉴
HN
Hacker News • 36일 전

AI 기업들이 실물 책을 파괴한다 — 희귀본을 늦기 전에 스캔하자

IMP
7/10
핵심 요약

AI 기업들이 중개업체를 통해 중고 서적을 대량 매입해 스캔한 뒤 훈련 데이터로 활용하고 물리적으로 파괴하는 것이 드러났으며, Anthropic의 '프로젝트 파나마'는 15억 달러 저작권 합의로 폭로되었습니다. 이 과정에서 인류의 지식이 사기업 서버에 영구적으로 독점되고 공공 영역에서 사라지는 심각한 문제가 발생하고 있습니다. 안나스 아카이브(Anna's Archive)는 전 세계 자원봉사자들이 이 문화유산이 사라지기 전에 책을 스캔·업로드할 것을 긴급히 촉구하고 있습니다.

번역된 본문

안나스 아카이브(Anna's Archive) 블로그 업데이트 — 인류 역사상 가장 큰 진정한 개방형 도서관입니다.

AI 기업들이 실물 책을 파괴한다 — 희귀본을 늦기 전에 스캔하자 (annas-archive.gl/blog, 2026-08-05)

안나스 아카이브 자원봉사자 'u'의 게스트 포스트(중국어에서 번역됨).

요약(TL;DR): AI 기업들이 수백만 권의 실물 책을 비밀리에 매입하고, 스캔한 뒤 파괴하여 모델 훈련에 사용하고 있습니다. 이로 인해 인류의 지식이 사기업 서버에 영구적으로 갇히게 됩니다. 안나스 아카이브는 이 문화유산이 영원히 사라지기 전에 전 세계 자원봉사자들이 책을 스캔하고 업로드할 것을 긴급히 촉구합니다.

여러 AI 기업들이 중개업체를 통해 대량의 중고 서적을 확보하고, 스캔한 뒤 파괴하고 있습니다. 이 모든 것은 2022년 이전의 '기계에 오염되지 않은' 훈련 데이터를 얻기 위해서입니다. Anthropic의 '프로젝트 파나마(Project Panama)'는 15억 달러 규모의 저작권 합의를 통해 폭로되었습니다. 2024년 초, 이들은 이 극비 프로젝트를 시작했으며, 수천만 달러를 들여 수백만 권의 종이책을 구매하고, 스캔해서 Claude LLM을 훈련시킨 다음, 전부 파괴했습니다. 분노할 일은 이것이 법적으로 허용된다는 점입니다. 하지만 윤리적으로 이는 인류에 대한 극히 심각한 범죄입니다.

그렇다면 왜 실물 책을 파괴하는 걸까요? 그 배경에는 AI 경쟁과 자본의 이익이 있습니다:

  • 경쟁사가 이 책들을 스캔하여 훈련에 사용하는 것을 막기 위해서입니다.
  • 법적 리스크를 회피하기 위해서입니다.
  • 책을 파괴하는 것이 무손실 스캔보다 저렴하기 때문입니다.

AI 기업들이 대량으로 실물 책을 스캔하고 파괴한 후에는, 디지털 사본을 가진 세계 유일의 주체가 됩니다. 지식은 사유 서버에 영구적으로 독점됩니다.

이 오래된 책들을 둘러싼 전쟁은 하나의 역설을 드러냅니다. AI 기업들은 '인류의 지식을 접근 가능하게 만들겠다'고 약속하면서도, 동시에 인류 지식의 가장 견고한 물리적载体인 책을 해체하고 있다는 것입니다. 대중은 더 지능적인 AI 비서를 얻을지 모르지만, 그 대가로 방대한 양의 지식 자원이 공공 영역에서 사라지게 됩니다.

섀도우 라이브러리 세계 최대의 섀도우 라이브러리로서, 안나스 아카이브는 AI 기업의 실물 책 파괴에 맞서기 위한 계획이 필요합니다. 결국 섀도우 라이브러리의 등장은 21세기 지식 공유의 가장 위대한 기적입니다. 우리는 다른 섀도우 라이브러리들과 함께 꺼지지 않는 인류의 빛인 '디지털 알렉산드리아 도서관'을 건설하고 있습니다. 자료(책, 학술지 논문, 신문, 잡지, 고서적, 희귀본 등)를 스캔하기 위해 전 세계 자원봉사자들의 도움이 필요합니다.

원문 보기
원문 보기 (영어)
Anna’s Blog ☾ ☀ 🌐 Language ar - العربية - Arabic ast - asturianu - Asturian az - azərbaycan - Azerbaijani be - беларуская - Belarusian bg - български - Bulgarian bn - বাংলা - Bangla br - Brasil: português - Portuguese (Brazil) ca - català - Catalan ckb - کوردیی ناوەندی - Central Kurdish cs - čeština - Czech da - dansk - Danish de - Deutsch - German el - Ελληνικά - Greek en - English eo - Esperanto es - español - Spanish et - eesti - Estonian fa - فارسی - Persian fi - suomi - Finnish fil - Filipino fr - français - French gl - galego - Galician gu - ગુજરાતી - Gujarati ha - Hausa he - עברית - Hebrew hi - हिन्दी - Hindi hr - hrvatski - Croatian hu - magyar - Hungarian hy - հայերեն - Armenian id - Indonesia - Indonesian it - italiano - Italian ja - 日本語 - Japanese jv - Jawa - Javanese ka - ქართული - Georgian ko - 한국어 - Korean lt - lietuvių - Lithuanian ml - മലയാളം - Malayalam mr - मराठी - Marathi ms - Melayu - Malay ne - नेपाली - Nepali nl - Nederlands - Dutch no - norsk bokmål - Norwegian Bokmål (Norway) or - ଓଡ଼ିଆ - Odia pl - polski - Polish ps - پښتو - Pashto pt - Portugal: português - Portuguese (Portugal) ro - română - Romanian ru - русский - Russian sk - slovenčina - Slovak sl - slovenščina - Slovenian sq - shqip - Albanian sr - српски - Serbian sv - svenska - Swedish ta - தமிழ் - Tamil te - తెలుగు - Telugu th - ไทย - Thai tr - Türkçe - Turkish tw - 中文 (繁體) - Chinese (Traditional) uk - українська - Ukrainian ur - اردو - Urdu vec - veneto - Venetian vi - Tiếng Việt - Vietnamese yue - 粵語 - Cantonese zh - 中文 - Chinese Updates about Anna’s Archive , the largest truly open library in human history. AI companies destroy physical books — let’s scan rare books before it’s too late annas-archive.gl/blog, 2026-08-05 A guest post by Anna’s Archive volunteer “u” (translated from Chinese). TL;DR: AI companies are secretly buying, scanning, and destroying millions of physical books to train their models, permanently locking human knowledge inside private corporate servers. Anna’s Archive is urgently calling on volunteers worldwide to scan and upload books before this cultural heritage disappears forever. Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022. Anthropic’s “Project Panama” was exposed in a $1.5 billion copyright settlement. In early 2024, they launched this highly confidential project. The company has spent tens of millions of dollars purchasing millions of paper books, scanning them, training its Claude LLM, and then destroying them all. It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity. So why destroy physical books? Behind it lies the AI race and the interests of capital: It prevents these books from being scanned and used for training by competitors. It avoids legal risks. Destroying books is cheaper than lossless scanning. After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers. This battle for old books reveals a paradox: while promising to “make human knowledge accessible,” AI companies are dismantling the most solid carriers of human knowledge. The public may gain more intelligent AI assistants, but at the cost of a vast amount of knowledge resources disappearing from the public domain. Shadow libraries As the world’s largest shadow library, Anna’s Archive needs a plan to combat the destruction of physical books by AI companies. After all, the emergence of shadow libraries is the greatest miracle of knowledge sharing in the 21st century. Along with other shadow libraries, we’re building a digital library of Alexandria, an inextinguishable light of humanity. We need the help of volunteers worldwide to scan materials (including books, journal articles, newspapers, magazines, ancient books, rare books, and other materials) from every library and archive around the world and upload them to the shadow library for knowledge preservation, especially those that are easily lost. If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth. For small scans and uploads, we usually award recognition and lifetime membership to Anna’s Archive. For large-scale scans and uploads of books, we can help pay for the scanning fees and other rewards. Time is running out Since the beginning of 2025, AI-generated content has accounted for more than half of newly published internet content. A frightening reality emerges: if much of the future content consists of AI-generated books and papers, will humans be able to distinguish them? Once AI has absorbed even the last sentence written by humans on paper, all that will remain on the internet will be AI’s own words. In such a world, how can human civilization be preserved? Shadow libraries offer the best answer. If you want the memory of human civilization to no longer be monopolized, if you want future generations to be able to read all of humanity’s wealth for free, if you don’t want publishers making a fortune while authors receive little, then please help us. Please make any contribution you can, whether it’s scanning and uploading books, purchasing books and papers to scan and upload, or donating. With the efforts of all humanity, the monopoly on knowledge will be broken. Each of us can make history. This is a race against time. Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers. - Anna’s Archive volunteer “u” Relevant tickets for more information: #223 #187