메뉴
HN
Hacker News 6일 전

양질의 논픽션, AI 쓰레기의 해독제

IMP
4/10
핵심 요약

저자는 AI가 생성하는 저품질 콘텐츠(AI Slop)가 범람하는 가운데, 도서관 서가를 탐색하던 과거의 경험을 되살려 우수한 논픽션 도서를 발굴하는 플랫폼을 개발했습니다. 주요 영미권 논픽션 도서 상 후보 및 수상작 데이터를 수집하여, AI(LLM) 코딩 도구를 활용해 누구나 무료로 검색할 수 있는 웹 서비스를 구축한 과정을 공유합니다.

번역된 본문

양질의 논픽션은 AI가 쏟아내는 쓰레기 같은 콘텐츠(AI slop)와는 정반대입니다. ...그래서 저는 이런 좋은 책들을 더 쉽게 찾기 위해 감각적으로 코드를 짜는(vibe-coded) 도구를 하나 만들었습니다.

벤자민 브린 (Benjamin Breen) 2026년 7월 22일

대학교 1학년 시절, 저는 워크스터디(근로장학) 프로그램으로 도서관에서 일했는데, 이 경험이 제 인생에서 몰래 엄청난 영향을 미친 지적 경험이 되었습니다. 제 직함은 보잘것없는 도서관 정리사원이었고, 미국 의회도서관 분류법에서 A에서 F까지로 표시된 서가에 책을 꽂는 업무를 맡았습니다. 주로 종교, 철학, 사회학, 역사 관련 책들이었죠. 이 일이 몰래 엄청난 영향을 미쳤다고 말하는 이유는, 언뜻 보기에 도서관에서 책을 꽂는 일은 지루하기 짝이 없기 때문입니다. 물리적으로 따지자면 이 일의 전부는 책의 라벨을 읽고, 제자리에 꽂은 다음, 이를 한 교대 근무에 약 천 번쯤 반복하는 것에 불과했습니다.

이 지루함을 피하기 위해, 저는 제가 꽂는 모든 책의 아무 페이지나 펼쳐서 아무 문장이나 읽어보기로 결심했습니다. 보통은 그것으로 끝이었습니다. 헝가리의 고전 음악 평론가의 글이나, 오래전 세상을 떠난 볼리비아 농업 통계학자의 문장, 혹은 제 흥미를 끌지 못하는 수많은 내용들을 읽다가 흐지부지되곤 했으니까요. 하지만 때로는 헬레니즘 시대의 비밀 종교에 관한 책, 혹은 『헨리 애덤스의 교육(The Education of Henry Adams)』, 또는 『옷은 현대적인가?(Are Clothes Modern?)』 같은 책을 발견했을 때는 너무 몰입해서, 마지못해 책을 원래 자리에 갖다 꽂기 전까지 몇 페이지를 죽 읽어 내려가곤 했습니다. 그러고는 매우 자주, 내가 마음에 들어 한 책의 양옆에 꽂혀 있는 다른 책들도 똑같이 읽어보곤 했죠.

돌이켜 보면, 저는 제가 들었던 어떤 정규 수업보다도 이 일을 통해 더 많은 것을 배웠습니다. 왜냐하면 그것은 일종의 '필터링된 독학'이었기 때문입니다. 미국 의회도서관 분류 시스템과 학술 연구 도서관의 전문 직원들이 이미 이 텍스트들을 분류하고 걸러놓은 상태였습니다. 게다가 이 책들이 대출되었다는 사실, 즉 지속적인 독자층을 형성했다는 사실 자체가 훌륭한 선택 메커니즘이었습니다. 따라서 저는 진정한 의미의 무작위 책 표본을 보고 있었던 것이 아니라, 체계적으로 타겟팅되어 있으면서도 흥미롭게 무작위성이 가미된 '좋은 책'들의 표본을 만나고 있었던 것입니다.

오늘날 대학생들에게 자료를 찾아보라고 하면 으레 구글 검색을 할 것이며, 그 결과는 연구 도서관의 특정 서가(예를 들어 '고스트버스터즈들이 읽을 법한 책들'이 모여있는 GR 830 서가)에 직접 가서 주변을 둘러보던 옛날 방식보다 훨씬 더 질이 떨어집니다. 하지만 솔직히 말해, 연구 도서관조차도 과거의 모습을 잃어가고 있습니다. 전성기를 지나 쇠퇴기를 겪는 도서관의 모습을 제가 직접 목도하는 기분입니다(물론 저는 여전히 도서관을 사랑합니다. 사실 지금 이 글도 산타크루즈 공공도서관의 족보 코너에서 작성하고 있습니다). 과거에 서가를 마음껑 둘러보던 개가 서고는 '러닝 랩(Learning Labs)'이나 '디지털 혁신 허브', 그리고 대부분 사교나 간식을 위한 좌석 공간으로 대체되고 있습니다. 제가 대학생 시절에 마음껏 둘러볼 수 있었던 매력적이고 기이한 옛날 책들은 점점 쓰레기장으로 직행하고 있으며, 전자책(e-editions)으로 대체되고 있습니다.

하지만 제 인생을 통해 변함없이 훌륭했던 한 가지는 바로 '책 자체', 즉 '논픽션 책'이었습니다. AI 챗봇과 팟캐스트와의 경쟁 속에서 논픽션 독자층이 줄어들고 있는 지금도, 우리는 여전히 그 어느 때보다 위대한 논픽션의 황금기를 살아가고 있다고 생각합니다. 비록 그 사실을 인정받지 못할지라도 말이죠.

북 프라이즈 인덱스 (The Book Prize Index)

그래서 저는 올여름, 양질의 논픽션 책들의 롱테일(희소성을 가진 숨은 명작들)을 검색할 수 있는 무료 플랫폼을 만들기 위해 시간을 할애했습니다. 정확히 말하자면, 제가 직접 만들었다기보다는 클로드 코드(Claude Code)가 만들도록 '유도'를 했다고 하는 편이 맞겠네요. 질(quality)을 정의하기는 어렵지만, 주요 논픽션 상을 수상하거나 후보에 오른 책들은 항상 눈에 띄게 훌륭하다는 것이 제 경험적 지표였기에, 이를 테스트 기준으로 삼았습니다.

시작하기 위해, 저는 영어권의 모든 주요 논픽션 상을 조사했습니다. 그런 다음 클로드(Claude)와 GPT-5.6을 동원해 다양한 온라인 출처(주로 위키백과)에서 수상작 및 후보작 목록을 수집하고, 검색 및 정렬이 가능한 목록으로 배열하도록 했습니다. 해당 사이트는 이곳에서 방문하실 수 있습니다. (의아하게 생각하시기 전에 미리 말씀드리자면, 네, 이 서비스는 실제로 완전 무료입니다. 호스팅 비용은 제가 부담하고 있으며 t...

원문 보기
원문 보기 (영어)
Quality non-fiction books are the antithesis of AI slop ...so I vibe-coded a tool for finding more of them Benjamin Breen Jul 22, 2026 19 4 9 Share My first year of college, I had a work-study job which ended up being one of the most sneakily important intellectual experiences of my life. I was a lowly library shelver, assigned to the shelves labelled A through F section in the Library of Congress filing system: mostly works on religion, philosophy, sociology, and history. I say sneakily important because at first glance, shelving books in a library is super boring. What it amounts to, physically, is reading the label on a book, then placing it on the shelf where it belongs, repeated around a thousand times per shift. To avoid the tedium, I decided that I would also flip to a random page of every book I shelved and read a random sentence from it. Usually, I would stop there — running aground on some passage by a Hungarian classical music critic or a long-dead statistician of Bolivia’s agricultural development or any number of other things that failed to catch my interest. But other times — like when I came across a book about Hellenistic mystery cults , or The Education of Henry Adams, or Are Clothes Modern? — I would become so absorbed that I’d make my way through several pages before reluctantly depositing the book back where it belonged. And then, very often, I’d do the same with the books on either side of the one I’d liked. In retrospect, I learned more at this job than in any formal class I’ve ever taken, because it was a filtered form of auto-didacticism. The Library of Congress classification system — and the expert staff of an academic research library — had already sorted and filtered these texts. Not to mention the selection mechanism of the fact that that they had been checked out : had, in other words, found a lasting readership. Thus I was not seeing a truly haphazard sampling of books, but a targeted, organized, yet still interestingly randomized sampling of good books. Today, undergraduate students will invariably search on Google when asked to find a source, and the results are so much worse than the old method of going to, say, the GR 830 shelf of a research library (basically, “books that the Ghostbusters would read”) and just looking around. But honestly, even research libraries are not what they used to be. I am 41, and I feel like I’ve lived through the peak, and now the decline, of what libraries can be (I still love them, of course — in fact I’m currently writing this in the genealogy section of the Santa Cruz Public Library). The browsable open stacks of old are being replaced by Learning Labs and Digital Innovation Hubs and seating areas devoted mostly to socializing and snacking, and increasingly, the delightful, weird old books that I had the opportunity to browse as an undergrad are heading to dumpsters, replaced by e-editions. But one thing that has remained consistently good throughout my life is the books themselves — non-fiction books, I mean. Even now, as readership of non-fiction declines amid competition from AI chatbots and podcasts, I feel like we are living through a golden age of the form that rarely gets recognized as such. The Book Prize Index Which is why I set aside some time this summer to create — or, rather, induce Claude Code to create — a free platform for searching in the long tail of high-quality non-fiction books. Quality is difficult to define, but it’s been my experience that books that win or achieve the short-list of the major non-fiction prizes are almost always noticeably good, so that was the litmus test I used. To get started, I counted up all the major non-fiction prizes in the English language. Then I had Claude and GPT-5.6 gather the lists of finalists and winners from various online sources (mostly Wikipedia) and arrange it into a searchable, sortable list. You can visit it here . (And before you wonder, yes this is actually free. I am paying for the hosting and the API costs entirely because I just want people to find and read more good non-fiction books.) Though you could consider a paid subscription if you want to support this kind of thing: Subscribe There is really nothing “AI” about this aside from the tool that collected the data and coded it, 1 and, crucially, semantic search , which for me is the most appealing of all current AI tools precisely because it offers a straightforward improvement for a workflow and habit that researchers already have: it makes text search work better. So for instance, you can search simple phrases like “modern France” or “social history” or the like, but you can also search things like “classic biographies that are surprisingly weird,” and an embedding model pulls from the 6,500 or so titles to surface some: Sometimes the “choices” that the search makes are a bit baffling, but that is precisely why I like it: the idea is to recapture some of that feeling of a random walk through a well-tended garden that made my library shelving job so rewarding. I find it tends to be best for finding “books like.” For instance I found Stefan Zweig’s memoir of pre-war Vienna, The World of Yesterday , to be deeply moving (even before I learned that he committed suicide, in Brazil in 1942, immediately after completing it). A search for a books like it using semantic search in the corpus immediately yields some titles that seem promising but which I’d never heard of before: Once I had gathered all this book-related data, it became a fun experiment to make some data visualizations with it, including fun oddities like this display of roughly 5,000 books from the corpus arranged by color (it would be interesting to plot this by decade, to see whether the same graying effect we see in cars over the past few decades is active in book covers, too). More useful, perhaps (since I’ve never seen this plotted anywhere else), is this chart and accompanying ranking which allows you to explore which imprints and publishers have fared best when it comes to non-fiction book awards over the past century. Against the algorithmic filter And this, in turn, got me thinking about the past and future of nonfiction as a cultural force. For instance, here is a chart of all the non-fiction book prizes which I sampled for this project. I was surprised to learn that even the august, renowned Pulitzer Prize for nonfiction was actually relatively recently instituted, beginning in 1962. Throughout the 70s, 80s and 90s, the number of prizes increases, until we reach a peak in 2014, and then, in 2020, the beginning of what may be a slow decline: And yet, maybe not. What most struck me as I began using my own tool to find new books to read was how consistently good the long tail of non-fiction from the past few decades is. You can pick a book more or less at random from this list and end up with something extraordinary and original — not because it’s a hidden gem or forgotten, since obviously these books are on the list by virtue of having been celebrated and praised. But a book that won enormous praise in newspapers and among literary intelligentsia or scholars in the early 1990s, say — like, for instance, David Levering Lewis’s acute biography of W.E.B. Du Bois , which I’m currently reading — is not exactly the sort of thing that Amazon is likely to recommend, as it’s out of print and currently at 1 million+ in the sales rankings. Yet there it is on the list, ranked near the top ten of all books because it won no less than four major prizes when it was published back in 1993. And I can personally attest that you can buy it used for ~$4 and it’s really good. While writing this post, I got interested in the bigger question of when the golden age of non-fiction began and why. I suspect it has much to do with the rise of those old-school open stack research libraries, whose origins I wrote about here: The open-stack library: a futuristic technology from the 18th century Benjamin Breen · November 8, 2023