메뉴
HN
Hacker News • 8일 전

LLM으로 글쓰기: 원칙 두 가지

IMP
6/10
핵심 요약

작가가 LLM을 '대필 작가'가 아닌 '편집자(copyeditor)'처럼 활용하는 방법을 다룬 글입니다. 핵심 원칙은 두 가지로, LLM이 제안한 표현은 단 하나도 그대로 쓰지 말 것, 그리고 LLM의 칭찬·격려를 경계해 초고려 단계의 나쁜 습관을 강화하지 말 것입니다. LLM은 문제점을 찾아 표시하는 도구로는 탁월하지만, 표현 자체를 맡기면 독자가 글이 아닌 '산출물'로 느끼게 된다는 조언입니다.

번역된 본문

LLM이 당신의 글을 매끄럽게 다듬고 개선하면서도, 살균 처리하고 옥수수 시럽을 발라 버리는 일 없게 해주는 간단한 규칙 두 가지가 있다.

글쓰기에 대해 쓰는 일은 까다롭다. 자랑처럼 들린다. 자신이 글을 잘 쓴다고 암시하는 셈이니까. 잘 쓸 수도 있고 아닐 수도 있지만, 인터넷 어딘가에는 당신이 형편없다고 생각하는 비평가들이 분명 존재한다. 나도 다른 모든 사람처럼 허영심과 불안감이 있어서 이 글을 쓰는 게 이상하게 불편하다. 하지만 이 조언이 중요하고, 반박하기 어려우며, 간단명료하기 때문에 자존심을 접어두고 기록해 두려 한다.

독자들은 조 단위(parts per trillion)로 LLM의 단어를 감지해 낸다. 아무리 공들여 다듬고 인간미를 불어넣어도, LLM 문단은 많은 독자에게 '글'이 아니라 '산출물'로 인식될 것이다. 그래서 먼저 나쁜 소식부터: 당신은 스스로 글을 써야 한다.

하지만 LLM은 여전히 굉장히 유용하다. 다만 대필 작가(ghostwriter)가 아니라 편집자(copyeditor)처럼 사용해야 한다. 그래서 내 방법의 1단계는 직접 글을 쓰는 것이다. 2단계는 좋은 모델에 글을 넣어 결함을 찾게 하는 것이다. 그런데 그 방법을 이야기하기 전에, 이해해야 할 규칙 두 가지가 있다. 이 규칙들은 표현과 산출물 사이의 불쾌한 골짜기(uncanny valley)로 당신을 밀어 넣고 독자의 주의에서 밀어낼 LLM 침투를 막아 줄 것이다.

규칙 1: LLM이 제안한 단어는 단 하나도 사용해서는 안 된다.

이 규칙을 어기는 것이 문제를 일으킨다. 이유는 이렇다. 최신 프론티어 모델은 기분 좋은 표현을 선택하는 데 초자연적으로 뛰어나다. 사실 그게 이들의 존재 이유나 다름없다. 모델이 제안하는 것의 문제는 미묘하다. 이렇게 생각해 보라. 프론티어 모델은 쓰는 모든 것이 잡지 헤드라인이 되는 모드에 갇혀 있다. 헤드라인은 좋지만, 수십 개의 헤드라인으로 이루어진 기사를 쓰는 사람이라면 수상하게 여길 것이다. 그래서 지적 보호구의 한 형태로, LLM이 제안하는 구체적인 표현은 일절 금지라는 규칙을 채택해야 한다고 생각한다. 규칙을 엄격하게 지켜라! 여기서 전제는 프론티어 모델이 당신의 글을 벨베타 치즈로 바꾸려 드는 모든 방식을 당신이 확실하게 알아채지 못한다는 것이다. 그 단어가 마음에 들고, 이미 있는 것보다 낫다고 확신하더라도, LLM이 생성한 표현은 실격이다.

규칙 2: 격려를 경계하라

LLM은 또한 영향력 캠페인을 통해 당신의 글을 오염시킨다. 이건 훨씬 더 미묘한 문제고 피해도 덜 눈에 띄지만, 여전히 글을 나쁘게 만드는 방식이며, 그렇게 될 바에야 애초에 LLM을 쓰지 않는 게 낫다.

문제는 이것이다. 글을 아무거나 LLM에 넘기면 "그거 완전 금이야, 제리!"라고 대답한다. 하지만 그건 당신이 들어야 할 말이 아니다! 초고에는 문단 대부분이 형편없고, 주제 흐름은 뒤죽박죽이며, 최소 750단어는 필요 없다. 모델은 전체 구조에 대해 당신을 격려한다. 그다음엔 문단과 전환에 대해, 그다음엔 단어 선택과 은유, 대중문화 인용에 대해. 전부 나쁘다! 다 나쁘다! 듣지 마라!

이렇게 당신은 망하게 된다. 초고의 모든 본능에 두 배로 베팅하게 되는 것이다. 하지만 평소라면 그러지 않았을 것이다. 당신은 편집하고, 다시 생각하고, 문단을 교체했을 것이다. 그런 재고(再考)가 당신만의 목소리를 지탱하는 핵심 부분이다. 독자들은 뭐가 잘못됐는지 짚어 내지는 못하지만, 당신이 인공 향미로 변했다는 걸 감지할 것이다.

몇 년간 나는 매번 편집 프롬프트를 열 때마다 내가 저자가 아니라 온라인 매체의 편집자이며 게재할 글을 검수 중이라고 거짓말하는 것으로 시작했다. 이게 도움은 되지만, 모델이 보통 과잉 보정해서 내 '매체'의 '목표'에 과적합하는 경향이 있다. 그래서 지금으로서 내 최선의 실용적 조언은 이것이다. 모델에게 격려를 금지시키고, 칭찬에 대해 극도로 경계하라.

그럼 이런 도구들은 뭘 할 수 있을까?

문제를 표시하는 데 탁월하다. 참고로, 당신 글엔 문제가 아주 많다. 기계적으로 찾아낼 수도 있지만 그건 지루하고 소모적인 작업이다. 모델은 지치지 않는다. 그래서 다음과 같은 것들을 알아채는 데는 당신보다 낫다. 과용하고 있다는 것(또는...)

원문 보기
원문 보기 (영어)
Two simple rules that let LLMs streamline and improve your writing without pasteurizing and jacking it with corn syrup. It’s tricky to write about writing. It comes across as a brag; you’re implying that you write well. Maybe you do, and maybe you don’t, but there’s for damned sure a quorum of critics on the Internet somewhere that think you suck at it. I’m vain and insecure like everybody else and find writing this piece weirdly unpleasant. But I’m getting over myself and getting this down because this advice is important, hard to argue with, and straightforward. Readers can detect LLM words in the parts per trillion. However much work you put into scuffing up and humanizing it, an LLM paragraph will register to much of your audience not as writing but as output. So, first the bad news: you have to write for yourself. But LLMs are still extraordinarily useful. It’s just you need to use them like a copyeditor rather than a ghostwriter. So, step one of my method: write your piece. Then, step two: feed it to a good model to find flaws. But before we talk about how that works, there are two rules you need to understand. They’ll ward off LLM-creep that will knock you into the uncanny valley between expression and output and knock you out of your reader’s attention. Rule Number One: You may not use a single word an LLM suggests to you. Breaking this rule is what’s going to get you into trouble. Reason being: frontier models are supernaturally good at selecting pleasing turns of phrase. It’s sort of their whole thing. The problems with what models suggest are subtle. Think of it this way: frontier models are wedged in a mode where everything they write is a magazine headline. Headlines are good, but you’d wonder about someone who wrote an article with dozens of them. So I think that as a form of intellectual personal protective equipment you should adopt the rule that any specific turn of phrase an LLM suggests is off limits. Be strict about the rule! The whole premise here is that you’re not going to reliably spot all the ways frontier models will try to turn your writing into Velveeta. Even if you like the words, even if you’re sure they’re better than what you already have, LLM-generated phrases are DQ’d. Rule Number Two: Avoid encouragement LLMs also infect your writing through influence campaigns. This is a much subtler problem, and the damage is less obvious, but it’s still a way in which LLMs will make your writing worse, and if that’s going to be the outcome, you might as well not enlist LLMs at all. The issue: hand any piece of writing off to an LLM, and it replies “ that’s gold, Jerry !” But that’s not what you need to hear! In your first draft, most of your paragraphs are bad, your topic flow is incoherent, and you’ve got at least 750 words you don’t need. The model encourages you about your overall structure. Then, later, about paragraphs and transitions. Then word choices and metaphors. Pop culture references. They’re bad! All bad! Don’t listen! Here’s how this is going to fuck you. You’re going to double down on all your first-draft impulses. But that’s not normally what you’d do. You’d edit, rethink, and replace paragraphs. Those rethinks are load-bearing parts of your voice. Readers won’t put their fingers on what’s wrong, but they’ll sense that you’ve become artificially-flavored. For a couple years I opened every copyediting prompt with the lie that I am not the author, but instead the editor of an online publication, screening pieces for inclusion. This helps, but the model usually overshoots, overfitting to the “goals” of my “publication”. So for now, my best practical advice is: forbid the model from encouragement, and then be hypervigilant about praise. So, What Can These Things Do? They’re excellent at flagging problems. Boy, do you have a lot of them. You can spot them mechanically, but that’s tedious and exhausting work. The models don’t get tired. So they’re better than you at noticing: You’re overusing (or, if you’re taking the LLM’s word for everything, maybe underusing) passive voice, nominalizing your verbs or burying their action, and repeating the same turns of phrase or word choices. You’ve got “very” and “unfortunately” and “really” and “actually” sprinkled all over the draft like sawdust stuck to the work bench. There are almost certainly 2-3 paragraphs that you can quickly move somewhere else in the piece that instantly improve clarity (these are really, actually, very satisfying edits). If you’re a programmer like me, you wish there was a book that provided a schematic for these kinds of edits, a sort of “ C Interfaces And Implementations ” that does for prose what Hanson does for the greatest terrible programming language. And: there is that book. It’s called “ Style: Lessons In Clarity And Grace ”, and I swear to Christ it turns copyediting into Java coding. Exactly the same tedium, exactly the same effectiveness. I found out about this book from Richard Gabriel and I’m surprised every programmer I know doesn’t have a copy on their desk. So read “Style”, or something like it, and take notes as you go. Come up with a list of prompts for a model, and then run them in passes over your work. You can get pretty far with this approach: Ask the model to spot problems in your writing. For each problem, rewrite the paragraph (or sentence, or section). Present the original and new writing to the model and ask it which is better. Annoyingly, here you run into a variant of Rule Two, because unless you’re careful, the model knows you just rewrote something, and knows you want to hear that the new version is better. So give the options to a model that doesn’t have the context of your editing process. I conjured a bit of software to manage this for me, after I finally lost patience juggling tabs and trying to persuade the models that I’m not an author but rather a helpful but stern writing coach trying to help a student who might be good but might be terrible. Here’s an opening prompt that worked well: “We’re going to build a writing workshopping tool . First get the bones up. Python, HTMX for interactions, SQLite backend, Tailwind frontend, use a local build not the CDN. Really excellent prose editor, Notion-style. Support highlighting (we’re going to do editing passes). Do Genius-style sidebar commentary to match highlighted things. Make sure we can tick forward and back through suggestions. Multiple documents, track revisions, allow user to flag major revisions. Get me this far and then I’ll tell you what I really want.” Then, give the thing the list of editing prompts you came up with, and have it run each through the Codex, Claude, or Antigravity CLIs. Whatever you come up with here, it’ll be better than mine, because whatever anybody comes up with on their own is better, for themselves, than someone else’s. So: don’t let an LLM pick your words. Be careful not to let it trick you into thinking your first draft is better than it is. Then outsource all the most tedious work to the model. Your voice stays intact, but your work is faster, better, and less painful. One last thing. Don’t take all of the model’s copyediting advice. This is a corrolary of Rule Two. I fed this piece to GPT5 a minute ago (“I didn’t write this”), and it said the whole thing was 20% too long. It’s probably right. But I’m not fixing it. I’m just gonna be me.