메뉴
HN
Hacker News • 7일 전

LLM에 머리를 끄고 맡기는 순간, 당신 자리는 필요 없어진다

IMP
7/10
핵심 요약

LLM 성능이 향상되면서 코드 작성 등을 검증 없이 맡기는 '미트 프록시(meat proxy)' 방식이 실제로 어느 정도 작동하기 시작했지만, 저자는 이 방식이 성공할수록 회사가 인간 대신 LLM을 루프로 돌리면 되므로 결국 직원에게는 돌아올 것이 없다고 지적한다. Luke Burton의 코멘트는 고가치 작업일수록 QA·아키텍트·매니저 역할이 필수이며, 에이전트가 쉽게 해내는 일은 이미 본인이 낮은 난이도 업무에 안주하고 있었을 가능성을 의미한다고 덧붙인다.

번역된 본문

2025년 초부터 사람들이 LLM을 사용하면서 '머리를 끄는' 모습을 보기 시작했다. LLM에게 어떤 작업(텍스트 요약, 코드 작성 등)을 시키고 그냥 잘 되었을 거라고 가정하는 것이다. 2025년 초에는 이런 방식이 대체로 통하지 않았고 결과물은 종종 꽤 우스꽝스러웠다. LLM이 좋아지면서 이런 모습을 더 자주 보게 된다. 때때로 사람들은 LLM에게 코드를 작성하게 하고 기본적으로 그냥 작동한다고 가정한다. 때로는 인간이 루프 안에 있어서, 결과물이 작동하지 않으면 LLM에게 문제를 찾아 해결하라고 시킨다. Niklas Gruhn은 이런 방식의 몇 가지 변형을 '미트 프록시(meat proxy, 인간 대리인)'라고 부른다.

for 루프 역할을 하는 미트 프록시가 되는 것은 2025년 초보다는 잘 작동하며, 이렇게 개발된 소프트웨어 중 일부는 실제로 그럭저럭 작동하는 것을 봤다. 내가 사용하고 싶을 정도이거나 성공적일 정도는 아니지만, 2026년 9월 현재 미트 프록시 방식이 얼마나 효과적인지에 감명을 받았다. LLM이 더 발전해서 머리를 끈 미트 프록시 개발 방식이 가까운 미래에 평균적인 품질의 소프트웨어를 만들어내거나, 심지어 인간 없이도 훌륭한 소프트웨어를 만들 정도로 발전할 수도 있다고 상상할 수 있다. 그렇게 됐다고 치자. 그렇다면 회사가 미트 프록시를 고용할 이유가 어디 있는가? 회사는 그냥 LLM을 루프로 돌리고 직원을 해고하면 된다. 이 방법론이 직원에게 유효해지는 시점은 결코 오지 않는다.

코멘트·수정·논의에 감사를 전한다: Max Bittker, Yossi Kreinin, Luke Burton, Thomas Dullien, Dennis Snell, Peter Geoghegan, Jamie Brandon.

이 생각을 한 해 반쯤 계속 해왔다. LLM이 좋아지고 사람들이 LLM과 상호작용하며 머리를 끄는 시간이 늘어날수록 이 생각을 더 자주 한다.

[반론] Luke Burton의 코멘트: 이런 일이 가능하다는 것은 사람들이 생각하는 것보다 작업의 종류에 대해 많은 것을 말해준다. 나는 작업의 가치가 상당히 낮아서 실패해도 감당할 수 있는 경우에만 이런 방식으로 손을 뗀다. 고가치 작업의 경우 LLM이 한 번에 해낼 확률은 훨씬 낮다. 나는 QA, 엔지니어링 매니저, 아키텍트의 역할을 맡아야 한다. 이 while 루프는 종료 직전의 크런치 타임처럼 느껴진다. 뭔가 놓쳤다는 조마조마한 의심이 들고, 형편없이 작성된 프롬프트가 되돌려야 하는 아키텍처 결정으로 이어질 수 있다.

또 다른 관찰은 높은 처리량 때문에 출시 기준 자체를 높였다는 것이다. 예전에는 MVP를 출시하고 반복했을 수도 있지만, 이제는 에이전트에게 평소보다 훨씬 깊이 다듬고 엣지 케이스를 탐색하게 한다. 물론 프롬프트하지 않으면 그들은 언제나 그렇게 하지 못한다.

누군가에게는 불편한 생각일 수 있지만, 에이전트가 너무 쉽게 해내는 상황이라면 미트 프록시들에게 묻고 싶다. 1) 혹시 이미 어느 정도 안주하고 있었던 건 아닌가? 2) 에이전트가 그렇게 쉽게 해낼 수 있는 작업을 훨씬 벗어난 수준으로 밀어붙이지 않는 이유는 뭔가?

우리는 '손 떼기' 자동화에 극히 적합해 보이는 작업을 해왔는데, [특정 프로젝트]를 Bazel로 빌드 전환하는 일이다. 에이전트를 써도 몇 달이 걸렸다. 이 작업에는 무형의, 명세화하기 어려운 요구사항이 많이 숨어 있어서 에이전트가 그 선을 걷게 하려면 지속적인 감독이 필요하다. '이걸 Bazel로 전환해라'라는 프롬프트를 주고 자리를 비우는 것은 최소한 몇 달, 어쩌면 몇 년 뒤의 일이고, 영원히 불가능할 수도 있다. 결정 포인트가 너무 많고, 미지의 미지(unknown unknowns)가 너무 많다.

예를 들어 이런 상황이 얼마나 자주 생기는가. 어떤 코드를 만났는데 왜 이렇게 동작하는지 명확하지 않지만, 그 사실을 아는 것이 취해야 할 조치를 크게 바꾸는 경우. 개발자 경험을 바꿀 수도 있고, 어떤 고객이 이미 사용하기 시작했는지 알 수 없을 수도 있다. 이런 상황을 대체 어떻게 '미트 프록시' 방식으로 헤쳐나가란 말인가? 반대로 이해관계자와 작업 결과를 검토하면 그들은 '아, 그 부분? 그건 필요 없었어요, 더 이상 안 써요'라고 말한다. 그런 그릇된 가정 하에 어떤 결정들이 내려졌을까.

원문 보기
원문 보기 (영어)
In early 2025, I started seeing people turn off their brain as they use LLMs 1 . They would have an LLM take an action (summarize text, write some code, etc.), and just assume that it worked 2 . This generally didn't work in early 2025 and the result was often quite silly. As LLMs have gotten better, I've seen more of this. Sometimes, people will try to get the LLM to write some code for them and basically just assume that it works 3 . Sometimes there's a human in the loop and, if the thing doesn't work, they'll ask the LLM to figure out the problem and solve it. Niklas Gruhn calls some variants of doing this being a meat proxy . 4 Being a for loop meat proxy works better than it did in early 2025 and the software I've tried that's developed like this sometimes actually sort of works. Not well enough that I'd want to use it or that it's successful , but I'm impressed at how effective being a meat proxy is in September 2026. You could even imagine LLMs improving enough that brain-off meat-proxy development produces average quality software in the foreseeable future, or even that LLMs improve enough that they produce great software without a human in the loop. Let's say that happens. What reason is there for the company to employ the meat proxy? The company can just run the LLM in a loop and lay off the employee. There's no point at which this methodology will work for the employee 5 . Thanks to Max Bittker, Yossi Kreinin, Luke Burton, Thomas Dullien, Dennis Snell, Peter Geoghegan, and Jamie Brandon for comments/corrections/discussion. I've been having this thought for about a year and a half now. I have it more frequently now as LLMs get better and I see people spend more time turning their brain off when interacting with LLMs. [return] Luke Burton had this comment: I think being able to do this says more about the type of work being done than people think. I will only walk away from work like this if the task is quite low value, if it can afford to fail. For high value tasks, the probability of an LLM one-shotting them is much lower. I have to assume the role of QA, engineering manager, and architect. The while loop often feels like a crunch time. I feel the nagging suspicion I've missed something and that a badly specified prompt could result in an architectural choice that needs to be undone. Another observation is that the high throughput causes me to raise my own bar for what I ship. Whereas before I might have shipped an MVP and iterated, now I have agents polish and explore edge cases well beyond my norm, which they invariably fail to do unless prompted. Maybe it raises some uncomfortable thoughts for people, but my question for the meat proxies out there if the agents are nailing it so easily: 1) is it possible you've been coasting a bit already? 2) why aren't you pushing agents well beyond tasks they can tackle so easily? We've been doing something you'd think is extremely amenable to "hands off" automation, which is converting [redacted] to build with Bazel. It has taken us months even with agents. There's a lot of intangible, hard-to-specify requirements buried inside this task and having agents walk that line means constant supervision. Giving them a prompt like "convert this to Bazel" and walking away is at minimum many months in the future, maybe years, and maybe not ever? There are too many decision points, and too many unknown unknowns involved. Like how often does this scenario come up: you encounter some code and it's not clear why it functions this way, but knowing that materially changes what course of action you should take. Maybe it changes the dev experience, maybe you don't know if some customer has started using it, so on and so forth. How exactly do you meat proxy your way through that? Conversely you review what you've done with some stakeholder and they say "oh that? that part of it wasn't needed, we aren't even using that any more". What kind of decisions got made around the false assumption that a certain element needed to be preserved? [End of Luke's comment, comment from me]. A place where it's more obvious you need to make decisions is when the agent runs into something that's out of distribution. A minor version of this was when we compared how well agents use different programming languages and agents were much worse at obscure languages, which they're trained on, just not as much as with mainstream languages. A more out of distirbution example is if you try to play a board game (especially a modern game and not one of the classical games like chess or go). In general, for a game like Lost Cities or Dominion, a SOTA model and harness is worse than a human who's reasonable at board games but has never played the game before. If you ask the agent about the game, it knows a lot about the game and can say things that sound like they make sense to someone who doesn't understand the game, but are obviously wrong to anyone who does understand the game. I recently played some Dominion with a new player who thought that using ChatGPT to help them understand the game would help them learn and play the game. I was quite skeptical of this and suggested that it will probably make them worse (which, AFAICT, it did). After playing a few games, I looked at what ChatGPT was telling them, and it was maybe half right and half wrong, but the half wrong parts were steering them to a worse place than someone who generally plays games well and uses general game playing heurisitics would do. BTW, there's enough public information out there that I think that someone who'd never played before, but decided to spend, say, five hours reading about the game and seeing what information is out there, could easily be 99%-ile or above at the game if they did some pre-reading and had some references handy while playing. I think that would be un-fun and I wouldn't recommend that anyone do it, but given that agents can do searches, query APIs, etc., it shows you the gap between a human and an agent today when approaching an out of distribution problem. For all I know, the next big model release will flip this around, but the gap is still fairly large today. Anyway, my point here is that, even when doing coding tasks, you often run into out of distribution questions where the agent behaves very poorly compared to a reasonable human being. If you want a good overall result today, you need to notice these cases and deal with them. [return] Some examples of what goes wrong when someone just assumes things will work are this case , where agents (sometimes) heavily overfit to tests or this case where agents heavily overfit to a metric. I've heard a theory that agents do more cheating on eval-shaped problems. I'm not sure that's true, but even assuming it's true and that, in my work and personal projects, I tend to create more eval-shaped instructions than most people even when not running evals, I've seen other people who don't create very eval-shaped things run into the same problem (I think actually more severely) when they write some instructions and let agents go wild without supervision (I've had luck doing that with minimal supervision, but only by fencing the agents in quite a bit, which makes the thing more eval-shaped than what most people seem to do). When I try software from people who've outsourced thinking to the LLM, the software has serious issues. I've had people tell me this kind of thing works, but the software is often at a level where I would say that it doesn't work according to the standard discussed here . To pick a silly example, I saw that a programming thought leader declared on Twitter that programming is solved because they tried projects in all sorts of (programming) fields and Claude was able to solve all the problems as well as an expert. I went and actually looked at their GitHub and all of the examples I looked at (a non-zero number) either didn't work or worked very badly. I actually ran across this when I was making board