메뉴
HN
Hacker News 5일 전

AI 소프트웨어 공장이 실패하는 이유

IMP
8/10
핵심 요약

최근 AI 코딩 에이전트 도입이 빠르게 늘면서 PR 리뷰 품질 저하와 버그 증가 등 시스템 장애가 속출하고 있습니다. 저자는 자동화된 '루프 엔지니어링(Loop Engineering)'만으로는 근본적인 모델 한계를 해결할 수 없다고 강조합니다. AI 코딩 툴 활성화에도 불구하고 코드베이스 유지보수와 품질 관리가 오히려 악화되는 실무적 딜레마를 짚어냅니다.

번역된 본문

제목: 소프트웨어 공장이 실패하는 이유 (또는: 하네스 엔지니어링(Harness Engineering)만으로는 부족하다)

참고: 저는 HumanLayer라는 회사를 운영하며 인간과 에이전트 협업 영역의 도구를 개발하고 있습니다. 따라서 아래 제가 할 말은 약간 편향되어 있을 수 있습니다. 그럼에도 불구하고 이 주제가 도움이 되길 바라며, 저만큼이나 흥미롭게 느끼셨으면 좋겠습니다. - Dex

자, 이제 우리 모두 루프(Loops)를 돌리는 시대가 되었습니다.

우리 모두는 AI 코딩을 실제 프로덕션에 적용하기 위해 경주하고 있습니다. 루프 엔지니어링(loop engineering)에 대해 많은 이야기가 오갔으며, 지배적인 통념은 아마도 '우리가 루프를 더 많이 작성해야 한다'는 것입니다. StrongDM은 인간이 코드를 읽거나 쓰지 않는 무인 소프트웨어 공장(lights-off software factory)에 대해 글을 썼습니다. 이들의 논리는 대략 이렇습니다. 당신이 병목이다. 모델은 충분히 훌륭하다. 코드는 공짜다. 그냥 더 많은 기능을 배포하라.

OpenAI의 라이언 로포폴로(Ryan Lopopolo)는 2월에 이에 대해 글을 썼고, 4월에는 OpenAI의 소프트웨어 공장인 '심포니(Symphony)'에 대한 강연을 했습니다. 제가 언급한 이 사람들은 정말 엄청나게 똑똑하며 저는 그들에게 깊은 존경을 표합니다. 하지만 가장 냉소적으로 보면, 이는 벤처캐피털(VC)의 돈을 더 많이 끌어다 대충 찍어내는 기계(슬롭 캐논)에 쏟아부어 달라는 또 다른 핑계일 뿐이라고 볼 수도 있습니다.

상황이... 음... 그렇게 되고 있습니다.

지인인 마리오는 'AI 엔지니어 유럽(AI Engineer Europe)' 행사에서 연단에 올라 우리에게 속도를 늦추라고 간청했습니다. 왜냐하면 코딩 에이전트의 실수로 인한 서비스 중단이 발생해서는 안 되는 기업들이 실제로 코딩 에이전트의 실수로 인해 서비스 중단을 겪고 있기 때문입니다. Matt Pocock이 지적했듯이, 코드베이스가 우리가 지금까지 본 그 어느 때보다도 빠르게 붕괴하고 있습니다.

저는 StrongDM의 그 어두운 공장(dark factory)이 어떻게 되었는지에 대한 명확한 데이터나 결과를 찾지 못했습니다. 올해 2월과 6월 사이에 몇몇 드문 업데이트가 있었을 뿐입니다. Faros AI의 사람들은 보고서를 하나 냈습니다. 우리 모두가 1월과 2월에 이런 AI 코딩 도구들을 도입한 이후 풀 리퀘스트(PR) 리뷰의 품질이 크게 떨어졌습니다. 댓글은 더 많아지고 길어졌으며, 리뷰 없이 병합되는 PR도 수두룩합니다. 인시던트(사고)는 급증했습니다. 개발자당 발생하는 버그 수도 크게 늘었습니다. 이 보고서는 검증된 명백한 증거라기보다는 상관관계 신호에 가깝고, 이 글의 요지는 쓰레기 같은 데이터(slop data)를 경계하라는 것이지만 제가 본 바에 비추어 볼 때 방향성은 맞는 것 같습니다.

'당신이 제대로 사용하지 않아서 그래요' (아닙니다)

많은 사람들이 이것이 역량 부족 문제라고 말할 것입니다. 좋은 결과를 얻지 못했다면 그것은 당신 잘못이라고요. 하지만 당신이 그것을 어떻게 사용하든, 토큰을 최대한 쥐어짜는(token-maxxing) 것이 효과가 없다면 그것은 당신의 역량 문제라고 확신할 것입니다. 그저 토큰을 더 많이 써야 합니다. 코드를 읽는 것은 포세요. 그리고 당신이 이제 막 거기에 도달했다면, 그것은 발전 과정의 일부라고 약속할게요. 저 역시 지난여름에는 이렇게 생각했습니다.

불행히도 제 자존심을 위해 말하자면, 제가 '어떻게 하면 더 잘 사용할 수 있는지'에 대해 떠든 약간의 어리석은 발언들이 녹음되어 지금은 유튜브에서 누적 100만 조회수를 기록하고 있습니다. 자랑하려는 것이 아닙니다. 제가 오랫동안 코딩 에이전트를 사용하는 최고의 방법들을 깊이 파고들어 왔으며, 많은 사람들이 실제로 유용하다고 생각하는 것들을 발견했다는 것을 보여주기 위해 이를 공유하는 것뿐입니다.

코딩 에이전트를 위한 고급 컨텍스트 엔지니어링 감(Vibe)은 허용되지 않음 - 복잡한 코드베이스에서 어려운 문제 해결하기 RPI에 대해 우리가 잘못 알고 있던 모든 것

아무튼, 우리가 감내해야 했던 온라인상의 '그냥 토큰을 더 써라'는 잔소리들의 약속은 간단히 말해 이것입니다. 충분한 하네스 엔지니어링(harness engineering)을 거치면 우리는 두 마리 토끼를 모두 잡을 수 있다는 것입니다. 10배에서 100배 더 빠르고, 고품질이며, 우리 모두가 싫어하는 '코드 리뷰'라는 골칫거리를 아무도 하지 않아도 된다는 것이죠.

우리가 해야 할 일이라고는 그저 린터(linter)를 더 많이 설정하고, 충분한 수의 PR 리뷰 봇에 '적대적 리뷰(adversarial review)'와 같은 마법의 단어를 뿌려주는 것뿐입니다. 그러면 우리의 소프트웨어는 아무 사고 없이 스스로 즐겁게 빌드될 것입니다.

이것은 역량(Skill)의 문제가 아닙니다.

제가 여러분을 설득하고자 하는 것은, 아무리 많은 하네스 엔지니어링이나 루프 최적화(loopsmaxxing)를 해도 근본적인 모델 훈련 문제를 해결할 수는 없다는 것입니다. 이 문제를 제대로 다루기 위해 저는 코딩 모델이 실제로 어떻게 훈련되고 평가받는지 파고들어야 했습니다. RLVR(Reinforcement Learning with Verifiable Rewards)과 벤치마크 측면 모두를 고려해서 말입니다.

이 글에서는 다음과 같은 내용을 다룰 예정입니다. 소프트웨어 공장은 1968년으로 거슬러 올라가며, 어떻게 발전해 왔는지 등등...

원문 보기
원문 보기 (영어)
Why Software Factories Fail or: harness engineering is not enough Note I run a company ( HumanLayer ) building tools in the human/agent collaboration space, so what I'm gonna say below may be a tad biased. Perhaps in spite of that, I hope you find the subject helpful or at the very least that you find it as interesting as I do. -Dex i guess we doin loops now We're all racing to put AI coding into production. A lot has been said about loop engineering, and the prevailing wisdom is that we should probably write more loops. 1 StrongDM wrote about their lights-off software factory where no human reads code and no human writes code. The narrative goes something like this: You are the bottleneck. The models are good enough. Code is free. Just ship more stuff. Ryan Lopopolo of OpenAI wrote about this in February and gave a talk in April about OpenAI's software factory, Symphony. These people are all really dang smart and I have a ton of respect for them. But the most cynical take here would be to call this yet another excuse to pump more VC money into the slop cannon. it's uh...it's going Our friend Mario got up at AI Engineer Europe and begged us to slow down -- because companies that have no business having outages due to coding-agent mishaps, are, well... having outages due to coding-agent mishaps . As Matt Pocock put it, codebases are falling apart faster than they ever have before . I haven't been able to dig up any definitive data/findings from StrongDM on how that whole dark factory went. The weather-report has a few sparse updates between February and June of this year. The folks at Faros AI put out a report: since we 2 all picked up these AI coding tools back in January and February, pull-request review quality is way down. More comments, longer comments, and tons of PRs getting merged with no review at all. Incidents are way up. Bugs per developer are way up. This report is more of a correlation signal than a verifiable smoking gun 3 , and the whole point of this post is to be wary of slop data, but it feels directionally valid based on what I've seen. "You're holding it wrong" (you're not) A lot of people will tell you that this is a skill issue -- that if you're not getting good results, that's your fault. But however you're choosing to...erhm...hold it, I guarantee you're being told that if token-maxxing isn't working for you, it's a skill issue. You just need to spend more tokens. Let go of reading the code. And if you're just getting there, I promise it's part of the progression. I thought this way last summer too . Unfortunately for my ego, some dumb stuff I decided to say about "how to hold it better" got recorded and now has about a million cumulative views on YouTube. I am not trying to brag here, I share this only to establish that I've been going deep on the best ways to use coding agents for a long time now, and have discovered some things that many others have found genuinely useful . Advanced Context Engineering for Coding Agents No Vibes Allowed -- Solving Hard Problems in Complex Codebases Everything We Got Wrong About RPI Anyhow, The promise of all this online "just token harder" yapping we've been forced to endure is, succinctly: with enough harness engineering, we can get the best of both worlds: 10 to 100x faster, high quality, and nobody ever has to do that thing we all hate called code review All we have to do is configure more linters and sprinkle some magic words like "adversarial review" onto enough PR review bots, and our software will happily build itself without incident. This is not a skill issue What I'm gonna try to convince you is that no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue. To grapple with this, I had to dig into how coding models are actually trained and evaluated - with respect to both the RLVR and the benchmark side of things. In this post I'm gonna run through: Software factories date back to 1968, how have they evolved, and how has AI changed them Why models can generate mountains of slop despite ace-ing benchmarks (even the brand new "frontier" benchmarks) In spite of this , you can move pretty fast without setting your codebase on fire I'm gonna try to cut through the hype of every daily-emerging skills plugin and the ai-psychosis-tokenmaxxing advice pandemic, and talk in general terms about the types of things that work without referencing any particular skill or framework. Video Version: this post is based on (and expands upon) my keynote at AI Engineer World's Fair 2026 . Thanks to @addyosmani , @CyrusNewDay , @HamelHusain , @zeeg , @dillon_mulroy , @nayshins , and @jeffreyhuber for feedback on this post. An aside: this has nothing to do with vibe coding Addy Osmani detangled this thing that is worth highlighting: A developer vibe-coding a side project a dozen people will ever run, and a team keeping a ten-year-old enterprise system alive for another quarter, share almost no constraints worth naming, and most of the advice in circulation is really one of those two people telling the other how to live. If you love vibe coding, please, go on vibing. I still vibe code lots of things, I just also maintain lots of production software (and through HumanLayer, help 1000s of other engineers do the same), so the rest of this is aimed at folks solving hard problems in complex codebases. I hear the word brownfield a lot to talk about this split. Historically that meant some ten-year-old Java thing, but at the pace we can ship now, it feels like an agent-built codebase starts to struggle after maybe three to six months -- you start to slow down, and the way you approach adding new things has to change. A brief history of the software factory I've been building and studying software factories my whole career, but I only learned this recently: the term traces all the way back to a NATO conference in 1968 -- the same one that gave us "software engineering." The only other bit I find super interesting since then is that the US Department of Defense wrote a 31-page pdf about how the DoD needs to start using jenkins better or something . The 2022 software factory Let's ground our "software factory" definition around 2022, right before AI. In a typical software factory: People decide what to build -- engineers, PMs, leadership driving the vision It goes in a tracker -- Linear, Jira, whatever: a state machine of what needs to happen Someone grabs a ticket and builds it -- probably does some manual/automated testing while they're at it Pull request -- automated checks, a human reviews the code, maybe someone pulls it down to test Anything wrong? Loop back to "someone builds the thing" Ship to prod -- and it makes contact with users Add monitoring -- there's an entire industry built around paging an engineer at 3am when something breaks Users complain -- ask for things, find bugs, file feature requests → back to the team to add to the tracker wsff-boxes-2x.mp4 And on and on. We haven't even hit AI yet, and there are already several loops in this picture. front-loading alignment The thing teams figured out decades ago: building takes hours or days, and so does review. So we front-load the work -- planning, architecture proposals, sprint planning -- together, as a team. That means: less rework , because we aligned before anyone wrote code less time reviewing every line , if you've ever read a long-but-well-done PR, you know how fast the review goes when it's close-to-perfect We'll come back to this later - let's look at what happens when you bring agentic coding into the picture. The agentic software factory Now every company and their mother -- Ramp Stripe WorkOS Brex has spent the better part of this year explaining how they built an agent factory that ships on the order of 75% of their code . The agentic factory looks mostly like swapping "someone builds the thing" → "an agent builds the thing" -- there's some stuff here like orchestration, a harness, a sandbo