메뉴
HN
Hacker News • 29일 전

소형 모델의 시대가 왔다

IMP
7/10
핵심 요약

gpt-5.6-luna 같은 소형 고속 모델이 놀라운 능력과 저렴한 비용(~$0.10)을 제공하면서 소비자용 AI 앱의 경제성이 finally 성립하게 되었다는 분석입니다. 저자는 '천재형' 업무보다 기업 업무의 약 95%를 차지하는 '빠르고 반응성 좋은 토큰 뿜어내기'형 업무에 소형 모델 수요가 폭발할 것으로 전망합니다.

번역된 본문

소형 모델의 시대가 왔다 (2026년 8월 26일) 최근 몇 주간 저는 gpt-5.6-luna를 가지고 실험해왔습니다. 놀라울 정도로 유능하고, 빠르고, 똑똑합니다. 약 100 tps(초당 토큰) 속도를 정기적으로 목격하며, 제 코드베이스와 이메일, 지식 베이스를 빠르게 훑고 다닙니다. 물론 루나의 가장 큰 특징은 비용입니다. 꽤 복잡한 리서치 작업을 실행해봐도 큰 청구 금액이 나오기가 어려울 정도입니다. 수천 통의 이메일을 검색하게 해도 API 비용이 수십 센트 수준에 그칩니다. (artificialanalysis.ai 제공) GLM 5.3의 등장으로 파레토 프론티어(최적 균형선)에 새로운 선택지까지 생겼습니다. 코딩 작업을 할 때 저는 거의 항상 가장 비싸고 강력한 모델(Fable 5, 5.6 Sol)을 찾곤 했습니다. 그래서 작고 빠른 모델들이 이룬 진보를 놓치기 쉬웠습니다. 제가 대화한 몇몇 투자자들이 이런 말을 했습니다: "소비자용 AI 회사가 더 많이 나오지 않는 게 이상하다. 왜 그럴까?" 여기엔 간단한 답이 있습니다: 토큰 비용입니다. AI 이전 시대의 대형 소비자 앱 플레이북은 대략 이렇습니다... 운영 비용이 비교적 저렴한 매력적인 웹사이트 만들기 사용자 대량 유치 (보통 바이럴을 통해) 자금 조달 후 더 많은 사용자로 확장 광고 마켓플레이스 구축 이것이 대부분의 대형 소비자 기업(Google, Facebook, Snapchat 등)을 설명합니다. 그런데 제품에 AI를 추가하려면 어떻게 될까요? 이제 모든 요청마다 실제 추론 비용이 발생합니다! 갑자기 필요한 자본 규모가 극적으로 커집니다. 제가 애용하는 평가 테스트 중 하나는 저에게 맞춤화된 일일 뉴스 사이트를 만드는 것입니다: 인터넷에서 @calvinfo를 리서치하고 그 사람이 좋아할 만한 뉴스를 파악 오늘의 주요 뉴스를 개인화한 마이크로 사이트 제작 hn, reddit, twitter 등 검색 이전 세대 모델(Sonnet급)이라면 결과를 얻는 데 약 $1이 들었습니다. 소비자 앱에 월 $30을 청구하는 건 지속 불가능합니다. 물론 최적화 여지는 많지만, WSJ나 이코노미스트가 받는 요금을 청구하려면 그에 걸맞은 가치를 제공해야 합니다. 그런데 루나를 보면 결과가 꽤 괜찮은 데다 평균 비용이 약 $0.10입니다. 이제 얘기가 다릅니다! 제 생각에 이것이 더 흥미로워지는 지점은 비즈니스 세계입니다. 세그먼트(Segment) 공동 창업자인 피터와 최근 하이킹을 하며 이야기를 나눴습니다. 피터는 자신의 여러 스타트업에서 두 종류의 업무를 발견했습니다: 'IQ 180' 업무. 천재 과학자 같은 사람이 상상도 못한 미친 해결책을 내놓는 것. '토큰 분사기(token spewer)' 업무. 초고 반응성으로 수십 개 전선에서 일을 밀어붙이는 것. 피터는 여러 회사를 운영합니다. Segment 외에도 Charm Industrial을 위해 1억 달러 이상을 조달했고, 최근에는 Revoy의 시리즈 A를 마쳤습니다. 시간 관리가 엄청나게 체계적이고 효율적인 사람입니다. 그럼에도 피터는 자기 업무의 약 95%가 두 번째 유형에 속한다고 말했습니다. 전화를 걸고, 사람들을 재촉하고, 잡일을 처리하는 것이죠. 분명히 말하자면, 피터는 깊은 문제를 해결하는 IQ 180급 기술 두뇌가 없었다면 자기 회사들은 이미 물 건너갔을 것이라고 합니다. 다만 대부분의 업무가 두 번째 유형에 속한다는 것입니다. '프론티어급' 모델에 대한 수요는 계속 복리로 증가할 것이라고 봅니다. 특히 새로운 돌파구나 발견이 필요한 분야(엔지니어링, 하드 과학, 모델 훈련)에서요. 하지만 '빠르고/저렴하고/충분히 좋은' 모델에 대한 수요도 이제 막 폭발하려 한다고 생각합니다. 매일 상호작용하는 사람들을 생각해보세요: 동료, 벤더, 고객. 십중팔구 우리가 원하는 건 초고 반응성으로 일을 그냥 처리해주는 사람입니다. 오늘날 기업의 대부분의 '인간 토큰'이 이런 방식으로 소비되고 있습니다. 채용도 빠르고/저렴하고/충분히 좋은 유형에 크게 치우쳐 있죠. 빠르고/저렴하고/충분히 좋은 모델을 비즈니스 현실로 만들려면 해야 할 일이 많습니다. 새로운 하니스(실행 환경), 프롬프트 인젝션 방어, 역할 및 권한 체계 등입니다. 하지만 우리가 해결해낼 것이라 확신합니다. 여러분도 실험 중이라면...

원문 보기
원문 보기 (영어)
Small Models Have Arrived AUG 26, 2026 For the past few weeks, I've been playing with gpt-5.6-luna . It is shockingly capable, fast, and smart. I regularly see it do ~100 tps, and rip around my codebase, email, and knowledge base. Of course, the biggest thing with luna is the cost . I've tried running some fairly complicated research threads, and it's pretty tough to run up a large bill. Even having it search across thousands of emails, I end up with an API cost in the tens of cents. Courtesy of artificialanalysis.ai With GLM 5.3, we even have a new option at the Pareto frontier. When doing coding work, I almost always reach for the most expensive and capable models (Fable 5, 5.6 Sol). So it's been easy to miss the progress the small fast models have made. One thing a few investors I've talked with have mentioned: "It's weird we're not seeing more consumer AI companies. Why is that?" There's a straightforward answer: token costs. In the times before AI, the playbook for big consumer apps looked like this... create some sort of compelling website which is fairly cheap to run attract a bunch of users (typically with some virality) raise money, scale to more users create an ads marketplace This roughly describes most of the big consumer companies (Google, Facebook, Snapchat, etc.). 1 But what if you want to add AI to your product? Well, now you have some real inference costs on every request! Suddenly the amount of capital required increases dramatically. A pet eval of mine is to build a daily news site, personalized to me: research @calvinfo on the internet. figure out what news they might like. build a micro-site with today's top stories, personalized for them. search hn, reddit, twitter, etc. With the previous generation of models (Sonnet class), you'd spend ~$1 to get anywhere. Charging $30/mo is untenable for a consumer app. There's obviously a lot we can optimize here, but if you're charging what the WSJ or The Economist charges, you'd better be delivering similar value. But looking at luna, the results are pretty decent, and the average cost is ~$0.10. Now we're talking! Where I think this gets even more interesting is in the world of business. My Segment co-founder Peter and I were recently comparing notes on a hike. Across his various startups, Peter has seen two kinds of work: the "IQ 180" work. some mad scientist genius type comes up with some crazy solution you've never thought of. the "token spewer" work. being ultra responsive, pushing the ball forward across dozens of different fronts. Peter runs multiple companies. Beyond Segment, he's raised $100m+ for Charm Industrial , and just recently closed a Series A for Revoy . He's incredibly organized and efficient with his time. And yet, Peter mentioned that ~95% of the work he does falls into bucket 2. It's hopping on calls. Nudging people. Blocking and tackling. To be clear, Peter says his companies would be dead-in-the-water today without an IQ 180 technical mind solving the deep problems . Just that most of his work falls in bucket 2. 2 I think demand for "frontier-level" models is going to keep compounding. Especially for fields that require novel breakthroughs or discovery (engineering, hard science, model training). But I also think the demand for "fast/cheap/good-enough" models is just about to take off. Think of the people you interact with on a daily basis: coworkers, vendors, and customers. Nine times out of ten, you want someone who is super responsive, and just handles things for you. Most of the "human tokens" at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype. There's a lot of work that needs to happen to make fast/cheap/good-enough models a reality for business. New harnesses, prompt injection safety, roles, and permissions. But I'm confident we'll figure that out. If you're also experimenting with making small models useful, please drop me a line. Footnotes Amazon and Netflix are the notable exceptions ↩ Peter is also being modest here. He's sharp as a tack. ↩