메뉴
BL
TechCrunch AI 49일 전

AI 기업, 저렴한 소형 모델로 눈 돌리다

IMP
8/10
핵심 요약

기존의 '모델이 클수록 좋다'는 AI 업계의 통념이 깨지고, 기업들은 비용 압박으로 인해 소형 모델로 눈을 돌리고 있습니다. 코인베이스 공동 창업자는 80%의 작업이 99% 저렴한 모델로 대체될 것이라 예측했으며, 실제 테스트에서도 품질 손실 없이 비용을 크게 절감한 사례가 확인되었습니다. 이는 AI 경제학을 근본적으로 바꾸고, 막대한 자본을 투자받은 대형 AI 연구소들에게 큰 재정적 타격을 줄 수 있는 중요한 변화입니다.

번역된 본문

AI 붐은 '모델이 클수록 더 강력하며, 가장 강력한 모델이 승리한다'는 기본적인 가정을 바탕으로 구축되었습니다. 이제 산업계는 이 가정이 무너지기 시작할 때 어떤 일이 발생하는지 알아가려 하고 있습니다. 증가하는 비용은 이미 사용자들이 더 작고 저렴한 모델을 재고려하게 만들었습니다. 비용을 의식하는 이러한 모델 선택은 새로운 현상이며 산업에 어떤 영향을 미칠지 불분명하지만, 그 파급력은 상당할 것입니다. 코인베이스 공동 창립자 브라이언 암스트롱(Brian Armstrong)이 가장 잘 설명한 한 가지 예측은, 작업의 대부분이 더 저렴한 모델로 이동하게 될 것이라는 점입니다. 암스트롱은 X(구 트위터)에 "지능에 대한 수요는 거의 무한하지만, 12~18개월 내에 80%의 작업량은 99% 더 저렴한 모델에서 실행될 것"이라고 적었습니다. "나머지 20%의 작업량은 최고 수준의 지능(IQ)이 중요한 최신 세대 모델에서 계속 실행될 것입니다." 암스트롱의 예측이 사실이 된다면 이것이 AI 산업에 얼마나 큰 변화를 가져올지 결코 과장될 수 없습니다. 그동안 대부분의 AI 기업들은 품질로 경쟁해 왔으며, 이는 가장 진보된 사용 가능한 모델을 기본으로 선택한다는 것을 의미했습니다. 만약 동일한 작업을 품질에 영향을 주지 않고 더 저렴한 모델로 처리할 수 있다면, AI 경제학에 있어 엄청난 패러다임의 전환이 될 것입니다. 그리고 결정적으로, 이렇게 절약되는 비용의 상당수가 대형 연구소들의 주머니에서 나오게 되어, 기업공개(IPO)를 앞둔 OpenAI와 Anthropic에게 재정적인 타격을 입히게 될 것입니다. 이는 '기업들이 소형 모델로 전환할 준비가 되었는가?'라는 하나의 기본적인 질문에 달려 있는 산업계의 잠재적인 지각변동입니다. 초기 테스트에 따르면 시스템이 올바르게 구성되었을 때, 더 저렴한 모델이 품질 저하 없이 투입될 수 있음을 시사합니다. 법률 AI 도구인 하비(Harvey)의 최근 테스트에서는 품질을 떨어뜨리지 않고 추론(Inference) 비용을 3배나 줄일 수 있었습니다. 추론 플랫폼인 파이어웍스 AI(Fireworks AI)와 공동으로 진행된 이 테스트는 클로드 오푸스(Claude Opus)와 파이어웍스의 GLM 5.1을 결합하고, 가장 고강도의 작업을 오푸스에 맡기는 방식으로 진행되었습니다. 그 결과 서버 시간과 전체 비용 측면에서 유의미하게 부하가 줄었습니다. 하비의 공동 창립자 게이브 페레이라(Gabe Pereyra)는 자신의 스타트업이 제공하는 AI 법률 서비스와 관련하여 테크크런치(TechCrunch)에 "품질이 최우선이며, 법률 분야에서는 항상 그럴 것"이라고 말했습니다. "하지만 품질의 정의가 모든 것에 가장 강력한 모델을 사용하는 것에서, 올바른 답을 가장 효율적으로 얻어내는 최적의 모델을 사용하는 것으로 진화하고 있습니다." 이러한 추세는 주로 대형 연구소와 중국산 모델 또는 오픈 웨이트(Open-weight) 모델 간의 대결로 묘사되곤 하지만, 이는 본질을 놓친 것입니다. 진정한 구분은 독점 모델과 오픈 소스 모델 간의 차이가 아니라 대형 모델과 소형 모델 간의 차이입니다. GPT-5.5에서 딥시크(DeepSeek)의 V4 Flash로 전환하여 비용을 절약할 수도 있지만, GPT-5.4-mini로 전환해도 마찬가지의 효과를 얻을 수 있습니다. 현재 대형 연구소의 자체 추론 서비스와 독립적으로 제공되는 오픈 웨이트 모델 간의 치열한 가격 경쟁이 진행되고 있습니다. 소형과 대형 모델 중 무엇이 이기든 상관없이, 소형 모델이라는 더 큰 흐름에는 크게 영향을 미치지 않습니다. 이 모든 것이 너무 당연해 보일 수 있습니다. 불필요하게 많은 컴퓨팅 파워를 사용할 이유가 없다는 점은 자명합니다. 하지만 이는 지금까지 산업을 지배해 온 '스케일링 우선(Scaling-first)' 접근 방식에 위배됩니다. '쓰라린 교훈(Bitter Lesson)'에 영감을 받은 연구소들은 가능한 한 가장 많은 컴퓨팅을 요구하는 모델을 학습시키는 데 전력을 다해왔고, AI 모델의 한계를 계속 밀어붙여 왔습니다. 가격이 투자자들에 의해 많은 부분 보조금 형태로 지원되었기 때문에, 고객들은 가장 진보된 옵션 외에는 선택할 이유가 없었습니다. 하지만 토큰(Token) 가격이 상승하고 보조금 지원이 줄어들면서, 사용자들은 처음으로 비용 압박에 직면하고 있습니다. 이러한 새로운 비용 압박이 실제로 기업 사용자들을 소형 모델로 이끌지는 아직 미지수입니다. 기업들은 API 호출을 줄이거나, 적은 컨텍스트(Context)를 사용하거나, 단순히 가망이 없는 배포를 포기하는 방식으로 비용을 절감할 수도 있습니다. 하지만 대부분의 배포가 소형 모델에서도 똑같이 잘 실행될 수 있다는 것이 밝혀진다면, 증가하는 추론 수요에 찬물을 끼얹을 수 있습니다. 그리고 한계선을 넘나드는 최신 프론티어 모델(Frontier model)의 학습 비용을 어떻게 정당화할 것인지에 대한 새로운 의문을 제기하게 될 것입니다.

원문 보기
원문 보기 (영어)
The AI boom has been built on a basic assumption: bigger models are more powerful, and the most powerful models win. Now, the industry is about to learn what happens if that assumption starts to break. Mounting costs have already pressured users to give smaller and cheaper models a second look. This cost-conscious model-shopping is new and it’s unclear how it will affect the industry, but the impact is likely to be significant. One prediction, laid out best by Coinbase co-founder Brian Armstrong, is that it will result in the vast majority of tasks shifting to cheaper models. "Demand for intelligence is near infinite, but 80% of workloads will be running on 99% cheaper models within 12-18 months,” Armstrong wrote on X . “20% of workloads will still run on latest gen models where IQ maxing is important.” It’s hard to overstate what a significant shift it will be for the AI industry if Armstrong’s prediction comes true. Before now, most AI companies have competed on quality, which has meant defaulting to the most advanced available model. If those same jobs can be handled by cheaper models without affecting quality, it would mean a massive shift in the economics of AI. And critically, much of the savings would be coming out of the pockets of the big labs, dealing a financial blow to OpenAI and Anthropic just as they’re heading for their IPOs. It’s a potentially seismic change in the industry, resting on one basic question: Are companies ready to switch to smaller models? Initial tests suggest that, when the system is arranged right, cheaper models could sub in without any sacrifice in quality. In a recent test by the legal AI tool Harvey, the company was able to reduce inference costs by 3x without reducing quality. The test, performed in partnership with the inference platform Fireworks AI, combined Claude Opus and Fireworks’ GLM 5.1, and shifted to Opus for the most intensive tasks. The result was a significantly lower load in terms of server time and overall cost. "Quality comes first, and in legal it always will," Harvey co-founder Gabe Pereyra told TechCrunch, referring to the AI legal services his startup provides. "However, the definition of quality is evolving from simply using the most powerful model for everything, to using the best model that gets the right answer most efficiently." This trend is often framed in terms of major labs versus Chinese models or open-weight ones, but that misses the bigger point. The real divide isn’t between proprietary and open models; it’s between large models and small ones. You can save money by switching from GPT-5.5 to DeepSeek’s V4 Flash, but switching to GPT-5.4-mini works just as well. There’s an active price war going on between in-house inference from the big labs and independently served open-weight models. For the bigger question of small versus large, it doesn’t really matter which kind of small model wins out. All of this might seem obvious — of course you shouldn’t use more compute than necessary — but it runs counter to the scaling-first approach that has dominated the industry until now. Inspired by the bitter lesson , labs have leaned hard into training the most compute-intensive models possible, pushing the frontier of what AI models can do. With prices heavily subsidized by investors, clients had no reason to choose anything but the most advanced option. With token prices rising and subsidies slowing down, users are facing cost pressure for the first time. We don’t know whether the new cost pressure will actually drive enterprise users to smaller models. They could just as easily economize by making fewer calls, using less context, or simply giving up on the least promising deployments. But if it turns out that most deployments can be run just as well on a smaller model, it could put a serious damper on the growing demand for inference – and raise new questions about how to justify the cost of training a frontier model. Topics AI , ai models , Anthropic , Harvey , OpenAI , TC When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio June 18 Los Angeles Get an inside look at what it takes to scale and succeed from leaders at Mach Industries, Founders Fund, and Shinkei Systems. Through candid fireside chats and high-impact networking, you'll walk away with valuable insights and new connections. REGISTER NOW Most Popular WWDC 2026: Everything announced on Siri AI, iOS 27, Apple Intelligence, and more Morgan Little Aisha Malik Microsoft's open source tools were hacked to steal passwords of AI developers Zack Whittaker Is this the dawn of the Tokenpocalypse? Anthony Ha Founders share VC horror stories, and some are naming names Julie Bort Google will pay SpaceX $920M per month for compute Sean O'Kane Mira Murati steps back into the spotlight, carefully Connie Loizos Ahead of its IPO, Anthropic's Daniela Amodei shrugs off doubts about AI's returns Marina Temkin