메뉴
HN
Hacker News 33일 전

현재 LLM 비용이 지속 불가능한 이유

IMP
8/10
핵심 요약

기업들의 천문학적인 AI 비용 부담이 가중되고 있으나, 모델 성능의 한계 도달, 오픈소스 모델의 약진, 전용 칩셋 발전, 그리고 교체 비용 제로(0)라는 시장 특성으로 인해 AI 모델의 고비용 구조는 곧 붕괴할 수밖에 없습니다. 이는 실무자와 기업들이 비용 효율적인 대체 모델으로 빠르게 전환할 수 있는 근거가 되며, 향후 AI 인프라 도입 전략 수립에 매우 중요한 시사점을 던집니다.

번역된 본문

AI는 비용 문제를 안고 있습니다. 앞으로 등장할 해결책은 우리가 예상하는 것보다 더 단순할 것입니다. (Aditya Patadia, 2026년 6월 25일)

많은 기업들이 높은 AI 비용으로 고통받고 있습니다. 우버(Uber)는 단 4개월 만에 1년 치 AI 예산을 모두 소진했으며, 마이크로소프트, 세일즈포스, 깃허브(Github)는 직원들의 AI 지출을 줄이기 위한 조치를 취하고 있습니다. 반면에 AI는 많은 프로그래밍 작업을 매우 쉽게 만들고, 데이터 해석, 멋진 슬라이드 제작, 앱 및 웹사이트 디자인 등 다른 분야에서도 계속해서 큰 도움을 주고 있습니다. 현재 주요 AI 연구소들은 이른바 '프론티어 모델(Frontier models)'을 보유하고 있으며, 이 모델들은 다양한 작업에서 탁월한 성능을 발휘합니다. 프론티어 AI 연구소들은 연구와 호스팅을 모두 자체적으로 수행하므로 해당 모델의 비용이 가장 높습니다. 예를 들어, GPT 5.5는 백만 입력 토큰당 5달러, 백만 출력 토큰당 30달러의 비용이 듭니다. 이는 현재 OpenRouter 기준으로 사용 가능한 가장 비싼 모델입니다. 예를 들어 오늘 오후, 이 모델을 사용하여 50개 파일에서 타입스크립트(TypeScript) 타입 수정만 한 번 진행했을 뿐인데 54달러가 지출되었습니다.

모델 성능 정체, 오픈 웨이트(Open weight) 모델 출시, 칩셋 및 모델의 발전, 교체 비용 제로(0), 그리고 로컬 모델(내부망 모델) 등은 AI 연구소들이 현재 요구하는 높은 가격을 유지할 수 없게 만드는 이유들입니다.

모델 성능 정체 요즘 모델이 출시될 때마다 성능 향상이 있지만, 그 향상 폭은 갈수록 줄어들고 있다는 것이 분명합니다. 완전히 새로운 돌파구가 발명되지 않는 한, 현재의 학습 및 추론 능력은 한계까지만 성장할 수 있습니다. 학습 데이터 문제도 있습니다. 대부분의 AI 연구소는 모델 학습을 위해 디지털 및 인쇄 매체에서 사용 가능한 모든 데이터를 이미 소비했을 가능성이 높습니다. 따라서 학습 데이터셋을 개선하는 것은 매우 어려울 것입니다. 이는 더 나은 성능을 이유로 모델 가격을 계속 인상하는 추세가 쉽지 않을 것임을 의미합니다. Claude Opus 4.8의 가격이 Claude Opus 4.7과 동일하게 책정된 것에서 이미 그 증거를 보았습니다. 모델의 발전이 크게 둔화되고 학습 데이터와 방법이 서로 비슷해지면, 경쟁으로 인해 모델 가격은 하락할 것입니다.

오픈 웨이트 모델 OpenAI가 2022년 ChatGPT를 출시했을 때 압도적인 우위를 점했지만, 그 격차는 서서히 줄어들고 있었고 우리는 2025-26년에 Anthropic이 정상을 차지하는 것을 보았습니다. 이제 GLM-5.2와 같은 오픈 웨이트 모델이 코딩 벤치마크에서 GPT와 Opus를 이기고 있습니다. 해당 모델의 비용은 GPT 5.5에 비해 10분의 1 수준입니다. 여기서 일어나는 일은, 최고 수준의 AI 연구소들이 단순히 추론 비용뿐만 아니라 모델 아키텍처 연구, 학습 데이터 수집 및 가공, 모델 학습 비용(수천만 달러에서 수억 달러에 달할 수 있음), 직원 급여 및 마케팅 비용 회수 등을 모두 가격에 포함시켜 청구하고 있다는 것입니다. 반면에, 오픈 웨이트 모델이 출시되고 나면 어떤 추론 제공업체든 쉽게 이를 호스팅하고 추론 비용에 약간의 마진만 붙여 서비스할 수 있습니다. 이는 프론티어 AI 연구소를 운영하는 것보다 훨씬 저렴하다는 것이 증명되었습니다.

칩셋 및 모델의 발전 Cerebras, Groq, Google 등 많은 기업들은 AI에 맞는 전용 실리콘이 필요하며 일반적인 GPU로는 한계가 있다는 것을 깨달았습니다. 특수 설계된 칩을 개발하는 것은 매우 비싸지만, 일단 아키텍처가 준비되고 나면 수백만 개를 쉽게 제작할 수 있어 추론 비용이 훨씬 저렴해집니다. 예를 들어 TPU는 엔비디아(Nvidia) H100 GPU보다 30~70% 더 저렴할 수 있습니다. 이러한 기술 발전은 계속해서 등장할 것이며 토큰당 가격을 계속 떨어뜨릴 것입니다. 모델 아키텍처 또한 진화하고 있습니다. 우리는 캐싱을 기본적인 개선 사항으로 보았고, 이제는 전문가 혼합(MoE) 모델과 다른 접근 방식을 통해 모델의 정확도는 유지하면서 속도는 더 빠르게 만들고 있습니다.

교체 비용 제로(0) Windows OS, MS Office, Adobe 제품군과 같은 전통적인 소프트웨어나 Salesforce, Hubspot, Figma와 같은 SaaS 제품에는 AI 모델에는 없는 매우 중요한 해자(진입 장벽)가 있었습니다. 개발된 모든 소프트웨어는 서로 교체할 수 없는 특성이 있었습니다. CRM(고객 관계 관리) 시스템을 오후 반나절 만에 교체할 수는 없었고, 수개월이 걸렸습니다. 더 많은 AI 연구소가 이 시장에 진입하고 더 많은 오픈 웨이트 모델을 사용할 수 있게 되면, 이 요인(교체 비용이 없다는 특성)으로 인해 AI 모델의 가격이 매우 빠르게 폭락하는 현상이 발생할 것입니다.

원문 보기
원문 보기 (영어)
AI and Cloud Costs AI has a cost problem. The solution that will emerge will be simpler than we expect. Aditya Patadia Jun 25, 2026 Share A lot of companies are getting bitten by high AI costs. Uber burned through the entire year’s AI budget in just 4 months and Microsoft, Salesforce and Github are taking steps to reduce AI spend by employees. On the other hand, AI is making many programming tasks very easy and also keeps helping in other domains like data interpretation, making beautiful slides and designing apps and websites. Currently, big AI labs have what we call frontier models and those models perform exceptionally well for a wide variety of tasks. Frontier AI labs are doing research and hosting both on their own and hence, the costs of those models are the highest. GPT 5.5, for example, costs $5 per million input tokens and $30 per million output tokens. This is currently the costliest model available as per OpenRouter . To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon. Model performance plateau, Open weight model releases, Chip and model improvements, Zero switching costs and local models are the reasons the AI labs might not be able to sustain the high price that they are asking right now. Model performance plateau We are seeing improvements with each model release these days but it’s clear that the improvements are getting smaller and smaller. Unless a completely new breakthrough is invented, current learning and inference capabilities can only scale so much. There is a problem of training data as well. Most AI labs have likely ingested everything available in digital and print media for the model training. Improving the training dataset is going to prove very difficult. This means the continuing trend of hikes in model price due to better performance is not going to be easy. We saw evidence of it where Claude Opus 4.8 costs the same as Claude Opus 4.7. Once models stop improving big time and the training data and methods are similar, the model prices will likely drop due to competition. Open weight models OpenAI had a massive lead when they launched ChatGPT in 2022 but slowly that lead is fading and we saw Anthropic take top spot in 2025-26. Now models like GLM-5.2 which is an open-weight model, beat GPT and Opus in coding benchmarks. That model has a 1/10th cost compared to GPT 5.5. What is happening here is that leading AI labs are charging not only for inference but also for research in model architecture, training data collection and curation, model training cost (which can be tens or even hundreds of millions of dollars), paying their employees and recovering the marketing costs. On the other hand, once an open weight model is released, any inference provider can easily host it and just do some markup on inference cost. This proves way cheaper than running a frontier AI lab. Chip and model improvements Companies like Cerebras, Groq, Google and many other companies have realised that AI needs its own silicon and normal GPUs are not cutting it. Specialised chips are very expensive to design but once the architecture is ready, making millions of them is easy and inference cost becomes much cheaper. A TPU for example can be 30-70% cheaper than an Nvidia H100 GPU. Such advancements will keep coming and keep dropping the price per token. Model architecture is also evolving. We saw caching as a basic improvement and now MoE models and other approaches are making models faster while keeping the same accuracy levels. Zero switching costs Traditional Software like Windows OS, MS Office, Adobe Suite and SaaS like Salesforce, Hubspot, and Figma had a very important moat that AI models don’t have. Every single software that was built was not interchangeable. You could not swap a CRM in an afternoon; it took months. When more AI labs enter the space and more open weight models are available, this factor is going to be responsible for a very quick price crash. AI gateway providers like OpenRouter.ai are making it extremely easy to switch models. It can happen in seconds and in fact, we can program it to change providers on the fly. Zero switching costs mean that if a better model comes along, consumers can switch to it without any time investment. Local models Last but not least and in fact the most important factor, is the ability of users to run local models. So far, almost everyone is using cloud-hosted models and local models are either too big to deploy or too slow to work with. With advancements in chips, this will change in 4-5 years’ time. Newer chips will run models locally and almost certain crash in RAM prices will make it easy to deploy models on computers and smartphones. I predict most operating systems will provide a way to deploy a model and they will also provide an interface so apps running locally can connect to the model. When this happens, cloud models will only be used for the most complex of the tasks and simple tasks like code tab completion, proofreading and fact checking will be done locally. This means customers will no longer need that $20 or $200 subscription. Closing thoughts This is my first blog on a personal level and I have made some bold predictions here. Only time will tell how they turn out but one thing is certain. The price pressure will come due to one or more reasons listed above and in the end, it’s all good for consumers. Thanks for reading Founder's Notes! Subscribe for free to receive new posts and support my work. Subscribe Share