메뉴
HN
Hacker News 9일 전

중국 AI 모델, 그리고 돌아온 한계비용

IMP
8/10
핵심 요약

최근 중국의 오픈웨이트(open weights) AI 모델인 Kimi K3가 최고 수준의 성능에 근접하면서 AI 산업의 경제적 구조가 변화하고 있습니다. 과거 소프트웨어 산업의 핵심이었던 '제로(0) 한계비용' 시대는 끝났으며, 이제 AI 추론(inference)에 수반되는 매출원가(COGS)가 실질적인 비용 지표로 작용하게 되었습니다. 엔비디아의 '토큰(Token)' 공장 관점과 함께, 향후 AI 비즈니스의 올바른 단가 산정 및 산업 구조를 이해하기 위해 필수적인 분석입니다.

번역된 본문

원문 제목: 중국 모델이 두려운가? (Who's Afraid of Chinese Models?) 2026년 7월 20일, 월요일

내가 켈로그 경영대학원(Kellogg School of Management)의 STRT-431 수업, 즉 모든 1학년 MBA 학생들이 의무적으로 들어야 하는 기초 과목 첫 날에 겪었던 일화가 있다. 나는 배부된 리딩 자료와 사례 연구를 넘겨보다가 일정표에 기술 기업이 단 하나도 없다는 사실에 당황했다. 평소 나답게 수업이 끝난 후 교수님께 찾아가 그 이유를 물었고, 교수님은 이 과정의 목표가 특정 산업을 배우는 것이 아니라 어떤 산업의 어떤 기업에든 적용할 수 있는 보편적인 원리를 깨닫는 것이라고 대답해 주셨다. 내가 흔히 하듯 이야기하자면, 나는 이 대답에 전혀 만족하지 못했다. 내게 있어 기술의 본질, 특히 소프트웨어와 유통에 있어 한계비용(그리고 거래비용)이 0이라는 사실은 근본적으로 다른 무언가였기 때문이다. 수식에 0을 넣으면 공식이 완전히 붕괴되고 말 것이다!

하지만 나는 곧 그것이 나에게 기회가 될 수 있음을 깨달았다. '집단화 이론(Aggregation Theory)'을 뒷받침하는 핵심 통찰은 한계비용이 0이라는 점이 사람들이 인터넷에 대해 예상했던 것과는 근본적으로 다른 가치 사슬을 만들어낸다는 것이다. 즉 공급을 유통하는 것보다 수요를 통제하는 것이 더 중요해지는 세계에서의 중앙 집중화와 규모의 경제를 의미한다.

하지만 AI가 흥미로운 이유는 이 낡은 보편적 원리들이 얼마나 다시 전면으로 부상하고 있는가에 있다. 이보다 더 명확했던 적은 지난 주말밖에 없었다. 중국의 또 다른 오픈웨이트(Open Weights) 모델인 Kimi K3가 기능과 성능 면에서 최고 수준(SOTA)에 근접함에 따라, X(옛 트위터)에서는 그 의미에 대해 격렬한 논쟁이 벌어졌다. 결론부터 말하자면 이렇다. 한계비용이 다시 크게 부활했다는 것이다. 최고 수준의 무료 모델이 가져오는 단기적인 영향과 산업의 장기적 구조 양 측면 모두에서 말이다.

매출원가(COGS) 대 연구개발(R&D) 오픈웨이트 모델에 대한 논의를 뒷받침하는 가장 흔한 오해 중 하나는 그것이 더 저렴하다는 것, 심지어 무료라는 것이다. 당연히 가중치(weights)를 다운로드하여 자체 모델을 만드는 데 필요한 시간과 비용, 역량을 건너뛸 수 있으니 말이다. 물론 이는 사실이지만, 이 경우 '무료'라는 말은 당신이 연구개발(R&D)에 지출해야 하는 금액을 의미한다. R&D는 당신이 창출하는 수익과 관계없는 고정 비용이다. R&D에 100만 달러를 지출했다면, 10만 달러의 수익을 내든 1억 달러의 수익을 내든 상관없이 당신은 여전히 R&D에 100만 달러를 지출한 것이다. (물론 수익성에는 영향을 미친다.)

수익과 관련된 것은 COGS, 즉 매출원가(Cost of Goods Sold)이며, AI에 있어서 COGS는 오랫동안 소프트웨어가 경험하지 못했던 방식으로 현실적으로 존재한다. 구체적으로 말해, Kimi든 Fable이든 모델을 실행하고 추론(inference)하는 데는 비용이 든다. 그리고 AI 제공자가 추론에 지출하는 비용은 적어도 대부분의 비즈니스 모델에서 수익과 직접적인 비례 관계에 있다. 위의 예시를 다시 사용해 보자면, 10만 달러가 아닌 1억 달러의 수익을 내려면 아마도 1,000배의 COGS가 필요할 것이다.

실제 상황으로 예를 들자면, 1달러의 수익을 창출하는 토큰을 생성하는 데 50센트가 든다면, 1억 달러의 수익에는 5천만 달러의 COGS가 발생할 것이다. 반면 10만 달러의 수익은 단 5만 달러의 COGS만을 발생시킬 뿐이다. 오픈웨이트 모델과 관련하여 핵심은 서비스를 제공하는 데 있어 결코 무료가 아니라는 점이다. Kimi K3는 백만 입력 토큰당 3달러, 백만 출력 토큰당 15달러의 비용이 든다. 이는 백만 입력 토큰당 5달러, 백만 출력 토큰당 30달러를 청구하는 Sol보다 저렴하지만, 이것조차 올바른 측정 기준이 아닐 수 있다.

토큰(Tokens) 대 지능(Intelligence) 엔비디아(Nvidia)의 젠슨 황(Jensen Huang) CEO는 엔비디아가 구축하는 것을 '토큰 공장(Token Factories)'이라고 묘사했으며, 엔비디아의 관점에서 이러한 프레이밍은 타당하다. 엔비디아 GPU는 모델에 구애받지 않는다. 그들은 가장 빠르고 효율적인 방식으로 토큰을 생성한다. 이로 인해 초당 토큰 수, 첫 토큰까지 걸리는 시간(time-to-first-token), 와트당 토큰 수, 토큰 비용 등과 같은 측정 항목이 파생되었으며, 황 CEO는 이러한 지표들이 의사 결정의 기준이 될 것이라고 주장한다. 이러한 관점은 분명 첫 단계 동안에는 타당했다.

원문 보기
원문 보기 (영어)
Who's Afraid of Chinese Models? Monday, July 20, 2026 Listen to Podcast Listen to this post : Log in to listen There's a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I leafed through the readings and case studies and was dismayed that there weren't any tech companies on the docket. Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries, but rather to uncover broadly applicable universal principles that could be applied to any company in any industry. I did not, as I usually tell the story, find this very satisfactory: to me the nature of tech, particularly the fact that software and distribution had zero marginal costs (and zero transaction costs), was something fundamentally different; putting in zeroes in formulas tends to wreak havoc! I soon realized, however, that that was my opportunity. The fundamental insight undergirding Aggregation Theory is that zero marginal costs leads to fundamentally different value chains than people once expected from the Internet: centralization and scale in a world where controlling demand mattered more than distributing supply. What is fascinating about AI, however, is the extent to which those old universal principles are coming back to the forefront. That was never more apparent than this past weekend, when arguments raged on X about the implications of Kimi K3, another open weights model out of China, approaching the state-of-the-art in terms of capabilities. The long and short of it is this: marginal costs are back in a big way, both in terms of short-term implications of state-of-the-art free models, and in terms of the long-term structure of the industry. COGS Versus R&D One of the most common misconceptions undergirding discussion of open weights models is that they are cheaper — free, even. After all, you can just download the weights, and skip the time and expense and capabilities necessary to create your own model. That is, of course, true, but the "free" in this case is a reference to the amount you need to spend on research and development; R&D is a fixed expense that is independent of the revenue you generate. If you spend $1 million in R&D, it doesn't matter if you do $100 thousand in revenue or $100 million; you still spent $1 million on R&D (it does, of course, impact your profitability). What is related to revenue is COGS — cost of goods sold — and COGS is real for AI in a way it hasn't been for software for a very long time. Specifically, running inference on a model — whether that model be Kimi or Fable — costs money, and the amount of money an AI provider spends on inference is, at least in most business models, directly correlated to revenue. To reuse the above example, generating $100 million versus $100 thousand in revenue will likely require 1,000x COGS. In concrete terms, if it costs 50 cents to generate the tokens that drive $1 in revenue, then $100 million in revenue will have $50 million in COGS; $100 thousand in revenue will only have $50 thousand in COGS. The point in terms of open weight models is that they are not free to serve. Kimi K3 costs $3 per million input tokens, and $15 per million output tokens; that is cheaper than Sol's $5 per million input tokens and $30 per million output tokens, but that might not even be the right measurement. Tokens Versus Intelligence Nvidia CEO Jensen Huang has described what Nvidia is building as "token factories", and from Nvidia's perspective that framing makes sense. Nvidia GPUs are model agnostic: they generate tokens, and do so in the fastest and most efficient way possible. That leads to measurements like tokens-per-second, time-to-first-token, tokens-per-watt, token cost, etc., and Huang argues that these metrics will be the basis for decision-making. This is a framing that definitely made sense during the first paradigm of AI, the ChatGPT era, when tokens were delivered straight to the end user. The second paradigm of AI, however, the reasoning era, confounds this measurement. Reasoning entails an explosion in chain-of-thought tokens, and different models need different amounts of reasoning tokens to arrive at the right answer. Kimi, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot. Agents introduce a similar dynamic: some models are more efficient than others in terms of the number of tokens they need to execute agentic workflows. What this means is that tokens are not a commodity. The defining characteristic of a commodity is that it is fungible: a gallon of oil is a gallon of oil; a ton of copper is a ton of copper; a bushel of wheat is a bushel of wheat. A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence. In other words, if both Kimi and Sol generated the right answer, then that answer is fungible; the difference in tokens generated to get to that right answer is a contributor to a difference in COGS. The COGS for intelligence is a function of a few different factors: Model footprint: The weights and runtime state determine how much expensive memory and how many accelerators are required to host each serving replica. Inference efficiency: Architectural choices (e.g. Mixture-of-Experts) reduce computation per generated token. Memory efficiency: Architectural choices can reduce KV cache requirements, allowing more concurrent requests and better GPU utilization. Serving efficiency: Batching, scheduling, prefix caching, and other inference optimizations maximize utilization and share work across requests. Token efficiency: The fewer tokens required to reach a correct answer, the lower the inference cost. The reason this matters is that we are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic CRUD app , for example, can likely do so using models from multiple providers. And, in a commodity market, the route to profitability is not through charging higher prices — again, you can (or will soon be able to) make the exact same app using multiple models — but rather through having a superior cost structure. Understanding Commodity Markets It's worth stepping through the mechanics here, because, as I noted a few months ago in Amazon's Durability , the dynamics of commodity markets are not something people in tech are generally familiar with: In commodity markets, everyone charges the same price, because everyone is selling the same thing; that price is determined by supply and demand. The demand for a commodity is a function of price elasticity: the cheaper the commodity, the more demand there is for it, and vice-versa. The supply for a commodity is a function of the marginal cost of producing the commodity. The key thing to understand is that the marginal cost of producing the commodity differs by supplier. What this means in practice is that the supplier with the worst cost structure ends up selling the commodity at their marginal cost (if they can produce at all); the profits of everyone else depend on the extent to which their cost structure is better than the marginal supplier. As an example: Supplier A can produce 10 units of the commodity for $10 each Supplier B can produce 10 units of the commodity for $15 each Supplier C can produce 10 units of the commodity for $20 each Let's assume the price elasticity is such that there is demand for 25 units of the commodity at $20. That means: Supplier A will sell 10 units of the commodity for $20, earning $10/unit Supplier B will sell 10 units of the commodity for $20, earning $5/unit Supplier C will sell 5 units of the commodity for $20, earning $0/unit This isn't precisely right: the reason why Supplier C will bear the shortfall is because Suppl