메뉴
HN
Hacker News 22일 전

GLM 5.2와 다가오는 AI 수익률 붕괴

IMP
8/10
핵심 요약

최근 공개된 Z.ai의 GLM 5.2는 오픈소스 가중치(open weights) 모델임에도 최상위 상용 모델들과 필적하는 성능을 보여줍니다. 하지만 AI 비즈니스의 핵심은 막대한 선투자를 바탕으로 한 높은 추론(inference) 수익률인데, 이와 같은 강력한 대안 모델의 등장은 향후 AI 시장의 수익률(margin)을 압박하는 결정적 계기가 될 것입니다.

번역된 본문

이 글은 AI 경제학에서 아마도 가장 이해가 부족한, 다가오는 변화에 초점을 맞춘 2부작 시리즈입니다. 이 글이 마음에 드셨고 두 번째 글의 알림을 받고 싶으시다면 제 뉴스레터를 구독해 주세요.

진짜 DeepSeek 충격이 다가오고 있다 마치 수십 년 전처럼 느껴질 법한 일입니다만, 시장은 DeepSeek의 R1 모델에 큰 충격을 받았습니다. 기저가 되는 V3 모델의 학습 비용이 600만 달러 미만이라는 보도가 나오면서, 시장은 모델 학습을 위한 막대한 자본적 지출(Capex) 투자가 이제 끝났다고 판단했고, 그 결과 엔비디아 등의 주가가 하룻밤 사이에 폭락했습니다. 물론 이는 AI의 실제 비용 구조가 어디에 있는지에 대한 매우 잘못된 읽기였습니다. 모델 학습은 의심할 여지 없이 막대한 자본이 투입되지만, 고정적이고 선행되는 비용입니다. 수억 달러를 들여 모델을 학습시킨 후에는 학습이 "완료"됩니다. [1]

반면 추론(Inference)은 수요에 비례하여 증가합니다. 여기에는 실질적인 한계 비용(marginal cost)이 존재합니다. 지난 1년 동안 저는 이에 대해 꽤 많이 글을 써왔습니다. 다시 한번 강조하지만, API 제공업체가 부과하는 비용이 그들의 실제 비용이라는 대중의 이해는 틀렸습니다. 실제로 Anthropic이나 OpenAI가 추론을 위해 MTok(백만 토큰)당 25달러를 청구할 때, 제 대략적인 계산에 따르면 컴퓨팅 원가 대비 약 90%의 매출 총이익률(gross margin)을 기록하는 것으로 보입니다. 이보다 높거나 낮을 수는 있습니다(유출된 OpenAI의 재무제표에 따르면 매출 대비 약 60%의 매출 총이익률을 보이지만, 여기에는 지원, 결제 처리 및 기타 서비스 비용이 포함되어 있을 것입니다). 하지만 최첨단 AI 연구소들의 전체 비즈니스 모델은 한마디로 모델 학습과 컴퓨팅 및 인건비에 막대한 자금을 지출한 뒤, 매우 수익성 높은 추론 과정을 통해 해당 비용을 상각하는 것입니다. 충분히 많은 추론을 통해 이 비용을 상각할 수 있다면, 단순 매출원가(COGS) 기준의 흑자에서 실질적인 영업 이익으로 넘어갈 수 있습니다.

GLM 5.2 지난 몇 주 동안 저는 Z.ai의 GLM 5.2를 테스트해 보았습니다. 저는 GLM 5.2가 Claude Opus 및 GPT(이 글을 작성하는 시점의 최신 버전은 5.5이며, 물론 미래의 모델은 이를 능가할 것입니다)와 경쟁할 수 있는 진정한 '오픈 웨이트(Open Weights, 모델 가중치 공개)' 경쟁자의 수준에 도달한 첫 번째 모델이라고 생각합니다. 이 모델은 정말 훌륭하며, 제가 매일 사용하는 Opus와 구분하기 어려울 정도입니다. 하지만 생각하는 과정이 길어져 속도가 다소 느립니다. 백그라운드에서 PR(Pull Request)을 검토하는 것과 같이 시간이 중요하지 않은 비대화형 에이전트 작업에는 문제가 되지 않지만, 대화형으로 사용하기에는 확실히 제 주의를 유지하기엔 약간 느립니다. 이는 모델의 비용 효율성을 다소 떨어뜨리기도 합니다(더 많은 생각 = 더 많은 토큰 = 비용 증가). 또한 시각(Vision) 기능을 지원하지 않습니다. 과거에는 시각 인식 기능의 부정확성 때문에 거의 사용하고 싶지 않았지만, Opus 4.7이 훨씬 더 높은 해상도의 시각 능력을 도입한 이후로 지금은 항상 사용하고 있습니다. 따라서 이미지 기반 PDF, 스크린샷 및 디자인 파일을 읽을 수 없다는 것은 정말 답답한 일입니다. 물론 그들 역시 더 강력한 멀티모달 모델을 준비 중이겠지만, 현재로서는 최첨단 모델들에 비해 분명한 약점입니다.

두 번째 약점은, 개인적으로 전혀 예상하지 못했던 부분이기도 한데, 웹 검색 기능의 부재 혹은 성능 저하입니다. 알고 보니 거의 모든 에이전트 세션이 정보를 찾기 위해 아주 많은 웹 검색을 수행합니다. Z.ai는 웹 검색을 위한 대체 MCP를 제공하지만, 상당히 형편없고 느립니다. 파이어웍스(Fireworks)는 어떠한 기능도 제공하지 않으며, 항상 제품 개선을 모색하고 있다는 모호한 답변만 주었습니다. 개인적으로는 당분간 계획이 없는 것으로 보지만, 지켜봐야 할 것 같습니다. 저는 에이전트에게 ddgr과 같은 CLI 기반 웹 검색을 사용하라고 지시해 어느 정도 우회해서 해결했습니다만, 이는 현재 진짜 약점입니다. 저는 서드파티 웹 검색 API의 잠재력을 매우 긍정적으로 봅니다. 이는 사실 오픈 웨이트 모델 제공업체들이 제공할 수 있는 역량에 있어 아주 큰 공백이며, 훌륭한 웹 검색 기능이 다양한 에이전트 작업에 필수적이라는 사실이 확인되었습니다. 그럼에도 불구하고 시간이 지나면 분명 해결될 것입니다. 이를 해결할...

원문 보기
원문 보기 (영어)
This is a two part series focusing on what I believe is perhaps the least understood upcoming shift in AI economics. If you've enjoyed this and want to be notified about the second post, please feel free to sign up for my newsletter . The real DeepSeek moment is upon us What feels like decades ago, markets recoiled at DeepSeek's R1 model. The theory being that given the underlying V3 model reportedly cost under $6m to train, the market therefore thought the huge investment in capex for model training was over, and thus the stock price of Nvidia et al collapsed overnight . Of course, this was a hugely poor read of where the costs actually lie in AI. Training - while no doubt capex intensive - is a fixed, up-front cost. You spend hundreds of millions to train a model, then you are "done". [1] Inference, on the other hand, scales with your demand. It has genuine marginal costs. I've written about this at length over the past year or so. Again, the mainstream understanding of this - that the API costs the providers charge are their real costs is mistaken. Indeed, when Anthropic/OpenAI charge $25/MTok for inference, my napkin maths suggests that this is probably something like 90% gross margin on the cost of compute vs the rack rate. It may be a bit higher, or a bit lower (OpenAI's leaked financials suggest a ~60% gross margin on revenue, but this no doubt includes a lot of other costs like support, payment processing and other services they offer), but the whole business model of frontier AI labs is in short to spend a large amount of money on salaries on compute to train a model, then amortise that cost over a lot of very profitable inference. If you can amortise that cost over enough inference you turn from profitable on a COGS basis to... actually profitable. GLM 5.2 I have been playing around with GLM5.2 from Z.ai for the last couple of weeks. I believe GLM5.2 is the first model that reaches the "bar" of a genuine open weights competitor to Opus and GPT (at the time of writing, the latest version of GPT was 5.5 - future models no doubt will exceed this). It's genuinely very good and hard for me to tell the difference between Opus - my daily driver and it. I've found that it is slow because of the amount of thinking it tends to do. For non interactive agentic tasks (like reviewing PRs in the background) which aren't time critical this is a non issue, but for interactive use it is definitely a tad too slow to keep my attention. This also somewhat reduces the cost effectiveness of it (more thinking means more tokens, which increases costs). It also doesn't have vision support. It's funny how quickly I've gone from basically never wanting to use vision (because it was so inaccurate, I'd often pause sessions when I caught it using vision), to using it all the time - since Opus 4.7 introduced far higher resolution vision capabilities. It's genuinely frustrating it not being able to read image-based PDFs, screenshots and design files. I'm sure they have a more multimodal model in the works, but this is a significant weakness against the frontier labs. Secondly, and something I really didn't expect to be a blocker, is the lack of/poor web search capabilities. It turns out that nearly every agentic session does a lot of web searching for looking up items. Z.ai provides a replacement MCP for web search, but it's pretty awful and slow. Fireworks doesn't provide any, though they gave me a very vague answer saying they are always looking to improve products. I would take that as no plans personally, but let's see. I've managed to somewhat work around this by telling the agent to use a CLI based web search like ddgr , but this is a real weakness right now. I am very bullish on the potential of 3rd party web search APIs. This is actually a huge gap in what open weights model providers can offer, and it turns out great web search capabilities are essential for many agentic tasks. Regardless, this no doubt will be solved with time - there are many people building web search indexes and it just requires the right partnerships and plumbing in place. Drop in replacement Where it gets really scary for the frontier labs is how easy it is to migrate to open weights models. Both Z.ai and Fireworks offer both an OpenAI compatible and Anthropic compatible endpoint. This makes it absolutely trivial to use with Claude Code and Codex. You just set the base URL to point to your inference provider, give it the API key and tell it to use GLM5.2. Given Anthropic recently announced (then backtracked) on charging API rates for claude -p non interactive agentic use, you will find for many/most of those use cases you can just drop in GLM instead. And for interactive use, apart from the lack of vision and slow(er) speed [2] , it was genuinely almost impossible for me to realise I wasn't using Opus in Claude Code. This is not Microsoft or Salesforce like lock in, where you need to spend years planning a migration. The switching costs are incredibly low, and I would argue that are actually far less than trying to keep up on all the policy and term changes that the frontier lab models tend to scramble around with. It's possible that Claude Code will make it harder to use 3rd party providers, but there are many good open source options (like Codex itself and OpenCode, amongst dozens). One concern I do hear from enterprise is data privacy and security. There is no doubt that using Z.ai's official API and subscription is almost certainly a non-starter, with their terms being at best weak and the deep connection to Mainland China. But of course, with open weights being open there are many other providers in the market, many with proper contractual provisions. And, if that isn't enough, you can of course host in on premises yourself, which actually opens up even more sensitive data - that couldn't be sent to any third party - to Opus-quality agentic workflows. Cost savings The going rate for GLM5.2 seems to be around the $4.40/MTok mark. This is less than 20% of the retail price of Opus and ~15% the cost of GPT5.5. Now, given it does use more tokens for a given task, this isn't a totally apples to apples comparison. But I'd be very surprised if it wasn't more than 50% cheaper for nearly all workflows, for a very similar level of quality. In terms of subscriptions, Z.ai offers a "coding plan" subscription which mirrors the plans you'd see from Anthropic and OpenAI, but with a higher claimed usage limit. I expect for most professional use the very lax terms around training and data retention will make this a difficult sell, but if the frontier labs were to try and increase pricing substantially I can see it being a credible option for those that are budget-conscious. I expect these costs for GLM5.2 to come down significantly over the coming months as well, as more optimisation is done to the serving stack(s). Wafer wrote an interesting write up of their efforts to run it on AMD hardware. They suggest that it is 2.75x cheaper per token to run inference on AMD vs Nvidia Blackwell. Part two is where this gets interesting - what a collapse in inference margins actually does to the industry, and who is likely to win and lose. I'd keep Bezos's famous "your margin is my opportunity" line in mind. If you'd like me to drop it in your inbox the moment it's out, sign up to the newsletter - or grab the RSS feed if that's more your thing. Disclosure - Fireworks kindly gave me some free credit to experiment with GLM to help write this article. This is a simplification - the frontier labs are effectively training new models constantly to stay competitive, so it's really a rolling cost rather than a true one-off. The key distinction still holds though: unlike inference, that cost doesn't scale with how much customers actually use the product. ↩︎ To be fair, the slowness is mostly the model thinking a lot rather than the serving itself - Fireworks launched GLM5.2 at genuinely quick tokens/sec, which was a huge i