메뉴
BL
The Decoder • 30일 전

Z.ai, 최고 수준 성능에 저렴한 'GLM-5.3-Flash' 공개…엔비디아 없이 구동

IMP
8/10
핵심 요약

Z.ai가 3,200억 파라미터와 100만 토큰 컨텍스트 창을 갖춘 GLM-5.3-Flash를 공개했습니다. 이 모델은 지능 지수에서 더 큰 GLM-5.3에 거의 버금가지만 작업당 비용은 약 7.5배 저렴하며, 특히 전적으로 중국산 AI 칩에서 구동되어 엔비디아 의존 없이도 동급의 하드웨어 효율을 달성했다는 점이 주목할 만합니다.

번역된 본문

Z.ai의 새로운 GLM-5.3-Flash 모델은 가격 대비 뛰어난 성능을 제공하며, 명확한 약점도 있고, 주목할 만한 인프라 측면의 의미를 담고 있습니다.

Z.ai에 따르면 GLM-5.3-Flash는 GLM-5 시리즈 최초의 네이티브 멀티모달 모델입니다. 총 3,200억 개의 파라미터를 갖추었으며 이 중 활성 파라미터는 180억 개에 불과하고, MIT 라이선스로 공개되었으며 100만 토큰의 컨텍스트 창을 제공합니다. 모델 가중치는 Hugging Face에서 이용할 수 있습니다.

Artificial Analysis의 측정에 따르면 이 모델은 최대 추론 노력에서 지능 지수(Intelligence Index) 57점을 기록했습니다. 이는 60점을 받은 더 큰 GLM-5.3에 단 3점 뒤지는 수치이며, GPT-5.6 Terra 및 Muse Spark 1.2와 동급입니다.

가격이 특히 눈에 띕니다. 지수 기준 작업당 비용은 0.09달러로, GLM-5.3의 0.68달러와 비교하면 약 7.5배 저렴합니다. Artificial Analysis에 따르면 이로써 이 모델은 지능과 비용의 파레토 프론티어에 위치하며, 최근 서구 공급업체에 강한 가격 압박을 가해온 중국 모델 목록에 추가되었습니다.

Z.ai의 API에서 GLM-5.3-Flash는 백만 입력 토큰당 0.15달러, 백만 출력 토큰당 0.50달러로, GLM-5.3 가격의 10% 조금 넘는 수준입니다.

에이전트 작업에서 이 모델은 더 큰 형제 모델에 뒤지지 않습니다. GDPval-AA v2에서 약 1770의 Elo 점수를 기록해 GLM-5.3 및 Grok 4.6과 동급이며, Claude Opus 5에만 뒤집니다. 다만 토큰 효율성은 여전히 떨어집니다. Artificial Analysis는 이 모델이 소비한 출력 토큰의 약 90%가 추론에 사용된 것을 확인했습니다.

엔비디아 대신 중국 칩 사용 출시 전에 Z.ai는 이 모델을 'ox-alpha'라는 익명 이름으로 OpenCode와 OpenRouter에서 테스트했고, 해당 주의 가장 인기 있는 모델이 되었습니다. 흥미롭게도 Z.ai에 따르면 그 모든 트래픽이 중국산 AI 칩에서 처리되었습니다. SemiAnalysis는 하루 100조 토큰을 서빙했다고 보도했는데, 이는 지금까지 프론티어 랩에서만 가능하다고 여겨졌던 수준의 용량입니다. Z.ai는 자사의 하드웨어 효율과 토큰당 비용이 일반적인 엔비디아 GPU와 동등하다고 밝혔습니다.

SemiAnalysis는 이를 최근 OpenAI의 새 칩 결과에 이어 'CUDA 해자(moat)'에 대한 또 다른 시험이라고 해석합니다. CUDA는 엔비디아의 AI 소프트웨어와 그래픽 카드 사이의 프로그래밍 계층으로, 거의 20년 동안 발전해 왔으며 거의 모든 AI 프레임워크가 이에 맞춰 최적화되어 있습니다. 다른 칩으로 전환하면 이 작업을 다시 해야 하며, 연산 작업을 재프로그래밍하고 메모리 접근을 조정하며 병목을 찾아 해결해야 합니다.

이에 Z.ai는 SGLang 위에 자체 서빙 소프트웨어를 구축하고, 처리 과정을 독립적으로 확장 가능한 단계로 분할했습니다. 팀에 따르면 이를 통해 동일한 하드웨어에서 첫 시도 대비 처리량이 3배 증가했습니다. GLM-5.3 기반 에이전트가 이 최적화 작업을 도왔습니다.

원문 보기
원문 보기 (영어)
GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 27, 2026 Nano Banana Pro prompted by THE DECODER Key Points Z.ai released GLM-5.3-Flash, a model with 320 billion parameters and a context window of one million tokens. On the Intelligence Index it nearly matches the larger GLM-5.3, but at 0.09 dollars per task it's about 7.5 times cheaper. The model ran entirely on Chinese AI chips, and Z.ai's own software delivered efficiency on par with Nvidia GPUs. Ask about this article… Search Z.ai's new GLM-5.3-Flash model offers strong value for the money, comes with clear weaknesses, and brings a noteworthy infrastructure angle. GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, according to Z.ai . It has 320 billion total parameters, of which only 18 billion are active, ships under an MIT license, and offers a context window of one million tokens. The weights are available on Hugging Face. Measurements from Artificial Analysis put the model at 57 points on the Intelligence Index at maximum reasoning effort. That's just three points behind the larger GLM-5.3, which scores 60, and level with GPT-5.6 Terra and Muse Spark 1.2. Ad The price is what stands out. Cost per task on the index runs 0.09 dollars, against 0.68 dollars for GLM-5.3, roughly 7.5 times cheaper. That puts the model on the Pareto frontier of intelligence and cost, according to Artificial Analysis, and adds it to a growing list of Chinese models that have recently put heavy price pressure on Western providers . Ad On Z.ai's API, GLM-5.3-Flash costs 0.15 dollars per million input tokens and 0.50 dollars per million output tokens, a little over ten percent of the price of GLM-5.3. On agentic tasks the model keeps pace with its bigger sibling. On GDPval-AA v2 it hits an Elo score of about 1770, matching GLM-5.3 and Grok 4.6, and trails only Claude Opus 5. It's still less token-efficient, though. Artificial Analysis found that roughly 90 percent of the output tokens it burned went to reasoning. Chinese chips instead of Nvidia Before launch, Z.ai tested the model anonymously as "ox-alpha" on OpenCode and OpenRouter, where it became the most popular model of the week. Interestingly, all of that traffic ran on Chinese AI chips, according to Z.ai. Ad SemiAnalysis reports it served 100 trillion tokens a day, a level of capacity that until now was thought possible only for frontier labs. Z.ai puts its hardware efficiency and cost per token on par with common Nvidia GPUs. SemiAnalysis reads this as another test of the "CUDA moat," coming after the recent results from OpenAI's new chip . CUDA is Nvidia's programming layer between AI software and the graphics card. It has grown for nearly 20 years, and just about every AI framework is tuned for it. Switching to other chips means redoing that work, reprogramming compute operations, adjusting memory access, and hunting down bottlenecks. Ad So Z.ai built its own serving software on top of SGLang and broke processing into stages that scale independently. The team says this tripled throughput over its first attempt on the same hardware. An agent based on GLM-5.3 helped with the optimization. Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Z.ai / GLM-5.3-Flash | Artificial Analysis / Benchmark | SemiAnalysis / Chinese AI chips
관련 소식