메뉴
HN
Hacker News 42일 전

GLM-5.2, 오픈웨이트 모델 중 1위 등극

IMP
9/10
핵심 요약

Z ai의 GLM-5.2가 AI 벤치마크 플랫폼인 Artificial Analysis의 지능 지수에서 51점을 기록하며 새로운 최고 수준의 오픈웨이트(Open weights) 모델로 등극했습니다. 이 모델은 과학적 추론 및 실제 에이전트 작업 수행 능력에서 대폭 향상되어 폐쇄형 모델인 GPT-5.5와 맞먹는 성능을 보여줍니다. 작업 당 생성하는 토큰량은 많아졌지만, 동급의 지능을 가진 모델들 중에서는 가장 낮은 비용으로 매우 효율적인 성능을 제공합니다.

번역된 본문

Artificial Analysis 전체 기사 - 2026년 6월 17일

GLM-5.2, Artificial Analysis 지능 지수에서 새로운 최고의 오픈웨이트 모델 등극

Z ai의 GLM-5.2가 Artificial Analysis 지능 지수(Intelligence Index)에서 51점을 기록하며 새로운 최고 수준의 오픈웨이트(Open weights) 모델이 되었습니다. 이 모델은 '지능 대 작업 당 비용(Intelligence vs Cost per Task)' 파레토 프론티어(Pareto frontier)에 위치해 있습니다.

GLM-5.2는 GLM-5.1과 동일한 크기(총 744B / 활성 파라미터 40B)를 유지하면서도 지능 지수 v4.1에서 11점 높은 점수를 받아 MiniMax-M3(44점)와 DeepSeek V4 Pro(max, 44점)를 앞섰습니다. 자체(First-party) API 가격은 GLM-5.1과 동일한 수준으로, 1백만 토큰당 입력 $1.4, 출력 $4.4, 캐시 적중(Cache hit) $0.26입니다.

주요 결과: ➤ GLM-5.2는 지능 지수 v4.1 기준 최고의 오픈웨이트 모델입니다. 51점으로 MiniMax-M3(44점), DeepSeek V4 Pro(max, 44점), Kimi K2.6(43점)을 리드합니다. ➤ 대부분의 평가에서 향상, 특히 과학적 추론 능력이 돋보임: GLM-5.2는 대부분의 평가에서 GLM-5.1 대비 향상되었습니다. 특히 CritPt(+16점, 21%) 및 HLE(+12점, 40%)의 과학적 추론에서 두각을 드러내며, AA-LCR(+9점, 71%), tau3 banking(+15점, 27%), SciCode(+7점, 50%) 성적을 기록했습니다. TerminalBench v2.1 역시 향상되었으며(+16점, 78%), GPQA Diamond는 3점 증가한 89%를 기록했습니다. ➤ GDPval-AA v2 기준 최고의 오픈웨이트 모델이자 폐쇄형 모델과 경쟁 가능: GLM-5.2는 GDPval-AA v2에서 1524점을 기록하여 MiniMax-M3(1418점)와 DeepSeek V4 Pro(max, 1328점)를 앞섭니다. 이 놀라운 결과는 GLM-5.2를 GPT-5.5(xhigh 추론)와 같은 폐쇄형(Proprietary) 모델들과 동등한 수준에 위치시킵니다. 기존 GDPval-AA를 발전시킨 v2는 인간의 성능을 1000 Elo로 베이스라인을 설정하고, 최첨단 모델 평가단 패널을 회전식으로 도입하며, 더 긴 에이전트 궤적을 위해 턴 제한을 100에서 250으로 늘렸습니다. ➤ 작업 당 출력 토큰 수 증가: 이 모델은 지능 지수 작업 당 43k의 출력 토큰을 사용하며, 이는 GLM-5.1(26k), MiniMax-M3(24k), Kimi K2.6(35k), DeepSeek V4 Pro(max, 37k)보다 높은 수치입니다. ➤ 지능 대비 작업 당 비용 파레토 프론티어: GLM-5.2는 동일한 지능 수준의 모델들 중 작업 당 비용이 가장 낮습니다. GLM-5.2의 작업 당 비용은 약 $0.46으로, GLM-5.1($0.25), Kimi K2.6($0.31), MiniMax-M3($0.18), DeepSeek V4 Pro(max, $0.05)와 비교됩니다.

추가 모델 세부 정보: ➤ 라이선스: MIT ➤ 크기: 총 744B 파라미터, 활성 파라미터 40B (GLM-5.1과 동일) ➤ 컨텍스트 윈도우: 1M 토큰 (GLM-5.1의 200K에서 상향) ➤ 가격: 1백만 토큰당 입력 $1.4 / 캐시 적중 $0.26 / 출력 $4.4 ➤ 가용성: Z ai의 자체 API 외에도 DeepInfra, Novita, Nebius, Parasail, Siliconflow, GMI Cloud, Baseten, Fireworks 등 타사 제공업체를 통해 이용 가능합니다.

GLM-5.2는 실제 에이전트 성능을 측정하는 핵심 지표인 GDPval-AA v2에서 모든 오픈웨이트 모델을 이끌고 있습니다. 1524점으로 MiniMax-M3(1418점) 및 DeepSeek V4 Pro(max, 1328점)를 앞서고 있으며, 사실상 GPT-5.5(xhigh, 1514점)와 동등한 수준입니다.

저희는 GLM-5.2의 출력물을 다양한 GDPval-AA 작업에서 직접 검토했으며, 그 결과의 일부를 아래에 첨부했습니다.

GLM-5.2는 AA-옴니사이언스(AA-Omniscience) 지수에서 4점을 받아 GLM-5.1(2점)보다 향상되었습니다. 이는 더 높은 정확도(24.2% 대 25.1%)와 더 낮은 환각률(29.4% 대 28.1%)에서 비롯되었으며, 시도율(attempt rate)은 47%로 동일했습니다.

GLM-5.2는 지능 지수 작업 당 43k의 출력 토큰을 사용하며, 이 중 37k가 추론(Reasoning)에 사용됩니다. 이는 GLM-5.1(26k)보다 증가한 수치이며 MiniMax-M3(24k) 및 Kimi K2.6(35k)과 같은 오픈웨이트 동료 모델들보다 높아, 동일한 지능 수준에서는 토큰 효율성이 상대적으로 떨어지는 모델군에 속합니다. 따라서 '지능 대 출력 토큰' 차트에서 GLM-5.2는 가장 매력적인 사분면에서는 벗어나 있습니다.

Artificial Analysis 지능 지수 v4.1의 개별 평가 항목 세부 분석은 다음에서 확인하실 수 있습니다. 다른 최고 수준 모델들과 비교하기: https://artificialanalysis.ai/models/glm-5-2 최신 Artificial Analysis 지능 지수 v4.1 읽기: 에이전트 워크로드로의 전환 Artificial Analysis 지능 지수 v4.1 발표: 전환을 향하여

원문 보기
원문 보기 (영어)
Artificial Analysis All articles June 17, 2026 GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index Z ai’s GLM-5.2 is the new leading open weights model on the Artificial Analysis Intelligence Index scoring 51 and it sits on the Pareto frontier of Intelligence vs Cost per Task GLM-5.2 is the same size as GLM-5.1 (744B total / 40B active parameters) but scores 11 points higher on the Intelligence Index v4.1, placing ahead of MiniMax-M3 (44) and DeepSeek V4 Pro (max, 44). On the first-party API it is priced in line with GLM-5.1 at $1.4/$4.4/$0.26 per 1M input/output/cache hit tokens Key results: ➤ GLM-5.2 is the leading open weights model on the Intelligence Index v4.1. At 51, it leads MiniMax-M3 (44), DeepSeek V4 Pro (max, 44) and Kimi K2.6 (43) ➤ Improvements across most evaluations, particularly scientific reasoning: GLM-5.2 gains over GLM-5.1 on most evaluations, led by scientific reasoning on CritPt (+16 points to 21%) and HLE (+12 points to 40%), alongside AA-LCR (+9 points to 71%), tau3 banking (+15 points to 27%) and SciCode (+7 points to 50%). TerminalBench v2.1 also improves (+16 points to 78%) and GPQA Diamond gains 3 points to 89% ➤ Leading open weights model on GDPval-AA v2 and competitive with proprietary models: GLM-5.2 scores 1524 on GDPval-AA v2, ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro (max, 1328). This impressive result places GLM-5.2 in-line with proprietary models including GPT-5.5 (xhigh reasoning). GDPval-AA v2 builds on the original GDPval-AA by baselining Elo to human performance at 1000, introducing a rotating panel of frontier-model judges, and raising the turn limit from 100 to 250 for longer-horizon agent trajectories ➤ GLM-5.2 uses more output tokens per task than other leading open weights models: the model uses 43k output tokens per Intelligence Index task, up from GLM-5.1 (26k) and above MiniMax-M3 (24k), Kimi K2.6 (35k) and DeepSeek V4 Pro (max, 37k) ➤ On the Intelligence vs. Cost per Task Pareto Frontier: GLM-5.2 is on the Pareto frontier of the Intelligence vs Cost per Task chart, with the lowest cost per task among models at its intelligence level. GLM-5.2 costs ~$0.46 per task, compared to GLM-5.1 ($0.25), Kimi K2.6 ($0.31), MiniMax-M3 ($0.18) and DeepSeek V4 Pro (max, $0.05) Additional Model Details: ➤ License: MIT ➤ Size: 744B total parameters, 40B active parameters, equivalent to GLM-5.1 ➤ Context window: 1M tokens, up from 200K on GLM-5.1 ➤ Pricing: $1.4/$0.26/$4.4 per 1M input/cache hit/output tokens ➤ Availability: Alongside Z ai's first-party API, GLM-5.2 is available across third-party providers including DeepInfra, Novita, Nebius, Parasail, Siliconflow, GMI Cloud, Baseten, and Fireworks GLM-5.2 leads all open weights models on GDPval-AA v2, our primary metric for real-world agentic performance. At 1524 it places ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro (max, 1328), and is effectively level with GPT-5.5 (xhigh, 1514). We visually inspected GLM-5.2's outputs across a range of GDPval-AA tasks. We have attached a selection below. GLM-5.2 scores 4 on the AA-Omniscience Index, up from GLM-5.1 (2). The gain comes from both higher accuracy (25.1% vs 24.2%) and a lower hallucination rate (28.1% vs 29.4%), with attempt rate flat at 47%. GLM-5.2 uses 43k output tokens per Intelligence Index task, of which 37k is reasoning. This is up from GLM-5.1 (26k) and higher than open weights peers MiniMax-M3 (24k) and Kimi K2.6 (35k), placing it among the less token-efficient open weights models at its intelligence level. GLM-5.2 sits off the most attractive quadrant on the Intelligence vs Output Tokens chart. Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.1. Compare GLM-5.2 with other leading models at: https://artificialanalysis.ai/models/glm-5-2 Read the latest Artificial Analysis Intelligence Index v4.1: a shift toward agentic workloads Announcing Artificial Analysis Intelligence Index v4.1: a shift toward agentic workloads, featuring upgraded benchmarks and new per-task metrics June 16, 2026 Claude Fable 5 Launches at #1 on the Artificial Analysis Intelligence Index Anthropic is nearly 5 points ahead of any other lab’s best model June 10, 2026 Claude Fable 5: the first public Mythos-class model Anthropic has released Claude Fable 5, the first publicly available Mythos-class model that ranks #1 in our agentic real-world knowledge work benchmark GDPval-AA June 9, 2026