메뉴
BL
The Decoder 34일 전

스노우플레이크 CEO, GLM-5.2 성능은 오피스 4.7 맞먹고 비용은 극히 저렴

IMP
8/10
핵심 요약

스노우플레이크의 실사용 코딩 벤치마크 결과, 중국 AI 모델인 GLM-5.2가 안스로픽의 Opus 4.7과 거의 동등한 문제 해결 능력을 보여주었습니다. 첫 번째 시도의 정확도나 토큰 소비량 등 효율성 측면에서는 Opus가 우세했지만, GLM-5.2의 압도적으로 저렴한 사용 비용은 오픈AI 등 서구 AI 기업들의 높은 기업 가치를 위협하는 강력한 요인으로 작용하고 있습니다.

번역된 본문

원문 제목: 스노우플레이크 CEO, 저렴한 비용으로 Opus 4.7과 경쟁 가능한 GLM-5.2를 발견하다 출처: 블로그

스노우플레이크 CEO, 저렴한 비용으로 Opus 4.7과 경쟁 가능한 GLM-5.2를 발견하다 Matthias Bastian이 작성함 | 2026년 6월 24일

핵심 요약 스노우플레이크가 진행한 실사용 프로그래밍 벤치마크에서, 중국 AI 모델 GLM-5.2와 안스로픽의 Opus-4.7은 각 작업당 3번의 시도가 주어졌을 때 각각 66%, 67%의 문제를 해결하며 거의 동일한 성능을 보여주었습니다. Opus는 첫 번째 시도 정확도(GLM 47.6% 대비 53.7%)에서 앞서며 전체적인 효율성도 높습니다. 반면 GLM은 작업당 평균 99회의 반복 횟수를 필요로 하고(Opus는 80회), 거의 두 배에 달하는 토큰을 소비합니다. 이러한 효율성 격차에도 불구하고 GLM-5.2는 백만 출력 토큰당 4.40달러로 압도적으로 저렴하여, 오픈AI와 같은 서구 AI 기업들의 높은 기업 가치를 위협할 수 있는 상당한 가격 압박을 만들어내고 있습니다.

본문 기사 스노우플레이크는 실사용 벤치마크에서 GLM-5.2와 Opus 4.7을 비교했습니다. 그 결과 중국산 모델이 선전하는 것으로 나타났습니다. 이 테스트는 DuckDB와 스노우플레이크 모두에서 작동하는 코드를 작성해야 하는 103개의 작업으로 구성되었으며, 각 작업은 3번씩 실행되었습니다. 모델이 각 작업당 3번의 시도를 받았을 때, 두 모델은 66% 대 67%의 문제 해결률을 기록하며 막상막하였습니다. 하지만 첫 번째 시도 정확도에서는 차이가 드러납니다. Opus는 53.7%를 기록한 반면, GLM은 47.6%에 그쳐 출력의 일관성이 떨어짐을 보여주었습니다. 중국 모델은 또한 작업당 평균 99회의 실행을 기록한 반면(Opus는 80회), 8억 6천만 개의 토큰을 소모해 4억 3천 9백만 개를 사용한 Opus의 거의 두 배에 달했습니다.

GLM의 강점은 두 플랫폼(DuckDB 및 스노우플레이크)에서 코드의 유효성을 동시에 안정적으로 검증하는 것입니다. 스노우플레이크의 스리다르 라마스와미(Sridhar Ramaswamy) CEO에 따르면, 이것이 특정 작업을 GLM만 해결할 수 있었던 이유입니다. 반면 이 모델의 약점은 너무 일찍 포기하거나, 잘못된 부분을 집착하여 반복적으로 확인하는 것입니다. 한 작업에서 GLM은 24분 동안 411번의 도구 호출을 발생시키며 행 수, 분포, Null 값 및 열 유형을 확인했음에도 3번의 시도를 모두 실패했습니다. 반면 Opus는 9분 만에 49번의 호출로 동일한 작업을 해결했습니다. 라마스와미 CEO는 GLM이 더 깔끔한 코드를 생성한다는 주장은 사실이 아니었다고 밝혔습니다. 더 많은 확인이 항상 더 올바른 결과로 이어지는 것은 아니기 때문입니다. 그럼에도 불구하고 그는 팀이 GLM-5.2의 가능성을 보고 흥분하고 있으며, 고객들에게 이 모델을 제공할 계획이라고 덧붙였습니다.

중국의 가격 정책, 서구 AI 거품에 가하는 현실적인 압박 이러한 결과는 가격 측면에서 볼 때 가장 중요한 의미를 갖습니다. 지푸(Zhipu)의 공식 가격표에 따르면, GLM-5.2는 백만 입력 토큰당 1.40달러, 백만 출력 토큰당 4.40달러입니다. 일부 서드파티 제공업체는 지푸의 가격보다도 훨씬 더 낮은 가격을 제시하고 있습니다. 반면, Claude Opus 4.7은 입력 토큰당 5달러, 출력 토큰당 25달러입니다. GPT-5.5는 입력 5달러, 출력 30달러를 청구합니다.

모델 | 입력 | 캐시된 입력 | 출력 GLM-5.2 | $1.40 | $0.26 | $4.40 Claude Opus 4.7 | $5.00 | $0.50 (캐시 적중 시) | $25.00 GPT-5.5 | $5.00 | $0.50 | $30.00 GPT-5.4 | $2.50 | $0.25 | $15.00

물론 GLM의 더 높은 토큰 사용량은 이러한 가격 차이를 어느 정도 좁혀줍니다. 하지만 안스로픽과 오픈AI는 심각한 가격 압박에 직면해 있으며, 이는 두 서구 AI 연구소가 미래를 걸고 있는 핵심 사용 사례인 코딩 분야에서 직접적인 타격을 주는 것입니다. 만약 이러한 압박이 수익 성장을 늦추거나, 최악의 경우 수익을 감소시킨다면, 이미 과도하게 부풀려진 AI 시장은 현실적인 스트레스 테스트에 직면하게 될 것입니다. 오픈AI와 안스로픽의 높은 기업 가치는 수익이 계속해서 빠르게 증가할 것이라는 가정에 기반하고 있습니다. 이러한 평가액은 데이터센터부터 칩 주문에 이르기까지, AI 인프라 구축에 투자된 수십억 달러의 베팅과 직결되어 있습니다.

원문 보기
원문 보기 (영어)
Snowflake CEO finds GLM-5.2 competitive with Opus 4.7 at a fraction of the cost Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 24, 2026 Key Points In a real-world programming benchmark conducted by Snowflake, the Chinese AI model GLM-5.2 and Anthropic's Opus-4.7 performed nearly identically when given three attempts per task, solving 66 and 67 percent of problems, respectively. Opus holds an edge on first-attempt accuracy at 53.7 percent versus GLM's 47.6 percent, and is more efficient overall—GLM requires an average of 99 iterations per task compared to 80 for Opus and consumes nearly twice as many tokens. Despite these efficiency gaps, GLM-5.2 is dramatically cheaper at $4.40 per million output tokens, creating significant price pressure that could challenge the high valuations of Western AI companies like OpenAI. Ask about this article… Search Snowflake compared GLM-5.2 and Opus 4.7 in a hands-on benchmark. The Chinese model held its own. The test covered 103 tasks, each run three times, where models had to write code that works on both DuckDB and Snowflake. When each model got three attempts per task, the two were neck and neck: 66% vs. 67% of tasks solved. First-attempt accuracy diverges: Opus hit 53.7%, GLM only 47.6%, showing GLM's output is less consistent. The Chinese model also averaged 99 runs per task versus Opus's 80 and burned through 860 million tokens, nearly double Opus's 439 million. Ad GLM's strength is validating code reliably across both platforms (DuckDB and Snowflake) at the same time. According to Snowflake CEO Sridhar Ramaswamy , that's why only GLM could solve certain tasks. Ad DEC_D_Incontent-1 Its weaknesses are giving up too early and obsessively checking the wrong things. On one task, GLM fired off 411 tool calls in 24 minutes, checking row counts, distributions, null values, and column types, and still failed all three attempts. Opus solved the same task with 49 calls in 9 minutes. The claim that GLM produces cleaner code didn't hold up, Ramaswamy said. More checks don't lead to more correct results. Still, the team is excited about GLM-5.2 and wants to make it available to customers. Ad China's pricing puts real pressure on the Western AI bubble The results matter most in the context of price. GLM-5.2 costs $1.40 per million input tokens and $4.40 per million output tokens, according to Zhipu's official price sheet . Some third-party providers undercut Zhipu's price even further. Claude Opus 4.7 runs $5 input and $25 output. GPT-5.5 costs $5 input and $30 output. Model Input Cached Input Output GLM-5.2 $1.40 $0.26 $4.40 Claude Opus 4.7 $5.00 $0.50 (Cache Hit) $25.00 GPT-5.5 $5.00 $0.50 $30.00 GPT-5.4 $2.50 $0.25 $15.00 GLM's higher token usage eats into that price gap somewhat . But Anthropic and OpenAI are facing serious pricing pressure, and right in coding, the flagship use case both Western AI labs are betting on. Ad DEC_D_Incontent-2 If that pressure slows revenue growth, or worse, shrinks it, the already inflated AI market faces a real stress test. OpenAI's and Anthropic's valuations rest on the assumption that revenue keeps climbing fast. Those valuations are tied to billions in bets on AI infrastructure buildout, from data centers to chip orders. Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: via X