메뉴
BL
The Decoder 20일 전

데이터브릭스, 저렴한 비용의 中 오픈소스 모델 GLM 5.2를 기본 코딩 엔진으로 선택

IMP
8/10
핵심 요약

데이터브릭스가 자체 코드베이스 테스트 결과, 중국의 오픈소스 모델인 GLM 5.2가 앤스로픽의 오푸스 4.8(Opus 4.8)과 통계적으로 동등한 성능을 내면서도 비용은 훨씬 저렴하다는 사실을 확인했습니다. 이에 따라 데이터브릭스는 개발자들의 일상적인 업무용 코딩 모델로 GLM 5.2를 기본 도입할 계획이며, 코인베이스 등 다른 기업들 또한 비용 절감을 위해 중국산 모델로 빠르게 전환하고 있습니다.

번역된 본문

데이터브릭스, 저렴한 비용으로 오푸스와 맞먹는 성능 입증된 중국 오픈소스 모델 GLM 5.2를 기본 코딩 엔진으로 채택 작성자: Matthias Bastian (2026년 7월 9일)

핵심 요약: 자체 코드베이스를 활용한 내부 벤치마크에서 데이터브릭스는 중국의 오픈소스 모델인 GLM 5.2가 성능 면에서 앤스로픽의 Opus 4.8과 통계적으로 동등한 수준이지만, 작업당 비용은 훨씬 저렴하다는 사실을 발견했습니다. 따라서 이 회사는 앞으로 개발자들의 일상적인 업무용 모델로 GLM 5.2를 사용할 계획입니다. 분석 결과 또한 테스트된 모델들이 3개의 성능 등급으로 나뉘며, 최상위 등급에는 다양한 공급업체의 모델들이 포함되어 있음을 보여주었습니다. 팀은 퍼블릭 데이터셋이 자사의 코드베이스를 제대로 반영하지 못하고 모델이 학습 데이터의 사전 지식을 활용해 '속임(cheat)'을 쓸 수 있기 때문에, 실제 작업을 기반으로 한 자체 벤치마크를 개발했습니다.

본문: 데이터브릭스는 수백만 줄에 달하는 자체 코드베이스에서 GLM 5.2를 벤치마크한 결과, 이 중국 오픈소스 모델이 더 낮은 비용으로 앤스로픽의 Opus 4.8과 통계적으로 동등한 성능을 보이는 것을 확인했습니다. 이에 따라 회사는 이를 개발자들을 위한 일상적인 핵심 작업 모델로 삼을 계획입니다. GLM 5.2는 작업당 1.28달러의 비용으로 최고 성능 그룹에 진입했으며, 이는 Opus의 작업당 1.94달러와 비교되는 수치입니다.

데이터브릭스의 공동 창립자인 마테이 자하리아(Matei Zaharia)를 포함한 블로그 게시물 작성자들은 "증거는 이제 이 모델들을 코딩을 위한 일상용 모델로 배포할 때가 되었음을 보여줍니다"라고 작성했습니다. 내부 파일럿 테스트에서 얻은 개발자들의 피드백이 이러한 결과를 뒷받침했으며, 회사 측은 이미 GLM이 최고의 성능을 발휘할 수 있도록 실행하는 작업을 진행 중이라고 밝혔습니다.

이러한 움직임은 데이터브릭스만의 일은 아닙니다. 암호화폐 거래소 코인베이스(Coinbase)는 GLM-5.2 및 Kimi 2.7을 포함한 중국 모델로 전환하여, 토큰 사용량은 계속 증가하는 데에도 불구하고 AI 지출을 절반으로 줄였습니다. 린디(Lindy)는 클로드(Claude)를 완전히 폐기하고 Deepseek v4를 채택하여 수백만 달러를 절약했습니다. 스노플레이크(Snowflake)는 Opus 4.7과 GLM-5.2를 테스트하고 두 모델이 극히 일부의 비용으로 거의 동등한 성능을 내는 것을 확인했습니다. AI 모델 API 라우터인 OpenRouter에 따르면, 2026년 2월부터 중국 모델의 주간 트래픽 비중이 작년의 11%에서 30%를 넘어섰으며, 서구권 대안 모델들보다 60~90% 더 낮은 비용을 보여주고 있습니다.

광고

3개의 성능 등급을 아우르는 단일 연구소의 독점은 없다 데이터브릭스에 따르면, 전반적으로 테스트된 모델과 구성은 3개의 클러스터로 나뉩니다. 8290%의 통과율을 보인 최상위 그룹에는 특정 구성의 Opus 4.8, GLM 5.2 및 GPT 5.5가 포함됩니다. 7182%의 통과율을 가진 중간 그룹에는 Sonnet 4.6, Sonnet 5 및 GPT 5.4 등이 속합니다. 51~60%의 통과율을 기록한 하위 등급에는 GPT 5.4-mini와 Haiku 4.5가 자리 잡고 있습니다.

광고

Unity AI 게이트웨이를 통한 분석 결과, 데이터브릭스 엔지니어들의 코딩 작업 중 61%는 중간 수준의 복잡도를 가지며, 약 19%는 낮은 복잡도, 오직 12%만이 높은 복잡도를 요구하는 것으로 나타났습니다. 이전까지는 가장 비싼 모델이 기본값으로 사용되었습니다. 이제 회사는 작업의 복잡도에 따라 더 많은 작업을 저렴한 등급의 모델로 분배할 계획입니다. 최고의 품질 대비 비용 효율성을 나타내는 파레토 최전선(Pareto frontier)은 OpenAI, Anthropic, 그리고 오픈소스 등 3개 공급업체의 모델에 의해 형성되고 있습니다. 데이터브릭스는 오직 이들을 혼합하여 사용할 때만 최고 수준의 성능을 낼 수 있다고 말합니다.

광고

데이터브릭스는 또한 토큰 가격과 실제 작업 비용이 동일하지 않다는 점을 지적합니다. 자동차의 연비와 마찬가지로 토큰 효율성 역시 매우 중요하며, 이는 소프트웨어 환경에 따라 크게 달라집니다. 한 테스트에서 Pi 하네스는 Claude Code보다 약 3배 적은 컨텍스트를 전송했습니다. '높은 노력(high effort)' 설정의 Opus 4.8 기준, Pi는 비슷한 품질(87% 대 85%)에서 2.08배 더 저렴했습니다. GPT 5.5도 유사한 패턴을 보였는데, Codex는 123만 개의 토큰을 사용한 반면 Pi는 66만 5천 개의 토큰만을 사용했습니다.

퍼블릭 데이터셋 대신 실제 작업 사용 데이터브릭스는 SWE-Bench와 같은 기존 퍼블릭 대안에 의존하는 대신, 실제 풀 리퀘스트(Pull Request)를 바탕으로 자체적인 벤치마크를 구축했습니다. 시간이 지남에 따라 솔루션이 학습 데이터에 유출될 수 있고, 기존의 작업들은 파이썬, 고(Go), 타입스크립트, 스칼라, 러스트(Rust) 등 10개 이상의 언어에 걸친 자사의 기술 스택과 맞지 않기 때문입니다. 최근 OpenAI도 비슷한 이유로 SWE-Bench-Pro 사용을 경고한 바 있습니다.

광고

각 작업은 최근에 작성되어야 하고, 사람이 직접 작성한 것이며, 고품질 테스트와 짝을 이루고, 전체 스택을 대표할 수 있어야 했습니다. 모든 테스트는...

원문 보기
원문 보기 (영어)
Databricks makes Chinese open-source model GLM 5.2 its default coding engine after it matched Opus at lower cost Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 9, 2026 Key Points In an internal benchmark using its own codebase, Databricks found that the Chinese open-source model GLM 5.2 is statistically on par with Anthropic’s Opus 4.8 in terms of performance, but is significantly less expensive per task. The company therefore plans to use GLM 5.2 as the day-to-day working model for its developers going forward. The analysis also revealed that the tested models fall into three performance classes, with the top tier consisting of models from various providers. For the test, the team developed its own benchmark using real-world tasks, since public datasets are often not representative of their own codebase and models can "cheat" by leveraging prior knowledge from training data. Ask about this article… Search Databricks benchmarked GLM 5.2 on its own multi-million-line codebase and found the Chinese open-source model statistically tied with Anthropic's Opus 4.8 at lower cost. The company now plans to make it a daily workhorse for its developers. GLM 5.2 hit the top performance cluster at $1.28 per task versus $1.94 for Opus. "The evidence shows it's time to start deploying these as daily drivers for coding," write the authors of the blog post, including Databricks co-founder Matei Zaharia. Developer feedback from internal pilots backed up the results, and the company says it's already working on running GLM at peak performance . Databricks isn't alone. Coinbase moved to Chinese models including GLM-5.2 and Kimi 2.7, cutting AI spending in half while token usage kept climbing. Lindy ditched Claude entirely for Deepseek v4 and saved millions. Snowflake tested GLM-5.2 against Opus 4.7 and found them nearly tied at a fraction of the cost. On OpenRouter , Chinese models have topped 30 percent of weekly traffic since February 2026, up from 11 percent last year, at 60 to 90 percent lower cost than Western alternatives. Ad No single lab dominates across three performance tiers Overall, the tested models and configs fell into three clusters, according to Databricks. The top group, with an 82 to 90 percent pass rate, includes Opus 4.8, GLM 5.2, and GPT 5.5 in certain configs. A middle group at 71 to 82 percent includes Sonnet 4.6, Sonnet 5, and GPT 5.4, among others. The bottom tier at 51 to 60 percent holds GPT 5.4-mini and Haiku 4.5. Ad DEC_D_Incontent-1 An analysis through Unity AI Gateway found that 61 percent of coding tasks from Databricks engineers are medium complexity, about 19 percent low, and only 12 percent high. The most expensive models had been the default. Now the company plans to route more work to cheaper tiers based on task complexity. The Pareto frontier, the best quality-to-cost ratio, is shaped by models from three providers: OpenAI, Anthropic, and open source. Only a mix delivers frontier-level performance, Databricks says. Ad Databricks also points out that token price and actual task cost aren't the same . Token efficiency matters just as much, like fuel economy in a car, and varies widely by software environment. In one test, the Pi harness sent about three times less context than Claude Code. For Opus 4.8 at "high effort," Pi was 2.08x cheaper at comparable quality (85 versus 87 percent). GPT 5.5 showed a similar pattern: Codex used 1,235,000 tokens versus 665,000 for Pi. Real tasks instead of public datasets Databricks built its own benchmark from real pull requests rather than relying on public alternatives like SWE-Bench. Solutions leak into training data over time, and the tasks don't match a stack spanning more than ten languages, including Python, Go, TypeScript, Scala, and Rust. OpenAI recently warned against SWE-Bench-Pro for similar reasons. Ad DEC_D_Incontent-2 Each task had to be recent, human-written, paired with high-quality tests, and representative of the full stack. All were reviewed by hand, with tests partly rewritten to allow alternative implementations. Scoring relied solely on passing tests, not an LLM judge, which Databricks says tends to reward answers that sound good rather than ones that are correct. Ad The team also hit a cheating problem : models searched the Git history for the correct solution instead of working it out. Databricks fixed this by truncating the entire Git history for each run. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Databricks