메뉴
BL
The Decoder • 36일 전

프론티어 레이더 #4: 중국이 따라잡았다, 서구 AI의 우위는 무엇이 남았나

IMP
8/10
핵심 요약

Kimi K3, GLM-5.3 등 최신 중국 오픈웨이트 모델이 거의 모든 주요 벤치마크에서 미국 최상위 모델과 근접한 성능을 보이며 수개월로 여겨지던 격차가 투자자 우려로 이어지고 있습니다. 서구 연구소들은 증류(distillation)와 벤치마크 점수 최적화(benchmaxxing)를 지적하지만, 결론은 모델 자체의 우위는 방어할 수 없다는 것입니다. 이에 따라 경쟁 우위의 축이 개별 모델에서 다음 모델을 만들어내는 전체 시스템으로 이동하고 있으며, 유럽은 두 경쟁에서 동시에 뒤처지고 있다고 분석합니다.

번역된 본문

프론티어 레이더 #4: 중국이 따라잡았다, 그렇다면 서구 AI의 우위에는 무엇이 남았는가? 막시밀리안 슈라이너, 2026년 8월 20일

Kimi K3와 GLM-5.3이 이제 미국 최고 수준 모델과 근접한 성능에 도달했습니다. 서구 연구소들은 증류(distillation) 탓이라고 주장하며, 이에 대한 실질적인 증거도 있습니다. 하지만 잘못이 있든 없든 결론은 같습니다. 모델 자체의 우위는 방어할 수 없다는 것입니다. 이번 호에서는 무엇이 방어될 수 있는지 살펴봅니다.

THE DECODER 편집팀은 1년에 여섯 번 '프론티어 레이더'를 통해 하나의 핵심 AI 주제를 심층적으로 다룹니다. 뉴스레터 형식으로 운영되며 THE DECODER 구독자를 위해 이 사이트에서 독점 공개됩니다. 4호는 지금까지 중 가장 방대한 분량이었고 예상보다 오래 걸렸습니다. 주제 탓입니다. 중국 AI 모델이 어떻게 따라잡았는지, 서구 연구소가 어디에서 여전히 차별화할 수 있는지, 업계의 해자(moat)가 왜 이동했는지, 그리고 유럽이 왜 두 경쟁에서 동시에 지고 있는지 다룹니다. 1호는 AI 에이전트의 현재 상황, 2호는 AI의 생산성 효과 측정, 3호는 신흥 AI 토큰 경제를 다뤘습니다.

1년 반 전, DeepSeek R1이 충격을 일으켰습니다. 중국 연구소가 갑자기 최초의 상업적 추론 모델인 OpenAI의 o1과 경쟁했고, 훨씬 적은 비용으로 이를 해냈다고 알려졌습니다. 시장은 불안해했고, 며칠 만에 수조 원의 시가총액이 증발했습니다. 계획된 인프라 투자가 과장된 것이었을까요?

당시 상황은 헤드라인이 시사하는 것보다 불분명했습니다. DeepSeek 자체 보고서에 따르면 R1은 AIME 2024 같은 개별 테스트에서는 o1을 이겼지만, 사실 지식(SimpleQA) 등 다른 항목에서는 명확히 뒤처졌습니다. 이후 벤치마크에서 더 많은 격차가 드러났습니다. 중국 모델은 전반이 아니라 개별 분야에서만 정상에 도달했습니다. 이번 호 작업을 시작했던 6월 말만 해도 Z.ai의 GLM-5.2는 같은 패턴을 보였습니다.

그러다 최신 중국 오픈웨이트 모델들이 나왔습니다. Moonshot의 Kimi K3, Alibaba의 Qwen3.8-Max, 그리고 GLM-5.3입니다. 중국 모델들은 이제 거의 모든 폭넓고 까다로운 평가에서 상위권에 위치합니다. 장기 지식 과제를 처리하고, 여러 단계에 걸쳐 코드를 작성하며, 이전 세대보다 훨씬 안정적으로 도구를 조율합니다. 일반적인 벤치마크로 측정하면 자주 언급되던 수개월 격차는 투자자 문제가 될 만큼 줄어들었습니다.

월스트리트 저널에 따르면 Anthropic은 다가오는 IPO를 앞두고 불편한 질문을 받고 있으며, 최상위에서의 남아있는 우위를 내세워 방어하고 있습니다. 그 아래 등급에서는 시장이 대부분 중국의 개방형이고 훨씬 저렴한 모델들의 차지입니다. 투자자들은 순수 모델 성능만으로는 더 이상 비즈니스를 지탱하기 어렵다고 우려합니다. 오늘 어떤 모델이 독점적으로 할 수 있는 일은 몇 달 후면 무료로 다운로드 가능한 모델이 할 수 있게 되기 때문입니다.

두 가지 비난이 제기되고 있습니다. 첫째, 중국 연구소들이 서구 모델을 교사로 활용했다는 소위 증류 관행입니다. 둘째, 폭넓은 실력 없이 벤치마크 점수만 높게 나오도록 모델을 조율했다는 이른바 '벤치맥싱(benchmaxxing)'입니다.

미국의 우위는 사라지지 않았습니다. 하지만 기술적으로 가능한 최전선인 이른바 '프론티어'의 점점 좁아지는 몇몇 영역으로 후퇴했습니다. 그 우위가 어디에 남아 있는지, 경제적으로 얼마나 가치 있는지가 이번 호의 첫 번째 질문입니다. 두 번째 질문은 여기서 이어집니다. 모델만으로 더 이상 뚜렷한 차이를 만들 수 없다면, 우위는 무엇에 기반하는가? 우리의 논점은 이렇습니다. 모델 자체에 기반하는 바는 점점 줄고, 그 주변의 전체 시스템, 즉 지속적인 작업을 통해 다음 모델을 만들어내는 시스템에 점점 더 기반하게 된다는 것입니다.

남아있는 우위 출시 당시 Artificial Analysis는 Kimi K3를 지능 지수(Intelligence Index)에서 57점으로 3위에 올렸고, 당시 선두였던 GPT-5.5와 Opus 4.8 바로 뒤를 따랐습니다. K3는 에이전트 과제, 즉 기업 사용에서 중요한 실무 작업에서 가장 크게 개선되었습니다. AutomationBench-AA에서는 K3가 1위로 데뷔했고, 이후 Anthropic이 Opus 5로 대응했습니다. 에이전트가 가상의 소프트웨어 회사를 500 시뮬레이션 일 동안 운영하는 CEO-Bench에서 K3는 2,215만 달러로 공개된 최고 단일 기록을 세웠습니다.

원문 보기
원문 보기 (영어)
Frontier Radar #4: China has caught up, so what's left of the Western AI lead? Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 20, 2026 Nano Banana Pro prompted by THE DECODER Kimi K3 and GLM-5.3 are now within striking distance of the best US models. Western labs blame distillation, and there's real evidence for it. But guilty or not, the conclusion is the same: a model lead can't be defended. This issue looks at what can. Six times a year, the THE DECODER editorial team takes a close look at one core AI topic in the "Frontier Radar." It runs as a newsletter and exclusively here on the site for THE DECODER subscribers . Issue #4 is our most extensive yet and took longer than expected. Blame the topic: We look at how Chinese AI models caught up, where Western labs can still stand out, why the industry's moat has shifted, and why Europe is losing two races at once. Issue #1 covered the current state of AI agents. Issue #2 examined the measurable effects of AI on productivity. Issue #3 covered the emerging token economy of AI. A year and a half ago, DeepSeek R1 caused a shock. A Chinese lab was suddenly competing with OpenAI's o1, the first commercial reasoning model, and had reportedly done it for far less money. Markets got nervous. Billions in market value evaporated within days. Was the planned infrastructure buildout overblown? The picture back then was murkier than the headlines suggested. In DeepSeek's own report , R1 beat o1 on individual tests like AIME 2024 but trailed clearly on others, such as factual knowledge (SimpleQA). Later benchmarks exposed more gaps. Chinese models only reached the top in individual disciplines, not across the board. As recently as late June, when we started working on this issue, Z.ai's GLM-5.2 still showed the same pattern. Then came the latest Chinese open-weights models: Moonshot's Kimi K3 , Alibaba's Qwen3.8-Max , and GLM-5.3 . Chinese models now sit near the top of almost every broad, demanding evaluation. They handle long knowledge tasks, code across many steps, and coordinate tools far more reliably than their predecessors. Measured by common benchmarks, the often-cited gap of a few months has shrunk enough to become an investor problem. According to the Wall Street Journal, Anthropic is fielding uncomfortable questions ahead of its upcoming IPO and points to its remaining lead at the top in its defense. Below that tier, the field belongs largely to open, far cheaper models from China. Investors worry that raw model performance can barely carry a business anymore. Whatever a model can do exclusively today, a freely downloadable one can do a few months later. Two accusations are in play: Chinese labs allegedly tapped Western models as teachers, a practice known as distillation. And they allegedly tune their models for strong benchmark scores without the broad capabilities to match - so-called benchmaxxing. The American lead hasn't disappeared. But it has retreated to a few, ever-narrower areas of the so-called frontier, the leading edge of what's technically possible. Where that lead still sits, and what it's worth economically, is the first question of this issue. The second follows from it. If the model alone can't make a clear difference anymore, what does the lead rest on? Our thesis: less and less on the model, and more and more on the overall system around it, meaning the system where ongoing work produces the next models. What's left of the lead At launch, Artificial Analysis had K3 in third place on its Intelligence Index with 57 points , right behind then-leaders GPT-5.5 and Opus 4.8. K3 improved most on agentic tasks, so exactly the hands-on work that matters in enterprise use. On AutomationBench-AA, K3 even debuted in first place until Anthropic answered with Opus 5. On CEO-Bench , where an agent runs a fictional software company for 500 simulated days, K3 posted the best published single run at $22.15 million. Chinese models, including its predecessor K2.7, had regularly failed these long hauls. Qwen3.8-Max reaches a similarly high overall level. One caveat remains: Newer Chinese models sometimes burn far more tokens than Western ones, which eats into part of their price advantage. The cost-per-task math still applies. The edge of current model capabilities has been growing at uneven speeds for a while, what researchers call the "jagged frontier." Individual capabilities advance at different rates, which is why the much-quoted months-long gap always depended on who measured what. What's new is where a Western lead is still measurable at all. Three areas remain. The first is abstract specialty tests. Shortly after K3's launch, Opus 5 retook the top of the index with 61 points , a small gap over K3's 57. On ARC-AGI-1, a test of abstract pattern recognition using small puzzle grids , K3 and Fable 5 - Anthropic's flagship line for coding and agent work - sit practically even at 94.5 and 98.5 percent . Only on ARC-AGI-2 does the gap widen: 60.4 versus 89.2 percent. But ARC-AGI deliberately measures abstract pattern recognition far removed from everyday tasks. There's no guarantee this gap predicts practical differences. It could simply stay economically irrelevant in most cases. The second is reliability. The AA-AnalystAgent benchmark , launched August 12, tests agentic data analysis on real tables and documents. It only counts a task as solved if a model gets it right in five out of five independent runs, a measure the operators call pass^5. Opus 5 leads with 54 percent, ahead of GPT-5.5 at 50. K3 is the best open model at 39 percent. The gap comes almost entirely from poor repeatability. K3 solves 73 percent of tasks at least once in five attempts, practically even with Opus 5 at 74. Reliability doesn't follow the usual intelligence-index rankings either. Even GPT-5.6 Sol falls behind its own predecessor here. Commercially, reliability weighs heavily, as the operators explain in their launch article . An analyst agent only saves work when its answers hold up without review. Multiple runs and checkers can compensate, but they drive up the cost per accepted result. The third is cybersecurity . Here the gap is best documented. A joint assessment by the UK's AISI and the US CAISI found that K3 lags far behind leading US models in offensive cyber capabilities. On ExploitBench, a test for developing exploits, K3 scored 32 percent. The top US models averaged about 76. K3 failed all 41 tasks that required executing code on a target system; leading US models solved 20 on average. And in a simulated attack, K3 made it to step 17 of 32, the US leaders to 28.5 on average. But this gap is shrinking too. GLM-5.3 , unveiled August 14, scores 54.4 percent on ExploitBench by Z.ai's own measurement, landing on more than double of its predecessor GLM-5.2. That cuts the distance to the US leaders roughly in half within a month. On finding and validating vulnerabilities in source code (CyberGym), GLM-5.3 even edges past the leading US models. At the same time, cybersecurity is the one area where providers don't openly sell their strongest capabilities. Anthropic's best cyber model, Mythos 5 , which hits 78 percent on ExploitBench, is only available under controlled conditions through Project Glasswing. Its public sibling Fable 5 effectively stays at the level of the older Opus 4.8 at 40 percent, because upstream safeguards intercepted 407 of 410 test episodes. OpenAI takes a similar approach with Daybreak and specialized cyber variants. GPT-5.6 Sol reaches 73.5 percent on ExploitBench in OpenAI's own evaluation but stays below Critical, the highest level on OpenAI's risk scale. With the upcoming Astra, OpenAI can't rule out for the first time that one of its own models crosses that threshold. If it does, the Preparedness Framework kicks in much harder, up to a full development halt. OpenAI has already paused some internal Astra work. And with GLM-5.3, a Chinese lab is adopting this pattern fo