메뉴
BL
The Decoder 11일 전

중국 무버스터 키미(Kimi K3), 서구 컴퓨팅 우위 흔들다

IMP
9/10
핵심 요약

중국의 문샷 AI가 최근 공개한 '키미(Kimi K3)' 모델가 서구의 최고 수준 AI 모델과 맞먹는 성능을 보여주며 큰 화제를 모으고 있습니다. GPU 부족이라는 한계를 혁신적인 아키텍처와 강화학습 설계로 극복하여 탄생한 이 모델은, 서구 진영의 '컴퓨팅 우위' 전략이 과연 유효한지에 대한 깊은 의문을 던지고 있습니다.

번역된 본문

디프시크(Deepseek)와 마찬가지로, 중국의 키미(Kimi K3)는 서구 AI 연구소들이 자신들의 컴퓨팅 파워 우위에 대해 의심하게 만들고 있습니다.

문샷 AI(Moonshot AI)가 서구의 최고 수준 모델들과 거의 맞먹는 것으로 알려진 '키미 K3' 모델을 출시했습니다. 이번 발표는 미국의 수출 통제가 실제로 효과를 발휘하고 있는지에 대한 새로운 의구심을 불러일으키고 있으며, 심지어 OpenAI의 전략가마저도 이에 깊은 인상을 받았습니다.

단 일주일 전까지만 해도 연구 기관인 세미애널리시스(SemiAnalysis)는 중국 연구소들이 "최첨단 기술에 도달하기에는 컴퓨팅 자원이 너무 부족하다"고 평가했으며, 딥마인드(DeepMind) 직원인 아니카 소마이아(Anika Somaia)도 이 의견에 동의했습니다. 불과 며칠 후, 직원 수 약 300명의 스타트업인 문샷 AI는 '키미 K3'를 공개했습니다. 초기 평가에 따르면 이 모델은 앤스로픽(Anthropic)의 Opus 4.8과 동등한 수준이지만, 앤스로픽의 Fable 5나 OpenAI의 GPT-5.6 Sol과 같은 최상위 최첨단 모델에는 아직 미치지 못합니다. 하지만 실제로 그 격차가 어느 정도인지는 아직 불분명합니다.

소마이아는 수출 통제부터 초대형 기업들의 수천억 달러 투자 경쟁, 그리고 '컴퓨팅 해자' 투자 명분에 이르기까지 서구의 합의는 단 하나의 가정, 즉 '컴퓨팅 파워가 AI 능력을 결정한다'는 것에 기반하고 있다고 주장합니다. 하지만 자원의 부족은 혁신을 강제했습니다. 소마이아는 문샷 AI가 자체적인 AI 훈련용 문케이크(Mooncake) 스택을 구축한 것은 스타트업에 GPU가 충분하지 않았기 때문이라고 설명합니다. "비용 때문에 모델을 상용화하여 서비스할 여력이 없더라도, 실력 있는 소규모 연구소는 최첨단 모델을 만드는 데 필요한 컴퓨팅 양을 압축할 수 있습니다."

하드웨어 분석 기업 세미애널리시스의 창립자인 딜런 파텔(Dylan Patel) 역시 동의합니다. 그는 "초일류 인재로 구성된 소규모 팀과 강화학습(RL), 아키텍처, 데이터에 대한 깊은 연구가 컴퓨팅 부족의 상당 부분을 만회해 준다"고 말했습니다. 하지만 그는 중국 기업들이 중국 외부에서 GPU를 쉽게 대여할 수 있어 수출 통제의 상당 부분이 무의미해졌다는 점도 지적했습니다.

한 구글 딥마인드 연구원은 키미 K3를 "미친 듯이 뛰어나다(insanely good)"고 평가했습니다.

서구 AI 연구소들은 종종 중국 기업들이 지식 증류(Distillation)라는 형태의 데이터 도용을 한다고 비난합니다. 이는 더 작은 AI 모델이 더 큰 모델의 출력 결과를 학습하며 본질적으로 무임승차를 하여 서구 AI 연구소의 비즈니스 모델을 위협하는 방식입니다. 지금까지 서구 진영은 중국 연구소들이 적은 컴퓨팅 파워에도 불구하고 경쟁력을 유지할 수 있는 이유로 지식 증류를 꼽았습니다. 하지만 키미 K3에 대해서는 이러한 설명이 통하지 않는 것으로 보입니다. MIT와 구글 딥마인드의 AI 연구원인 미히엘 바커(Michiel Bakker)는 "이 결과는 단순한 지식 증류만으로는 설명할 수 없는 것 같다"며 이 모델을 "미친 듯이 뛰어나다"고 평가했습니다.

반면 블룸버그(Bloomberg)에 따르면 구글의 자체 플래그십 모델인 제미나이(Gemini) 3.5 Pro는 특히 주된 활용처인 코딩 분야에서 성능 목표를 달성하지 못해 몇 달째 출시가 지연되고 있습니다. 구글의 AI 전략은 다시 비판을 받고 있으며, 특히 독일을 중심으로 한 AI 검색 분야에서도 규제적 역풍에 직면해 있습니다.

OpenAI의 전략 미래 담당 책임자이자 전 정부 자문관인 딘 W. 볼(Dean W. Ball)은 키미를 "매우 훌륭한 모델"이라 평가하며, 에이전트 기반 코딩 세션에서는 "2026년 1분기 최고의 공개 모델들과 일치한다"고 말했습니다. 하지만 그는 또한 이 모델이 "토큰을 매우 많이 소비하는 것처럼 보여 실제로 운영 비용이 그렇게 저렴한지는 확신할 수 없다"고 지적했습니다. 그의 말이 틀리지 않습니다. 분석 기업 Artificial Analysis에 따르면, 키미 K3는 작업당 평균 0.94달러의 비용이 듭니다. 이는 작업당 1.04달러인 GPT 5.6 Sol과 비슷하고, 1.80달러인 Opus 4.8의 약 절반 수준입니다. 여전히 서구의 최고 수준 모델들보다는 저렴하지만, 이전 버전에 비해서는 격차가 줄어들었고 기존 중국산 오픈 웨이트(Open-weight) 모델들보다는 훨씬 비쌉니다.

OpenAI 전략가는 "완전한 AI 공산주의"를 경고합니다.

볼(Dean W. Ball)은 중국 정부가 이렇게 강력한 모델을 오픈소스로 공개하도록 허용한 것에 놀라움을 표했습니다. 그는 이유의 75%를 전략적 무지(無知)에 돌리며, 중국 공산당이 AI의 위험을 평가하는 방식이 "양 르쿤(Yann LeCun, 구글 뇌과학자 출신 AI 학자) 스타일"이어서 어떠한 실존적 위협도 보지 못하고 있다고 말했습니다. 나머지 이유는 기기 내 추론(Inference)을 위한 컴퓨팅 용량 부족으로, 이는 미국의 수출 통제로 인해 오픈 웨이트 전략이 의도치 않은 부산물로 탄생하게 만들었습니다. 중국 기업들 역시 거의 아무도 [이 막대한 초기 컴퓨팅 비용을 지불하려 하지 않는다는 사실을] 잘 알고 있기 때문입니다.

원문 보기
원문 보기 (영어)
Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 17, 2026 GPT-Image-2 prompted by THE DECODER Moonshot AI has released Kimi K3, a model reportedly close to matching top Western models. The launch raises fresh doubts about whether U.S. export controls are actually working. Even an OpenAI strategist is impressed. Just a week ago, research firm SemiAnalysis wrote that Chinese labs are "simply too compute poor to truly reach the frontier," a line flagged by Deepmind employee Anika Somaia. Days later, Moonshot AI, a startup with roughly 300 employees, released Kimi K3 , which by early assessments is on par with Anthropic's Opus 4.8 but still falls short of top frontier models like Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol . How large that gap actually is remains unclear. Somaia argues that the entire Western consensus, from export controls to the hyperscalers' hundreds-of-billions investment race to the "Compute Moat" investment thesis, rests on a single assumption: that computing power determines capability. But scarcity has forced innovation. Moonshot AI's in-house Mooncake stack for AI training was built precisely because the startup didn't have enough GPUs, Somaia says . "A small lab with taste can compress the compute needed to make a frontier model, even if it can't afford to serve one." Dylan Patel, founder of hardware analysis firm SemiAnalysis, agrees. "What they did with an extremely talented small team, strong research in RL, arch, data helps make up for lot of the compute deficit," he writes . But Patel also points out that Chinese companies can easily rent GPUs outside of China, which makes a portion of the export restrictions pointless. A Google Deepmind researcher calls Kimi K3 "insanely good" Western AI labs often accuse Chinese companies of a form of data theft through distillation , where a smaller AI model learns from the output of a larger one and essentially free-rides, threatening Western AI labs' business models. Until now, distillation has been the go-to explanation for how Chinese labs stay competitive despite having less compute. For Kimi K3, that explanation apparently doesn't hold up. "These results seem impossible to explain through distillation alone," writes Michiel Bakker , an AI researcher at MIT and Google Deepmind, calling the model "insanely good." Google's own flagship model, Gemini 3.5 Pro, meanwhile, has been delayed for months according to Bloomberg because it isn't hitting performance targets, especially in coding, its main use case. The company's AI strategy is drawing criticism again, and Google is also facing regulatory headwinds in AI search, particularly from Germany . Dean W. Ball , Head of Strategic Futures at OpenAI and a former government advisor, calls Kimi a "very good model" that in agent-based coding sessions matches "the best public models from Q1 2026." But he also notes that it seemed "very token hungry," making it "not obvious to me that this model is actually that cheap to run." He's not wrong. According to Artificial Analysis , Kimi K3 costs an average of $0.94 per task. That's close to GPT 5.6 Sol at $1.04 but roughly half the cost of Opus 4.8 at $1.80. It's still cheaper than the top Western models, but the gap has narrowed compared to the previous version, and it's much pricier than earlier open-weight Chinese models. OpenAI strategist warns against "full AI communism" Still, Ball says he's surprised the Chinese government allows such powerful models to be released as open-source. He attributes 75 percent of it to strategic blindness, saying the CCP is "very Yann LeCun-y" in how it assesses AI risks and doesn't see any existential threats . The rest comes down to a lack of computing capacity for client-side inference, which makes the open-weight strategy an unintended byproduct of U.S. export controls . The companies also know that hardly anyone would pay for Chinese models below the frontier, Ball claims. Open-weight models are "inherently decelerationist," Ball argues, because they slow down further AI investment. One possible outcome of a world dominated by them would be "full AI communism," with AI as a public good provided by the state as digital infrastructure. That's what China is proposing, according to Ball, who calls this scenario a "dystopian hellscape." That an OpenAI strategist is criticizing open-weight models this sharply is, of course, not without self-interest. His company relies on a closed business model and faces growing price pressure from providers like Moonshot AI and Deepseek. Regulatory fog instead of outright bans Ball predicts the Trump administration will create regulatory risk around using Chinese open-weight models. There's no need to ban open source, he argues, calling it "one of the dumber motifs of AI policy discussion." Authorities would only need to create enough uncertainty through "soft law," like having the Federal Reserve issue warnings about potential backdoors in Chinese AI models. The rationale wouldn't even need to be well-founded. The goal is a middle ground with enough risk to deter regulated companies from using Chinese models, without spooking the hyperscalers so badly that startups migrate to less reputable providers. Ball expects the government to roll out some version of this strategy . More efficient AI can still mean more demand for compute Kimi's progress doesn't necessarily mean less computing power is needed, but if it did, U.S. tech companies' massive infrastructure buildouts would look unnecessary, likely triggering a stock market crash. The opposite is more likely, though. The Jevons paradox suggests that more efficient models lead to more AI being deployed, which could actually drive even more demand for computing power. According to SemiAnalysis , Kimi K3 has 2.8 trillion parameters and is so large it doesn't fit on a single Nvidia DGX B200, even with FP4 quantization. It needs more powerful systems like the GB300 NVL72 or B300, each with 288 GB of memory per GPU. Again, the parallels to Deepseek are hard to miss . Back then, skeptics predicted a compute surplus and briefly rattled the markets. Instead, demand for computing power climbed as reasoning models gained traction, ironically driven in part by Deepseek's own models. Or as Google Deepmind CEO Demis Hassabis puts it, "Nobody in the world knows what happens next." AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->
관련 소식