메뉴
BL
The Decoder 10일 전

문샷 Kimi K3, 프론트엔드 코딩 1위...수학은 부진

IMP
7/10
핵심 요약

중국 문샷(Moonshot)의 AI 모델 Kimi K3가 프론트엔드 코드 벤치마크에서 서방 선도 모델들을 제치고 1위를 차지하며 인간 선호도 기준 뛰어난 코딩 성능을 입증했습니다. 그러나 전문가 수준의 고난도 수학 벤치마크에서는 약 39%의 정확도에 그쳐 오픈AI와 앤스로픽 등 최상위권 모델들(약 90%)에 크게 뒤처지는 편향된 성능을 보여주었습니다.

번역된 본문

문샷(Moonshot)의 Kimi K3, 프론트엔드 코드에서는 Fable 5를 압도하지만 복잡한 수학에서는 크게 뒤처져 Matthias Bastian이 작성함 | 2026년 7월 19일

중국 문샷의 AI 모델 Kimi K3가 서구 AI 커뮤니티에서 많은 관심을 받고 있습니다. 가장 큰 관심사는 이 모델이 실제로 최고 수준의 서구권 모델들에 얼마나 근접했느냐입니다. 새로 발표된 두 가지 데이터는 엇갈린 결과를 보여줍니다.

사용자의 선호도 평가를 기반으로 모델을 순위를 매기는 코드 아레나(Code Arena) 프론트엔드 벤치마크에서 Kimi K3는 1,679점을 기록하며 클로드 페이블 5(Claude Fable 5, 1,631점), GPT-5.6 솔(GPT-5.6 Sol, 1,618점) 및 기타 모든 테스트 대상 모델을 큰 차이로 누르고 1위를 차지했습니다. 중국 모델이 해당 벤치마크에서 정상을 차지한 것은 이번이 처음입니다.

하지만 고난도 수학 분야에서는 전혀 다른 결과가 나타납니다. Epoch AI의 데이터에 따르면, Kimi K3는 해당 벤치마크 중 가장 어려운 전문가 수준의 수학 과제인 '프론티어매스 티어 4(FrontierMath Tier 4)'에서 약 39%의 정확도만을 기록했습니다. 이는 특정 경우 오픈AI와 앤스로픽 모델들이 동일한 테스트에서 90%에 가까운 점수를 기록한 것과 비교하면 현저히 낮은 수치입니다.

원문 보기
원문 보기 (영어)
Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 19, 2026 Moonshot's AI model Kimi K3 is getting a lot of attention in the Western AI community. The big question is how close it actually gets to the best Western models. Two new data points paint a mixed picture. In the Code Arena: Frontend benchmark, which ranks models based on human preference ratings , Kimi K3 scores 1,679, beating Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and every other tested model by a wide margin. It's the first time a Chinese model has claimed the top spot on this benchmark. The picture looks different for hard math. According to data from Epoch AI , Kimi K3 hits only about 39 percent accuracy on FrontierMath Tier 4, the benchmark's hardest expert-level math tasks. Models from OpenAI and Anthropic score close to 90 percent there in some cases. Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: via X | Frontier Math Ask about this article… Search
관련 소식