메뉴
BL
The Decoder 12일 전

키미(Kimi) 오픈 모델 K3, GPT-5.6에 맞먹는 성능…중국 AI 초저가 시대는 끝났다

IMP
8/10
핵심 요약

중국의 AI 기업 키미(Kimi)가 최고 수준의 폐쇄형 모델과 맞먹는 성능을 지닌 2.8조 매개변수(Parameters) 규모의 멀티모달 오픈 모델 'K3'를 공개했습니다. 이 모델은 코딩 및 에이전트 작업에서 GPT-5.6 및 클로드 포블 5(Claude Fable 5)와 필적하는 성능을 보여주지만, 할루시네이션 비율이 증가했고 가격 역시 크게 상승하여 중국 AI의 '초저가' 시대가 저물었음을 시사합니다.

번역된 본문

키미(Kimi)가 공개한 오픈 모델 K3는 GPT-5.6 Sol 및 Fable 5에 근접한 성능을 보여주며, 동시에 중국산 AI의 초저가 시대가 끝났음을 알립니다. Matthias Bastian (2026년 7월 16일)

핵심 요약:

  • 키미(Kimi)는 896명의 전문가(Experts), 2.8조 개의 매개변수(Parameters), 100만 토큰의 컨텍스트 윈도우를 갖춘 전문가 혼합(Mixture-of-Experts) 기반의 오픈 웨이트(Open-weight) 멀티모달 모델인 K3를 출시했습니다. 전체 가중치(Weights)는 7월 말에 공개될 예정입니다.
  • 키미의 자체 벤치마크에서 K3는 클로드 포블 5(Claude Fable 5) 및 GPT 5.6 Sol에 근접한 성능을 보였으며, 테스트된 다른 모든 시스템을 큰 차이로 압도했습니다.
  • 독립 테스트 기관인 Artificial Analysis의 결과 역시 이를 대체로 확인해 주었지만, K3의 환각(할루시네이션, Hallucination) 비율은 이전 모델에 비해 증가했습니다.
  • 입력 100만 토큰당 3달러, 출력 100만 토큰당 15달러로 책정된 K3의 가격은 이전 모델들보다 훨씬 비싸지만, Sonnet 5와 같은 서구권 중급 모델과 비슷한 수준입니다.
  • 작업 당 비용은 약 0.94달러로, GPT-5.6 Sol과 유사하며 Opus 4.8의 절반 가격입니다.

키미(Kimi)는 2.8조 개의 매개변수와 100만 토큰의 컨텍스트 윈도우를 지원하는 멀티모달 모델인 K3를 출시했습니다. 이 모델은 회사의 자체 벤치마크에서 최고 수준의 기존 폐쇄형 모델들과 동등한 성능을 보여줍니다.

키미에 따르면, 새로운 플래그십 모델 K3는 총 2.8조 개의 매개변수를 보유하고 있으며, 이미지와 비디오를 네이티브로 처리하고 100만 토큰의 컨텍스트 윈도우를 지원합니다. 키미는 K3를 약 3조 매개변수 범위 내의 최초의 오픈 모델이라고 부릅니다. 전체 모델 가중치는 7월 27일까지 공개될 예정입니다. 이 모델은 오랜 시간이 걸리는 프로그래밍 작업, 지식 노동 및 복잡한 추론을 타겟팅합니다.

키미의 자체 벤치마크에서 K3는 최고 수준의 폐쇄형 모델인 클로드 포블 5(Claude Fable 5)와 GPT 5.6 Sol에는 뒤처지지만, 클로드 오퍼스(Claude Opus) 모델과 중국 경쟁사인 GLM-5.2를 포함한 테스트된 모든 다른 시스템을 물리쳤습니다. 회사에 따르면, 모든 결과는 키미에서 나온 것이며 최대 또는 높은 사고 강도(thinking intensity)를 통해 달성되었습니다.

총 35개의 테스트 중에서 K3는 약 7번 1위를 차지했고, 나머지 대부분에서 2위 또는 3위를 기록했습니다. 개별 테스트에서는 Fable 5가 가장 많이 이겼습니다. 거의 모든 벤치마크에서 K3는 Opus 4.8, GPT 5.5, GLM 5.2를 큰 차이로 이겼습니다. 벤치마크에 따라 KimiCode, Claude Code 또는 Codex 중 하나의 에이전트 시스템이 사용되었습니다. 즉, 모든 결과가 동일한 조건에서 수집된 것은 아닙니다.

Artificial Analysis, K3의 뛰어난 성능 확인... 단, 할루시네이션 비율은 증가

독립 테스트 연구소인 Artificial Analysis는 키미 K3에 대한 첫 번째 평가를 발표했습니다. 이 모델은 Artificial Analysis 지능 지수에서 57점을 획득하여 Opus 4.8 및 GPT-5.5와 동등한 수준이지만 여전히 Fable 5 및 GPT-5.6 Sol에는 뒤처집니다. 이는 키미의 주장과 대체로 일치합니다.

에이전트 작업에서 K3는 GDPval v2에서 1,668의 엘로(Elo) 점수를 기록했으며, 이는 K2.6의 1,190점에서 크게 도약한 수치입니다. GLM-5.2(1,514), GPT-5.5(1,494), 클로드 오퍼스 4.8(1,600)을 뛰어넘었지만 여전히 클로드 포블 5(1,760)에는 미치지 못합니다. K3는 또한 Zapier의 에이전트 SaaS 워크플로 평가 버전인 AutomationBench-AA에서 53%의 점수로 1위를 차지했습니다.

개인용 장기 지식 노동 평가인 AA-Briefcase에서 K3는 전체 엘로 1,547점을 기록하여 K2.6에서 732점이나 상승했습니다. 오직 클로드 포블 5만이 더 높은 점수를 받았습니다. Artificial Analysis는 K3가 뛰어난 두루 갖춘(Well-rounded) 모델이라고 평가하며, 루브릭 평가 및 분석 품질이 Fable 5의 수준에 근접한다고 밝혔습니다. 다만 발표 퀄리티 부분에서는 여전히 GPT-5.6 Sol이 선두를 달리고 있습니다.

K3의 정확도는 AA-Omniscience Index에서 33%에서 46%로 향상되어 전체 점수를 +6에서 +18로 끌어올렸습니다. 하지만 할루시네이션 비율은 39%에서 51%로 증가하여, 더 많은 질문에 올바르게 답변함에도 불구하고 거짓 정보를 생성하는 경우가 많아졌습니다.

K3, 코드 조각을 넘어선 전체 개발 프로젝트를 타겟팅

키미에 따르면, 이 모델의 주요 사용 사례는 최소한의 인간 개입으로 이루어지는 장기적인 소프트웨어 개발입니다. K3는 대규모 코드베이스를 분석하고, 터미널 도구를 조정하며, 수많은 작업 단계에 걸쳐 작업에 집중할 수 있도록 구축되었습니다.

원문 보기
원문 보기 (영어)
Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 16, 2026 Kimi (eigener Screenshot) Key Points Kimi has released K3, a multimodal open-weight model built on a mixture-of-experts architecture with 896 experts, 2.8 trillion parameters, and a context window of one million tokens. Full weights are expected by the end of July. In Kimi's own benchmarks, K3 comes close to Claude Fable 5 and GPT 5.6 Sol but beats all other tested systems by a wide margin. Independent testing by Artificial Analysis largely confirms these results, though K3's hallucination rate increased compared to its predecessor. At $3 per million input tokens and $15 per million output tokens, K3 is much pricier than its predecessor but comparable to Western mid-range models like Sonnet 5. Per-task costs land around $0.94, similar to GPT-5.6 Sol and about half the price of Opus 4.8. Ask about this article… Search Kimi is launching K3, a multimodal model with 2.8 trillion parameters and a context window of one million tokens. In the company's own benchmarks, it performs on par with leading proprietary models. According to Kimi, the new flagship model K3 has 2.8 trillion total parameters, processes images and video natively, and supports a context window of one million tokens. Kimi calls K3 the first open model in the roughly 3 trillion parameter range. Full model weights are scheduled for release by July 27. The model targets long-running programming tasks, knowledge work, and complex reasoning. In Kimi's own benchmarks , K3 still trails the top proprietary models Claude Fable 5 and GPT 5.6 Sol but beats every other system tested, including the Claude Opus models and Chinese rival GLM-5.2. All results come from Kimi and were achieved at maximum or high thinking intensity, according to the company. Ad Across all 35 tests, K3 took first place about seven times and landed second or third in most of the rest. Fable 5 won the most individual tests. In nearly every benchmark, K3 beat Opus 4.8, GPT 5.5, and GLM 5.2 by a wide margin. Depending on the benchmark, one of three agent systems was used: KimiCode, Claude Code, or Codex. That means the results weren't all collected under identical conditions. Ad DEC_D_Incontent-1 Artificial Analysis confirms K3's strong performance but flags higher hallucination rate Independent testing lab Artificial Analysis has published its first evaluation of Kimi K3. The model scores 57 on the Artificial Analysis Intelligence Index, putting it on par with Opus 4.8 and GPT-5.5 but still behind Fable 5 and GPT-5.6 Sol. That largely lines up with Kimi's own claims. On agentic tasks, K3 reaches an Elo rating of 1,668 on GDPval v2, a big jump from K2.6's 1,190. It beats GLM-5.2 (1,514), GPT-5.5 (1,494), and Claude Opus 4.8 (1,600), though it still falls short of Claude Fable 5 (1,760). K3 also takes the top spot on AutomationBench-AA, Artificial Analysis's version of Zapier's agentic SaaS workflow evaluation, with a score of 53 percent. Ad On AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reaches an overall Elo of 1,547, up 732 points from K2.6. Only Claude Fable 5 scores higher. Artificial Analysis calls K3 well-rounded, with rubric scoring and analytical quality close to Fable 5's level. GPT-5.6 Sol still leads on presentation quality, though. K3's accuracy rate improved from 33 percent to 46 percent on the AA-Omniscience Index, pushing the overall score from +6 to +18. But its hallucination rate climbed from 39 percent to 51 percent, meaning K3 fabricates more answers even as it gets more questions right. Ad DEC_D_Incontent-2 K3 targets full development projects beyond code snippets According to Kimi, the model's primary use case is long-running software development with minimal human oversight. K3 is built to analyze large codebases, coordinate terminal tools, and stay focused on a task across many work steps. Ad The model pairs programming with visual feedback: it examines screen captures, modifies code, then checks the visible output. Kimi calls this closed-loop system "Vision in the Loop" and positions it as a foundation for game development, UI design, and CAD. As demos, Kimi shows off a procedurally generated 3D open-world game that K3 reportedly built entirely in the browser using Three.js, WebGPU, and GPU Compute, along with an interactive black hole visualization . For the open-world demo, K3 procedurally generated the environment and used an external tool to create the 3D rider and horse models. Other demos include a simulation of the Long March 10 rocket launch and return, plus a Game Boy Advance emulator. K3 uses a mixture-of-experts architecture that activates only 16 of 896 experts at a time. It's paired with a new attention architecture called Kimi Delta Attention , which Kimi says enables up to 6.3x faster decoding for million-token contexts. "Attention residuals" reportedly boost training efficiency by about 25 percent while adding less than 2 percent in extra compute overhead. Chinese providers are raising prices for frontier models too According to the Kimi API docs , one million input tokens cost $0.30 with a cache hit and $3.00 without. One million output tokens, including reasoning, cost $15.00. These prices apply regardless of context length. Caching happens automatically, which makes unmodified long prefixes especially useful for agents and large codebases. That puts K3 well above the price level of its predecessor K2.6, which officially costs $0.16 per million tokens with a cache hit, $0.95 without, and $4.00 for output. Chinese providers aren't offering their frontier models at rock-bottom prices anymore either. Still, K3 is much cheaper than the top Western models and sits more in the upper midrange. Anthropic's new Sonnet 5 , for example, also costs $3 per million input tokens and $15 for output but delivers lower performance. Model Input (cache hit) Input (no hit) Output Kimi K3 $0.30 $3.00 $15.00 Kimi K2.6 $0.16 $0.95 $4.00 Claude Sonnet 5 $0.30 $3.00 $15.00 Claude Fable 5 $1.00 $10.00 $50.00 GPT 5.6 Sol $0.50 $5.00 $30.00 According to Artificial Analysis , K3 averages $0.94 per task on the Intelligence Index, close to GPT-5.6 Sol at $1.04 and about half the price of Opus 4.8 at $1.80. It's well above open-weight peers like GLM-5.2 ($0.32) and DeepSeek V4 Pro ($0.04), though. K3 also uses fewer tokens than its predecessor. It needed about 132 million output tokens to complete all nine evaluations, down from roughly 166 million for K2.6, a 21 percent reduction while scoring 13 points higher. Because of the much higher token prices, K3 will likely still cost more per task than K2.6 in most cases. Availability K3 is already available through Kimi.com , the mobile app for iOS, Android, and HarmonyOS, the Kimi Work desktop client (version 3.1.0 and later), and Kimi Code . On OpenRouter, the model is listed under the identifier "moonshotai/kimi-k3," though it's currently served there only through Moonshot itself. The open weights are expected by the end of July. For businesses, Kimi offers a separate version with member management and the ability to split personal and business accounts. A planned platform called Kimi Hosted Agent will provide isolated environments and runtimes for long-running tasks. Interested users can sign up for the waitlist now. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Kimi Blog
관련 소식