메뉴
BL
TechCrunch AI 6일 전

전문가들 "Kimi K3의 성능, 단순 기술 도용 때문은 아냐"

IMP
8/10
핵심 요약

미국 백악관은 중국의 문샷(Moonshot)이 수출 통제를 받는 칩을 사용하고 미국 기업 앤스로픽(Anthropic)의 모델을 불법적으로 증류(Distillation)하여 오픈소스 모델인 Kimi K3를 만들었다고 비난했습니다. 하지만 AI 전문가들은 Fable 모델이 공개된 지 얼마 되지 않았으며, Kimi K3의 압도적인 성능은 단순한 데이터 증류나 지도 미세조정(SFT)만으로는 불가능하다고 반박했습니다. 대규모 강화학습(RL)과 자체적인 고도화된 훈련이 있었을 것으로 분석하며, 증류 기법의 한계와 최근 중국 모델들의 실제 기술력을 보여주는 사례로 주목받고 있습니다.

번역된 본문

현재 사용 가능한 가장 큰 오픈 웨이트(Open-weight) 대형 언어 모델(LLM)인 Kimi K3를 개발한 중국 기업 문샷(Moonshot)이 미국으로의 수출이 승인되지 않은 칩을 사용하면서 앤스로픽(Anthropic)의 Fable LLM을 복사하여 모델을 구축했다고 미국 백악관 과학 자문관 마이클 크라치오스(Michael Kratsios)가 밝혔다. AI 업계를 뒤흔든 중국 오픈 웨이트 모델 금지 논의가 보도되는 가운데, 크라치오스는 "미국의 독점 기술을 훔치고 미국의 연구를 훼손하려는 대규모의 은밀한 산업 증류(distillation) 행위는 용납될 수 없다"고 적었다. 문샥은 훈련 과정에 대한 질문에 답하지 않았으며, 크라치오스도 주장의 출처에 대해 더 자세한 내용을 공유하지 않았다.

크라치오스의 게시물은 "우리는 많은 중국 모델에서 미국 대형 언어 모델의 워터마크를 발견하고 있으며, 이는 용납할 수 없는 일이다"라는 스캇 베센트(Scott Bessent) 미국 재무부 장관의 발언과 맥락을 같이한다. 이 워터마크가 정확히 무엇으로 구성되어 있는지는 불분명하며, 재무부 또한 조회에 대한 응답을 내놓지 않았다.

하지만 전문가들은 LLM의 내부 작동 원리를 파악하고 그 기능을 복사하기 위해 모델에 쿼리(질의)를 날리는 과정인 '증류(Distillation)'가 Kimi K3가 보여주는 고도화된 성능을 책임졌을 것이라는 점에 대해 회의적이다. 스노클 AI(Snorkel AI)의 공동 창립자이자 라우드 연구소(Laude Institute)의 연구원인 브레이든 핸콕(Braden Hancock)은 테크크런치(TechCrunch)에 "Fable이 공개된 직후 단순한 증류만으로 이렇게 강력한 모델을 이토록 빨리 만들어냈을 거라고는 생각하지 않는다"며 "솔직히 시간이 전혀 부족했다. Fable은 7월 1일에야 공개되었다. 불과 2주 만에 그 많은 양의 데이터를 증류하고, 모델을 훈련시킨 뒤 출시하는 것은 불가능하다"고 지적했다.

앨런 AI 연구소(Allen Institute for AI)의 AI 연구원 네이선 램버트(Nathan Lambert)는 어제 공개된 팟캐스트에서 "중국 모델들이 최고 수준(frontier)에 근접하고 훈련 방식이 강화학습으로 전환됨에 따라, 증류의 효과는 시간이 지날수록 점점 줄어들고 있다는 생각이었다"고 말했다. 그는 "만약 (증류로 쉽게 따라잡을 수 있는 것이라면) 누구나 GLM이나 K3의 데이터를 증류에 사용해 쉽게 따라잡을 수 있을 것이다. 하지만 우리는 지도 미세조정(Supervised Fine-Tuning, SFT)만으로는 이를 달성하지 못했으며, 앞으로도 그럴 것이다"라고 덧붙였다.

증류를 수행하려면 연구실은 사후 훈련(Post-training)에 사용할 데이터를 생성하기 위해 대상 모델에 체계적으로 쿼리를 보내야 한다. 때로는 모델이 문제를 어떻게 해결하는지 이해하기 위해 모델에게 사고 과정(Chain-of-thought)을 명시적으로 표현해달라고 요구하기도 한다. 다른 경우에는 감독 미세조정(SFT)이라는 과정을 통해 모델의 프롬프트와 응답이 새로운 모델을 훈련시키는 데 사용된다. 이 미세조정 과정으로 인해 제3자가 만든 것으로 보이는 모델이 자신을 클로드(Claude)라고 주장하는 상황이 발생할 수 있다. 램버트의 견해에 따르면 모델은 미세조정 과정에서 '예의 범절'을 배우게 된다.

하지만 램버트는 모델이 점점 더 복잡해짐에 따라 SFT의 이점은 덜 중요해지고 있다고 말한다. Fable과 유사한 수준의 능력을 증류하려면 강화학습(Reinforcement Learning) 기술이 필요할 가능성이 높다. 많은 경우 이는 대형 모델의 에이전트(Agent)가 소형 모델의 응답을 평가하고 그 점수를 바탕으로 조정하는 것을 의미한다. 더욱 진보된 기술은 더 막대한 인프라를 요구한다. 대규모 강화학습 실행에는 수천만 개의 에이전트가 필요할 수 있다. 최고 수준의 연구실 API를 사용해 이 작업을 수행하는 것은 "미친 듯이 비쌀 것이며, 이 모델들이 상당히 느리기 때문에 아마 시간 병목 현상이 발생할 것이다. 솔직히 말해 성능 향상을 보장하지도 않을 것이다."

이전의 최고 수준 모델들이 Kimi의 발전에 기여했을 가능성은 있어 보인다. 앤스로픽은 올해 초 문샷, 딥시크(DeepSeek), 미니맥스(MiniMax)가 자사 모델을 체계적으로 증류했다고 공개적으로 비난한 바 있다. 앤스로픽은 IP 주소 및 기타 메타데이터를 통해 자사 모델과 해당 기업들로 식별된 사용자 간의 수백만 건의 교환 내역을 발견했다고 밝혔다. 이러한 쿼리들은 "정상적인 사용 패턴과는 달리 합법적인 사용이 아닌 의도적인 기능 추출을 반영하고 있었다"고 앤스로픽은 전했다. 앤스로픽은 Fable 증류와 관련된 테크크런치의 질의에 응답하지 않았다. 하지만 중국뿐만 아니라 전 세계 AI 기업들 사이에서 증류는 흔한 관행으로 여겨진다. 일론 머스크는 증언에서...

원문 보기
원문 보기 (영어)
White House science advisor Michael Kratsios said that Moonshot, the Chinese company behind the Kimi K3, the largest available open-weight LLM, built its model by copying Anthropic's Fable LLM while using chips that aren't cleared for export to China. "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote , amid reported discussions about banning Chinese open-weight models that have roiled the AI sector. Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations. Kratsios' tweet echoed comments from Treasury Secretary Scott Bessent that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." It's not clear what those watermarks consist of, and the Treasury Department did not respond to a query. However, experts are skeptical that distillation—the process of querying an LLM to determine its inner workings and copy its capabilities—is responsible for the advanced capabilities that Kimi K3 displays. "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. "There's just not even frankly time, right? Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks." "I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]," Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday. "[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won’t see this, from supervised fine-tuning alone." Performing distillation requires a lab to systematically query its target model in order to generate data that can be used for post-training. Sometimes this explicitly involves asking the model to articulate its chain-of-thought to understand how it solves problems. Other times, the prompts and responses from a model are used to train a new model in a process called supervised fine-tuning, or SFT. It's this fine-tuning process that can result in a model ostensibly created by a third party claiming that it is Claude. Fine tuning is where, in Lambert's view, the "model picks up its manners." But Lambert says that the benefits of SFT are becoming less important as models become more complex. To distill Fable-like capabilities would likely require reinforcement learning techniques. In many cases, that means having an agent of the larger model grade the smaller model's responses, and adjusting based on the grade. The more advanced techniques also require more significant infrastructure. Large reinforcement learning runs can require tens of millions of agents. Using a frontier lab's API to do that "would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift." It seems likely that previous frontier models might have contributed to Kimi; Anthropic publicly accused Moonshot, DeepSeek and MiniMax of systematically distilling its models earlier this year. Anthropic said it discovered millions of exchanges between its models and users it identified at those companies through IP addresses and other meta data. Those queries were "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." Anthropic didn't respond to TechCrunch's queries about Fable distillation. However, distillation is seen as common among AI companies, not just in China. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, and that the practice was common in the industry. The line between distillation and developing synthetic data sets, for example, can be fairly blurry. "[I]n general, Americans are understating the technical expertise of these Chinese teams," Hancock said. "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China's progress would slow, but would still continue. They're not just riding coattails here." It's also hard to disentangle distillation from the second part of Kratsios' comment — that Moonshot had obtained advanced Nvidia Chips, Grace Blackwell 300s, and also accessed GB300 equipped-servers in Thailand. Those chips are banned from export to China, but a black market exists, according to Sam Bresnick, a research fellow at Georgetown's Center for Security and Emerging Technology. In May, the founder of Supermicro, a US server builder, was indicted for smuggling advanced chips into China. "I am a proponent of know your customer laws for data centers across the world," Bresnick said. "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they're doing." President Joe Biden's Department of Commerce proposed federal know-your-customer rules for data centers in 2024, but no further progress appears to have been made under Donald Trump. Exporters shipping advanced chips abroad, however, are supposed to ensure they are only used for approved purposes. Topics AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Tim Fernholz Senior Reporter Tim Fernholz is a journalist who writes about technology, finance and public policy. He has closely covered the rise of the private space industry and is the author of Rocket Billionaires: Elon Musk, Jeff Bezos and the New Space Race. Formerly, he was a senior reporter at Quartz, the global business news site, for more than a decade, and began his career as a political reporter in Washington, D.C. You can contact or verify outreach from Tim by emailing tim.fernholz@techcrunch.com or via an encrypted message to tim_fernholz.21 on Signal. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 toda y! REGISTER NOW Most Popular Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents Amanda Silberling Light made a flip phone — it's colorful and it's cheap Amanda Silberling AI music generator Suno breach affects 55M users, per Have I Been Pwned Zack Whittaker Anthropic's landmark $1.5B copyright settlement is approved Kirsten Korosec Google is working on a new AI chip designed to make Gemini more efficient Lucas Ropek X relaunches a rebuilt Android app after year-long effort Sarah Perez Judge pauses $110B Paramount-Warner Bros. merger Aisha Malik