메뉴
BL
The Decoder • 3일 전

샤오미 저가 플래그십 AI, 오픈 모델 1위 등극… 앤스로픽은 '클로ude 데이터 무단 추출' 주장

IMP
8/10
핵심 요약

샤오미가 공개한 MiMo-V2.6-Pro는 분석업체 Artificial Analysis의 지능 지수에서 46점으로 오픈 모델 중 최강을 기록했으며, 태스크당 약 0.13달러라는 압도적으로 낮은 비용으로 지능-비용 효율의 파레토 프론티어에 올랐습니다. 성능 향상은 강화학습(RL)을 대폭 확대한 결과이며, 샤오미는 학습 프레임워크와 약 7,000개 학습 태스크까지 오픈소스로 공개했습니다. 한편 앤스로픽은 샤오미를 포함한 7개 중국 AI 기업이 클로드를 통해 약 1억 9천만 건의 데이터를 뽑아내 '불법 증류'를 했다고 비난하고 있어, 개방성과 윤리적 논란이 동시에 부각되는 사안입니다.

번역된 본문

샤오미의 저가 플래그십 AI가 오픈 모델 집계를 석권했고, 앤스로픽은 그 성과에 클로드가 사용됐다고 주장한다 (막시밀리안 슈라이너, 2026년 9월 22일)

핵심 요점

  • 샤오미의 새로운 MiMo-V2.6-Pro 모델은 경쟁 모델보다 훨씬 낮은 비용으로 현재 오픈 AI 모델 순위를 선도하고 있다.
  • 샤오미는 확대된 강화학습을 통해 성능 향상을 이뤄냈으며, 자사 도구와 학습 태스크를 공개했다.
  • 동시에 앤스로픽은 샤오미가 자체 모델 학습을 위해 클로드 모델에서 데이터를 부적절하게 빼내갔다고 비난하고 있다.

샤오미가 MiMo-V2.6 라인업을 공개했으며, 플래그십 모델은 태스크당 비용이 경쟁사의 극히 일부 수준으로 공개 가능 모델 중 정상에 올랐다. 이런 성과는 대폭 확대된 강화학습에서 나왔지만, 앤스로픽의 클로드에서 '차용'했다는 비난도 함께 받고 있다.

샤오미에 따르면 두 신모델 중 큰 버전인 MiMo-V2.6-Pro는 분석업체 Artificial Analysis의 지능 지수(Intelligence Index)에서 46점을 기록했다. 이는 Kimi K3, Qwen 등 경쟁 모델을 제치고 현재 공개된 AI 모델 중 가장 강력하다는 의미다. 진짜 핵심은 가격이다. 백만 입력 토큰당 0.435달러, 백만 출력 토큰당 0.87달러다. Artificial Analysis 계산으로는 단일 테스트 태스크에 약 0.13달러만 들어, 유사한 성능의 모델들보다 훨씬 싸다. 이로써 이 모델은 지능과 비용의 '파레토 프론티어'에 위치하게 됐다.

Pro는 혼합 전문가(MoE) 모델로 총 1.02조 개 파라미터를 갖지만 요청당 활성화되는 파라미터는 420억 개뿐이다. 이와 함께 더 작고 효율적인 MiMo-V2.6-Flash도 함께 공개됐다.

강화학습에서 나온 성과 샤오미는 성능 도약을 대폭 확대된 강화학습(RL), 즉 시행·피드백·보상을 통한 학습 덕분이라고 설명했다. 회사는 이 단계를 세 축으로 확장했다: 학습 단계당 더 많은 데이터, 더 다양한 태스크 환경, 솔루션 채점에 더 많은 컴퓨팅이다. 학습은 6일도 채 걸리지 않았으며, 비용은 Pro 약 262만 달러, Flash 약 85만 달러였다고 샤오미는 밝혔다. 코딩 벤치마크 DeepSWE에서는 Pro의 점수가 58.4에서 72.6으로, Flash는 48.8에서 65.7로 상승했다.

이런 규모에서 학습을 안정적으로 유지하기 위해 샤오미는 모델의 내부 분배 메커니즘을 고정하고, 과제를 실제로 해결하지 않은 채 보상을 얻어내는 모델의 꼼수인 '리워드 해킹'에 대한 여러 층의 방어 장치를 추가했다.

자동 채점기가 붙은 약 7,000개 태스크 모델과 함께 샤오미는 출력 속도가 최대 20배 빠른 Pro-UltraSpeed라는 초고속 변형 모델도 공개했다. 가장 주목할 점은 회사가 RL 툴킷을 오픈한다는 것이다. 기술 보고서, 완전한 학습 프레임워크, 추가 학습용 소형 모델, 그리고 소프트웨어 개발·사이버보안·사무 작업·웹 디자인 분야의 자동 채점기가 딸린 약 7,000개의 기성 학습 태스크, 그리고 음악 작곡용 약 1,000개 태스크까지 공개된다.

태스크는 다양한 출처에서 수집됐다. 일부 코드는 직원과 사용자 쿼리의 실제 GitHub 풀 리퀘스트에서 왔고, 다른 태스크 설명은 언어 모델이 생성했다. 사이버 태스크는 수만 건의 실제 소프트웨어 취약점 모음인 OSS-Fuzz에 기반하며, 사무 환경은 합성적으로 재구성됐다.

노골적인 개방성, 그리고 앤스로픽의 표적이 되다 이런 개방성 선언은 불과 2주 전 앤스로픽이 제기한 비난과 뚜렷한 대비를 이룬다. 앤스로픽은 위협 인텔리전스 보고서에서 2025년 12월부터 2026년 8월 사이 발견된 클로드 남용 사례를 분석하며, 이 모델을 겨냥한 캠페인과 연결된 7개 중국 연구소를 지목했다: 알리바바, 문샷 AI, 딥시크, 즈푸(Zhipu), 샤오미, 미니맥스, 센스타임이다. 이들 연구소는 자체 모델 학습을 위해 클로드의 능력을 빼내기 위해 약 1억 9천만 건의 대화를 생성한 것으로 알려졌으며, 앤스로픽은 이를 '불법 증류(illegal distillation)'라고 부른다. 샤오미는 이름이 명시적으로 거론됐다.

원문 보기
원문 보기 (영어)
Xiaomi's affordable flagship AI leads the open models, and Anthropic says Claude helped get it there Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 22, 2026 Xiaomi Key Points Xiaomi's new MiMo-V2.6-Pro model leads current rankings of open AI models while costing far less than the competition. Xiaomi achieved the performance jump through expanded reinforcement learning, and the company released its tools and training tasks openly. At the same time, Anthropic accuses Xiaomi of improperly siphoning off data through the Claude model to train its own model lineup. Ask about this article… Search Xiaomi has released its MiMo-V2.6 lineup, and the flagship model tops the charts among openly available models while costing a fraction of the competition per task. The gains come from a massively expanded round of reinforcement learning, though the company also stands accused of borrowing from Anthropic's Claude. According to Xiaomi, the larger of the two new models , MiMo-V2.6-Pro, scores 46 points on the Intelligence Index from analysis firm Artificial Analysis . That makes it the strongest openly available AI model right now, ahead of rivals like Kimi K3 and Qwen. The real kicker is the price: $0.435 per million input tokens and $0.87 per million output tokens. By Artificial Analysis's math, a single test task costs only about $0.13, a fraction of what similarly capable models charge. That puts the model on what's called the Pareto frontier of intelligence and cost. Ad Pro is a mixture-of-experts model with 1.02 trillion parameters, only 42 billion of which are active per request. Alongside it sits the smaller, more efficient MiMo-V2.6-Flash. Ad The gains come from reinforcement learning Xiaomi credits the jump in performance to heavily expanded reinforcement learning (RL), meaning training through trial, feedback, and reward. The company scaled this phase along three axes: more data per training step, more varied task environments, and more compute for grading the solutions. The run took less than six days, Xiaomi says, and cost about $2.62 million for Pro and $0.85 million for Flash. On the DeepSWE coding test, Pro's score climbed from 58.4 to 72.6, while Flash rose from 48.8 to 65.7. Ad To keep training stable at this scale, Xiaomi froze the model's internal distribution mechanism and added several layers of protection against "reward hacking," the tricks a model uses to game rewards without actually solving the task. About 7,000 tasks with automatic graders Along with the models, Xiaomi is shipping an especially fast variant called Pro-UltraSpeed, with up to 20 times the output speed. What stands out most is that the company is opening up its RL toolkit, including the technical report, the full training framework, a smaller model for further training , and about 7,000 ready-made training tasks with automatic graders for software development, cybersecurity, office work, and web design, plus roughly 1,000 tasks for music composition. Ad The tasks come from a mix of sources. Some of the code comes from real GitHub pull requests by employees and user queries, while other task descriptions are generated by a language model. The cyber tasks draw on OSS-Fuzz, a collection of tens of thousands of real software vulnerabilities, and the office environments are rebuilt synthetically. Ad Pointedly open, and in Anthropic's crosshairs This show of openness sits in sharp contrast to accusations Anthropic raised just two weeks earlier . In its threat intelligence report, Anthropic examined cases of Claude abuse discovered between December 2025 and August 2026, and named seven Chinese labs tied to campaigns against the model: Alibaba, Moonshot AI, DeepSeek, Zhipu, Xiaomi, MiniMax, and SenseTime. All told, the labs are said to have generated about 190 million exchanges to siphon off Claude's capabilities for training their own models, a technique Anthropic calls illegal distillation. Xiaomi shows up by name. In a case tagged GTG-16008, Anthropic tracked more than 400,000 exchanges over 20 days in March and April 2026, in which Xiaomi passed user conversations and coding sessions from its own MiMo models through OpenClaw and OpenCode to Claude, aiming to enrich training data for future models. The report offers almost no documentation of where the earlier training and teacher data for the internal distillation of teacher models came from. Put another way, the report claims Xiaomi recorded user conversations with its MiMo models and then fed them into Claude to extract training data, the same data that has now pushed it to the top of the open-model rankings. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: MiMo