메뉴
HN
Hacker News 50일 전

샤오미, 1T 모델 1000TPS 달성 'MiMo-V2.5-Pro-UltraSpeed' 공개

IMP
8/10
핵심 요약

샤오미와 TileRT는 1조(1T) 파라미터를 갖춘 대규모 언어 모델에서 초당 1,000토큰(1000 TPS)이라는 전례 없는 디코딩 속도를 달성한 'MiMo-V2.5-Pro-UltraSpeed'를 공개했습니다. 단순한 속도 향상을 넘어, AI가 실시간으로 의사결정에 개입하고 스스로 검증 및 교정하는 등 산업 패러다임을 근본적으로 전환시키는 의미를 지닙니다. 오는 6월 9일부터 23일까지 제한된 기업 및 개발자 대상으로 API 및 채팅 체험이 무료로 제공됩니다.

번역된 본문

블로그 참여하기 영어 간체중국어

블로그 참여하기 영어 간체중국어

2026년 6월 8일

MiMo-V2.5-Pro-UltraSpeed: 1조 파라미터(1T) 모델의 생성 속도를 1000 TPS로 끌어올리다

  1. 샤오미 MiMo-V2.5-Pro-UltraSpeed: 속도가 곧 궁극적인 경쟁력입니다 내연기관 시대의 포효하는 경주용 자동차부터 음속 장벽을 깨뜨린 소닉붐까지, 인류의 속도에 대한 갈망은 우리 DNA 자체에 각인되어 있습니다. AI 추론 속도 역시 마찬가지입니다. 이는 지능의 경계를 정의합니다. 모델이 충분히 빨라지면, 더 이상 기다려야 하는 도구가 아니라 여러분의 생각 그 자체를 연장하는 것이 됩니다. 실시간으로 대응하고, 즉각적으로 반복하며, 아무런 마찰 없이 협업하는 것이죠.

오늘 저희는 TileRT와의 협력을 통해 '샤오미 MiMo-V2.5-Pro-UltraSpeed'를 출시하게 되어 매우 기쁩니다. 이를 통해 1조 파라미터 모델에서 처음으로 초당 1,000 토큰(1000 TPS)의 디코딩 속도를 돌파했습니다!

  1. 기간 한정 액세스 · 신청제 MiMo-V2.5-Pro-UltraSpeed API가 한정된 기간의 프로모션 가격으로 동시에 출시됩니다. 기존 MiMo-V2.5-Pro 대비 3배의 비용이 들지만, 약 10배의 생성 속도를 제공합니다! 3배의 가격으로 10배의 출력 경험을 누리실 수 있습니다. (API 전용, Token Plan은 지원되지 않습니다.)

고속 추론 리소스가 제한되어 있어, MiMo-V2.5-Pro-UltraSpeed는 신청제를 통해 기간 한정으로 제공됩니다. 승인된 사용자는 평가 기간 동안 API에 액세스할 수 있으며, 이는 2026년 6월 9일부터 6월 23일 23:59(베이징 시간, UTC+8 / 08:59 PDT)까지만 가능합니다.

신청 방법 API 플랫폼: platform.xiaomimimo.com/ultraspeed 체험 슬롯은 제한되어 있으며, 신청이 승인을 보장하지는 않습니다. 실제 비즈니스 니즈가 있는 기업 및 전문 개발자를 우선적으로 승인할 예정입니다. 표준 모델 액세스는 MiMo-V2.5 모델 시리즈를 참조해 주십시오. UltraSpeed 모델에 대한 심층적인 비즈니스 파트너십은 business-mimo@xiaomi.com으로 문의해 주십시오.

채팅 체험 (평가 기간 무료) 승인된 사용자는 2주 유효한 무료 채팅 액세스 권한을 받게 됩니다. 접속 주소: ultraspeed.xiaomimimo.com 리소스 제약 하에서 품질과 공정성을 보장하기 위해 다음 규칙이 적용됩니다. 각 계정은 하루에 최대 10번까지 대기열에 들어갈 수 있습니다. 각 세션은 30분으로 제한됩니다. 5분 이상 유휴 상태인 세션은 자동으로 해제됩니다.

  1. 1000 tokens/s: 단순히 빠른 것이 아니라, 패러다임의 전환입니다 1조 파라미터(1T) 규모에서 1000 TPS를 돌파하는 것은 단순히 더 빠른 타자기가 되는 것 이상입니다. 이는 AI 애플리케이션 패러다임을 근본적으로 뒤흔듭니다.

첫째, 속도 그 자체가 지능으로 변모하기 시작합니다. 이전에는 어려운 문제에 직면했을 때 '단 하나의 정답을 기다리고 그것이 맞기를 바라는' 수밖에 없었습니다. 이제 동일한 실제 시간 내에 모델이 수십 개의 추론 경로를 병렬로 실행(Best-of-N / Tree Search)하고 백그라운드에서 자동으로 검증 및 자체 수정을 수행할 수 있습니다. 즉, 원시적인 속도를 사용해 사고의 깊이를 생성하고 추론 품질을 직접적으로 끌어올리는 것입니다.

둘째, 코딩 에이전트(Coding Agents)의 생산성 상한선을 완전히 해방합니다. 이전에는 AI가 코드를 작성하게 하는 것이 곧 개발자가 화면 앞에서 추론 지연 시간의 병목 현상을 겪으며 고통스럽게 기다리는 것을 의미했습니다. 1000 TPS에서는 코드 생성 속도와 생산 효율성이 패러다임 수준의 가속도를 경험하게 됩니다.

가장 중요한 것은, 1조 파라미터 모델이 이제 실시간 의사결정 루프에 진입할 수 있다는 점입니다. 밀리초 수준의 '생각-응답' 주기를 통해 1T 플래그십 모델이 시간이 촉박한 시나리오에 매끄럽게 연결될 수 있습니다. 고빈도 정량적 트레이딩 신호 생성, 즉각적인 사기(Fraud) 차단, 지능형 입찰, 실시간 대화형 인터랙션 등이 그 예입니다. 그리고 이러한 힘이 생사가 오가는 상황에서 수술 보조 및 의료 영상 분석에 적용될 때, AI의 속도는 더 이상 단순한 효율성의 지표가 아닙니다. 죽음과의 경주에서 기술이 인류를 돕는 강력한 무기가 됩니다. 수술대 위에서 AI가 병변 분석 및 위험 예측을 완료하는 데 절약하는 매 초는 외과의사에게 한층 더 높은 자유도를 부여합니다.

이러한 사실은 속도의 궁극적인 의미는 단순히 생산성을 높이는 것이 아니라, 기술이 인류의 삶을 보존하고 돕도록 하는 데 있다는 우리의 확신을 더욱 굳건하게 만듭니다.

원문 보기
원문 보기 (영어)
Blog Join us English 简体中文 Blog Join us English 简体中文 June 8, 2026 MiMo-V2.5-Pro-UltraSpeed: Pushing 1T-Parameter Model Generation Speed to 1000 TPS 1. Xiaomi MiMo-V2.5-Pro-UltraSpeed: Speed is the Ultimate Edge From the first roaring racer of the combustion age to the sonic boom that shattered the sound barrier, humanity's hunger for speed is written into our very DNA. The speed of AI reasoning is no different — it defines the boundaries of intelligence itself. When a model is fast enough, it ceases to be a tool you wait on and becomes an extension of your own thinking: responding in real time, iterating in an instant, collaborating without friction. Today, we are thrilled to release Xiaomi MiMo-V2.5-Pro-UltraSpeed in collaboration with TileRT, breaking the 1000 tokens/s decode speed on a 1-trillion-parameter model for the first time! 2. Limited-Time Access · Application-Based The MiMo-V2.5-Pro-UltraSpeed API launches simultaneously at a limited-time promotional price — 3× the cost of MiMo-V2.5-Pro, but delivering approximately 10× the generation speed! 3× the price, 10× the output experience. (API only; Token Plan not supported.) Due to limited high-speed inference resources, MiMo-V2.5-Pro-UltraSpeed will be available through an application-based, limited-time window. Approved users can access the API during the trial period, available only from June 9 to June 23, 2026, 23:59 (Beijing Time, UTC+8 / 08:59 PDT) . How to Apply API platform: platform.xiaomimimo.com/ultraspeed . Trial slots are limited — submission does not guarantee approval. We will prioritize enterprises and professional developers with genuine business needs. For standard model access, please follow the MiMo-V2.5 model series. For in-depth business partnerships for the UltraSpeed model, contact business-mimo@xiaomi.com . Chat Experience (Free During Trial) Approved users will receive free Chat access valid within the two-week window. Entry point: ultraspeed.xiaomimimo.com To ensure quality and fairness under resource constraints, the following rules apply: each account may enter the queue up to 10 times per day; each session is capped at 30 minutes; sessions idle for more than 5 minutes will be automatically released. 3. 1000 tokens/s: Not Just Fast, But a Paradigm Shift At the trillion-parameter (1T) scale, breaking 1000 tps is far more than a faster typewriter — it fundamentally disrupts AI application paradigms. First, speed itself begins to transmute into intelligence. Previously, when facing a hard problem, you could only "wait for one answer and pray it's correct." Now, within the same wall-clock time, the model can run dozens of reasoning paths in parallel (Best-of-N / Tree Search), automatically verifying and self-correcting in the background — using raw speed to generate depth of thought, directly elevating reasoning quality. Second, it completely unleashes the productivity ceiling of Coding Agents. Before, having AI write code meant developers painfully waiting in front of screens, bottlenecked by inference latency. At 1000 tps, code generation speed and production efficiency undergo a paradigm-level acceleration. Most importantly, trillion-parameter models can now enter real-time decision loops. Millisecond-level "think-respond" cycles allow 1T flagship models to seamlessly plug into time-critical scenarios — high-frequency quantitative trading signal generation, instant anti-fraud interception, intelligent bidding, and real-time interactive dialogue. And when this power is brought to surgical assistance and medical imaging analysis in life-or-death situations, AI speed is no longer just a metric of efficiency — it becomes a chip in the race against death. On the operating table, every second AI saves in completing lesion analysis and risk prediction gives the surgeon one more degree of freedom. This deepens our conviction that the ultimate significance of speed is not merely boosting productivity, but enabling technology to help humanity live better. 4. Extreme Model-System Codesign Achieving 1000+ tokens/s generation speed with a 1T flagship model is not the breakthrough of a single technique — it is the product of deep collaboration and extreme Codesign between the MiMo model team and the TileRT system team. The industry's current approach to similar extreme speeds typically relies on specialized hardware — Cerebras's Wafer-Scale integration or Groq's pure on-chip SRAM custom architecture. We chose a different path: achieving even more impressive inference speed on commodity GPUs through model-system codesign alone. On the model side, we applied FP4 quantization targeting the bandwidth bottleneck of commodity hardware, dramatically shrinking model size and reducing memory-access overhead; simultaneously, we introduced DFlash, an efficient speculative decoding method based on block-level masked parallel prediction , substantially increasing the accepted token length per verification step. On the system side, TileRT perfectly adapts to the dynamic characteristics of these algorithms, delivering a tailor-made compilation engine and compute kernels optimized specifically for the novel quantization and speculative decoding pipeline. Through this extreme Codesign, we achieved 1000+ tokens/s output from a 1T model using just a single standard 8-GPU commodity node. 3.1 FP4 Quantization At the trillion-parameter (1T) scale, traditional 8-bit (FP8 / INT8) or even 16-bit inference imposes prohibitive memory footprint and bandwidth pressure. Reducing parameter bit-width directly contributes to decoding speed. We therefore adopt the widely validated, virtually lossless FP4 (MXFP4) quantization format [1] . However, naively applying FP4 across the entire model causes degradation in complex reasoning, logic, and code generation. Given the MoE (Mixture of Experts) architecture of Xiaomi MiMo-V2.5-Pro — where Experts constitute the vast majority of parameters and exhibit the highest tolerance to quantization — we selectively quantize only the MoE Experts to FP4 while preserving original precision for all other modules. Through FP4 QAT (Quantization-Aware Training), we dramatically reduce model size and maximize hardware bandwidth utilization while keeping the model's overall capability essentially on par with the original, as shown below: 3.2 DFlash Speculative Decoding Traditional Speculative Decoding relies on a small draft model to "guess" subsequent tokens, which the large model then verifies. This transforms autoregressive generation (1 token per forward pass) into parallel multi-token generation, with rejection sampling during verification ensuring lossless output quality. However, its bottleneck lies in the draft model's quality determining the acceptance rate, while a stronger draft model incurs higher compute overhead — a fundamental tension. To break this deadlock, we adopt DFlash , an innovative block-level masked parallel prediction method from the research community [2] : the draft model fills an entire block of masked positions in a single forward pass, fundamentally eliminating the serial constraint of "autoregressive drafting." We deployed this approach on MiMo-V2.5-Pro with custom optimizations tailored for trillion-scale MoE and long-context scenarios. Using the Muon second-order optimizer and model self-distillation, we ensure that compact mask blocks still deliver ideal acceptance rates while compressing draft-stage overhead to near its theoretical minimum: The draft model exclusively uses Sliding Window Attention (SWA), naturally aligning with the SWA design of the MiMo-V2 series. This eliminates dependency on complete prefixes, reducing per-prediction compute from context-length-linear to constant. During training, mask-signal sampling is pushed down to GPU-local shards, enabling a single sequence to produce tens of thousands of independent training signals covering diverse context positions in one step — aligning with the long-c