메뉴
HN
Hacker News 56일 전

마이크로소프트 AI 추론 모델 MAI-Thinking-1 공개

IMP
9/10
핵심 요약

마이크로소프트 AI가 소프트웨어 엔지니어링 및 수학 추론 벤치마크에서 최고 수준의 성능을 기록한 중간 크기 모델 'MAI-Thinking-1'을 발표했습니다. 이 모델은 타사 모델의 지식 증류 없이 클린 데이터로 처음부터 학습되었으며, 블라인드 평가에서 Sonnet 4.6보다 높은 선호도를 보였습니다. 적은 추론 자원으로도 코딩 에이전트 및 복잡한 작업 수행에 탁월한 효율성을 보여주는 것이 가장 큰 특징입니다.

번역된 본문

Copilot --> Models: MAI-Thinking-1 소개 (초지능 팀, 2026년 6월 2일)

오늘 저희는 Microsoft AI의 추론 모델인 MAI-Thinking-1을 소개합니다. 이 모델은 동일한 파라미터 규모(Weight class)의 모델들 중 가장 강력한 성능을 자랑하는 중간 크기(Medium-sized) 모델입니다. 핵심 소프트웨어 엔지니어링 벤치마크에서 최고 수준 모델들과 동등한 성능을 발휘하며, 고도화된 수학적 추론 능력을 입증했고, 블라인드 인간 평가에서 Sonnet 4.6보다 더 높은 선호도를 받았습니다.

이 모델은 타사 모델의 지식 증류(Distillation) 없이, 기업용 등급의 깨끗하고 상업적으로 라이선스가 허가된 데이터를 바탕으로 처음부터(From the ground up) 학습되었습니다. MAI-Thinking-1은 인간과 조직을 대체하는 것이 아니라 그들을 돕기 위해 설계된 고급 AI 역량인 '휴머니스트 초지능(Humanist Superintelligence)'을 구축하려는 저희의 광범위한 노력의 일환입니다. 이 모델은 '무엇을 할 수 있는가'와 '어떻게 구축되었는가'라는 두 가지 측면에서 모두 중요한 의미를 갖습니다.

언덕 오르기 기계 (The Hill-Climbing Machine)

저희는 단순한 단일 모델 그 이상을 기쁘게 생각합니다. 모델 개발의 모든 구성 요소를 지속적으로 개선할 수 있도록 공동 설계된 파이프라인인 '언덕 오르기 기계(Hill-Climbing Machine)'를 소개합니다. 이는 시간이 지남에 따라 역량이 지속적이고 안정적으로 향상되도록 만드는 시스템입니다. 더 나은 데이터, 더 강력한 보상, 더 유능한 환경, 그리고 더 많은 컴퓨팅 자원을 흡수할 수 있는 반복 가능한 시스템을 구축하는 것이 목표입니다. 저희의 철학을 이끄는 세 가지 주요 기둥은 다음과 같습니다.

첫째, 역량은 물려받는 것이 아니라 학습해야 합니다(Capabilities should be learned, not inherited). 지식을 물려받는 것이 더 빠르게 습득할 수 있는 방법일 수 있지만, 이는 실제 상황에서 필수적인 제어 가능성(Steerability)이 부족합니다. 모방자는 근본적으로 교사의 설계 선택에 얽매여 있으며 새로운 상황에 적응하는 데 어려움을 겪습니다. MAI-Thinking-1은 타사 모델로부터의 증류 없이 학습되었으며, 이를 통해 모델이 주어진 작업을 진정으로 학습하도록 강제했습니다.

둘째, 깨끗한 데이터입니다. MAI-Thinking-1은 적절히 라이선스가 부여된 깨끗한 데이터로 학습되었으며, 사전 학습(Pre-training) 과정에서 AI 생성 콘텐츠를 배제했습니다. 이는 품질, 출처, 제어 측면에서 매우 중요합니다. 모델을 형성한 요인을 파악할 수 없다면 모델의 동작을 온전히 이해하거나 신뢰할 수 있게 개선할 수 없습니다.

셋째, 전체 기술 스택의 자급자족입니다. 마이크로소프트(MSFT)의 자체 가속기와 모델을 공동 설계하는 것부터 강화 학습 프레임워크에 이르기까지, 저희는 인하우스 학습 인프라에 집중해 왔습니다. 이는 엔드투엔드로 시스템을 완전히 최적화하고 형성하여 요구 사항을 가장 잘 충족할 수 있도록 보장하는 '언덕 오르기 기계'를 구축하는 데 핵심적인 부분입니다.

강력한 소프트웨어 엔지니어링 성능을 갖춘 중간 크기 모델

MAI-Thinking-1은 350억(35B)개의 활성 파라미터와 약 1조(T)개의 총 파라미터를 가진 희소 전문가 혼합(Sparse Mixture of Experts) 모델로, 훨씬 큰 모델들보다 더 작은 추론 공간(Inference footprint)을 차지합니다. 그럼에도 불구하고 저희 모델은 SWE-Bench Pro에서 Claude Opus 4.6과 필적하는 성능을 보여줍니다. 모델 크기는 고급 코딩 지원이 어디에 배포될 수 있는지, 얼마나 자주 사용될 수 있는지, 예외적인 작업에서 일상적인 워크플로로 전환될 수 있는지를 결정하므로 개발자와 기업에게 매우 중요합니다.

저희는 에이전트 코딩(Agentic coding)에 필요한 학습 환경에 막대한 투자를 했습니다. 검증된 각 환경은 결정론적이고 실행 가능하며, 실제 테스트 스위트(Suite)를 통해 평가됩니다. 이를 통해 모델이 개발자가 실제로 수행하는 다단계 작업(코드 읽기, 파일 편집, 테스트 실행, 오류 관찰 및 중간 실수 복구 등)을 연습할 수 있도록 했습니다.

고도화된 수학적 추론 능력

MAI-Thinking-1은 AIME 2025에서 97.0%, AIME 2026에서 94.5%의 정확도를 달성하여 동일 규모 모델들 중 강력한 수학 및 과학적 추론 능력을 입증했습니다. 이러한 뛰어난 성능은 저희의 학습 루프가 자체 데이터, 보상 및 평가 프로세스를 바탕으로 처음부터 끝까지 올라가 실제 추론 성능 향상을 만들어낼 수 있으며, 이러한 지능이 시간이 지남에 따라 다른 도메인으로 일반화될 수 있다는 확신을 줍니다.

Sonnet 4.6 대비 블라인드 인간 평가에서 높은 선호도 획득

사람들은 모델이 작업을 이해하고 지시를 따르며 적절한 수준의 세부 사항을 사용하고 명확하게 작성하며 시간을 존중하는지 여부를 중요하게 생각합니다. 저희는 파트너 중 한 곳인 Surge의 전문 평가자 풀(Pool)을 활용하여 블라인드 인간 병렬 평가(Side-by-side evaluation)를 진행했습니다.

원문 보기
원문 보기 (영어)
Copilot --> Models Introducing MAI-Thinking-1 Superintelligence team June 2, 2026 Models Superintelligence team LI X FB Today we are introducing MAI-Thinking-1, Microsoft AI's reasoning model. It is a medium-sized model that stands among the strongest models in its weight class. It matches leading models on key software engineering benchmarks, demonstrates advanced mathematical reasoning capabilities, and is preferred to Sonnet 4.6 in our blind human side-by-side evaluations. We trained it from the ground up on enterprise grade, clean and commercially licensed data, without distillation from third-party models. MAI-Thinking-1 is a step in our broader work to build towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not to replace them. The model matters on both axes: what it can do, and how it was built. The Hill-Climbing Machine More than a single model, we are excited to introduce our Hill-Climbing Machine: a co-designed pipeline built to make every component of model development climbable, so capabilities improve continually and reliably over time. The aim is a repeatable system that can absorb better data, stronger rewards, more capable environments, and more compute. Three main pillars guide our philosophy. First, capabilities should be learned, not inherited. Although faster to acquire, inherited intelligence lacks the steerability essential for real world usage: an imitator is fundamentally tied to the design choices of its teacher and struggles to adapt to new situations. MAI-Thinking-1 was trained without distillation from third party models, forcing our model to truly learn the tasks at hand. Second, clean data. MAI-Thinking-1 was trained on clean and appropriately licensed data, with AI-generated content excluded from pre-training. This matters for quality, provenance, and control. If we cannot account for what shaped a model, we cannot fully understand its behavior or credibly improve it. Third, self-sufficiency across the entire stack. All the way from co-design of our models with MSFT’s own accelerators through to our reinforcement learning framework, we have focused efforts on in-house training infrastructure. This is a crucial part of building our hill-climbing machine, to ensure we can fully optimize and shape our systems end-to-end to best serve our needs. Medium-sized model, with strong software engineering performance MAI-Thinking-1 is a 35B-active, ~1T-total parameters, sparse Mixture of Experts model, a smaller inference footprint than much larger models. Despite this, our model is toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro. That matters for developers and enterprises because model size determines where advanced coding assistance can be deployed, how often it can be used, and whether it can move from exceptional tasks into daily workflows. We have invested heavily in the training environments needed for agentic coding. Each verified environment is deterministic, executable, and graded by real test suites. This gives the model practice on the kind of multi-step work developers actually do: reading code, editing files, running tests, observing failures, and recovering from intermediate mistakes. Advanced mathematical reasoning capabilities MAI-Thinking-1 reaches 97.0% on AIME 2025, and 94.5% on AIME 2026, showing strong mathematical and scientific reasoning for its weight class. Strong performance here gives us confidence that our training loop can create real reasoning gains – climbing all the way from the ground up - from our own data, rewards, and evaluation process, enabling this intelligence to generalize to other domains over time. Preferred in human side-by-sides vs. Sonnet 4.6 People care about whether a model understands the task, follows instructions, uses the right level of detail, writes clearly, and respects their time. We built a blind side by side human evaluation with one of our partners, Surge, using their pool of professional raters to measure various models on these traits. This comprised of 1,276 evaluations designed to test the model’s capabilities in a large variety of tasks across both single-turn and multi-turn conversations, with a focus on measuring how helpful the response to the user is and whether it actually advances the user’s goals. In these evaluations, users preferred MAI-Thinking-1 over Claude Sonnet 4.6. This has been a core focus of post-training. We want the model to be capable without being brittle, concise without being incomplete, and helpful without overreaching. Human preference data gives us a direct signal on whether benchmark improvements translate into better experiences for users. Enterprise ready MAI-Thinking-1 is built with enterprise readiness in mind. It supports long context with a 256k token window (enough to fit a 600 page document), function calling, and the flexibility to add developer instructions. We trained the model to follow multiple layers of instructions and aligned its default style to enterprise needs. It's compatible with the widely used Chat Completions API. All MAI models come with enterprise-grade security and compliance through Microsoft Foundry. Results We report results in two views: post-trained MAI-Thinking-1 evaluations, and pre-training metrics for our base model. Table 1. MAI-Thinking-1 metrics Post-trained model evaluation results on public STEM and agentic coding benchmarks. Other model numbers are taken from respective official model cards. Scores are percentages unless otherwise noted; dashes indicate unavailable model values. Table 2. Pre-training metrics Putting humans first We are building towards Humanist Superintelligence: advanced AI capabilities designed to serve people and organizations, not replace them. Our models must remain subordinate technologies under human control with the goal of upholding human autonomy and being helpful. That means our models must not refuse legitimate requests under the guise of safety and compliance as then they are not truly serving humans. Striking the delicate balance between being helpful and safe is not easy. For MAI-Thinking-1, we aimed to achieve this balance by treating unsafe compliance and unnecessary refusal as defects in the same reward construction where aggregation is based on severity of potential of harm. Safety is trained with the same reinforcement learning infrastructure used for capability, so safety rewards are part of the same hill-climbing loop ensuring safety is always aligned to the capabilities and not incidental. As a result, we see that our model can balance ensuring a safety bar on sensitive unsafe requests while also being helpful on non-sensitive content. Availability and access MAI-Thinking-1 is available in private preview on Microsoft Foundry today. It will be available in public preview on MAI Playground soon. Read the paper Build the Future With Us We’re a lean, fast-moving lab made up of some of the world’s most talented minds. We have an exciting roadmap of compute at MAI, which is ramping quickly and extensively. And we have an ambitious mission we truly believe in. We’re also fortunate to partner with incredible product teams giving our models the chance to reach billions of users and create immense positive impact. If you’re a brilliant, highly-ambitious and low ego individual, you’ll fit right in—come and join us as we work on our next generation of models! Explore all jobs Related Stories Building a hill-climbing machine: Launching seven new MAI models announcements Our values in operation: Health partnerships Towards Humanist Superintelligence announcements