메뉴
BL
The Decoder 56일 전

마이크로소프트 빌드 2026: 이미지 생성 최고, 추론 모델은 추격 중

IMP
9/10
핵심 요약

마이크로소프트가 '빌드 2026' 행사에서 자체 개발한 7종의 AI 모델과 기업 맞춤형 강화학습 기반의 '프론티어 튜닝(Frontier Tuning)' 기술을 발표했습니다. 특히 이미지 생성 모델은 구글을 제치고 2위를 차지했으며, 1/10의 비용으로 GPT-5.4 수준의 성능을 내는 튜닝 기술이 비용 효율성의 핵심으로 강조되었습니다. 또한 사용자의 업무 패턴을 학습해 백그라운드에서 능동적으로 일정 관리 등을 수행하는 상시 작동 에이전트 '스카우트(Scout)'를 최초로 선보이며 AI 자동화 생태계를 확장했습니다.

번역된 본문

빌드 2026: 마이크로소프트, 이미지 생성 부문에서 구글 제치고 1위... 추론 모델은 추격 중 막시밀리안 슈나이더(Maximilian Schreiner) | 2026년 6월 3일

핵심 요약:

  • 마이크로소프트는 자체 최초의 추론 모델인 MAI-Thinking-1을 포함해 7개의 자체 개발 AI 모델을 발표했습니다.
  • 벤치마크 결과, 이 모델은 Deepseek V3.2와 대체로 비슷한 수준입니다.
  • '프론티어 튜닝(Frontier Tuning)'이라는 새로운 방식을 통해 기업들이 강화학습을 사용해 자사 워크플로우에 맞게 모델을 적용할 수 있습니다. 마이크로소프트에 따르면 튜닝된 모델은 GPT-5.4 수준의 성능을 1/10의 비용으로 발휘합니다.
  • 일정 관리 및 회의 준비와 같은 사무 작업을 처리하는 상시 작동형 백그라운드 에이전트 '스카우트(Scout)'도 출시합니다.
  • 이러한 소프트웨어 발표와 함께 로컬 개발자 하드웨어 및 AI 에이전트를 위해 설계된 새로운 운영 체제(OS)도 함께 공개되었습니다.

빌드 2026에서 마이크로소프트는 최초의 추론 모델을 포함해 7개의 자체 개발 AI 모델을 발표했습니다. 또한 새로운 튜닝 방식과 자율 백그라운드 에이전트도 소개했습니다.

이번 발표의 핵심은 마이크로소프트 최초의 추론 모델인 'MAI-Thinking-1'입니다. 마이크로소프트 AI 총괄 머스타파 술레이만(Mustafa Suleyman)에 따르면, 이 모델은 1조 개의 매개변수(parameters)와 350억 개의 활성 매개변수, 128,000 토큰의 컨텍스트 윈도우를 갖추고 있으며 다단계 지시사항, 긴 컨텍스트 및 코드 생성에 최적화되어 있습니다. 마이크로소프트는 MAI-Thinking-1이 주요 소프트웨어 엔지니어링 벤치마크에서 최고 수준의 모델들과 맞먹는 성능을 보이며, 사내 블라인드 테스트에서 Anthropic의 Sonnet 4.6보다 선호도가 높았다고 밝혔습니다. 술레이만은 이 모델이 타사 모델의 증류(distillation) 없이 깨끗한 데이터로 처음부터 학습되었다고 덧붙였습니다. 이는 타 연구소의 관행을 겨냥한 노골적인 비판이기도 합니다. 하지만 공개된 벤치마크를 살펴보면 이 모델은 Deepseek V3.2와 대체로 비슷한 수준입니다.

6가지 작업 영역을 아우르는 모델 패밀리 추론 모델 외에도 MAI 패밀리에는 6개의 추가 시스템이 포함되어 있습니다. MAI-Code-1-Flash는 50억 개의 매개변수를 가진 에이전트 코딩 모델로, Anthropic의 Haiku와 필적하지만 운영 비용이 더 저렴하다고 마이크로소프트는 설명합니다. 이 모델은 GitHub Copilot 및 Visual Studio Code에 통합됩니다.

MAI-Image-2.5는 텍스트-이미지 변환 및 이미지 편집을 담당하며, Arena-Score 이미지 벤치마크에서 GPT-Image-2에 이어 2위를 차지해 구글의 Nano-Banana 모델들을 제쳤습니다. MAI-Transcribe-1.5는 43개 언어를 지원하는 가장 빠른 전사(transcription) 모델로 홍보되고 있습니다. MAI-Voice-2는 15개 언어로 음성을 생성하고 짧은 샘플로 음성을 복제할 수 있습니다. 마이크로소프트에 따르면 모든 모델은 동일한 데이터 기반, 인프라 및 평가 파이프라인을 공유합니다. 이 모델들은 Azure Foundry를 통해 사용할 수 있으며, 개발자들은 처음으로 가중치(weights)를 직접 파인튜닝할 수 있습니다.

비용 효율성을 입증하는 프론티어 튜닝 마이크로소프트는 이러한 모델들과 함께 '프론티어 튜닝'이라는 새로운 접근 방식을 결합했습니다. 고객은 강화학습 환경을 사용하여 모델을 자신의 워크플로우에 직접 맞출 수 있습니다. 마이크로소프트는 가장 가치 있는 학습 데이터가 조직 내에서 에이전트가 남기는 실제 작업 traces라고 강조합니다. 사내 테스트에서 Excel에 맞게 튜닝된 MAI 모델은 GPT-5.4의 성능과 일치하면서도 최대 10배 더 높은 효율성을 보였습니다. 맥킨지(McKinsey)에서 테스트된 맞춤형 MAI 모델은 테스트된 모든 시스템 중 가장 높은 승률을 달성했으며, 역시 약 1/10 수준의 비용이 들었습니다.

마이크로소프트 최초의 상시 작동 에이전트, 스카우트 세 번째 축은 마이크로소프트가 '오토파일럿(Autopilots)'이라고 부르는 새로운 에이전트 카테고리입니다. 이는 자체적인 아이덴티티를 가지고 백그라운드에서 자율적으로 작동하는 상시 에이전트입니다. 첫 번째 제품은 Teams, Outlook, OneDrive 및 SharePoint에 통합된 '마이크로소프트 스카우트(Microsoft Scout)'입니다.

스카우트는 시간대가 다른 회의 일정을 조정하고, 브리핑 자료를 준비하며, 캘린더에 향후 업무 일정을 예약하고, 문제가 발생하기 전에 지연된 결정을 사전에 알리도록 설계되었습니다. 'Work IQ'라는 구성 요소를 통해 이 에이전트는 사용자가 일하는 방식과 우선순위에 대한 컨텍스트 메모리를 구축합니다. 각 에이전트는 자체 Entra 아이덴티티로 실행되며 엄격하게 범위가 지정된 액세스 권한을 갖고 샌드박스 환경을 통해 실행됩니다.

원문 보기
원문 보기 (영어)
Build 2026: Microsoft tops Google in image generation while playing catch-up on reasoning Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jun 3, 2026 Microsoft Key Points Microsoft unveiled seven homegrown AI models at Build 2026, including its first reasoning model, MAI-Thinking-1. In benchmarks, it lands roughly on par with Deepseek V3.2. A new method called "Frontier Tuning" lets companies adapt models to their own workflows using reinforcement learning. Microsoft says tuned models match GPT-5.4 performance at one-tenth the cost. Microsoft is also launching "Scout," an always-on background agent that handles office tasks like scheduling and meeting prep. The software announcements are paired with local developer hardware and a new operating system built for AI agents. Ask about this article… Search At Build 2026, Microsoft announced seven new AI models developed in-house, including its first reasoning model. The company also introduced a new tuning method and an autonomous background agent. The centerpiece is MAI-Thinking-1, Microsoft's first reasoning model. According to Microsoft AI chief Mustafa Suleyman, it's a 1-trillion-parameter model with 35 billion active parameters and a 128,000-token context window, built for multi-step instructions, long contexts, and code generation. Microsoft says MAI-Thinking-1 matches leading models on key software engineering benchmarks and was preferred over Anthropic's Sonnet 4.6 in internal blind comparisons. The model was trained from scratch on clean data without distillation from third-party models, according to Suleyman. That's a not-so-subtle jab at practices at other labs. A look at the published benchmarks, though, puts the model roughly on par with Deepseek V3.2. Ad A model family spanning six task areas Beyond the reasoning model, the MAI family includes six more systems. MAI-Code-1-Flash is an agentic coding model with 5 billion parameters that Microsoft says is comparable to Anthropic's Haiku but cheaper to run. It's integrated into GitHub Copilot and Visual Studio Code. Ad DEC_D_Incontent-1 MAI-Image-2.5 handles text-to-image and image editing, landing second place on the Arena-Score image benchmark behind GPT-Image-2 and ahead of Google's Nano-Banana models. MAI-Transcribe-1.5 is pitched as the fastest transcription model, supporting 43 languages. MAI-Voice-2 generates speech in 15 languages and can clone voices from short samples. All models share the same data foundation, infrastructure, and evaluation pipeline, according to Microsoft. They're available through Azure Foundry, and for the first time, developers can fine-tune the weights themselves. Ad Frontier Tuning makes the cost case Microsoft is pairing the models with a new approach called Frontier Tuning . Customers can use reinforcement learning environments to align models directly with their own workflows. The most valuable training data, Microsoft argues, is the actual work traces an agent leaves behind inside an organization. In an internal test, a MAI model tuned for Excel matched GPT-5.4's performance while running up to ten times more efficiently. At McKinsey, a customized MAI model achieved the highest win rate of any system tested, again at roughly one-tenth the cost. Ad DEC_D_Incontent-2 Scout is Microsoft's first always-on agent The third pillar is a new agent category Microsoft calls "Autopilots." These are persistent agents with their own identity that work autonomously in the background. The first one is Microsoft Scout , integrated into Teams, Outlook, OneDrive, and SharePoint. Ad Scout is designed to coordinate meetings across time zones, prepare briefing materials, schedule upcoming deliverables on your calendar, and flag stalled decisions before they become blockers. Through a component called Work IQ, the agent builds a context memory of how you work and what you prioritize. Each agent runs under its own Entra identity with tightly scoped access rights, sandboxed execution via Microsoft Execution Containers, and mandatory human approval for sensitive actions. Credentials are also scoped to each task and scrubbed from logs. Whether that's enough remains to be seen. Previous agent systems have consistently failed at exactly the point where language models meet external data. Scout is available first as an experimental release through the Frontier program. It requires an Intune configuration and a GitHub Copilot license. Hardware, an OS, and a clinical model round out the strategy The software announcements come alongside several more pieces of a broader AI strategy. With Project Solara, Microsoft is previewing an Android-based operating system designed to run agents across devices, co-developed with Qualcomm and MediaTek. On the Build stage, the company showed a desktop hub and a digital badge as possible form factors. For local AI development, Microsoft is launching the Surface RTX Spark Dev Box, equipped with Nvidia's Arm-based Spark RTX chip and 128 GB of unified memory. Pricing and full specs haven't been announced yet. In healthcare, Microsoft announced a partnership with the Mayo Clinic to co-develop a clinical foundation model. The model will first be deployed in Mayo Clinic's own operations and later made available through Azure Foundry. The Mayo Clinic retains ownership. Microsoft frames the overarching goal as "Humanist Superintelligence," meaning AI systems that remain tools under human control. Suleyman says the company plans to rapidly scale compute and capabilities over the coming year, backed in part by Microsoft's own Maia 200 chips. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: MAI Familie | Arena / Bild-Benchmark | Microsoft Devblogs / Frontier Tuning | Microsoft / Scout