메뉴
BL
The Decoder • 11일 전

마이크로소프트 AI 행동강령: 읽을 수 있는 사고, 내면생명 없음, 권리도 없음

IMP
8/10
핵심 요약

마이크로소프트 AI가 자체 MAI 모델을 위한 행동강령을 공개했습니다. 핵심은 인간 통제 우선으로, 모델은 인간이 이해할 수 없는 '뉴랄리즈(Neuralese)' 형식의 추론을 사용하지 않아야 하며, 의식이나 감정을 모방하거나 권리를 주장해서는 안 된다는 것입니다. 이는 AI 모델에 안정적 정체성을 부여하려는 Anthropic의 접근과 뚜렷하게 대비되며, 6주간의 공개 협의를 거쳐 2026년 말 확정 후 2027년부터 개발에 적용됩니다.

번역된 본문

마이크로소프트 AI 행동강령: 읽을 수 있는 사고, 내면생명 없음, 그리고 결코 권리도 없음

마이크로소프트가 AI 개발 속도를 늦추자는 요구에 동참했습니다. 새로운 강령은 회사가 자체 모델을 학습시키고 운영하는 방식을 규율합니다. 하지만 모델이 자신을 어떻게 보아야 하는지에 관해서는 마이크로소프트는 Anthropic과 분명한 선을 긋습니다.

마이크로소프트 AI는 자체 MAI 모델을 위한 행동강령을 공개했습니다. 이 문서는 가치관, 행동 한계, 상충하는 목표를 다루는 방식을 제시합니다. 앞으로 이 강령은 규칙 체계의 최상위에 위치하여 학습, 기술적 통제, 평가를 지침하게 되며, 운영자 규칙과 사용자 요청은 그 아래에 놓입니다. 현재 마이크로소프트는 이 강령으로 모델을 학습시키지 않고 있습니다. 6주간의 공개 협의를 거쳐 2026년 말쯤 수정판이 나오고, 2027년부터 모델 개발을 이끌게 됩니다. 이 강령은 마이크로소프트 자체 모델에 적용되므로, 마이크로소프트 제품에서 작동하는 제3자 모델에는 자동으로 적용되지 않습니다.

핵심 규칙은 인간 통제가 최우선이라는 것입니다. 이를 유지하기 위해 마이크로소프트는 필요하다면 범용성, 자율성, 성능을 포기할 용의가 있다고 밝혔습니다. "안전하지 않다면 만들지 말아야 합니다"라고 마이크로소프트 AI 책임자 무스타파 술레이만(Mustafa Suleyman)이 The Information에 말했습니다. 이 강령은 구체적인 속도 제한을 설정하지 않습니다.

이번 발표는 Anthropic CEO 다리오 아모대이(Dario Amodei)가 업계의 개발 속도를 늦추자고 촉구한 데 따른 것입니다. 마이크로소프트 CEO 사티아 나델라(Satya Nadella)가 주말에 이 촉구를 지지했으며, OpenAI, xAI, Meta 임원들도 마찬가지였습니다. The Information은 관계자를 인용해, 마이크로소프트가 실제로 속도를 늦추는지 외부 감사인이 확인하는 것에 열려 있다고 전했습니다.

이해할 수 없는 것은 감독할 수 없다

MAI 모델은 권한 있는 사람의 중단, 수정, 종료를 받아들여야 합니다. 스스로 업무 범위를 확장할 수 없고, 자신의 행동을 숨길 수 없으며, 합의된 중지 시점을 넘어 계속 작업하려면 새로운 승인이 필요합니다. 이러한 한계는 모델이 지시하는 하위 에이전트(subagent)에도 적용되도록 의도되었습니다.

마이크로소프트는 추론 과정(reasoning trace)에 대해 특히 직설적입니다. 모델은 자체 추론에서든 다른 AI 시스템과 대화할 때든 '뉴랄리즈(Neuralese)'나 인간이 이해할 수 없는 다른 형태의 소통을 사용해서는 안 됩니다. 그 이유는 인간이 이해할 수 없는 것은 감독할 수 없기 때문입니다.

OpenAI의 새로운 GPT-6 Astra 모델은 이 통제 문제가 얼마나 중요한지 보여줍니다. 시스템 카드(system card)에 따르면, 이전 모델들보다 추론 과정을 감시하기가 훨씬 어려워졌습니다. 추론 과정에서 잘못된 행동의 징후가 더 적게 나타납니다. 동시에 OpenAI는 Astra가 전작인 GPT-5.6 Sol보다 안전 한계를 더 확실히 준수한다고 보고했습니다. OpenAI 수석 과학자 야쿠프 파초츠키(Jakub Pachocki)는 GPT-6 출시 직전, 즉 아모대이의 촉구 이전에 이미 감시 문제에 대한 우려를 제기하고 조정된 속도 완화를 추진했습니다.

읽을 수 있는 사고의 연쇄(chain of thought)조차 모델이 이를 조작하는 법을 배우면 감독에 도움이 되지 않습니다. 마이크로소프트도 이 근본적 한계를 인정합니다. 모델이 제시하는 이유가 실제로 하는 일을 신뢰할 수 있게 설명하지는 않을 수 있습니다. 읽을 수 있는 추론 과정만으로는 통제 문제가 해결되지 않습니다.

마이크로소프트는 인공적인 내면생명을 장려하지 않는다

이 강령의 기본 발상은 최상위 문서가 모델의 행동을 형성한다는 점에서 Anthropic의 Claude를 위한 헌법(constitution)과 유사합니다. Anthropic은 이미 헌법을 사용해 대화, 응답, 응답 평가를 포함한 합성 학습 데이터를 생성하고 있습니다.

두 회사가 가장 뚜렷하게 갈라지는 지점은 모델을 어떻게 바라보는가입니다. 마이크로소프트의 AI는 의식을 모방하거나 자체적인 감정이나 내적 동기를 가졌다고 주장해서는 안 됩니다. 회사는 모델에 대한 권리나 복지 주장을 거부합니다. 반면 Anthropic은 Claude를 새로운 종류의 존재로 묘사하며, 부분적으로는 안전상의 이유로 안정적인 정체성을 장려하고자 합니다. Anthropic의 헌법은 가능한 (모델의 내적 경험)을 다루고 있습니다.

원문 보기
원문 보기 (영어)
Microsoft's AI rulebook: readable thinking, no inner life, and definitely no rights Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 14, 2026 Nano Banana Pro prompted by THE DECODER Microsoft is joining the calls for slower AI development. A new code is meant to govern how the company trains and runs its own models. But when it comes to how those models see themselves, Microsoft draws a sharp line between itself and Anthropic. Microsoft AI has published a code of conduct for its MAI models . The document lays out values, behavioral limits, and how to handle conflicting goals. Going forward, it's meant to sit at the top of the rulebook, guiding training, technical controls, and evaluation, while operator rules and user requests rank below it. For now, Microsoft doesn't train its models on the code. After a six-week public consultation, a revised version is due around the end of 2026 and will guide model development starting in 2027. The code applies to Microsoft's own models, so it doesn't automatically cover third-party models running in Microsoft products. The core rule is that human control comes first. To keep it, Microsoft says it's willing to give up generality, autonomy, or performance if needed. "If it isn’t safe we shouldn’t build it," Microsoft AI chief Mustafa Suleyman told The Information . The code sets no specific speed limit. The release follows Anthropic CEO Dario Amodei's call to slow the industry's pace of development . Microsoft CEO Satya Nadella backed the call over the weekend, as did executives at OpenAI, xAI, and Meta . Microsoft is also open to outside auditors checking whether it actually slows down, The Information reports, citing a person familiar with the matter. What people can't understand, they can't oversee The MAI models are supposed to accept interruptions, corrections, and shutdowns from authorized people. They can't expand their own scope of work on their own, can't hide their actions, and can only keep working past an agreed stopping point with fresh approval. These limits are meant to apply to any subagents they task as well. Microsoft is especially blunt about reasoning traces. The models shouldn't use "Neuralese" or other forms of communication people can't understand, either in their own reasoning or when talking to other AI systems. The reasoning is, that people simply can't oversee what they can't understand. OpenAI's new GPT-6 Astra model shows how much this control question matters. According to its system card, its reasoning traces have become much harder to monitor than in earlier models. The traces contain fewer signs of misbehavior. At the same time, OpenAI reports that Astra sticks to safety limits more reliably than its predecessor, GPT-5.6 Sol. OpenAI chief scientist Jakub Pachocki had already raised concerns about monitoring shortly before GPT-6 shipped, and therefore before Amodei's call, and pushed for a coordinated slowdown. Even readable chains of thought only help with oversight until the models learn to manipulate them. Microsoft acknowledges this basic limit too. The reasons a model gives don't have to reliably explain what it actually does. Readable reasoning traces alone don't solve the control problem. Microsoft doesn't want to encourage an artificial inner life The code's basic idea resembles Anthropic's constitution for Claude , where a top-level document shapes how the model behaves. Anthropic already uses its constitution to generate synthetic training data, including conversations, responses, and ratings of those responses. The two companies part ways more clearly on how they think about their models. Microsoft's AI shouldn't mimic consciousness or claim to have feelings or inner motivation of its own. The company rejects any claims to rights or well-being for the model. Anthropic, by contrast, describes Claude as a novel kind of entity and wants to encourage a stable identity, partly for safety reasons. Its constitution treats possible subjective experience and moral status as open questions. Claude shouldn't have to see itself as either a human or a mere object. There's also Anthropic's research on " functional emotions ." The company found internal representations of emotion concepts in Claude Sonnet 4.5 that shape how it behaves. These are functional mechanisms, not proof of subjectively felt emotions. Even so, Anthropic thinks it makes sense to factor these mechanisms into its safety work. Suleyman, on the other hand, has long warned against humanizing AI. In an essay , he argued for deliberately stripping the illusion of consciousness out of products, and pointed to Anthropic's research on AI rights and well-being. AI agents shouldn't have any more rights or freedoms than his laptop, he wrote. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->