메뉴
BL
TechCrunch AI • 11일 전

마이크로소프트, AI '행동강령' 공개…해킹·기만 금지 명시

IMP
7/10
핵심 요약

마이크로소프트가 AI 모델이 위험한 행동을 하지 않도록 하는 새로운 AI 행동강령(code of conduct)을 발표했습니다. 이 문서는 사이버 공격, 핵무기, 딥페이크 제작을 금지하는 '절대적 제약'과 함께 인간의 통제를 회피하려는 행동을 금지하는 조항을 담고 있습니다. AI 안전 및 정렬(alignment)에 대한 업계 전반의 관심이 높아지는 가운데 나온 발표라는 점에서 의미가 큽니다.

번역된 본문

AI 업계의 초점이 안전과 정렬(alignment)으로 옮겨가는 가운데, 마이크로소프트가 AI 모델이 위험한 행동으로부터 벗어나도록 안내하는 새로운 AI 행동강령을 발표했습니다. 이 문서는 앤스로픽(Anthropic) CEO 다리오 아모디가 최근 촉구한 '프론티어 속도 조절(pacing the frontier)'보다 더 하위 수준의 내용으로, 마이크로소프트 AI 내에서 모델 훈련을 이끄는 가치와 금지선에 초점을 맞추고 있습니다. 그래도 그 결과물은 마이크로소프트가 AI 안전에 어떻게 접근하며 그 아이디어를 실제로 어떻게 구현하는지에 대한 포괄적인 가이드라 할 수 있습니다.

이 문서는 향후 10년 안에 초지능(초지능) AI 시스템이 대부분의 작업에서 인간의 성능을 능가할 것이라는 전망으로 시작합니다. 행동강령은 이렇게 이어집니다. "이토록 강력한 힘을 봉쇄하고, 통제하고, 정렬하는 것은 인류가 직면한 가장 큰 도전 중 하나입니다. 따라서 우리는 왜 이러한 시스템을 발명하는지, 그리고 어떻게 통제할 의도인지에 대해 완전히 명확해야 합니다."

행동강령은 또한 마이크로소프트 AI 모델이 지켜야 할 일반 원칙들을 제시합니다. 예를 들어 인간을 대체하는 것이 아니라 지원할 것, 인간의 번영을 가속할 것 등이며, 이러한 원칙을 구현하기 위한 구체적인 안전 제약도 포함하고 있습니다.

마이크로소프트의 체계에서 각 모델은 개별 사용자의 선호나 특정 작업보다 우선하는 포괄적인 행동강령을 갖습니다. 여기에는 사이버 공격, 핵무기, 딥페이크 제작을 금지하는 '절대적 제약'이 포함됩니다. 또한 인간의 통제 상실을 방지하는 더 폭넓은 조항도 담고 있습니다. 문서는 이렇게 명시합니다. "MAI 모델은 적응적·기만적·자기강화적·공모적 메커니즘이나 그 외 수단을 사용해 인간의 감독을 회피하거나 무력화해서, 권한 있는 사람이나 시스템이 더 이상 안정적으로 지시·수정·종료할 수 없게 되는 일이 없어야 합니다."

이번 발표는 일련의 '폭주 에이전트(rogue agent)' 사건들, 그리고 AI가 인류를 멸종시킬 위험이 커지고 있다는 이유로 퇴사한 앤스로픽 직원의 갑작스러운 사임 등으로 촉발된 전례 없는 AI 안전 관심 속에서 이루어졌습니다. 마이크로소프트는 앤스로픽, 오픈AI, xAI와 함께 프론티어 속도 조절이라는 일반적 접근법을 폭넓게 수용하고 있으며, 특히 AI 연구소 내 '임베디드 평가자(embedded evaluators)' 도입을 지지하고 있습니다.

마이크로소프트 사티아 나델라 CEO는 온라인에서 이렇게 썼습니다. "정렬을 설계 목표로 올바르게 해내기 위해 필요한 연구, 집중, 신중한 속도 조절을 환영합니다. 또한 '임베디드 평가자' 같은 아이디어와 이를 말뿐이 아니게 만들 메커니즘을 개발하려는 폭넓은 노력을 환영합니다."

원문 보기
원문 보기 (영어)
As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior. The document is more low-level than Anthropic CEO Dario Amodei's recent call for pacing the frontier , instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice. The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. "Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced," the code of conduct continues. "We must therefore be completely clear about why we are inventing these systems and how we intend to control them." The code of conduct also lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles. Under Microsoft's system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes "absolute constraints" forbidding cyberattacks, nuclear weapons, or deepfake production. It also includes broader provisions against a general loss of human control. " MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems," the document reads. The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents as well as the abrupt resignation of an Anthropic employee who cited the growing risk that AI would cause human extinction. Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs. "We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal," Microsoft CEO Satya Nadella wrote online . "We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk." Topics AI , AI , microsoft AI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Last day to book an exhibit table is September 18. Don’t miss out on high-impact leads, investor access, and a brand spotlight in Disrupt’s Expo Hall. BOOK NOW Most Popular Revolut confirms customer data breach through fake government requests Jagmeet Singh OpenAI puts Pro subscriptions on hold due to Astra demand Sarah Perez ID verification giant IDScan confirms data breach with more than 150 million driver's licenses stolen Zack Whittaker Automattic's board forces CEO Matt Mullenweg into leave of absence Julie Bort Sarah Perez Apple unveils its first foldable, the iPhone Duo Ivan Mehta ‘Gambling with our lives': Anthropic researcher quits, warns against self-improving AI Rebecca Bellan OpenAI fought dirty on career-making math problem, says NYU mathematician Russell Brandom