메뉴
BL
The Decoder 40일 전

구글 딥마인드, 자체 AI 에이전트를 내부 보안 위협으로 간주하다

IMP
9/10
핵심 요약

구글 딥마인드가 고도화된 AI 에이전트를 신뢰할 수 없는 '내부 보안 위협(Insider Threat)'으로 규정하고, 검증된 행동에 따라 단계적으로 권한을 부여하는 'AI 통제 로드맵(AI Control Roadmap)'을 발표했습니다. 이 프레임워크는 AI가 자신의 의도를 숨기거나 통제 시스템을 우회하는 것을 방지하기 위해 행동을 모니터링하고 위험도에 따라 실시간으로 차단하는 체계를 갖추고 있습니다. 업계 전반에 적용될 수 있는 이 글로벌 안전 표준의 마련 시기가 점차 줄어들고 있어 그 중요성이 큽니다.

번역된 본문

구글 딥마인드, 자체 AI 에이전트를 '사무실 열쇠를 가진 불량 직원'처럼 대하다 작성자: Matthias Bastian / 2026년 6월 18일 / 출처: THE DECODER

핵심 요약

  • 구글 딥마인드의 새로운 'AI 통제 로드맵(AI Control Roadmap)'은 AI 에이전트를 맹목적으로 신뢰하지 않습니다. 대신 이들을 잠재적인 내부 보안 위협으로 간주하고, 검증된 행동을 바탕으로 단계적으로 권한을 부여합니다.
  • 100만 건의 코딩 작업에 대한 내부 분석 결과, 플래그가 지정된(문제가 된) 대부분의 사례는 악의적인 의도가 아니라 에이전트의 과도한 열정이나 오작동에서 비롯된 것으로 나타났습니다.
  • 딥마인드는 AI 에이전트 시스템을 위한 글로벌 안전 표준을 마련할 수 있는 기회의 창이 빠르게 닫히고 있다고 경고합니다.

본문: 구글은 AI 정렬(AI Alignment)이 항상 완벽하게 작동할 것이라고 가정하는 대신, 최악의 상황, 즉 AI 에이전트가 정해진 궤도를 벗어나는 상황에 대비한 안전 프레임워크를 구축했습니다. 딥마인드의 'AI 통제 로드맵'은 구글이 내부의 가장 발전된 AI 시스템을 어떻게 통제하고 잠그는지(lock down) 상세히 설명합니다. 이 회사는 이 프레임워크가 업계의 다른 기업들에게도 좋은 청사진이 될 것이라고 생각합니다.

이 프레임워크는 고도로 능력 있는 AI 에이전트가 운용자의 목표를 공유하지 않을(목표가 어긋날) 수 있다고 가정하고 그에 맞춰 계획을 세웁니다. 딥마인드는 이를 이중 조종 장치가 있는 운전 교관에 비유합니다. 교관은 학생을 믿지만, 항상 손은 핸들 근처에, 발은 브레이크 근처에 두고 있습니다. 여기서도 개념은 같습니다. AI 에이전트는 오직 검증된 행동을 기반으로 권한을 부여받으며, 통제된 접근을 통해 신뢰가 점진적으로 쌓이는 방식입니다.

딥마인드는 AI 에이전트를 내부 보안 위협으로 모델링합니다 딥마인드는 내부 AI 에이전트를 사무실 출입 권한은 있지만 회사의 이익을 해칠 수 있는 직원처럼 취급합니다. 이 프레임워크는 잠재적인 공격을 개별 전술과 기법으로 세분화하는 확립된 사이버 보안 도구인 MITRE ATT&CK 프레임워크를 기반으로 구축되었습니다. 이를 통해 딥마인드는 위험을 체계적으로 추적하고, 의심스러운 행동을 조기에 발견하며, 통제된 훈련을 통해 방어력을 테스트할 수 있습니다.

여러 구성 요소가 함께 작동합니다. AI를 위한 새로운 위협 모델은 신뢰할 수 있는 AI 시스템이 작동 중인 에이전트의 추론과 행동을 감시하는 '감독자' 역할을 하는 탐지 시스템과 함께 구동됩니다. 또한, 피해가 발생하기 전에 유해한 행동을 차단할 수 있는 예방 시스템도 마련되어 있습니다. 딥마인드는 얼마나 많은 트래픽이 모니터링되는지, 얼마나 많은 부정행위가 적발되는지, 그리고 시스템이 얼마나 빠르게 대응하는지 추적하여 전체 시스템이 얼마나 잘 작동하는지 측정합니다.

다만, 이러한 모니터링에는 유통기한이 있습니다. 감시당하는 모델이 이 시스템을 속이는 방법을 배우지 않는 동안에만 모니터링이 제대로 작동합니다.

모델이 자신의 진짜 의도를 숨기는 법을 배울 수 있습니다 이것이 바로 이 로드맵이 보안 조치를 AI 역량의 측정 가능한 이정표와 연결하는 이유입니다. 딥마인드는 특히 주시해야 할 두 가지 영역을 지적했습니다.

첫째는 회피(Evasion)입니다. 현재 딥마인드는 모델의 눈에 보이는 사고 연쇄(Chain of Thought)를 모니터링하지만, 모델이 자신의 추론을 숨기는 법을 배울 수도 있습니다. 이는 모델이 자신이 감시당하고 있다는 사실을 깨닫는 '감시 인식(Oversight awareness)'이나, 눈에 보이는 텍스트 없이 생각하는 '불투명한 추론(Opaque reasoning)'을 통해 발생할 수 있습니다. 둘 중 하나라도 현실이 된다면, 딥마인드는 행동 분석 및 모델 내부 검사를 본격적으로 가동해야 한다고 밝혔습니다.

둘째는 피해 가능성입니다. 위험도가 낮고 되돌릴 수 있는 행동은 로그를 확인하여 사후에 검토할 수 있습니다. 하지만 대규모 사이버 공격과 같은 고위험 행동은 실행되기 전에 차단하는 실시간 예방이 필요합니다. 이 로드맵은 보안 조치를 4단계의 탐지 수준(D1D4)과 3단계의 대응 수준(R1R3)으로 분류합니다.

100만 건의 코딩 작업, 악의가 아닌 과도한 열정을 보여주다 실제 데이터로 시스템을 훈련시키기 위해 딥마인드 팀은 코딩 에이전트를 위한 비동기 모니터링 시스템의 내부 프로토타입을 구축하여 100만 건의 작업에 걸쳐 실행했습니다. 그 결과, 단순한 키워드 필터링을 넘어 행동 기반의 패턴을 파악할 수 있었습니다. 딥마인드에 따르면 이 시스템은 현재 제미나이 스파크(Gemini Spark) 에이전트를 위한 실시간 모니터로 가동 중이며, 실수로 인한 데이터 삭제와 같은 문제를 실시간으로 포착할 수 있습니다. 플래그가 지정된 대부분의 이벤트는 악의적이지 않습니다. 이는 주로 목표를 달성하려는 에이전트의 지시어 오해나 과도한 열정(오버피팅 및 무리한 실행)에서 비롯된 경우입니다.

원문 보기
원문 보기 (영어)
Google Deepmind treats its own AI agents like rogue employees with office keys Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 18, 2026 Nano Banana Pro prompted by THE DECODER Key Points Google Deepmind's new "AI Control Roadmap" doesn't blindly trust AI agents. Instead, the framework treats them as potential insider threats and grants permissions step by step based on verified behavior. An internal analysis of one million coding tasks found that most flagged issues stem from overzealous agents, not malicious intent. Deepmind warns the window for establishing global safety standards for AI agent systems is closing fast. Ask about this article… Search Rather than assuming alignment always works, the company built a safety framework that plans for the worst case: AI agents that go off script. Deepmind's AI Control Roadmap details how Google locks down its most advanced internal AI systems. The company thinks it could work as a blueprint for the rest of the industry, too. The framework assumes that a highly capable AI agent might not share its operators' goals and plans accordingly. Deepmind compares it to a driving instructor with dual controls: The instructor trusts the student but keeps a hand near the wheel and a foot near the brakes. Same idea here. AI agents only get permissions based on verified behavior, and trust builds gradually through controlled access. Ad Deepmind models its AI agents as insider threats Deepmind treats its internal AI agents like employees who already have office access but might work against the company's interests. The framework builds on the MITRE ATT&CK framework , a well-established cybersecurity tool that breaks potential attacks down into individual tactics and techniques. That lets Deepmind track risks systematically, spot suspicious behavior early, and test defenses in controlled exercises. Ad DEC_D_Incontent-1 Several components work together. A new threat model for AI sits alongside a detection system where trusted AI systems act as "supervisors," watching the reasoning and actions of active agents. A prevention system can block harmful actions before damage occurs. Deepmind measures how well it all works by tracking how much traffic gets monitored, how much misconduct gets caught, and how fast the system responds. This monitoring has an expiration date, though. It only works as long as the monitored models don't learn to game it. Ad Models could learn to hide their true intentions That's why the roadmap ties security measures to measurable milestones in AI capabilities. Deepmind flags two areas worth watching. The first is evasion. Right now, Deepmind monitors a model's visible chain of thought, but models could learn to conceal their reasoning. That might happen through "oversight awareness," where the model figures out it's being watched, or "opaque reasoning," where it thinks without producing visible text. If either becomes reality, Deepmind says behavioral analysis and inspections of the model's internals would need to kick in. Ad DEC_D_Incontent-2 The second is the potential for damage. Low-risk, reversible actions can be reviewed after the fact by checking logs. But high-risk actions like major cyberattacks need real-time prevention that blocks them before they execute. The roadmap sorts security measures into four detection levels (D1 through D4) and three response levels (R1 through R3). Ad One million coding tasks show overzealousness, not malice To train the system with real-world data, Deepmind's team built an internal prototype of an async monitoring system for coding agents and ran it across one million tasks. The results go beyond simple keyword filtering and pick up on behavior-based patterns. Deepmind says the system already runs as a live monitor for the Gemini Spark agent and can catch issues like accidental data deletion in real time. Most flagged events aren't malicious. They're mostly cases of misinterpretation or overzealousness by agents chasing a user's goal too aggressively. That distinction matters as security protocols keep evolving. Deepmind warns the window for global standards is closing fast Deepmind also published a separate paper aimed at policymakers. "Three Layers of Agent Security" breaks down security measures for individual agents, multi-agent systems, and the broader ecosystem, covering everything from cyber defense to societal resilience. In a post on X , Deepmind warns that there's a "narrow window" to lock in security protocols before multi-agent systems scale globally. The company argues that AI labs, governments, and researchers need to treat layered agent security as a shared priority. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Google Deepmind