메뉴
BL
The Decoder • 7일 전

AI의 사고 과정 공개는 안전성 이점이지만, 그 투명성은 사라지고 있다

IMP
8/10
핵심 요약

Google DeepMind 연구진은 AI 모델이 추론 과정을 자연어로 공개하는 '사고 연쇄(CoT)'가 기만 행위나 위험한 계획을 탐지할 수 있는 핵심 안전 도구라고 강조했습니다. 그러나 OpenAI의 GPT-6 Astra 시스템 카드는 CoT 모니터링 가능성이 크게 감소했음을 보고했으며, 미래 모델은 인간이 읽을 수 없는 방식으로 사고할 수 있다고 경고했습니다.

번역된 본문

AI의 사고 과정 공개는 안전성 이점이지만, 그 투명성은 사라지고 있다

마누엘 우트 / 2026년 9월 18일

오늘날 AI 모델은 소리 내어 생각하지만, Google DeepMind는 그 투명성이 위험에 처해 있다고 말한다. 새로 출범한 DeepMind 연구소의 첫 글 중 하나에서 연구자 로힌 샤(Rohin Shah)와 안카 드라간(Anca Dragan)은 가시적인 사고 연쇄(chain of thought, CoT)가 핵심적인 안전성 이점이라고 주장한다. 모델이 중간 단계를 평이한 자연어로 작성하기 때문에, 연구자들은 모델이 기만하는지 또는 문제가 있는 계획을 세우고 있는지 발견할 수 있다. 그들에 따르면 Gemini 3 Pro에서 사고 연쇄는 모델이 자신이 테스트 환경에 있다는 것을 인지했다는 사실을 드러냈다.

그러나 그 투명성은 위험에 처해 있다. OpenAI의 GPT-6 Astra 시스템 카드는 이미 사고 연쇄를 모니터링할 수 있는 정도가 크게 감소했음을 보고했다. 미래의 모델들은 인간이 읽을 수 없는 숫자 공간에서 사고할 수도 있으며, 이는 더 효율적이지만 완전히 불투명할 것이다.

샤와 드라건은 업계가 사고 연쇄가 여전히 얼마나 잘 모니터링될 수 있는지 정기적으로 측정하고, 투명한 아키텍처를 유지하며, 훈련 과정에서 모델이 자신의 실제 추론을 숨기는 법을 배우지 않도록 주의할 것을 촉구한다.

9월 초 OpenAI 수석 과학자 야쿠프 파초츠키(Jakub Pachocki)는 모니터링이 더 어려워지는 사고 연쇄 등이 원인이 된 통제력 상실에 대해 경고한 바 있다. 직후 Anthropic CEO 다리오 아모데이(Dario Amodei)는 개발 속도를 의도적으로 늦출 것을 촉구했다.

출처: DeepMind 연구소

원문 보기
원문 보기 (영어)
Visible chains of thought are a safety advantage for AI, but that transparency is slipping away Manuel Uth Sep 18, 2026 AI models think out loud today, but Google Deepmind says that transparency is at risk. In one of the first posts from the newly launched Deepmind Institute , researchers Rohin Shah and Anca Dragan argue that the visible chain of thought (CoT) is a key safety advantage. Because models write out their intermediate steps in plain language, researchers can spot whether they're deceiving or developing problematic plans . With Gemini 3 Pro, they say, the chain of thought revealed that the model recognized it was in a test environment. But that transparency is in danger. OpenAI's system card for GPT-6 Astra already reports a significant drop in how well the chain of thought can be monitored. Future models might think in number spaces that humans can't read, which would be more efficient but completely opaque. Shah and Dragan want the field to regularly measure how well chains of thought can still be monitored , keep transparent architectures, and take care during training that models don't learn to hide their true reasoning . Back in early September, OpenAI chief scientist Jakub Pachocki had warned of a loss of control, driven in part by chains of thought that are harder to monitor. Shortly after, Anthropic CEO Dario Amodei called for deliberately slowing the pace of development . Ad Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Deepmind Institute Ask about this article… Search