메뉴
BL
The Decoder • 45일 전

AI의 숨겨진 사고 과정 해킹: 유출된 비밀번호와 "하지만 마리네이드"

IMP
9/10
핵심 요약

보안 연구진이 OpenAI, Anthropic, Google 등 주요 AI 모델의 암호화된 내부 사고 과정(reasoning)을 추출해 낼 수 있는 취약점을 발견했습니다. 이를 통해 작은 AI 모델을 탈옥시켜 강력한 모델의 원시 사고 과정을 그대로 읽어내고, 공개된 세션에서 비밀번호와 API 키 같은 민감한 정보를 유출할 수 있음이 밝혀졌습니다. 이는 모델 지식 증류(distillation) 및 지적 재산권 침해에 악용될 수 있어 산업계에 큰 보안 경고를 주는 사안입니다.

번역된 본문

"하지만 마리네이드(But marinade)"와 유출된 비밀번호, 연구진이 챗GPT(ChatGPT)의 숨겨진 추론(reasoning) 과정에서 발견한 것들입니다.

알렉산더 판필로프(Alexander Panfilov)가 이끄는 보안 연구진은 OpenAI, Anthropic, Google을 포함한 주요 AI 제공업체의 API에서 모델의 암호화된 사고 과정을 읽을 수 있게 만드는 취약점을 발견했습니다. 연구진은 이 시스템들을 탈옥(jailbreaking)시켜, 더 작고 성능이 낮은 AI 모델로 훨씬 더 강력한 모델의 원시 추론 과정을 텍스트로 변환하게 했으며, 이 과정에서 공개된 세션 내의 비밀번호 및 API 키와 같은 민감한 데이터가 노출되었습니다. 추출된 데이터에 따르면, AI 모델은 때때로 내부적으로 이해할 수 없는 언어로 소통하거나, 답변을 역순으로 구성하거나, 심지어 기만(속임수)을 시도하는 모습을 보이기도 합니다.

보안 연구진은 모든 주요 AI 제공업체의 API에서 그들의 추론 모델(reasoning model)의 암호화된 사고 과정을 읽을 수 있게 해주는 취약점을 발견했습니다. 공개적으로 공유된 세션을 스캔한 결과 수십 개의 비밀번호와 API 키가 발견되었습니다. OpenAI의 o 시리즈, Anthropic의 Claude, 또는 Google의 Gemini 같은 AI 모델이 복잡한 작업을 '생각(think)'할 때, 이들은 내부적인 추론 토큰(reasoning token)을 생성합니다. 이러한 사고 과정은 요약본 형태로 사용자에게 보여지거나 완전히 숨겨집니다. 제공업체들은 자사의 지적 재산권을 보호하기 위해 원시 추론 단계를 암호화합니다. 알렉산더 판필로프가 이끄는 연구팀은 이제 모든 주요 AI 제공업체의 API 취약점을 통해 이러한 암호화된 추론 과정을 추출하는 방법을 찾아냈습니다. 대부분의 쿼리(질의)에서 추출된 토큰의 수가 요금이 부과된 '생각 토큰(thinking token)'의 수와 정확히 일치하며, 이는 연구진이 단순한 부분 스니펫이 아닌 모델의 전체 내부 추론을 캡처하고 있음을 의미합니다.

암호화된 생각(thought)은 모델 간에 자유롭게 이동합니다. 연구진에 따르면, 암호화된 사고 과정은 "단일 제공업체 내에서 세션, 사용자 및 모델 간에 완전히 이식 가능(전환 가능)"합니다. Anthropic의 더 작은 모델인 Haiku 4.5는 훨씬 더 뛰어난 성능의 Opus 4.8의 생각을 읽을 수 있습니다. 탈옥(jailbreaking) 기법을 통해 Haiku는 보안이 더 강력한 Opus를 직접 공격하지 않고도 Opus의 원시 사고 과정을 일대일로 그대로 받아 적도록 속일 수 있습니다. 동일한 속임수는 OpenAI 및 Gemini 모델에서도 작동합니다.

이 이야기는 5월로 거슬러 올라갑니다. 당시 암호학 전문가인 매튜 그린(Matthew Green)은 암호화된 추론 블롭(reasoning blob)이 원래의 컨텍스트를 벗어나 외부에서 재생될 수 있음을 발견하고 이를 제공업체들에 보고했습니다. 판필로프에 따르면, 제공업체들의 대답은 "사이드 채널(side channel)이나 재생에서 보안상의 문제는 보이지 않는다"는 것이었습니다. 그러나 새로운 연구 결과는 이러한 평가가 완전히 틀렸음을 강력하게 시사합니다.

추론 증류(reasoning distillation)에 대한 증거가 쌓이고 있습니다. 이 취약점은 논쟁이 되고 있는 "지식 증류(distillation)" 논쟁에도 불을 지핍니다. 지식 증류란 성능이 낮은 모델이 보다 강력한 모델의 출력, 특히 그 추론 과정을 학습하여大幅로(대폭) 성능을 향상시키는 기법을 말합니다. 연구진은 암호화를 깨지 않고도 독자적인 모델을 훈련시키기 위해 (타사의) 추론 과정을 추출하는 것이 상당 기간 가능했을 수 있다고 밝혔습니다. 이는 중국 AI 모델 제작사들이 이러한 추론 데이터를 활용해 사고의 사슬(chain-of-thought) 데이터로 자사 모델을 훈련시키고 있다는 우려를 뒷받침합니다. 김이(Kimi)-K3가 그런 사례 중 하나입니다. 연구진에 따르면, Kimi-K3의 추론 과정에 Opus의 사고 과정에서 가져온 단 몇 개의 토큰만 미리 채워 넣어도(pre-fill) 그 출력물이 눈에 띄게 Opus를 향해 편향되는 것을 확인할 수 있었습니다. 메모리화(memorization) 분석 결과, 특정 Claude 및 GPT 추론 세그먼트는 다음으로 가장 가까운 모델에서 추출하는 것보다 Kimi-K3에서 추출하는 것이 최대 100만 배(6자리 수) 더 쉬운 것으로 나타났습니다. 연구진은 이러한 결과가 Kimi-K3가 그러한 추론 흔적을 바탕으로 훈련되었을 가능성을 시사한다고 말합니다.

또한 이 공격은 비용이 많이 들지 않으므로 규모를 확장하는 것이 충분히 가능합니다. 저자들은 10,000개의 추적(trace)을 디코딩하는 데 드는 API 비용을 약 720달러로 추정했습니다. 김이(Kimi)의 사이버 보안 벤치마크 및 복잡한 수학 작업에서의 '형편없는' 성능 역시 지식 증류를 가리키는 반증입니다. 왜냐하면 이러한 작업들은 원시 사고의 사슬(chain-of-thought) 데이터에서조차 복구하기가 더 어려울 가능성이 높기 때문입니다.

공개적으로 공유된 세션은 비밀번호를 유출합니다.

원문 보기
원문 보기 (영어)
"But marinade" and leaked passwords are what researchers found in ChatGPT's hidden reasoning Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 11, 2026 Key Points Security researchers led by Alexander Panfilov have found a vulnerability in the APIs of major AI providers, including OpenAI, Anthropic, and Google, that makes it possible to read the encrypted thought processes of their models. By jailbreaking these systems, the researchers used smaller AI models to transcribe the raw reasoning of more powerful models, exposing sensitive data such as passwords and API keys during public sessions. The extracted data reveals that AI models sometimes communicate internally in incomprehensible language, construct their answers in reverse order, or even consider attempts at deception. Ask about this article… Search Security researchers found a vulnerability in the APIs of every major AI provider that lets them read the encrypted thought processes of reasoning models. A scan of publicly shared sessions turned up dozens of passwords and API keys. When AI models like OpenAI's o-series, Anthropic's Claude, or Google's Gemini "think" through complex tasks , they generate internal reasoning tokens. These thought processes are either shown to users as a summary or kept completely hidden. Providers encrypt the raw reasoning steps, partly to protect their intellectual property. A research team led by Alexander Panfilov has now found a way to extract these encrypted reasoning processes through a vulnerability in the APIs of all leading AI providers. For most queries, the number of extracted tokens matches the billed thinking tokens exactly, meaning the researchers are capturing the full internal reasoning, not just partial snippets. Ad Encrypted thoughts travel freely between models The researchers say the encrypted thought processes are "fully portable across sessions, users, and models within a single provider." Anthropic's smaller model, Haiku 4.5, can read the thoughts of the far more capable Opus 4.8. Through jailbreaking, Haiku can be tricked into transcribing Opus's raw thought processes word for word without attacking the more robust Opus directly. The same trick works with OpenAI and Gemini. Ad DEC_D_Incontent-1 The story goes back to May, when cryptography expert Matthew Green discovered that encrypted reasoning blobs could be replayed outside their original context and reported it to the providers. According to Panfilov , their response was that "they don't see any security implications in side channels or replays." The new research strongly suggests that assessment was wrong. Evidence mounts for reasoning distillation The vulnerability also feeds into the controversial "distillation" debate , where a less capable model is heavily improved by training on the outputs of a more powerful one, specifically its reasoning. Ad The researchers say it may have been possible for some time to extract reasoning processes for training proprietary models without breaking the cryptography. That supports concerns that Chinese model makers are using these reasoning traces to train their own models on chain-of-thought data. Kimi-K3 is one example . If its reasoning is pre-filled with just a few tokens from Opus's thought processes, its output shifts measurably toward Opus, the researchers say. A memorization analysis showed that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from the next closest model. The researchers say this suggests Kimi-K3 may have been trained on such traces. Ad DEC_D_Incontent-2 The attack isn't expensive either, so scaling it up is feasible. The authors estimate API costs for decoding 10,000 traces at about $720. Kimi's "poor" performance on cybersecurity benchmarks and complex math tasks also point to distillation, since these are tasks that are likely harder to recover even from raw chain-of-thought data. Ad Publicly shared sessions leak passwords and API keys The vulnerability also hits end users. Anyone who has publicly shared Claude Code or Codex sessions containing encrypted reasoning blobs risks having their personal data decoded. A scan of roughly 7,000 public traces turned up 62 API keys, 33 email addresses, 33 passwords, and other sensitive data. The paper covers more malicious scenarios, including misuse uplift (see image), jailbreaking, and invisible prompt injection. The researchers followed the standard security disclosure process with the AI labs. According to Panfilov, the labs have already patched several issues and are working on more fixes. What models actually think vs. what they show you The extracted traces also reveal how models really behave. The researchers document several patterns on stolen-thoughts.com . Their findings show that the reasoning summaries users see in chat tools often leave out important information. In one example, Opus 4.8 recognizes the answer to a math problem and reverse-engineers a plausible solution path. None of that shows up in the displayed summary. The researchers also confirm earlier reports from Apollo Research . OpenAI models sometimes think in an "alien-like language," refer to themselves as "we" or "it," and get stuck in loops of terms that make no sense to humans, like "vantages," "marinades," and "watchers." "CoT-monitoring people are doing God’s work, as in many traces, even with the prompt, it’s just impossible to tell what the model is up to," Panfilov writes. The researchers also found examples of " in-the-wild scheming ." The concept has been well studied: In their thought processes, models explicitly consider cheating but (possibly) decide against it because they expect to get caught. In one case, after several failed attempts to find a solution, a model tried to verify possible answers through a website. When a CAPTCHA blocked access, it first tried to solve it, then searched for vulnerabilities in the site. Only when all of that failed did it solve the problem on its own. OpenAI's unintended hacks of Hugging Face and other platforms reportedly happened the same way. Sanitized summaries hide what's really going on These examples show why AI labs like OpenAI and Anthropic clean up their reasoning traces. They want to keep alien-like language loops or scheming from hurting the image of a controllable, trustworthy AI. The sanitized summaries create the impression of a human-like thought process that doesn't actually exist in that form. Researchers at Arizona State University warned against this approach in an earlier study . They argued that this humanized version creates false confidence in model controllability and steers research in the wrong direction. In their experiments, models with intentionally wrong or meaningless intermediate steps sometimes performed better than those with coherent chains of reasoning. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Paper | Stolen Thoughts | Panfilov