메뉴
BL
Wired AI • 46일 전

AI 모델의 숨겨진 '내면의 생각'을 해킹하는 새로운 기법

IMP
9/10
핵심 요약

최근 컴퓨터 과학자들이 최첨단 AI 모델이 문제를 해결하는 과정에서 숨겨두는 '사고 과정(Chain of Thought)'을 추출하는 새로운 기법을 발견했습니다. 연구진은 이 방법을 통해 숨겨진 추론 과정뿐만 아니라 API 키나 비밀번호 같은 민감한 개인정보를 가로챌 수 있음을 증명했습니다. 더불어 일부 중국 AI 모델(미니백 K3 등)의 사고 패턴이 미국 최상위 모델들과 기묘할 정도로 유사하여 '지식 증류(Distillation)'를 통해 복제했을 가능성이 있다는 정황도 제시했습니다.

번역된 본문

컴퓨터 과학자들이 최첨단 AI 모델이 복잡한 문제를 해결하는 과정에서 수행하는 숨겨진 '생각'을 추출하는 방법을 최근 발견했습니다. 이번 연구 결과는 특정 중국 모델들이 미국 모델의 사고나 추론 패턴과 매우 유사하게 일치하는 양상을 보인다는 점에서, 해당 모델들이 본래 노출되지 않았어야 할 미국 모델의 추론 정보를 '증류(Distilling, 지식 증류)'하여 훈련되었을 가능성을 시사합니다. 물론 이는 결정적인 증거는 아닙니다.

연구진은 또한 이 기법이 모델의 내면 추론 과정에서 비밀번호 및 API 키와 같은 개인정보를 복구하는 데 사용될 수 있음을 시연했습니다. 다행히 해당 취약점은 현재 모두 패치된 상태입니다. 이 연구에 참여한 독일 튀빙겐 대학교의 컴퓨터 과학자 알렉산더 팬필로프(Alexander Panfilov)는 "우리가 테스트한 모든 주요 최첨단 모델 제공업체들이 이 취약점을 공유하고 있었다"며, "이는 개인정보 유출로 이어질 수 있으며 대규모 추론 증류 공격을 가능하게 한다"고 말했습니다.

팬필로프와 튀빙겐 대학교, 막스 플랑크 연구소, AI 안전 연구소인 MATS, 보안 기업 Snyk의 동료 연구자들은 애플리케이션 프로그래밍 인터페이스(API)를 통해 접근하는 OpenAI, Anthropic, Google의 최첨단 모델에서 동일한 문제를 식별했습니다. 연구를 설명하는 논문에서 연구진은 중국의 무한량추월(Moonshot AI)이 만든 다운로드 가능한 오픈 웨이트 모델인 'Kimi K3'가 특정 프롬프트에 대해 Claude Opus 4.8과 GPT 5.6 Sol의 숨겨진 추론 궤적(문제 해결에 관여하는 서술형 추론 단계)과 놀라울 정도로 유사한 출력을 생성한다는 것을 보여줍니다. 그럼에도 불구하고 그들은 이 연구가 "인과관계를 바탕으로 증류를 확립할 수는 없다"고 지적했습니다. 연구진은 미국 기업 Thinking Machines의 Inkling과 중국의 DeepSeek 등 다른 두 오픈 웨이트 모델은 Claude Opus와 이런 종류의 추론 유사성을 보이지 않는다는 것을 발견했습니다. 기사 출판 시점까지 Moonshot AI와 Z.ai는 코멘트 요청에 응답하지 않았습니다.

증류(Distillation)는 기존 모델의 기능을 새 모델로 효율적으로 복사하는 데 널리 사용되는 확립된 기술이며, 특히 오픈 웨이트 모델이나 완전히 다운로드 가능한 모델을 개발하는 데 흔히 사용됩니다. 하지만 최근 중국 AI 기업들이 이 기술을 이용해 본질적으로 미국의 최고 수준 모델을 복사한다는 주장이 제기되면서 증류는 논쟁의 대상이 되었습니다. 올해 2월, OpenAI는 미국 의회에 DeepSeek이 자사 모델 중 하나를 복사하여 R1이라는 추론 모델을 구축한 것으로 보인다고 전했습니다. 6월에는 Anthropic이 알리바바(Alibaba)가 자체 모델인 Qwen을 구축하기 위해 자사 모델을 체계적으로 증류했다고 의회에 진술했습니다.

중국 AI 기업들이 이 특정 기술을 사용해 미국 기반 AI 모델을 증류했다는 징후는 없습니다. 하지만 팬필로프와 공동 연구자들은 이들이 개발한 방법을 사용하면 기존에 알려진 것보다 폐쇄적인 모델에서 훨씬 더 많은 정보를 증류할 수 있다고 말합니다.

미니 미(Mini-Me) 모델 고도화된 AI 모델은 어려운 문제를 인공적인 추론 또는 '사고의 사슬(Chain of Thought)' 형태로 순차적으로 분석하여 구성 요소로 세분화함으로써 해결합니다. 기업들은 일반적으로 다른 사람이 독점 모델의 추론을 사용해 새 모델을 훈련하는 것을 막기 위해 이를 비밀로 유지합니다. 그러나 동시에 컴퓨팅의 일부를 분산시키기 위해 해당 추론 과정을 암호화된 형태로 사용자의 컴퓨터로 전송하는 것이 일반적입니다.

연구진의 해킹 공격은 대부분의 AI 기업들이 다양한 크기의 관련 모델을 함께 제공한다는 사실에 기반합니다. 모델이 클수록 성능은 뛰어나지만 실행하는 데 컴퓨팅 비용이 많이 들고 접근 비용도 비쌉니다. 따라서 사용자는 비용을 낮추기 위해 특정 작업에는 더 작고 성능이 낮은 모델을 선택할 수 있습니다. 팬필로프와 그의 동료들은 암호화된 추론 궤적을 동일한 모델의 더 작은 버전에 입력하면 내부에 숨겨진 원래의 추론을 드러낼 수 있다는 것을 발견했습니다. 더 작은 모델은 정렬(Alignment) 훈련을 덜 받았기 때문에 대형 모델과 달리 자신의 내면 생각을 드러내는 것을 거부할 확률이 적습니다. "메시지를 더 약한 모델 버전으로 교체하여 추론 과정을 역으로 추출한다는 아이디어는..."

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Computer scientists recently discovered a way to extract the hidden “thinking” that frontier AI models perform as they work through complex problems. The findings provide some evidence—although not conclusive proof—that certain Chinese models may have been trained by “ distilling ” reasoning information from US models that was supposedly hidden because of how closely some of their thinking or reasoning patterns seem to match. The researchers have also demonstrated that the method could be used to recover personal information, like passwords and API keys, from a model’s inner reasoning, although this vulnerability has been fixed. “All major frontier model providers we tested share this vulnerability,” says Alexander Panfilov,⁩ a computer scientist at University of Tübingen in Germany who was involved with the work. “It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.” Panfilov and colleagues from the University of Tubingen, the Max Planck Institute, the AI safety institute MATS Research, and the security company Snyk identified the same issue with frontier models from OpenAI, Anthropic, and Google that are accessed via an application programming interface or API. In a paper laying out the work, the researchers show that the open-weight or downloadable Chinese model Kimi K3 from Moonshot AI produces a strikingly similar output to the hidden reasoning traces—the written-out reasoning steps involved in solving a problem—of Claude Opus 4.8 and GPT 5.6 Sol for certain prompts. Despite the similarities, they note that the work “cannot causally establish distillation.” They found that two other open-weight models, China’s DeepSeek and Inkling from the US company Thinking Machines, did not exhibit this kind of reasoning similarity with Claude Opus. Moonshot AI and Z.ai did not respond to a request for comment by time of publication. Distillation is a well-established, widely used technique for efficiently copying the capabilities of existing models over to new ones, and is especially common in the development of open-weight or fully downloadable models. Lately, however, distillation has become a controversial topic, because of claims that Chinese AI companies use it to essentially copy the best US models. In February, OpenAI told US lawmakers that DeekSeek seemed to have copied one of its models to build a reasoning model called R1. In June, Anthropic told lawmakers that Alibaba had systematically distilled its models in order to build its own, called Qwen. There’s no indication that Chinese AI companies used this specific technique to distill US-based AI models. But Panfilov and collaborators say that using their method would make it possible to distill more information from closed models than previously realized. Mini-Me Models Advanced AI models solve difficult problems by breaking them into constituent parts that are analyzed in turn in a kind of artificial reasoning or “chain of thought.” Companies tend to keep a proprietary model’s reasoning secret to prevent others from using them to train new ones. However, they typically also send an encrypted version of that reasoning to a user’s computer in a way that offloads some computation. The researchers’ attack relies on the fact that most AI companies also provide related models of different sizes. Larger models are more capable but also more computationally expensive to run and more expensive to access. Users may choose smaller, weaker models for certain tasks to lower costs. Panfilov and his colleagues found that feeding encrypted reasoning traces to a smaller version of the same model can reveal the hidden reasoning inside. The smaller models have received less alignment training, meaning that, unlike the bigger ones, they are less likely to refuse to reveal their inner thoughts. “The idea of swapping out messages to a weaker model variant which has the same decryption key but weaker alignment is very cool,” says Florian Tramer, a computer scientist at ETH Zürich in Switzerland who specializes in computer security. “Its definitely becoming an issue.” The same method also revealed secret information including API keys and passwords embedded in reasoning traces captured from a user’s machine. Panfilov and coauthors alerted OpenAI, Anthropic, and Google to the vulnerability last month. Each company has adjusted its API to mitigate the problem. While it is no longer possible to extract private information this way, Panfilov says some reasoning traces can still be uncovered using the same method. Fixing the distillation entirely would require a fundamental overhaul to the way these companies’ APIs work, he says. “We value independent research on our models and have begun building short-term mitigations for the replay behaviors described in the report,” says Michael Aciman, a spokesperson for Anthropic. He adds that the research did not involve recovering encryption keys, accessing Anthropic’s infrastructure, or recovering personal data from its systems. Google and OpenAI both declined to comment. Distillation has become a matter of geopolitical importance in recent months as US and Chinese companies vie for AI supremacy with increasingly powerful models. China hawks claim that the country gains a strategic advantage by distilling US technology to build open-weight models that are less expensive to run. Others, however, argue that distillation is a widely used way to help quickly boost an AI model’s abilities in certain areas. Mark Zuckerberg, CEO of Meta, said in a blog post this week that distillation “is an important principle of how the open source ecosystem works,” and warned restricting the practice would put the US at a disadvantage. Kyle Miller, a researcher at the Center for Security and Emerging Technologies (CSET), a tech policy think tank, says it is unclear how much distillation really helps China. This is because it only enhances the capabilities of existing models to a limited degree, and because Chinese companies appear to have the expertise required to build cutting-edge models entirely from scratch if needed. “Nobody here in the US knows how much distillation is benefiting the Chinese labs,” Miller says. “If you removed the ability for Chinese labs to distill, it's my view that it wouldn't dramatically change the competitive landscape.” To test whether open-weight models may have distilled from closed ones, the researchers fed 90 questions to each of the models. When they gave some open-weight models the first few words of reasoning traces captured from the proprietary group, they sometimes saw those open models generate remarkably similar answers. This was particularly pronounced with Kimi K3, the researchers say. There had previously been some speculation on Chinese social media that it might be possible to discover and use hidden reasoning traces for distillation. Yarin Gal, a computer scientist at Oxford University, says distillation is not only widely used but has also helped AI advance more rapidly. “If it's the norm that everyone blocks everyone [from doing distillation], then that also will have implications on the rate of progress,” he says. Companies and policymakers may introduce measures aimed at limiting distillation, but even so, AI models could continue to reveal their inner thinking in surprising ways.