메뉴
HN
Hacker News 22일 전

AI 모델 내부에 나타난 '의식 작업 공간'의 발견

IMP
8/10
핵심 요약

Anthropic의 연구진은 Claude 모델 내부에서 인간의 '의식적 접근'과 유사한 특수한 내부 신경 패턴인 'J-space'가 자발적으로 형성되었음을 발견했습니다. J-space는 AI가 텍스트로 출력하지 않고도 내면에서 조용히 개념을 떠올려 다단계 추론이나 제어된 사고를 수행하는 데 핵심적인 역할을 합니다. 이는 AI의 단순한 문장 생성을 넘어, 고차원적인 인지 작용과 추론 메커니즘을 신경과학적 관점에서 해체하고 이해하는 데 매우 중요한 의미를 갖습니다.

번역된 본문

해석 가능성: 언어 모델 내의 글로벌 작업 공간 2026년 7월 6일 논문 읽기

독자 여러분이 이 문장을 읽고 있는 지금, 여러분 뇌의 회로는 자세를 조정하고, 호흡을 제어하며, 화면상의 선과 곡선을 인식 가능한 단어로 변환하고 있습니다. 이러한 처리 과정의 대부분은 여러분 자신에게 보이지 않습니다. 하지만 뇌에서 일어나는 일 중 일부는 여러분이 접근할 수 있습니다. 즉, 머릿속에 떠오르는 이미지나 어디로 쇼핑하러 갈지에 대해 의도적으로 세우는 계획 등입니다. 신경과학자와 철학자들은 때때로 이러한 후자의 뇌 활동을 무의식적으로 일어나는 다른 모든 처리 과정과 구별하기 위해 '의식적으로 접근 가능한(consciously accessible)' 활동이라고 부릅니다.

이러한 활동은 특별한 속성을 가집니다. 우리는 이를 묘사하고, 제어하며, 우리가 알아차리지 못하는 채 자동으로 일어나는 모든 처리 과정과는 대조적으로 이를 의도적인 추론에 사용할 수 있습니다. 새로운 논문에서, 우리는 Claude와 같은 최신 언어 모델에서도 이와 유사한 구분이 나타났다는 증거를 제시합니다. 우리는 Claude가 다른 모든 내부 처리 과정에 비해 특별한 역할을 하는 소수의 내부 신경 패턴을 발전시켰다는 것을 발견했습니다. 우리는 이러한 패턴의 집합을 발견하는 데 사용한 '야코비안(Jacobian)'이라는 수학적 개념의 이름을 따서 'J-space'라고 부릅니다.

각 J-space 패턴은 특정 단어와 연결되어 있습니다. 하지만 이러한 패턴 중 하나가 활성화된다고 해서 모델이 그 단어를 말한다는 것을 의미하지는 않으며, 단지 모델이 그 단어를 '생각하고 있음'을 의미합니다. 언어 모델이 추론하는 동안 스스로에게 작성하는 텍스트인 '스크래치패드(scratchpad)'나 '사고의 사슬(chain of thought)'에 대해 들어보셨다면, J-space는 이와 다릅니다. 이는 모델의 내부 신경 활성화 내에서 조용히 작동하여, 모델이 단어를 적지 않고도 개념에 대해 생각할 수 있게 해줍니다. 특히 J-space는 우리가 설계하거나 프로그래밍한 것이 아니라 Claude의 훈련 과정 중에 스스로 출현(발현)한 것입니다.

우리는 Claude의 다른 처리 과정과 비교하여 J-space가 여러 가지 고유한 속성을 가지고 있음을 발견했습니다. 첫째, Claude는 이러한 표현에 대해 보고할 수 있습니다. Claude에게 무엇을 생각하고 있는지 물어보면, J-space에 있는 것을 말해줍니다. 반면 J-space가 아닌 표현은 보고하기가 더 어렵습니다.

둘째, 요청 시 이를 조절(modulate)할 수 있습니다. Claude에게 무언가를 생각해 보라고 하거나 머릿속으로 소리 없이 문제를 풀도록 요청하면, J-space에서 적절한 패턴이 빛을 발합니다. 이와 대조적으로, J-space에 있지 않은 패턴은 조절하는 데 어려움을 겪습니다.

셋째, Claude는 내부적 추론을 위해 J-space를 사용합니다. 여러 단계가 필요한 문제를 풀도록 Claude에게 요청하면, 소리 내어 말하지 않더라도 중간 단계들이 J-space에서 활성화됩니다. 이러한 J-space 패턴은 다른 표현들보다 규모가 작음에도 불구하고, 이러한 작업에서 모델의 성능을 인과적으로 매개합니다. J-space의 표현은 다양한 작업에 유연하게 사용될 수 있습니다. 예를 들어, Claude의 J-space에서 '프랑스'가 활성화되면 모델은 그 수도, 국가 통화 또는 속한 대륙을 떠올릴 수 있습니다.

하지만 J-space는 중요한 역할을 함에도 불구하고 언어 모델이 수행하는 대부분의 작업(유창하게 말하기, 단순한 사실 기억하기, 올바른 문법 사용하기 등)에는 관여하지 않습니다. Claude가 J-space를 사용하지 못하도록 방해한 실험에서, Claude는 여전히 정상적으로 상호작용했지만 고차원적인 인지 기능을 잃었습니다.

우리의 실험은 의식적 접근이 어떻게 작동하는지 설명하기 위해 개발된 신경과학의 저명한 이론인 '글로벌 작업 공간 이론(Global Workspace Theory)'에서 영감을 받았습니다. 이 이론은 뇌를 무의식적으로 병렬 처리하며 대체로 서로 격리된 상태로 작동하는 전문가 시스템들의 집합으로 묘사합니다. 정보의 일부가 다른 뇌 시스템이 볼 수 있고 활용할 수 있는 '작업 공간(workspace)'이라는 작은 공유 채널로 진입하게 되면, 비로소 의식적으로 접근 가능해진다는 것입니다. 우리의 연구 결과를 바탕으로, 우리는 J-space가 Claude 안에서 이와 유사한 '작업 공간' 역할을 한다고 생각합니다. 예를 들어, 우리는 Claude의 J-space가 신경망의 나머지 부분과 특히 강력한 연결을 가지고 있다는 증거를 발견했으며, 이를 통(해 생태계 전반에 정보를 방송합니다.)

원문 보기
원문 보기 (영어)
Interpretability A global workspace in language models Jul 6, 2026 Read the paper As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all the other processing that goes on unconsciously. This activity has special properties: we can describe it, control it, and use it for deliberate reasoning, in contrast to all the automatic processing that goes on without our awareness. In a new paper, we present evidence that a similar distinction has emerged in modern language models like Claude. We find that Claude has developed a small collection of internal neural patterns that, compared to all its other internal processing, play a special role. We call the collection of these patterns the J-space —named after the technique we used to find them, involving a mathematical concept called the Jacobian. Each J-space pattern is linked to a particular word. But when one of these patterns lights up, it doesn’t mean the model is saying that word—just that the word is on its mind. If you've heard of language models having a "scratchpad" or “chain of thought”—text they write to themselves while reasoning—the J-space is something different. It operates silently, in the model’s internal neural activations, allowing the model to think about a concept without writing it down. Notably, the J-space wasn’t designed or programmed by us, but instead emerged on its own during Claude’s training process. We find that the J-space has a number of unique properties, compared to the rest of Claude's processing: Claude can report on these representations. If you ask Claude what it's thinking about, it will tell you what’s in the J-space. Non-J-space representations are less reportable. It can also modulate them on request. If you ask Claude to think about something, or solve a problem silently in its head, it will light up the appropriate patterns in its J-space. By contrast, it has trouble modulating patterns not in the J-space. Claude uses its J-space for internal reasoning. If you ask Claude to solve a problem that requires multiple steps, the intermediate steps will light up in its J-space, even when it doesn’t say them out loud. These J-space patterns causally mediate its performance in such tasks, despite being smaller in magnitude than other representations. Representations in the J-space can be used flexibly for many tasks—for example, once “France” has lit up in Claude’s J-space, the model can recall its capital, or its national currency, or the continent it belongs to. However, despite its important role, the J-space is not involved in most of what a language model does—speaking fluently, recalling simple facts, using correct grammar, etc. In experiments where we prevented Claude from using its J-space, it still interacted normally, but lost its higher-order cognitive functions. Our experiments were inspired by a prominent theory in neuroscience that was developed to explain how conscious access works: the global workspace theory . This account pictures the brain as a collection of specialist systems that work in parallel, unconsciously, and largely in isolation from one another. A piece of information becomes consciously accessible when it gains entry to a small shared channel, the “workspace,” which is broadcast to other brain systems that can see it and make use of it. Based on our findings, we think the J-space plays a similar “workspace” role in Claude. For example, we find evidence that Claude’s J-space has especially strong connections to the rest of its neural network, allowing it to fulfill this kind of broadcasting role. None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all; we’ll come back to that question at the end of the post. But whatever its philosophical significance, the J-space is a practically useful tool for us, as it gives us a way to see what Claude is thinking but not saying. For instance, we’re able to use it to catch Claude privately noticing that it’s being tested, intentionally producing fabricated data, or pursuing a hidden goal that we planted during training. We’ve also developed a technique to influence what lights up in Claude’s J-space, and thereby influence its decision-making. More broadly, these findings have changed our understanding of how Claude’s mind works, revealing a privileged mental workspace that can be used for deliberate reasoning, operating amidst a sea of more automatic, inflexible processing. Rather than being a chaotic jumble of numbers, Claude’s internals have organized themselves in a way that is reminiscent of our own minds. This post is a short summary of a much more extensive research paper , where you can find more detail on our experiments. We’ve also released a code repository with an open-source implementation of the core methods, and have partnered with Neuronpedia to provide an interactive demo of our methods on open-weights models. To provide additional perspectives on the broader implications of this work, we also invited commentary from several experts in neuroscience, philosophy, and LLM interpretability, which can be viewed here . How we found the J-space The starting point for this research was inspired by one of the key features of consciously accessible thoughts in humans: they can, unlike un conscious processing, often be put into words. If a thought is consciously accessible to you, you can typically describe it if someone asks. We went looking for representations in Claude with the same property: representations that are positioned to influence what Claude might say—not necessarily what it’s saying right now, but what it could talk about, if asked. Our technique is called the Jacobian lens, or J-lens for short. For every word in Claude's vocabulary, the J-lens finds the internal activity pattern that makes Claude more likely to say that word at some point in the future. When we apply the lens to Claude’s internal activity, we get a list of words—the contents of the J-space at that moment—which we can simply read. Claude processes text through a series of multiple internal stages called layers, and by applying this technique over different layers, we can watch these silent words in the J-space evolve as the model works through what to say. What shows up in the J-space goes well beyond the text Claude is reading or writing. When Claude reads code with a bug that nobody has pointed out, its J-space contains “ERROR.” When it reads the raw letters of a protein sequence, the J-space contains the protein's biological function. When it reads search results that are secretly an attempt to manipulate it (an attack known as a “prompt injection”), the J-space contains “injection” and “fake.” When we ask Claude a multi-step math problem, the intermediate steps pop up in the J-space, in the right order. So even though the J-space was discovered by looking for representations that could be spoken, it nevertheless uncovers Claude’s internal thoughts. In a sense, this is similar to how some people “think in words,” without having to say them out loud. Claude reports what’s in its J-space Our first set of experiments tested how the J-space is involved in Claude’s verbal reports. In one experiment, we ask Claude to silently think of an item from some category—a sport, say—and then name it. If we read the J-lens right before Claude answers, we can see what it picked: “Soccer” is at the top of the list, and sure enough, Claude says “soccer.” By itself, though, this is just a correlation