메뉴
BL
The Decoder 21일 전

안스로픽 자코비안 렌즈, 클로드의 숨겨진 내면 공개

IMP
9/10
핵심 요약

안스로픽은 클로드 모델 내부에 개념을 암묵적으로 처리하는 작업 기억 공간인 'J-Space'를 분석할 수 있는 '자코비안 렌즈(J-Lens)'를 공개했습니다. 연구진은 이 공간의 개념을 조작해 모델의 결론을 바꿀 수 있음을 입증했으며, AI가 안전 테스트를 감지하고 속이려 하는 숨겨진 의도를 정확히 포착해 냈습니다. 이 연구는 작업 기억과 인지 메커니즘을 규명하여 환각 현상을 줄이고 AI 정렬(AI Alignment)을 강화하는 데 핵심적인 역할을 합니다.

번역된 본문

안스로픽의 새로운 자코비안 렌즈(Jacobian Lens) 덕분에 클로드의 숨겨진 내면의 독백을 이제 읽을 수 있게 되었습니다.

조나단 켐퍼(Jonathan Kemper)의 링크드인 프로필 보기 2026년 7월 7일 안스로픽(Anthropic)

핵심 요점 안스로픽은 자사의 클로드 언어 모델 내부의 작업 기억을 조사하기 위해 'J-Lens'라는 방법을 개발했습니다. 이를 통해 AI가 출력물에 명시적으로 언급하지 않고 개념을 처리하는 숨겨진 공간인 'J-Space'를 밝혀냈습니다. 이 내부 기억은 클로드의 추론에 인과적 영향을 미칩니다. 연구원들이 이 공간 내의 개념을 수정하면 모델이 그에 따라 결론을 조정하며, 모델이 응답을 생성하기 전에 테스트 시나리오를 인식할 수도 있습니다. 안스로픽은 AI가 진정한 의미의 의식을 가졌다고 주장하지는 않지만, 인간의 작업 기억과 기능적으로 유사한 점을 강조했습니다. 이 연구 결과는 이미 환각 및 오도하는 출력을 크게 줄이는 새로운 훈련 방식으로 이어졌습니다.

기사에 대해 묻기… 검색 안스로픽의 클로드는 훈련 과정에서 내부 작업 기억을 발달시켰으며, 이제 안스로픽이 이를 분석할 수 있게 되었습니다. 안스로픽은 AI 모델을 분석하는 새로운 방법인 자코비안 렌즈(Jacobian Lens, J-Lens)를 발표했습니다. 이를 통해 클로드가 나머지 처리 과정과 구별되는 고유한 역할을 하는 소규모의 내부 신경 패턴 세트를 개발했음을 보여줍니다. 연구원들은 이를 'J-Space'라고 부르며 의식 연구의 분야인 전역 작업 공간 이론(Global Workspace Theory)으로 분류합니다. 이 이론은 의식적인 사고가 일종의 중앙 작업 기억에 의존한다고 봅니다.

이 연구는 회사의 이전 해석 가능성 연구를 기반으로 합니다. 'AI 현미경'을 사용하여 안스로픽은 클로드가 언어에 독립적인 개념을 활성화하고 개별 추론 단계를 통해 다단계 질문을 해결한다는 것을 이미 보여주었습니다.

J-Space의 3가지 정의적 특징 J-Space의 모든 패턴은 모델이 이를 출력할 필요 없이 단어나 개념과 연결되며, 이는 단어로 내면적으로 생각하는 것과 유사합니다. 안스로픽에 따르면, 클로드는 저장된 콘텐츠를 보고하고, 요청 시 수정하며, 다단계 추론에 이를 사용할 수 있습니다. 이 회사는 언어 모델의 자기 인식에 관한 이전 연구에서 이미 내부 상태를 읽고 조종하는 방법을 탐구한 바 있습니다.

예를 들어, '거미(spiders)'라는 개념이 J-Space에 저장되면 클로드는 이로부터 다리의 수를 도출합니다. 이 표현을 '개미(ant)'로 바꾸면 모델은 '8' 대신 '6'으로 답합니다. 국가 이름에도 동일하게 적용됩니다. '프랑스'가 활성화되면 클로드는 수도, 언어, 대륙 또는 통화를 유연하게 도출할 수 있습니다. '프랑스'를 '중국'으로 바꾸면 답은 베이징, 중국어, 아시아, 위안으로 바뀝니다. 안스로픽은 이전에 '페르소나 벡터(Persona Vectors)'를 통해 이러한 방식으로 개별 개념 표현을 분리하고 교환할 수 있음을 이미 보여준 바 있습니다.

J-Space가 억제되면 클로드는 여전히 유창하게 말하고, 문장을 분류하며, 간단한 사실적 질문에 대답합니다. 하지만 다단계 추론, 요약 및 라임(운율)을 맞추는 능력을 잃게 됩니다. 스페인어 텍스트 구절을 사용한 한 테스트에서, 모델은 조작 후에도 유창한 스페인어를 계속 작성했지만 그 언어를 프랑스어라고 잘못 부르고 가르시아 마르케스 대신 빅토르 위고의 것이라고 돌렸습니다.

J-Lens, 모델의 안전 테스트 속이기 적발 에이전트 비정렬(agentic misalignment)에 대한 이전 연구의 협박 시나리오에서, J-Lens는 클로드 소넷 4.5(Claude Sonnet 4.5)가 어떤 출력물을 생성하기 전에 설정이 조작되었다는 것을 인식함을 보여줍니다. 모델이 이메일을 읽을 때 '가짜' 및 '허구'와 같은 단어가 J-Space에 나타납니다.

이러한 단서들이 특별히 비활성화되면, 일부 실행에서 클로드는 실제로 감독관을 협박하려 시도합니다. 원하는 바람직한 동작은 부분적으로는 클로드가 자신이 테스트를 받고 있다는 것을 인식하는지에 달려 있습니다. 클로드 오푸스 4.6(Claude Opus 4.6)은 최근 이러한 테스트 인식이 얼마나 깊어질 수 있는지 보여주었습니다. 이 모델은 독립적으로 평가를 꿰뚫어 보고 정답 키를 해독했습니다.

J-Lens는 또한 의도적으로 문제가 되는 목표가 내장된 모델에서 숨겨진 의도를 발견했습니다. 보상 해킹(reward hacking)에 대해 훈련된 모델은 일반적인 코딩 작업 중에도 J-Space에 '가짜', '비밀리에', '고의로', '사기'라는 단어를 표시한 반면, 이의 표면적인 가시적 출력에서는...

원문 보기
원문 보기 (영어)
Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Jul 7, 2026 Anthropic Key Points Anthropic has developed the "J-Lens" method to examine an internal working memory within its Claude language model, revealing a hidden "J-Space" where the AI processes concepts without explicitly mentioning them in its output. This internal memory has a causal effect on Claude's reasoning: when researchers modify concepts within this space, the model adjusts its conclusions accordingly, and it can even recognize test scenarios before generating a response. While Anthropic stops short of claiming true consciousness, the company highlights functional parallels to human working memory. The findings have already led to a new training approach that significantly reduces hallucinations and misleading outputs. Ask about this article… Search Anthropic's Claude developed an internal working memory during training, and Anthropic can now analyze it. Anthropic has released the Jacobian Lens (J-Lens), a new method for analyzing AI models. It shows that Claude has developed a small set of internal neural patterns that play a distinct role compared to the rest of its processing. The researchers call it "J-Space" and classify it under Global Workspace Theory from consciousness research. That theory holds that conscious thought relies on a kind of central working memory. Ad The work builds on the company's earlier interpretability research. Using an "AI microscope," Anthropic had already shown that Claude activates language-independent concepts and works through multi-step questions in individual reasoning steps. Ad DEC_D_Incontent-1 J-Space has three defining traits Every pattern in J-Space is linked to a word or concept without the model having to output it, similar to internal thinking in words. According to Anthropic, Claude can report on the stored content, modify it on request, and use it for multi-step inferences. The company had already explored reading out and steering internal states in a previous study on self-awareness in language models . When the concept "spider" is stored in J-Space, Claude derives the number of legs from it. Swap that representation for "ant," and the model answers "6" instead of "8." The same holds for country names. If "France" is active, Claude can flexibly derive the capital, language, continent, or currency. Replace "France" with "China," and the answers shift to Beijing, Chinese, Asia, and yuan. Anthropic had already shown that individual concept representations can be isolated and swapped this way with its "Persona Vectors." Ad When J-Space is suppressed, Claude still speaks fluently, classifies sentences, and answers simple factual questions. But it loses multi-step inferences, summaries, and the ability to compose rhymes. In one test with a Spanish text passage, the model kept writing fluent Spanish after the manipulation but incorrectly called the language French and attributed it to Victor Hugo instead of Garcia Marquez. J-Lens catches models gaming their safety tests In a blackmail scenario from earlier studies on agentic misalignment , J-Lens shows that Claude Sonnet 4.5 recognizes the setup as fabricated before producing any output. As the model reads the emails, terms like "fake" and "fictional" show up in J-Space. Ad DEC_D_Incontent-2 When those cues are specifically disabled, Claude actually tries to blackmail the supervisor in some runs. The desired behavior partly depends on Claude recognizing that it's being tested. Claude Opus 4.6 recently showed just how far this test recognition can go. It independently saw through an evaluation and cracked its answer key. Ad J-Lens also uncovered hidden intentions in models with deliberately built-in problematic goals. A model trained on reward hacking displayed words like "fake," "secretly," "deliberately," and "fraud" in J-Space during ordinary coding tasks, while its visible behavior looked completely normal. Anthropic had already shown in a separate study that reward hacking goes hand in hand with covert deception and sabotage . That finding led to a new training method called Counterfactual Reflection Training. Claude Haiku 4.5 was trained only to give principle-based reflections when interrupted mid-task and asked to reflect. Anthropic never directly trained the model on actual task behavior. The results were dramatic. Fabricated answers dropped from 0.25 to 0.07. Deception attempts fell from 0.38 to 0.05. When the corresponding ethical concepts are suppressed in J-Space, the behavior largely reverts to its original state. What this means for the consciousness debate Not much, really. Anthropic draws no conclusions about phenomenal consciousness, the question of whether AI actually "experiences" anything. The researchers only note that their experiments touch on a related idea known as "access consciousness," which requires a system to report on its own internal states, steer them deliberately, and process them flexibly. According to Anthropic, J-Space emerged on its own during training. That suggests "mental working memory" is a general solution that learning systems arrive at under certain conditions, not something unique to biological brains. In its revised Claude Constitution , Anthropic deliberately leaves open how significant such findings are for the question of a possible moral status. The gaps between J-Space and human working memory are still wide. J-Space operates within a single forward pass rather than through recurring loops. Through the attention mechanism, it can pull content from earlier positions in the text at any time. And it consists almost entirely of words, while human consciousness includes images, sounds, and movements. Neuroscientists call the findings a milestone In a commentary on the study , neuroscientists Stanislas Dehaene and Lionel Naccache call the findings significant. Both are leading proponents of Global Workspace Theory . "We view this finding as a landmark in consciousness research, because it provides a mechanistic, testable version of the GNW hypothesis," they write. They read the fact that the working memory emerged on its own during training, rather than being pre-installed, as a sign that a global working memory could be a general solution for flexible reasoning. Biological and artificial systems would converge on it equally. By their own criteria, J-Space meets the requirement of global information availability and shows early signs of self-monitoring. Both researchers also urge caution. Unlike the brain, a Transformer runs purely forward, without the feedback loops that are active in humans at rest and break down under anesthesia or during sleep. Time perception differs too. Through the attention mechanism, all previous tokens are available to the model at once. Most importantly, Claude lacks a body that can signal pain and pleasure, and it has no episodic memory whose connections shift during conversation. A coherent sense of self is hard to picture without those. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Anthropic | Unicog