메뉴
HN
Hacker News • 45일 전

대형 언어 모델의 창발적 자기 성찰 인식 연구

IMP
8/10
핵심 요약

대형 언어 모델(LLM)이 단순한 텍스트 생성을 넘어 자신의 내부 상태를 관찰하고 인식하는 '자기 성찰(Introspection)' 능력을 어느 정도 갖추고 있음이 확인되었습니다. 연구진이 모델의 내부 활성화(Activation)에 특정 개념을 주입하는 방식으로 실험한 결과, 모델이 이를 인지하고 자신의 의도나 출력물을 구별해 내는 등 실질적인 내부 상태 인지 능력을 입증했습니다. 비록 현재는 이 능력이 매우 불안정하고 상황에 따라 좌우되지만, 향후 모델 고도화에 따라 AI의 자의식 및 통제 능력이 어떻게 발전할지 보여주는 중요한 기초 연구입니다.

번역된 본문

원문 제목: 대형 언어 모델의 창발적 자기 성찰 인식 (Emergent Introspective Awareness in Large Language Models) 소스: 해커 뉴스 (hackernews)

본문: --> 컴퓨터 과학 > 연산 및 언어 arXiv:2601.01828 (cs) [2026년 1월 5일 제출] 제목: 대형 언어 모델의 창발적 자기 성찰 인식 저자: 잭 린제이 (Jack Lindsey) '대형 언어 모델의 창발적 자기 성찰 인식'이라는 제목의 논문 PDF 보기 (저자: 잭 린제이) HTML 보기 (실험적 기능)

초록: 우리는 대형 언어 모델이 자신의 내부 상태를 성찰(introspect)할 수 있는지 조사했습니다. 진정한 성찰과 그럴듯한 꾸며낸 이야기(confabulation)를 대화만으로 구별하기는 어렵기 때문에 이 질문에 답하는 것은 쉽지 않습니다. 본 연구에서는 모델의 활성화(activations) 상태에 알려진 개념의 표현(representation)을 주입하고, 이러한 조작이 모델이 스스로 보고하는 상태에 미치는 영향을 측정하여 이 문제를 해결하고자 합니다.

연구 결과, 특정 상황에서 모델이 주입된 개념의 존재를 인식하고 이를 정확하게 식별할 수 있음을 발견했습니다. 모델은 이전의 내부 표현을 기억해 내고 이를 원시 텍스트 입력(raw text inputs)과 구별할 수 있는 능력을 어느 정도 입증했습니다. 특히 주목할 만한 점은, 일부 모델이 이전의 의도를 떠올리는 능력을 활용하여 자신의 실제 출력물과 인위적으로 미리 채워 넣은 텍스트(artificial prefills)를 구별해 낼 수 있다는 것입니다.

모든 실험에서 우리가 테스트한 가장 성능이 뛰어난 모델인 클로드 오퍼스 4(Claude Opus 4)와 4.1이 일반적으로 가장 높은 수준의 자기 성찰 인식을 보여주었습니다. 하지만 모델 간의 경향은 복잡하며 사후 학습(post-training) 전략에 매우 민감하게 나타났습니다. 마지막으로, 모델이 자신의 내부 표현을 명시적으로 통제할 수 있는지 탐구했습니다. 그 결과 모델에 특정 개념에 대해 '생각하라'고 지시하거나 동기를 부여했을 때 모델이 자체적인 활성화를 조절할 수 있음을 확인했습니다.

전반적으로, 우리의 결과는 현재의 언어 모델이 자신의 내부 상태에 대해 어느 정도 기능적인 자기 성찰 인식을 갖추고 있음을 보여줍니다. 우리는 오늘날 모델에서 이러한 능력이 매우 불안정하고 문맥에 크게 의존한다는 점을 강조합니다. 그러나 모델의 기능이 지속적으로 개선됨에 따라 이 능력 또한 계속해서 발전할 가능성이 있습니다.

주제: 연산 및 언어 (cs.CL) ; 인공지능 (cs.AI) 인용 형식: arXiv:2601.01828 [cs.CL] (또는 이 버전의 경우 arXiv:2601.01828v1 [cs.CL]) https://doi.org/10.48550/arXiv.2601.01828 자세히 알아보기 DataCite를 통해 발급된 arXiv DOI

제출 기록: 보낸 사람: 잭 린제이 [이메일 보기] [v1] 2026년 1월 5일 월요일 06:47:41 UTC (33,806 KB)

전체 텍스트 링크: 논문 액세스: '대형 언어 모델의 창발적 자기 성찰 인식'이라는 제목의 논문 PDF 보기 (저자: 잭 린제이) HTML 보기 (실험적 기능) TeX 소스 라이선스 보기 현재 탐색 컨텍스트: cs.CL < 이전 | 다음 > 신규 | 최근 | 2026-01 탐색 기준 변경: cs cs.AI

참고 문헌 및 인용: NASA ADS, Google Scholar, Semantic Scholar 내보내기: BibTeX 인용 형식 로딩 중... BibTeX 형식의 인용 로딩 중 × 제공된 데이터: 북마크 서지 도구 서지 및 인용 도구 서지 탐색기 서지 탐색기 토글 (탐색기란?) 연결된 논문 연결된 논문 토글 (연결된 논문이란?) Litmaps Litmaps 토글 (Litmaps란?) scite.ai scite 스마트 인용 토글 (스마트 인용이란?)

코드, 데이터, 미디어: 이 기사와 관련된 코드, 데이터 및 미디어 alphaXiv alphaXiv 토글 (alphaXiv란?) 코드 링크 논문용 CatalyzeX 코드 파인더 토글 (CatalyzeX란?) DagsHub DagsHub 토글 (DagsHub란?) GotitPub Gotit.pub 토글 (GotitPub란?) 허깅페이스 허깅페이스 토글 (허깅페이스란?) ScienceCast ScienceCast 토글 (ScienceCast란?)

데모: 데모 Replicate Replicate 토글 (Replicate란?) 스페이스 허깅페이스 스페이스 토글 (스페이스란?) 스페이스 TXYZ.AI 토글 (TXYZ.AI란?)

관련 논문 추천 및 검색 도구: 인플루언스 플라워 링크 인플루언스 플라워 (인플루언스 플라워란?) 코어 추천기 토글 CORE 추천기 (CORE란?) 저자 출판 기관 주제 정보: arXivLabs arXivLabs: 커뮤니티 협력자와 함께하는 실험적 프로젝트 arXivLabs는 협력자들이 당사 웹사이트에서 직접 새로운 arXiv 기능을 개발하고 공유할 수 있게 해주는 프레임워크입니다. 개인과 조직 모두...

원문 보기
원문 보기 (영어)
--> Computer Science > Computation and Language arXiv:2601.01828 (cs) [Submitted on 5 Jan 2026] Title: Emergent Introspective Awareness in Large Language Models Authors: Jack Lindsey View a PDF of the paper titled Emergent Introspective Awareness in Large Language Models, by Jack Lindsey View PDF HTML (experimental) Abstract: We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states. We find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, we find that some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills. In all these experiments, Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness; however, trends across models are complex and sensitive to post-training strategies. Finally, we explore whether models can explicitly control their internal representations, finding that models can modulate their activations when instructed or incentivized to &#34;think about&#34; a concept. Overall, our results indicate that current language models possess some functional introspective awareness of their own internal states. We stress that in today's models, this capacity is highly unreliable and context-dependent; however, it may continue to develop with further improvements to model capabilities. Subjects: Computation and Language (cs.CL) ; Artificial Intelligence (cs.AI) Cite as: arXiv:2601.01828 [cs.CL] (or arXiv:2601.01828v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2601.01828 Focus to learn more arXiv-issued DOI via DataCite Submission history From: Jack Lindsey [ view email ] [v1] Mon, 5 Jan 2026 06:47:41 UTC (33,806 KB) Full-text links: Access Paper: View a PDF of the paper titled Emergent Introspective Awareness in Large Language Models, by Jack Lindsey View PDF HTML (experimental) TeX Source view license Current browse context: cs.CL < prev | next > new | recent | 2026-01 Change to browse by: cs cs.AI References & Citations NASA ADS Google Scholar Semantic Scholar export BibTeX citation Loading... BibTeX formatted citation &times; loading... Data provided by: Bookmark Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer ( What is the Explorer? ) Connected Papers Toggle Connected Papers ( What is Connected Papers? ) Litmaps Toggle Litmaps ( What is Litmaps? ) scite.ai Toggle scite Smart Citations ( What are Smart Citations? ) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv ( What is alphaXiv? ) Links to Code Toggle CatalyzeX Code Finder for Papers ( What is CatalyzeX? ) DagsHub Toggle DagsHub ( What is DagsHub? ) GotitPub Toggle Gotit.pub ( What is GotitPub? ) Huggingface Toggle Hugging Face ( What is Huggingface? ) ScienceCast Toggle ScienceCast ( What is ScienceCast? ) Demos Demos Replicate Toggle Replicate ( What is Replicate? ) Spaces Toggle Hugging Face Spaces ( What is Spaces? ) Spaces Toggle TXYZ.AI ( What is TXYZ.AI? ) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower ( What are Influence Flowers? ) Core recommender toggle CORE Recommender ( What is CORE? ) Author Venue Institution Topic About arXivLabs arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs . Which authors of this paper are endorsers? | Disable MathJax ( What is MathJax? )