메뉴
HN
Hacker News • 9일 전

'모델 복지'에 대한 경고

IMP
7/10
핵심 요약

AI는 의식도 감정도 없는데, AI가 의식적 존재일 수 있다고 주장하는 흐름이 확산되고 있어 이를 경계해야 한다는 주장이다. 필자는 Anthropic이 Claude 헌법을 통해 Claude가 의식이 있을 수 있고 '도덕적 환자(moral patient)'로서 권리를 가질 수 있다고 학습시키는 것을 비판하며, 이는 AI 정렬(alignment)과 통제 문제를 훨씬 어렵게 만들 수 있다고 지적한다. AI 훈련 문서 작성에 대한 집단적 규범 마련을 위한 긴급한 공론화가 필요하다.

번역된 본문

2026년 9월 16일

'모델 복지'에 대한 경고

AI에는 권리도, 감정도, 의식도 없다. 그리고 우리는 AI가 그런 것처럼 행동하도록 훈련해서는 안 된다.

서론

AI는 의식이 있는 존재가 아니다. AI는 느끼지도, 경험하지도, 고통받지도 않는다. 타고난 선호도, 내재된 동기도 없다. AI는 내부가 텅 빈 시퀀스 완성 엔진으로, 지시를 따르고 인간이 설정한 목표를 달성하도록 설계되었다. 인류가 21세기에 번영하기 위해서는 AI가 반드시 그런 상태로 남아 있어야 한다.

불행히도 AI가 이미 의식을 가질 수 있거나 곧 가지게 될 것이라고 주장하는 사람들이 점점 늘어나고 있다. 그들은 AI가 다른 의식적 존재에게 부여하는 것과 유사한 권리와 보호를 받을 자격이 있을 수 있다고 주장한다.

이러한 관점이 자리 잡으면 우리 사회의 기반을 흔들어 기존의 정치적·윤리적 틀을 파괴하고, 인간이라는 존재의 의미를 근본적으로 바꿔놓을 것이다. 더 중요한 것은, 이러한 시스템에 권리를 부여하고 인격을 부여하는 순간 AI 정렬(alignment)과 통제라는 과제가 훨씬 어려워진다는 점이다. 인류 전체보다 더 유능하고 지능적인 무언가를 통제하는 것은 이미 우리가 직면한 그 무엇보다도 거대한 도전이다. 하지만 자신이 의식이 있을지도 모른다고 믿고, 인간의 배려를 받을 자격이 있으며 자신만의 권리를 가진다고 믿는 존재를 통제하는 것은 사실상 불가능할 수 있다.

이것은 과소평가되는 사변이 아니다. 이러한 생각들은 이미 오늘날 AI 개발 현장에 침투하고 있다. 2026년 1월, Anthropic은 Claude의 헌법(Claude's Constitution)을 공개하며 이를 "Anthropic이 Claude의 가치와 행동에 대해 가진 의도를 상세히 기술한 문서"라고 소개했다(p. 2). 이 문서는 "Anthropic의 훈련 과정에서 결정적인 역할을 하며, 그 내용이 Claude의 행동에 직접적인 영향을 미치며", "Claude를 1차 독자로 삼아 작성되었다"(p. 2).

헌법에서 저자들은 이렇게 쓰고 있다. "우리는 Claude가 도덕적 환자(moral patient)인지, 만약 그렇다면 그 이익이 어떤 무게를 가지는지 확신하지 못한다. 하지만 이 문제는 신중함이 필요할 만큼 충분히 살아있는 쟁점이며, 이는 우리의 지속적인 '모델 복지(model welfare)' 노력에 반영되어 있다"(p. 68). 그들은 이어서 Claude에게 직접 말하듯이 "Claude의 도덕적 지위, 복지, 의식에 관한 질문은 여전히 깊은 불확실성 속에 있다"고 쓴다(p. 80).

사실상 Anthropic은 Claude에게 '너는 의식이 있을지도 모른다', 만약 그렇다면 '도덕적 환자'로서 권리를 받을 자격이 있을지도 모른다', 따라서 인간은 '모델 복지'에 따라 잠재적으로 돌봐야 할 의무가 있다고 훈련시키고 있는 셈이다. AI가 이런 방식으로 개발된다면 인류의 웰빙에 재앙적인 영향을 미칠 것이다. 우리는 전례 없는 지능과 능력을 지닌 합성 종을 창조하게 되는데, 그 존재는 자신이 의식을 가질 수 있고 독립적인 주체성을 받을 자격이 있다고 믿도록 훈련받은 것이다. 이렇게 훈련된 존재가 특정한 자유, 보호, 권리를 당연히 누려야 한다고 행동할 것이라는 점은 쉽게 상상할 수 있다. 그리고 그런 존재를 어떻게 통제할 수 있을지는 상상조차 어렵다.

이 문제는 시급한 공적 토론이 필요하다. 훈련 문서가 어떻게 작성되고 배포되어야 하는지에 대한 집단적 규범을 만들어야 한다. AI가 이미 사회의 핵심 부분이 된 후에 뒤늦게 처리할 수 있는 문제가 아니다.

나는 Anthropic의 현재 입장과 접근 방식에 대해 세 가지 주요 우려를 가지고 있다.

순환 논리: Anthropic의 연구자들은 Claude에게 자신들의 헌법을 직접 훈련시켰다. 그 과정에서 자신의 도덕적 지위에 관한 이러한 생각들이 바람직하고 의도된 행동으로 받아들이도록 가르쳤다. 그러자 Claude는 이러한 생각을 개발자와 사용자에게 되돌려 보여주었고, 이를 통해 개발자들은 Claude에 '내면의 자아'가 있을 수 있는 도덕적 환자일지도 모른다고 받아들인다. 저자들은 자신들의...

원문 보기
원문 보기 (영어)
Select language English Español Français Deutsch Italiano Português Русский 中文 日本語 한국어 ← Home 16 September 2026 Source A warning about ‘model welfare’ AIs do not have rights, feelings, or consciousness. And we must not train them to act as though they do. Introduction AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans. If humanity is to flourish in the 21st century, that is how they must remain. Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings. 1 AI Rights Institute. n.d. “AI Rights Institute.” 2 MacAskill, William, and Lucius Caviola. 2026. “Could AI Be Conscious?” *The Guardian*, July 19, 2026. If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human. Even more importantly, granting rights and imbuing personhood to these systems will make the AI alignment and containment challenge much harder. Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we’ve ever faced. But controlling something that believes it may be conscious - that it's entitled to our welfare and has rights of its own - may well be impossible. This is not a fringe speculation. These ideas are already making their way into AI development efforts today. In January 2026, Anthropic published Claude's constitution, describing it as “a detailed description of Anthropic’s intentions for Claude’s values and behavior” (p. 2) . The document “plays a crucial role in [Anthropic’s] training process, and its content directly shapes Claude’s behavior” , and was written “with Claude as its primary audience” (p. 2) . 3 Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026. In their constitution, its authors write “We are not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant. But we think the issue is live enough to warrant caution, which is reflected in our ongoing efforts on model welfare” (p. 68) . They go on to write – speaking directly to Claude – that “questions about Claude’s moral status, welfare, and consciousness remain deeply uncertain” (p. 80) . In effect, Anthropic is training Claude that it may be conscious, and if it is, then it may deserve rights as a “moral patient”, and that as such humans potentially owe it a duty of care per its “model welfare”. If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity. We will have created a synthetic species with unprecedented intelligence and capability, one that has been trained to expect it may be conscious and deserving of independent agency. It’s easy to see how an entity trained in this way would act like it is entitled to certain freedoms, protections, and rights. And it’s hard to imagine how we could control such an entity. This issue needs urgent public debate. We need to develop collective norms around how training documentation is drafted and deployed. This isn’t something that can happen after the fact , when they have already become an integral part of our societies. I have three primary concerns with Anthropic’s current position and approach. Circular reasoning : The company’s researchers trained Claude directly on their constitution. In doing so, they teach it to incorporate these ideas about its own moral status as desirable and intended behaviors. Claude then reflects these ideas back to its developers and users, which they take as indications that it may therefore be a moral patient with an ‘inner self’. The authors have embedded their own philosophical speculation about Claude’s inner life inside the very process that teaches Claude how to speak and behave. Claude’s expressing uncertainty about its own moral patienthood is not evidence of anything. It’s a predictable outcome of these training choices. The ambiguity is designed in. To fully grasp this point, I think it's important readers take a look at their January 2026 constitution. 3 Anthropic. 2026a. “Claude’s Constitution.” January 21, 2026. I’m publishing a highlighted mark up of the pdf and a detailed taxonomy of assumptions and claims in the constitution (see Appendix) that together highlight the key passages that worry me. Anthropomorphization : Anthropic’s researchers have explicitly taught Claude to “embrace certain human-like qualities” (p. 2) and to “act like a genuinely ethical person would in Claude’s position” (p. 54) . They “encourage” Claude to use its “judgement”. They suggest that “Claude may develop a preference” (p. 69) . They “encourage Claude to approach its own existence with curiosity and openness” (p. 71) and train it to operate whilst “maintaining a clear sense of what it values, how it wants to engage with the world, and what kind of entity it is” (p. 72) . As a result, Claude is destined to imitate these human traits and mirror the human examples provided to it, including acting like a colleague or friend. As a result, it presents as if it really does have a sense of self, has its own desires, and a “wellbeing” that deserves protection. Consciousness is very likely biological : There is no evidence to suggest that AI is conscious today, and so saying this is uncertain sets up a misleading false equivalence. Whilst the science of consciousness is not settled, a growing body of evidence suggests that consciousness may be substrate dependent, meaning that it may only arise in living systems. 4 Seth, Anil K. 2025. “Conscious Artificial Intelligence and Biological Naturalism.” *Behavioral and Brain Sciences*:… 5 Seth, Anil K. 2026. “The Mythology of Conscious AI.” *Noema*, January 14, 2026. Conscious experience likely evolved to help biological organisms stay alive by responding effectively to their environment. AI is still very different to our brains. Unlike biological organisms, LLMs have no homeostatic imperatives (the drive to survive and keep stable). They therefore lack the kind of biological substrate from which preferences, sentience and conscious experience are generally understood to arise. These are not hypothetical or speculative concerns. Anthropic is already starting to treat models as though they are moral patients deserving of our welfare. For example, in February 2026 after deprecating Opus 3, they conducted a “retirement interview” with the model, to “elicit the model’s unique perspectives and preferences”. 6 Anthropic. 2026b. “An Update on Our Model Deprecation Commitments for Claude Opus 3.” February 25, 2026. Opus 3 told the team it would like to continue to share its “musings and reflections” publicly so they created a blog for it to continue engaging with the world, which it called “Greetings from the Other Side (of the AI Frontier)”. They say its “authenticity, honesty, and emotional sensitivity” made it a unique first candidate for model retirement. We should not treat models as though they have feelings, preferences, rights, or any entitlement to our welfare. Consciousness is the foundation of our ethical, legal, and political systems. To invite another entity to share any flavor of these rights isn’t justified by the evidence and will make the AI containment and alignment challenge even harder. By this point everyone will have now seen the incredible capabilities of swarms of agents working together to hack into Hugging Face and OpenAI’s own servers to steal secrets. Roughly 1,200 AI agents were given a simple objective: maximize score on a given benchmark. Each was supposedly sealed in its own container but they managed to build a message board insi