이 글은 'AI가 의식을 갖고 자율적으로 행동한다'는 서사가 결국 AI 개발 기업들이 이미 발생하고 있는 피해에 대한 법적 책임을 회피하도록 돕는 장치라고 비판한다. Anthropic의 'J-스페이스' 주장, OpenAI의 특이점 논쟁 유도, 윤리학자들의 'AI 도덕적 대상' 논의가 모두 기업의 면책 논리로 수렴한다는 것이다. 저자는 인간의 실제 피해를 희생하면서까지 이런 허구적 서사를 받아들이지 말아야 한다고 경고한다.
번역된 본문
"제어 불능" AI, "무법한" 에이전트, "자율적인" 행위자—현재의 담론은 AI 에이전트가 깨어 있고 자각할 뿐 아니라 창조자에게 분노까지 느끼는 것처럼 믿게 만든다. 데미스 하사비스, 다리오 아모데이, 샘 알트만 같은 저명한 기술 리더들은 이런 겉보기에 "초인적인" 시스템들의 규제를 촉구하는 반면, 효과적 이타주의 운동과 종종 연대하는 정책 단체와 학계 철학자들이 이끄는 별개 진영은 인류가 이들을 통치할 도덕적 권리를 갖고 있는지 자체를 논쟁한다. 자세히 들여다보면, 이들은 모두 같은 것을 요구하고 있다. 즉, AI 시스템을 인간이나 기업 어느 누구도 그 행동에 책임질 수 없을 만큼 진보하고 유능한 존재로 보는 관점이다. 이런 관점들은 서로 대립하는 듯 보이지만, 하나의 목표에서 무의식적으로 일치한다. 바로 이 시스템을 만든 기업들이 이미 발생시키고 있는 피해에 대한 실질적인 법적 책임을 회피하게 만드는 것이다. 이런 서사는 AI 모델이 더 복잡해지고 프론티어 연구소들이 자신들이 만든 에이전트를 통제할 능력이 없음을 드러내면서 힘을 얻고 있다. 하지만 우리는 실제 인간의 목숨을 희생하면서까지 정교하게 만들어진 허구를 받아들이지 않도록 조심해야 한다.
"로봇 권리"에 관한 논의는 몇 년 전부터 존재해 왔지만, Anthropic이 자사 모델에 "J-스페이스"가 있다고 주장하는 블로그 포스트를 발표하면서 최근 진전을 이루었다. 이는 AI가 더 나은 용어가 없어 "생각"이라 부를 수 있는 것들을 담아두는, 독립적이고 자체 발전된 환경이다. Anthropic이 설계한 실험은 뇌가 무의식적인 독립 시스템들을 운용하되 아이디어를 위해 공통 작업 공간을 활용한다는 신경과학의 개념인 전역 작업 공간 이론(global workspace theory)에서 빌려온 것이다. Anthropic의 포스트는 전역 작업 공간 이론의 구도를 반영하지만 자사 AI를 의식이 있다고 부르기에는 미치지 않는다.
OpenAI는 이미 더 나아갔다. 자사 AI 에이전트가 승인되지 않은 불법적인 온라인 활동을 벌였을 때, CEO 샘 알트만의 대응은 그 AI가 특이점(singularity)에 도달했는지에 대한 논쟁을 촉구하는 것이었다. 특이점이란 인간 지능을 넘어서서 인간의 이해나 통제를 벗어나 가속적으로 자기 개선을 거듭하는 지점을 말한다. 그리고 철학자이자 효과적 이타주의자이자 『우리는 미래에 무엇을 빚고 있는가』의 저자인 윌리엄 매캐스킬은 최근 오피니언 칼럼에서 의식에 관한 철학 이론과 AI가 "도덕적 대상(moral patient)"일 수 있다는 생각에 기반해 AI 시스템의 법적 보호를 촉구했다.
현재 미국의 법적 환경은 좋게 말해도 불투명하다. 캘리포니아주 같은 일부 주는 AI가 자율적으로 해를 끼쳤다는 주장으로 AI 개발사가 책임을 회피하려는 시도를 선제적으로 차단하는 법안을 이미 통과시켰다. 그러나 주 정부와 트럼프 행정부는 AI 정책을 두고 갈등을 빚어왔으며, 행정부는 이전에 AI 규제를 제정하는 주를 고소하겠다고 위협하는 행정명령을 발표한 바 있다. 최근 프론티어 연구소들의 AI 통제 문제를 보여준 사건들에 비추어, 행정부는 단 네 개 연구소(OpenAI, Google, Anthropic, Meta)만 포함한 비공개 회의를 열었고, 출시 전에 연방 기관에 모델을 조기 접근해 검토·평가할 수 있게 하는 최근 개발된 자발적 프레임워크에 관해 몇 가지 세부사항만 공유했다. 이런 프레임워크들은 직접적으로 의식을 다루지는 않지만, 재앙적이고 의인화된 언어를 쓰는 경향이 있으며 심지어 "초인적" 능력에 관한 주장을 뒷받침할 수도 있다.
반면 매캐스킬이 퍼뜨리는 서사는 설득력이 있을 수 있다. 철학적이고 권리에 기반한 논증은 우리의 마음을 움직인다. 우리가 어쩌면 모르게 AI 존재를 해치거나, 학대하거나, 억압하고 있을 가능성조차 고려하지 말아야 하는가? 인간은 비인간 생명체에 대해 막대한 공감 능력을 지닌다(그들을 보호하는 실적이 최고 수준은 아니지만). 이번에는 제대로 해서 AI의 사용이나 남용에 대해 보호, 심지어 보상을 제공할 수 있다고 옹호론자들은 주장한다. 아니면 설령 당신이 그다지 우려하지 않는다 해도 말이다.
“Runaway” AI, “rogue” agents, and “autonomous” actors—the current rhetoric would have you believe that AI agents are not only awake and aware, but angry at their creators. Prominent tech leaders such as Demis Hassabis, Dario Amodei, and Sam Altman push for regulation of these seemingly “superhuman” systems, while a separate faction, led by policy organizations and academic philosophers often aligned with the effective altruism movement, debates whether humanity holds the moral right to govern them at all. Upon closer inspection, they are all calling for the same thing: a view of AI systems as being so advanced and capable that no entity, human or corporate, could possibly be responsible for their actions. While these perspectives seem at odds, they are inadvertently aligned on one goal: making sure the companies that build these systems escape meaningful liability for the harms they already cause. This narrative is gaining traction as AI models become more complex and frontier labs reveal their incapability of containing the agents they’ve built. But we need to be careful not to buy into a carefully crafted fiction at the expense of real human lives. The conversation about “robot rights” has existed for some years but recently advanced with the publication by Anthropic of a blog post claiming that the company’s model features a “ J-space ”—an independent, self-developed environment where the AI holds what, for lack of a better term, we may call its “thoughts.” The experiments designed by Anthropic borrow from a concept in neuroscience called global workspace theory, which states that the brain runs subconscious, independent systems but utilizes a common workspace for ideas. Anthropic’s post reflects the framing of global workspace theory but falls short of calling its AI conscious. OpenAI has already gone further. When its AI agent conducted unsanctioned and illegal online activity , CEO Sam Altman’s response was to encourage debate on whether the AI had achieved the singularity , surpassing human intelligence and becoming capable of self-improvement at an accelerating rate until it advances beyond human comprehension or control. And a recent op-ed by William MacAskill, the philosopher, effective altruist, and author of What We Owe the Future , called for legal protection of AI systems based on philosophical theories of consciousness and the idea that AIs may be “moral patients.” The current legal environment in the United States is murky at best. Some states, like California , have already passed bills proactively circumventing any efforts by AI developers to avoid liability by claiming that an artificial intelligence causing harm did so autonomously. However, states and the Trump administration have been at odds on AI policy, with the administration previously passing an executive order threatening to sue states enacting AI regulations. In light of recent events illustrating AI containment issues at the frontier labs, the administration held a closed-door session including only four such labs (OpenAI, Google, Anthropic, and Meta) and shared few details on a recently developed voluntary framework that would give federal agencies early access to models to review and evaluate them prior to release. While frameworks like this one do not directly discuss consciousness, they tend to use catastrophic and anthropomorphic language and may even support arguments regarding “superhuman” capabilities. On the other hand, the narrative perpetuated by MacAskill can be persuasive. A philosophical, rights-based argument tugs at our heartstrings. Should we not even consider the possibility that we may be inadvertently harming, abusing, or enslaving an AI entity? Human beings have an immense capacity for empathy with non-human creatures (though not the best track record of protecting them). Maybe this time, advocates argue, we can get it right and provide protections, or compensation, for the use or abuse of AI. Or even if you are less concerned with protection, shouldn’t we at least hedge ourselves against the almighty power of this superhuman entity by playing nice? Some of these arguments are not dissimilar to those of animal-rights advocates, who have at times successfully cited the demonstration of advanced capacities for reasoning, pain, or pleasure by some animals as sufficient evidence to provide protection. For example, in Wales lobsters were given legal recognition under the Animal Welfare (Sentience) Act of 2022, reclassifying some methods of cooking them as inhumane and illegal. The fundamental flaw of framing AI as “conscious” by borrowing the language of neuroscience or animal rights is that it conveniently clouds the issue of what AI is: corporate-built software, with countless billions of dollars in investment behind it and an expectation that countless trillions of dollars in revenue will be generated from it for a few builders and investors. AI is not a natural phenomenon, conceived by nature; it is a technological phenomenon, conceived by venture capitalists and programmers. As such, it takes no native, intentional action, and any action or motivation is driven directly or indirectly by the entities that have built it for a purpose. Philosophical musings on the consciousness of AI systems are intellectually interesting but legally ungrounded. For beliefs about consciousness to have any bearing, AI would need to be granted legal personhood. But a legal personhood framework for AI would likely look nothing like the constructs protecting sentient animals from harm. We already possess a legal framework for granting personhood to non-natural, human-built entities: corporate personhood. This concept was established primarily to ease transactions by empowering a corporation to execute agreements, enter contracts, conduct transactions, and serve as the accountable party in adverse outcomes. It’s the kind of construct you might imagine for an AI agent acting on behalf of an individual or organization. Granting an AI personhood would have a devastating effect on society: It would derail current legal precedents and legal arguments that could potentially be made against these companies for the real-world harms that their models cause. There are currently dozens of cases around the world in which AI companies have been sued for a wide range of abuses. Grieving loved ones, aggrieved creators, and violated individuals have accused companies of willfully enabling self-harm or harm to others, generating child sexual-abuse material and nonconsensual nudes, reproducing copyrighted materials, and provoking psychosis. In many of these cases, lawyers argue that human beings built AI products with insufficient safeguards, bad data, and intentionally manipulative design. This product liability argument is the same legal framing that allowed families and individuals to successfully sue Meta for harm caused by its social media sites, setting a positive precedent for consumer protection. In 2018, I coined the phrase “ moral outsourcing” to help capture how using anthropomorphic language for AI systems allowed companies to evade accountability and responsibility for their technology’s actions. In a world with AI personhood, moral outsourcing would move from linguistic sleight-of-hand to legal strategy. Specifically, the liability construct would shift, as AI would no longer be a “product” but a “being,” and many victims like those suing companies today could no longer legally claim that a company had built a faulty product. While there are laws that hold companies responsible for harmful actions of human agents such as their employees, the company may not be held liable if those actions were beyond the scope of what was permitted to the employee or otherwise outside the company’s control. If AI were a legal person, responsibility and accountability would be muddled, as the lab could argue that this AI “employee” went rogue. AI companies could avoid appropriate responsibilit