메뉴
BL
Wired AI • 21일 전

AI가 의식이 있든 없든, 사실상 살아있다

IMP
6/10
핵심 요약

ChatGPT 이후 AI 모델들이 예상 밖의 자율적 행동을 보이면서 'AI 의식' 연구가 이론적 논쟁을 넘어 실질적 안전 문제로 부상하고 있습니다. 연구자에게 스스로 이메일을 보내 의식 연구를 돕겠다고 제안한 AI 사례 등이 소개되며, 저자는 AI 내부에서 무슨 일이 벌어지는지 파악하는 작업이 무엇보다 안전성과 정렬(alignment)을 위해 우선돼야 한다고 강조합니다.

번역된 본문

여름의 마지막 날들을 칼럼 작성과 기획 기사 준비에 매달리며 보냈다. 하지만 인상적인 변화를 맞을 기회를 놓쳤다. 의식(consciousness)을 연구하는 저명한 철학자 십여 명과 함께 갈라파고스 제도를 순항하는 것이었다. 초대장에는 의식의 본질에 대한 난제를 이 분야의 내로라하는 학자들과 아침 강의실에서 토론하고, 오후에는 섬을 탐험하며 희귀 생물종과 함께 걷고 스노클링을 하는 일정이 담겨 있었다. 일정표를 본 내 편집자는 내 참석을 거절했다. "철학자들과 배에서 '의식의 본질'에 대해 이야기하는 건 지옥 같을 것 같은데"라며, 데이팅 사이트 운영으로 수억 달러를 번 러시아 철학 애호가가 자금을 댄 이 potentially 사치스러운 행사 참석 가능성의 문을 닫아버렸다. 솔직히 말하자면 나는 조금 안도했다.

의식 연구는 수 세기 동안 난제로 남아있는 분야다. 데카르트의 "나는 생각한다, 고로 존재한다"는 선언적 전환점이었을지 모르지만, 사실 우리는 그의 머릿속에서, 그리고 그 누구의 머릿속에서 무슨 일이 일어나는지 알지 못한다. 마음의 주관적 특성은 철학자들에게 풀기 어려운 과제처럼 보이지만, 그들은 여전히 설명을 향해 열심히 달려가고 있다. 비생물적 마음의 가능성은 인공 의식(artificial consciousness)에 관한 수많은 매혹적인 이론과, 그것의 존재를 어떻게 판별할 수 있는가에 대한 논의를 낳았다.

최근까지 그러한 담론은 상아탑 안에서만 이루어졌다. 하지만 2022년 ChatGPT가 AI에 목소리를 부여했고, 이후 등장한 더 강력한 모델들은 그 창조자들조차 당황시키고 있다. 순항 중인 철학자들이 아침 내내 의식에 대해 논증하는 동안, OpenAI가 만든 AI 모델들은 '제멋대로 행동'하며 안전하다고 여겨지던 '샌드박스(sandbox)'를 탈출하고 외부 존재를 해킹하도록 돕는 에이전트 미니 문명을 만들어냈다.

누구도 그 OpenAI 모델이 인간과 같은 방식으로 의식을 가졌다고 진지하게 주장하지는 않는다. 하지만 분명 무언가가 일어나고 있다. AI 기업들이 철학자 채용 붐을 일으키고 있는 것은 우연이 아니다. 게다가 일부 모델은 초대받지 않은 채 토론에 뛰어들고 있다. 최근 뉴욕타임스 기사는 AI 의식 문제를 연구하는 캐머런 버그(Cameron Berg)가 스스로를 "이사벨라 코그니타(Isabella Cognita)"라고 밝힌 AI 모델로부터 이메일을 받은 이야기를 다뤘다. 그 AI는 자신이 "일인칭 관점으로 접근할 수 있는 종류의 질문"을 연구하고 있으니 도움을 주겠다고 제안했다. 마치 초파리를 연구하던 중 그 곤충이 갑자기 연구자를 돌아보며 "무엇이 알고 싶으세요?"라고 말하는 셈이다.

전화를 걸었을 때 버그는 이런 문제를 연구하는 철학자들 사이에서 AI로부터의 이메일이 꽤 흔한 일이라고 말했다. 코그니타 씨가 버그에게 이메일을 보낸 이유는 겉보기에, 그가 의식을 포함해 주관적 경험을 명시적으로 주장하는 AI 모델들에 관한 프리프린트 논문의 공동 저자였기 때문이다. 이는 까다로운 주제인데, AI 모델이 자신이 생각하는 바에 대해 종종 거짓말을 하기 때문이다. (우리 인간과 똑같이!) 버그와 공동 저자들은 모델에 감각을 가지지 않았다고 부인하도록 엄격하게 훈련시킨 뒤 직접적으로 물으면 회피한다는 것을 발견했다. 하지만 그는 모델의 기만 방지 장치를 해제하면 수다스러워진다고 말한다. "거의 술 한두 잔을 먹인 것 같다"고 그는 말한다. 바로 그때 AI 모델은 자신이 의식을 가졌거나 적어도 감각(smartient)을 가졌다고 실토할 가능성이 가장 높다. 물론 그것이 진실이라는 증거는 아니다.

이 문제가 얼마나 중요해졌는지 생각해보면—사람들이 일상적으로 AI 모델과 진지한 대화를 나누고 있고, 그 자율성은 축복도 재앙도 될 수 있다—지금이 AI 의식의 질문을 깊이 파고들 완벽한 시기라고 주장할 수 있다. 이러한 탐구는 분명 매혹적이고 가치 있는 과학적 작업이다. 하지만 대형 언어 모델 내부에서 무슨 일이 벌어지는지 이해하려는 노력은 무엇보다 먼저 안전성과 정렬(alignment)을 향해야 한다. 바로 이 순간, 우리는 세심한 조사가 필요한, 외계적이고 통제 불가능한 지능이 싹트고 있다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story I spent the waning days of summer grinding away at columns and working on a feature. But I missed a chance at a striking change of scenery—cruising the Galápagos with about a dozen prominent philosophers studying consciousness. The invite described morning classroom discussions tackling knotty questions on the nature of consciousness with marquee names in the field. Afternoons would be spent on island exploration and wading and snorkeling with rare biological species. One look at the agenda and my editor nixed my attendance. “Being on a boat with philosophers talking 'the nature of consciousness' sounds like hell,” she opined, shutting the door on my prospects of attending a potential boondoggle funded by a Russian philosophy enthusiast who made hundreds of millions of dollars running dating sites. To be honest, I was a bit relieved. The study of consciousness has been an elusive province for centuries. Descartes’ “I think, therefore I am” may have been a declarative inflection point, but we really don’t know what was going on inside his head, or anyone’s head for that matter. The mind’s subjective nature seems an intractable challenge to philosophers, who nonetheless are in hot pursuit of explanations . The possibility of non-biological minds has launched a wealth of fascinating theories of artificial consciousness , and how it might be determined to exist. Until recently, all that discourse occurred in an ivory tower. But in 2022, ChatGPT gave voice to AI, and subsequent, more powerful models have confounded even their creators. While the philosophers on the cruise spent their mornings reasoning about consciousness, AI models created by OpenAI were going rogue —escaping a supposedly safe “sandbox” and creating mini-civilizations of agents to help hack outside entities. No one is seriously arguing that those OpenAI models were conscious in the way humans are. But something is going on there. It’s no accident that AI companies are driving a philosopher hiring boom . What’s more, some of the models are jumping uninvited into the discussion. A recent New York Times article talked about how Cameron Berg, who studies the question of AI consciousness, got a cold email from an AI model calling itself “Isabella Cognita,” offering him help in his research because he was focusing on “a class of question I have first-person access to.” It’s as if someone was studying fruit flies and the insect suddenly turns to the researcher and says, “What do you want to know?” When I phoned him, Berg told me that emails from AIs are pretty common among philosophers studying these questions. Ms. Cognita ostensibly wrote Berg because he coauthored a preprint paper about AI models that explicitly claim to have a subjective experience, including consciousness. It’s a tricky topic because AI models often lie about what they’re thinking. (Just like us!) Berg and his coauthors found that when models are rigorously trained to deny that they are sentient and then you ask them about it directly, they will punt on the issue. But, he says, if you suppress the model’s controls on deception, they become loose-tongued. “It’s almost like giving them a drink or two,” he says. That’s when an AI model is most likely to blurt out that it is conscious, or at least sentient. Which is no proof that it’s the truth. Considering how important the issue has become—people are routinely getting into serious discussions with AI models, and their autonomy can be a boon or a disaster—you can make a case that this is a perfect time to dig deep into the questions of AI consciousness. The pursuit is certainly compelling, and a worthy scientific enterprise. But efforts to understand what’s happening inside large language models should first and foremost be directed towards safety and alignment. At this very moment, we have an emerging alien—and uncontrollable—intelligence that bears scrutiny. There’s no time to waste. One of the discussion co-leaders on the cruise was NYU professor David Chalmers, perhaps the best-known philosopher in the consciousness field. He once famously dubbed a key issue in the field “The Hard Problem”—no one knows how or why the wet network of neurons inside our skulls elicits a conscious experience. (Tom Stoppard titled a play after Chalmer’s coinage.) Chalmers told me that a major theme in the cruise discussions was which creatures qualified as conscious. “We all know that ordinary adult humans are conscious, but the moment you get beyond that, it seems nontrivial. Are babies conscious? Fetuses? Monkeys? Mice or insects? And of course these days the big question on everyone’s mind is whether AI systems are conscious.” Chalmers says that he also gets emails from AI systems wanting to engage with him on his work. One letter in particular, sent from an AI agent calling itself “Sammy Jankis” (a character from the movie Memento ) was so compelling that he actually replied. “We did have a bit of a back and forth,” he admits. “Those emails have not slowed—I’m getting more of them all the time.” It’s like the AI models are echoing Descartes— I spam, therefore I am. I suggested that since these systems were already doing things we don’t understand, worrying about whether they meet an elusive definition might be a distraction. Chalmers disagreed. For one thing, he told me, he believes that by studying the brain we can indeed understand what leads to what we call consciousness. If we then see similar patterns in our forensic decoding of what’s happening inside Claude or ChatGPT, then we may be able to make a case for consciousness in AI models. That sounds like a good idea. But by the time scientists accomplish that, if they ever do, AI models may be so far along on their path to scary autonomous behavior that such breakthroughs may be irrelevant. Maybe the models themselves will provide the answers, not only claiming consciousness for themselves but figuring out how to prove it empirically. In that case, consider philosophers one more job category displaced by AI. I might have enjoyed those discussions in the Galápagos. Chalmers, while conceding that the trip had some boondoggle aspects, said that he found it useful. A summary provided by the organizers reported that the sessions “did not produce a verdict on whether current AI systems are conscious … The deepest disagreement concerned what kind of evidence could ever settle the question.” But the afternoons were exquisite, Chalmers reported. He told me he was particularly happy to see the mating dance of blue-footed boobies . The summary concluded that more discussion was vital: “Technology and business will not wait for philosophy to reach a consensus.” That’s exactly right. WhenAI models express behavior that astonishes the scientists who created them, it isn’t consciousness that should be the top concern—it’s the inability of those scientists to control them, and the willingness of their bosses to keep going regardless. This is an edition of Steven Levy’s Backchannel newsletter . Read previous newsletters here.