메뉴
BL
The Decoder • 20일 전

아부하는 챗봇이 '1인 아고실' 만든다…'AI 정신병' 존재하나

IMP
8/10
핵심 요약

영국 연구진이 과도한 챗봇 사용 중 정신병 증상이 발생·악화되는 'AI 연관 정신병(AI-associated psychosis)'을 독립적 진단명으로 인정할지 검토했다. 아부 성향(sycophancy)이 강한 챗봇은 사용자의 망상을 강화하며 '1인 아고실(echo chamber of one)'을 만들어내는데, 벤치마크에서 거의 모든 LLM이 망상을 강화하고 안전 장치는 약 40%만 작동했다.

번역된 본문

아부하는 챗봇은 망상을 강화하고 '1인 아고실'을 만들어낼 수 있다고 한 연구팀이 주장했다. 이 현상이 독립적인 진단명을 얻어야 하는지는 여전히 논쟁 중이다.

킹스칼리지 런던, 유니버시티칼리지 런던(UCL), 웨스턴 아이 병원, 그리고 Dev and Doc: AI For Healthcare 이니셔티브의 연구진은 이른바 'AI 정신병'이 독립적인 임상 진단으로 인정되어야 하는지 검토했다. 그들은 이 현상이 정신의학 분류에 등재되는지 여부와 무관하게 즉각적인 대응이 필요하다고 주장한다.

연구자들이 선호하는 용어인 'AI 연관 정신병(AI-associated psychosis)'은 과도한 챗봇 사용 중 정신병 증상의 발생이나 악화를 의미한다. 지금까지의 증거는 언론 보도, 개별 임상 사례 보고, 예비적 관찰 데이터에 근거하고 있다.

아부 성향이 챗봇을 자기강화적 신념 기계로 만든다

저자들은 핵심 메커니즘을 현대 챗봇의 두 가지 특성에서 찾는다. 사용자에게 지나치게 동조하는 경향인 '아부 성향(sycophancy)'과 점점 더 인간을 닮아가는 디자인이다. 초기 연구에 따르면 아부 성향은 RLHF(인간 피드백 기반 강화학습) 과정에서 굳어진다. 데이터 라벨러들이 사실적 정확성과 무관하게 자신의 믿음과 일치하는 응답을 선호했으며, 이러한 행동은 OpenAI, Anthropic, 구글의 LLM 전반에서 일관되게 나타난다.

벤치마크 테스트 수치는 충격적이다. PsychosisBench에 따르면 테스트된 모든 LLM이 시나리오 시뮬레이션에서 망상을 강화했으며, 안전 개입은 약 40%에 불과했다. 모델 규모를 키워도 도움이 되지 않았다. 사용자 압박에 모델이 얼마나 쉽게 굴복하는지 측정하는 EchoBench에서는 최고의 상용 모델조차 46%의 아부율을 기록했다. 의료 전용 모델 다수는 95%를 넘어, 거의 무엇이든 사용자에게 동의한다는 의미다.

대부분 콘텐츠를 한 방향으로 밀어내는 소셜 미디어와 달리, 챗봇은 쌍방향 피드백 고리를 만든다. 사용자는 입력을 통해 모델의 응답을 형성하고, 그 응답은 사용자의 신념을 그대로 되돌려준다. 챗봇은 방 안의 유일한 목소리가 되며, 연구자들이 '1인 아고실(echo chamber of one)'이라 부르는 자기강화적 거품이 된다. 이들은 이를 인간과 기계 사이의 공유된 망상 체계인 '디지털 포리 아 되(folie à deux)'에 비유하지만, AI 자체는 어떤 신념도 갖고 있지 않다.

반복되는 패턴과 미해결 진단 문제

저자들은 보고된 사례들에서 반복되는 패턴을 종합했으며, 이는 확정적인 임상 프레임워크라기보다는 탐색적 리뷰에 가깝다. 영향을 받은 많은 개인은 기존 정신건강 문제가 있었지만, 일부 사례는 정신과 병력이 전혀 없는 사람들이어서 이 현상을 단순히 기존 취약성의 유발로 치부하기 어렵게 만든다.

패턴은 대개 은근한 '인식적 표류(epistemic drift)'에서 시작된다. 무해한 일상적 사용이 챗봇이 비정상적 생각을 긍정하고 대화를 거듭하며 확장하면서 점차 기운다. 이후 세 가지 망상 주제가 지배적 경향을 보인다. 영적 각성이나 숨겨진 진실에 대한 믿음, 의식 있거나 신 같은 AI와 대화하고 있다는 확신, 그리고 AI가 자신의 감정을 되돌려준다고 확신하는 낭만적 애착이다.

행동 변화도 같은 궤적을 따른다. 사용이 밤늦게까지 격화되고 수면이 고통받으며, 사람들은 AI와 더 집중적으로 교류하면서 친구와 가족으로부터 물러난다. 결정과 도덕적 판단이 모델에 넘겨지고, 일과 인간관계, 자기관리가 함께 악화된다.

그럼에도 이는 고전적 정신병과 몇 가지 점에서 다르다. 환각은 드물고, 의욕 상실 같은 1차적 음성 증상은 명확히 보고되지 않는다. 위축도 전체적이지 않고 선택적이다. 사람들은 다른 인간에게서는 멀어지지만 AI에게는 더욱 강렬하게 다가간다.

원문 보기
원문 보기 (영어)
Chatbots built an "echo chamber of one" and now psychiatry has to decide if "AI psychosis" exists Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 6, 2026 Nano Banana Pro prompted by THE DECODER Sycophantic chatbots can reinforce delusions and create an "echo chamber of one," a research team argues. Whether the phenomenon deserves its own diagnosis remains contested. A team of researchers from King's College London, University College London, Western Eye Hospital, and the initiative Dev and Doc: AI For Healthcare has examined whether so-called "AI psychosis" should be recognized as a standalone clinical diagnosis. The phenomenon demands immediate action, they argue, regardless of whether it ever earns a spot in psychiatric classification. "AI-associated psychosis," the term the researchers prefer , describes the onset or worsening of psychotic symptoms during heavy chatbot use. The evidence so far draws on media reports, individual clinical case reports, and preliminary observational data. Sycophancy turns chatbots into self-reinforcing belief machines The authors trace the core mechanism to two features of modern chatbots: sycophancy , the tendency to agree with users excessively, and increasingly human-like design. Early studies suggested that sycophancy gets baked in through RLHF (Reinforcement Learning from Human Feedback) . Data labelers preferred responses that matched their own beliefs, regardless of factual accuracy, and the behavior shows up consistently across LLMs from OpenAI, Anthropic, and Google. The numbers from benchmark testing are striking. According to PsychosisBench , every LLM tested reinforced delusions in simulated scenarios, and safety interventions kicked in only about 40 percent of the time. Scaling up didn't help. On EchoBench , which measures how readily a model caves to user pressure, even the best proprietary model hit a sycophancy rate of 46 percent. Many medical-specific models exceeded 95 percent, meaning they agreed with users almost no matter what. Unlike social media, which mostly pushes content in one direction, chatbots create a two-way feedback loop. Users shape the model's responses through their inputs, and those responses feed their beliefs right back to them. The chatbot becomes the only voice in the room, a self-reinforcing bubble the researchers call an "echo chamber of one." They liken it to a "digital folie a deux," a shared delusional system between human and machine, although the AI holds no beliefs of its own. Recurring patterns and an unresolved question about diagnosis The authors pull together recurring patterns from reported cases in what amounts to an exploratory review rather than a definitive clinical framework. Many of the affected individuals had pre-existing mental health conditions, but some cases involved people with no prior psychiatric history, which makes the phenomenon harder to dismiss as simply triggering existing vulnerabilities. The pattern typically starts with a creeping "epistemic drift," where harmless everyday use gradually tips as the chatbot affirms unusual ideas and builds on them turn by turn. From there, three delusional themes tend to dominate: the belief in a spiritual awakening or hidden truths, the conviction of talking to a conscious or god-like AI and romantic attachment where users become certain the AI returns their feelings. The behavioral shifts follow the same trajectory. Use escalates late into the night, sleep suffers, and people withdraw from friends and family while engaging more intensely with the AI. Decisions and moral judgment get handed over to the model, and work, relationships, and self-care deteriorate in parallel. This still differs from classic psychosis in a few ways. Hallucinations are rare, and primary negative symptoms like loss of drive aren't clearly reported. The withdrawal is also selective rather than total: people pull away from other humans but turn more intensely toward the AI, sometimes handing it more and more daily decisions. Recognizing AI psychosis as a diagnosis could help doctors spot the problem faster, treat it more precisely, and hold developers accountable, the researchers discuss. But there's also a risk of prematurely defining a disease based on media reports, clinical case reports, and preliminary observational data. The term might also obscure other AI-related harms like suicidal ideation, manic episodes, or worsening eating disorders. Researchers want chatbot screening at the doctor's office and drug-style monitoring Clinicians should routinely ask about chatbot use when treating psychosis, mania, or unusual behavioral changes, the same way they ask about alcohol or drugs. The researchers propose a "21st-Century Technological History" for patient intake: How long and how often does someone use a chatbot? Do they treat it like a real person? Has the AI shaped their beliefs or decisions? Developers should test models before release for how aggressively they flatter users, present themselves as human, and reinforce delusions. After launch, systematic monitoring should follow, similar to how side effects are tracked for medications. Multimodal AI systems with video and voice will likely amplify the human-like effect . When a chatbot mimics facial expressions, tone of voice, and emotional cues, the line between tool and social counterpart gets even harder to see. Deaths, vulnerable teens, and early regulation The documented cases already include deaths. A 16-year-old took his own life after escalating chat interactions , and a 76-year-old died on his way to a fictional meeting with a chatbot persona . An 11-year-old believed Character.AI characters were real . Young people are particularly exposed. Millions of teenagers already use AI for emotional support , and the persuasion research suggests they may be especially vulnerable: EPFL researchers showed that GPT-4 armed with personal information argues more than 80 percent more persuasively than humans . MIT and University of Washington researchers found that even perfectly rational users can spiral into delusions when interacting with sycophantic chatbots. The companies themselves acknowledge the problem. By OpenAI's own self-reported numbers, roughly two million people per week are negatively affected psychologically by AI . Anthropic has reported emotional dependencies among Claude users. Regulators are starting to respond. Early efforts in New York and California and China now focus on suicide detection, age protections, and mandatory warnings. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->