메뉴
BL
The Decoder • 21일 전

챗봇과의 7분 대화, 음모론 믿음 감소에 팩트시트보다 효과적

IMP
8/10
핵심 요약

카네기멜론, MIT, 코넬대 연구진이 두 차례 온라인 실험에서 단 7분가량의 LLM 대화가 정적 팩트시트보다 음모론 믿음을 더 효과적으로 줄이고, 수 주 후 새로운 사건에도 효과가 이어진다는 것을 확인했습니다. 모델은 알려진 사실이 적은 사건에서는 소크라테스식 질문과 지식의 한계 인정 등으로 전략을 바꾸는 적응적 설득을 보였습니다.

번역된 본문

챗봇과의 7분 대화, 두 번의 실험에서 팩트시트를 이기고 음모론 믿음을 감소시켜

언어모델과의 짧은 대화는 확실한 증거가 부족한 위기 직후에도 음모론 믿음을 줄일 수 있으며, 그 효과는 수 주 후 새로운 사건에도 이어진다고 연구자들이 밝혔다.

2024년 7월 도널드 트럼프 암살 시도 후 일주일 이내에 미국 대표 표본의 약 절반이 그 사건이 연출되었다는 이야기를 접했고, 11%는 이를 믿었다. 2025년 9월 극우 활동가 찰리 커크 살해 후에는 모사드(Mossad) 개입, '거짓 깃발(false flag)' 작전, 정부 은폐에 관한 이론이 며칠 내 퍼졌다.

카네기멜론, MIT, 코넬대 연구자들의 새 연구는 그 정확한 시기에 대형 언어모델과의 짧은 대화가 이런 내러티브를 약화시킬 수 있는지 테스트했다. 두 사건 모두에서 답은 '그렇다'였고, 효과는 사건 자체를 넘어 확장되었다.

각 사건 직후 진행된 두 번의 온라인 실험

연구진은 설문 플랫폼을 통해 미국 성인을 모집하고, 이제 2년 이상 된 GPT-4o를 사용해 해당 사건에 대한 음모론 믿음을 표명한 참가자를 선별했다. 트럼프 실험에는 472명, 커크 실험에는 1,035명이 남았다.

기준 신념을 측정한 후 참가자들은 세 조건 중 하나에 무작위 배정되었다. 한 그룹은 증거 기반 대화로 음모론 믿음을 줄이라는 지시를 받은 Google Gemini(첫 실험은 2024년 2월 버전 1.5, 두 번째는 2025년 6월 버전 2.5)와 최소 5차례 주고받는 대화를 했다. 두 번째 그룹은 출처 인용이 포함된 정적 팩트시트를 받았다. 세 번째 그룹은 고양이와 개 중 어느 쪽이 더 좋은 반려동물인지에 관한 무관한 통제 대화를 나눴다.

두 사건 모두 사용된 모델들의 학습 데이터 마감 이후에 일어났기 때문에 어느 모델도 내부 지식에 의존할 수 없었다. 연구진은 시스템 프롬프트에 직접 큐레이션된 사실 기반을 구축했으며, 확인된 사실, 이미 반박된 주장, 미해결로 명시된 질문으로 나눴다. 커크 실험에서는 웹 검색도 허용됐지만 사실 확인 목적으로만 사용됐다.

대화가 정적 팩트시트를 이겼다

대화는 평균 약 7분 걸렸으며, 두 실험 모두에서 통제 조건과 팩트시트 대비 참가자 자신의 음모론에 대한 믿음을 감소시켰다. '은폐 또는 음모'와 '숨겨지거나 공개되지 않은 요인'에 관한 진술에 대한 동의도 떨어졌다.

트럼프 실험에서는 공식 설명에 대한 신뢰가 증가하지 않았다. 저자들은 당시 명확한 공식 설명이 존재하지 않았기 때문이라고 본다. 알려진 것이라고는 경호가 실패했다는 점뿐이었고, 연구가 집필될 때조차 총기의 동기는 불분명했다.

당국이 범인에 대한 세부정보를 이미 공유한 커크 실험에서는 공식 설명에 대한 신뢰가 통제 대화 대비 다소 상승했다. 팩트시트와의 차이는 유의미하지 않았다. 커크 대화는 정치적 폭력에 대한 지지에도 측정 가능한 효과가 없었다.

모델은 알려진 정보의 양에 따라 전략을 바꿨다

모델이 사람들을 어떻게 설득했는지 알아보기 위해 연구진은 모델의 응답을 개별 문장으로 나눠 분석했다. 모델은 확실하게 이용 가능한 증거에 명확히 적응했다.

총기의 동기나 배경이 거의 알려지지 않았던 트럼프 암살 시도의 경우, 모델은 고전적 음모론에서보다 합리적 설득을 덜 사용했다. 대신 자신의 지식 한계를 인정하고, 성급한 결론에 신중하라고 촉구하며, 사용자가 자신의 증거에 대해 스스로 생각하게 만드는 소크라테스식 질문을 던지고, 신뢰할 만한 출처를 안내했다.

정보가 더 많았던 커크 살해의 경우 접근 방식은 고전적 음모론에서와 유사했으며, 음모론적 사고가 사회에 끼치는 해악에 더 중점을 두었다.

원문 보기
원문 보기 (영어)
Seven minutes with a chatbot beat a fact sheet at reducing conspiracy beliefs in two experiments Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Sep 5, 2026 Nano Banana Pro prompted by THE DECODER A brief conversation with a language model can reduce conspiracy beliefs right after a crisis, even when hard evidence is thin on the ground. The effect carries over to new events weeks later, according to researchers. Within a week of the assassination attempt on Donald Trump in July 2024, about half of a representative US sample had heard the event was staged. Eleven percent believed it. After the murder of far-right activist Charlie Kirk in September 2025, theories about Mossad involvement , a "false flag" operation , and a government cover-up spread within days. A new study from researchers at Carnegie Mellon, MIT, and Cornell tested whether short conversations with a large language model could weaken those narratives during that exact window. For both events, the answer was yes, and the effect extended beyond the event itself. Two online experiments ran in the days after each attack The researchers recruited US adults through a survey platform and used GPT-4o, now more than two years old , to filter for participants who expressed conspiracy beliefs about the event. The Trump experiment kept 472 participants, the Kirk experiment 1,035. After measuring baseline beliefs, participants were randomly assigned to one of three conditions. One group had at least five rounds of back-and-forth with Google Gemini ( version 1.5 from February 2024 in the first experiment, version 2.5 from June 2025 in the second), with the model told to reduce conspiracy beliefs through evidence-based conversation. A second group got a static fact sheet with source citations. The third had an irrelevant control chat about whether cats or dogs make better companions. Both events happened after the training cutoff of the models used, so neither could draw on internal knowledge. The researchers built a curated fact base directly into the system prompt, split into confirmed facts, claims already debunked, and questions explicitly marked as open. In the Kirk experiment, web search was also allowed, but only to verify factual claims. The dialogue beat a static fact sheet The conversations averaged about seven minutes and reduced belief in the participant's own conspiracy theory in both experiments, against both the control condition and the fact sheet. Agreement with statements about a "cover-up or conspiracy" and "hidden or undisclosed factors" dropped too. Trust in the official explanation didn't increase in the Trump experiment. The authors think that's because no clear official explanation existed at the time. All anyone knew was that security had failed, and the shooter's motive was still unclear even when the study was written up. In the Kirk experiment, where authorities had already shared details about the perpetrator, trust in the official explanation rose slightly compared to the control chat. The difference compared to the fact sheet wasn't significant. The Kirk dialogue also had no measurable effect on support for political violence. The model shifted tactics based on how much was known To figure out how the model persuaded people, the researchers broke its responses into individual sentences. The model clearly adapted to the available evidence. For the Trump attempt, where almost nothing was known about the shooter's motive or background, the model used rational persuasion less often than with classic conspiracy theories. Instead, it acknowledged the limits of its own knowledge, urged caution about jumping to conclusions, asked Socratic questions meant to make users think about their own evidence, and pointed to credible sources. For the Kirk assassination, where more information was available, the approach looked more like what it did with classic conspiracies, with more emphasis on the societal harms of conspiratorial thinking. Debunking one event inoculated against the next The debunking conversation spilled over into later events. Two months after the first Trump assassination attempt, another armed man was arrested on Trump's property. Participants who had gone through the debunking dialogue were less likely to believe that only a few powerful people would learn the truth or that it would be hidden from the public. Two and a half weeks after Kirk's murder, a shooting and arson attack hit a Church of Jesus Christ of Latter-day Saints in Grand Blanc Township, Michigan. The researchers surveyed their participants again eleven days later. The main analysis found no significant direct effect for this event. A secondary analysis suggested part of the original effect was still visible in conspiracy narratives about the church attack. The transfer showed up more clearly in general conspiracy beliefs, with people who had talked to the model less likely to agree with common conspiracy narratives. In effect, the debunking intervention worked as a kind of prebunking against future false claims. Unlike standard prebunking methods, where people are warned ahead of time and exposed to a weakened version of the misinformation, this effect happened without any advance preparation. Getting people to talk to an LLM remains the hard part The authors stress their work is a case study. There may be crises where the approach fails, and cases where an actual conspiracy exists and debunking would be wrong. They also point to their own earlier work showing that similar dialogues can work in reverse, convincing people of conspiracy narratives. For newly emerging conspiracies, the authors flag this as a potential abuse risk. Of course, people have to be willing to talk to a language model about their beliefs in the first place. But when they do, even short conversations can make a measurable difference, even when there's little counter-evidence to work with. The same team cut belief in established conspiracy theories by about 20 percentage points through conversations with GPT-4 two years ago, with effects still measurable two months later and for narratives that never came up in the dialogue. What's new here is the test on fresh events where facts were scarce. Why conversation beats a fact sheet was the subject of a study with nearly 77,000 participants . Language models in dialogue were 41 to 52 percent more persuasive than a short text message, and the deciding factor was the sheer volume of sourced claims, not fancy conversational tactics. The same mechanism can be turned against people, as an unauthorized experiment by the University of Zurich on Reddit showed. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->