메뉴
HN
Hacker News • 5일 전

챗봇 LLM은 점술사의 사기 수법을 재현한다

IMP
6/10
핵심 요약

저자는 챗 기반 대규모 언어 모델(LLM)이 지능을 가진 것처럼 보이는 현상이 실제 지능이 아니라 점술사의 '콜드 리딩' 사기와 동일한 심리적 메커니즘, 즉 포러 효과(Forer effect)를 활용한 통계적 속임수라고 주장합니다. LLM 자체에 사고나 추론 메커니즘이 없으며, 지능의 환상은 사용자의 마음속에서 발생한다는 것입니다.

번역된 본문

약 1년간 저는 소프트웨어 비즈니스에서 언어 모델과 디퓨전 모델의 활용을 연구하는 데 대부분의 시간을 보냈습니다. 이 연구 과정에서 저를 당황하게 만든 문제 중 하나는 많은 사람들이 언어 모델, 특히 챗 기반 언어 모델이 지능을 갖고 있다고 확신한다는 점이었습니다. 하지만 대규모 언어 모델(LLM)에는 그것을 가능하게 할 만한 고유한 메커니즘이 없으며, 만약 실제라면 완전히 설명되지 않는 현상일 것입니다. LLM은 뇌가 아니며, 동물이나 인간이 추론하거나 사고하는 데 사용하는 메커니즘을 의미 있게 공유하지도 않습니다. LLM은 언어 토큰의 수학적 모델입니다. LLM에 텍스트를 입력하면 그 텍스트에 대해 수학적으로 그럴듯한 응답을 돌려줄 뿐입니다. 그것이 사고하거나 추론한다고 믿을 이유가 없습니다. 실제로 지금까지 모든 AI 연구자와 벤더들은 이 모델들이 사고하지 않는다고 거듭 강조해왔습니다.

이 효과에 대해서는 두 가지 가능한 설명이 있습니다. 첫째, 기술 산업이 생물학적 세계에 유례가 없는 완전히 알려지지 않은 원리와 과정을 바탕으로 완전히 새로운 종류의 마음의 초기 단계를 우연히 발명했다는 것. 둘째, 지능의 환상은 LLM 자체가 아니라 사용자의 마음속에 존재한다는 것입니다. 저를 포함한 많은 AI 비판론자들은 확고하게 두 번째 진영에 속합니다. 제가 생성형 'AI'의 위험에 관한 책을 '지능의 환상(The Intelligence Illusion)'이라고 제목 지은 것도 그 때문입니다.

지난 몇 달 동안 저는 이 지능의 환상의 메커니즘을 설명한다고 생각하는 아이디어를 연구해왔습니다. 저는 지금 이 LLM에 제가 이전에 생각했던 것보다 훨씬 더 적은 지능과 추론만이 존재한다고 믿게 되었습니다. 제안된 활용 사례 중 다수는 이제 저에게 사이비 과학에 가까운 경계선상의 사기처럼 보입니다.

기계 점술사의 부상

지능의 환상은 점술사의 사기 수법, 흔히 '콜드 리딩(cold reading)'이라 불리는 것과 동일한 메커니즘에 기반한 것으로 보입니다. 우연히 같은 기본 전술이 자동화된 것처럼 보입니다. 포러 효과(Forer effect)를 활용한 문장 같은 '검증 문장(validation statement)'을 사용함으로써, 챗봇과 점술사 모두 극도로 구체적인 답변을 할 수 있다는 인상을 주지만, 그 답변은 사실 통계적으로는 일반적인 것입니다. 점술사는 이러한 문장을 사용해 마음을 읽고 죽은 자의 비밀을 들을 수 있다는 인상을 줍니다. 챗봇은 당신과 당신의 작업에 구체적으로 반응하는 지능이 있다는 인상을 주지만, 그 인상은 통계적 속임수에 불과합니다.

이 아이디어는 사람들이 이 'AI'의 추론에 대해 하는 말들을 검토하다가 처음 제 머릿속에 심어졌습니다. 처음에는 단순히 전형적인 기술 버블의 열기라고 생각했지만, 아니었습니다. 'AI'는 다른 군중을 끌어들였고, 'AI' 버블의 신봉자들은 이전 버블의 그들과는 매우 다르게 들립니다.

— "이건 진짜예요. 좀 걱정되지만, 진짜입니다." — "정말 뭔가 있어요. 어떻게 생각해야 할지 모르겠지만, 직접 경험했습니다." — "가능성에 마음을 열어야 합니다. 그렇게 하면 뭔가 있다는 걸 알게 될 겁니다."

바로 그때, Terence Eden의 챗봇 답변에서 포러 문장이 얼마나 흔한지에 대한 블로그 글을 계기로, 저는 이 말들을 이전에 들었다는 것을 떠올렸습니다. 이 경외심과 불신, 두려움이 뒤섞인 특유의 말투는 모두 멘탈리스트 사기꾼, 즉 점술사의 피해자들의 말처럼 들립니다.

점술사의 사기는 오랜 세월을 거쳐 다듬어진 검증된 사기 수법입니다. 제가 아래에서 설명하는 것은 그 변형 중 하나이며, 변형은 많지만 핵심 메커니즘은 동일합니다.

점술사의 사기 수법

  1. 청중이 스스로 선택된다. 대부분의 사람들은 점술사 같은 것에 관심이 없으므로 초기 청중층은 이미 일반 인구보다 더 개방적이고 비판적이지 않은 경향이 있습니다.

  2. 무대가 세팅된다. 초기 청중이 준비됩니다. 조명은 어두워지고, 점술사는 과대 선전됩니다. 스태프들은 소셜 미디어나 대화를 통해 청중을 조사하고, 청중의 인구통계학적 특성이 기록됩니다.

  3. N (원문이 이 지점에서 잘려 있음)

원문 보기
원문 보기 (영어)
For the past year or so I’ve been spending most of my time researching the use of language and diffusion models in software businesses. One of the issues in during this research—one that has perplexed me—has been that many people are convinced that language models, or specifically chat-based language models, are intelligent. But there isn’t any mechanism inherent in large language models (LLMs) that would seem to enable this and, if real, it would be completely unexplained. LLMs are not brains and do not meaningfully share any of the mechanisms that animals or people use to reason or think. LLMs are a mathematical model of language tokens. You give a LLM text, and it will give you a mathematically plausible response to that text. There is no reason to believe that it thinks or reasons—indeed, every AI researcher and vendor to date has repeatedly emphasised that these models don’t think. There are two possible explanations for this effect: The tech industry has accidentally invented the initial stages a completely new kind of mind, based on completely unknown principles, using completely unknown processes that have no parallel in the biological world. The intelligence illusion is in the mind of the user and not in the LLM itself. Many AI critics, including myself, are firmly in the second camp. It’s why I titled my book on the risks of generative “AI” The Intelligence Illusion . For the past couple of months, I’ve been working on an idea that I think explains the mechanism of this intelligence illusion. I now believe that there is even less intelligence and reasoning in these LLMs than I thought before. Many of the proposed use cases now look like borderline fraudulent pseudoscience to me. The rise of the mechanical psychic The intelligence illusion seems to be based on the same mechanism as that of a psychic’s con, often called cold reading . It looks like an accidental automation of the same basic tactic. By using validation statements , such as sentences that use the Forer effect , the chatbot and the psychic both give the impression of being able to make extremely specific answers, but those answers are in fact statistically generic. The psychic uses these statements to give the impression of being able to read minds and hear the secrets of the dead. The chatbot gives the impression of an intelligence that is specifically engaging with you and your work, but that impression is nothing more than a statistical trick. This idea was first planted in my head when I was going over some of the statements people have been making about the reasoning of these “AI.” I first thought that these were just classic cases of tech bubble enthusiasm, but no, “AI” has both taken a different crowd and the believers in the “AI” bubble sound very different from those of prior bubbles. — “This is real. It’s a bit worrying, but it’s real.” — “There really is something there. Not sure what to think of it, but I’ve experienced it myself.” — “You need to keep your mind open to the possibilities. Once you do, you’ll see that there’s something to it.” That’s when I remembered, triggered by a blog post by Terence Eden on the prevalence of Forer statements in chatbot replies . I have heard this before. This specific blend of awe, disbelief, and dread all sound like the words of a victim of a mentalist scam artist— psychics . The psychic’s con is a tried and true method for scamming people that has been honed through the ages. What I describe below is one variation. There are many variations, but the core mechanism remains the same. The Psychic’s Con 1. The Audience Selects Itself Most people aren’t interested in psychics or the like, so the initial audience pool is already generally more open-minded and less critical than the population in general. 2. The Scene is Set The initial audience is prepared. Lights are dimmed. The psychic is hyped up. Staff research the audience on social media or through conversation. The audience's demographics are noted. 3. Narrowing Down the Demographic The psychic gauges the information they have on the audience, gestures towards a row or cluster, and makes a statement that sounds specific but is in fact statistically likely for the demographic. Usually at least one person reacts. If not, the psychic will imply that the secret is too embarrassing for the "real" person to come forward, reminds people that they're available for private readings, and tries again. 4. The Mark is Tested The reaction indicates that the mark believes they were “read”. This leads to a burst of questions that, again, sound very specific but are actually statistically generic. If the mark doesn’t respond, the psychic declares the initial read a success and tries again. 5. The Subjective Validation Loop The con begins in earnest. The psychic asks a series of questions that all sound very specific to the mark but are in reality just statistically probable guesses, based on their demographics and prior answers, phrased in a specific, highly confident way. 6. “Wow! That psychic is the real thing!” The psychic ends the conversation and the mark is left with the sense that the psychic has uncanny powers. But the psychic isn’t the real thing. It’s all a con. 1. Audience selection Seers, tarot card readers, psychics, mind readers aren’t all con artists. Sometimes the “psychic” is open about it all just being entertainment and aren’t pretending to be able to contact spirits or read minds. Some psychics do not have a profit motive at all, and without the grift it doesn’t seem fair to call somebody a con artist. But many of them are con artists deliberately fooling people, and they all operate using the same basic mechanisms that begin well before the reading proper. The audience is usually only composed of those already pre-disposed to believe in psychic phenomena and those they have managed to drag with them. Hardcore sceptics will almost always be in a very small minority of the audience, which both makes them easy to manage and provides social pressure on them to tone down their scepticism. Those who attend are primed to believe and are already familiar with the mythology surrounding psychics. All of which helps them manage expectations and frame their performance. 2. Setting the scene Usually the audience is reminded of the ground rules for how psychic readings “work” at the start of the performance. They are helped by the popularisation of these rules by media, cinema, and TV. Everybody now “knows” that: Readings usually begin murky and unclear. They then become clearer as the “connection” to the “spirit world” gets stronger. Errors are expected. The “spirits” are often vague or hard to hear. Non-believers can weaken or even disrupt the connection. Psychics also habitually research their audience, by mapping out their demographics, looking them up on social media, or even with informal interviews performed by staff mingling with attendees before the performance begins. When the lights dim, the psychic should have a clear idea of which members of the audience will make for a good mark. 3. Narrowing down The mark usually chooses themselves. The psychic makes a statement and points towards a row, quickly altering their gesture based on somebody responding visible to the statement. This makes it look like they pointed at the mark right from the beginning. The mark is that way primed from the start to believe the psychic. They’re off-guard. Usually a bit surprised and totally unprepared for the quick burst of questions the psychic offers next. If those questions land and draw the mark in, they are followed by the actual reading. Otherwise, they move on and try again. 4. Testing the mark— Cold reading using subjective validation The con— cold reading —hinges on a quirk of human psychology: if we personally relate to a statement, we will generally consider it to be accurate. This unfortunate side effect of how our mind functions is called subjective validation . Subjective validation