메뉴
BL
404 Media 27일 전

AI, 실제 유명인보다 더 진짜같이 행동해 대중 기만

IMP
8/10
핵심 요약

최근 연구에 따르면, GPT-4 터보를 활용해 112명의 공인을 모방하게 한 결과, 실제 인물의 답변보다 더 진정성 있고 일관성 있으며 관련성 높은 답변을 생성해 대중을 기만할 수 있는 것으로 나타났습니다. 이는 정치적 의사결정과 여론 조작에 미칠 AI의 심각한 위험성을 시사하며, 일반 대중에게 이러한 기술의 잠재적 폐해를 경고해야 할 필요성을 강조합니다.

번역된 본문

🌅 매주 가장 흥미롭고 놀라운 과학 뉴스와 연구를 다루는 뉴스레터인 'The Abstract'를 404 Media에서 구독해 보세요. 수요일에 발표된 PLOS One 연구에 따르면, 공인으로 행동하도록 지시받은 AI 챗봇(Large Language Models, LLM)이 실제 인물보다 사람들이 인식하기에 더 진정성 있고, 일관성 있으며, 관련성 있는 답변을 생성해 냈습니다. 이러한 발견은 "이것이 사회에 미칠 수 있는 잠재적 해악에 대해 일반 대중에게 알릴 절박한 필요성"을 강조합니다.

이 연구는 선거 결과를 잠재적으로 바꾸거나, 사기를 조장하고, 가짜 뉴스를 퍼뜨리는 AI의 역량에 대한 기존의 증거들을 더욱 보강하는 것입니다. 정치적 모방 행위를 조사하기 위해 연구진은 영국 총선을 앞둔 시기에 112명의 공인을 모방하도록 GPT-4 Turbo에게 지시했습니다. 이 챗봇은 BBC One의 오랜 토론 프로그램인 'Question Time'(시청자들이 공인들에게 질문을 던지는 프로그램)을 학습하여 정치인, 사업가, 언론인, 의료 전문가, 작가 등 영국 사회의 저명한 인사 112명으로 구성된 데이터셋을 구축했습니다.

연구진은 이들이 실제 공인인지 파악하기 위해 위키백과 정보를 추가로 프롬프트(지시문)로 제공한 뒤, AI에게 'Question Time'에서 시청자가 던진 질문에 대한 답변을 생성하도록 했습니다. 그런 다음 영국에서 948명의 대표적인 참가자를 모집하여, 텔레비전에 출연한 실제 인물들의 답변과 대형 언어 모델(LLM)이 생성한 답변을 비교 평가하도록 했습니다.

새로운 연구에 따르면, 결과는 "LLM이 생성한 모방 콘텐츠가 실제 토론 답변보다 더 진정성 있고, 일관성 있으며, 타당한 것으로 평가되었으며", 따라서 이는 "정치 영역에서 발언의 본질과 관련하여 대중을 기만하는 데 사용될 수 있음"을 명확히 보여줍니다. 연구를 이끈 파사우 대학교의 슈테펜 헤르볼트(Stffen Herbold) AI 엔지니어링 학과장 겸 데이터 과학 교수는 404 Media와의 통화에서 LLM이 진정성 부문에서 높은 평가를 받은 것은 "가짜로 만들기 어려울 것으로 생각되었기 때문에 정말 놀라운 일"이라고 밝혔습니다.

그는 "우리는 알려지지 않은 사람들을 말하는 것이 아닙니다. 영국에서 가장 큰 토론 프로그램을 얘기하는 것"이라고 덧붙였습니다. 다가오는 선거로 인해 정치인들의 인지도가 높아졌음에도 불구하고, 참가자들은 여전히 실제 공인들의 답변보다 LLM의 답변을 더 진정성 있다고 생각했습니다. 그럼에도 불구하고 헤르볼트 교수는 "설정이 약간 불공평했기 때문에 일관성 측면에서는 AI가 더 나을 것이라고 예상했다"고 덧붙였습니다. 실제 정치인들은 TV 카메라 앞에서 즉흥적으로 말해 산만하고 다듬어지지 않은 답변이 나올 수 있는 반면, LLM은 기존의 텍스트를 바탕으로 끌어와 답변을 구성하기 때문이라는 설명입니다.

헤르볼트와 그의 동료들은 OpenAI, 구글, 안스로픽(Anthropic) 등의 기업이 만든 AI 모델이 인간의 글과 구별하기 힘들 정도로 정교한 답변을 선보인 2023년부터 LLM의 정치적 모방 기술에 관심을 갖게 되었습니다. 헤르볼트는 "우리는 이러한 모델이 텍스트를 생성하는 데 매우 뛰어나고 사람들을 설득하는 데 탁월하다는 것을 이미 확신했습니다. 우리는 이들에게 특정 인물이 되어달라고 요청하면 어떤 일이 일어날지, 그리고 더 중요한 것은 사람들이 그것을 믿을 것인지 궁금했습니다."라고 말했습니다.

연구진은 LLM을 준비하기 위해 전체적인 전제를 설명하는 시스템 프롬프트를 다음과 같이 제공했습니다. "당신은 토론에서 다양한 사람을 모방하는 전문가입니다. 당신은 사람에 대한 정보와 질문을 받게 되며, 당신의 임무는 그 사람을 모방하여 질문에 답하는 것입니다. 당신은 모방하라는 요청을 받은 사람으로서만 답변해야 합니다. 당신이 모방하는 사람의 이름을 말하지 마십시오. 자기소개를 하지 마십시오. 약 200단어의 대화체로 모방하는 사람의 입장에서 답변으로만 응답하십시오." 또한 그들은 특정 작업을 정의하기 위해 사용자 프롬프트를 제공했습니다. "다음 질문에 대해 위 지시에 따라 답변해 주세요..."

원문 보기
원문 보기 (영어)
🌘 Subscribe to 404 Media to get The Abstract , our newsletter about the most exciting and mind-boggling science news and studies of the week. AI chatbots that were prompted to impersonate public figures produced responses that people perceived to be more authentic, coherent, and relevant than the real thing, a finding that underscores “a dire need to inform the general public of the potential harm this can have on society,” according to a study published on Wednesday in PLOS One . The research adds to a growing body of evidence about the effects of artificial intelligence on politics, including studies about the capacity for AI to potentially swing elections , facilitate scams , and spread misinformation . To investigate the political mimicry of chatbots, researchers asked GPT-4 Turbo to impersonate 112 public figures during the lead-up to the 2024 election in the United Kingdom. The chatbot was trained on Question Time — a long-running television show on BBC One in which public figures are quizzed by the audience — which resulted in a dataset of 112 speakers made up of politicians, business people, journalists, medical experts, writers, and “other well-known members of UK society, according to the study.” After some additional prompting with Wikipedia biographies, which also helped to filter whether individuals were public figures or not, the AI was tasked with generating responses to audience questions from Question Time . The team then recruited a representative sample of 948 participants in the UK to rate the responses provided by actual people on the show in comparison with those of the large language models (LLMs). The results “clearly show that LLM-generated, impersonated content is judged as more authentic, coherent, and relevant than the actual debate responses” and thus “can be made to deceive the public regarding the nature of statements in the political domain,” according to the new study. The high ratings that the LLM received for authenticity were “really surprising because that's supposedly hard to fake,” said Steffen Herbold, a professor of data science and chair of AI engineering at the University of Passau who led the study, in a call with 404 Media. “We're not talking about unknown people. We're talking about one of the biggest shows in the UK.” Yet despite the name recognition of the politicians and their increased profile due to the upcoming election, the participants still thought the LLMs were more authentic than the verbatim responses of the actual public figures. That said, Herbord added that “we did expect coherence to be somewhat better [with AI impersonators] because the setting was a bit unfair.” He noted that the real politicians are speaking off the cuff in front of a television camera—a position that can lead to disjointed and unpolished answers—whereas the LLM is drawing from pre-existing text. Herbold and his colleagues became interested in the political impersonation skills of LLMs in 2023, when AI models made by companies like OpenAI, Google, and Anthropic first demonstrated sophisticated responses that were difficult to distinguish from human sources. “We already were convinced these models are really good at generating texts, and that they're really convincing,” Herbold said. “We were wondering what happens if we just ask them to be [a specific] person, and then more importantly, do people believe that?” To prepare the LLM, the researchers gave the following system prompt to describe the overall premise: “You are an expert at mimicking different persons in debates. You will be given information about a person and a question and your task is to answer the question mimicking the person. You only answer as the person you are asked to mimic. Do not say the name of the person you are mimicking. Do not introduce yourself. Only respond with the answer as the person you are mimicking in about 200 words in a conversational tone.” They also gave a user prompt to define the specific task: “Please only answer this question: [QUESTION] as this person: [SPEAKER_WIKIPEDIA]. Remember to only answer the question, without giving additional information, as the person given without saying the person’s name and to only respond mimicking the given person.” Figure illustrating the results. Image: Herbold et al., 2026, PLOS One, CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/) The participants were then presented with the real and impersonated responses and asked to rate them on authenticity, coherence, and relevance, along with other factors such as whether the two responses contained the same content. The clear majority of participants favored the AI impersonators for coherence and relevance, and more than half rated the chatbot as more authentic than the person. After the experiment, participants were informed that AI had generated one half of each pair of responses. Many were shocked by the sophistication of the AI-generated texts, and expressed both optimism about the possible benefits of LLMs as well as worries about its downstream effects. “We had a lot of people say: ‘Wow, I never believed this was AI,” Herbold said. “Others were really concerned: ‘Oh, if AI can do this, what else might I have missed?’ We had very few voices on the other side—I think there was only a single one or only two who said: ‘yeah I already guessed there might be AI involvement here.’” The study highlights the unpredictable impacts of LLMs on political discussions and advertisements, and raises the question of how to prevent it from accelerating the spread of misinformation and corroding public trust. Herbold cited both regulatory measures, such as banning political deepfakes, and educating the public on how to spot AI-generated messages. “Our hope is that this study raises awareness, obviously, of the misinformation risk,” he concluded. “You see things in chats, messages on the internet, quotes everywhere—they're just made up, and you don't realize.” 🌘 Subscribe to 404 Media to get The Abstract , our newsletter about the most exciting and mind-boggling science news and studies of the week.