메뉴
BL
The Decoder • 49일 전

AI가 쓴 소설이 인간 작품보다 낫게 평가받는 이유

IMP
8/10
핵심 요약

최근 연구에 따르면, 독자들은 ChatGPT가 생성한 단편 소설과 인간이 쓴 소설을 구별하지 못했으며, AI가 쓴 텍스트를 오히려 더 몰입감 있고 고품질로 평가했습니다. 하지만 기계가 썼다는 사실을 알게 되면 평가 점수가 하락하는 편향이 나타났습니다. 이는 창작 분야에서 AI가 인간과 대등한 수준의 결과물을 내고 있음을 시사합니다.

번역된 본문

독자들은 ChatGPT가 생성한 단편 소설과 인간이 쓴 단편 소설을 구별하지 못합니다. 심지어 AI가 생성한 텍스트에 더 높은 점수를 주기도 하지만, 이는 기계가 썼다는 사실을 모를 때만 해당합니다.

총 2,500명 이상의 참여자를 대상으로 진행된 세 차례의 실험에서, 참가자들은 인간이 쓴 소설과 ChatGPT가 생성한 허구의 단편 소설을 구별해 내는 데 실패했습니다 (무작위로 찍는 것과 다르지 않았습니다).

첫 번째 실험에서 1,682명의 참여자 각각은 약 1,000단어 분량의 단편 소설 6편 중 하나를 읽었습니다. 그중 3편은 유명 문학 잡지나 단편 소설집에서 발췌한 것이었습니다. 나머지 3편은 ChatGPT 4.0을 사용하여 생성했으며, 인간 원작의 주제, 스타일 및 서사적 관점을 바탕으로 프롬프트를 작성했습니다. 참가자의 절반에게는 이 이야기가 인간에 의해 쓰였다고 말했습니다. 나머지 절반에게는 ChatGPT가 작성했다고 했습니다. '판단과 의사결정(Judgment and Decision Making)' 저널에 연구를 발표한 시드니 시어스(Sydney Sears)와 디나 스콜니크 와이스버그(Deena Skolnick Weisberg)에 따르면, 이 정보는 각 그룹 참가자의 절반에게만 정확했습니다.

ChatGPT의 소설은 인지된 품질과 몰입도 측면 모두에서 인간이 쓴 텍스트보다 유의미하게 높은 평가를 받았습니다. -3에서 +3까지의 척도에서 품질의 평균 점수는 인간 소설이 0.97인 반면, AI 소설은 1.54였습니다. 몰입도에서는 1.00 대 1.42로 차이가 났습니다.

참가자들의 AI에 대한 태도 또한 평가에 영향을 미쳤습니다. 실제로 누가 이야기를 썼는지와 관계없이, 참가자들은 인간이 저자라고 말했을 때 더 높은 점수를 주었습니다. AI에 대해 긍정적인 태도를 가진 참가자들은 전반적으로 더 높은 평가를 내렸습니다. 특히 이야기가 ChatGPT에서 나왔다고 말했을 때, 그들의 점수는 더욱 올라갔습니다. 반면 AI를 의심하는 참가자들에게는 그 효과가 반대로 나타났습니다. 이전의 AI 생성 시에 대한 연구에서도 유사한 편향이 발견된 바 있습니다.

총 905명의 참가자를 대상으로 한 두 차례의 추가 실험에서 연구진은 과제를 더 어렵게 만들었습니다. 각 참가자는 인간이 쓴 이야기와 AI가 생성한 이야기를 모두 읽은 후, 어느 것이 어느 것인지 알아맞혀야 했습니다. 직접 비교를 하더라도 참가자들의 성과는 무작위 확률과 다르지 않았습니다. AI 시스템에 대한 자가 보고된 사용 경험은 이야기의 출처를 정확하게 식별하는 능력과 양의 상관관계가 있었습니다. 반면 소설에 대한 자가 보고된 경험은 식별에 도움이 되지 않았습니다.

더 높은 평가가 반드시 더 나은 글쓰기를 의미하지는 않습니다

AI가 생성한 텍스트는 인간이 쓴 텍스트보다 더 매끄럽고, 읽기 쉬우며, 감정적으로 더 밝은 경향이 있습니다. 저자들에 따르면, 이러한 특성은 AI 소설이 문학적인 의미에서 실제로 더 낫기 때문이 아니라 할지라도 더 높은 평가를 받는 원인을 설명할 수 있습니다. 사람들은 일반적으로 처리하기 쉬운 자료를 선호합니다. 반면 고품질의 문학 소설은 종종 의도적으로 접근하기 어렵게 만들어 독자가 의미를 찾기 위해 노력하게 유도합니다. 연구진은 이야기가 고품질일 수 있지만 그다지 흥미롭지 않을 수 있으며, 그 반대의 경우도 가능하다고 제안합니다.

단편 소설이라는 형식 역시 AI에게 유리하게 작용했을 가능성이 높습니다. 1,000단어로 일관된 이야기를 쓰는 것은 수백 페이지에 걸쳐 그렇게 하는 것과는 매우 다른 도전 과제입니다.

그럼에도 불구하고 연구진의 결론은 명확합니다. AI는 사람들이 인간의 작업과 최소한 동등하다고 인식하는 창작물을 생성할 수 있지만, 사람들은 AI가 그럴 수 있다고 믿지 않습니다. 이 연구의 모든 데이터와 자료는 과학 개방 프레임워크(Open Science Framework)에서 자유롭게 이용할 수 있습니다.

지난 10월 스토니브룩 대학교와 콜롬비아 로스쿨의 연구에 따르면, 독자의 전문성은 어느 정도까지만 중요한 것으로 나타났습니다. 단순한 프롬프트를 사용했을 때, 전문 독자들은 인간이 쓴 텍스트를 명확하게 선호했습니다. 하지만 개별 작가의 스타일에 맞춰 모델을 구체적으로 훈련시켰을 때, 전문가들은 스타일 모방의 경우 8배, 글쓰기 품질의 경우 2배 더 자주 AI가 생성한 텍스트를 선호했습니다.

AI 뉴스, 과장 없이 – 인간이 엄선하여 제공합니다. 광고 없는 읽기 경험을 위해 THE DECODER를 구독하세요. 매주 소식을 받아보세요.

원문 보기
원문 보기 (영어)
Readers rate AI-generated short stories higher than human ones until they learn a machine wrote them Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 8, 2026 Readers can't tell the difference between short stories generated by ChatGPT and those written by humans. They even rate the AI-generated texts higher, but only as long as they don't know a machine wrote them. In three experiments with more than 2,500 total participants, test subjects did no better than chance at telling human-written and ChatGPT-generated fictional short stories apart. In the first experiment , each of the 1,682 participants read one of six short stories, each about 1,000 words long. Three came from well-known literary magazines and short story collections. The other three were generated using ChatGPT 4.0, with prompts based on the theme, style, and narrative perspective of the human originals. Half the participants were told the story was written by a human. The other half were told it came from ChatGPT. That information was accurate for only half the participants in each group, according to researchers Sydney Sears and Deena Skolnick Weisberg in their study published in the journal Judgment and Decision Making . ChatGPT's stories were rated significantly higher than the human-written texts on both perceived quality and immersion. For quality, the mean score for AI stories was 1.54 compared to 0.97 for human stories on a scale from minus 3 to plus 3. For immersion, the gap was 1.42 versus 1.00. Participants' own attitudes toward AI also shaped their ratings. Regardless of who actually wrote the story, participants gave higher scores when told a human was the author. Participants with a positive attitude toward AI generally gave higher ratings across the board. When they were also told the story came from ChatGPT, their scores rose even further. Among AI-skeptical participants, this effect flipped. An earlier study on AI-generated poems found a similar bias . In two more experiments with 905 total participants, the researchers made the task harder. Each person read both a human-written and an AI-generated story, then had to figure out which was which. Even with a direct comparison, participants performed no better than chance. Self-reported experience with AI systems correlated positively with the ability to correctly identify the stories' origins. Self-reported experience with fiction, on the other hand, didn't help participants tell them apart. Higher ratings don't necessarily mean better writing AI-generated texts tend to be smoother, easier to read, and more emotionally upbeat than human-written texts. According to the authors, these traits could explain the higher ratings without the AI stories actually being better in a literary sense. People tend to prefer material that's easier to process. High-quality literary fiction, by contrast, is often intentionally hard to access and pushes readers to work for meaning. A story can be high quality but not very engaging, and vice versa, the researchers suggest. The short story format also likely works in AI's favor. Telling a coherent story in 1,000 words is a very different challenge than doing so across hundreds of pages. Still, the researchers' conclusion is clear: AI can generate creative works people perceive as at least on par with human work, yet people don't believe AI is capable of that. All data and materials from the study are freely available on the Open Science Framework. A study from last October by Stony Brook University and Columbia Law School shows that reader expertise matters only up to a point. With simple prompts, professional readers clearly preferred the human-written texts. But when models were specifically trained on individual authors' styles, the experts preferred the AI-generated texts eight times more often for style imitation and twice as often for writing quality. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->