메릴랜드 대학교와 구글 딥마인드의 연구진은 AI가 생성한 소설이 단순한 문체를 넘어 서사적 구조의 한계로 인해 인간의 창작물과 명확히 구분된다고 분석했습니다. 연구진은 문체가 아닌 플롯, 캐릭터, 시간적 구조 등 서사적 특징을 분석하는 도구인 '스토리스코프(StoryScope)'를 개발하여 AI 창작물을 탐지했습니다. 이는 단순한 텍스트 탐지를 넘어 AI가 구성하는 이야기의 근본적인 구조적 한계를 증명한다는 점에서 중요합니다.
번역된 본문
메릴랜드 대학교 칼리지파크 캠퍼스와 구글 딥마인드 연구진의 사전 출판 논문에 따르면, 인공지능이 작성한 소설은 복잡한 이야기 구조를 다루는 데 어려움을 겪고 어색하게 교훈을 강조하는 경향이 있어 쉽게 탐지할 수 있다고 합니다. 연구진은 AI 소설이 전형적인 긴줄표(em-dash, —)의 과도한 사용이나 뻔한 AI 클리셰를 넘어, 텍스트 자체의 공식화된 성격과 더 깊은 관련이 있는 '특징(tells)'을 가지고 있다는 사실을 발견했습니다. 5만 편 이상의 AI 생성 단편 소설을 분석한 이 연구는 “AI 이야기들은 주제를 과도하게 설명하고 깔끔하고 단일한 전개의 플롯을 선호하는 반면, 인간의 이야기들은 주인공의 선택을 더 도덕적으로 모호하게 구성하고 시간적 복잡성이 더 높다”고 밝혔습니다. 또한 “클로드(Claude)는 사건의 전개가 현저히 단조로우며, GPT는 꿈(몽환) 장면에 지나치게 집착하고, 제미나이(Gemini)는 기본적으로 캐릭터의 외모 묘사에만 의존한다. AI가 생성한 이야기들은 서사적 공간의 특정 영역에 뭉쳐 있는 반면, 인간이 쓴 이야기들은 훨씬 다양성을 보인다. 더 나아가 이러한 결과는 단순한 글쓰기 스타일뿐만 아니라 기저의 서사 구조 차이를 통해 인간의 창작물과 AI 소설을 분리해 낼 수 있음을 시사한다”고 설명했습니다. 기본적으로 AI가 생성한 소설은 현재 품질이 떨어져 탐지하기가 매우 쉽습니다. 일반적인 탐지 방법은 긴줄표의 과도한 사용, 'delve'라는 단어의 남용, 또는 '고블린'에 대한 집착과 같은 문체적 표지를 찾는 것이지만, 이 프로젝트는 조금 다른 접근을 시도했습니다. 이 프로젝트의 저자 중 한 명이자 AI 탐지 기업 팬그램(Pangram)에서 인턴으로 일하고 있는 메릴랜드 대학교의 연구원 제나 러셀(Jenna Russell)은 404 매체(404 Media)에 “이 프로젝트의 아이디어는 단순한 텍스트 탐지 수준을 넘어, 인간의 아이디어와 AI가 생성한 아이디어를 분리할 수 있는 어떤 공간으로 나아가기를 바라는 마음에서 출발했다”고 말했습니다. 러셀과 그녀의 팀은 AI 소설에서 이른바 '서사적 특징(narrative features)'을 탐지하려는 시도를 하기로 결정했습니다. 이 탐지기는 '스토리스코프(StoryScope)'라고 불리며, 소설의 서사적 특징에 대한 분류 체계를 제안했던 2025년 벤치마크인 '나라벤치(NarraBench)'를 기반으로 구축되었습니다. 스토리스코프는 소설이 플롯 전개, 캐릭터 묘사, 배경 설정, 그리고 시간 구조를 어떻게 처리하는지를 분석하여 해당 작품이 인간이 쓴 것인지 AI가 쓴 것인지 판별했습니다. 러셀은 “이것은 겉으로 드러난 부분 아래를 파고들어 아이디어 자체에 더 집중하려는 나의 첫 번째 시도였다”며 “우리는 오직 서사적 특징에만 의존해서 일반적인 AI 탐지기에 얼마나 근접할 수 있는지 확인하여 이러한 구조적 차이가 실제로 존재하는지 알아내길 원했다. 이 방법은 또한 탐지 과정에 일종의 해석 가능성을 부여하는데, 이는 해당 분야의 미해결 과제다. 서사적 특징을 사용하면 우리는 (이야기에 포함된 하위 플롯의 수와 같은) 특정한 실체 있는 특징들을 명확히 짚어낼 수 있다. 최근 이 연구가 사람들의 공감을 이끌어낸 이유가 바로 이것이라고 생각한다. 사람들은 '아, 이것이 AI가 소설을 쓰는 근본적인 방식의 특징이구나'라고 확실히 말할 수 있기 때문이다”라고 설명했습니다. 스토리스코프를 테스트하기 위해 연구진은 10,272편의 인간 작성 소설을 선정한 다음, 제미나이 2.5(Gemini 2.5)를 사용해 이를 역설계하여 글쓰기 프롬프트로 만들었습니다. 그런 다음 수천 개의 프롬프트를 제미나이 3 플래시(Gemini 3 Flash), 딥시크 V3.2(DeepSeek V3.2), 클로드 소넷 4.6(Claude Sonnet 4.6), 키미 K2.5(Kimi K2.5), GPT 5.4에 투입했습니다. 프롬프트와 그 결과로 나온 AI 스토리를 포함한 모든 데이터는 허깅 페이스(Hugging Face)에서 확인할 수 있습니다. 소설의 출처 데이터로 연구진은 불법 복제된 전자책에서 수집된 18만 3,000권의 책이 담긴 데이터베이스인 'Books3' 데이터셋을 사용했습니다. 이 데이터셋은 여러 건의 소송의 대상이 되었으며 알려지지 않은 수의 대형 언어 모델(LLM)을 훈련하는 데 사용되었습니다. 스토리스코프 연구에는 조이스 캐럴 오츠, 스티븐 킹, 루이스 라무어, 샬럿 퍼킨스, 할런 엘리슨 등 유명 단편선에서 발췌한 역사상 가장 유명한 단편 소설 1만 편 이상이 포함되었습니다. 이 소설들은 모두 AI에 의해 가장 기본적인 요소로 분해된 뒤, 다른 대형 언어 모델에 투입되어 이를 복제할 수 있는지 확인하는 과정을 거쳤습니다. 러셀은 이 데이터셋이 논란의 여지가 있다고 말했습니다. 그녀는 “그래서 우리는 이 데이터를 공개하지 않는다”고 덧붙였습니다. 이 연구 논문 자체에도 관련 공시 사항이 포함되어 있습니다.
Fiction written by artificial intelligence is easy to detect because it struggles with complex story structure and tends to moralize in clunky ways, according to a preprint study from researchers at University of Maryland, College Park and Google DeepMind. They found that AI fiction has tells that go beyond stereotypical overuse of em-dashes and other obvious AI tropes and have more to do with the formulaic nature of the text itself. “AI stories over-explain themes and favor tidy, single-track plots while human stories frame protagonists’ choices as more morally ambiguous and have increased temporal complexity,” the study, which looked at more than 50,000 AI-generated short stories, found. “Claude produces notably flat event escalation, GPT over-indexes on dream sequences, and Gemini defaults to external character description. We find that AI-generated stories cluster in a shared region of narrative space, while human-authored stories exhibit greater diversity. More broadly, these results suggest that differences in underlying narrative construction, not just writing style, can be used to separate human-written original works from AI-generated fiction.” Basically, AI-generated fiction sucks and at the moment is easy to detect. The typical method of detection involves looking for stylistic markers such as an abundance of em-dashes, the overuse of the word “delve,” or an obsession with goblins , but this project tried something different. “The idea for this project came because we are hoping to eventually move past plain text detection, into some sort of space where we can separate human ideas from AI-generated ideas,” Jenna Russell, a University of Maryland researcher and one of the study’s authors, told 404 Media. Russell is also an intern at the AI-detection company Pangram. Russell and her team decided to attempt to detect what she called “narrative features” in AI- generated fiction. The detector is called StoryScope and it builds on NarraBench, a 2025 benchmark that suggested a taxonomy of narrative features in fiction. StoryScope looked at how fiction handled plot development, character descriptions, setting, and temporal structure to determine if something was written by a human or an AI. “It was my first attempt at getting 'under the surface' and focusing more on ideas,” Russell said. “We wanted to see how close to typical AI-detection we could get by only relying on the narrative features, to understand if this sort of structural difference really even exists. This method also adds some interpretability to detection, which is an open question in the field. Using narrative features, we can point to certain tangible features (such as the number of subplots included in a story). I think this is why it's struck a chord recently, people can really say ‘ah these are some of the underlying traits of how AI writes fiction.’” To test StoryScope, the researchers selected 10,272 human-written stories then reverse engineered them into writing prompts using Gemini 2.5. Then it took those thousands of prompts and fed them into Gemini 3 Flash, DeepSeek V3.2, Claude Sonnet 4.6, Kimi K2.5, and GPT 5.4. All of the data — including the prompts and the resulting AI stories — are available on Hugging Face . To source the stories, the researchers used the Books3 dataset — a database of 183,000 books collected from pirated ebooks . The dataset is the subject of several lawsuits and has been used to train an unknown number of LLMs. The StoryScope study included more than 10,000 of some of the most famous short stories ever written, many of them pulled from popular anthologies. There’s Joyce Carol Oates, Stephen King, Louis L'Amour, Charlotte Perkins, and Harlan Ellison. All have been rendered down to their base elements by AI and then regurgitated into a different LLM to see if it can replicate them. Russell told me the dataset was controversial. “Hence why we do not release it to the public,” she said. The study itself contained a disclosure. “We acknowledge the copyright issues related to the Books3 dataset and do not endorse its use for model training or commercial text generation,” it said. “The use of the dataset in our paper is restricted to academic purposes only and is meant to understand the narrative differences in human-written and AI-generated text to help inform discussions on AI-detection, authorship, and copyright policy.” The various AIs, of course, can’t possibly replicate the prose of O. Henry. So what, according to StoryScope, are the narrative quirks of LLM-written simulacra of English’s grand works of fiction? AI tools tend to over explain themes, for one. “Narrators explicitly explain the story’s theme 77% of the time, versus 52% for humans: a grieving character’s arc will typically end with the narrator stating the lesson learned. AI dialogue serves philosophical debate more often (59% vs. 34%), and references to other works tend to be vague allusions (72% vs. 50%) rather than specific, named references. The pattern is one of over-determination: AI spells out meaning rather than trusting the reader to infer,” the study said. AI also more often avoids subplots and fails to play with time jumps and flashbacks. The systems overwrite passages about the body and senses. “Where a human author might write that a character ‘felt afraid,’ AI renders fear as a tightening chest, cold sweat, and dimming lamplight,” the study said. Humans also spin more complicated narratives involving more characters and locations than AI can handle. Humans also reference other works of fiction, specific people and places in a way that AI struggles with. A disclosure caught my eye at the bottom of the StoryScope study. “Large language models and coding agents (Claude Code and Codex) are used to aid with and polish writing and generate some tables and plots,” it said. “I believe it's important to disclose AI use (and ideally think it should be more in-depth than I wrote in the paper),” Russell told me. “Most researchers are using AI, a lot of it seemingly 'slop' [...] but a lot of it is high-effort, good research. Also, technically you are supposed to disclose AI use for conference submissions, but most people don't. I want to help change that norm!” She also explained a bit more about how AI agents helped shape the project. “I use AI agents to help implement the code (using the claude code / codex interfaces). I also use them as an editor during the writing process! They have access to the project codebase and the paper latex, so the agents can implement graphics for me much more quickly than I could,” she said. “They write comments and add to the paper draft, but I keep it all in different colors so I can manually review and accept/reject/edit any suggestions from AI. I am a big believer that AI can help or hurt writing, but usually helps when not used to create more internet 'slop'.” I kept thinking about Harlan Ellison and Robert Silverberg’s story “Ship-Shape Pay-Off” being turned into an AI prompt and then spit back out by an LLM. Ellison died in 2018 and was notoriously protective of his work to the point of violence. He successfully sued James Cameron for plagiarism over The Terminator . I have a hard time imagining he’d be happy to see his story pumped into a machine, no matter the results. “A lot of people, like teachers or readers, don't really care if AI was used in the writing process, but do care if the human is the one behind the heart of it,” Russell said. “A teacher wants to know if their student understood the lesson, and a reader wants to know that the creativity behind a touching story was truly the work of the human author.” About the author Matthew Gault is a writer covering weird tech, nuclear war, and video games. He’s worked for Reuters, Motherboard, and the New York Times. More from Matthew Gault