메뉴
BL
404 Media • 29일 전

AI가 만든 '유령 저자'들이 학술 출판을 오염시키고 있다

IMP
7/10
핵심 요약

삼성과 바르샤바 대학 연구진은 실존하지 않는 인물인 '엘레나 바스케즈'와 '마커스 첸' 등 대형 언어 모델(LLM)이 반복적으로 생성하는 이름들이 수백 편의 AI 생성 논문과 문서에 공동 저자로 등장하며 학술 기록을 오염시키고 있다고 밝혔습니다. 이들 유령 논문은 CERN이 운영하는 Zenodo에서 실제 DOI를 발급받아 Google Scholar와 Semantic Scholar 등 학술 검색 플랫폼에 인덱싱되고 있어, 대규모 학술 기록 오염의 인프라가 이미 갖춰져 있다는 지적입니다.

번역된 본문

삼성과 바르샤바 대학의 새로운 프리프린트 연구 논문에 따르면, "엘레나 바스케즈(Elena Vasquez)와 마커스 첸(Marcus Chen)은 화산 전문가, 우주비행사, 스릴러 주인공, 팟캐스트 진행자, 학술 공동 저자 등으로 수백 건의 서로 독립적으로 제작된 AI 생성 문서에 등장했지만, 실제로 살아본 적이 없는 인물"이다. 이 논문은 수백 편의 AI 생성 학술 논문, 기사, 책에 공동 저자로 등장한 다수의 이름을 식별해냈다. 이 저자들은 실제로 존재하지 않으며, 특정 분야의 전문가를 생성하라는 요청을 받았을 때 대형 언어 모델(LLM)이 반복적으로 만들어내는 이름들이다.

'고스트 커플: 상관된 LLM 이름 선호도와 웹·학술 출판에 대한 그들의 출몰(The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing)'이라는 제목의 이 논문은 특정 LLM이 특정 맥락에서 계속 같은 이름을 만들어내는 알려진 현상을 활용했다. 예를 들어, 6월에 샘(Sam)은 ChatGPT, Gemini, Claude가 생성하는 소설에서 '엘리아스 손(Elias Thorne)'이라는 이름을 사용할 가능성이 높으며, 이 캐릭터가 종종 등대지기로 등장한다는 글을 썼다. 마찬가지로 사용자들은 ChatGPT에게 소프트웨어 개발자를 생성해달라고 요청하면 이름이 종종 '마커스 첸'이 된다는 것을 발견했다.

연구진은 LLM이 종종 같은 이름을 기본값으로 사용할 뿐만 아니라 '상관된 캐릭터 앙상블(correlated character ensembles)'을 만들어낸다는 것, 즉 일부 이름들이 함께 나타날 가능성이 더 높다는 것을 보여줬다. AI 모델들이 일관되게 생성한 다른 이름들로는 Claude의 '엘레나 아마라 오카포르(Elena Amara Okafor)', Gemini의 '아리스 손(Aris Thorne)'과 '레나 페트로바(Lena Petrova)', ChatGPT의 '엘라라 보스(Elara Voss)' 등이 있다.

이달 초, 필자는 '리서치 골드(Research Gold)'라는 회사에 관한 기사를 보도했다. 이 회사는 인간이 작성한 의학 연구라고 주장하는 콘텐츠를 판매했지만, 실제로는 전부 AI가 생성한 것이었다. 그 회사의 창립자이자 수석 방법론자는 '엘레나 바스케즈'라는 이름이었다. 리서치 골드는 기사가 게재된 후 사이트에서 엘레나 바스케즈를 삭제했다.

논문의 주 저자인 미하우 브조조프스키(Michał Brzozowski)는 이 이름들을 구글에서 검색하면 AI가 생성한 인물의 다른 사례들이 나타났다고 말했다. 예를 들어, 1월에 미국 국경순찰대원들에게 살해당한 알렉스 프레티(Alex Pretti) 사건 이후, 페이스북에 그가 비위행위 혐의로 간호사 직장에서 해고되었다는 소문이 퍼졌다. 스놉스(Snopes)는 이 허위 주장이 존재하지 않는 '집행 이사 엘레나 바스케즈 박사'에게 귀속되었다고 보도했다.

브조조프스키는 LLM이 자주 생성하는 것으로 알려진 이름으로 학술 논문 데이터베이스를 검색할 수 있었다. "CERN이 운영하며 실제 DataCite DOI를 발급하는 저장소인 Zenodo에서, 우리는 존재하지 않는 저널과 조작된 발행일을 주장하는 1,655건의 유령 저자 명의 기록을 식별했다." DOI(Digital Object Identifier, 디지털 객체 식별자)는 학술 논문을 식별하는 데 사용되는 문자와 숫자의 나열이다. 연구진은 이 AI 이름들이 저자로 된 논문 다수가 소급 날짜가 기입되어 있음, 즉 Zenodo에 업로드된 날짜와 발행일이 다르다는 것을 발견했다.

누구나 무료 계정으로 Zenodo에서 DOI를 만들 수 있지만, DOI를 가진 AI 생성 논문의 존재는 웹과 학술 출판의 다른 부분에 영향을 미친다. "이들 [AI 생성 논문]은 어떤 학술 애그리게이터든 수집할 수 있는 실제 DOI를 지니고 있다. 대규모 학술 기록 오염을 위한 인프라는 이미 갖춰져 있다. 유령 이름들은 ResearchGate에도 등장하여 여러 모델 계열에서 뽑힌 공동 연구자들과 가짜 연구 그룹을 형성하고 있으며, Google Scholar와 Semantic Scholar가 검증 없이 인덱싱하고 있다. 학술 기록이 조용히 유령들에게 시달리고 있는 것이다." ResearchGate와 Google Scholar는 모두 검색 결과에 나타날 가능성이 높은 학술 출판 애그리게이터다.

연구진은 서로 다른 LLM과 그 버전들이 특정 이름을 너무 일관되게 생성하기 때문에, 이를 활용해 AI 쓰레기 콘텐츠(AI slop)의 출처를 파악할 수 있다고 말한다. 예를 들어 '엘레나 바스케즈'라는 이름은 Claude Sonnet 4가 생성한 콘텐츠에서 특히 흔했기 때문에, 그녀를 저자로 명시한 논문은 해당 LLM이 생성했을 가능성이 높다. 다만 브조조프스키는 이 수준의 정확도가...

원문 보기
원문 보기 (영어)
“Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived,” a new preprint research paper from Samsung and the University of Warsaw said. The paper identified a number of names that co-authored hundreds of AI-generated academic papers, articles, and books. The authors don’t actually exist, but are instead names that large language models repeatedly produce when tasked with generating experts in certain fields. The paper, titled “ The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing ,” utilized a known phenomenon where certain LLMs will keep coming up with the same names in certain contexts. For example, In June, Sam wrote about how ChatGPT, Gemini, and Claude were likely to use the name Elias Thorne in fiction they generated, and that the character was often a lighthouse keeper. Similarly, users noticed that if they ask ChatGPT to generate a software developer, their name will often be Marcus Chen . The researchers were able to show not only that LLMs often default to the same names, but that they produce “correlated character ensembles,” meaning some names were more likely to appear together. Other names that were consistently generated by AI models include Elena Amara Okafor from Claude, Aris Thorne and Lena Petrova from Gemini, and Elara Voss from ChatGPT. Earlier this month, I reported a story about Research Gold, a company that offered what it claimed was human medical research, but that was in fact entirely AI generated . The founder and lead methodologist for that company was named Elena Vasquez. Research Gold removed Elena Vasquez from its site after I published the story. Michał Brzozowskim, the lead author of the paper, told me that searching for these names on Google turned up other instances of AI generated personalties. For example, following the killing of Alex Pretti at the hands of U.S. Border Patrol agents in January, a rumor spread on Facebook that he was fired from his nursing job for misconduct allegations. Snopes reported that the false statement was attributed to “executive director Dr. Elena Vasquez,” who does not exist. Brzozowskim was able to search databases of academic papers for the names they knew LLMs often generated. “On Zenodo, a CERN operated repository that mints real DataCite DOIs, we identify 1,655 ghost-authored records claiming nonexistent journals with fabricated publication dates.” A DOI, or a Digital Object Identifier, is a string of characters and numbers used to identify academic papers. The researchers saw that many of the papers authored by these AI names were backdated, meaning their publication dates were different from the date they were uploaded to Zenodo. Anyone with a free account can create a DOI on Zenodo, but the existence of AI generated papers with DOIs has impacts on other parts of the web and academic publishing. “These [AI generated papers] carry real DOIs harvestable by any scholarly aggregator; the infrastructure for large-scale scholarly record contamination is already in place. Ghost names additionally appear on ResearchGate, forming synthetic research groups with collaborators drawn from multiple model families, and are indexed without verification by Google Scholar and Semantic Scholar [...] The academic record is being quietly haunted.” Research Gate and Google Scholar are both aggregators of academic publishing that are likely to come up in search results. The researchers say that different LLMs and versions of those LLMs generate certain names so consistently, they think they can use them to determine the provenance of AI slop. For example, the name Elena Vasquez was particularly common in content generated by Claude Sonnet 4, so papers that list her as an author were likely generated by that LLM. However, Brzozowskim said that this level of accuracy might not hold for long now that AI generated content with these names is flooding the internet and feeding back into all AI models that are scraping the internet for training data. The fact that AI academic publishing is struggling to deal with the load of AI generated content isn’t new. We’ve previously reported that scientific journals have published AI generated text , that AI is impacting the peer review process , and that the open-access repository for preprint academic research Arxiv will now ban authors for a year if they are caught submitting AI generated work . The potential upside of this research is that these names might allow us to detect this AI generated content more easily. About the author Emanuel Maiberg is interested in little known communities and processes that shape technology, troublemakers, and petty beefs. Email him at emanuel@404media.co More from Emanuel Maiberg