메뉴
BL
The Decoder 33일 전

AI 탐지기 성능 편차 심화: 완벽 탐지 vs 전량 오탐

IMP
7/10
핵심 요약

미국작가조합(Authors Guild)의 테스트에 따르면, Pangram과 Grammarly의 AI 탐지 도구는 사람이 작성한 글을 100% 정확히 구별했으나, Sidekicker와 ZeroGPT는 모든 글을 AI가 쓴 것으로 오탐지하는 치명적인 실패를 보였습니다. 완벽해 보이는 도구조차 내부 구조가 '블랙박스'이며 오탐지 발생 시 작가가 계약을 잃을 수 있으므로, 테스트 결과를 맹신해서는 안 된다는 경고가 중요합니다.

번역된 본문

미국작가조합(Authors Guild)의 테스트에 따르면, Pangram(팬그램)과 Grammarly(그래머리)의 AI 탐지기들은 인간이 작성한 모든 텍스트를 사람이 쓴 것으로 완벽하게 식별해 냈습니다. Originality.ai(오리지낼리티 닷 AI) 역시 우수한 성능을 보였습니다. 이번 테스트는 생성형 AI가 대중화되기 전인 2020년에서 2022년 사이에 발행된 조합의 칼럼 10개를 사용해 진행되었습니다. 반면 Sidekicker(사이드키커)는 최악의 결과를 기록했습니다. 모든 글이 대부분 AI가 생성한 것으로 지정되었으며, 그중 두 편은 100% AI가 작성한 것으로 나타났습니다. ZeroGPT(제로지피티) 또한 신뢰할 수 없었으며, 인간이 작성한 모든 텍스트에 대해 때때로 매우 높은 AI 작성 비율을 보고했습니다.

[세부 테스트 결과 표 - AI가 작성했다고 탐지된 비율(%)]

  • 외설 청원 기각 (ZeroGPT 14.3% / Originality.ai 0.0% / Sidekicker.ai 85.0% / Grammarly 0.0% / Pangram 0.0%)
  • 반독점 소송 및 출판물 (5.3% / 0.0% / 100.0% / 0.0% / 0.0%)
  • 워홀 공정사용 서한 (40.7% / 0.0% / 79.0% / 0.0% / 0.0%)
  • 저작권 청구 위원회 (28.1% / 0.0% / 96.0% / 0.0% / 0.0%)
  • 금서 클럽 (64.5% / 1.0% / 71.0% / 0.0% / 0.0%)
  • 키스 라이브러리 불법 다운로드 소송 (26.5% / 1.0% / 71.0% / 7.0% / 0.0%)
  • 조안 디디온 사망 기사 (66.0% / 0.0% / 82.0% / 9.0% / 0.0%)
  • 얼드리치 퓰리처상 수상 (76.3% / 0.0% / 100.0% / 0.0% / 0.0%)
  • 작가 및 문학 예술 지지 (50.6% / 0.0% / 92.0% / 0.0% / 0.0%)
  • 2020년 12월 라운드업 (18.1% / 0.0% / 96.0% / 0.0% / 0.0%)

오탐지는 작가의 계약을 앗길 수 있다 그럼에도 불구하고 미국에서 가장 오래되고 규모가 큰 작가 단체는 가장 성능이 좋은 도구라 할지라도 어떤 결정을 내릴 유일한 근거가 되어서는 안 된다고 경고합니다. 이러한 도구들은 끊임없이 변화하며 그 정확도를 당연하게 여겨서는 안 되기 때문입니다.

Pangram의 CEO 맥스 스페로(Max Spero)는 최근 자사의 탐지기가 본질적으로 블랙박스이며, 텍스트가 왜 AI가 생성된 것으로 지정되었는지 상세히 설명할 방법이 없다고 밝혔습니다. 하지만 언어 모델, 특히 논리를 구성하는 방식에서 획일성을 보인다는 점에서 스스로 정체를 드러낸다고 그는 덧붙였습니다. 스페로는 인간은 훨씬 더 다양하게 글을 쓴다고 말했습니다.

미국작가조합에 따르면, 전문적으로 작성된 텍스트는 AI 출력물과 동일한 많은 통계적 패턴을 공유합니다. 단순히 언어 모델이 바로 그런 글쓰기 방식으로 훈련되었기 때문입니다. 잘못된 결과는 작가의 계약과 명성을 앗아갈 수 있으므로, 출판사는 자신들의 검사 방식을 투명하게 공개하고 항상 작가에게 반박할 기회를 주어야 합니다.

이는 골치 아픈 역설을 만들어냅니다. 수십 년 동안 명확성, 간결성, 정확성을 연마하는 데 매달려 온 작가는 정의상 AI가 생성하는 방식과 겹치는 방식으로 글을 쓸 수밖에 없습니다. 탐지 도구는 장인 정신을 마스터한 인간 작가와 그것을 모방하는 법을 배운 기계를 구별할 수 없습니다. 이 도구들이 작동하는 수준에서는 둘의 차이를 찾기 어렵기 때문입니다.

그렇다 하더라도 Pangram과 Originality가 인간이 작성한 텍스트를 사람이 쓴 것으로 신뢰성 있게 식별했다는 사실이, 그들이 AI가 생성한 텍스트를 잡아내는 데에도 동일하게 뛰어나다는 것을 의미하지는 않습니다. 이 결과는 주로 해당 도구들이 오탐지를 최소화하도록 조정되어 인간의 글이 AI로 잘못 지정되는 것을 방지하고 있음을 보여줍니다. AI가 작성했거나 AI의 도움을 받은 수많은 글은 여전히 들키지 않고 통과할 수 있습니다. 이 테스트에서 나타난 신뢰성은 가장 먼저, 그리고 무엇보다 '인간의 글쓰기를 올바르게 인식하는 것'에 국한됩니다.

탐지 논쟁 이면의 문화적 배경 오류는 계속해서 발생할 것이며, 이것이 바로 이러한 탐지기들의 유용성이 끊임없이 의문시되는 이유입니다. 특히 AI는 진정으로 유용한 글쓰기 도구가 될 수 있으며, 더 광범위한 논쟁은 종종 '글을 쓰기 위해 AI를 사용하는 것'과 '생각하기 위해 AI를 사용하는 것'을 혼동하기 때문입니다.

Pangram의 CEO 맥스 스페로와 같은 탐지기 옹호자들은 작가와 독자 간의 사회적 계약을 언급하며 자신들의 비즈니스 모델을 정당화합니다. 작가는 아이디어를 다듬기 위해 시간과 노력을 투자하고, 독자는 그것에 몰입하기 위해 시간을 투자합니다. 스페로는 만약 AI가 글쓰기의 비용을 0으로 만든다면 잘못된 인센티브가 따르게 되며, 필자가 생산하는 데 걸린 시간보다 독자가 읽는 데 더 많은 시간을 낭비하게 만드는 무가치한 콘텐츠로 인터넷이 넘쳐나게 될 것이라고 경고합니다.

원문 보기
원문 보기 (영어)
Authors Guild test finds some AI detectors perfectly identify human writing while others fail on every single text Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 25, 2026 Nano Banana Pro prompted by THE DECODER Ask about this article… Search In a test by the Authors Guild, AI detectors from Pangram and Grammarly correctly identified every human-written text as human. Originality.ai also performed well. The test used ten Guild articles published between 2020 and 2022, before generative AI went mainstream. Sidekicker delivered the worst results. Every single article was flagged as mostly AI-generated, with two scoring 100 percent. ZeroGPT was also unreliable, reporting sometimes high AI percentages for all the human-written texts. Articles ZeroGPT Originality.ai Sidekicker.ai Grammarly Pangram Obscenity Petitions Dismissed 14.3% 0.0% 85.0% 0.0% 0.0% Antitrust Litigation & Publications 5.3% 0.0% 100.0% 0.0% 0.0% Warhol Fair Use Letter 40.7% 0.0% 79.0% 0.0% 0.0% Copyright Claims Board 28.1% 0.0% 96.0% 0.0% 0.0% Banned Books Club 64.5% 1.0% 71.0% 0.0% 0.0% Kiss Library Piracy Lawsuit 26.5% 1.0% 71.0% 7.0% 0.0% Obituary: Joan Didion 66.0% 0.0% 82.0% 9.0% 0.0% Erdrich Pulitzer Prize 76.3% 0.0% 100.0% 0.0% 0.0% Support Authors & Literary Arts 50.6% 0.0% 92.0% 0.0% 0.0% The Roundup 12/2020 18.1% 0.0% 96.0% 0.0% 0.0% False positives can cost authors their contracts Still, the oldest and largest professional organization for writers warns that even the best-performing tools should never be the sole basis for any decision. These tools change constantly, and their accuracy can't be taken for granted. Ad Pangram CEO Max Spero recently explained that his detector is essentially a black box, with no way to explain in detail why a text gets flagged as AI-generated. Language models do give themselves away through uniformity, though, especially in how they build arguments. Humans write with far more variety, Spero said. Ad DEC_D_Incontent-1 Professionally written texts share many of the same statistical patterns as AI output, according to the Authors Guild , simply because language models were trained on exactly that kind of writing. False results can cost authors their contracts and their reputations, so publishers should disclose their methods and always give authors a chance to defend themselves. This creates a troubling paradox. A writer who has spent decades honing clarity, economy, and precision is, by definition, writing in a way that overlaps with what AI has learned to produce. Detection tools cannot distinguish between a human writer who has mastered the craft and a machine that has learned to imitate it, because at the level these tools operate, there may be little difference to find. Ad Author’s Guild That said, the fact that Pangram and Originality reliably identify human-written texts as human doesn't necessarily mean they're equally good at catching AI-generated ones. The results mainly show that these tools are tuned to minimize false positives, avoiding cases where human text gets wrongly flagged as AI. Plenty of texts written by or with AI could still slip through undetected. The reliability shown in this test applies first and foremost to correctly recognizing human writing. The cultural debate behind detection Errors will keep happening, and that's why the usefulness of these detectors keeps getting questioned. This is especially true since AI can be a genuinely useful writing tool, and the broader debate often conflates using AI to write with using AI to think. Ad DEC_D_Incontent-2 Detector advocates like Pangram CEO Max Spero justify their business model by pointing to a social contract between writer and reader. The writer invests time and effort to shape an idea; the reader invests time to engage with it. If AI drops the cost of writing to zero, bad incentives follow, and people flood the internet with worthless content that takes readers more time to consume than it took the author to produce, Spero said. Ad Whether a piece of writing gets its value from the typing, though, or from the topic selection, the idea, the perspective, the story, the research, the argument, and the judgment behind it, that's a different question entirely. So is whether AI text detection can actually do anything about the flood of worthless content. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Authors Guild