메뉴
BL
The Decoder • 32일 전

퓨 조사: 챗GPT 출시 이후 웹에서 AI 작성 텍스트 급증 확인

IMP
7/10
핵심 요약

퓨 리서치 센터가 영어 웹페이지 약 50만 개를 분석한 결과, 챗GPT 출시 이후 게시된 페이지의 3분의 1 이상이 AI 생성 텍스트의 징후를 보였습니다. 상업용 .com 도메인은 .edu나 .gov 사이트보다 약 10배 더 높은 비율로 AI 텍스트를 포함했으며, 'delve' 같은 AI 선호 단어와 대시(-), 옥스포드 콤마 사용이 크게 늘었습니다. 다만 AI 검출 도구가 완전 자동 생성과 부분적 AI 보조 작성을 구분하지 못한다는 한계도 있습니다.

번역된 본문

퓨 조사, 챗GPT 출시 이후 웹에서 AI 작성 텍스트 급증 확인 Matthias Bastian, 2026년 8월 24일, THE DECODER

핵심 요점

  • 퓨 리서치 센터가 영어 웹페이지 약 50만 개를 분석한 결과, 챗GPT 출시 이후 게시된 페이지의 3분의 1 이상이 AI 생성 텍스트의 징후를 보였습니다.
  • Common Crawl 웹 아카이브의 텍스트가 Open Pangram 검출 도구를 사용해 분석되었습니다.
  • 상업용 .com 도메인은 .edu나 .gov 사이트보다 약 10배 더 자주 AI 생성 텍스트를 포함하고 있습니다.
  • 다만 현재 검출 도구가 완전 자동 생성 텍스트와 부분적으로 AI 보조를 받은 글을 거의 구분하지 못하기 때문에 이 분석에는 한계가 있습니다.

퓨 리서치 센터는 영어 웹페이지 약 50만 개에서 AI 생성 콘텐츠를 분석했습니다. 챗GPT가 출시된 이후 온라인에서 기계가 작성한 텍스트의 비중이 급격히 상승한 것으로 나타났습니다.

분석 대상 텍스트는 Common Crawl 웹 아카이브에서 가져왔으며, AI 검출 도구 Open Pangram을 사용해 기계 작성 여부 징후를 확인했습니다.

2026년 7월 표본에서는 검토된 전체 페이지의 약 10%가 AI 작성의 명확한 징후를 보였습니다. 그러나 표본을 챗GPT 출시 이후 게시된 페이지로만 한정하면 상황이 크게 달라집니다. 분석에 따르면 이러한 최신 페이지의 3분의 1 이상이 AI 작성 징후를 보였습니다.

이 흐름은 2022년 말 챗GPT와 함께 시작되었으며, 이후 AI 생성 웹 콘텐츠로 보이는 비중은 꾸준히 상승해 왔습니다.

.com 도메인 페이지 중 약 10%가 AI 작성 징후를 보인 반면, .org 도메인은 4.6%, .edu와 .gov 도메인은 각각 약 1%에 불과했습니다. 이는 상업용 웹사이트가 학교나 정부기관 페이지보다 AI 작성 텍스트를 포함할 확률이 약 10배 높다는 의미입니다.

"Delve", 대시(-), 옥스퍼드 콤마가 급증

퓨의 분석은 2023년 이후 웹에서 훨씬 더 흔해진 여러 언어 패턴을 발견했습니다. 대시(em dash)는 2023년보다 약 2배 자주 등장하며, 옥스퍼드 콤마 사용은 63% 증가했습니다. "delve", "interplay", "testament", "pivotal", "landscape", "tapestry", "bolstered", "crucial", "meticulous", "vibrant"처럼 AI가 선호하는 특정 단어들은 빈도가 2배 이상 늘었습니다. "it's not just X, it's Y(단순히 X가 아니라 Y이다)" 패턴의 부정적 병렬구조는 절대적 수치로는 여전히 드물지만 거의 3배 증가했습니다. 기업 홍보 문서를 살펴본 별도 연구에서는 이 특정 표현이 2022년 이후 4배 증가한 것으로 나타났습니다.

2026년 4월 임페리얼 칼리지 런던, 인터넷 아카이브, 스탠퍼드 대학의 연구도 비슷한 결론에 도달했으며, 신규 게시된 모든 웹사이트의 약 35%가 전체 또는 부분적으로 AI 생성된 것으로 추정했습니다. 연구진은 또한 AI 텍스트 간 의미론적 유사성이 33% 더 높고 전반적으로 훨씬 긍정적인 어조를 보인다는 점을 발견했지만, 부정적 영향에 대한 대중의 인식이 실제 데이터가 뒷받침하는 수준을 훨씬 뛰어넘는 경우가 많다고 지적했습니다.

"AI 텍스트"의 정의는 여전히 모호

이 연구와 유사한 연구들, 그리고 대중 논쟁에도 문제가 있습니다. "AI 텍스트"가 무엇을 의미하는지 아무도 합의하지 못한다는 점입니다. 그 범위는 완전 자동화된 콘텐츠부터 AI로 다듬은 인간 초안, 모델이 몇 문장에만 개입한 텍스트까지 아우릅니다.

제가 수백 개의 자체 텍스트를 테스트한 경험에 비추어 보면, Open Pangram과 유사한 모델들은 인간 또는 기계가 작성했을 가능성이 높은지에 대해 대략적인 판단만 내릴 수 있습니다. AI가 얼마나 관여했는지, 어느 단계에서 관여했는지 신뢰할 수 있게 알려주지 못하며, 여전히 잘못 판단하는 경우가 잦습니다. 그런데 이들은 서로 매우 다른 작업 방식입니다.

AI 텍스트를 둘러싼 대중 논쟁은 점점 더 양극화되고 있으며, 최근 Anthropic이 Claude 출력에 워터마크를 적용하려는 계획을 둘러싼 논의에서 이 점이 분명히 드러났습니다. AI 도구 사용에는 이미 양방향으로 작용하는 낙인이 존재하며, 직장 관련 연구들에서도 이 점이 기록되어 있습니다. 어느 진영도 절충의 여지를 많이 남겨두지 않고 있습니다.

원문 보기
원문 보기 (영어)
Pew study confirms sharp rise of AI-written text on the web since ChatGPT's launch Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 24, 2026 Nano Banana Pro prompted by THE DECODER Key Points In an analysis of nearly half a million English-language web pages, the Pew Research Center found that more than a third of pages published after ChatGPT's launch show signs of AI-generated text. Texts from the Common Crawl web archive were analyzed using the Open Pangram detection tool. Commercial .com domains contain AI-generated text roughly ten times more often than .edu or .gov sites. The analysis has limits, though, mainly because current detection tools can barely tell the difference between fully automated text and writing that was only partly AI-assisted. Ask about this article… Search The Pew Research Center analyzed nearly half a million English-language web pages for AI-generated content. Since ChatGPT launched, the share of machine-written text online has climbed sharply. The texts came from the Common Crawl web archive and were checked for signs of machine authorship using the AI detection tool Open Pangram . In a sample from July 2026, about 10 percent of all pages examined showed clear signs of AI authorship. Filtering the sample to only include pages published after ChatGPT's release changes the picture dramatically. More than a third of those newer pages show signs of AI authorship, according to the analysis . The trend kicked off with ChatGPT in late 2022, and the share of likely AI-generated web content has climbed steadily ever since. Ad About one in ten pages with a .com domain shows signs of AI authorship, while .org domains sit at 4.6 percent and .edu and .gov domains come in at only about 1 percent each. That makes commercial websites roughly ten times more likely to contain AI-written text than pages from schools or government agencies. Ad "Delve," em dashes, and Oxford commas are booming Pew's analysis found several language patterns that have become much more common on the web since 2023. Em dashes now show up about twice as often as they did in 2023, and Oxford comma usage has jumped 63 percent. Certain AI-favorite words like "delve," "interplay," "testament," "pivotal," "landscape," "tapestry," "bolstered," "crucial," "meticulous," and "vibrant" have more than doubled in frequency. Negative parallelisms following the "it's not just X, it's Y" pattern have nearly tripled, though they remain rare in absolute numbers. A separate study looking at corporate PR documents found that this particular phrase quadrupled since 2022. Ad A study by Imperial College London, the Internet Archive, and Stanford University from April 2026 reached a similar conclusion, finding that roughly 35 percent of all newly published websites were fully or partly AI-generated. The researchers also found 33 percent higher semantic similarity between AI texts and a much more positive tone overall but cautioned that public perception of negative effects often goes well beyond what the data actually supports. What counts as "AI text" remains fuzzy There's a problem with this and similar studies, and with the public debate too. Nobody agrees on what "AI text" even means. The spectrum runs from fully automated content to human drafts polished with AI to texts where a model only stepped in for a few sentences. Ad Open Pangram and similar models can, in my experience from testing hundreds of my own texts, only make a very rough call on whether a human or a machine likely wrote something. They can't reliably tell you how much AI was involved or at what stage, and they still misfire regularly. Yet these are very different ways of working. Ad Public debate around AI text is growing more polarized, as the recent discussion about Anthropic's planned watermark for Claude output made clear. Using AI tools already carries a stigma that cuts both ways, something workplace studies have documented as well. Neither camp leaves much room for the messy reality of how people actually write with these tools. And since AI adoption isn't slowing down, figuring out that middle ground is going to matter more and more. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Pew Research