메뉴
BL
404 Media 19일 전

AI가 만든 콘텐츠, 링크드인과 X를 잠식하다

IMP
7/10
핵심 요약

AI 작문 탐지 기업 Pangram의 연구에 따르면, 사용자가 링크드인과 X에서 보는 장문 콘텐츠의 최대 41%가 AI에 의해 생성된 것으로 나타났습니다. 이는 AI가 생성한 스팸성 콘텐츠가 단순히 외딴 스팸 사이트에 머물지 않고, 사람들이 매일 사용하는 주요 소셜 미디어 플랫폼마저 심각하게 오염시키고 있음을 보여줍니다.

번역된 본문

AI 작문 탐지 기업의 데이터에 따르면, 대중적인 소셜 미디어 웹사이트에서 사용자들이 마주하는 콘텐츠 중 충격적으로 많은 양이 AI가 생성한 것일 수 있습니다. 해당 데이터에 따르면, 링크드인(LinkedIn)에서 사용자가 본 장문 콘텐츠의 최대 41%가 완전히 AI가 생성한 것일 확률이 높으며, X(구 트위터)의 긴 게시물의 약 3분의 1이 AI가 생성한 것입니다. 또한 레딧(Reddit)과 서브스택(Substack)의 긴 게시물 중 약 10분의 1 정도가 AI가 작성한 것으로 나타났습니다.

이 데이터는 AI가 생성한 글을 탐지하는 기업인 팽그램(Pangram)의 크롬 확장 프로그램을 통해 수집되었습니다. 팽그램의 크롬 확장 프로그램은 사용자가 인터넷을 탐색하는 동안 접하는 글을 스캔하여 특정 게시물이 AI가 생성했는지, 아니면 사람이 작성했을 가능성이 높은지 판단합니다. 팽그램은 사용자가 인터넷을 서핑하는 동안 백그라운드에서 수동적으로 작동하기 때문에, 실제로 사용자가 보는 게시물만 스캔합니다. 이를 통해 AI 쓰레기(slop)가 인터넷 전체를 광범위하게 오염시키는 것인지, 아니면 실제로 인간이 사용하는 인터넷을 정말로 독성으로 물들이고 있는지에 대한 질문에 답을 얻을 수 있었습니다. 그 대답은 명확합니다. AI가 생성된 쓰레기 글들은 아무도 읽지 않는 인기 없는 자동화된 SEO 팜이나 스팸 사이트에만 고립되어 있지 않으며, 사람들은 인기 있는 대형 사이트에서 정기적으로 AI 쓰레기 더미를 헤치고 나아가고 있습니다.

팽그램의 CEO인 맥스 스페로(Max Spero)는 전화 인터뷰에서 "사람들이 실제로 얼마나 많은 AI 콘텐츠를 보고 있는지에 대한 연구는 이전에는 없었던 것입니다. AI 콘텐츠는 독자의 시간을 뺏는 세금과도 같습니다."라고 말했습니다. (팽그램은 이전에 404 Media에 광고를 게재한 적이 있습니다. 저는 AI가 생성된 콘텐츠가 소셜 미디어를 장악하고 알고리즘을 무자비하게 조작하는 것에 대해 여러 기사를 써왔으며, 이런 '쓰레기 콘텐츠'의 실제 범람 정도를 측정하려는 다른 데이터를 본 적이 없기에 이 데이터를 취재하고 있습니다.)

이 연구를 위해 팽그램은 특별히 크롬 확장 프로그램 사용자들에게 자신의 브라우징 결과를 회사와 공유하는 것에 동의(옵트인)해 달라고 요청했습니다. 회사는 2개월 동안 링크드인, 미디엄(Medium), X, 레딧, 서브스택 전반에 걸쳐 사용자가 자연스럽게 스크롤하여 지나친 약 100만 개의 게시물을 분석했습니다. 팽그램은 모든 플랫폼에서 예외 없이 짧은 게시물보다 긴 게시물이 AI로 생성되었을 가능성이 훨씬 높다는 사실을 발견했습니다.

회사는 분석한 콘텐츠를 '단문(50~250단어)'과 '장문(250단어 이상)'으로 나누었습니다. 어쩌면 당연하게 들릴 수도 있지만, 데이터에 따르면 링크드인과 X의 새로운 기사 형식의 장문 게시물 중 상당 부분이 완전히 AI가 생성되었거나 AI의 보조를 받은 것(AI가 초안을 작성, 편집, 재작성하고 일부 인간의 요소가 가미된 형태)으로 나타났습니다. 회사 측은 데이터상으로 분석된 링크드인 장문 게시물의 40%가 완전히 AI가 작성한 것이었으며, X의 기사 중 4분의 1은 완전히 AI가 작성한 것이었지만 추가로 23%는 AI의 보조를 받은 것이라고 밝혔습니다. 사람들은 보통 몇 마디 되지 않는 답글이나 인용 트윗의 날카로운 댓글을 굳이 AI로 생성하게 하지 않기 때문에 장문 콘텐츠일수록 AI가 생성되었을 가능성이 높다는 것은 직관적으로 이해가 됩니다. 또한 AI는 말이 많기로(verbose) 유명하므로, AI가 생성한 콘텐츠가 긴 게시물에 더 많이 나타날 가능성이 높습니다.

회사는 블로그 게시물에서 "우리의 데이터는 AI가 생성한 콘텐츠가 모든 플랫폼에 걸친 문제이며, 특히 장문 콘텐츠에 큰 타격을 주고 있음을 보여줍니다. 예상과 달리 사람들은 자신의 실명이 연결된 전문적인 환경에서 AI가 자신을 대신해 말하도록 하는 것을 압도적으로 기꺼워하며, 캐주얼하고 익명인 플랫폼에서는 그럴 가능성이 적습니다."라고 작성했습니다. 또한 이 연구는 링크드인과 레딧의 최상위 게시물이 원본 게시물 아래에 달리는 댓글보다 AI로 생성되었을 가능성이 훨씬 높다는 것을 발견했습니다.

저는 '당신의 AI 사용이 내 뇌를 망치고 있다'라는 제목의 기사를 쓰기 위해 스페로와 인터뷰한 후 몇 달 동안 팽그램의 크롬 확장 프로그램을 사용해 왔습니다. 그 기사에서 저는 인터넷을 서핑할 때 특정 글이 AI가 생성한 것인지 아닌지를 판단하려고 끊임없이 평가하면서 겪는 인지적 부담에 대해 썼습니다. 그 기사를 작성한 후, 저는 팽그램의 크롬 확장 프로그램을 사용해 그들의 평가가 어떠한지 확인해 보기로 결심했습니다.

원문 보기
원문 보기 (영어)
A shocking amount of the content that users encounter on popular social media websites is likely AI generated, according to data from a company that detects AI writing. As much as 41 percent of longform written content seen by users on LinkedIn is likely to be fully AI-generated and roughly a third of longer posts on X are AI-generated; roughly one-in-ten longer Reddit and Substack posts are AI, according to the data . The data was collected using a Chrome extension from Pangram, a company that detects AI-generated writing. Pangram’s Chrome extension scans writing that users encounter while browsing and determines if any given post is likely AI-generated or likely human written. Because Pangram works passively in the background while a user is browsing the internet, it only scans posts that its users actually see. This helps answer the question of whether AI slop is actually poisoning the internet that humans actually use, versus polluting the internet more broadly. The answer is unequivocal: AI slop writing is not just sequestered off on unpopular automated SEO farms or spam sites that no one reads; humans are regularly wading through AI dreck on hugely popular sites. “This isn’t something that had really been studied before—how much AI content people are actually seeing,” Max Spero, the CEO of Pangram, told me in a phone interview. “AI content is a tax on readers’ time.” (Pangram formerly advertised on 404 Media. I am covering this data because I have written many articles about how AI-generated content is taking over social media and is brute forcing social media algorithms , and I have not seen other data that attempts to measure the actual popularity of slop.) For this research, Pangram specifically asked users of its Chrome extension to opt-in to share Pangram browsing results with the company. The company analyzed roughly a million posts that its users organically scroll through across LinkedIn, Medium, X, Reddit, and Substack over a two-month period. Pangram found that, universally, longer posts on all platforms are more likely to be AI-generated than shorter posts. The company split the content it analyzed into “shortform” (between 50 and 250 words) and “longform” (longer than 250 words). The data suggests, perhaps unsurprisingly, that a huge portion of longform posts on LinkedIn and X’s new article format are fully AI-generated or AI-assisted (meaning drafted, edited, or rewritten by AI with some human elements). Forty percent of longform LinkedIn posts analyzed in the data were fully AI-written; a quarter of X articles were fully AI written, but another 23 percent of X articles were AI-assisted, the company said. It intuitively makes sense that longer form content is more likely to be AI-generated, because people usually won’t bother to AI-generate a few word response or a pithy comment on a quote tweet, for example. AI is also famously verbose , meaning AI-generated content is more likely to show up in longer posts. “Our data shows that AI-generated content is a problem across all platforms, and it is hitting longform content especially hard,” the company wrote in a blog post. “Contrary to what one might expect, people are overwhelmingly willing to use AI to speak on their behalf in professional settings that are associated with their real identity, and less likely to use it on casual and anonymous platforms.” The study also found that top-level posts on LinkedIn and Reddit are far more likely to be AI-generated than the comments underneath an original post. I have been using the Pangram Chrome extension for several months now, after interviewing Spero for an article I wrote called “ Your AI Use Is Breaking My Brain .” In that article, I wrote about the cognitive weight of the constant assessments I am doing when I’m browsing the internet, trying to determine whether a piece of writing is AI-generated or not. After writing that article, I decided to try the Pangram Chrome extension to see whether its assessments of likely AI-generated writing aligned with my own brain’s assessments. After using the extension for nearly two months, my experience has largely aligned with what Pangram’s data suggests: Many of the longform articles I see on X are obviously AI generated, and are detected by Pangram as such. A huge amount of the LinkedIn posts I see are obviously AI-generated. Because of the way the study worked, by passively detecting AI generated content that people see in their normal browsing, the data is potentially more useful than other studies that have sought to estimate the raw percentage of AI-generated content on the internet, but not whether anyone was actually seeing that content. These prior studies, which found that as many as a third of new sites are AI , allowed for the possibility that AI-generated content was flooding the internet but that it was of such a low quality that actual people may not have been seeing it. The Pangram data raises questions about what platforms are doing to promote or disincentivize AI slop. LinkedIn, for example, had for years built AI writing tools into its platform meaning that it has been incredibly easy to post AI-generated content on the platform and that AI-generated content became incredibly common on the platform. In May, the company announced that it is trying to disincentivize AI content in the name of “keeping conversations real,” and the AI writing assistant is no longer built into the post button. Reddit, meanwhile, has become a vector for companies trying to game LLM tools by promoting their products on the site because AI search tools often scrape Reddit. But Reddit’s moderators are also overwhelmingly anti AI, and the company has worked to delete AI-generated posts and ban accounts that spam. On Monday, Reddit published a blog post saying that “in the age of AI, spam, bot activity, and inauthentic content are top of mind for people who love Reddit (and humans).” In the last few weeks, Reddit launched an ad campaign called “people are best” specifically highlighting that its users are human. A Reddit spokesperson referred us to the blog post when asked for comment. As we have reported before, no AI detector is 100 percent foolproof , and Pangram certainly has both false positives (human content detected as AI) and false negatives (AI content detected as human). Spero said that the company is constantly working on minimizing both, and that it estimates its false positive rate at roughly one in 10,000. He said he believes the Pangram data is likely a “lower bound” and that the actual problem is likely worse, because people who are willing to install AI detectors on their browsers are likely trying to avoid AI-generated content. “I think the data generalizes out [to non Pangram users], but that it’s a lower bound on AI content because someone with the Pangram extension probably cares more about seeing AI content than the average person and would be more likely to block or mute AI posters,” he said. A LinkedIn spokesperson told 404 Media in a statement that “Professionals come to LinkedIn to hear from real people and their unique insights and perspectives. We actively work to reduce low quality, automated or generic content, and while AI can be used to beat the blank page problem, our focus is on surfacing professional conversations that help people advance their careers.” Substack and X did not respond to a request for comment. About the author Jason is a cofounder of 404 Media. He was previously the editor-in-chief of Motherboard. He loves the Freedom of Information Act and surfing. More from Jason Koebler