메뉴
BL
MIT Tech Review • 39일 전

사람들이 AI를 실제로 어떻게 쓰는지 우리는 아직 모른다

IMP
8/10
핵심 요약

스탠퍼드 STAIR 연구실이 중심이 된 'AI Observatory' 프로젝트가 기존 7개 데이터셋의 실제 AI 대화를 독립적으로 분석해, 앤스로픽·OpenAI 등이 공식 발표하는 사용 통계의 맹점을 드러냈다. 특히 앤스로픽의 경제지수 분석법을 적용하면 대화의 48%가 '업무 외'로 필터링되며, 실제 사용에는 건강·관계, 성인 콘텐츠, 괴롭힘 등 민감한 사용이 훨씬 많이 포함된다는 것이다. 이는 AI 정책 수립의 기반이 되는 데이터에 왜곡이 있을 수 있음을 시사한다.

번역된 본문

요약: AI 연구자들에 따르면, 앤스로픽과 OpenAI 같은 AI 기업들은 클로드와 ChatGPT 같은 제품이 어떻게 사용되는지에 관한 보고서를 정기적으로 발표하지만, 공개하고 싶은 데이터만 내놓는다고 한다. 스탠퍼드 신뢰 가능 AI 연구(STAIR) 랩의 컴퓨터과학 박사과정생인 안카 로이얼(Anka Reuel)은 "이를 검증할 수 있는 독립적인 출처가 없다"고 말한다. 로이얼은 이 공백을 메우기 위한 새 연구 프로젝트 'AI Observatory(AI 관측소)'의 공동 책임자다. 이것은 사용자 동의 아래 수집된 7개 기존 데이터셋을 통해 모은 클로드, 제미나이 등 인기 모델과의 실제 AI 대화를 집계·분석하는 공개 플랫폼이다. 이 관측소의 목적은 사람들이 생성형 AI를 어떻게 사용하는지 평가할 수 있도록 연구자와 정책입안자에게 독립적인 정보원을 제공하는 것이다. 로이얼에 따르면 이해당사자들은 매우 제한된 데이터를 바탕으로 AI의 편익과 위험에 관한 중대한 결정을 내리고 있다.

AI Observatory는 AI 사용이 모델별로 크게 다르고 시간이 지나며 변화해 왔음을 발견했다. 그 연구는 주요 AI 기업들의 보고서에 포착되는 것보다 훨씬 많은 민감한 행위들을 보여주는데, 기업 보고서들은 개인적 사용보다 업무에 더 초점을 맞추고 있다고 연구진은 말한다. 앤스로픽의 '이코노믹 인덱스(Economic Index)'는 AI 사용 데이터의 출처로 가장 잘 알려지고 가장 널리 인용되는 자료이지만 맹점이 있다. 이름에서 알 수 있듯, 이 지수는 클로드 AI의 업무·생산성 관련 사용에 초점을 맞추어 이와 무관한 대화는 걸러낸다. AI Observatory 팀이 앤스로픽의 방법론을 자신들의 데이터셋에 적용해 보니, 대화의 거의 절반에 해당하는 48%가 필터링되었을 것이다. 걸러진 업무 외 대화에는 건강과 관계(앤스로픽 분석의 31.2% 대비 44.2%), 성인 또는 불법 주제(2.1% 대비 7.9%), 괴롭힘과 혐오(5.66% 대비 27.5%), 성적인 콘텐츠(2.4% 대비 16.7%)가 포함될 가능성이 더 높았다. (마찬가지로 OpenAI의 2025년 ChatGPT 사용 보고서도 소비자 사용의 30%만 업무와 관련된 것으로 밝혔다.)

앤스로픽은 사람들이 클로드를 지지·동반자 관계나 심지어 CSAM(아동성착취물) 생성에 사용하는 방식에 관한 별도의 블로그 포스트를 발표한 바 있지만, AI 시스템과 사람의 상호작용을 연구하며 AI Observatory에는 참여하지 않은 UT 오스틴 정보학과 데이비드 위더(David Widder) 조교수는 "별도 보고서로 구분하기보다 [AI Observatory 같은] 조감도 분석이 있으면" 연구자들이 다양한 사용 양상을 더 일관되게 이해하는 데 도움이 된다고 말했다.

AI Observatory가 살펴본 데이터셋은 2023년부터 2025년 사이의 대화를 포함하며, 사람들의 AI 사용 방식과 각 AI 플랫폼의 대응 방식 모두에서 차이를 발견했다. AI Observatory 연구에 포함된 가장 크고 상세한 데이터셋 중 하나인 WildChat의 대화는 시간이 지날수록 프롬프트 토큰, 응답 토큰, 대화 턴 수가 증가한 것에서 알 수 있듯 더 길고 정교해졌다. 또한 스몰토크(가벼운 잡담)도 시간이 지나며 크게 늘었다. 이는 AI 동반자적 사용이 증가하고 있음을 시사하며, 동시에 AI 어시스턴트의 자기 폭로(즉, 자신이 챗봇이라는 밝힘)는 감소했다. 또한 연구진이 '민감한 사용'—성적 괴롭힘과 혐오 발언 등 잠재적으로 유해하거나 제한적인 콘텐츠가 포함된 대화—으로 분류한 교환은 줄어들었다. 이는 플랫폼들이 전반적으로 더 효과적인 안전장치를 배치하고 있음을 보여줄 수 있다.

AI Observatory는 또한 모델에 따라 AI 사용 양상이 크게 달라짐을 발견했다. 도구에 따라 사용자의 주제, 상호작용 스타일, 대화 구조, 그리고 민감한 사용 사례의 발생 가능성과 유형이 달랐다. 예를 들어 연구진은 사람들이 Grok과 제미나이를 정보 검색에 더 자주 사용한다는 것을 발견했다. 특히 Grok은 뉴스와 정치 정보에 특히 인기가 높았지만, 동시에 허위정보가 집중되는 경향도 있었다. (이는 허위정보가 얼마나 쉽게 퍼지는지 보여준 다른 연구들과도 일치한다.)

원문 보기
원문 보기 (영어)
EXECUTIVE SUMMARY AI companies like Anthropic and OpenAI regularly publish reports on how people are using products like Claude and ChatGPT, but they only release the data they want us to see, AI researchers say. “There is no independent source to corroborate it,” says Anka Reuel, a Computer Science PhD candidate at the Stanford Trustworthy AI Research (STAIR) Lab. Reuel is co-lead of a new research project, called the AI Observatory, that aims to fill in the gap. It’s a public platform that aggregated and analyzed real AI conversations with popular models like Claude and Gemini that were collected with users’ consent through seven existing datasets. The Observatory's intent is to provide independent sources of information for researchers and policymakers to assess how people are using generative AI. Stakeholders are currently making highly consequential decisions about AI’s benefits and risks based on very limited data, says Reuel. The AI Observatory found that AI use differs significantly across models, and has changed over time. Its research shows many more sensitive behaviors than are captured in reports from major AI companies, which they say focus more on work than on personal use. Anthropic Economic Index is one of the best known and most widely cited sources of AI usage data but it has blind spots. As its name suggests, it focuses on work- and productivity-related uses of Claude AI—filtering out conversations that are unrelated to these uses. When the AI Observatory team applied Anthropic’s methods to their dataset, they found that nearly half of the conversations—or 48%— would have been filtered out. Those non-work-related conversations that were filtered out were more likely to include health and relationships (44.2% versus 31.2% in Anthropic’s analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). (OpenAI’s 2025 report on ChatGPT use, similarly, found that only 30% of consumer use was related to work.) Anthropic has released separate blog posts on how people use Claude for support or companionship , and even to generate CSAM , “but having [the AI Observatory’s] bird eye view analysis rather than sectioned off into a separate report helps” researchers understand the different uses more consistently, says David Widder, an assistant professor at UT-Austin’s School of Information, who researches how people interact with AI systems and is not involved with the AI Observatory. The datasets the AI Observatory looked at include conversations that took place between 2023 and 2025, and it found differences both in how people were using AI and how various AI platforms responded. Conversations within WildChat, one of the largest and most detailed datasets included in the AI Observatory’s study, got longer and more elaborate over time, indicated by increases in prompt tokens, response tokens, and conversation turns. There was also significantly more small talk over time. That suggests that AI companionship was increasing; meanwhile the AI assistants’ self-disclosure (i.e. that it was a chatbot) decreased. Additionally, exchanges that the researchers labeled as sensitive use—meaning ones with potentially harmful or restricted content, including sexual harassment and hate speech—dropped. That might suggest that platforms were generally deploying more effective safeguards. The AI Observatory also found that AI use looked significantly different depending on the model. Depending on the tool, users ranged in topics, interaction styles, conversation structures, as well as both the likelihood and type of sensitive use cases. For example, the researchers found that people used Grok and Gemini more frequently for information retrieval. Grok, in particular, was especially popular for information on news and politics, but it was also where misinformation tended to concentrate. (This is consistent with other research that has also shown how readily misinformation proliferates on Grok. xAI did not respond to a request for comment.) Meanwhile, people were more likely to turn to Anthropic for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. There were even differences among different versions of the same model. Researchers found that people had shorter conversations with ChatGPT when it was powered by GPT-3.5, and longer and more iterative ones with GPT-4o—which makes sense given that that version became known for leading to emotional addiction . Company reports, however, didn’t tend to capture these nuances across or even within their own models. “No single company report tells the whole story,” says Shayne Longpre, a recent PhD graduate from the MIT Media Lab who co-led the research with Reuel. To create the AI Observatory, Reuel and researchers from MIT, Stanford, the Data Provenance Initiative , and other institutions, aggregated 24,521 conservations across 85,633 conversational turns (that is, the user prompt and corresponding AI response) from seven real-world datasets collected by previous research. These conversations came from 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025. But these conversations are a drop in the proverbial bucket compared to the data that the big labs themselves have access to. The latest Anthropic Economic AI Index , for example, is based on analysis of 1 million Claude conversations; OpenAI’s report on how people are using ChatGPT analyzed 1.5 million conversations. An Anthropic representative said that their published research reflects their research teams’ specific questions and interests, and the importance of supporting external independent research. OpenAI did not respond to requests for comment. Additionally, the fact that the AI Observatory’s dataset draws from voluntarily-provided sources means that it’s likely underrepresenting sensitive uses, which people may be less likely to share. Thus, the researchers caution that its findings are not indicative of all AI use. The Observatory’s work, though, broadens access for the research community. AI companies don’t typically share their chat data for analysis, which means their reports tend to focus on the findings that paint their companies in the best light, independent researchers, like Reuel and Widder, say. “When we want to ask, for example: is Anthropic's general-purpose AI system…used mostly for good, or mostly for bad…we don't have a way of answering that question because that information is proprietary,” explains Widder, the assistant professor at UT-Austin’s School of Information. The AI Observatory’s data will be available to researchers for analysis, and the team hopes to expand its datasets over time. Ideally, Reuel says, the AI companies would share their data with independent researchers—in ways that protect user privacy, of course. But as it currently stands anyone making decisions based on AI usage data risks “completely operating in the wild and making these really consequential decisions without knowing what's actually happening beyond those company narratives,” says Reuel. Deep Dive Artificial intelligence A startup claims it broke through a bottleneck that’s holding back LLMs Subquadratic has now shared more details about its new model. But some are still skeptical. By Will Douglas Heaven archive page A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven archive page Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. By Will Douglas Heaven archive page Claude Science is Anthropic’s newest flagship product The company is doubling down on AI for science. By Grace Huckins archive page Stay connected Illustration b