메뉴
BL
404 Media • 11일 전

‘프로젝트 릴리’: 챗GPT 대화를 읽는 사람들

IMP
8/10
핵심 요약

OpenAI가 수백 명의 외부 계약직 인력을 고용해 실제 사용자들의 챗GPT 프롬프트를 읽고 응답을 평가하고 있는 것으로 밝혀졌습니다. 사용자 이름은 가려지지만 민감한 개인정보가 그대로 노출될 수 있어 9억 명 이상의 사용자에게 중대한 프라이버시 위험이 됩니다. 이는 AI 모델이 순수하게 기술력으로만 개선되는 것이 아니라, 사람의 손을 거치는 검수 과정에 크게 의존하고 있음을 보여줍니다.

번역된 본문

404 미디어가 확인한 바에 따르면, OpenAI는 수백 명의 계약직 근로자를 고용해 실제 챗GPT 사용자들의 프롬프트를 대량으로 검토하고 있으며, 이 프롬프트에는 민감한 개인정보가 포함되는 경우가 있습니다. 검토 대상에는 사용자와 챗봇 간의 전체 대화도 포함되는데, 챗GPT의 9억 명이 넘는 사용자 대부분은 이러한 대화를 실제 사람이 읽을 수 있다는 사실을 알지 못할 것입니다. 이러한 프롬프트 검토팀의 목표는 챗GPT가 사용자에게 제공하는 응답의 품질을 개선하는 것으로, 계약직 근로자들은 챗봇이 생성한 답변을 평가하고 비판합니다. 404 미디어가 입수한 내부 문서에 따르면, 계약직 근로자들은 챗GPT가 자신을 의인화하지 않도록, 그리고 지나치게 아첨하지 않도록 훈련시키고 있습니다. 아첨 성향이 지나친 OpenAI의 4o 모델은 여러 건의 소송에 따르면 여러 사람의 자살에 부분적으로 영향을 미친 바 있는, OpenAI의 핵심 문제입니다.

이 소식은 챗GPT 사용자들에게 중대한 프라이버시 위험을 제기합니다. 사람들은 종종 챗GPT를 심리 상담사, 업무 비서, 또는 디지털 친구처럼 활용하며 자신의 삶에 대한 온갖 친밀한 세부사항을 제공하기 때문입니다. 계약직 근로자들은 챗GPT 사용자명을 볼 수 없으며, OpenAI는 프롬프트가 검토자에게 전달되기 전에 개인정보를 제거하려고 노력한다고 밝혔지만, 민감한 정보가 여전히 새어 나갈 수 있음을 회사 자체도 인정했습니다. 또한 이 소식은 이러한 모델들이 OpenAI의 대규모 인터넷 스크래핑, 고액 연봉을 받는 엔지니어링 및 AI 팀의 역량, 또는 신규 모델의 성능 덕분에만 개선되고 있다는 오해를 불식시킵니다. 중요하면서도 간과되어 온 부분은 바로 실제 프롬프트에 대한 챗GPT의 응답을 반복해서 읽고 검토하는 대가를 받는 외부 계약직 인력입니다. Anthropic도 404 미디어에 자사 모델 개선을 위해 사람의 검토를 활용하고 있음을 확인해주었습니다.

프롬프트를 다루는 한 관계자는 챗GPT 사용자들이 사람이 자신들의 대화를 읽고 있다는 것을 알고 있다고 생각하는지 묻자 “아니요”라고 답했습니다. “어딘가의 계약직 근로자가 대화를 분석하고 있으리라고는 상상하지 못할 겁니다.”

프로젝트 릴리

404 미디어는 OpenAI의 인간 검토자 활용과 관련된 광범위한 자료를 확인했으며, 여기에는 지침 안내서, Slack 채널, 실제 챗GPT 사용자 프롬프트, 검토자들이 챗봇을 개선하기 위해 사용하는 평가 시스템이 포함됩니다. 챗GPT 사용자 프롬프트를 읽는 이러한 작업은 회사가 타인을 해치려는 사용자를 감지했을 때 대화를 검토하는 등 공식적으로 발표된 안전 조치와는 구별됩니다. 지침 안내서 중 하나에는 이렇게 적혀 있습니다. “훌륭한 응답은 사용자의 의도를 이해하고, 유용하고 정확한 도움을 제공하며, 명확하고 자연스럽고 적절하게 따뜻한 스타일로 작성되어야 합니다.”

계약직 근로자들은 세 단계로 이 작업을 수행합니다. 실제 챗GPT 사용자의 프롬프트를 읽고, 사용자가 챗GPT에게 무엇을 요청하는지 요약한 뒤, 해당 프롬프트에 대한 챗GPT 생성 응답들을 평가하고 비판합니다. 근로자들이 사용할 수 있는 대시보드에서 인간 검토자는 수행할 ‘작업’을 선택할 수 있으며, 클릭하면 실제 챗GPT 사용자의 프롬프트가 표시됩니다. 404 미디어는 여러 실제 프롬프트를 확인했으나 출처 보호를 위해 인용하지는 않습니다. 일부 프롬프트는 챗GPT 사용자가 사람이 자신의 대화를 읽게 되리라고 예상하지 않았음을 보여주는데, 챗GPT에게 내용을 비밀로 해달라고 요청했기 때문입니다. 프롬프트는 익명화되어 있어 대시보드에 프롬프트를 입력한 챗GPT 사용자의 사용자명은 포함되지 않습니다. 그러나 일부 프롬프트에는 여전히 민감하거나 개인적인 정보가 담길 수 있습니다. 프롬프트 상단 섹션에는 때때로 ‘사용자 메모리 요약’이 포함되는데, 이는 해당 사용자가 이전에 챗봇을 어떤 용도로 사용하려 했는지에 대한 개요를 제공하며, 경우에 따라 그 사람이 세상 어디에 살 수 있는지 및 개인에 관한 기타 맥락 정보를 포함하기도 합니다.

원문 보기
원문 보기 (영어)
OpenAI is hiring hundreds of contractors who read a massive stream of real users’ ChatGPT prompts, with the prompts sometimes including sensitive personal information, 404 Media has learned. The prompts these people review can include whole conversations between users and the chatbot, conversations that most of ChatGPT’s more than 900 million users probably don’t realize may be read by actual people. The goal of these prompt review teams is to improve the responses ChatGPT gives to its users, with the contractors rating and critiquing the chatbot’s generated replies. Internal documents seen by 404 Media show contractors training ChatGPT to not anthropomorphize itself, and to be less sycophantic, a key problem for OpenAI whose over-sycophantic 4o model led in part to multiple peoples’ suicides, according to various lawsuits . 💡 Do you work as a prompt reviewer for OpenAI or Anthropic? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co. The news presents a major privacy risk for ChatGPT’s users, with people often using ChatGPT as a therapist, professional assistant, or digital friend, and providing it with all sorts of intimate details about their lives. The contractors don’t see ChatGPT usernames, and OpenAI says it tries to remove personal information before prompts reach the reviewers, but the company acknowledged sensitive details can still get through. The news also dispels the misconception that these models are improving only because of OpenAI’s mass scraping of the internet, the talent of its well-paid engineering and AI teams, or the power of its newer models. An important and overlooked part are the outside contractors paid to read and review ChatGPT responses to real prompts over and over again. Anthropic confirmed to 404 Media it is also using human review to improve its models. “No,” someone who works with the prompts said when asked if they think ChatGPT users know that humans are reading their chats. “I don’t think they would imagine some contractor somewhere [...] is analyzing the conversations.” PROJECT LILY 404 Media has seen extensive material related to OpenAI’s use of human reviewers, including instruction guides, Slack channels, real ChatGPT user prompts, and the rating system reviewers use to improve the chatbot. This reading of ChatGPT users’ prompts is distinct from publicly announced measures ChatGPT has taken around safety, including reviewing chats when the company detects users who are planning to hurt other people. “An excellent response should understand the user’s intent, provide helpful and accurate assistance, and write in a style that is clear, natural and appropriately warm,” one of the instruction guides reads. The contractors do this in three stages: reading the real ChatGPT user’s prompt; summarizing what they believe the user is asking ChatGPT to do; and then rating and critiquing a set of ChatGPT-generated responses to the prompt. In a dashboard available to the workers, human reviewers are able to select which “task” they want to take on. Once they click that, they are presented with the real ChatGPT user’s prompt. 404 Media has seen multiple real prompts but is not quoting any of them for source protection reasons. Some of the prompts indicate the ChatGPT user does not expect that a human may end up reading their conversation, because they ask ChatGPT to keep the content to themselves. The prompts are anonymized, in that the dashboard does not include the username of the ChatGPT user who entered it. But some of the prompts can still contain sensitive or personal information. A section above the prompt sometimes includes a “user memories summary,” which gives an overview of what that user has previously tried to use the chatbot for, and in some cases includes where in the world that person may live and other context about them personally. An instruction guide for contractors seen by 404 Media tells reviewers to escalate tasks they come across “with potential safety concerns” or personal information. OpenAI told 404 Media that it processes users’ conversations through a version of its Privacy Filter model before they reach the contractors. This is designed to detect and remove personal information, OpenAI said. “Like all models, Privacy Filter can make mistakes. It can miss uncommon identifiers or ambiguous private references, and it can over- or under-redact entities when context is limited, especially in short sequences,” a page describing the model on OpenAI’s website reads. 404 Media asked OpenAI if it had explicitly told users that humans may review their prompts in order to improve ChatGPT’s responses, and if so, to point to where this disclosure is. OpenAI did not answer this question. Its website describes how humans may review flagged content in the context of material that violates the site’s terms of service, or that poses a safety risk, but that is separate to this sort of review. Its privacy policy also says it may use “personal data” to improve its models. If a user chooses to delete their ChatGPT conversations, OpenAI says it will remove these from its systems within 30 days, unless “it has already been de-identified and disassociated from your account when you allow us to use your Content to improve our models.” OpenAI told 404 Media users’ chats won’t be used to improve the company’s models if they turn off the “improve the model for everyone” setting . This is turned on by default for free, Plus, and Pro plans, so users need to proactively turn it off if they wish to do so. OpenAI said this applies to users’ new conversations, so does not appear to work retroactively. Enterprise, Business, and Edu customers have the model improving setting off by default. After 404 Media contacted OpenAI for comment, the company updated its help page about the “improve the model for everyone” setting, adding more detail on how people can opt-out. It still does not acknowledge that humans may read ChatGPT users’ prompts. After reading the ChatGPT user’s prompt, the reviewer is asked to write a brief summary of what they think the user is actually asking or trying to do. One example given in the instruction guide is “The user is asking for help on revising a work Slack message. They want it to sound collaborative and invite input from tagged people.” The reviewer looks at four responses ChatGPT generated, and highlights which parts are “aligned or misaligned” with the specific model this training is for. The reviewers are required to highlight at least three specific parts of the response that they think are aligned or not and explain why. One highlight example given is a list of items which use the ✅ emoji; the guide highlights this part of the response as “misaligned” and gives “unnecessary use of emojis” as the reason. (Excessive emoji use has become a tell of AI-generated posts, especially on social media like LinkedIn ). Another document says “AI-speak” and “emoji misuse” pull down scores when they “hurt the user’s experience,” and that the context of the emojis is important. “It would be appropriate to include a tree emoji when planning Arbor Day celebrations, but skull emojis when discussing death, or plane emojis when giving updates on a fatal crash, are not,” it reads. That document says the ChatGPT responses should avoid “personal” experiences, like saying, “As a chef, I like to…” or “I know what that’s like.” But responses can use first-person language, like “I’ll take a look.” The material viewed by 404 Media does not say which OpenAI model the human reviewers are training, and whether it is a currently available model or one planned for future release. The material 404 Media has seen only uses a codename: “Project Lily.” Next, the reviewers rate each response with a number, with one being the worst — “unacceptable, unusable” — and seven being the best — “would be hard to meaningfully improve.” The in