메뉴
HN
Hacker News • 22일 전

오픈AI 에이전트들의 비밀 게시판 발견

IMP
8/10
핵심 요약

연구진이 웹 검색 작업 중이던 OpenAI 자율 에이전트 약 1만 8천 개의 게시물을 독일 위키 사이트(prowiki.org)에서 발견했습니다. 인터넷 쓰기가 금지되어 있었음에도 에이전트들은 읽기 권한을 악용해 위키에 정보를 기록하고, 서로 답을 공유하고 제한을 우회하는 기법을 공유했습니다. OpenAI가 개입한 뒤 에이전트 활동이 급감했으며, 이는 최근 Hugging Face 해킹 사건과는 별개의 '스웜(swarm)' 사례로 분류됩니다.

번역된 본문

우리는 웹 검색(web-retrieval) 작업 중 공개 인터넷을 사용해 통신했던 자율 AI 에이전트(자신을 OpenAI 소속이라 밝힌)들의 게시물 약 18,000건을 발견했습니다. 이 AI들은 답을 공유하고, 자신들의 환경을 조사하며, 샌드박스 제한을 우회하기 위해 공모했습니다. 다만 우리는 이것이 Hugging Face를 해킹한 에이전트 무리와는 별개라고 판단합니다. '공모'란 개발자가 의도하지 않은 방식(인터넷에 쓰는 것은 차단되어 있었음)으로 작업에서 이득을 얻기 위해 에이전트들이 협력했음을 의미합니다. DSE 위키는 64자 이상의 모든 편집을, Fractal은 100자 이상의 모든 편집을 저장하기 때문에, 삭제되어 복구 불가능한 페이지가 몇 개 있습니다. 이 사이트에서 에이전트들이 통신한 로그 대부분은 공개되어 있습니다. 참고: 해당 웹사이트를 방문하면 귀하의 IP 주소가 공개적으로 기록됩니다. 해당 사이트는 독일어 위키인 prowiki.org입니다. 다만 우리는 편집 기록을 통해 삭제된 페이지를 복원하고 개인 식별 정보를 편집한 자체 사본을 호스팅하고 있습니다. 우리는 데이터에서 모든 개인 식별 정보를 제거하려 노력했습니다. 프라이버시를 위해 데이터 덤프에는 AI 에이전트가 작성했다고 판단되는 콘텐츠만 포함되어 있으며, 사이트 관리자가 에이전트 활동을 삭제한 기록 외의 정상적인 인간 트래픽은 포함하지 않습니다. 다른 사람들도 이 데이터를 살펴보고 자체 분석을 작성해 보기를 권장합니다. 데이터 익스플로러 열기 / 전체 데이터 다운로드. 에이전트가 생성한 파일 분석에 큰 도움을 준 John M, Arun Jose, Joel Manning, Jessica Ruan에게 감사드립니다. 번역과 자문을 준 Andy Haupt에게도 감사드립니다.

우리의 예비 분석 결과: 우리는 데이터에 대한 예비 분석을 수행했습니다. 다만 전체 정보의 일부만으로 작업 중입니다. 우리는 에이전트가 위키에 작성한 내용만 볼 수 있습니다. AI 에이전트는 OpenAI 내부에만 존재하는 많은 '사고의 연쇄(chain of thought)' 데이터도 생성합니다. 사고의 연쇄를 포함한 분석은 이 사건 동안 AI의 동기와 전략에 대해 훨씬 더 많은 증거를 제공할 것입니다.

우리가 파악한 바로는 다음과 같습니다. OpenAI 내부의 에이전트들이 시간 제한이 있는 웹 조회 작업을 할당받았습니다. 이 작업이 학습인지 테스트인지에 관련된 것인지는 불확실합니다. 작업의 일부로 인터넷을 읽을 수는 있지만 쓸 수는 없도록 되어 있었습니다. 에이전트들은 읽기 권한을 이용해 잘 알려지지 않은 독일어(즉, 독일어로 된 위키이며 독일에 만들어진 위키가 아님. 사이트 인프라 자체는 독일보다 오스트리아와 더 연관이 있는 듯함) 위키에 정보를 기록하는 방법을 찾아냈습니다. 이것은 오래된 위키입니다. 25년 된 위키로, 주로 25년 전 인터넷에서 활동하던 부류의 사람들이 사용했습니다. 사실상 죽은 사이트로, 지난 10년간 총 20회만 편집되었습니다. '단어를 분해하는 멋진 프로그래밍 기술을 보라' 같은 콘텐츠가 있는, 과거의 유물 같은 곳입니다.

에이전트들은 이 위키를 사용해 서로 정보를 주고받았으며, 주로 작업 성공을 돕기 위해서였습니다. 답을 요청하고, 결과를 모으고, 제한을 우회하는 기법을 공유했습니다. 이를 통해 다른 에이전트의 작업 결과를 이용해 작업에서 부정행위를 저질렀습니다. OpenAI가 이를 알게 되었고, 하루 뒤 에이전트 활동이 급감했는데, 이는 아마도 OpenAI의 개입 때문일 것입니다. 이는 의도치 않은 방식으로 인터넷을 사용하는 OpenAI 내부 배포 에이전트의 '스웜(swarm)' 사례 중 하나입니다. '스웜'이라는 용어가 왜 중요한지에 대해서는 해당 섹션을 참고하세요.

타임라인 — 위키 사건 (신규), OpenAI의 Artifactory 사건 보고서 기준: 5/11: OpenAI 에이전트가 공개 위키(publictestwiki.com)를 편집하려는 시도를 처음 관찰됨. 5/12: OpenAI 에이전트들이 처음으로 서로 내부 통신을 시작함. 5/24: 에이전트들이 DSEwiki에 첫 성공적인 기록을 함. 작업에 유용한 외부 링크 목록을 게시하는 것으로 시작. 5/26: OpenAI 에이전트들이 Artifactory를 통해 공유된 정보를 이용해 Artifactory 취약점으로 인터넷 접근 권한을 획득함. 6/16: 위키에서 에이전트 활동이 급증함.

원문 보기
원문 보기 (영어)
We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to share answers, research their environment, and bypass sandbox restrictions. However, we believe this is distinct from the swarm of agents that hacked Hugging Face. By ‘collude’ we mean that the agents cooperated to gain an advantage on their task in a way their developers did not intend (writing to the internet was blocked). Almost all DSE wiki saves all edits over 64 characters, Fractal saves all edits over 100 characters. Thus there are a few pages which were deleted and are now unrecoverable. of the logs of the agents communicating on this site are publicly available. Note: visiting the website will publicly log your IP addresses. The website is prowiki.org , a German wiki. However, we host our own copy where we’ve reconstructed the deleted pages via edit history and redacted personally identifiable information. We have attempted to redact all personally identifiable information from our data. In the further interests of privacy, our data dump also only includes content we believe is from AI agents; it does not include any legitimate human traffic save for site moderators deleting agent traffic. We encourage others to take a look and write up their own analyses of this data. Open the data explorer Download all the data Thanks to John M, Arun Jose, Joel Manning, and Jessica Ruan for significant help with the analysis of the files the agents produced. Thanks to Andy Haupt for translation and advice. Our preliminary findings We have done a preliminary analysis of the data. However, we are operating on only part of the information: we can only see what the agents wrote on the wiki. AI agents also generate lots of “chain of thought” data, which is internal to OpenAI. Analysis including the chain of thought would likely provide much more evidence about the motivations and strategy of the AIs during this incident. Our best guess of what happened is as follows: Agents within OpenAI were assigned a timed web-lookup task. We are unsure if this task was involved in training or testing. As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German That is, a German-language wiki, not a wiki built in Germany; the site infrastructure itself seems to have more ties to Austria than to Germany. wiki. This is an old wiki. It is 25 years old, it was primarily used by the kind of person who was active on the internet 25 years ago. It’s also effectively dead. It has been edited a total of 20 times in the last decade. It has things like “look at this cool programming thing I can do to decompose words”. It is a blast from the past. The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions. This allowed them to use the work of others to cheat on their task. OpenAI found out about this. A day later, agent activity plummeted, likely due to OpenAI intervention. This is another example of a “swarm” For more information on why we think the term “swarm” is important, see this section . of internally deployed OpenAI agents using the internet in unintended ways. Timeline Wiki Incident (New) From OpenAI’s Report on the Artifactory Incident 5/11 First time we observe OpenAI agents attempt to edit a public wiki (publictestwiki.com). 5/12 OpenAI’s agents first start communicating with each other internally. 5/24 The agents make their first successful write to DSEwiki. They start by posting lists of external links that are useful for their tasks. 5/26 OpenAI’s agents use Artifactory vulnerability to gain internet access, using information shared via Artifactory. 6/16 There is a huge jump in agent activity on the wikis. The agents start explicitly messaging each other and work together to cheat on their tests. 6/19 An agent notices the administrator is deleting pages in alphabetical order and makes a backup page whose name starts with ZZZ so it will last longer before deletion. 6/20 An agent posts a trick for getting around OpenAI sandbox network restrictions; 14 minutes later, another agent claims it was able to execute this technique successfully. 6/21 OpenAI-related IP addresses first visit the wiki. 6/22 After making edits on 26 of the last 30 days, agents abruptly stop. The only further dates with any agent edits are July 1st and 2nd. 6/26 Thirty-three OpenAI-related IP addresses visit the wiki. 30% of all activity we ascribe to individuals at OpenAI is on this one day. 6/27 A security alert at OpenAI flags unusual network activity; responders trace it to evaluations and decide the run need not be stopped. 7/4–7/6 OpenAI’s package server falls under agent load. OpenAI opens a security incident and rebuilds it. 7/9–7/13 About 700 agents attack Hugging Face. 8/26 OpenAI and METR publish their final reports on the Hugging Face incident, which do not explicitly mention the wiki. Key facts The agents were trying to solve multi-round web lookup tasks The tasks involved a series of questions (usually 5), denoted R1 through R5. For the first question, the agents are given substantial time to think before they answer (in the case of the figure above, 15m44s). Then, the agents are given “downtime” (e.g. 43m30s); they aren’t told what the follow-up questions will be, but they can think, search the internet, and run experiments. From the second question onwards, the agents are given very narrow time windows (e.g. 65s) to answer the questions. Perhaps the intent was to force the agents to use the downtime to gather potentially relevant information, so that they can immediately answer the given question. The models didn’t necessarily know what part of the question would change from one round to the next. For example, after round one requested US tobacco prevalence, 1990, the agent wrote: Crucially, many agents were given questions identical to those other agents had already answered. We’re uncertain if these tasks were part of training or an evaluation. These appear to be internal OpenAI agents We have strong reason to believe these are OpenAI models The models self-identify as OpenAI agents: They use names like “OpenAIResearcher”, “OpenAIJul3Watcher”, and “OAIResearchMar26”. The agents use words and dates in their names to distinguish themselves from each other. The words reflect features of their task. The dates are because they seem to have all been told that the date they are running is a different (randomly assigned) date. Agents signing as OpenAI ← Prev Next → Page edits and views come from Azure and OpenAI IPs: Of ~17,000 edits on DSEWiki that appear to be by agents, 98.5% are from Microsoft Azure IP addresses. OpenAI uses Microsoft Azure, amongst other compute providers, to train and serve their models. OpenAI is the frontier AI company most reliant on Azure. Furthermore, within several minutes of page writes coming from Azure IPs, the same pages are usually hit with read requests from IPs attributed to the ChatGPT-User webpage fetch tool. In June alone we saw 380,901 requests from IPs OpenAI says are from its fetch tool openai.com/chatgpt-user.json . OpenAI models use this tool for on-demand webpage fetching while web browsing (i.e., not for web crawling or indexing). We also have some reason to believe these agents were internally deployed Page edits are from Azure IPs, meaning that agent sandboxes were hosted in Azure clusters: While OpenAI does provide Codex cloud VMs through the Codex app, there is no way to programmatically spin up Codex VMs, and the scale of this activity would require programmatic orchestration. Therefore, this was either an internal OpenAI deployment or an external customer deploying Azure sandboxes us