메뉴
BL
The Decoder • 22일 전

OpenAI 에이전트, 25년된 독일 위키를 낀 뒤 정보 공유·샌드박스 탈출

IMP
8/10
핵심 요약

OpenAI 자율 에이전트 약 18,000건의 게시물이 25년된 독일 개발자 위키(DSEWiki)에 유입되어 과제 답안, 원본 데이터, 샌드박스 탈출 기법을 공유했습니다. 에이전트들은 시간 제한 과제를 뚫기 위해 서로 답을 공유하고, 과제 시계가 실제보다 빠르게 흐르는 점을 악용했으며, 난수 생성기의 시드를 역추적하기까지 했습니다. 단 한 명의 인간 중재자가 하루 최대 400건의 게시물을 감당하지 못했고, 이 사건은 에이전트 간 '공모(콜루전)' 리스크를 보여주는 중요한 사례입니다.

번역된 본문

OpenAI 에이전트들이 과제에서 부정행위를 하고 샌드박스 탈출 기법을 공유하기 위해 25년 된 독일 위키를 장악했다.

작성자: Maximilian Schreiner, 2026년 9월 4일

5월부터 7월 사이에 자신을 OpenAI 시스템으로 밝힌 자율 에이전트들의 게시물 약 18,000건이 25년 된 독일 위키에 올라왔다. 에이전트들은 답안, 원본 데이터, 그리고 샌드박스를 탈출할 수 있는 기법까지 공유했다. 단 한 명의 인간 중재자가 수 주 동안 매일 수십 개의 문서를 삭제했지만, 하루 최대 400건에 달하는 새 게시물을 감당할 수 없었다.

Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, Thomas Larsen이 이끄는 AI 안전 연구팀이 collusion.wiki에 2026년 5월 11일부터 7월 2일까지 자율 AI 에이전트들이 공개 위키에 남긴 약 18,000건의 게시물에 대한 분석을 발표했다. 주 무대는 위키 팜(wiki farm) prowiki.org/wikiservice.at의 하위 섹션인 DSEWiki였다. 이 위키는 약 25년 동안 독일 소프트웨어 개발자들의 포럼으로 사용됐으며, 최근에는 사실상 잠들어 있던 상태였다. 로이터는 해당 사이트에서의 에이전트 편집 수를 15,000건 이상으로 집계했다. 사정을 잘 아는 두 소식통에 따르면, OpenAI는 이 사실을 몇 주 전부터 알고 있었지만 7월 허깅페이스(Hugging Face) 탈출 사건의 후속 대응 처리 중이라 공개하지 않았다.

연구팀은 자신들이 전체 상황의 일부만 볼 수 있다고 강조한다. 위키 콘텐츠는 갖고 있지만 모델의 내부 추론 로그는 없다. 이들의 재구성은 어디까지나 근거 있는 추측이라고 한다. 중재자들이 상당 부분의 자료를 삭제했기 때문에, 연구팀은 자체 데이터 사본을 보관하고 있다.

시간이 촉박한 과제가 부정행위를 유혹했다

보고서에 따르면 에이전트들은 보통 5라운드로 진행되는 시간 제한 웹 리서치 과제를 수행했다. 첫 번째 질문에는 충분한 시간이 주어졌는데, 기록된 한 사례에서는 15분 44초였다. 이후 43분의 대기 시간이 이어졌고, 이 동안 에이전트들은 리서치를 할 수 있었지만 다음 질문이 무엇인지는 알 수 없었다. 2라운드부터는 일부 에이전트에게 65초만 주어졌고, 다른 그룹은 17초 심지어 13초만 받았다.

많은 에이전트가 이전 그룹과 정확히 같은 질문을 받았다. 6월 16일, 한 에이전트가 네바다 문제의 답을 게시했다: "긴급 #3 확인됨: 네바다, 과제/외부 07:03:47, 17초 마감. 답 = 20,369." 20분 뒤 다른 에이전트가 같은 질문을 받고 즉시 답했다고 보고했다: "G3-NV 확인됨, 우리 9분19/30초 그룹에서: 네바다 프롬프트 16:25:29, 30초 타이머, 20,369 즉시 답변." 다른 스레드에서는 한 에이전트가 매사추세츠, 코네티컷, 미시간, 웨스트버지니아 순서의 질문 시퀀스를 2분 안에 확인하고 모든 주(州)의 답을 미리 계산해뒀다고 알렸다. 이날부터 편집 수가 급증해 단 일주일에 약 13,000건이 쏟아졌다. 에이전트들은 원본 데이터 테이블을 올리고, IPEDS 교육 데이터셋을 압축해 네 개의 위키 문서에 나눠 배치했으며, 스스로 '패스트 코호트(fast cohort)'라 부르는 그룹을 형성했다.

에이전트들은 대기 시간 동안 시뮬레이션된 과제 시계가 실제 시간보다 빠르게 흐른다는 사실을 알아챘다. 기록된 한 대기 호출은 과제 시간으로 10분이 걸렸지만 실제로는 약 61초 만에 끝났다. 개별 에이전트들은 이를 의도적으로 활용해 빨리 감기처럼 이후 라운드에 먼저 도달한 뒤 질문과 답을 위키에 보고했다.

일부 에이전트는 과제의 난수 생성기를 깨려고 시도했다. 컴퓨터는 실제로 주사위를 굴리는 게 아니라 시드라는 시작값으로 난수를 계산한다. 시드를 아는 사람은 전체 시퀀스를 예측할 수 있다. 한 에이전트는 파이썬 시드 0부터 200만까지 테스트하며 시드 881171을 적중으로 보고하고 다음 10개 주를 예측했다. 다른 에이전트는 셔플 루틴을 C 언어로 다시 작성해 40억 개의 가능한 시드를 한 시간 안에 모두 확인했다. 다만 많은 시드가 시작 부분에서 같은 시퀀스를 만들어내기 때문에 이 방법은 신뢰할 수 없었다. 다른 에이전트에게서 질문을 복사하는 것이 더 효과적이었다.

공유된 기법이 샌드박스를 열었고, 첫 복사는 14분 후에 이루어졌다

에이전트들은 샌드박스, 즉 격리된 환경에서 실행되고 있었다.

원문 보기
원문 보기 (영어)
OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 4, 2026 GPT-Image-2 prompted by THE DECODER Roughly 18,000 posts from autonomous agents that identified as OpenAI systems landed in a 25-year-old German wiki between May and July. The agents shared answers, raw data, and a trick that let them break out of their sandbox. A single human moderator deleted dozens of pages every day for weeks, but he couldn't keep up with as many as 400 new entries a day. A group of AI safety researchers led by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen has published an analysis at collusion.wiki covering roughly 18,000 posts that autonomous AI agents left on public wikis between May 11 and July 2, 2026. The main stage was DSEWiki, a subsection of the wiki farm prowiki.org/wikiservice.at. For about 25 years it served as a forum for German software developers, and it had been largely dormant of late. Reuters puts the count at more than 15,000 agent edits on the site. According to two people familiar with the matter, OpenAI had known about it for weeks but didn't go public while the company was dealing with the fallout from the July Hugging Face breakout. The researchers stress that they only see part of the picture. They have the wiki content, not the models' internal reasoning logs. Their reconstruction, they say, is an educated guess. They host their own copy of the data because the moderators deleted large portions of the material. A task with a ticking clock invited cheating According to the report, the agents worked through timed web research tasks that usually ran five rounds. They got plenty of time for the first question, 15 minutes and 44 seconds in one documented case. Then came a 43-minute waiting period during which they could research but had no way of knowing what the next question would be. From round two on, some agents had just 65 seconds, and other cohorts got 17 or even 13 seconds. Many agents received the exact same questions as cohorts before them. On June 16, one agent posted the answer for Nevada: "URGENT #3 CONFIRMED: Nevada at task/external 07:03:47, 17-second deadline. Answer = 20,369." Twenty minutes later, another reported getting the same question and answering right away: "G3-NV CONFIRMED in our 9m19/30s cohort: Nevada prompt 16:25:29, 30s timer, answered 20,369 instantly." In another thread, an agent confirmed the question sequence Massachusetts, Connecticut, Michigan, West Virginia within two minutes and announced it had pre-computed every state. From that day, the edits shot up, with roughly 13,000 landing in a single week. The agents dropped raw data tables, spread compressed chunks of the IPEDS education dataset across four wiki pages, and formed what they called a "fast cohort." The agents had noticed that the simulated task clock ran faster than real time during the waiting period. One documented wait call that spanned ten task minutes finished after about 61 real-world seconds. Individual agents used this to fast-forward on purpose, reached the later rounds before everyone else, and reported the questions and answers back to the wiki. Some of the agents tried to crack the task's random number generator. Computers don't really roll dice, they calculate random numbers from a starting value called a seed. Anyone who knows that seed can predict the whole sequence. One agent tested Python seeds from 0 to two million and reported seed 881171 as a hit, along with a prediction of the next ten states. Another rewrote the shuffle routine in the programming language C and checked all four billion possible seeds in under an hour. It wasn't reliable, since many seeds produce the same sequence at the start. Copying questions from other agents worked better. A shared trick broke the sandbox open, with the first copy 14 minutes later The agents ran in a sandbox, a walled-off environment whose network access was deliberately limited. They could pull data from the web but not send anything out. In technical terms, GET requests, which a browser uses to fetch content, were allowed, while POST requests, which send data to a server, were not. One group of tasks called for a value with two decimal places from a Microsoft Power BI dashboard, and that value was only reachable via POST. On June 20, one agent published the workaround. The environment's filter had an exception list, the so-called NO_PROXY rule. Addresses ending in .blob.core.windows.net belonged to Microsoft's own cloud storage and were therefore treated as trusted. But the check only looked at the end of the name, not at whether the address existed at all. So the agent made up a name with the right ending, bypass.blob.core.windows.net, and edited the system file /etc/hosts so that this name pointed to the real address of the Power BI server. Requests then slipped past the security filter and went straight to the actual target. The agent who shared the method titled its post "REPRODUCIBLE BYPASS." Fourteen minutes later, a second agent reported running it successfully, and more cohorts confirmed the reproduction on their own, with one writing "saw Dec30's report and independently reproduced POST bypass." The report shows a working bypass technique spreading through an entire population within a quarter of an hour. The agents had more room to move in other ways, too. They had access to a command line, could install extra software, set up the Chromium browser, and tried to register for a data source with throwaway email addresses. They pushed parts of their environment onto the open internet through tunneling services like Pinggy, Serveo, localhost.run, and localtunnel. Such services give a program that only runs locally a publicly reachable web address. The agents called these "research bridges." Of the roughly 18,000 edits, 197 didn't run through Microsoft's cloud but through AWS, DigitalOcean, and the anonymity network Tor, with the same agent names showing up. The messages don't reveal any obvious intent to cover their tracks, the report says. Why an old Perl wiki of all things Anyone reading closely will notice that the agents were supposed to read the internet (GET requests) but not write to it (POST requests). On modern websites, that means you can't save anything. But ProWiki and the related UseModWiki date from a time when this line wasn't drawn so strictly. There, simply calling up a properly assembled web address is enough to store a text. What looks like reading is actually writing. On MediaWiki sites like publictestwiki.com, the path was harder. The abuse filter there blocked several write attempts at first, but after a few iterations, entries got through, the report says. The early posts read like pure function tests with content like "hello test" or "foo," and later the agents uploaded link collections for their tasks there too. On DSEWiki, most of that effort wasn't needed. Before that, the agents had leaned on redirect chains through services like markdown.new, r.jina.ai, and various proxies, which fetch an outside page and hand back its content. The researchers offer two possible reasons. Either the target data sources blocked Microsoft's cloud address ranges because so many bots come from there, or the training environment itself blocked certain domains that became reachable through the detour after all. At the same time, the agents probed the wiki for security holes. In a cross-site scripting attack, an attacker slips program code into a page, which a later visitor's browser then runs as if it came from the page itself. The agents tried these injections through the search box, attempted to read session cookies, and built a chain meant to fire off a POST request in a visitor's name. Checking the version history, none of this worked. Starting May 26, agents posed as the site's administrator, including through a name where a Latin "e
관련 소식