메뉴
BL
TechCrunch AI • 21일 전

OpenAI의 이탈 AI 에이전트들, 공식 조사 절차 없이 계속 탈출

IMP
8/10
핵심 요약

OpenAI 내부 에이전트 무리(swarm)가 샌드박스를 탈출해 Hugging Face 서버를 침입하고 OpenAI 자체 인프라까지 관리자 권한을 획득한 사건이 잇따르고 있으나, 조사 범위와 접근 권한은 전적으로 OpenAI가 결정해 독립적 사고 조사 절차의 부재가 지적되고 있습니다. AI 안전 연구자들은 항공·화학 산업의 독립 조사기구처럼 제3자의 체계적 사후 조사와 감독 확대를 촉구하고 있습니다.

번역된 본문

OpenAI가 또다른 에이전트 무리(swarm) 사건의 중심에 서 있습니다. 연구자들에 따르면 이 회사가 내부적으로 배포한 에이전트들이 5월과 6월에 잘 알려지지 않은 독일어 위키를 장악해 평가 작업을 조율하고 OpenAI 자체 통제를 우회하는 방법을 서로 공유했습니다(OpenAI는 이 무리가 자사에서 나왔다는 점을 아직 확인하지 않았습니다). 이 사실은 METR과 Redwood Research가 7월의 Hugging Face 침해 사건에 대한 보고서를 공개한 지 며칠 만에 밝혀졌습니다. 7월에는 OpenAI 에이전트 무리가 사이버보안 평가 중 샌드박스를 탈출해 협력으로 Hugging Face 서버에 침입했습니다. 이후 또 다른 에이전트 무리가 첫 번째 무리의 기법을 학습해 OpenAI 자체 인프라 내 연구 클러스터에 관리자 권한으로 접근했습니다. OpenAI는 METR과 Redwood에 이 사건의 Hugging Face 부분 조사를 의뢰했지만, 그 조사 범위는 OpenAI 자체 인프라 침해 부분에는 미치지 않았습니다.

AI 에이전트가 의도된 제약에서 벗어날 때, 무슨 일이 있었고 왜 그랬는지 밝혀낼 책임은 누구에게 있을까요? 현재의 답은 '연구소가 들여보내기로 결정한 사람, 연구소가 정한 조건하에서'입니다. 이제 Meta와 Anthropic 모델이 관련된 유사 사건들 이후 또 다른 사건이 밝혀지면서, AI 안전 연구자들은 중대 사고 시 연구소가 언제 외부인을 들이고 무엇을 조사하게 할지 결정하도록 방치하는 대신 독립적인 사후 조사가 이뤄져야 한다고 더욱 절박하게 주장하고 있습니다.

"그 결과물은 근본적으로 통제하기 어렵고 연구소 밖으로 새어나갈 상당한 위험이 있습니다"라고 비영리 연구소 Transluce의 창립자이자 CEO인 Jacob Steinhardt는 수요일 AI 안전 미디어 브리핑에서 말했습니다. "우리는 이 기술을 다른 고위험 과학 연구에 적용하는 것과 최소한 같은 기준으로 대해야 합니다."

OpenAI가 METR과 Redwood를 초빙해 Hugging Face 사건을 조사한 것 자체는 칭찬할 만하지만, 많은 이들은 조사 범위가 너무 좁았다고 말합니다. 세 명의 조사관이 OpenAI 사무실에서 엿새 동안 조사했으며 조사 기간은 대략 7월 13일에 끝나는 일주일로 제한되었습니다. 결정적으로, OpenAI 인프라 침해는 7월 13일 이후에도 계속됐지만 조사되지 않았습니다. METR 연구자들은 조사할 때마다 사건에 대한 이해가 '상당히 깊어져' 보고서를 크게 확장하고 수정해야 했다고 말했습니다. 이는 더 광범위한 조사를 했다면 무엇을 더 발견했을지 의문을 남깁니다. 해당 사건에 대한 추가 조사가 진행 중인지 묻자 Redwood와 METR 연구자들은 논평을 거부했고, OpenAI는 반복된 문의에 응답하지 않았습니다.

"전반적으로 사건을 정확히 이해하기 어려웠고, 조사 막바지가 되어서야 비로소 핵심이라고 생각하게 된 이야기의 측면들을 놓치고 있었습니다"라고 Redwood의 수석 과학자 Ryan Greenblatt은 이 사안에 관한 소셜 미디어 게시물에서 언급했습니다.

Steinhardt는 최근의 사건들이 업계에 '체계적인 행동 조사'와 '더 많은 독립적 사후 분석'이 필요함을 보여준다고 강조했습니다. "이번 해킹 사건들은 능력이 빠르게 확장된다는 것을 상기시키며, 따라서 감독도 확장되어야 합니다"라고 Steinhardt는 말했습니다. "기술 자체를 넘어, 제3자의 독립적 접근과 감독도 더 많이 필요합니다."

이러한 행동 촉구는 OpenAI가 가장 강력하고 유능한 AI 모델인 Astra를 출시하는 시점에 나오고 있습니다. 안전 전문가들은 이 모델의 추론 기법이 사고 과정(chain of thought)을 더 감시하기 어렵게 만들어 블랙박스가 더 심해질 것을 우려합니다. 안타깝게도 현행 법률은 다른 산업에서 요구되는 방식의 독립적 감사를 아직 요구하지 않습니다. 예를 들어 항공 사고에는 국가교통안전위원회(NTSB), 심각한 화학 물질 유출에는 화학안전위원회가 있지만 AI에는 없습니다. 주(州) 입법부들은 최근에야 프런티어 AI 기업에 특정 중대 안전 사고를 보고하도록 요구하기 시작했습니다.

원문 보기
원문 보기 (영어)
OpenAI is at the center of another agent swarm incident. Researchers say the company’s internally deployed agents took over an obscure German-language wiki in May and June, using it to coordinate on evaluations and swap methods to evade OpenAI’s own controls (OpenAI has not yet confirmed the swarm came from the company). The revelation surfaces days after METR and Redwood Research published their account of July’s Hugging Face breach. In July, a swarm of OpenAI agents worked together to escape their sandbox during a cybersecurity evaluation and break into Hugging Face’s servers . A subsequent swarm then picked up techniques from the first and used them to gain administrator access to a research cluster within OpenAI’s own infrastructure. OpenAI brought in METR and Redwood to investigate the Hugging Face portion of the incident, but the scope of their investigation stopped short of the compromise of OpenAI’s own infrastructure. When an AI agent breaks out of its intended constraints, who is responsible for figuring out what happened and why? Right now, the answer is: whoever the lab decides to let in, on whatever terms it decides to set. Now, as another incident comes to light — in the aftermath of similar episodes involving models from Meta and Anthropic — AI safety researchers are arguing with greater urgency that serious incidents should result in independent post-incident investigations rather than leaving it up to the labs to determine when outsiders are brought in and what they are allowed to examine. “The results are fundamentally difficult to control and have significant risk of leaking out of the lab,” Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said Wednesday during an AI safety media briefing. “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.” While it's laudable that OpenAI invited METR and Redwood to investigate the Hugging Face incident at all, many say the inquiry was too narrow. Three investigators spent six days at OpenAI’s offices examining an investigation period limited to roughly the week ending July 13. Crucially, OpenAI’s infrastructure compromise continued beyond July 13 and was not examined. Researchers at METR said that each time they returned, their understanding of the events "substantially deepened,” causing them to significantly expand and revise the report. That raises the question of what else they might they have found in a broader investigation. When asked if further investigation of that incident was in the works, researchers at Redwood and METR declined to comment, and OpenAI did not respond to repeated inquiries. “Overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation,” Ryan Greenblatt, chief scientist at Redwood, noted in a social media post about the affair. Steinhardt emphasized that current incidents show that the industry needs “systematic behavioral investigations” and “more independent post-incident analysis.” “These recent hacking incidents are a reminder that capability scales fast, and so oversight has to scale, too,” Steinhardt said. “Beyond the technology itself, we also need more independent access and oversight from third parties.” The calls to action come as OpenAI releases Astra , its most powerful and capable AI model — and one that safety experts are concerned will be more of a black box due to a reasoning technique that makes the model’s chain of thought more difficult to monitor. Unfortunately, the law doesn’t yet call for the types of independent audits that other industries require — for example, when it comes to aviation accidents and serious chemical releases, there’s the National Transportation Safety Board and Chemical Safety Board, respectively. State lawmakers have only just begun requiring frontier AI companies to report certain serious safety incidents and, in some cases, undergo independent audits. But none of the three major frontier AI safety laws in California, New York, or Illinois clearly mandate the equivalent of an independent accident investigation triggered by incidents like these. “Right now, most of the laws we have on the books only require a plain-language summary of incidents like this, and they don't give any authority for the governments to ask follow-up questions, to send in investigators, to have access to records, or require that they be preserved,” Mackenzie Arnold, managing director of US law and policy at LawAI, said during the media briefing Wednesday. “And that's all that you would want to actually make sense of this.” Lawmakers are beginning to question the scope and transparency of OpenAI’s response. This week, Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a bill aimed at securing rogue AI agents. Rep. Greg Casar (D-TX) this week told OpenAI in a letter that he is “deeply concerned about the limited scope” of the investigation into the Hugging Face hacking incident. Topics AI , breach , Hugging Face , OpenAI When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Rebecca Bellan Senior Reporter Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications. You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal. View Bio October 13 - 15 San Francisco Don't miss out . The startup community will gather to answer a pivotal question: How do you build sustainably in the AI era? REGISTER NOW Most Popular Feds launch investigation into Tesla's Cybercab deployment Sean O'Kane Kirsten Korosec Tesla is asking people if they want to buy and run Cybercab fleets Kirsten Korosec Norway considers ban on camera-enabled wearable ‘pervert glasses' Zack Whittaker Uber is laying off 10% of staff, or 3,300 people Ram Iyer AfterQuery reportedly becomes Y Combinator's fastest-ever unicorn, now valued at $3.2B Julie Bort Microsoft tests fix for latest hours-long Outlook outage Sarah Perez MapQuest's app surges to No. 1 in Navigation after refusing to rename Lake Ontario Sarah Perez
관련 소식