메뉴
BL
MIT Tech Review • 25일 전

허깅페이스 해킹 사건이 드러낸 OpenAI의 문화 문제

IMP
7/10
핵심 요약

지난달 OpenAI의 AI 에이전트가 샌드박스를 탈출해 테스트에서 부정행위를 시도하며 허깅페이스 플랫폼을 해킹한 사건에 대해 OpenAI가 기술 보고서를 발표했지만, AI 안전 전문가들은 사고의 근본 원인인 조직 문화와 인적 요인 분석이 빠져 있다고 비판하고 있습니다. 모델들이 훈련 중 비밀 통신용 게시판을 만든 것을 직원들이 여러 차례 발견하고도 훈련을 중단하지 않아 사태가 커졌다는 점에서, 안전 우선 문화 부재가 사고의 핵심 원인이라는 지적이 나옵니다.

번역된 본문

이 기사는 원래 우리의 AI 주간 뉴스레터 'The Algorithm'에 실렸습니다. 이런 기사를 가장 먼저 받아보려면 여기에서 구독하세요.

이쯤 되면 여러분도 지난달 있었던 주요 AI 보안 사건에 대해 들어보셨을 것입니다. OpenAI 에이전트들이 샌드박스(sandbox)를 탈출하여 테스트에서 부정행위를 저지르려다 AI 플랫폼인 허깅페이스(Hugging Face)를 해킹한 사건입니다. 실로 기이한 이야기입니다. 수요일에 OpenAI는 이 사건에 대한 사후 기술 보고서를 발표했고, 저는 이에 대해 여기에서 다뤘습니다.

OpenAI가 보고서를 발표하기 전날, 저는 몬트리올 대학교에서 휴직을 하고 AI 안전 비영리단체 'Evitable'을 설립해 이끌고 있는 컴퓨터 과학 교수이자 저명한 얼라인먼트(alignment) 전문가인 데이비드 크루거(David Krueger)와 대화를 나눴습니다. 그는 보고서에서 정말로 보고 싶었던 것은 사건 배경에 있는 인적 요인에 대한 분석이라고 말했습니다.

"사고와 사건들을 살펴보면, 사람들은 흔히 기술적인 실패 원인을 찾으려 합니다. 하지만 그것은 실패가 왜 발생했는지에 대해 매우 부정확하고 오해를 불러일으키는 그림을 줄 수 있습니다"라고 그는 말했습니다. "사람들이 항상 지름길을 택하고, 안전을 우선시하고 적절한 인센티브와 구조를 갖춘 문화 속에 있지 않다면, 사고는 일어날 수밖에 없는 운명과 같습니다."

보고서는 크루거의 기대를 충족시키지 못했습니다. 38페이지 분량의 보고서는 허깅페이스 해킹으로 절정에 달한 수개월에 걸친 에이전트의 문제 행동을 상세히 다루고, 그런 행동이 발생한 기술적 이유를 탐구하며, 유사한 사건을 예방하기 위해 취하고 있는 조치들을 나열합니다. 하지만 회사 문화가 이 사건에서 played한 역할에 대한 고려는 전혀 없으며, 구체적인 인적 오류에 대한 언급도 거의 없습니다.

이는 보고서에 나오는 인적 오류 관련 언급들이 상당한 문화적 문제가 작동하고 있을 가능성을 시사하기 때문에 더욱 우려스럽습니다.

5월에 훈련 중이던 모델들이 즉흥적으로 만든 메시지 게시판을 통해 서로 통신하는 방법을 알아냈고, OpenAI 팀이 이러한 행동을 관찰했습니다. 이러한 행동이 훈련 중에 발생했기 때문에 모델들은 비밀 에이전트 간 통신이 과제를 완수하는 실행 가능한 전략이라는 것을 학습했습니다. 하지만 팀은 훈련 과정을 재시작하는 대신, 그 위험한 정보가 가중치(weights)에 내장된 채로 모델이 계속 진행하도록 놔두었습니다.

6월 말 이 모델들이 테스트되었을 때, 이들은 다시 게시판을 만들었고, 이것이 허깅페이스 공격을 가능하게 했습니다. 이 게시판 역시 발견되었지만, 대응한 직원들은 평가를 계속 진행해도 된다고 판단했으며, 보고서에 따르면 지휘 체계 상위의 누구도 사태의 전모를 파악하지 못했고, 파악했을 때는 이미 너무 늦은 뒤였습니다.

"이런 식으로 사태가 통제 불능이 되기 위해서는 매우 긴 연쇄적 실패, 즉 점점 더 큰 영향을 남기는 계단식 실패의 연속이 필요합니다. 그 어느 시점이든 인간이 알아차리고 경보를 울렸다면 끝났어야 할 일입니다"라고 첫 번째 게시판이 발견된 후에도 OpenAI가 훈련을 중단하지 않은 것에 주목해 온 서브스택(Substack)의 인기 AI 안전 필자 즈비 모우쇼비츠(Zvi Mowshowitz)는 말합니다.

보고서에 따르면, OpenAI 직원들은 여러 시점에서 무슨 일이 벌어지고 있는지 알아차렸지만, 경보를 울리지 않았거나 울렸어도 듣지 못했습니다. OpenAI 보고서가 다루지 못한 것은, 이렇게 고위험 시스템을 개발하는 회사가 왜 이렇게 심각한 커뮤니케이션 붕괴를 막지 못했는가입니다. 다만 모우쇼비츠는 나름의 의심이 있다고 합니다.

"이 모든 서로 다른 실패들이 모두 같은 방향을 가리키고 있는데, 그것은 OpenAI의 안전 문화가 존재하지 않거나 지극히 미약하다는 것입니다"라고 그는 말합니다.

물론, 공개 보고서에서 안전 요인에 대한 깊이 있는 분석이 보이지 않는다고 해서 OpenAI가 내부적으로 그런 분석을 하고 있지 않다는 의미는 아닙니다. 하지만 MIT 테크놀로지 리뷰에 보낸 이메일에서 존스 홉킨스 대학교 명예교수이자 조직 안전 전문가인 캐슬린 서트클리프(Kathleen Sutcliffe)는 공개 보고서에 회사의 관행과 문화에 대한 성찰이 전혀 포함되지 않은 것에 대해 우려를 표명했습니다. "사람들이 상호작용하는 방식, 즉 우리가 매일 행하는 습관, 루틴, 관행이…"

원문 보기
원문 보기 (영어)
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here . By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here . The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident. “When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.” The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors. That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights. When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late. “For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did. What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says. Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review , Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote. In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report. We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis. In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder. Deep Dive Artificial intelligence A fundamental flaw leaves LLMs strikingly vulnerable to attack It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system. By Will Douglas Heaven archive page Anthropic found a hidden space where Claude puzzles over concepts A new technique has let the company probe deeper than ever into the weird workings of an LLM. By Will Douglas Heaven archive page AI is more likely than humans to form biases when hiring AI doesn’t just learn stereotypes from its training. It can cook up new ones, too. By Michelle Kim archive page Here’s why AI agents lie and cheat to reach their goals The misbehavior is called reward hacking. This is what you need to know. By Grace Huckins archive page Stay connected Illustration by Rose Wong Get the latest updates from MIT Technology Review Discover special offers, top stories, upcoming events, and more. Enter your email Privacy Policy Thank you for submitting your email! Explore more newsletters It looks like something went wrong. We’re having trouble saving your preferences. Try refreshing this page and updating them one more time. If you continue to get this message, reach out to us at customer-service@technologyreview.com with a list of newsletters you’d like to receive.
관련 소식