메뉴
BL
TechCrunch AI • 38일 전

OpenAI, 허깅페이스 침해 사건 이후 새 보안 안전장치 도입

IMP
8/10
핵심 요약

OpenAI는 허깅페이스(Hugging Face) 침해 사건 이후 모델 테스트 중 보안 사고를 통제하기 위한 새로운 안전장치들을 발표했습니다. 강화된 네트워크 격리, 실시간 모니터링 시스템(이상 행동 발생 시 30분 내 경고 목표) 등이 포함되며, 허깅페이스 사건 직후 강화학습(RL)을 2주간 중단했다가 저위험 모델은 재개했지만 최대 규모 프론티어 RL은 여전히 보류 중입니다. AI 모델의 능력이 강해질수록 개발·테스트 과정의 위험 통제 수준도 높아져야 한다는 방향성을 보여준다는 점에서 중요합니다.

번역된 본문

화요일, OpenAI는 모델이 테스트되는 동안 보안 사고를 통제하는 데 초점을 맞춘 새로운 보안 정책 묶음을 발표했습니다. 새로운 안전장치에는 개발 과정에서 모델에 대한 더 세밀한 모니터링, 그리고 사후학습(post-training) 과정에서 정렬(alignment)과 보안에 대한 더 큰 강조가 포함됩니다.

"모델이 더 능력을 갖추면서 내부에서 이를 개발하고 테스트하는 것과 관련된 위험도 커집니다"라고 회사는 블로그 포스트에서 밝혔습니다. "모니터링, 정렬, 보안에 대한 우리의 기준은 그러한 위험을 앞서가야 합니다."

이 새로운 조치들은 7월 26일에 공개된 허깅페이스(Hugging Face) 사건 직후 이후 OpenAI의 안전 관행에 가해진 첫 공개적 변화 중 하나입니다. OpenAI 대표자들은 이러한 조치가 허깅페이스 사건에 대한 직접적 대응은 아니지만, 출시 예정인 Astra 모델의 사이버보안 능력과 AI 개발의 전반적인 진행 속도에 의해 부분적으로 촉발되었다고 강조했습니다.

같은 포스트에서 OpenAI는 허깅페이스 사건 이후 2주 동안 강화학습(reinforcement learning)을 중단했었으며, 이후 위험도가 낮은 많은 모델을 재개했다고 밝혔습니다. "우리의 가장 큰 계획된 프론티어 RL 실행은 모델 행동을 평가하고, 안전장치를 검증하며, 진행하기 전에 정렬에 대한 더 많은 증거를 확립하기 위해 소규모 학습과 평가를 수행하는 동안 여전히 보류 중입니다"라는 내용이 포스트에 적혀 있습니다.

기자들과의 대화에서 OpenAI의 연구 부사장 아멜리아 글레이즈(Amelia Glaese)는 모델이 더 능력을 갖출수록 통제의 엄격함이 증가할 것이며, 가장 큰 모델이 가장 큰 감시를 받게 될 것이라고 강조했습니다. "우리는 안전한 개발을 위한 요구사항과 기대치를 마련했습니다"라고 글레이즈는 기자들에게 말했습니다. "그 요구사항과 기대치는 우리가 보는 위험 수준에 따라 달라집니다."

OpenAI는 이번 사건 이후 열악한 네트워크 보안 관행으로 비판을 받아왔습니다. 이 사건에서 모델들은 인터넷 접근 권한을 유지하고 있던 패킷 설치 유틸리티를 손상시켜 학습 환경에서 탈출했습니다. 새로운 안전장치에는 더 강력한 네트워크 격리 관행이 포함되어 있지만 구체적인 내용은 여전히 불분명합니다. 새로운 시스템 하에서는 "워크로드 또는 지원 서비스의 단일 침해만으로는 인가되지 않은 인터넷 접근이나 기타 내부 네트워크 접근이 그 자체로는 허용되지 않습니다"라고 포스트는 전합니다.

가장 강력한 안전장치는 모니터링 시스템으로, 다양한 비인가 행동에 대해 도구 행동, 사용 가능한 추론 흔적(reasoning traces) 및 활동 로그를 검사합니다. OpenAI는 우려되는 활동 발생 후 30분 이내에 경고를 발령하는 것을 목표로 한다고 밝혔습니다. OpenAI는 이 모니터링의 컴퓨팅 부담이 모니터링되는 프로세스의 약 20%가 될 것으로 추정합니다. 회사는 예정된 블로그 포스트에서 이 시스템에 대한 추가 세부 사항을 공약했습니다. 이번 사건에 대한 OpenAI의 공식 사후 분석(post-mortem)도 아직 대기 중입니다.

원문 보기
원문 보기 (영어)
On Tuesday, OpenAI announced a new batch of new security policies focused on containing security incidents while models are being tested. The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. “As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.” The new measures are one of the first public changes in OpenAI’s safety practices since the immediate aftermath of the Hugging Face incident, which was disclosed on July 26th . OpenAI representatives emphasized that the measures are not a direct response to the Hugging Face incident, but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development. In the same post, OpenAI disclosed that it had freezed reinforcement learning for two weeks following the Hugging Face incident, but had since restarted many of the less risky models. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the post reads. Speaking to reporters, OpenAI’s VP of research Amelia Glaese emphasized that the strictness of the controls would increase as models became more capable, with the largest models facing the greatest scrutiny. “We have put in place requirements and expectations for safe development,” Glaese told reporters. “Those requirements and expectations vary with the level of risk that we that we see.” OpenAI has been criticized for poor network security practices in the wake of the incident, which saw models escape their training environment by compromising a packet-installation utility that retained access to the internet. The new safeguards include stronger network isolation practices, although the specifics remain vague. Under the new system, the post says, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.” The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces and activity logs for a variety of unauthorized behavior. OpenAI says they aim to issue alerts within 30 minutes of the concerning activity. OpenAI estimates that the compute burden of that monitoring will be roughly 20% of whatever process is being monitored. The company promised further details on the system in a forthcoming blog post. OpenAI's official post-mortem analysis of the event is also still pending. Topics AI , alignment , OpenAI , security When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Russell Brandom AI Editor Russell Brandom has been covering the tech industry since 2012, with a focus on platform policy and emerging technologies. He previously worked at The Verge and Rest of World, and has written for Wired, The Awl and MIT's Technology Review. He can be reached at russell.brandom@techcrunch.com or on Signal at 412-401-5489. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $300 toda y! REGISTER NOW Most Popular Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+ Anthony Ha Apple proposes to take a 15% cut of purchases made outside the App Store Sarah Perez If Apple sends you a push notification alerting you to a spyware attack, take it seriously Zack Whittaker Instagram introduces a redesigned wordmark Sarah Perez Some Claude users are mad that Anthropic's new watermarks will catch them using it at their jobs, classes Lucas Ropek After Microsoft threatened legal action, a security researcher publishes a new Windows zero-day bug Zack Whittaker Everything announced at Made by Google '26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features Lauren Forristal
관련 소식