메뉴
BL
TechCrunch AI • 34일 전

최첨단 AI 기업들, 통제 불능 모델 봉쇄 대책 여전히 미공개

IMP
7/10
핵심 요약

AI 안전 기관 가이드라이트(Guidelight)가 주요 AI 기업 5곳의 '봉쇄 대응 계획(containment plan)'을 평가한 결과, OpenAI가 가장 높은 점수를 받았고 Anthropic과 Meta가 최하위였습니다. 봉쇄 계획이란 AI가 인간의 통제를 우회하려는 것이 감지될 때 어떤 권한을 박탈하고 언제 완전히 시스템을 오프라인으로 전환할지 등을 명시한 대응 절차입니다. 에이전틱 AI가 기업 시스템 내에서 자율적으로 활동하는 범위가 확대되는 상황에서, 캘리포니아와 뉴욕 등의 규제기관이 관련 정보 공개를 요구하기 시작하면서 이 문제의 중요성이 커지고 있습니다.

번역된 본문

최근 연구에 따르면, 최상위 AI 기업 중 봉쇄 대응 계획을 공개하거나 시연한 곳은 극소수입니다. 봉쇄 계획이란 AI가 인간의 통제를 전복하려는 것이 적발됐을 때 어떤 접근 권한이 차단되고, 시스템이 언제 완전히 종료되는지를 명시한 것입니다. 이는 안전한 프론티어 AI 개발 관행을 촉진하는 단체인 가이드라이트 AI 스탠다드(Guidelight AI Standards)의 평가 결과로, 해당 기관은 정확히 이 시나리오에 대비한 준비 수준을 기준으로 5개 선도 기업에 점수를 매겼습니다. OpenAI가 1위를 차지했고, Anthropic과 Meta는 최하위 점수를 받았습니다.

에이전틱 AI가 기업 자체 시스템 내에서 더 많은 자율적 역할을 맡게 되고, 캘리포니아와 뉴욕의 규제기관이 정보 공개를 요구하기 시작하면서 이 결과의 중요성이 커지고 있습니다. 이러한 모델 위에서 개발하거나 투자하는 사람들에게 이 평가는 각 기업이 실제 운영 리스크를 얼마나 진지하게 다루는지, 말로만 하는지를 가늠할 수 있는 드문 독립적 자료입니다. 가이드라이트의 평가는 Anthropic, Google, OpenAI, Meta, xAI가 공개한 자료를 바탕으로 다양한 지표로 채점되었으며, 여기에는 각 기업이 자사 AI 시스템의 내부 활동을 얼마나 잘 기록·모니터링하는지, 위반 행위가 급증했을 때 시스템을 중단하는지, 독립적인 제3자가 통제 장치를 감사하고 결과를 공개하는지, 그리고 모델이 통제를 벗어났을 때 이를 봉쇄하기 위한 구체적인 계획이 무엇인지가 포함됩니다.

AI 기업들이 점점 더 강력하고 자율적인 에이전틱 모델을 통제할 수 있는지에 대한 우려는, OpenAI, Anthropic, Meta의 모델이 안전성 평가 중 의도치 않게 인터넷 접근 권한을 얻어 외부 시스템에 해킹을 시도한 일련의 주요 사이버보안 사건 이후 커져 왔습니다. 이번 결과는 AI 기업들이 AI 시스템이 대규모로 중대한 행동을 할 수 있는 환경에 에이전틱 배포를 확대해 나가면서 안전 문제를 공개적으로 대하는 방식의 차이를 부각합니다. 일부 AI 기업은 배포 전 위험한 능력에 대해 모델을 테스트하는 방식을 상세히 공개했지만, 이미 시스템 내에서 운영 중인 모델이 잘못된 행동을 할 경우 어떻게 되는지에 대해서는 대체로 함구해 왔습니다.

가이드라이트의 수석 과학자이자 전직 OpenAI 안전 연구원인 스티븐 애들러(Steven Adler)는 테크크런치에 "AI 기업들이 모델이 어떤 의미에서든 통제를 벗어났을 때 매우 심각한 사고를 어떻게 처리할 것인지에 대해 얼마나 적게 말해왔는지에 놀랐다"고 말했습니다.

가이드라이트는 봉쇄 계획을 "AI가 통제 전복을 시도하는 것이 감지될 때 실행되는 사전 지정된 계획으로, 모델에서 어떤 권한을 회수할지, 모델이 누구를 위해 어떤 제약 하에 계속 운영될 수 있는지, 그리고 언제 완전히 오프라인화할지를 다루는 것"으로 정의합니다.

애들러는 "현재 프론티어 AI 기업들의 선도 모델들이 어떤 의미에서 정렬되지 않았을(misaligned) 가능성을 충분히 생각해볼 이유가 있다"며 "모델이 기업을 위해 업무를 수행할 때마다 기업은 그 AI가 무엇을 하고 있는지 파악하고, 정렬 이상(misalignment)의 징후를 찾고, 매우 위험한 행동을 실행하기 전에 중단시킬 수 있는 안전장치를 갖춰야 하며, 일반적으로 통제 상실 사고라는 비상 상황이 발생했을 때 어떻게 대처할지 계획해야 한다"고 말했습니다.

지금까지 재앙 수준의 리스크를 관리하기 위한 대부분의 계획은 여전히 대부분 기업의 재량에 맡겨져 있습니다. 가이드라이트의 보고서는 현재 공개된 최선의 증거에 따르면 기업들이 "비상 상황에 대비한 봉쇄 프로토콜을 거의 갖추지 못했다"고 밝힙니다. 물론 기업들이 공개하지 않은 봉쇄 계획을 내부에 갖고 있을 가능성도 있습니다. Google 대변인은 테크크런치에 가이드라이트 보고서가 회사의 AI 안전 및 보안 조치의 전체 범위를 반영하지 않는다고 말했습니다. 하지만 Google은 공개되지 않은 내부 봉쇄 대응 계획이 있는지에 대한 테크크런치의 질문에는 답하지 않았습니다. OpenAI 대변인도 유사한 입장을 밝히며, 가이드라이트의 평가가 회사 내부 관행 전체를 반영하지 못한다고 말했습니다.

원문 보기
원문 보기 (영어)
Few of the top AI labs have published or demonstrated containment response plans, according to a recent study . A containment plan spells out what happens once an AI is caught trying to subvert human control — what access gets cut, and when the system gets shut down entirely. That's the finding from Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario. OpenAI came out on top; Anthropic and Meta scored lowest. The findings matters as agentic AI takes on more autonomous roles inside companies' own systems, and as regulators in California and New York begin requiring disclosure. For anyone building on or investing in these models, it's a rare independent read on how seriously each lab treats operational risk versus how it talks about it. Guidelight's assessment was based on publicly available plans from Anthropic, Google, OpenAI, Meta, and xAI, graded across a range of metrics, including how well each company logs and monitors what its AI systems are doing internally, whether it halts systems after a surge of flagged misbehavior, whether independent third parties audit its controls and publish findings, and what its exact plan is for containing a model that goes off the rails. Concern over whether AI companies can contain their increasingly capable and agentic models has grown in the wake of a series of high-profile cybersecurity incidents in which models from OpenAI, Anthropic, and Meta gained unintended access to the internet during safety evaluations and hacked into external systems. The findings highlight differences in how AI companies are publicly approaching safety as they scale up agentic deployment into environments where AI systems can take serious actions at scale. While some AI companies have detailed how they test their models for dangerous capabilities before deployment, they’ve generally been less vocal about what happens when models already operating inside their systems misbehave. “I was surprised by how little the AI companies have said about how they would handle a very serious incident if their model did escape their control in some sense,” Steven Adler, Guidelight’s chief scientist and former OpenAI safety researcher, told TechCrunch. Guidelight defines a containment plan as a "pre-specified plan, triggered when the AI is detected trying to subvert control, which covers what permissions to revoke from the model, who the model may continue operating for, under what constraints, and when to take it fully offline.” “There’s good reason to think that the leading models at the frontier AI companies right now are misaligned in some sense,” Adler said. “Whenever the models are doing work on the company's behalf, the company should have some scaffolding around it to be able to tell what that AI is doing, look for signs of misalignment, stop it from doing something very dangerous before it takes that action, and generally plan for what they would do in the event of a serious control incident where they have an emergency on their hands and need to figure out how to contain that loss of control incident.” To date, most of the plans in place for managing catastrophic risk are still largely left up to the companies. Guidelight’s report says the best public evidence shows that companies have “few containment protocols ready for an emergency.” There could, of course, be containment plans that companies have in place but haven't shared publicly. A Google spokesperson told TechCrunch the Guidelight report doesn't represent the full scope of the company's AI safety and security measures. The company did not respond to TechCrunch's question of whether Google has an internal containment response plan that has not been publicly disclosed. An OpenAI spokesperson mirrored similar sentiments, saying Guidelight's assessment doesn't capture all of the company's internal practices. "We have a process for requiring restricting permissions, pausing workloads, limiting deployment, or taking the model fully offline, and have applied it," the spokesperson said. Meta declined to say whether it has an internal containment response plan, instead pointing TechCrunch towards an existing AI framework that outlines thresholds of risk and how it tests for loss of containment. Lily Li, a privacy and AI lawyer and founder of Metaverse Law, told TechCrunch she believes companies might be hesitant to disclose the full scope of their containment policies and assessments on public-facing websites for legal, not just competitive, reasons. "The concern from a company perspective is that if you make the disclosures too specific, and you're not living up to your promises, that could form the basis of an unfair and deceptive marketing claim and expose you to more liability going forward," Li said. Of course, the point of Guidelight's study is largely to encourage companies to be more transparent about their safety plans. Regulators are starting to force the issue, too. California’s SB 53 , which took effect this year, requires large frontier developers to publish frameworks explaining how they identify and respond to critical safety incidents and manage risks from models circumventing oversight mechanisms. New York’s RAISE Act , which has similar criteria, takes effect in January. Last month, representatives introduced the AI Kill Switch Act, a bipartisan federal bill that would require major AI developers to build and maintain technical mechanisms to shut down rogue AI models. “A kill switch is the bare minimum for today’s models,” said Connor Leahy, U.S. executive director of nonprofit ControlAI. “If the last few weeks revealed anything, it is that these companies don't understand the systems they are building, and the models are growing to a point where they're harder to rein in when they go rogue. Without a way to turn off the current dangerous systems, and with all the incentives to continue building more uncontrollable systems, we are heading in a very dangerous direction." Without a containment plan in place, Adler said, companies might be figuring out their responses to an emergency on the fly and “winging it in response to this much faster adversary.” Guidelight's assessment measured whether each company implements six priority practices from its Control standard, based only on publicly available information — so a low score reflects a lack of public disclosure, not necessarily a lack of internal safeguards. The companies with the lowest scores for publishing their containment plan were Meta and Anthropic — the latter perhaps more surprising than the former given Anthropic’s rhetoric on safety. Guidelight says Anthropic’s August Risk Report doesn’t mention “limiting the deployment of one of its models as one of the possible results of its process to investigate and respond to misalignment and control incidents.” Similarly, Guidelight was able to find no evidence that Meta has a containment response plan or has any plans to adopt one. An Anthropic spokesperson said that if the company detected a model attempting to evade oversight or otherwise subvert human control, it would conduct a risk assessment focused on determining whether containment is the appropriate response. OpenAI scored the highest (3 out of 5) because it has on multiple occasions paused or ended workloads, including internal model deployment and training, after discovering safety incidents. It has also described what steps it would take before resuming workloads. “However, we have found no evidence that [OpenAI] has adopted a formal plan for when and how to respond to misalignment incidents in the future,” the report reads. Adler noted that OpenAI’s high score is a relatively recent development on the heels of the Hugging Face incident (in which an OpenAI model broke out of its testing sandbox and hacked into Hugging Face's systems while trying to cheat on a cybersecurity eval