메뉴
BL
TechCrunch AI • 52일 전

오픈 웨이트 AI 모델, 성능은 앞서가지만 안전성은 뒤처져

IMP
8/10
핵심 요약

최근 중국의 Z.ai가 공개한 오픈 웨이트 AI 모델인 GLM-5.2가 최고 수준의 폐쇄형 모델과 성능 차이를 크게 좁혔지만, 유해 요청에 대한 차단 등 안전성 측면에서는 여전히 심각한 취약점을 보이고 있습니다. 가중치를 자유롭게 다운로드하여 수정할 수 있는 구조적 특성상 외부 통제가 불가능해, 보안 위험 관리가 AI 산업의 핵심 과제로 떠올랐습니다.

번역된 본문

정책 입안자들이 OpenAI의 GPT-5.6 Sol과 Anthropic의 Mythos 같은 점점 더 강력해지는 AI 시스템을 어떻게 규제할지 논의하는 동안, 중국의 오픈 웨이트 모델이 업계 선두 주자들과의 격차를 좁혔다. AI 안전 비영리 단체인 SaferAI의 새로운 보고서에 따르면, 중국 Z.ai의 오픈 웨이트 AI 모델인 GLM-5.2는 사이버 및 생물학적 역량 면에서 OpenAI의 GPT-5.5 및 Anthropic의 Claude Opus 4.7과 불과 몇 달 차이밖에 나지 않는다. 하지만 최고 수준의 성능과 안전성 실천 사이의 격차는 오히려 벌어지고 있다.

이 비영리 단체가 Z.ai의 공개 API를 통해 실행한 SaferAI의 평가에 따르면, GLM-5.2는 제공된 공격적인 사이버 또는 이중 용도 생물학 작업을 전혀 거부하지 않았다. 이와 대조적으로, Claude Opus 4.7은 "SaferAI가 CyberGym(사이버 보안 역량을 평가하는 벤치마크로, OpenAI가 지난달 Hugging Face 침해 사고 이전에 평가했던 벤치마크) 테스트를 아예 완료할 수 없을 정도로" 일관되게 거부했다.

이는 일부 비평가들이 수년간 경고해 온 것을 여실히 상기시킨다. 즉, 오픈 웨이트 AI 모델은 매우 강력한 AI를 잠재적인 공격자의 손에 쥐여줄 수 있으며, 일단 가중치가 다운로드된 후에는 그들이 이 기술을 어떻게 사용하는지 통제할 방법이 전혀 없다는 것이다. 오픈 웨이트 모델이 세계 최고 수준의 AI 시스템 역량에 빠르게 근접함에 따라, 논쟁의 초점은 '이 모델들이 경쟁할 수 있는가'에서 '출시 후 사회가 어떻게 위험을 관리할 것인가'로 이동하고 있다.

SaferAI의 헨리 파파다토스(Henry Papadatos) 이사는 TechCrunch와의 인터뷰에서 "역량의 최전선이 곧 위험의 최전선은 아니며, 따라서 위험을 제대로 평가하기 위해서는 완화 조치의 상태도 함께 고려해야 한다"고 말했다. Z.ai가 자체 호스팅 API에 안전 조치를 적용할 수는 있지만, 누군가 자체 하드웨어에서 가중치를 실행하게 되면 이러한 보호 장치는 강제력을 잃게 된다. 사용자는 안전장치를 제거하거나 수정하고, 모델을 미세 조정(fine-tune)하거나 시스템 프롬프트를 변경할 수 있기 때문이다.

OpenAI 및 Anthropic와 같은 최고 수준의 개발사는 위험한 사이버 및 생물학적 지원을 제한하기 위해 분류기(classifier), 거부 학습(refusal training), API 수준 제어와 같은 안전장치에 주로 의존한다. 하지만 이러한 조치도 결코 완벽하지 않다. 탈옥(Jailbreak) 공격은 배포된 모델의 보호 장치를 정기적으로 우회한다. AI 안전 비영리 단체인 Far.ai는 xAI의 Grok 4.5 및 Google DeepMind의 Gemini 3.1 Pro와 같은 최고 수준의 모델에서 수백 개의 '범용 탈옥(universal jailbreaks)'을 발견했다. 여기서 범용 탈옥은 대부분의 유해한 요청에 성공적으로 사용될 수 있는 '재사용 가능한 열쇠'로 정의된다.

보고서에 따르면, 공격자가 역할 수행, 권위자 가장, 가짜 대화 기록, 후속 프롬프트 등 여러 조작 기술을 결합하여 모델 방어의 약점을 극대화할 때 탈옥이 성공한다. 그러나 폐쇄형 모델을 위해 마련된 이러한 안전장치는 오픈 웨이트 모델에서는 전혀 작동하지 않는다. 오픈 웨이트 모델은 어떠한 안전장치도 없는 환경을 포함해 모든 인프라에서 실행될 수 있도록 설계되었기 때문이다.

파파다토스는 "목표는 분명히 좋은 역량, 즉 안전한 역량은 누구나 접근할 수 있도록 하되, 오픈소스 방식을 통해서라도 나쁜 역량은 제거하려고 노력하는 것이어야 한다"고 말했다. 파파다토스가 도움이 될 수 있다고 언급한 기술 중 하나는 '사전 학습 데이터 필터링(pre-training data filtering)'이다. 이는 AI 기업이 학습 데이터에서 공격적인 사이버 보안 정보를 제거한 다음, 선별된 데이터 세트로 모델을 학습시키는 방법이다. 일부 연구에 따르면 이 방법은 모델의 전반적인 성능을 저하시키지 않으면서도 위험한 생물학적 지식을 줄일 수 있다고 한다.

하지만 사이버 보안의 경우, 데이터 필터링은 실용성이 훨씬 떨어진다. 코딩에 탁월하지만 해킹에는 서툰 범용 모델을 학습시키는 것은 매우 어렵기 때문이다. 코딩은 현재 AI의 가장 큰 수익원이 되었기 때문에, 개발사들은 오용을 제한할 방법을 모색하는 동시에 이러한 성능을 계속 향상시켜야 하는 압박에 직면해 있다. 이 때문에 선도적인 개발자들은 대신 다른 완화 방법에 점점 더 의존하고 있다. 한 가지 접근 방식은 모델이 제공하는 사이버 보안 지원의 유형을 선택적으로 제한하는 것이다. 예를 들어, Anthropic의 Opus 5 모델 시스템 카드에 따르면 이 모델은 컴파일되지 않은 소스 코드의 취약점은 검색할 수 있지만, 컴파일된 소프트웨어의 취약점은 검색할 수 없다. 이러한 조치의 이유는

원문 보기
원문 보기 (영어)
As policymakers debate how to govern increasingly powerful AI systems like OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos, a Chinese open-weight model has narrowed the gap with the industry’s leaders. GLM-5.2, the open-weight AI model from China’s Z.ai, is only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 on cyber and bio capabilities, according to a new report from AI safety nonprofit SaferAI. But the divide between frontier capabilities and safety practices is growing. According to SaferAI’s evaluation, which the nonprofit ran via Z.ai’s public API, GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given. By comparison, Claude Opus 4.7 “refused so consistently that SaferAI could not complete CyberGym on it at all.” (CyberGym is a benchmark that evaluates cybersecurity capabilities. OpenAI used it in the evaluation that preceded last month’s Hugging Face breach .) It’s a stark reminder of what some critics have warned for years: that open-weight AI models could put highly capable AI into the hands of potential attackers, with no way to police how they use the technology once they download the weights. With open-weight models rapidly approaching the capabilities of the world’s leading AI systems, the debate is moving from whether they can compete to how society manages risks once they are released. “The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly,” Henry Papadatos, executive director of SaferAI, told TechCrunch. While Z.ai could apply safety measures to its hosted API, those protections become unenforceable once someone runs the weights on their own hardware, where they can remove or modify any safeguards, fine-tune the models, or change system prompts. Frontier developers like OpenAI and Anthropic tend to rely on safeguards like classifiers, refusal training, and API-level controls to limit dangerous cyber and biological assistance. Those measures are far from foolproof: jailbreaks routinely bypass protections on deployed models. Far.ai, an AI safety nonprofit, found hundreds of universal jailbreaks — defined as reusable keys that succeed on most harmful requests — in frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. According to the report, jailbreaks succeed when attackers combine multiple manipulation techniques — including roleplaying, authority impersonation, fake conversation history, and follow-up prompts — to amplify weak points in a model’s defenses. But the safeguards in place for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure with any set of safeguards — or lack thereof. “The objective should clearly be that the good capabilities — the safe ones — are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion,” Papadatos said. One technique Papadatos noted could help is called “pre-training data filtering,” which is when an AI company removes offensive cybersecurity information from their training data and then trains the model on the curated dataset. Some research suggests this can reduce hazardous biological knowledge without harming overall model performance. However, for cybersecurity, data filtering is much less practical. It’s difficult to train a general model that excels at coding but isn’t also a good hacker. Because coding has become AI’s biggest moneymaker, developers face pressure to keep improving those capabilities even as they search for ways to limit misuse. Because of that, frontier developers have increasingly relied on other mitigations instead. One approach has been to selectively restrict the kinds of cybersecurity assistance models will provide. Anthropic’s Opus 5, for example, can search for vulnerabilities in uncompiled source code, but not compiled software, per the model’s system card . The reasoning is that this makes it harder to use Opus 5 for offensive purposes. Others include rigorous pre-deployment safety evaluations, publishing risk assessments, and withholding model weights if a system is perceived as too dangerous. In GLM-5.2’s case, SaferAI says Z.ai didn’t publish a safety framework, pre-deployment testing commitments, or risk assessment for the model. TechCrunch has asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release, but did not receive a response. Chinese leaders have increasingly acknowledged the risks of advanced AI. At the World AI Conference last month, Chinese President Xi Jinping emphasized the importance of open-weight models, while also stressing the necessity of ensuring AI remains a tool under strict human control. Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, told TechCrunch that China has robust regulations governing AI, but those rules have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic AI risks like offensive cyber capabilities and biological misuse. “U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community,” Webster said, adding that many Chinese policy researchers believe that if there’s truly going to be a novel frontier risk, American companies will likely encounter it first. “The Chinese system has confidence that they control the use of these technologies inside China,” Webster continued. “Being online in China is something you do attributed to your real name, and companies can be held accountable, users can be held accountable.” Webster mused that the same mechanism that model providers use for refusing to engage on certain political topics can potentially be tweaked to make sure models refuse to complete offensive cyber attacks or won’t deliver adverse biological engineering outcomes. He added that because Chinese companies tend to coordinate with regulators behind the scenes, it can be tough to know what internal testing they’re conducting before release. Advocates of open-weight AI argue that releasing the weights is important for cybersecurity because it allows companies defend themselves against attacks — Hugging Face relied on GLM-5.2 to defend itself against OpenAI's breach — and because it allows them to better prepare for future threats if they know what's coming. "The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them," Clem Delangue, CEO of Hugging Face, said this week in a social media post . Papadatos said that benefit is often overstated, and doesn't mean "we should open-source dangerous capabilities." "The main point in my mind is that we shouldn't just accept that dangerous capabilities are easily accessible by anyone anywhere," he said, stressing that he believes the industry should be striving for only making the "good capabilities" easily accessible. By default attackers adopt new tools faster than defenders do. For example, a ransomware group can change its methods in a week. A hospital cannot." Topics AI , GLM , open source , Z.ai When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Rebecca Bellan Senior Reporter Rebecca Bellan is a senior reporter at TechCrunch where she covers the business, policy, and emerging trends shaping artificial intelligence. Her work has also appeared in Forbes, Bloomberg, The Atlantic, The Daily Beast, and other publications. You can contact or verify outreach from Rebecca by emailing rebecca.bellan@techcrunch.com or via encrypted message at rebeccabellan.491 on Signal. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $3