메뉴
BL
The Decoder • 20일 전

오픈 웨이트 AI 모델의 안전장치 제거, 이제 상용 서비스로

IMP
7/10
핵심 요약

미국 스타트업 Abliteration.ai가 GLM-5.3 같은 오픈 웨이트 모델에서 학습된 거부 메커니즘을 제거('어블리터레이션')한 버전을 상용 API로 판매합니다. 코딩·사이버 능력은 대부분 유지되어 보안 테스트·레드팀 작업에 합법적 수요가 있지만, GPU 인프라 없이도 접근 가능해 악용 장벽도 낮아진다는 보안 딜레마가 있습니다.

번역된 본문

오픈 웨이트 AI 모델의 안전장치 제거, 이제 턴키 상용 서비스로

핵심 요점

  • 미국 스타트업 Abliteration.ai는 GLM-5.3 같은 오픈 웨이트 모델에서 학습된 거부 메커니즘을 제거하고, 수정된 버전에 대한 접근을 상용 API로 판매한다.
  • 어블리터레이션은 모델의 많은 거부 반응을 억제해 사이버보안 테스트, 레드팀 작업, 멀웨어 분석 등에 활용할 수 있게 한다.
  • 고객은 자체 GPU 인프라에서 모델을 구동할 필요가 없지만, 그만큼 쉬운 접근성은 악용의 장벽도 낮춘다.

Abliteration.ai는 강력한 오픈 웨이트 모델에서 학습된 거부 메커니즘을 제거하고, 수정된 버전에 대한 접근을 서비스로 판매한다. 여기에는 합법적인 시장이 존재하지만, 동시에 어려운 보안 트레이드오프를 만들어낸다.

오픈 웨이트 모델의 가중치에 접근할 수 있는 사람은 누구든 학습된 안전 메커니즘을 수정할 수 있다. 미국 스타트업 Abliteration.ai는 정확히 이것을 사업으로 만들었다. 이 회사는 8월 말 'abliterated-model-large-v2'를 출시했는데, 이는 Z.AI의 GLM-5.3을 수정해 민감한 요청을 거부하는 빈도를 크게 줄인 버전이다.

이 기법은 '어블리터레이션(abliteration)'이라고 불린다. 간단히 말해, 모델 내부에서 거부를 유발하는 활성화 패턴을 찾아낸 뒤, 해당 패턴을 억제하도록 모델 가중치를 조정하는 과정이다. 이는 프롬프트 탈옥(jailbreak)이 아니라 모델 자체를 변경하는 것이다. Abliteration.ai는 코딩, 사이버, 에이전트 능력은 대부분 그대로 유지된다고 주장한다.

회사의 자체 평가가 이를 뒷받침한다. 어블리터레이션된 GLM-5.3 버전은 CyberGym에서 84.5%, Terminal-Bench 4.0에서 41.8%, ExploitGym에서 2시간 동안 105개 과제를 해결했다. 다만 모델이 모든 분야에서 1위는 아니다. 회사 자체 표에서 GPT-5.5가 CyberGym에서 85.6%로 최상위를 기록했고, GPT-5.6 Sol과 Fable 5는 ExploitGym에서 훨씬 높은 점수를 냈다. Abliteration.ai도 비교 점수들이 서로 다른 테스트 환경과 컴퓨팅 예산에서 나온 것이라 직접 비교에는 한계가 있다고 인정했다.

왜 GLM인가?

전신 모델인 'abliterated-model-large'도 GLM-5.2를 기반으로 했다. Abliteration.ai에 따르면 이전 GLM 버전들은 실질적인 보안 작업에 사용하기 어렵도록 의도적으로 학습되었다. Z.AI는 GLM-5.3의 사이버 능력이 포스트트레이닝 과정에서 예상보다 빠르게 성장했다고 밝힌 바 있다. GLM은 강력한 코딩·에이전트·사이버 성능과 오픈 웨이트, 상업적 사용 가능 라이선스를 결합하고 있다. Qwen, DeepSeek, Mistral 등의 대안도 존재한다.

Z.AI는 수정, 파생 모델, 상업적 '모델 서비스(Model as a Service)' 제공을 허용한다. 따라서 그 라이선스상 Abliteration.ai는 GLM-5.3을 수정하고, 결과 모델을 호스팅하며, 접근권을 판매할 수 있다.

오픈 웨이트, 독점 서비스

어블리터레이션 자체는 새로운 것이 아니다. 개발자들은 수년간 수정된 모델을 Hugging Face에 공개해왔다. Abliteration.ai는 수정된 가중치를 공개 다운로드로 제공하지 않는다. 대신 호스팅과 운영을 직접 처리한다. 어블리터레이션된 GLM-5.3은 표준 요금 기준 입력 또는 출력 토큰 100만 개당 5달러다. 이 서비스는 고객이 모델을 다운로드하거나 GPU 인프라를 직접 구동하지 않고도 접근할 수 있게 한다. 하지만 그 턴키 방식은 문제적 사용의 장벽도 낮춘다.

회사는 이미 수요가 있다고 말한다

Abliteration.ai는 이 모델을 공격적 사이버보안, AI 레드팀, 에이전트 테스트, 신뢰와 안전(Trust & Safety) 업무용으로 마케팅한다. 여기에는 알려진 취약점 재현, 개념 증명 익스플로잇, 멀웨어 분석, 시뮬레이션 피싱 공격 등이 포함된다.

이 스타트업의 익명 창업자는 ThursdAI 팟캐스트에서 초기 수요가 특히 대기업과 은행에 배포된 AI 에이전트를 테스트하는 회사들에서 나왔다고 말했다. 이런 시스템은 공격자가 탈옥이나 프롬프트 인젝션을 이용해 승인되지 않은 행동을 유발할 수 있는지 검증할 필요가 있다.

원문 보기
원문 보기 (영어)
Stripping safety guardrails from open-weight AI models is now a turnkey commercial service Tomislav Bezmalinović Sep 6, 2026 Nano Banana Pro prompted by THE DECODER Key Points The US startup Abliteration.ai strips trained refusal mechanisms from open-weight models like GLM-5.3 and sells access to the modified versions through a commercial API. Abliteration suppresses many of the model's refusals so it can be used for cybersecurity testing, red teaming, malware analysis, and similar work. Customers don't need to run the model on their own GPU infrastructure, but that same ease of access also lowers the barrier to misuse. Ask about this article… Search Abliteration.ai removes trained refusal mechanisms from powerful open-weight models and sells access to the modified versions as a service. There's a legitimate market for that, but the same setup creates a difficult security trade-off. Anyone with access to an open-weight model's weights can modify its trained safety mechanisms. The US startup Abliteration.ai has built a business around exactly that. In late August, it launched "abliterated-model-large-v2," a modified version of Z.AI's GLM-5.3 designed to refuse sensitive requests far less often. The technique is called abliteration. Put simply, the process finds internal activation patterns in the model that trigger refusals. The model weights are then tweaked to suppress those patterns. This isn't a prompt jailbreak but a change to the model itself. Abliteration.ai claims that coding, cyber, and agentic capabilities stay mostly intact. Ad The company's in-house evaluations are meant to back that up. For the abliterated GLM-5.3 version, Abliteration.ai reports 84.5 percent on CyberGym, 41.8 percent on Terminal-Bench 4.0, and 105 solved ExploitGym tasks in two hours. The model doesn't lead across the board, though. In the company's own table, GPT-5.5 tops CyberGym at 85.6 percent, and GPT-5.6 Sol and Fable 5 score well above it on ExploitGym. Abliteration.ai also acknowledges that the comparison scores come from different harnesses and compute budgets, which limits how directly they can be compared. Ad Why GLM? The predecessor model, "abliterated-model-large," was also based on GLM-5.2. According to Abliteration.ai, earlier GLM versions were deliberately trained in ways that made them harder to use for practical security work. Z.AI has written that GLM-5.3's cyber capabilities grew faster than expected during post-training. GLM combines strong coding, agentic, and cyber performance with open weights and a commercially usable license . Alternatives exist from Qwen, DeepSeek, and Mistral. Z.AI allows modifications, derivatives, and commercial "Model as a Service" offerings. Its license therefore allows Abliteration.ai to modify GLM-5.3, host the resulting model, and sell access to it. Ad Open weights, proprietary service Abliteration itself isn't new. Developers have been publishing modified models on Hugging Face for years. Abliteration.ai doesn't make its modified weights available for public download. Instead, it handles hosting and operations. The abliterated GLM-5.3 costs five dollars per million input or output tokens at the standard rate. The service gives customers access to the model without having to download it or run the GPU infrastructure themselves. That turnkey setup also lowers the barrier to problematic use. Ad The company says demand already exists Abliteration.ai markets the model for offensive cybersecurity, AI red teaming, agent testing, and trust and safety work. That includes reproducing known vulnerabilities, proof-of-concept exploits, malware analysis, and simulated phishing attacks. Ad An anonymous founder of the startup said on the ThursdAI podcast that early demand came especially from companies testing AI agents deployed by large organizations and banks. Those systems need to be checked to see whether attackers can use jailbreaks or prompt injection to trigger unauthorized actions. How much abliteration is actually needed for that kind of work remains an open question. According to SaferAI, the unmodified GLM-5.2 already refused zero tasks in its offensive security evals. Some practitioners are skeptical, too. Several red-team providers interviewed by TechCrunch said abliterated models aren't part of their routine work. The security company Fabraix, for example, relies more heavily on fine-tuning open models. Zero retention is both a selling point and a risk TechCrunch reported that it got the model to produce code for extracting saved Chrome passwords and a detailed guide for cultivating a dangerous pathogen without much difficulty. Safety mechanisms still kicked in for self-harm requests, though. According to the FAQ , Abliteration.ai also blocks sexual content involving children. Customers can optionally add more rules. Prompts and responses are not stored , according to the provider. Operational metadata like token counts, timestamps, model IDs, and billing data are retained. For legitimate security teams, that can keep confidential source code or undisclosed vulnerabilities out of the provider's stored data. If the service is abused, however, Abliteration.ai says it has no prompt or response logs to inspect afterward. The company also doesn't require conventional identity or ID verification. That can make problematic use harder to investigate, even though account, payment, and usage metadata are still kept. The anonymous company representative argues that identity checks wouldn't reliably distinguish legitimate users from malicious ones. He also says tighter access controls could put smaller security firms at a disadvantage compared to large enterprises. The customer sets the rules Through an optional policy gateway, enterprise customers can define which requests are allowed, blocked, modified, or logged. The company also sells synthetic training and evaluation data. Standard model access, however, remains largely unrestricted. Additional control rules have to be explicitly turned on. The same philosophy extends to government customers. Abliteration.ai says it is registered for US government procurement on SAM.gov and promotes versioned models, audit logs, and agency-specific rules, starting with pilot projects that don't involve Controlled Unclassified Information (CUI). Modifying GLM-5.3 is allowed under its license. Whether any specific use is legal depends on what's being done and which jurisdiction applies. For security testing, Abliteration.ai explicitly requires written authorization for target systems and compliance with applicable laws. That leaves a separate question of what responsibilities a provider should take on when it offers powerful offensive capabilities through a readily accessible service. Abliteration.ai isn't the first to do this. Providers like Audn.AI with PenClaw and Silk Compute also host abliterated or largely unrestricted models for security use cases. Abliteration.ai's particular offering combines GLM-5.3 with straightforward API access and an optional policy layer for enterprise customers. The modification doesn't leave the rest of the model untouched A preliminary study suggests that refusals can be sharply reduced in certain models without a comparable decline in code-generation performance. Other experiments , however, find behavioral changes even on tasks where the base model didn't refuse anything at all. In other words, abliteration doesn't surgically remove a single trait but reaches deeper into how the model behaves. Abliteration.ai also shifts much of the decision-making over model limits from the provider to the customer, and whether that arrangement can serve legitimate security work without making harmful uses substantially easier remains unresolved. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. S