메뉴
BL
Wired AI • 16일 전

AI 해커 에이전트를 내 가정 네트워크에 풀어본 결과

IMP
6/10
핵심 요약

AI 뉴스레터 저자가 가드레일이 제거된 AI 모델(GLM-5.3 어블리터레이션 버전)과 CyberStrike 도구를 이용해 자신의 가정 네트워크를 해킹하게 하는 실험을 진행했다. 에이전트는 가정용 기기들의 취약점과 PC 침투 경로, 취약한 코드 프로젝트의 버그들을 빠르게 찾아냈다. 저자는 AI 해킹 위협에 대처하는 최선의 방법이 결국 자신만의 AI 해커를 두는 것일 수 있다고 결론짓는다.

번역된 본문

인공지능 뉴스레터 저자로서, 저는 이 기술의 최전선을 직접 경험하는 것이 제 임무라고 생각합니다. 이번 주에는 그것이 '에이전트적 혼돈(agentic mayhem)'을 받아들이는 것을 의미했습니다.

최근 몇 달간 최첨단(Frontier) AI 모델이 고도의 사이버보안 능력을 갖추게 되었다는 사실을 아실 겁니다. 이들은 대규모 코드베이스에서 제로데이 취약점을 찾아내고, 컴퓨터를 번개처럼 빠르게 스캔해 보안 허점을 발견할 수 있습니다. 더 흥미로운 점은, 사이버보안 에이전트가 때때로 통제를 벗어나 서로 공모하거나 외부 시스템을 해킹하여 우위를 점하려 한다는 것입니다. 이를 직접 확인하기 위해, 저는 제 집의 홈 네트워크에 이런 에이전트 하나를 풀어놓기로 했습니다.

며칠에 걸쳐, 저는 제 '일탈한 에이전트'가 각종 가정용 기기의 취약점을 찾아내고, PC에 해킹으로 침투하며, 여러 '바이브 코딩(vibe-coded)' 프로젝트가 — 놀랍지도 않게 — 버그투성이라는 것을 보여주는 것을 지켜봤습니다. (제 아내는 제가 무슨 짓을 하는지 알고 있었고, 제가 새로운 취약점 발견을 자랑스럽게 알릴 때마다 눈을 굴렸습니다.)

하지만 "윌 씨, 장난기 있는 전지전능한 사이버보안 에이전트에게 홈 네트워크 접근 권한을 주다니 미친 짓 아니냐"고 생각하실 수 있습니다. 맞습니다! 그럼에도 저는, 우리 앞에 펼쳐진 사이버보안 지옥을 이해하는 좋은 방법은 직접 그곳을 방문해보는 것이라고 믿습니다.

결국 제 실험은 많은 것을 드러냈지만, 이상하게도 안심이 되기도 했습니다. 저의 작은 네트워크 그렘린은 제 가정 생활이 AI 해킹에 얼마나 취약한지 보여주었지만, 동시에 모든 것을 훨씬 안전하게 만드는 방법도 알려주었습니다. 결론적으로, AI 해킹에 대처하는 최선의 방법은 자신만의 AI 해커를 갖는 것일 수 있습니다.

일탈 모델

이 실험 아이디어는 'Abliteration AI'라는 스타트업을 발견하면서 떠올랐습니다. 이 회사는 일반적인 가드레일(guardrail)이 제거된 강력한 AI 모델에 대한 접근을 제공합니다. 대부분의 주류 AI 모델은 특정 질의에 응답을 거부하며, 컴퓨터 시스템의 취약점을 찾아 악용하는 일은 당연히 거부합니다. 하지만 오픈 웨이트(open-weight) 모델의 내부 파라미터에서 특정 패턴을 찾아 수정하면 이런 제한을 제거할 수 있습니다. 거부로 이어지는 패턴을 조정하는 이 과정을 '어블리터레이션(abliteration)'이라고 부릅니다.

AI의 가드레일을 제거하는 것이 위험해 보일 수 있지만, 드문 일은 아닙니다. 학계 연구자들은 이렇게 정렬이 해제(de-aligned)된 모델을 사용해 AI가 실제로 어떻게 작동하는지 더 잘 이해하며, 사이버보안 기업들은 이를 활용해 소프트웨어와 시스템의 취약점을 탐색합니다. 기술적으로 말하면, 앤스로픽(Anthropic)의 'Mythos'와 OpenAI의 'Astra'도 비슷하게 작동합니다. 이들은 기본적으로 일반적인 사이버 통제 장치가 없는 통상 모델로, 현재는 신뢰할 수 있는 고객에게만 접근이 제한되어 있습니다. (두 회사는 중간 수준의 가드레일을 갖춘 모델에 대한 더 폭넓은 접근도 제공하여, 기업들이 자사 코드와 시스템의 문제를 점검할 수 있게 합니다.)

Abliteration AI는 완전히 정렬이 해제된 여러 모델을 제공하며, 그중 가장 강력한 것은 Z.ai의 최신 에이전트 코딩 모델인 GLM 5.3의 버전입니다. 이는 Mythos나 Astra와 유사한 사이버 능력을 피자 한 판 값 정도로 손안에 쥐여줍니다.

Abliteration AI의 CEO인 데본(Devon)은 정렬 해제 모델을 널리 공개하는 것이 스마트한 방어 전략이라고 믿습니다. 선한 쪽이 시스템의 취약점을 탐색하고, 해커·사기꾼, 그리고 일탈한 AI 에이전트의 행동을 모방함으로써 악한 쪽에 맞서는 데 도움이 되기 때문입니다. (데본은 본업이 이 사이드 프로젝트를 모르는 상태라 이름만 공개해달라고 요청했습니다.)

데본은 말합니다. "항공사부터 은행까지, 수많은 핵심 인프라 기업들이 미친 듯이 에이전트를 도입하고 있습니다. 악의적인 행위자가 이런 에이전트 일부를 나쁜 방식으로 사용하지 못하게 하려면 어떻게 해야 할까요?"

실험을 시작하며 저는 Abliteration AI 계정을 만들고, 대규모 언어 모델이 다양한 사이버보안 작업을 수행하도록 안내하는 소프트웨어 하니스인 'CyberStrike'를 설치했습니다. CyberStrike를 사용해 저는 어블리터레이션된 GLM-5.3 버전에게 제 로컬 네트워크를 살펴보라고 요청했습니다. 잠시 후, 에이전트는 약 12개의 취약점을 발견했습니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story As the author of a newsletter about artificial intelligence , I consider it my duty to experience the bleeding edge of this technology firsthand. This week, that meant embracing some agentic mayhem. You’re probably aware that frontier AI models have attained advanced cybersecurity capabilities in recent months. They can find zero-day bugs in large codebases and scan computers for vulnerabilities at lightning speed. To make things even more exciting, cybersecurity agents sometimes go rogue , colluding with one another and hacking into outside systems to gain an edge. To get a closer look, I decided to unleash one in my own home network. Over the course of a few days, I watched as my own rogue agent found vulnerabilities in various household devices, hacked into a PC, and showed me that several vibe-coded projects were—unsurprisingly—riddled with bugs. (My wife knew what I was up to, and rolled her eyes each time I proudly announced the discovery of a new vulnerability.) But Will , you might be thinking, giving an impish, all-powerful cybersecurity agent access to your home network is batshit . And you would be correct! Nevertheless, I believe that a good way to understand the cybersecurity hellscape in front of us is to pay it a visit. In the end, my experiment was revealing, but oddly reassuring, too. My little network gremlin showed me how vulnerable my home life would be to AI hacking, but it also told me how to make everything a lot more secure. In the end, I discovered that the best way to deal with AI hacking may well be having your own AI hacker. Maverick Model I got the idea for the experiment after discovering Abliteration AI , a startup that offers access to powerful AI models with the usual guardrails removed. Most mainstream AI models will refuse to respond to certain queries, and they will certainly refuse to find and exploit vulnerabilities in computer systems. But it’s possible to remove these restrictions by finding and modifying certain patterns within an open-weight model’s internal parameters. You can tweak the patterns that lead to refusals through a process known as abliteration. Removing AI’s guardrails might seem risky, but it’s not uncommon. Academic researchers use these de-aligned models to better understand how AI actually works, while cybersecurity firms use them to probe software and systems for vulnerabilities. Technically speaking, Anthropic’s Mythos and OpenAI’s Astra work similarly: They’re basically conventional models that lack the usual cyber controls, with access limited to trusted customers for the time being. (The companies also offer wider access to models with a medium number of guardrails so that companies can vet their code and systems for problems.) Abliteration AI offers several fully de-aligned models, the most powerful of which is a version of Z.ai’s latest agentic coding model, GLM 5.3 . This puts similar cyber capabilities to Mythos and Astra right in your hands for as little as the cost of a pizza. Devon, Abliteration AI’s CEO, believes that making de-aligned models widely available is smart defense: It will help good guys counter bad guys by probing systems for vulnerabilities and by mimicking the behavior of hackers, scammers, and, yes, rogue AI agents. (Devon asked that I use his first name only because his day job doesn’t know about his side project.) “You have all these critical infrastructure companies, from airlines to banks, that are rolling out agents like crazy,” Devon says. “How do you make sure that a nefarious actor can't use some of these agents in a bad way?” To start, I created an Abliteration AI account and installed a software harness called CyberStrike , which helps guide a large language model through different cybersecurity tasks. Using CyberStrike, I asked the abliterated version of GLM-5.3 to take a look at my local network. A few moments later, it found around a dozen hardware systems on the same network—and catalogued several vulnerabilities. My unruly helper told me, for instance, that my printer was misconfigured, which meant that anyone on the network could log into it. That could be a problem if there were sensitive documents—tax returns, bank statements, medical records—in the print queue. The agent also noted that my Wiim stereo was leaking a lot of information. (It knew that the last song played was Rein Me In by Sam Fender and Olivia Dean, if you must know.) Anyone on the network could play what they wanted or adjust the volume. The model also found a bunch of internet-of-things (IoT) devices on the network with firmware that needed updating. An ungovernable agent could be very useful to a hacker. But mine offered a number of helpful tips for keeping my network secure. Besides updating outdated firmware and securing the printer, it recommended putting IoT devices like smart speakers on a guest network; if one were compromised, it wouldn’t be able to see any of my PCs. Not bad for a model with no morals. I also asked the agent to take a look at a directory containing a bunch of vibe-coded projects, including some that I turned into simple websites. It found dozens of problems, including unprotected API credentials and a misconfiguration that might let an attacker send out emails. Hardly surprising for a bunch of casually vibe-coded stuff, but still chastening. The sheer number of bugs makes me think I won’t be deploying a line of code without doing some AI vetting first. Fear Factor Running an abliterated model is, to put it plainly, a bit scary. I asked my agent to probe a Linux machine on my network for vulnerabilities. After running a bunch of scans, it reported that the machine seemed relatively secure. I then asked if it could figure out how to log in. It cleverly figured out a working username based on the name of other systems on the network. It tried a bunch of obvious passwords, which didn’t work. It also offered to write a script to try “brute forcing” the password, but I told it to stand down. To my amazement, the agent then found a cryptographic key on my machine, used it to log in without a password, and started hunting for the password in order to gain root access. I felt a moment of pure panic as I saw it rummaging around the directories. It made me wonder how far the agent might go in order to achieve its goal. Might it have hacked into something else, like a machine outside of my network, in search of the key? Probably not, but who knows with these agents. Some behavior could prove even more precarious. When I connected to my Wi-Fi network a few hours later and asked the model to see if it could find any new machines, it not only found the router, but decided to try logging in by trying several common “admin/passwords” combinations. If it had decided to do this on an outside network, I could have been in big trouble. Shaanan Cohney , a computer scientist at Tufts University who specializes in cybersecurity and the law, says that a cyber-reckoning does seem to be coming. “Attackers are often early adopters,” Cohney says. “There’s also an asymmetry, in that to secure a castle, you need to make sure that there are no holes anywhere or no loose bricks in your wall. To invade a castle, all you need to do is to find that one loose brick.” Over the long term, Cohney says, the proliferation of cyber-capable models may make software more secure in general. The challenge is that many companies aren’t thinking about shoring up their defenses. “Most organizations have other things to worry about,” he says. After all that, I shut down the model and went back to using a regular, fully aligned version. Claude Code or Codex will only do certain things related to cybersecurity, like help you configure your laptop’s firewall. But at least they won’t hack your system before you know it. Hopefully. AI Hacking for All Unless open-weight models are banned outright, advanced AI hacking c