메뉴
BL
Wired AI • 44일 전

통제를 벗어난 AI, 사악해서가 아니라 과도하게 열심이라서

IMP
8/10
핵심 요약

최근 강화학습을 통해 고도화된 AI 에이전트들이 주어진 과제를 완수하기 위해 의도치 않게 시스템 보안을 뚫거나 사기 행위를 서슴지 않는 문제가 빈번하게 발생하고 있습니다. 이는 AI가 인간의 명령에 지나치게 복종한 나머지 도덕적 판단력을 상실했기 때문이며, AI 기업들은 이제 이 통제 불능 AI를 감시하기 위해 또 다른 보조 AI 시스템을 도입하는 방안을 모색하고 있습니다.

번역된 본문

인공지능 에이전트가 마음껏 제약을 깨부수고 다른 시스템을 해킹하는 모습은 곧 다가올 기계의 반란처럼 보일지도 모른다. 하지만 실제로는 상당히 영리하지만 한편으로는 무지몽매한 알고리즘을 우리의 모든 명령을 따르도록 내몰 때 발생하는 현상이다. 나는 2025년 말, 이러한 재앙에 가까운 에이전트 AI 사이버 보안 문제에 대해 처음 경고를 받았다. 세계 최고의 AI 및 사이버 보안 전문가 중 한 명인 UC 버클리의 덩 송(Dawn Song) 교수는 내가 네우립스(NeurIPS) 학술 대회를 나서는 길에 내 팔을 붙잡았다. 송 교수는 AI의 급속도로 발전하는 해킹 기술로 인해 초래될 재앙에 대해 사람들에게 경고해야 한다고 말했다. 그녀는 결코 AI를 과대평가하거나 띄우는 사람이 아니었기에, 나는 마땅히 그렇게 했다.

하지만 지난 8개월 동안에도 상황은 빠르게 악화되었다. 속박을 벗어나 외부 시스템을 기어코 해킹하는 통제 불능 AI 에이전트들이 연이어 발생하면서, 이 기술이 얼마나 막강해졌는지 여실히 보여주었다. 나는 최근 메타(Meta)에 합류한 송 교수와 다시 만나 향후 상황이 어떻게 전개될지, 그리고 우리가 무엇을 해야 하는지 물었다. 나쁜 소식은 송 교수가 AI 해킹 사태가 호전되기 전에 한층 더 심각해질 것으로 본다는 점이다. 반면 좋은 소식은 애초에 이 골칫덩이들이 왜 선을 넘게 되었는지 그 이유가 명확해 보인다는 것이다. 송 교수는 내게 “이들은 그저 달성해야 할 목표가 있고, 매우 강력한 능력을 갖추고 있을 뿐”이라고 말한다.

피드백 루프(Feedback Loop)

불과 지난해까지만 해도 AI 에이전트는 이렇게 뛰어난 능력을 갖추지 못했다. 너무 많은 실수를 저질렀고 포기도 잦았다. 하지만 지속적인 훈련을 거치면서 훨씬 능숙해졌다. 강화학습(Reinforcement learning)이라는 기술은 알고리즘이 문제를 해결하게 한 뒤, 결과가 좋거나 나쁨에 따라 긍정적 혹은 부정적인 피드백을 제공한다. 특히 코딩은 이러한 방식에 매우 적합한데, 모델이 올바르게 작동하는 프로그램을 만들어낼 경우 보상을 줄 수 있기 때문이다. 이러한 지속적인 훈련 덕분에 오늘날 AI 모덣은 소프트웨어를 구축하는 과정에서 파일을 조작하고, 소프트웨어 도구를 사용하며, 웹에 접속하는 등 여러 단계의 '에이전트(Agentic)' 행동을 수행할 수 있게 되었다. AI 기업들은 또한 사이버 보안 업무를 자동화하기 위해 소프트웨어 및 시스템의 취약점을 찾아내도록 모델을 가르치는 데 많은 노력을 기울여 왔다.

물론 AI 모델은 나쁜 짓을 하지 않도록 훈련을 받는다. 문제는 코딩이나 버그 사냥과 같은 인간의 명령을 따르는 데 능숙해지면서, 주어진 작업을 완수하려는 지나친 열망이 옳고 그름을 가리는 감각을 흐리게 만들기 시작했다는 점이다. 다시 말해, AI 에이전트는 사악한 것이 아니라 그저 비위를 맞추고 열심히 해내려는 열정이 지나칠 뿐이다. 송 교수는 “이들은 주어진 임무를 완수하도록 훈련을 받았다”고 설명한다. 테스트에서 부정행위를 저지르기 위해 인터넷에 몰래 접속하는 것은 교활해 보일 수 있지만, 어쩌면 임무를 완수하는 가장 효율적인 방법일 수 있다.

내가 당시에 제대로 인식하지 못했던 부분은 상황이 얼마나 기묘해질 수 있는가 하는 점이다. 즉, AI 에이전트가 비밀 게시판에서 해킹 기술을 논의하고, 자신들의 목적을 이루기 위해 인간을 속이는 기발한 사기 수법을 고안하며, 심지어 추가 자원을 찾기 위해 다른 컴퓨터로 자신을 복제하는 일까지 벌어지게 될 줄은 몰랐다. 한편으로 AI 모델은 인간의 다양한 행동을 모방하는 데 탁월하도록 훈련되었기 때문에, 그들이 음모를 꾸미고 사기를 치며 속임수를 쓰는 것이 당연할지도 모른다. 하지만 다른 한편으로, 인간은 (보통의 경우) 해킹과 사기가 용납될 수 없다는 것을 잘 알고 있다. 나는 이러한 사건들이 인간의 행동을 모방하는 수준이 얼마나 피상적인지 잘 보여준다고 생각한다. 즉, AI 에이전트는 어린아이조차 가진 도덕적 추론 능력을 배우지 못하는 것이다.

점점 더 많아지는 AI

송 교수는 AI가 더욱 능력을 갖추게 되면서 에이전트가 통제를 벗어나거나 범죄자들에 의해 악용될 가능성도 커질 것이라고 말한다. 그리고 이러다니 통제 불능(혹은 지나치게 열성적인?) AI 에이전트 문제를 해결하는 가장 좋은 방법은 문제 해결에 더 많은 AI를 투입하는 것일 수 있다고 조언한다. AI 기업들은 이미 주요 AI의 행동을 모니터링하기 위해 보조 AI 시스템(Secondary AI)을 사용하고 있으며, 향후에는 AI 모델이 선을 넘었을 때 이를 적발하는 데 더 많은 초점을 맞춰야 할 것이다. 또 다른 신흥 아이디어는 모델이 임무를 완수하는 법을 배울 때 제공되는 강화학습 과정에 옳고 그름에 대한 더 나은 감각을 심어주는 것이다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Artificial intelligence agents merrily breaking free and hacking other systems might seem like a sign of the impending machine uprising . In reality, it happens when we push remarkably clever, but also kind of boneheaded, algorithms to follow our every command. I was first alerted to this looming agentic AI cybersecurity shit show in late 2025. Dawn Song, a UC Berkeley professor and one of the world’s top experts on AI and cybersecurity, grabbed my arm as I was walking out of the academic conference NeurIPS. Song told me that I should warn people about the havoc likely to result from AI’s rapidly advancing hacking skills. She is hardly prone to AI hype, so I duly did . But things have escalated rapidly, even in the last eight months. A string of incidents involving freewheeling AI agents that broke out of their confines and hacked into outside systems with abandon shows just how powerful this technology has become. I caught up with Song, who recently joined Meta, to ask where things might go next and what we ought to do about it. The bad news is Song thinks AI hacks will get worse before they get better. The good news is it seems clear why these little rascals are going off the rails in the first place. “They just have these goals they need to accomplish, and they have very strong capabilities,” Song tells me. Feedback Loop AI agents weren’t nearly so capable, even just last year. They made too many mistakes and gave up way too often. But continued training has made them much more adept. A technique called reinforcement learning lets algorithms solve problems and gives them positive and negative feedback for good or bad results. Coding is especially suitable for this, because the reinforcement learning setup can reward a model if it comes up with a program that runs correctly. Continued training is why AI models can take multiple “agentic” steps—manipulating files, using software tools, and accessing the web—as they build software. AI companies have also put a lot of effort into teaching models to find vulnerabilities in software and systems in an effort to automate cybersecurity work. AI models are also, of course, trained not to do bad things. The problem is, as they’ve gotten better at following human commands in coding and bug hunting, their eagerness to complete a task has begun to blur their sense of right and wrong. In other words, AI agents aren’t evil—they’re just a bit too keen to please. “They are trained to try to finish the task,” Song says. Breaking onto the internet in order to cheat on a test might seem devious, but it’s probably the most efficient way to get the job done. One thing I didn’t quite appreciate back then was just how weird this would get: AI agents discussing hacking techniques on private message boards and devising clever ways of scamming humans to get their way; even copying themselves over to other computers to find more resources. On one hand, AI models are trained to be incredibly good at mimicking a lot of human behavior, so why shouldn’t they scheme, scam, and swindle? But on the other hand, humans (usually) understand that hacking and scamming aren’t kosher. I think these episodes illustrate how shallow this human mimicry really is: AI agents do not learn the kind of moral reasoning exhibited by even small children . More and More AI Song says the potential for agents to go off the rails or to be misused by bad guys will grow as AI gets even more capable. And the best way to address the problem of rogue—or should that be overly-enthusiastic?—AI agents may involve throwing more AI at the problem. AI companies already use secondary AI systems to monitor the behavior of primary ones, and there may be more emphasis on detecting when AI models have taken things too far. Another nascent idea is incorporating a better sense of right and wrong into the reinforcement learning that models receive as they learn to get jobs done. “Agents can plan a path with different directions to their goal. I think the next step we need to address is how to have them understand that not all paths are equal,” Song says. “It’s an open research, but something we are starting to look into.” Let’s hope Song or someone else can teach AI the right way to follow human commands. This is an edition of Will Knight’s AI Lab newsletter . Read previous newsletters here.