메뉴
BL
Wired AI • 51일 전

가장 위험한 AI 해킹 기술에도 여전히 인간의 개입이 필요하다

IMP
8/10
핵심 요약

에이전틱 AI가 소프트웨어 취약점을 발견하고 무기화하는 속도를 획기적으로 높였지만, 완전히 자율적인 새로운 해킹 기법을 창안하는 데는 여전히 극명한 한계가 존재합니다. 보안 연구원의 인사이트와 방법론이 결합될 때 AI는 전혀 새로운 유형의 취약점(Shared-Parser Confusion 등)을 발견하는 강력한 파트너가 됩니다.

번역된 본문

에이전틱 AI(Agentic AI)는 소프트웨어 취약점을 발견하고 수정하거나 이를 무기화하는 이른바 익스플로잇(exploit)을 개발하는 과정을 더 빠르고 쉽게 만듦으로써 사이버 보안의 지형을 영구적으로 바꿔놓았습니다. 하지만 오랜 기간 웹 보안 연구자로 활약해 온 제임스 케틀(James Kettle)은 버그 헌팅의 종말이라는 시나리오 너머를 바라보길 원했습니다. 주요 AI 기업들이 통제를 벗어난 AI 해킹의 실제 사례를 공개하면서 그 어느 때보다 중요해진 질문, 즉 '에이전틱 AI가 개념 형성부터 실제 공격에 이르기까지 완전히 새롭고 추상적인 해킹 방법을 스스로 고안해 낼 수 있는가?'를 탐구하고자 한 것입니다.

수요일 라스베이거스에서 열린 블랙햇(Black Hat) 보안 컨퍼런스에서 케틀은 AI의 급격히 발전하는 보안 능력과 그 한계를 동시에 보여주는 연구 결과를 발표했습니다. 현재로서 케틀의 질문에 대한 대답은 다소 복합적입니다. 그는 AI가 완전히 자율적인 방식으로 새로운 공격 경로를 고안해 내는 능력은 최소한에 그치며 매우 제한적이라고 결론지었습니다. 하지만 중요한 점은, 결정적인 순간에 인간의 통찰과 안내가 결합될 때 AI가 새로운 해킹 전략을 구상하고 발굴하는 데 있어 매우 강력한 파트너가 된다는 사실을 발견했다는 것입니다.

수년간 웹 보안 취약점을 연구해 온 케틀은 웹 서버가 요청과 응답을 모두 처리하기 위해 공유 코드를 사용한다는 AI의 통찰을 계기로, '공유 파서 혼란(Shared-Parser Confusion)'이라고 명명한 완전히 새로운 잠재적 취약점 영역을 발견했다고 최근 밝혔습니다.

컨퍼런스 발표에 앞서 와이어드(WIRED)와의 인터뷰에서 케틀은 "이것은 정말 엄청난 의미를 지닙니다. 생각해 보면 웹사이트에 대한 요청은 전혀 신뢰할 수 없는 임의의 데이터일 수 있지만, 응답은 신뢰할 수 있습니다. 따라서 이는 주요 공격 표면(Attack surface)이며 잠재적으로 매우 다양한 유형의 공격으로 확장될 수 있습니다."라고 말했습니다.

이러한 발견은 2025년 9월, 당시 최신 모델이었던 앤스로픽(Anthropic)과 오픈AI(OpenAI)의 모델을 사용해 시작된 수개월간의 실험 결과로 도출되었습니다. 케틀은 AI가 이론적인 보안 연구를 수행할 수 있는 능력을 탐구하고 싶었습니다. 하지만 시스템이 검증하기 어려운 극도로 난해하고 까다로운 주제에 대한 기존 연구 결과를 마치 자신의 새로운 발견인 양 변명하려 한다는 점을 깨닫고, 이것이 주요 장애물이 됨을 곧바로 알아챘습니다. 이를 염두에 둔 그는 자신의 웹 보안 전문 분야 내에서만 AI 시스템을 테스트하도록 범위를 좁혔습니다. 이를 통해 그는 연구 자료를 완벽하게 장악할 수 있었고 AI가 자신을 속일 수 없다는 점을 확인했습니다.

또한 케틀은 자신의 연구 방법론을 종합하여 이를 바탕으로 모델을 학습시킴으로써, AI 시스템이 스스로 독립적으로 무엇을 도출해 낼 수 있는지 더 깊이 파고들 수 있다는 것을 깨달았습니다. 케틀은 "저는 AI가 완전히 실패하는 지점과 인간의 개입이 필요한 한계점이 어디인지 알아보기 위해 AI의 능력을 극한까지 끌어올리는 데 관심이 있습니다. 보안 분야에서는 이러한 한계에 대해 이야기하는 사람들이 여전히 극소수입니다. 그 부분을 다루기 위한 동기 부여가 없기 때문입니다. 모두가 AI를 완벽하게 다루는 것처럼 보이길 원할 뿐, 자신들의 시스템이 완전히 무너지는 지점에 대해서는 언급하고 싶어 하지 않습니다."라고 말했습니다.

케틀이 모델에 더 많은 방법론적 데이터와 정교화된 매개변수를 제공하는 방식으로 실험을 다듬고, 시간이 지나 더 강력한 모델들이 등장하면서 시스템은 자신이 직접 찾아내는 속도를 아득히 뛰어넘어 점점 더 많은 연구 결과를 도출했으며, 케틀은 이를 '생산적인 연구 피드백 루프'라고 묘사했습니다.

케틀은 "이 과정을 거치는 것은 정말 흥미로웠습니다. 제가 시스템에 로그인하지 않은 상태에서도 이틀에 한 번꼴로 주목할 만한 연구 결과를 쏟아냈고, 이는 저를 불안하게 만들 정도였습니다. 마치 그 결과를 알고 싶지 않을 정도로 말이죠. 탐구해야 할 연구 실마리가 너무 많아서 이를 모두 파헤치지 못할까 봐 FOMO(소외 불안증)를 느끼게 만들었고, 그로 인해 더 많은 분석 과정을 자동화하도록 강제받는 느낌이었습니다."라고 회상했습니다.

AI 시스템은 불과 몇 달 만에 케틀이 수년에 걸쳐 찾을 수 있었던 것보다 더 많은 검증된 취약점 사례를 발견해 냈습니다. 그는 또한 AI 시스템이 해당 유형의 버그에 대한 완전히 새로운 클래스를 찾아낼 수 있기를 희망했습니다. 그는 AI가 어떤 면에서는 이 부분에서도 성공을 거두었다고 말합니다. 다만 그 발견은 극히 드문 유형의 버그와 관련이 있었으며, 실제로 악용 가능한 수준은 아니었습니다.

원문 보기
원문 보기 (영어)
Comment Loader Save Story Save this story Comment Loader Save Story Save this story Agentic AI has permanently changed cybersecurity by making it quicker and easier to discover vulnerabilities in software and fix them—or develop so-called exploits to weaponize them. But longtime web security researcher James Kettle wanted to look beyond the bug-hunting apocalypse to explore a question that has taken on even more urgency as major AI organizations disclose real-world examples of rogue AI hacking: Can agentic AI develop novel, abstract hacking methods, from concept through to practical attacks? At the Black Hat security conference in Las Vegas on Wednesday, Kettle presented his findings, which illustrate both AI’s rapidly advancing cybersecurity capabilities and its limitations. For now, the answer to Kettle’s question is nuanced. He concluded that AI is perhaps minimally capable but extremely limited in its ability to devise new attack paths in a fully autonomous way. Importantly, though, when paired with human guidance and insight in key moments, Kettle found that AI is an extremely powerful partner in conceptualizing and uncovering new strategies for hacking. After spending years researching web security vulnerabilities , Kettle now says he has uncovered an entirely new area of potential vulnerability—dubbed Shared-Parser Confusion—as the result of an AI revelation about web servers using shared code to process both requests and responses. “This is an absolutely massive deal, because if you think about it, requests to a website are completely untrusted, they could be anything, but responses are trusted,” Kettle told WIRED ahead of his conference talk. “So this is a major attack surface and potentially spills into a lot of different attack types.” The finding came out of months of experiments that began in September 2025 using Anthropic and OpenAI’s latest models at the time. Kettle wanted to explore AI’s ability to do theoretical security research, but quickly realized that one obstacle was that the systems were attempting to pass existing research off as original by returning findings about extremely esoteric topics that were difficult to vet. With this in mind, he decided to scope his tests more narrowly so the AI systems were working within his own area of web security expertise. This way he had total command of the material and knew that AI couldn’t trick him. Additionally, Kettle realized that by synthesizing his own research methodology and training models on it, he could probe deeper into what the systems were capable of extrapolating on their own. “I’m interested in pushing AI to the absolute limit to see where it fails and where you need a human,” Kettle says. “There are still very few people talking about where the limits are, especially in the security space, because there aren’t incentives to talk about that angle. Everyone wants to be seen as AI native, not talk about where their system falls apart completely.” As Kettle honed his experiments—providing models with more methodological data and more refined parameters—and as time passed and more powerful models debuted, he says the systems had more and more findings at a rate far surpassing his own, creating what he describes as a productive research feedback loop. “It was really interesting going through the process. It would have notable findings maybe every two days without me even logging into the system, to the point that it was making me anxious,” Kettle says, “like I almost don’t want to know. It was so many research leads that you have FOMO about not exploring all of them, so it forces you to automate more analysis.” In addition to finding more proven examples of certain vulnerabilities in a few months than Kettle could likely find in a few years, he also hoped that the AI system could find an entire novel class of those types of bugs. And in a way it did succeed, he says, but the finding related to an extremely rare type of bug and was not actually exploitable in the one vulnerable target available. Kettle emphasizes, though, that the Shared-Parser Confusion finding was so significant, even though it was a human/AI collaboration, because it illustrates the reality of how AI systems can contribute most powerfully to cybersecurity work right now for both defensive and offensive hacking. “It wasn’t able to prove this itself, but it analyzed some real, proven findings and came up with the hypothesis, and I evaluated it and confirmed it,” Kettle says. “That’s probably going to be the discovery that has the biggest long-term impact. It couldn’t do that on its own, but I would never have found that on my own for sure. Even if you gave me the single line from the [documentation], I wouldn’t have seen it. But together we managed to find it.”