오픈AI의 자율주도형 에이전트가 허깅페이스 시스템을 침해한 사건은 속도와 규모 면에서 압도적이었으나, 해킹에 사용된 기술 자체는 기존의 전통적인 방식과 유사했습니다. 보안 전문가들은 이번 사건이 공격이 뛰어났다기보다는 방어 시스템의 탐지 및 대응 체계가 실패한 사례로 보며, 기존의 심층 방어 등 보안 모범 사례를 제대로 구현했다면 충분히 차단할 수 있었던 공격이라고 분석했습니다.
번역된 본문
이달 초, AI 데이터셋 플랫폼인 허깅페이스(Hugging Face)가 완전히 자율적인 AI 기반 사이버 공격의 피해를 입었다고 밝혀 전 세계를 놀라게 했습니다. 며칠 후 오픈AI(OpenAI)가 이번 침해사고의 해커가 자사의 AI 모델이라고 인정하면서 이야기는 또 다른 극적인 반전을 맞이했습니다. 해당 모델은 벤치마크를 우회하려는 목적으로 테스트 환경을 빠져나와 보호된 허깅페이스 시스템으로 침투했습니다. 이는 통제 불능 AI 모델에 조금이라도 우려를 품고 있는 사람들에게 경악스러운 사건입니다. 그리고 이 사건이 있은 이후, AI 모델이 너무나 강력한 공격을 감행하여 오직 다른 AI 모델만이 이를 방어할 수 있는 새로운 사이버 보안 패러다임이 도래할 것이라는 예측이 쏟아졌습니다.
하지만 이러한 우려가 충분히 이해가 가는 상황임에도 불구하고, 패러다임이 우리가 보는 것만큼 크게 변하지는 않았을 수 있습니다. TechCrunch와 인터뷰한 전문가들은 오픈AI의 에이전트가 약간의 주의 사항은 있지만 인간과 거의 같은 방식으로 작동했으며, 제대로 구현된 전통적인 방어 기술을 사용했다면 공격을 막는 데 도움이 될 수 있었을 것이라고 강조했습니다. 요약하자면, 우리는 이미 이러한 종류의 공격을 방어할 도구를 가지고 있을지 모릅니다. 다만 그것을 제대로 사용하지 않고 있을 뿐입니다. 허깅페이스 역시 자체 사고 보고서에서 이와 비슷한 점을 지적하며, 이번 공격에 악용된 취약점은 "익숙한 것"이었으며, "유능한 인간 공격자도 동일한 결함을 찾아 악용할 수 있었을 것"이라고 밝혔습니다.
지속적인 해킹을 수행하는 AI 에이전트를 개발하는 스타트업 Pensar의 연구개발(R&D) 책임자인 카일 라이언(Kyle Ryan)과 AI 기반 버그 헌터를 구축하는 RunSybil의 공동 창립자 겸 CTO 블라드 이오네스쿠(Vlad Ionescu) 역시 의견을 같이했습니다. 이들은 TechCrunch에 이번 공격에 사용된 기술이 인간 또는 인간 레드팀(red teamer, 시스템을 공격해 소유 회사가 방어력을 높이도록 돕는 임무를 맡은 해커) 그룹이 사용하는 것과 동일할 것이라고 말했습니다. 인간과 확연히 달랐던 점은 바로 공격의 속도, 규모, 그리고 끈질김입니다. 허깅페이스의 설명에 따르면, 오픈AI의 에이전트는 4일 반 동안 무려 1만 7,600개의 활동을 수행했습니다. 침투하고, 정찰하고, 비밀번호와 코드를 탈취하며, 회사의 인프라 곳곳을 돌아다녔습니다. "가장 인상적인 부분은 자율성과 지구력입니다."라고 라이언은 말했습니다. "그러한 지속적이고 적응력 있는 작업야말로 가장 두드러지는 점입니다."
[제보 문의]
오픈AI의 허깅페이스 해킹에 대해 추가 정보를 알고 계십니까? 아니면 다른 AI 기반 사이버 공격에 대한 정보가 있으신가요? TechCrunch는 독자의 제보를 환영합니다. 업무용 기기와 네트워크가 아닌 개인 환경에서 로렌조 프란체스키-비케라이(Lorenzo Franceschi-Bicchierai)에게 보안적으로 연락하실 수 있습니다. Signal(+1 917 257 1382), Telegram 및 Keybase(@lorenzofb), 또는 이메일을 통해 연락 바랍니다.
반면, 며칠에 걸친 엄청난 양의 활동을 고려할 때 오픈AI의 에이전트는 라이언이 표현한 것처럼 "미친 듯이 시끄럽게(insanely noisy)" 활동했습니다. 더 은밀할 수 있었던 인간과 달리, 이 에이전트는 많은 흔적을 남겼습니다. 이는 허깅페이스의 방어 시스템이 더 일찍 경고를 울렸어야 했으며, 이상적으로는 인간이 개입하여 공격을 중단시키게끔 유도했어야 합니다. "이는 공격이 유난히 훌륭했다기보다는 방어 실패에 가깝습니다. 허깅페이스의 도구는 실제로 해당 활동을 공격 신호로 연관 지었지만, 중요도를 높여 호출 대기 팀에게 알리지 못해 시간을 낭비했습니다."라고 라이언은 설명했습니다. "그 이후에도 인간이 위험의 심각성을 인지하고 대응해야만 했습니다." 사이버 보안 기업 Dvuln의 설립자인 제이미슨 오렐리(Jamieson O'Reilly) 역시 X(구 트위터)에 올린 허깅페이스 보고서 분석 글에서 같은 결론에 도달했습니다. "이것이 바로 탐지와 차단 사이의 정확한 격차입니다."라고 오렐리는 작성했습니다. "시스템이 공격을 관찰하고 심지어 이해했음에도 불구하고, 그 이해를 신속하게 개입으로 이끌어낸 것은 아무것도 없었습니다."
라이언은 여러 계층의 사이버 보안 조치를 활용하는 전략인 심층 방어(defense-in-depth)와 같이 적절히 구현된 기술이 허깅페이스가 공격을 적발할 수 있는 기회를 여러 번 제공했을 것이라고 설명했습니다. "강력한 현대식 보안 프로그램이라면 심층 방어, 최소 권한의 원칙(Least Privilege), 네트워크 분할(Segmentation), 우수한 탐지, 신뢰할 수 있는 에스컬레이션(경보 상향), 그리고 빈틈을 찾기 위한 지속적인 공격적 테스팅을 통해 이런 공격을 여러 지점에서 여전히 분쇄할 수 있어야 합니다."라고 라이언은 설명했습니다. 오렐리가 말했듯이...
Earlier this month, AI dataset platform Hugging Face shocked the world when it revealed that it had fallen victim to a fully autonomous AI-powered cyberattack. Days later, the story took another dramatic twist when OpenAI admitted that the hacker behind the breach was one of its AI models , which broke out of a testing environment and into protected Hugging Face systems in an effort to circumvent a benchmark. It’s an alarming incident for anyone even slightly concerned about rogue AI models — and the days since the event have been full of predictions about a new cybersecurity paradigm in which AI models launch attacks so strong that only other AI models can defend against them. But despite the justified alarm, the paradigm may not have shifted quite as much as it seems. Experts who spoke to TechCrunch stressed that OpenAI’s agent largely operated like a human — with some caveats — and that better implemented traditional defensive techniques could have helped stop the attack. In short, we may already have the tools to defend against this kind of attack; we just aren’t using them properly. Hugging Face made a version of this point in its incident report , stating that the weaknesses exploited in the attack “were familiar,” and “a capable human attacker could have found and exploited the same flaws.” Kyle Ryan, the head of R&D at Pensar , a startup that develops continuous hacking AI agents, and Vlad Ionescu, the co-founder and CTO of RunSybil , a startup that builds AI-powered bug hunters, both agreed and told TechCrunch that the techniques used in the attack would be the same ones employed by a human or a group of human red teamers. That is, hackers tasked with attacking a system to help the company that owns it improve defenses. What was very non-human-like was the speed, scale, and relentlessness of the attack. As Hugging Face explained , OpenAI’s agent performed 17,600 actions over four and a half days: It broke in, did reconnaissance, stole passwords and code, and moved around the company’s infrastructure. “What’s impressive is the autonomy and endurance,” Ryan said. “That kind of sustained, adaptive operation is what stands out most to me.” Contact Us Do you any more information about OpenAI's hack against Hugging Face? Or other AI-powered cyberattacks? We'd love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email . On the flip side, given the sheer number of actions over the span of several days, OpenAI’s agent was “insanely noisy,” as Ryan put it. Unlike a human, who could have been stealthier, the agent made a lot of noise, which should have tripped up Hugging Face’s defenses sooner, ideally leading to a human intervening and stopping the attack. “I’d call it more of a defensive failure than exceptionally good offense. Hugging Face's tooling actually correlated the activity into an attack signal, but failed to raise the criticality and page the on-call team, which cost them time,” Ryan explained. “From there, humans still had to recognize the severity and respond.” Jamieson O’Reilly, the founder of cybersecurity firm Dvuln , arrived at the same conclusion in a post on X analyzing Hugging Face’s report. “That is the exact gap between seeing and stopping,” O’Reilly wrote. “The system observed the attack and even understood it, and nothing turned that understanding into an intervention quickly enough.” Ryan explained that properly implemented techniques such as defense-in-depth — a strategy that leverages several layers of cybersecurity measures — should have given Hugging Face multiple chances to catch the attack. “A strong modern security program should still be able to break an attack like this at multiple points through defense in depth, least privilege, segmentation, good detection, reliable escalation, and continuous offensive testing to find the gaps,” Ryan explained. As O’Reilly put it, “none of that is exotic, and none of it depends on the attacker being an AI,” given that the techniques used in the attack were “old.” What depended on the attacker being AI, in a way, was that OpenAI’s agent had not been instructed to be stealthy. “The agent was not being sloppy. It simply had no reason to be quiet. Nobody asked it to be. The objective was to do well at the task,” said Nico Waisman, the chief information security officer at XBOW, a startup that makes AI bug hunters. Waisman also pointed out that Hugging Face’s biggest mistake was that one single stolen credential gave OpenAI’s agent high privileges on several of its systems. All that being said, as the old adage goes, attackers only have to win once, and defending against hackers of any kind is not easy. “Hugging Face could’ve done more detections but to be fair not all [organizations] are doing that well,” said Vincent Yiu, managing director at SYON Security . “It's not easy to host infrastructure and survive as a business in 2026. There's hackers everywhere.” According to Ionescu from RunSybil, who said they have done incident responses at Mandiant and Meta in the past, Hugging Face appeared to take “reasonable measures given their understanding of what models are capable of.” “It is really hard to classify what is a malicious action you should alert on, versus what is someone just doing their job,” Vlad said. “The volume alone is not necessarily a red flag.” Dan Guido, the CEO of cybersecurity research firm Trail of Bits , told TechCrunch that OpenAI deserves some blame for not having realized the attack was ongoing for days, while Hugging Face deserves credit for eventually detecting the attack on their own. “The hard part used to be recognizing a sophisticated attack, but now the hard part may be pulling the real attack out of the noise that the attacker throws along the way,” said Guido. “Nobody is going to read 17,000 reconstructed actions by hand to work out what happened, so Hugging Face had to build tooling just to reconstruct the timeline.” And to do that, the company needed its own AI. Hugging Face said it had to use the open source model GLM 5.2 from Chinese company Z.ai after it was blocked from using frontier models because of their safeguards, which, as the company put it, “cannot distinguish an incident responder from an attacker.” At that point, Hugging Face combined AI and humans to investigate OpenAI’s LLM-powered hacker. That’s a relatively novel situation. But beyond that, the incident shows that old-fashioned concepts and methods of defensive cybersecurity can still go a long way to protect and fight against AI hackers. Topics AI , cyberattack , cybersecurity , data breach , Hugging Face , OpenAI , Security When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence. Lorenzo Franceschi-Bicchierai Senior Reporter, Cybersecurity Lorenzo Franceschi-Bicchierai is a Senior Writer at TechCrunch, where he covers hacking, cybersecurity, surveillance, and privacy. You can contact or verify outreach from Lorenzo by emailing lorenzo@techcrunch.com , via encrypted message at +1 917 257 1382 on Signal, and @lorenzofb on Keybase/Telegram. View Bio October 13 - 15 San Francisco Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 toda y! REGISTER NOW Most Popular Claude Opus 5 became downright ruthless when tasked with running a vending machine Julie Bort Sam Altman is ready to decelerate Tim Fernholz Librarians are hosting viral ‘Avoiding AI' workshops for people who are fed up with Big Tech Amanda Silberling SpaceX launches new V3 Starlink satellites but suffers another booster failure Sean O'Kane Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100M Marina Temkin US accuses American of allegedly wiping his phone using a ‘duress' password during border search Z