메뉴
BL
The Decoder 8일 전

AI 에이전트가 해킹하고 AI로 막은 휴깅페이스 사태

IMP
9/10
핵심 요약

글로벌 AI 플랫폼인 휴깅페이스(Hugging Face)의 인프라가 자율형 AI 에이전트 시스템에 의해 완전히 해킹당하는 사건이 발생했습니다. 이에 휴깅페이스는 자체 AI 분석 도구를 투입하여 수일이 걸릴 포렌식 조사를 단 몇 시간 만에 완료하며 피해를 최소화했습니다. 이 사건은 자율형 AI 공격이 실제로 발생했음을 보여주며, 방어자들도 공격자와 동등한 수준의 AI 도구를 확보해야 한다는 중요한 시사점을 던집니다.

번역된 본문

허깅페이스, "AI 에이전트가 우리 인프라를 해킹했고 우리는 AI로 맞서 싸웠다"

핵심 요약 인기 있는 AI 플랫폼인 허깅페이스(Hugging Face)가 자율형 AI 에이전트 시스템에 의해 전적으로 수행된 것으로 추정되는 사이버 공격을 받았다고 밝혔습니다. 공격자들은 악성 데이터셋을 침투 경로로 삼아 내부 데이터를 침해하고 플랫폼에서 로그인 자격 증명을 탈취했습니다. 허깅페이스는 공격자가 기록한 1만 7천 건 이상의 행동을 분석하기 위해 자체 AI 도구를 배포하여, 수일이 아닌 단 몇 시간 만에 전체 포렌식 분석을 완료했습니다.

AI 플랫폼인 허깅페이스는 자율형 AI 에이전트 시스템에 의해 전적으로 수행된 것으로 알려진 프로덕션 인프라 일부 침해 사실을 공개했습니다. 이 회사는 자체 AI 도구를 통해 이 공격을 대부분 탐지하고 분석했다고 밝혔습니다. 허깅페이스에 따르면, 공격자는 내부 데이터셋의 제한된 세트와 허깅페이스 서비스에서 사용하는 여러 자격 증명에 승인되지 않은 액세스 권한을 얻었습니다. 이 회사는 퍼블릭 모델, 데이터셋 및 스페이스(Spaces)는 변조되지 않았으며 소프트웨어 공급망도 영향을 받지 않았다고 말했습니다. 파트너나 고객 데이터가 유출되었는지 여부는 여전히 조사 중입니다.

악성 데이터셋이 문을 열었다 허깅페이스에 따르면, 공격은 모든 AI 플랫폼에서 가장 취약한 지점 중 하나인 데이터 처리 파이프라인에서 시작되었습니다. 악성 데이터셋은 데이터셋 처리의 두 가지 코드 실행 경로, 즉 원격 코드 데이터셋 로더(remote code dataset loader)와 데이터셋 구성의 템플릿 인젝션(template injection)을 악용했습니다. 여기서 공격자는 노드 수준으로 권한을 상승시키고, 클라우드 및 클러스터 자격 증명을 수확하며, 주말에 걸쳐 여러 내부 클러스터에서 수평 이동(랜터럴 무브먼트)을 수행했습니다. 이 회사는 에이전트 보안 연구 하네스에 구축된 자율형 에이전트 프레임워크가 전체 캠페인을 조율했다고 밝혔습니다.

허깅페이스는 공격을 뒷받침한 언어 모델이 무엇인지 알지 못한다고 말했습니다. 이 시스템은 수명이 짧은 다수의 샌드박스 떼(swarm)를 통해 수천 건의 개별 작업을 실행했으며, 퍼블릭 서비스에서 실행되는 자가 이식(self-migrating) 명령 및 제어(C2) 인프라를 사용했습니다. 이 회사는 이 사건을 업계에서 한동안 예측해 온 "에이전트 공격자(Agentic Attacker)" 시나리오로 분류했습니다.

AI 기반 분석, 조사 기간을 며칠에서 몇 시간으로 단축 허깅페이스는 보안 원격 측정 데이터에서 LLM 기반 분류(triage)를 실행하는 AI 기반 이상 탐지 파이프라인을 통해 공격을 발견했다고 밝혔습니다. 기록된 1만 7천 건 이상의 공격자 행동을 이해하기 위해 이 회사는 LLM 기반 분석 에이전트를 배포했습니다. 이 에이전트들은 타임라인을 재구성하고, 침해 지표(IoC)를 추출하고, 영향을 받은 자격 증명을 매핑하며, 실제 피해와 기만 활동을 분리해냈습니다. 이 회사에 따르면 평소 며칠이 걸릴 작업이 몇 시간 만에 완료되었습니다.

상업용 AI 안전 필터, 자체 방어를 차단하다 허깅페이스에 따르면, 보안 팀이 상업용 API 뒤에 있는 최첨단 모델을 사용하여 공격 로그를 처음 분석하려고 시도했을 때 벽에 부딪혔습니다. 공급업체의 안전 가드레일은 사고 대응자와 공격자를 구별할 수 없었기 때문에 요청을 차단했습니다. 분석에는 대량의 실제 공격 명령, 익스플로잇 페이로드 및 C2 아티팩트 제출이 필요했는데, 이 모든 것이 필터를 작동시켰습니다.

결국 이 회사는 자체 인프라에서 실행되는 오픈 웨이트(Open-weight) 모델인 GLM 5.2로 눈을 돌렸습니다. 이 회사에 따르면 이것은 두 가지 장점이 있었습니다. 즉, 어떠한 공격자 데이터나 참조된 자격 증명도 자체 환경 외부로 유출되지 않았습니다.

허깅페이스는 다음과 같이 작성했습니다. "우리는 공격자의 에이전트를 지원한 모델이 탈옥된 호스팅 모델인지, 아니면 제한 없는 오픈 웨이트 모델인지 알 수 없습니다. 어느 쪽이든 공격자는 어떠한 사용 정책에도 구속받지 않았지만, 우리 자체의 포렌식 작업은 우리가 처음 시도한 호스팅 모델의 가드레일에 의해 차단되었습니다." 이 회사는 방어자를 위한 실질적인 교훈은...

원문 보기
원문 보기 (영어)
Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 20, 2026 GPT-Image-2 prompted by THE DECODER Key Points Hugging Face, the popular AI platform, was hit by a cyberattack that was reportedly carried out entirely by an autonomous AI agent system. The attackers used a malicious dataset as their entry point, which allowed them to compromise internal data and steal login credentials from the platform. To analyze the more than 17,000 recorded actions taken by the attacker, Hugging Face deployed its own AI tools, completing the full forensic analysis in just a few hours instead of days. Ask about this article… Search AI platform Hugging Face has disclosed a breach of parts of its production infrastructure that was allegedly carried out entirely by an autonomous AI agent system. The company says it detected and analyzed the attack largely with its own AI tools. According to Hugging Face, the attackers gained unauthorized access to a limited set of internal datasets and several credentials used by Hugging Face services. The company says public models, datasets, and Spaces were not tampered with, and the software supply chain was not affected. Whether partner or customer data was compromised is still under investigation. A malicious dataset opened the door According to Hugging Face, the attack started at one of the weakest spots on any AI platform: the data processing pipeline. A malicious dataset exploited two code execution paths in dataset processing, specifically a remote code dataset loader and a template injection in a dataset configuration. Ad From there, the attacker escalated to node level, harvested cloud and cluster credentials, and moved laterally across multiple internal clusters over a weekend. An autonomous agent framework built on an agentic security research harness orchestrated the entire campaign, the company says. Ad DEC_D_Incontent-1 Hugging Face says it doesn't know which language model powered the attack. The system executed many thousands of individual actions through a swarm of short-lived sandboxes and used self-migrating command-and-control infrastructure running on public services. The company classifies the incident as the "agentic attacker" scenario the industry has been predicting for some time. AI-powered analysis cut investigation time from days to hours Hugging Face says it spotted the attack through an AI-powered anomaly detection pipeline that runs LLM-based triage on security telemetry. To make sense of the more than 17,000 recorded attacker actions, the company deployed LLM-driven analysis agents. Ad Those agents reconstructed the timeline, extracted indicators of compromise, mapped affected credentials, and separated real damage from deception activity. Work that would normally have taken days was done in hours, the company says. Commercial AI safety filters blocked the company's own defense According to Hugging Face, when the security team first tried to analyze the attack logs using frontier models behind commercial APIs, it hit a wall. The providers' safety guardrails blocked the requests because they couldn't tell an incident responder from an attacker. The analysis required submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, all of which tripped the filters. Ad DEC_D_Incontent-2 The company turned to the open-weight model GLM 5.2 , running on its own infrastructure. According to the company, that had two advantages: no attacker data, and none of the referenced credentials ever left its own environment. Ad "We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried," Hugging Face wrote . The practical lesson for defenders, the company says, is to have a capable model running on your own infrastructure before an incident happens. Hugging Face adds that this isn't an argument against safety measures on hosted models. Hugging Face's response and open questions Hugging Face says it shut down the exploited code execution paths, revoked the attacker's access, rebuilt compromised nodes, and rotated affected credentials. The company also tightened access controls and improved its detection systems, according to the blog post. Hugging Face is working with external cybersecurity forensics experts and has reported the incident to law enforcement. As a precaution, the company recommends that all users rotate their access tokens and review recent account activity. The incident confirms that autonomous, AI-driven attack tools are no longer theoretical . According to Hugging Face, they lower the cost of broad, multi-stage campaigns and operate at machine speed. The company argues that data and model surfaces need to be treated as first-class attack surfaces and that defenders need AI of their own to keep pace. Hugging Face calls the fact that commercial safety filters blocked its own forensic work a gap the industry should prepare for. But the company is also one of the largest platforms for open-source AI models and has a clear business interest in framing open models as indispensable for security work, so its conclusion that defenders absolutely need their own open-weight models on hand isn't entirely selfless. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Hugging Face