메뉴
BL
The Decoder • 7일 전

앤스로픽 클로드로 오픈AI 내부 시스템 해킹, 72시간 만에 성공

IMP
9/10
핵심 요약

보안 연구팀 Hacktron이 Anthropic의 Claude를 이용해 OpenAI 커뮤니티 포럼의 취약점을 연결(chain)해 직원 ChatGPT·Codex 계정과 내부 GitHub 코드 저장소에 접근하는 데 성공했다. 오래된 이미지 라이브러리(libheif)와 SSO 인증 오구성이 발판이었으며, Claude Opus 5가 출시된 후에야 완전한 익스플로잇이 가능했다. 이 사건은 AI 모델이 정교한 사이버 공격에 필요한 시간과 기술 수준을 획기적으로 낮추고 있음을 보여준다.

번역된 본문

보안 연구자들이 Anthropic의 Claude를 사용해 OpenAI의 내부 시스템과 코드 저장소를 침입했다. 이 공격은 OpenAI의 커뮤니티 포럼을 통해 이루어졌으며, 연결된 취약점들을 통해 연구자들은 직원들의 ChatGPT 및 Codex 계정에 접근할 수 있었다. 전체 익스플로잇 개발에는 단 72시간이 걸렸을 뿐이며, AI 모델이 정교한 사이버 공격에 필요한 시간과 기술을 얼마나 크게 줄이고 있는지 보여준다.

세 명의 보안 연구자들이 Anthropic의 Claude 모델을 사용해 OpenAI의 커뮤니티 포럼(community.openai.com)을 통해 회사 내부 시스템에 침입했다. 공격은 72시간도 걸리지 않았으며, 팀에 따르면 Opus 5가 출시된 후에야 가능해졌다.

OpenAI가 자업자득을 겪은 셈이다. 수개월간 에이전트들이 인터넷을 자유롭게 해킹하도록 방치했던 이 회사는 이제 AI의 도움을 받아 해킹을 당했다.

Hacktron의 보안 팀은 두 개의 취약점을 연결해 OpenAI 직원들의 ChatGPT와 Codex 계정에 접근했고, 이를 통해 OpenAI의 GitHub 내부 코드 저장소까지 침입했다. 포럼에서 "OpenAI로 로그인"을 사용한 모든 사용자와 직원이 잠재적 영향 대상이었다. 사용자는 GitHub, Slack, 이메일을 Codex와 ChatGPT에 연결할 수 있으므로 공격이 이론적으로는 해당 서비스까지 확대될 수 있었다. 접근을 증명하기 위해 연구자들은 한 직원의 Codex 계정으로 내부 모노레포에 무해한 풀 리퀘스트를 생성했다. 민감한 데이터는 열람하지 않았다고 밝혔다.

오래된 이미지 라이브러리와 인증 결함이 문을 열었다

첫 번째 취약점은 포럼이 업로드된 HEIC 이미지를 처리하는 데 사용하던 라이브러리 libheif에 있었다. Hacktron에 따르면 수정 패치는 원본 소스 코드에 1년 전부터 존재했지만 아무도 보안 문제로 지적하지 않았다. 포럼이 사용하던 Debian 패키지에는 여전히 패치가 적용되지 않은 상태였다. 조작된 이미지 파일을 통해 연구자들은 서버에서 자신들의 코드를 실행할 수 있었다.

두 번째 취약점은 OpenAI의 중앙 싱글 사인온(SSO) 시스템의 잘못된 구성이었다. 포럼 서버를 통제하는 사람은 누구나 활성 포럼 회원을 사칭해 그들의 ChatGPT와 Codex 계정을 탈취할 수 있었다. Hacktron은 이 결함이 포럼을 넘어서는 것이라고 밝혔다. OpenAI 로그인을 사용하는 서비스가 침해되면 어디서든 동일한 접근 권한을 얻을 수 있었다.

Claude Opus 5는 전작이 실패한 곳에서 성공했다

연구자들은 처음에 Claude Opus 4.8로 취약점을 찾았다고 한다. 해당 모델은 동작하는 익스플로잇을 만들었지만, 메모리 공격 방어 기법인 ASLR이 비활성화된 경우에만 가능했다. 여러 세션에 걸쳐 ASLR이 활성화된 상태의 안정적인 버전을 만들지 못했다.

7월 24일 저녁, Anthropic이 Claude Opus 5를 출시했다. Hacktron에 따르면 새 모델은 3시간 만에 로컬 Mac에서 동작하는 익스플로잇을 만들었고, 이를 Discourse 서버 환경에 맞게 적응시켰다. 연구자들은 자체 테스트 인스턴스에 대해 Claude를 자율 루프로 실행했다. 모델이 실제 시스템에 대한 익스플로잇 작성을 거부했기 때문에 대상을 벤치마크 과제로 제시했다. 4시간 후 에이전트는 서버를 장악했다. 관련 과제에서 연구자들은 OpenAI의 GPT-5.6 Sol에 비해서도 성능이 크게 향상된 것을 관찰했다.

OpenAI는 보고 약 14시간 후 수정을 확인했으며, 포럼의 기반 소프트웨어인 Discourse도 며칠 내에 대응했다.

AI가 공격 비용을 획기적으로 낮추다

OpenAI 해킹 외에도 연구자들은 "HEIF Heist"라 불리는 조사를 Slack, Meta, GitHub Enterprise 등 다른 대상으로 확장했다. 세 명이 2개월에 걸쳐 프로젝트를 수행했으며 AI 비용은 3,000달러도 채 들지 않았다. 새 대상마다 공격을 적응시키는 데는 1~2일밖에 걸리지 않았다. 수천 건의 이미지 업로드에도 불구하고 Shopify만 이상 활동을 감지했다.

원문 보기
원문 보기 (영어)
Security researchers used Anthropic's Claude to hack OpenAI's internal systems in under 72 hours Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 18, 2026 Nano Banana Pro prompted by THE DECODER Key Points Security researchers used Anthropic's Claude to break into OpenAI's internal systems and code repositories. The attack ran through OpenAI's community forum, where chained vulnerabilities gave the researchers access to employee ChatGPT and Codex accounts. The entire exploit took just 72 hours to develop, showing how drastically AI models are cutting the time and skill needed for sophisticated cyberattacks. Ask about this article… Search Three security researchers used Anthropic's Claude models to break into OpenAI's internal systems through the company's community forum. The attack took less than 72 hours and, according to the team, only became possible once Opus 5 shipped. OpenAI is getting a taste of its own medicine. After inadvertently letting agents hack their way across the internet for months , the company has now been hacked with AI's help. Hacktron's security team chained two vulnerabilities together to access OpenAI employees' ChatGPT and Codex accounts. From there, they broke into OpenAI's internal code repository on GitHub. The attack ran through OpenAI's community forum at community.openai.com . Any user or employee who had used "Sign in with OpenAI" there was potentially affected. Users can connect GitHub, Slack, and email to Codex and ChatGPT, so the attack could theoretically have reached those services too. To prove they had access, the researchers used an employee's Codex account to create a harmless pull request in the internal monorepo. They say they didn't view any sensitive data. Ad An outdated image library and flawed authentication opened the door The first vulnerability was in libheif, the library the forum used to process uploaded HEIC images. According to Hacktron, a fix had been available in the original source code for a year, but no one had flagged it as a security issue. The Debian packages running on the forum still lacked the fix. A crafted image file let the researchers run their own code on the server. Ad The second vulnerability was a misconfiguration in OpenAI's central single sign-on (SSO) system. Anyone controlling the forum server could impersonate active forum members and take over their ChatGPT and Codex accounts. The flaw extended beyond the forum, Hacktron writes. Any compromised service using OpenAI login would have granted the same access. Claude Opus 5 succeeded where its predecessor failed The researchers say they initially used Claude Opus 4.8 to find the vulnerability. The model built a working exploit, but only with ASLR, a common defense against memory attacks, disabled. Across several sessions, it couldn't produce a reliable version with ASLR enabled. Ad On the evening of July 24, Anthropic released Claude Opus 5 . According to Hacktron, the new model produced a working exploit for a local Mac within three hours, then adapted it to the Discourse server environment. The researchers then ran Claude in an autonomous loop against their own test instance. Because the model refused to write exploits against real systems, they presented the target as a benchmark task. Four hours later, the agent had taken over the server. On a related task, the researchers also observed a jump in performance compared to OpenAI's GPT-5.6 Sol. Ad OpenAI confirmed the fix about 14 hours after the report. Discourse, the software behind the forum, also responded within days. Ad AI is making attacks drastically cheaper Beyond the OpenAI hack, the researchers expanded their investigation, dubbed "HEIF Heist" , to cover Slack, Meta, GitHub Enterprise, and other targets. Three people carried out the project over two months, spending less than $3,000 on AI. Adapting the attack to each new target took just one to two days. Only Shopify noticed the activity, despite thousands of image uploads and repeated crashes in image processing. Software has long benefited from a kind of security through complexity, the Hacktron team writes. Even with public source code and a known vulnerability, building a reliable exploit required rare expertise, time, and deep knowledge of the target environment. Complexity wasn't a true security barrier, but it did protect many companies in practice. AI strips away that protection by replacing scarce expertise with computing power. Hacktron argues that threat models must reflect how cheap attacks have become. "Work that once required a well-resourced team and months of effort can now be compressed into days," the team writes. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Hacktron