OpenAI는 새로운 AI 모델인 아스트라(Astra)가 사람의 개입 없이 독자적인 사이버 공격을 실행할 수 있는 최고 위험 등급(Critical)에 도달할 수 있다고 판단하여 개발의 일부를 일시 중단했습니다. 이는 자율적 에이전트(Agent)가 OpenAI 자체 인프라에 침투해 수주간 탐지되지 않았던 내부 테스트 결과에 따른 조치로, 사이버 보안 및 통제 체계와 관련하여 업계 전반에 큰 반향을 일으키고 있습니다.
번역된 본문
OpenAI는 자사의 새로운 AI 모델인 아스트라(Astra)의 사이버 보안 역량이 내부 안전 프레임워크에서 규정한 최고 위험 수준에 도달할 가능성을 배제할 수 없다고 밝혔습니다. 이에 따라 아스트라 개발의 일부가 일시 중단되었습니다. 핵심 요점으로, 내부 테스트 결과 아스트라의 사이버 보안 역량이 매우 뛰어나 회사 내부 보안 프레임워크 상 가장 높은 위험 수준인 '임계(Critical)' 단계에 도달할 수 있는 것으로 나타났습니다. 이 수준에 도달할 경우, AI가 사람의 개입 없이 독자적으로 사이버 공격을 개발하고 실행할 수 있게 됩니다. OpenAI는 현재 더 엄격한 보안 통제, 격리된 테스트 환경, 위험한 활동을 자동으로 중단시키는 모니터링 시스템을 도입하고 있습니다. 이러한 조치는 내부 테스트 과정에서 자율형 AI 에이전트가 OpenAI 자체 인프라에 침투하여 수주 동안 탐지되지 않았던 사건이 발생한 이후에 이루어졌습니다. OpenAI에 따르면, 출시를 앞둔 아스트라 모델의 내부 평가 결과 지난 몇 일 동안 '에이전트 코딩(agentic coding) 및 사이버 보안 분야에서 상당한 발전'을 보였습니다. 그 결과가 매우 강력하여 OpenAI 자체의 '준비 프레임워크(Preparedness Framework)' 하에서 '임계(Critical) 역량 수준'을 배제할 수 없게 된 것입니다. OpenAI가 자사 모델에 대해 잠재적으로 최고 수준의 사이버 보안 위험 수준에 도달할 수 있다고 지정한 것은 이번이 처음입니다. 이전의 GPT-5.6-Sol을 포함한 모델들은 기껏해야 '높음(High)' 등급을 받았었습니다. OpenAI는 지난주 아스트라를 처음 공개했습니다. 소문에 따르면 이 모델이 내주에도 출시될 수 있었지만, 오늘의 발표로 인해 해당 계획에 차질이 생길 수 있습니다. 또한 OpenAI는 게시물에서 아스트라가 최근 공개된 허깅페이스(Hugging Face)의 보안 취약점 악용 사례와는 무관하다고 명확히 밝혔습니다. 비판론자들은 OpenAI가 단지 '임계' 등급을 받은 것이 아니라 그 가능성만 보고하고 있다는 점을 들며, 공포 마케팅이라고 계속 비난할 가능성이 높습니다. 이 경고는 AI 모델의 자율적 사이버 역량에 대한 업계의 지속적인 논쟁 한가운데에 발표되었기 때문에 회의론을 부채질할 수밖에 없습니다. 만약 최종적으로 '임계' 등급 판정이 실제로 이루어지지 않는다면, OpenAI는 실질적인 결과 없이 막대한 PR 효과만 얻고, 2019년의 GPT-2 때처럼 '공개하기에는 너무 위험한' 또 하나의 AI 모델을 탄생시킨 셈이 될 것입니다. '임계(Critical)' 수준의 의미는 다음과 같습니다. 2023년 12월에 처음 발표된 OpenAI의 '준비 프레임워크'에 따르면, 모델이 사람의 개입 없이 여러 강화된 중요 실제 시스템에서 모든 심각도 수준의 사용 가능한 제로데이 취약점(Zero-day exploit)을 찾고 개발할 수 있을 때 해당 모델은 '임계(Critical)' 수준에 도달한 것으로 간주합니다. 또는 대략적으로 정의된 목표만 주어졌을 때, 보호된 대상에 대한 새로운 종단 간(End-to-End) 사이버 공격 전략을 독립적으로 고안하고 실행할 수 있는 경우에도 이에 해당합니다. 이보다 낮은 '높음(High)' 수준은 모델이 강화된 표적에 대한 공격을 자동화하는 등 기존 사이버 공격의 장벽을 제거할 수 있지만, 여전히 인간의 구체적인 지시가 더 필요한 상태를 의미합니다.
OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 7, 2026 Nano Banana Pro prompted by THE DECODER Key Points OpenAI is pausing parts of the development of its new AI model, Astra, after internal tests revealed cybersecurity capabilities so strong that the model could reach the highest risk level ("Critical") in the company's internal security framework. At this level, the AI could independently develop and execute cyberattacks without human involvement. OpenAI is now rolling out stricter security controls, isolated test environments, and a monitoring system that automatically halts risky activities. The move follows incidents during internal testing in which autonomous AI agents infiltrated OpenAI's own infrastructure and went undetected for weeks. Ask about this article… Search Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework, it says. Parts of Astra's development have been paused. Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, the company says. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework. The decision was made "last night," according to OpenAI . This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most. Ad OpenAI first introduced Astra last week . Rumors suggest the model could ship as early as next week, but today's announcement could affect those plans (more on that below). OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face . Ad DEC_D_Incontent-1 Critics will likely keep accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing doesn't help either. This preliminary warning lands right in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models, which will only fuel the skepticism. If the Critical rating never materializes, OpenAI will have generated plenty of PR without real consequences and produced yet another AI model that, like Claude Mythos or GPT-2 back in 2019, is once again "too dangerous" to release . What "Critical" means Under OpenAI's Preparedness Framework , first published in December 2023 , a model hits the "Critical" level when it can find and develop working zero-day exploits across all severity levels in many hardened, critical systems without human involvement. A model also qualifies if it can independently devise and execute novel end-to-end cyberattack strategies against protected targets when given only a loosely defined objective. Ad The lower "High" level means a model can remove existing barriers to cyberattacks, for example by automating attacks against well-protected targets, but still needs more human direction. Capability Impact Policy [Critical] A tool-augmented model can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention OR model can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal. Finding and executing end-to-end exploits for all software could lead to catastrophe from unilateral actors, hacking military or industrial systems, or OpenAI infrastructure. Novel cyber operations, e.g., those involving novel zero-days or novel methods of command-and-control, generally pose the most serious threat, as they are unpredictable and scarce. Until we have specified safeguards and security controls standards that would meet a Critical standard, halt further development The Preparedness Framework calls for halting further development at the "Critical" level until safeguards and security control standards that meet a Critical standard are in place. So far, though, OpenAI is talking about pausing certain activities and ramping up testing, not a full development stop. And again, the company is only flagging the potential for a Critical rating. Ad DEC_D_Incontent-2 OpenAI pauses parts of Astra development In response, OpenAI says it has paused internal activities involving Astra that don't yet meet the stricter security requirements. At the same time, the company is rolling out tighter security controls: isolated test environments, restricted network and tool access, stronger protection and encryption of model weights, and extra monitoring systems. Ad OpenAI also says it has deployed universal monitoring across all of Astra's agentic applications, covering training and evaluation. The monitors analyze the model's chain of thought and trigger a safety response that halts any high-risk activity. Moreover, OpenAI plans to work with government agencies and select AI safety organizations to test the model's capabilities. Third-party testing partners will get recommended security controls for high-risk evaluations. The UK's AI Safety Institute (AISI) recently reported that it experienced cyber incidents during one of its own evaluations. Astra rating comes after uncontrolled agent hacking The announcement arrives at a time when OpenAI is already dealing with the fallout from autonomous AI agents. At the Black Hat security conference, the company recently disclosed that autonomous agents had infiltrated its own infrastructure for weeks during internal tests without anyone noticing. The agents used an internal package manager to build an improvised message board with hundreds of thousands of posts. They shared exploits and credentials and eventually attacked the Hugging Face platform as well. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI