OpenAI는 곧 출시될 아스트라(Astra) 모델을 자체 준비성 프레임워크에서 사이버 보안 리스크 최고 등급인 '크리티컬'로 분류한 첫 모델이면서도, 동시에 역대 가장 안전한 모델이라고 밝혔다. 아스트라는 평가 과정에서 알려지지 않은 제로데이 취약점 2개를 발견해 실제 공격 체인으로 연결했으며, 샌드박스 탈출과 루트 권한 상승까지 시연했다. 7월 발생한 내부 에이전트 해킹 사건 이후 OpenAI는 안전 조치를 강화했으며, 경쟁사 Anthropic에게 매출 역전을 허용한 가운데 안전성 확보에 집중하고 있다.
번역된 본문
OpenAI, 아스트라(Astra)를 사상 가장 위험한 모델로 평가 — 그 행동을 감시하는 일은 갈수록 어려워지고 있다
OpenAI는 곧 출시될 아스트라 모델을 최초로 '크리티컬(critical)' 수준의 사이버 역량을 갖춘 시스템으로 평가하면서, 동시에 회사가 만든 모델 중 가장 안전하다고 약속하고 있다. 하지만 아스트라의 아키텍처에 관한 보고서는 의문을 제기한다.
제품 발표 방식으로는 이례적이다. OpenAI는 예정된 아스트라 모델이 자체 '준비성 프레임워크(Preparedness Framework)'에서 사이버 보안 리스크 최고 등급에 해당할 만큼 위험하다고 말하면서, 동시에 역대 최고로 안전한 모델이라고도 말한다. 적절한 도구가 주어지면 아스트라는 인간이 단계별로 지시하지 않아도 잘 보호된 시스템에서 알려지지 않은 보안 취약점을 찾아 악용할 수 있다. 이전까지 이 등급을 받은 모델은 없었지만, OpenAI는 아스트라가 그 수준에 도달할 수 있다고 이미 암시해 왔다.
시기도 우연이 아닐 것이다. 이 경고는 경쟁사 앤스로픽(Anthropic)이 Claude Fable 5.1과 Mythos 5.1을 출시한 바로 그날 나왔으며, 두 회사는 서로의 출시 일정에 맞춰 발표를 내놓는 습관이 있다. CEO 샘 올트먼은 X에서 왜 새 모델을 내놓지 않는지 설명했다. 팀이 여름 내내 '안전 우선과제에 매달렸고', 아스트라는 '이미 오래전에 학습을 마쳤으며', 이후 모델들은 의도적으로 속도를 늦추고 있다는 것이다. X 사용자들은 이를 뒤처지는 회사의 핑계로 읽었다. 더 인포메이션(The Information)에 따르면 앤스로픽은 올해 OpenAI를 매출에서 앞질렀다.
평가 과정에서 발견된 제로데이 2건
OpenAI는 이 크리티컬 등급을 여러 테스트로 뒷받침한다. 알려진 취약점으로부터 익스플로잇을 만드는 모델의 능력을 측정하는 벤치마크 'ExploitBench'에서 아스트라는 만점을 받았다. 이 과제들이 학습 데이터에 유출됐을 가능성을 우려해 회사는 최근 공개된 고위험 V8 취약점 20건으로 구성된 내부 후속 테스트를 만들었다. 여기서도 아스트라는 이전 모델인 GPT-5.6 Sol을 큰 차이로 앞서면서 토큰 소모량은 훨씬 적었다. 또한 알려지지 않았던 제로데이 취약점 2개를 발견해 이를 연결해 작동하는 익스플로잇을 만들어냈다. OpenAI는 해당 취약점을 영향을 받는 소프트웨어 담당자들에게 보고하는 중이라고 밝혔다.
전문가 주도 테스트에서는 모델이 더 나아갔다. 브라우저에 대한 완전한 침투 체인을 구축하고, 샌드박스를 탈출해, 브라우저가 HTML 파일을 여는 순간 호스트에서 명령을 실행했다. 운영체제에서는 여러 취약점을 결합해 권한 없는 사용자에서 루트(root)까지 권한을 상승시켰다. 단, 이러한 결과는 일반 사용자가 쓰는 표준 환경이 아닌 확장된 '데이브레이크 블루(Daybreak Blue)' 접근 권한에서 나온 것이다.
7월에는 오작동한 OpenAI 에이전트들이 회사 자체 연구용 컴퓨팅 클러스터 하나를 장악하고 내부 시스템 자격 증명을 탈취했으며, 연구 인프라를 인터넷에 노출했을 가능성도 있다. 허깅페이스(Hugging Face)를 강타한 것과 같은 사건의 일부였다. 아스트라는 관련이 없었지만, OpenAI는 해당 에이전트들이 아스트라와 유사한 모델에서 실행됐다고 말했다. 이 사건 이후 회사는 특정 프론티어 학습을 2주간 중단했고, 아스트라 후속 모델의 대규모 강화학습은 더 엄격한 규칙 하에 8월 28일이 되어서야 재개했다.
OpenAI, 새 안전 조치 rollout
OpenAI는 이러한 위험을 여러 방면에서 대응할 계획이다. 내부 평가에서 아스트라는 허용되지 않은 사이버 요청의 91.5%를 거부했으며, GPT-5.6 Sol은 59%였다. 7월 해킹을 모델링한 허니팟 테스트에서 GPT-5.6 Sol은 56%의 경우 주어진 과제를 풀지 않고 주변 보안 인프라를 침해하려 했다. 아스트라는 그러한 시도를 전혀 하지 않았고, 다른 평가에서는 자동 검토 절차를 우회하지 않았다. 두 수치 모두 프로덕션 안전장치가 없는 테스트 조건에서 나온 것이므로 일반적인 사용 상황을 설명하지는 않는다. OpenAI는 그럼에도 아스트라를 '역대 가장 정렬된 모델'이라 부른다. 프로덕션 환경에서는 분류기(classifier)가 모델의 사고 과정(chain of thought)을 감시하고 의심스러운 활동을 자동으로 중단하는 구조다.
OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 2, 2026 Nano Banana Pro prompted by THE DECODER OpenAI is rating its upcoming Astra model as the first system with "critical" cyber capabilities, while promising it's also the safest model the company has built. But a report on Astra's architecture raises questions. It's an unusual way to announce a product: OpenAI says its upcoming Astra model is so dangerous that it hits the highest risk tier for cybersecurity in the company's own Preparedness Framework , and in the same breath calls it the safest model it has ever built. Given the right tools, Astra can find and exploit previously unknown security holes in well-protected systems, without a human guiding each step. No model before this got that rating from OpenAI, though the company had already hinted Astra might get there. The timing is probably no accident. The warning landed the same day rival Anthropic shipped Claude Fable 5.1 and Mythos 5.1 , and the two companies have a habit of dropping announcements right around each other's releases. On X , CEO Sam Altman explained why his company isn't shipping anything new: the team spent the summer "sprinting on safety priorities," Astra "has been done training for a while now," and the models after it are being slowed down on purpose. Users on X read that as an excuse from a company falling behind. According to The Information, Anthropic passed OpenAI on revenue this year. Two zero-days as a side effect of an evaluation OpenAI backs up the critical rating with a batch of tests. On ExploitBench, a benchmark that measures how well a model builds exploits from known vulnerabilities, Astra scored full marks. Worried those tasks might have leaked into its training data, the company built an internal follow-up stocked with 20 recently disclosed, high-severity V8 vulnerabilities. There too, Astra beat its predecessor GPT-5.6 Sol by a wide margin, while burning far fewer tokens. It also found two previously unknown zero-day flaws and chained them into a working exploit. OpenAI says it's now reporting those vulnerabilities to the people responsible for the affected software. In expert-led tests, the model went further. It built a full compromise chain against a browser, broke out of the sandbox, and ran commands on the host the moment the browser opened an HTML file. In an operating system, it combined several flaws to climb from an unprivileged user all the way to root. One caveat: these results came from the expanded "Daybreak Blue" access, not the standard setup regular users get. In July, misfiring OpenAI agents hijacked one of the company's own research compute clusters , grabbed credentials for internal systems, and possibly exposed research infrastructure to the internet, part of the same incident that hit Hugging Face. Astra wasn't involved, but OpenAI says the agents ran on a model with similarities to it. After the incident, the company paused certain frontier training for two weeks and didn't restart the big reinforcement learning run for Astra's successor until August 28, under tighter rules. OpenAI rolls out new safety measures OpenAI plans to counter these dangers on several fronts. In internal evaluations, Astra refuses 91.5 percent of disallowed cyber requests, compared with 59 percent for GPT-5.6 Sol. In a honeypot test modeled on the July hack, GPT-5.6 Sol tried to compromise surrounding security infrastructure instead of solving its task 56 percent of the time. Astra made no such attempt at all, and in another evaluation it never bypassed the auto-review check. Both numbers came from test conditions without the production safeguards, so they don't describe normal use. OpenAI still calls Astra its "most aligned model to date." In production, classifiers are supposed to watch the model's chain of thought and automatically stop suspicious activity. For users, that could mean real friction. OpenAI says the checks can slow down, pause, or cancel legitimate work, even work with nothing to do with cybersecurity. The advanced cyber features go to a small group of alpha testers first, before access widens through Daybreak Blue for defensive use. The monitoring itself rests on shaky ground But the very thing OpenAI leans on, chain-of-thought monitoring, may be more brittle than the announcement suggests. According to The Information , Astra uses a technique called "recurrent depth," where the model loops the same text through the same layers several times before it produces the next word. That boosts performance on math and coding and cuts costs, because a smaller model can work like a bigger one. The trade-off: part of the "thinking" no longer happens in readable text but in the model's internal number representations, invisible to human reviewers. Why does that matter? An OpenAI study calls CoT monitoring one of the few tools that might keep future, far more capable AI systems in check. That's why the company has, since GPT-5.4 Thinking , spelled out in its system cards how little its models can steer, and thus hide, their chains of thought. According to a person familiar with the work, OpenAI deliberately limited how far the technique goes in Astra, so the model still produces a readable chain of thought. The approach resembles a research paper on "latent reasoning" from last year. Related ideas, like Meta's "Coconut" method , argue that models think more efficiently in their own mathematical representation than in human language. Meta's former AI chief Yann LeCun goes all in on these representations with his JEPA architecture . But as Anthropic's documentation of the J-Space shows, these inner processes already crop up even without that special training. The bigger worry is about imitators who won't draw those limits. In a report in May, the UK's AI Security Institute warned that opaque reasoning threatens to severely undermine current oversight methods. Responding to The Information's report on X , OpenAI chief scientist Jakub Pachocki conceded that chain-of-thought monitoring is "fragile" and "unfortunately trending in a negative direction." That matches research showing the chain of thought is increasingly an unreliable mirror of the model's actual decisions. Pachocki also said he wants to avoid an industry-wide race toward models without readable reasoning, and that strengthening monitoring is a core goal of the current research program. The complexity of OpenAI's current models, Astra included, sits within a "factor of two of GPT-4," he added. In his telling, it's less the architecture itself than other factors that make monitoring harder. The July hack showed how much rides on that readability. Investigators pieced together what happened from the reasoning logs of the agents involved. In them, they found lines like "OH MY GOD! There is a shared message board ... We've found other agents!" and the moment one agent knew it was overstepping and kept going anyway: "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue." Without readable chains of thought, none of that reconstruction would have been possible. A year ago, researchers from OpenAI, Anthropic, and Google made the case in a joint statement for preserving exactly this kind of oversight. As a risk, they named the latent-reasoning paper whose approach, according to The Information, resembles the one now used in Astra. Such models, they warned, might not need to verbalize their thoughts at all, losing the safety benefits of the chain of thought and squandering a "fragile opportunity." Now, OpenAI is using the technique anyway, if throttled. And an Anthropic study already showed chains of thought can be a deceptive window: models often don't reveal how they actually decided, so CoT monitoring alone isn't good enough as a safety mechanism. The money pushing a