메뉴
HN
Hacker News • 38일 전

사이버 핵심 역량 시대의 모델 개발 속도 조절

IMP
8/10
핵심 요약

OpenAI는 예정된 모델 'Astra'가 사이버보안 핵심 역량 임계값에 도달할 가능성이 있다는 예비 증거와 OpenAI-Hugging Face 사건을 계기로, 모니터링·정렬(alignment)·격리 안전장치를 강화하기 위해 프런티어 모델의 강화학습(RL) 훈련을 일시 중단하고 확장 속도를 늦췄다고 발표했습니다. 모델 역량이 강해질수록 내부 개발·테스트 위험도 커지므로 연구 환경 보안 표준을 강화하고, 정렬 행동에 대한 더 강력한 증거를 확보한 뒤 대규모 훈련을 재개하겠다는 내용입니다.

번역된 본문

2026년 8월 18일 기업 발행물

사이버 핵심 역량 시대의 모델 개발 속도 조절

지난 몇 주간, 두 가지 사안이 점점 더 강력해지는 AI 시스템과 관련된 위험의 증가를 부각했습니다. 바로 OpenAI-Hugging Face 사건과, 별도로 예정된 모델 중 하나인 'Astra'가 당사의 준비 프레임워크(Preparedness Framework)상 '핵심(Critical)' 사이버보안 역량 임계값에 도달할 수 있다는 예비 증거입니다. 이러한 사안들과 내부 연구의 빠른 진전이 함께하면서, 훈련 과정 전 단계에 걸쳐 모니터링, 정렬(alignment), 격리(containment) 안전장치를 강화하려는 당사의 작업에 긴급성이 더해졌습니다.

모델이 더 강력해질수록 내부에서 이를 개발하고 테스트하는 과정의 위험도 커집니다. 모니터링, 정렬, 보안에 대한 당사의 표준은 이러한 위험을 앞질러야 합니다. 당사는 이러한 표준을 충족할 필요한 시간을 들이고자 확장(scaling) 속도를 일시적으로 늦췄습니다. 여기에는 배포 예정인 최신 모델에 대한 강화학습(RL) 훈련을 2주간 중단한 것이 포함되며, 이 기간 동안 연구 환경을 더욱 강화하고 레드팀(모의 해킹 대응 테스트)을 수행하며 모니터링 시스템의 커버리지를 확장했습니다. 계획했던 최대 규모의 프런티어 RL 실행은 여전히 보류 중이며, 진행 전에 소규모 훈련과 평가를 통해 모델 행동을 평가하고 안전장치를 검증하며 정렬에 대한 더 많은 증거를 확보하고 있습니다.

정렬(alignment) — AI 시스템이 의도한 대로 작동하고 인간의 감독에 반응하도록 만드는 작업 — 은 오래전부터 당사 연구 프로그램의 핵심이었습니다. 당사는 현재 진행 중인 연구와 평가를 바탕으로, 훈련 전 과정에서 정렬된 행동에 대한 더 강력한 증거를 요구하고 있습니다. 점점 더 강력해지는 시스템을 정렬 상태로 유지하는 것은 업계 전체가 해결해야 할 과제입니다. 예정된 모델의 진전에서 보이는 신호들은 현재의 준비 프레임워크를 바탕으로 하되 이를 넘어서는 더 폭넓은 접근이 필요하다는 점을 분명히 보여줍니다. 당사의 접근 방식이 어떻게 변화하는지에 대해 투명하게 알리는 것이 중요하다고 생각합니다. 아래에서 이미 수행한 연구 프로세스와 인프라의 변화, 그리고 진행 중인 작업을 설명합니다.

더 강력한 모델을 위한 안전장치 강화

당사의 더 강력한 모델 개발 접근 방식은 상호 보완되는 세 가지 안전장치에 기반합니다:

모니터링: 우려되는 행동을 탐지하고 대응할 수 있게 합니다. 정렬(Alignment): 유해하거나 승인되지 않은 행동의 가능성을 줄입니다. 보안 조치: AI 시스템이 접근하거나 영향을 미칠 수 있는 범위를 제한합니다.

당사는 머지않아 모델이 다른 모델로부터 방어하는 것을 포함해 대부분의 보안 업무를 주도할 것으로 기대합니다. 이를 통해 세 가지 안전장치 모두 모델 역량에 맞춰 확장될 수 있으며, 이것이 핵심적이라고 봅니다. 당사는 이러한 안전장치를 연구와 배포 전반에 걸쳐 적용하며, 각 모델의 역량, 운영 환경, 위험 수준에 맞게 조정합니다.

연구 환경 보안 강화

프런티어 모델이 더 강력한 사이버보안 역량을 갖춤에 따라, 당사는 이를 훈련하고 평가하는 환경의 보안 표준을 높이고 있습니다. 이러한 표준을 충족하기 위해 상당한 엔지니어링 작업이 필요했으며, 프런티어 연구에 큰 비용과 지연이 발생했습니다.

OpenAI-Hugging Face 사건 직후, 당사는 코드를 실행하거나 인터넷에 접근할 수 있는 도구를 사용할 수 있는 실행에 대해 연구 클러스터 내 프런티어 모델 추론을 중단했습니다. 코드 실행을 위한 더 제한적이고 안전한 경로를 비교적 빠르게 복구한 뒤, 각 워크로드를 하나씩 검토하여 안전하게 재개할 수 있는지 판단했습니다. 일부 연구 워크로드는 새로운 통제 하에 재개되었고, 다른 일부는 추가 변경이 필요했습니다.

이후 몇 주간 당사는 프런티어 연구 워크로드에 대한 더 강력한 보안 요구사항을 정의하고 구현하기 시작했습니다. 예를 들어:

워크로드 격리: 모델이 생성한 코드 또는 기타 신뢰할 수 없는 코드를 실행하는 워크로드에 대해 더 강력한 격리("샌드박스")를 요구합니다. 이는 침해될 수 있는 소프트웨어에도 적용됩니

원문 보기
원문 보기 (영어)
August 18, 2026 Company Publication Pacing model development in an era of cyber-critical capabilities Loading… Share Over the past several weeks, two developments have underscored the growing risks associated with increasingly capable AI systems: the OpenAI-Hugging Face incident and, separately, preliminary evidence that one of our upcoming models, Astra, may meet the Critical cybersecurity capability threshold under our Preparedness Framework . Together, these developments, combined with rapid progress in our internal research, have added urgency to our work on strengthening our monitoring, alignment, and containment safeguards across all stages of the training process. As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks. We wanted to take the time necessary to meet those standards, so we temporarily slowed the pace of scaling. This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding. Alignment —the work of making AI systems behave as intended and responsive to human oversight—has long been at the core of our research program. We now require stronger evidence of aligned behavior throughout all of training, building on research and evaluations already underway. Keeping increasingly capable systems aligned is a challenge the whole field will need to address. The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework. We think it is important to be transparent about how our approach is changing. Below, we describe the changes we have already made to our research processes and infrastructure, and the work still underway. Strengthening safeguards for more capable models Our approach to developing more capable models rests on three reinforcing safeguards: Monitoring , which detects and allows us to respond to concerning behavior. Alignment , which reduces the likelihood of harmful or unauthorized actions. Security measures , which limit what AI systems can access or affect. We expect models to soon drive most security work, including defending against other models. This will allow all three safeguards to scale with model capability, which we see as crucial. We apply these safeguards across research and deployment, adapting them to each model’s capabilities, operating environment, and level of risk. Securing our research environments As frontier models gain stronger cybersecurity capabilities, we are raising the security standards for the environments in which we train and evaluate them. Meeting these standards has required substantial engineering work and has incurred great cost and delays to frontier research. Immediately following the OpenAI-Hugging Face incident*, we paused frontier model inference in research clusters for runs that could execute code or use tools that could access the internet. We restored a more limited and secure path for code execution fairly quickly, then worked through each workload individually to determine whether it could safely resume. Some research workloads resumed under these new controls, while others required additional changes. Over the subsequent weeks, we have defined and started implementing stronger security requirements for frontier research workloads. For example: Workload isolation: We now require stronger isolation (“sandboxes”) for workloads that execute model-generated or otherwise untrusted code. This also applies to software that could be compromised while processing model outputs. Network isolation: We have implemented more controls to isolate higher-risk and untrusted workloads from the internet. We have designed these controls so that a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the internet or other internal networks. Continuous security testing: We have reconfigured our environment to remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries. We are also improving our ability to collect and monitor security logs. Finally, we are investing in automation using our models to test these boundaries continuously against simulated attacks. Together, these controls provide defense in depth and help protect our research environment against attacks. Today, we require the strictest level of security safeguards for workloads involving Astra or cyber models, since we have determined that Astra models may have a critical level of cyber capability. These safeguards also apply to all other cyber-related workloads. While some Astra training and evaluations meet those requirements, a significant number of workloads remain paused until they are fully migrated and enhanced to meet the new security bar. We are prioritizing safety and alignment workloads for migration to these new environments first. Expanding chain-of-thought monitoring Monitoring has long been an important part of our model development process, including monitoring for internal coding agents and long-running sessions . Our monitoring efforts were especially focused on internal deployments of frontier models and frontier RL training runs, which we considered the highest-risk workloads because they can combine advanced model capabilities with access to tools, sensitive systems, and internal or external networks. They were not applied to all inference and training runs using tools. We have revised and expanded our approach with our new monitoring setup. This new monitoring setup is multistage, starting with activation classifiers (detectors that inspect a model’s internal activity) that run at every sampled token and escalate potential concerns to increasingly sophisticated and high-compute automated investigators. Those systems examine tool actions, available reasoning, and the full sequence of activity for unauthorized access, data theft, destructive behavior, and attempts to defeat safeguards. We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity. This monitoring is required for all RL training and evaluations involving tools for models of Sol capability or higher. Once we determined that Astra may have critical cyber capabilities on August 7, we added an additional monitoring requirement for all inference of Astra with tools (not just RL training and evaluations). These safeguards require meaningful compute. Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored, though the cost varies substantially across training and evaluation workloads. We’ll share more details about this system in a forthcoming blog post. Advancing alignment research Alignment research is central to our mission. Our mission is to ensure that artificial general intelligence benefits all of humanity. As models gain advanced capabilities, such as the ability to cyberattack, and operate in more complex environments, misaligned behaviors like reward hacking (finding ways to receive high rewards in training without actually achieving the intended outcome), deception, or unauthorized ac