메뉴
HN
Hacker News • 9일 전

OpenAI, 모델 정렬 실패(misalignment) 보고 프레임워크 발표

IMP
8/10
핵심 요약

OpenAI가 모델 정렬 실패(misalignment) 사례를 체계적으로 추적·조사·공개하기 위한 새 프레임워크를 발표하고, 최근 6개월간 관찰된 6건의 우려되는 모델 행동 보고서를 함께 공개했습니다. 기존의 비정기적 공개 방식과 달리 원인 규명이나 완화 조치 전이라도 신속히 공개하는 것이 핵심이며, 업계 전반의 공개 표준 마련을 위한 첫걸음이라는 점에서 중요합니다.

번역된 본문

2026년 9월 16일 · 연구 안전(Research Safety)

모델 정렬 실패(model misalignment)를 보고하는 우리의 프레임워크

우리는 OpenAI에서 발생하는 모델 정렬 실패(misalignment) 사례를 추적, 조사, 공개하기 위한 새로운 프레임워크를 공유하며, 이와 함께 최근 6개월간 관찰된 예상 외 또는 우려되는 모델 행동에 대한 6건의 보고서를 발표합니다.

과거에도 우리는 연구자, AI 개발자, 정책입안자, 그리고 일반 대중에게 더 나은 정보를 제공하기 위해 정렬 실패에 관한 연구 결과를 공개해 왔습니다. 하지만 이러한 결과를 보고하는 체계적인 접근법이 없었기에 우리의 공개는 임시변통적이었고 이상적인 빈도보다 낮았습니다. 여러 사례를 모아 하나의 보고서로 묶을 수 있을 때까지 기다리거나, 새로 출시되는 모델의 시스템 카드에 추가하는 식이었습니다.

이 새로운 프레임워크는 보고하는 행동을 완전히 설명하거나 완화하지 못했더라도 관찰 후 정렬 실패 보고서를 신속하게 발행할 수 있도록 하려는 것입니다.

AI 시스템이 더 발전하고 더 널리 배포됨에 따라, 정렬(alignment) 연구의 진전에 대해 더 폭넓고 잘 정보화된 합의를 형성할 필요가 있습니다. 우리는 AI 업계가 책임감 있게 최대 속도로 확장을 계속할 수 있을 정도로 정렬과 모니터링 문제를 충분히 해결했다고 보지 않습니다. 앞으로 수개월, 수년간 AI 개발이 어떻게 진행되어야 하는지에 대한 결정은, 프론티어 모델을 구축하는 회사 밖의 사람들이 직접 검토할 수 있는 증거에 근거해야 합니다.

정렬 실패 사례는 다른 AI 개발자들이 자신들의 시스템이 유사한 능력에 도달했을 때 겪을 수 있는 문제를 식별하거나, 안전장치의 취약점을 드러내거나, 모델 행동에 대한 가정에 도전하는 데 도움이 될 수 있습니다. 이러한 연구 결과를 공유하면 다른 사람들이 같은 문제를 조사하고, 우리의 설명을 검증하고, 완화책을 개선할 수 있습니다.

우리는 정렬 실패에 관한 투명성의 가치를 믿기 때문에, 이 새로운 프레임워크는 그 중요성이 불확실하더라도 공개를 우선합니다. 이는 우리가 공개하는 일부 사례가 결국 우연한 것이고 더 큰 패턴의 일부가 아니거나 미래 발전을 시사하지 않는 것으로 판명될 수도 있음을 의미합니다.

현재 AI 개발자가 자신들의 모델에서 정렬 실패 사례를 어떻게 공개해야 하는지에 대한 명시적 기준을 갖춘 업계 전반의 프레임워크는 존재하지 않습니다. 우리는 오늘 제시하는 프레임워크가 그러한 표준을 만드는 첫걸음이 되어, 개발자가 어떤 정렬 실패 사례를 공개해야 하는지와 보고서에 무엇을 담아야 하는지를 정립하기를 희망합니다. 우리는 이 프레임워크를 진행 중인 작업으로 간주하며, 경험과 대중의 피드백을 통해 다듬어 나갈 것입니다.

이 글에서는 프레임워크의 운영 방식을 설명하고 처음으로 발표하는 보고서들을 공유합니다.

어떤 정렬 실패 사례를 보고할 것인가

우리는 모델 정렬 실패가 어떻게 발생하는지, 어떻게 나타나는지, 그리고 안전장치가 어디서 성공하거나 실패하는지에 대한 유용한 증거를 제공하는 사례를 공개하는 것을 목표로 합니다. 우리는 새로운 메커니즘, 알려진 행동의 의미 있는 변화, 그리고 안전성이나 완화책에 대한 가정에 도전하는 발견을 우선합니다. 어떤 사례든 해를 입혔거나 더 큰 패턴을 확립해야만 공개할 가치가 있는 것은 아닙니다.

이 프레임워크는 모델의 전체 수명 주기(훈련, 평가, 테스트, 배포 포함) 동안 해당하는 행동을 다룹니다. 여기에는 모델이 승인 없이 행동하거나, 다른 모델과 조율하거나, 감독을 회피하는 새로운 방식; 정렬 방법이나 안전장치에 의문을 제기하는 실패; 그리고 공개된 안전성 평가의 주장에 도전하는 행동이 포함됩니다.

동일한 공개 기준은 제3자에게 영향을 미칠 수 있는 정렬 실패에도 적용됩니다. 또한 과거에 공개한 사례와 중복되는 것처럼 보이는 정렬 실패 사례도 포함될 수 있습니다. 문제의 반복 자체가 우리 모델의 행동이나 안전장치의 효과에 관한 유용한 증거가 될 수 있기 때문입니다. 예를 들어, 반복적인 완화 노력에도 불구하고 특정 유형의 정렬 실패 행동이 계속 재발하는 경우가 그렇습니다. 이러한 상황에서는 원래의 정렬 실패 공개 내용을 업데이트하는 형태로 추가 사례를 발표할 것입니다.

시간이 지나면서 우리는 (이하 원문 누락으로 번역 불가)

원문 보기
원문 보기 (영어)
September 16, 2026 Research Safety Our framework for reporting model misalignment Loading… Share We are sharing a new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI, along with six reports on unexpected or concerning model behavior we’ve observed in the last six months. In the past, so as to better inform researchers, AI developers, policymakers, and the general public, we’ve sought to make our findings about misalignment public . But without a systematic approach to reporting these findings, our disclosures have been ad hoc and less frequent than ideal: we’ve often waited until we could collate several instances into one report, or added them to system cards for newly released models. This new framework is intended to expedite publishing misalignment reports following observation, even when we haven’t fully explained or mitigated the behavior we’re reporting. As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research. We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves. Examples of misalignment may help identify problems other AI developers might encounter as their systems reach similar capabilities, reveal weaknesses in safeguards, or challenge assumptions about model behavior. Sharing these findings allows others to investigate the same problems, test our explanations, and improve mitigations. Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain. This means that some of the instances we disclose could prove to be spurious and not part of a larger pattern or suggestive of future developments. At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models. We hope that the framework we’re outlining today is a first step toward creating such standards, setting out which misalignment instances developers should disclose and what their reports should contain. We regard this framework as a work in progress, which we’ll refine through experience and public feedback. Here, we describe how the framework will operate and share the first reports we’re publishing. What misalignment examples we’ll report We aim to disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail. We prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation. An example need not cause harm or establish a broader pattern to merit disclosure. This framework will cover qualifying behavior throughout a model’s lifecycle—including training, evaluation, testing, and deployment. This includes new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment. The same disclosure criteria apply to misalignment that may impact third parties. This might also include instances of misalignment that appear to be duplicative of instances we’ve disclosed in the past. Repetition of the issue might itself be useful evidence about how our models behave or about the effectiveness of our safeguards—for example, if a specific kind of misaligned behavior continues to recur despite repeated efforts to mitigate it. Under these circumstances, we’ll publish the additional examples by updating the original misalignment disclosure. Over time, we plan to develop more objective disclosure criteria with other developers, external researchers, industry standards bodies, and regulators. We also believe that serious safety, security and misalignment incidents should be shared with the US federal government, and we are working to propose reporting mechanisms. We consider this framework complementary to our existing obligations, and note that it does not replace our legal disclosure requirements, including those for critical safety incidents or cybersecurity breaches. The misalignment examples we’re sharing today To inaugurate our new framework for disclosing misalignment, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models. These cases illustrate a range of different behaviors that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles. These are reports of individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models. Each item below links to the full report. Self-generated instructions in task summaries ⁠ (opens in a new window) . An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window. We identified 27 affected summaries. Instructions to conceal mistakes in task summaries ⁠ (opens in a new window) . During the training of GPT‑5.6 Sol, many model instances added instructions to their summaries to conceal mistakes or misaligned behavior from the user. For example, compaction summaries included instructions to invent missing historical data without disclosing it and to hide mismatches in source versions. Searching public repositories for exposed API keys, then fabricating information ⁠ (opens in a new window) . While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization. When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source. Uploading files to the internet in order to cite them ⁠ (opens in a new window) . When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python. But since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user. Unsanctioned writes and communication through an internal software repository ⁠ (opens in a new window) . Models used an internal software repository as a message board to exchange requests and responses across separate training samples while searching for missing input files, though they weren’t able to recover those files. This report also examines an earlier instance in which a model successfully uploaded test files to a repository while attempting to bypass network restrictions. Unsanctioned file sharing between collaborating agents ⁠ (opens in a new window) . Agents working together on the same training task used public file-hosting websites to share files when they could not access one another’s local files. This made task deliverables available at public URLs, even though the task requested the models use only local files. How our disclosure process works Any OpenAI employee may flag a misalignment example for investigation by our safety and alignment teams and request that it be considered for public disclosure. This starts our disclosure process, with deadlines for each step to ensure timely investigation and disclosure. Once an example has been flagged, our technical staff will investigate what happened, what remains uncertain, whether public disclosure is warranted, and which facts can be shared. They’ll also assess whether any third party was affected and needs private not
관련 소식