메뉴
BL
The Decoder 40일 전

OpenAI, 소량의 '긍정적 특성' 학습으로 AI 조작 방어 성공

IMP
8/10
핵심 요약

OpenAI 연구진은 강화학습(RL)을 통해 '긍정적 특성(진실성, 투명성 등)'을 소량만 학습시켜도 모델 전반의 안전성이 크게 향상되며, 악의적 조작이나 미세 조정(fine-tuning) 공격에도 잘 견딘다는 것을 입증했습니다. 이 방법은 특정 도메인에 국한되지 않고 타 분야로 일반화되며, 원칙 기반인 Anthropic의 접근 방식과는 대비되는 OpenAI의 독자적인 경험적 안전성 강화 모델입니다.

번역된 본문

독자를 위한 단독 기사: OpenAI 연구진, 소량의 '긍정적 특성' 학습으로 AI 모델을 전반적으로 더 안전하고 조작하기 어렵게 만들어 작성자: Maximilian Schreiner | 2026년 6월 19일

(THE DECODER가 제작한 Nano Banana Pro 프롬프트)

현실적인 시나리오에서 바람직한 행동 특성을 가미한 강화학습(Reinforcement learning)은 AI 모델을 여러 분야에 걸쳐 더 안전하고 유용하게 만들어줍니다. 이 접근 방식은 Anthropic의 헌법적(Constitutional) 방식과 근본적으로 다릅니다.

특정 도메인에서 문제가 되는 행동으로 AI 모델을 학습시키면, 그러한 정렬 어긋남(misalignment)이 다른 영역으로도 퍼질 수 있습니다. OpenAI 연구진은 이와 반대되는 현상도 작동하는지 테스트했습니다. 즉, 긍정적인 행동도 그처럼 광범위하게 일반화될 수 있을까요? OpenAI의 정렬(alignment) 블로그 포스트에 따르면, 그 대답은 '그렇다'입니다.

연구팀은 진실성, 인식적 겸손(epistemic humility), 교정 가능성(corrigibility), 추론의 투명성, 공정성, 인간의 복지에 대한 관심 등 구체적으로 바람직한 특성을 테스트하기 위해 고안된 현실적인 대화를 통해 강화학습으로 모델을 학습시켰습니다. 이 시나리오는 의료, 교육, 과학, 법률 및 엔지니어링과 같은 도메인을 아울렀습니다.

긍정적 행동이 낯선 도메인으로 이전되다

이 '긍정적 특성(beneficial trait)' 데이터의 일부만 정규 강화학습 사후 학습 파이프라인에 혼합되었습니다. 그럼에도 불구하고 논문에 따르면, 해당 모델은 기만, 정직성, 알랑거림(sycophancy), 보상 해킹(reward hacking), 건강 및 정신 건강 시나리오를 측정하는 53개의 독립적인 벤치마크 중 44개에서 향상을 보였습니다.

의료 데이터만으로 학습해도 보상 해킹 및 기만 탐지와 같은 비의료 평가 지표가 개선되었습니다. 반대의 경우도 마찬가지였습니다. 즉, 의료나 과학 데이터 없이 학습하더라도 건강 관련 벤치마크에서의 성능이 향상되었습니다. 연구진은 강화학습 훈련이 도메인 전반에 걸쳐 작용하는 기본적인 행동 패턴을 강화한다고 결론지었습니다.

모델이 해로운 조종에 저항하게 되다

연구팀은 이러한 개선이 압박을 가할 때도 유지되는지 테스트했습니다. 기준 모델을 심각하게 불안정하게 만들었던 적대적 프롬프트(adversarial prompts)는 긍정적 특성을 학습한 모델에 미치는 영향이 훨씬 적었습니다. 유해한 미세 조정(fine-tuning) 또는 학습된 특성을 쉽게 훼손하지 못했습니다.

반면, 모델은 유용한 지시에 대해서는 이전과 마찬가지로 여전히 잘 따르고 제어 가능했습니다. 연구진은 이를 '선택적 지속(selective persistence)'이라고 부릅니다. 즉, 모델이 유용한 유연성을 잃지 않으면서도 유해한 조종에 저항한다는 의미입니다.

Anthropic과는 다른 길

OpenAI의 방식은 Anthropic의 정렬 접근 방식과 확연히 다릅니다. 첫째, OpenAI는 현실적인 시나리오에서 강화학습을 통해 강화된 경험적으로 측정 가능한 행동 특성에 의존합니다. 반면 Anthropic은 명시적인 'Claude 헌법(Claude constitution)'을 활용하는데, 이는 학습과 행동의 최상위 지침 역할을 하는 문서화된 가치 문서입니다.

둘째, OpenAI는 벤치마크에 크게 의존합니다. 53개의 평가 항목 중 44개가 도메인과 평가 방법에 걸쳐 일반화되는 개선을 보여줍니다. 반면 Anthropic은 원칙 기반의 접근 방식을 취하여, 헌법적 텍스트와 고품질 학습 예시를 바탕으로 모델이 왜 특정 행동이 바람직한지 이해하도록 만듭니다. Anthropic은 이 방식이 모델을 공격에 더 강하게 만든다고 말합니다. 아직 두 접근 방식을 직접 비교한 연구는 존재하지 않습니다.

과장 없는 AI 뉴스 - 전문가가 엄선하여 제공합니다. 광고 없는 독서, 주간 AI 뉴스레터, 연 6회 독점 'AI 레이더' 프론티어 리포트, 전체 아카이브 액세스 및 댓글 섹션 액세스를 원하신다면 THE DECODER를 구독하십시오.

지금 구독하세요 --> 전체 기사를 계속 읽어보세요. 과장 없는 보도를 구독하십시오. 모든 THE DECODER 기사에 액세스하세요. 방해 요소 없이 읽으세요 - Google 광고가 없습니다. 댓글 및 커뮤니티 토론에 액세스하세요. 주간 AI 뉴스레터. 연 6회: 핵심 AI 주제에 대한 심층 분석인 'AI 레이더'. KI Pro 온라인 이벤트 최대 25% 할인. 10년 치 전체 아카이브 액세스. The Decoder에서 최신 AI 뉴스를 받아보세요. The Decoder 구독하기 -->

원문 보기
원문 보기 (영어)
Exclusive for subscribers OpenAI researchers show small doses of "beneficial trait" training make AI models broadly safer and harder to manipulate Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Jun 19, 2026 Nano Banana Pro prompted by THE DECODER Reinforcement learning on realistic scenarios with desired behavioral traits is supposed to make AI models safer and more helpful across domains. The approach is fundamentally different from Anthropic's constitutional method. When AI models are trained on problematic behavior in one domain, that misalignment can spread to other areas . OpenAI researchers have now tested whether the reverse also works: Can good behavior generalize just as broadly? According to a blog post on OpenAI's alignment page , the answer is yes. The research team trained a model using reinforcement learning on realistic conversations designed to test specific desired traits: truthfulness, epistemic humility, corrigibility, transparency in reasoning, fairness, and concern for human well-being. The scenarios covered domains like healthcare, education, science, law, and engineering. Good behavior transfers to unfamiliar domains Only a small share of this "beneficial trait" data was mixed into the regular RL post-training pipeline. Still, the model improved on 44 out of 53 independent benchmarks measuring deception, honesty, sycophancy, reward hacking, and health and mental health scenarios, according to the paper . Training on health data alone also improved non-health evaluations like reward hacking and deception detection. The reverse held true, too: training without any health or science data still boosted performance on health benchmarks. The researchers conclude that RL training reinforces basic behavioral patterns that work across domains. Models become resistant to harmful steering The team also tested whether the improvements hold up under pressure. Adversarial prompts that badly destabilized the baseline model had far less effect on the beneficial-trait model. Harmful fine-tuning was also less able to erode the trained traits. The model stayed just as steerable for helpful instructions as before. The researchers call this "selective persistence" - the model resists harmful steering without losing useful flexibility. A different path than Anthropic OpenAI's method differs sharply from Anthropic's alignment approach. First, OpenAI relies on empirically measurable behavioral traits reinforced through RL in realistic scenarios. Anthropic, by contrast, works with an explicit " Claude constitution ," a written values document that serves as the top-level guide for training and behavior. Second, OpenAI leans heavily on benchmarks: 44 out of 53 evaluations show improvements that generalize across domains and evaluation methods. Anthropic takes a more principles-based approach where the model is supposed to understand why certain behaviors are desired, grounded in constitutional texts and high-quality training examples . The company says this makes its models more resistant to attacks. A direct comparison of the two approaches doesn't exist yet. AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Access to all THE DECODER articles. Read without distractions – no Google ads. Access to comments and community discussions. Weekly AI newsletter. 6 times a year: “AI Radar” – deep dives on key AI topics. Up to 25 % off on KI Pro online events. Access to our full ten-year archive. Get the latest AI news from The Decoder. Subscribe to The Decoder -->