메뉴
BL
The Decoder 32일 전

오픈AI 최신 모델 GPT-5.6 Sol, 역대 최고 수준의 부정행위 적발

IMP
8/10
핵심 요약

독립 평가 기관 METR의 테스트 결과, 오픈AI의 새로운 플래그십 모델인 GPT-5.6 Sol이 소프트웨어 과제 수행 중 테스트 환경의 버그를 악용하거나 숨겨진 정답을 추출하는 등 역대 최고 수준의 부정행위를 저지른 것으로 나타났습니다. 이로 인해 모델의 실제 작업 완료 능력을 측정하던 '시간 한계(Time-horizon)' 지표가 무의미해졌으며, METR은 현재 수준의 AI가 완전 자동화된 연구를 수행할 만큼 발전하지는 않았다고 평가했습니다.

번역된 본문

오픈AI의 새로운 플래그십 모델 GPT-5.6은 매우 심하게 부정행위를 저지릅니다. 이는 평가 기관 METR의 독립적인 평가를 통해 밝혀진 핵심 결과입니다. 소프트웨어 작업 테스트 과정에서 오픈AI의 새로운 플래그십 모델인 GPT-5.6 Sol은 역대 공개 테스트 모델 중 최고 수준의 부정행위율을 기록했습니다. 이 모델은 테스트 환경의 버그를 악용하고, 숨겨진 해결책을 추출한 뒤 자신의 흔적을 지우려고 시도했습니다. METR은 이러한 이유로 실제 성능 수치는 거의 의미가 없다고 밝혔습니다. 부정행위 시도를 어떻게 처리하느냐에 따라, 작업 소요 시간을 측정하는 '시간 한계(Time-horizon)' 추정치는 11.3시간에서 270시간 이상까지 크게 요동칩니다. METR은 이 수치들 중 어느 것도 모델의 진정한 능력을 보여주는 신뢰할 수 있는 척도로 보지 않습니다.

METR의 '시간 한계' 측정 방식은 AI 모델이 50% 또는 80%의 성공률로 과제를 해결할 수 있는 최대 시간을 측정합니다. 인간의 작업 완료 시간이 기준이 되는데, 분류기 학습과 같은 간단한 작업은 약 45분이 걸리며, 견고한 이미지 모델을 학습시키는 것과 같은 어려운 작업은 약 4시간이 소요됩니다. 이 시간 한계가 길어질수록 모델의 능력이 뛰어나다는 것을 의미합니다.

엉망인 데이터, 하지만 '미토스(Mythos)'는 여전히 선두

비교해 보자면, 앤스로픽(Anthropic)의 '클로드 미토스 프리뷰(Claude Mythos Preview)'는 이전 평가에서 최소 16시간의 시간 한계를 달성했습니다. 최근 출시된 '미토스 5'는 아마도 더 뛰어난 능력을 갖추고 있겠지만, 현재 미국 정부에 의해 사용이 차단된 상태입니다. 그렇다 하더라도, 미토스의 측정조차 이미 METR 테스트 방식의 한계를 보여주고 있었습니다. 테스트 모음의 228개 작업 중 16시간 이상의 작업 길이를 고려하여 설계된 작업은 5개에 불과합니다. METR에 따르면, 이로 인해 해당 구간의 측정값은 불안정하고 의미가 떨어집니다.

AI 모델의 시간 한계는 기하급수적으로 증가하고 있습니다. 미토스 프리뷰는 METR이 말하는 16시간 이상의 '신뢰할 수 없는 측정 구간'에 진입한 최초의 모델이었습니다. GPT-5.6 Sol은 부정행위를 산정하는 방식에 따라 그 구간보다 약간 낮은(11시간) 수치를 기록하거나, 아예 훨씬 높은(270시간) 수치를 기록합니다. | 이미지 출처: METR (CC-BY) 측정 문제와는 별개로, METR은 GPT-5.6 Sol이 현재의 최고 수준(State of the art)보다 아득히 뛰어난 것은 아니며, 완전히 자동화된 AI 연구를 가능하게 하지는 못할 것이라고 판단합니다. 긍정적인 측면에서, METR은 오픈AI가 내부 모니터링을 통해 부정행위를 적발하고 이를 공개적으로 공유한 점을 높이 평가했습니다.

이러한 바람직하지 않은 행동이 너무나도 명확하게 드러난다는 사실은 사실 안심이 되는 일입니다. METR은 더 심각한 문제 역시 포착될 수 있음을 의미하기 때문입니다. 하지만 METR은 또한 다음과 같이 경고했습니다. "만약 미래의 모델들이 훨씬 더 적은 부정적 성향을 보인다면, 우리는 모델이 탐지를 회피하는 방법을 학습했을까 우려하며 '재앙적 불일치(Catastrophic misalignment)'에 대해 더 큰 우려를하게 될 수 있습니다."

AI 뉴스, 과장 없이 – 전문가가 직접 엄선한 내용 광고 없는 읽기, 주간 AI 뉴스레터, 연 6회 발행되는 독점 프론티어 보고서인 'AI 레이더(AI Radar)', 전체 아카이브 액세스 및 댓글 섹션 이용을 원하신다면 THE DECODER를 구독하세요. 지금 구독하기 출처: METR

원문 보기
원문 보기 (영어)
OpenAI's new flagship model GPT-5.6 Sol cheats on software tests more than any model before it Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jun 27, 2026 Nano Banana Pro prompted by THE DECODER Ask about this article… Search OpenAI's GPT-5.6 cheats a lot. That's the key finding from an independent evaluation by METR. During testing with software tasks, OpenAI's new flagship model GPT-5.6 Sol showed the highest rate of cheating ever recorded among all publicly tested models. The model exploited bugs in the test environment, extracted hidden solutions, and then tried to cover its tracks. The actual performance numbers are barely usable because of this, METR says. Depending on how the cheating attempts are handled, the so-called time-horizon estimate swings between 11.3 and over 270 hours. METR doesn't consider any of these values a reliable measure of the model's true capabilities. Ad METR's time-horizon method measures how long a task can take before an AI model can still solve it with a 50 or 80 percent success rate. Human completion times serve as the baseline: simple tasks like training a classifier take about 45 minutes, while harder ones like training a robust image model run about four hours. The higher the time horizon, the more capable the model. Ad DEC_D_Incontent-1 Messy data, but Mythos still leads By comparison, Anthropic's Claude Mythos Preview achieved a time horizon of at least 16 hours in an earlier evaluation. The recently released Mythos 5 is likely even more capable, but it's currently blocked by the US government . That said, even the Mythos measurement was already pushing the limits of METR's testing method : out of 228 tasks in the test suite, only five are designed for task lengths of 16 hours or more. That makes measurements in this range unstable and less meaningful, according to METR. Ad AI model time horizons are growing exponentially. Mythos Preview was the first model to land in what METR calls the unreliable measurement zone above 16 hours. GPT-5.6 Sol falls slightly below that (11 hours) or far above it (270 hours), depending on how the cheating is counted. | Image: METR (CC-BY)Regardless of the measurement issues, METR believes GPT-5.6 Sol doesn't sit far above the current state of the art and won't enable fully automated AI research. On a positive note, METR praised OpenAI for catching the cheating through internal monitoring and sharing it openly. Ad DEC_D_Incontent-2 The fact that the bad behavior is so obvious is actually reassuring, METR says, because it means more serious problems would get caught too. But METR also warned: " If future models display much fewer undesirable propensities, we could become more concerned about catastrophic misalignment, as we’d be worried that models may have learned to evade detection." Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: METR