메뉴
BL
The Decoder • 58일 전

OpenAI, 자체 API 환경에서 GPT-5.6 Sol 성능이 Claude Opus 5 능가한다고 주장

IMP
7/10
핵심 요약

OpenAI가 자체 응답 API와 최적화된 환경을 적용했을 때 GPT-5.6 Sol 모델이 논리 벤치마크인 ARC-AGI-3에서 Anthropic의 Claude Opus 5보다 높은 점수를 기록했다고 주장했습니다. 하지만 공식 표준 환경에서는 점수가 크게 하락하여, 벤치마크 측정 방식과 API 기술적 세팅에 대한 공정성 논쟁이 촉발되었습니다. AI 모델 간의 벤치마크 비교가 단순한 모델 성능을 넘어 인프라 세팅에 따라 크게 좌우될 수 있음을 보여주는 중요한 사례입니다.

번역된 본문

OpenAI는 자사 모델이 ARC-AGI-3 벤치마크에서 경쟁사를 따라잡을 수 있다고 밝혔습니다. Anthropic의 Claude Opus 5가 논리 추론 벤치마크 기록을 4배나 끌어올린 직후, OpenAI는 GPT-5.6 Sol 모델이 두 가지 API 설정을 적용했을 때 38.3%의 점수를 기록하며 Opus 5의 30.2% 점수를 뛰어넘었다고 발표했습니다.

하지만 OpenAI가 이 성과를 낸 데에는 공식 테스트 환경이 아닌 자체 환경이 사용되었다는 점이 중요합니다. OpenAI는 단계별로 모델의 사고 과정(Chain of Thought)을 유지하는 'Retained Reasoning(추론 유지)' 기능과, 기존 문맥을 단순히 잘라내는 대신 요약하는 'Compaction(압축)' 기능이 포함된 자체 Responses API를 통해 GPT-5.6 Sol을 구동했습니다. 반면 공식 환경에서 GPT-5.6 Sol의 점수는 7.8%에 그쳤는데, 이는 모델이 행동을 취할 때마다 추론 과정이 초기화되기 때문입니다.

OpenAI는 벤치마크가 단순히 모델 자체의 성능만을 측정하는 것이 아니라 그것을 둘러싼 기술적 설정 환경도 함께 측정하는 것이라고 주장합니다. 이는 어느 정도 사실이지만, ARC-AGI-3의 설계 목적 자체는 순수한 모델의 성능을 평가하는 것입니다. ARC 측은 공정한 비교를 보장하기 위해 특정 업체의 맞춤형 설정 없이 표준화된 방식을 사용한다고 밝혔습니다. 현재 논쟁의 핵심은 Claude API가 이미 제공하던 기능들이 결여된 구버전의 'OpenAI 방식 completions API'를 ARC 측이 사용했느냐는 점입니다. 만약 그렇다면 OpenAI에게 불리한 불공정한 비교가 되었을 것입니다.

원문 보기
원문 보기 (영어)
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness Matthias Bastian View the LinkedIn Profile of Matthias Bastian Jul 30, 2026 OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmark , OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent. OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action. OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's true, and ARC-AGI-3 is designed to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons, ARC Prize said in response to OpenAI's results . The sticking point is whether ARC Prize used an older "OpenAI-style completions API" that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI. Ad DEC_D_Incontent-1 Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: OpenAI Ask about this article… Search
관련 소식