OpenAI의 GPT-6 Astra는 벤치마크 평가 기관마다 상반된 평가를 받고 있다. Epoch AI는 169점으로 1위로 꼽았지만 Artificial Analysis는 이전 모델과 동급인 61점에 그쳤다. 그러나 ARC-AGI-3에서 처음으로 평균 인간보다 효율적으로 과제를 수행하면서 ARC Prize의 프랑수아 촐레(François Chollet)는 예상보다 2배 빠른 진전이라며 AGI 도래 전망을 앞당겼다.
번역된 본문
OpenAI의 GPT-6 Astra가 상반된 벤치마크 평가 결과를 받고 있다. Epoch AI는 이 모델을 선두로 평가하는 반면, Artificial Analysis는 이전 세대보다 나을 게 없다고 본다. 가장 큰 놀라움은 ARC-AGI-3에서 나왔는데, Astra가 역대 처음으로 평균적인 인간보다 효율적으로 과제를 수행했다. ARC Prize 책임자 프랑수아 촐레(François Chollet)는 이러한 진전이 자신이 예상했던 것보다 "2배 빠르다"며 AGI 전망 시기를 앞당기고 있다.
두 독립 연구기관이 각각 수십 개의 개별 테스트를 하나의 종합 점수로 묶었지만 정반대의 결론에 도달했다. Epoch AI는 50개 이상의 벤치마크를 종합해 GPT-6 Astra를 169점으로 267개 모델 중 명확히 1위에 올렸다. 반면 Artificial Analysis는 지식, 코딩, 텍스트 이해력을 테스트해 GPT-6 Astra에 61점을 매겼는데, 이는 이전 모델과 정확히 동급이며 66점을 받은 Claude Fable 5.1보다 뒤지는 수치다.
Astra는 자사 이전 모델보다 확실히 비싸다. OpenAI는 처리 텍스트 단위당 2.5배의 비용을 청구하며, 이로 인해 과제 하나의 비용이 Sol 대비 약 75% 더 높다. 그러나 Anthropic와 비교하면 상황이 반전된다. Artificial Analysis 기준으로 코딩 과제에서 Astra는 Claude Fable 5와 같은 점수를 받았지만 과제당 비용은 절반 이하다. 그 이유는 이 모델의 절제된 사용량에 있다. Sol이 쓰는 컴퓨팅 단계의 3분의 1, Opus 5가 쓰는 5분의 1만 필요로 하기 때문이다.
벤치마크
Astra
Sol
Fable 5.1
Opus 5
Epoch ECI 종합 점수
169
162
163
162
AA 지능 지수
61
61
66
63
ARC-AGI-3 (낯선 게임 세계)
62.7%
7.8%
데이터 없음
30.2%
ARC-AGI-2 (추상 시각 퍼즐)
95.0%
92.5%
90.0%
90.4%
ARC-AGI-1 (구버전)
98.5%*
97.5%
97.5%
97.5%
FrontierMath Erdős (미해결 수학)
3%
0%
0%
데이터 없음
*xhigh 추론 강도 기준, max에서는 97.5% 달성. ARC-AGI-1은 현재 대부분 포화 상태로 간주된다.
Coding Agent Index에서는 Sol의 토큰 사용량의 약 3분의 1로 67점을 기록했으며, Fable 5.1이 70점으로 선두다. AA-Omniscience의 환각률은 92%에서 51%로 떨어졌다. 동시에 GDPval-AA v2에서는 약 80 Elo를 잃었고, 뱅킹 지원, SciCode, 롱컨텍스트 추론 과제에서도 하락했다.
Epoch AI에 따르면 새로운 FrontierMath Erdős에서 GPT-6 Astra는 시도당 300달러 예산으로 Lean으로 검증된 증명으로 68개의 미해결 에르되시 문제 중 2개를 푼 유일한 모델이었다. 22만 달 이상의 컴퓨팅을 소모한 비표준 추가 실행에서 3개의 풀이가 더 나왔지만, Epoch은 이는 점수에 반영되지 않는다고 밝혔다.
Epoch의 개별 수치를 자세히 보면 GPT-6 Astra와 경쟁 모델 Fable 5.1이 지금까지 서로 다른 조건에서 측정됐음을 알 수 있다. Astra는 수학, 지식, 퍼즐에서 앞서고 Fable 5.1은 거의 모든 코딩 테스트에서 최고점을 유지하고 있다. 하지만 Epoch은 지금까지 Astra의 코딩 점수를 단 하나만 기록했고, 그마저 중간 추론 수준에서 실행된 결과다.
ARC-AGI-3에서 큰 도약
가장 뚜렷한 도약은 ARC-AGI-3에서 나타났다. 이 테스트는 AI를 규칙과 목표를 아무도 설명해주지 않는 낯선 게임 세계에 투입한다. 모델은 시행착오를 통해 무엇을 해야 하는지 스스로 파악해야 한다. GPT-6 Astra는 약 2만 6천 달러의 테스트 비용으로 62.7%를 달성했다. 이전 모델인 GPT-5.6 Sol은 7.78%, 경쟁 모델 Claude Opus 5는 30.16%에 그쳤다. Fable 5와 Fable 5.1은 아직 이 벤치마크에 참여하지 않았다.
OpenAI가 보고한 99.9%는 다른 조건에서 나온 수치다. 그 설정에서 Astra는 OpenAI가 구축한 하니스(harness)를 사용할 수 있었는데, 이는 개별 요청 사이에 추론 체인을 유지하고 긴 실행을 자동으로 요약해준다. ARC Prize의 측정에 따르면 두 설정이 모두 푼 167개의 게임-추론 쌍을 비교했을 때, 이러한 실행은 자체 하니스 대비 약 3.66배 빠르고 토큰 사용량이 49% 적었다. 이러한 하니스의 사용과 그에 따른 성능 향상은 이미...
Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 4, 2026 GPT-Image-2 prompted by THE DECODER OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front, while Artificial Analysis rates it no better than its predecessor. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress "2x faster" than he expected and is moving up his AGI forecast. Two independent labs each roll dozens of individual tests into a single overall score, but they reach opposite conclusions. Epoch AI combines more than 50 benchmarks and puts GPT-6 Astra clearly in first place with 169 points, ahead of 267 models. Artificial Analysis tests knowledge, coding, and text comprehension, and rates GPT-6 Astra at 61 points, exactly level with its predecessor and behind Claude Fable 5.1 at 66 points. Astra is clearly more expensive than its own predecessor. OpenAI charges two and a half times as much per unit of processed text, which makes a task cost roughly 75 percent more than it did with Sol. Compared with Anthropic, the picture flips. On coding tasks, Astra hits the same score as Claude Fable 5 according to Artificial Analysis, but costs less than half as much per task. The reason is how sparing the model is. It needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses. Benchmark Astra Sol Fable 5.1 Opus 5 Epoch ECI, overall score 169 162 163 162 AA Intelligence Index 61 61 66 63 ARC-AGI-3, unfamiliar game worlds 62.7% 7.8% no data 30.2% ARC-AGI-2, abstract visual puzzles 95.0% 92.5% 90.0% 90.4% ARC-AGI-1, older version 98.5%* 97.5% 97.5% 97.5% FrontierMath Erdős, open math 3% 0% 0% no data *GPT-6 Astra at xhigh reasoning effort; at max it hits 97.5 percent. ARC-AGI-1 is now considered largely saturated. On the Coding Agent Index, it reaches 67 points at roughly a third of Sol's token usage, while Fable 5.1 leads with 70. The hallucination rate on AA-Omniscience drops from 92 to 51 percent. At the same time, the model loses about 80 Elo points on GDPval-AA v2 and slips on banking support, SciCode, and long-context reasoning tasks. Epoch AI reports that on the new FrontierMath Erdős , GPT-6 Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, on a budget of $300 per attempt. Three more solutions came out of non-standardized extra runs that burned through more than $220,000 in compute, but Epoch says those don't count toward the score. A closer look at Epoch's individual numbers shows that GPT-6 Astra and rival Fable 5.1 have so far been measured on different ground. Astra leads on math, knowledge, and puzzles. Fable 5.1 holds the top marks on nearly every coding test. But Epoch has recorded only a single coding score for Astra so far, and that one comes from a run at a medium reasoning level. ARC-AGI-3 shows a big jump The clearest jump comes on ARC-AGI-3 . The test drops an AI into unfamiliar game worlds whose rules and goals nobody explains to it. The model has to figure out what to do by trial and error. GPT-6 Astra reaches 62.7 percent at a test cost of roughly $26,000. Its predecessor GPT-5.6 Sol managed 7.78 percent, and rival Claude Opus 5 got 30.16 percent. Fable 5 and Fable 5.1 aren't on the benchmark yet. The 99.9 percent OpenAI reported came under different conditions. In that setup Astra got to use the harness OpenAI built, which keeps reasoning chains between individual requests and automatically summarizes long runs. By ARC Prize's measurements, those runs went about 3.66 times faster and used 49 percent fewer tokens than runs on the in-house harness, compared across 167 game-reasoning pairs that both setups solved. The use of these harnesses, and the performance jump that comes with them, was already a sticking point between ARC Prize and OpenAI with GPT-5.6 Sol . ARC Prize notes that only the lower figure of 62.7 percent, run on the internal ARC harness, allowed a fair comparison between vendors, though it plans to publish the numbers from vendor harnesses in the future as well. Here, more thinking lowers the bill The relationship between thinking effort and cost is unusual. Normally a higher reasoning level makes a test run more expensive. With Astra it's the opposite. On the standard ARC scaffold, costs drop from $49,791 with no reasoning to $26,098 at maximum reasoning, while the score climbs from 35.2 to 62.7 percent. According to ARC Prize, the reason is that Astra solves the games in fewer moves, which means fewer model calls and fewer tokens. One oddity stands out: the "low" level scores 17.5 percent, worse than running with no reasoning at all. ARC Prize doesn't comment on the outlier, but GPT-6 Astra in other benchmarks showed that it can solve longer-horizon tasks without reasoning due to its new architecture which presumably loops processing internally before generating the first token . For comparison, the human testers got $115 per 90-minute session plus $5 per game solved, so at about nine attempts that works out to roughly $12.78 per game. But that mostly pays for time and willingness to take part. Count only the metabolic energy of the brain as electricity instead, and ARC Prize arrives at 0.067 cents per game. More interesting than the raw score is the efficiency. Before the launch, ARC Prize had about 500 testers play with no pre-screening and recorded, for each level, the median number of moves among those who solved it. On the run with the OpenAI scaffold, Astra cleared 96 percent of levels in fewer moves than that median, on average with a little over half. Unlike the usual cost measures, this figure doesn't track compute consumed. It tracks how much experience with an environment the model needed before it mastered it. This is exactly where the organizers had expected humans to hold a lasting edge. That still holds for brute-force approaches, but with top models ARC Prize sees an almost binary pattern. Once the model has figured out the mechanics, its execution lands in the human efficiency range. Astra invents its own notation To get there, Astra keeps its own notes and works out a self-invented, algebra-like shorthand in which it records objects, coordinates, rules, and open plans, for example extend8 to3; retract10 to2 as an ordered sequence of moves or Turn 5: P=(24,20), empty, facing west as a state note. ARC Prize saw similar behavior from other models, but singles out Astra for its precision and information density. On the standard harness, that's an important skill, because everything the model doesn't save into its own visible notes is lost. ARC co-founder François Chollet describes it on X as "highly efficient, on-the-fly symbolic world modeling for each game and level." The model goes so far as "developing its own shorthand DSL to represent in-game situations," which at its core is "essentially a game-specific algebraic notation." What matters most to Chollet is where this behavior comes from: "Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses, so harness capabilities are increasingly shifting into the model itself." Astra creates a dense, compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: “L8: hub q2 (8↓). Lengths: 14=1…” - It mapped operations to exact controls:… pic.twitter.com/tMHP002mkB — ARC Prize (@arcprize) September 3, 2026 A third test environment demonstrates what Astra is capable of with external tools. PRO-LONG is an agent framework developed by a third party that the ARC Prize team deployed early on as a red-teaming partner for ARC-AGI-3—that is, to systematically explore the limits of the benchmark