메뉴
BL
The Decoder • 15일 전

GPT-6 Astra, 수학자들에게 숨돌릴 틈 줘…오픈AI는 의도된 결과라고 밝혀

IMP
7/10
핵심 요약

OpenAI의 GPT-6 Astra가 미해결 수학 문제 벤치마크 ErdosBench에서 1위를 차지했으나, 수석 과학자 Jakub Pachocki는 수학 최적화를 의도적으로 우선하지 않았다고 밝혔다. 이는 AI 능력이 최적화하는 영역에서만 집중적으로 성장하는 '뾰족한(spiky)' 발전 경향을 시사하며, OpenAI는 재귀적 자기개선(RSI)과 자동 정렬 연구에 자원을 쏟고 있다.

번역된 본문

GPT-6 Astra는 수학자들에게 숨돌릴 틈을 줬고, OpenAI는 이것이 의도된 설계라고 말한다

Matthias Bastian, 2026년 9월 10일

OpenAI의 GPT-6 Astra가 미해결 수학 문제 벤치마크인 ErdosBench에서 1위를 차지했지만, 수석 과학자 Jakub Pachocki는 회사가 수학을 우선하지 않았다고 말한다. 이는 'AGI'가 실제로 어느 수준에 있는지를 보여주는 대목이다.

수학자들은 Astra 출시 이후 짧은(아마 매우 짧은) 숨돌릴 틈을 얻었다. 수학계는 최근 실존적 질문과 씨름해왔는데, 새 모델이 수학 연구를 앞으로 밀어붙이긴 하지만 일부가 희망했고 다른 이들이 두려워했던 속도로는 움직이지 않고 있다.

ulam.ai의 ErdosBench에서 Astra는 1위를 차지했다. 이 벤치마크는 유명한 에르되시(Erdős) 문제에서 영감을 받은 226개의 미해결 수학 문제를 다룬다. Astra는 3.23점을 기록했으며 106개 문제를 풀었고, 그중 43개는 완전히 해결했다. 또한 27개의 문제를 반례를 통해 기각(부정)했다. 최대 추론 모드에서 78개 문제를 푼 Sol과 비교하면, Astra는 더 강력한 과학적 글쓰기 능력을 보였고 과장된 주장에 덜 치우쳤다. 오히려 자신의 결과를 과소 평가하는 경우도 있었다. 전체적으로 벤치마크 개발자 Przemek Chojecki는 "테스트된 다양한 수학 연구 기량에서 확실한 5%~10% 향상"이라고 평가했지만, 벤치마크는 여전히 포화 상태와는 거리가 멀다.

OpenAI는 수학 최적화를 선택하지 않았다

OpenAI는 Astra를 수학 연구에서 훨씬 더 강하게 만들 수 있었지만 그러지 않기로 했다. 첫 발표에서 수학 성과를 최전면에 내세웠음에도 말이다. OpenAI 수석 과학자 Jakub Pachocki는 에세이 'An Alien Mind'에서 "추가적인 집중을 통해 모델을 수학 연구에 특화해 더 뛰어나게 만들 수 있다고 믿지만, RSI(재귀적 자기개선)와 자동 정렬 연구의 긴급성 때문에 이 방향을 우선하지 않는다"고 썼다.

즉, 현재 가장 능력 있는 수학 모델은 목표한 최적화의 결과가 아니라 다른 우선순위의 부산물이다. OpenAI는 재귀적 자기개선과 미래 AI 시스템의 안전 확보에 자원을 쏟고 있다. Pachocki는 "이것이 앞으로 AI 연구의 최전선에 머무는 유일한 방법이라고 믿는다"고 썼다. (주석: 에세이 발표 이후 OpenAI는 더 뛰어난 내부 수학 모델을 학습시킨 것으로 알려졌는데, 이는 복잡한 이야기이며 RSI와 AI 안전성 주제는 폭발적으로 커졌다.)

Pachocki의 발언은 또 다른 이유에서도 흥미롭다. 들쭉날쭉한 AI 개발의 최전선에서 어려운 최적화 트레이드오프가 이루어지고 있음을 보여주기 때문이다. 수학을 지배하는 모델이 반드시 다른 모든 것을 지배하는 것은 아니다.

이는 케임브리지 연구자 Adam Hunt의 시각화를 떠오르게 한다. 그는 AI의 두 가지 가능한 경로를 대비시켰다. '주류 AGI' 가설은 모델이 모든 인간 과제를 포괄할 때까지 점진적이고 폭넓게 발전한다고 본다. 반면 대안 시나리오는 점점 더 '뾰족해지는(spiky)' 궤적을 묘사한다. 코딩이나 수학 같은 소수 영역에서는 극단적인 강점을 보이지만, 언어 품질이나 상식, 사회적 추론 같은 영역에서는 정체되거나 오히려 퇴보하는 것이다. 그렇다면 우리가 갖게 되는 것은 대부분의 사람들이 범용 인공지능이라 부를 만한 것이 아니라, 고도로 특화된 모델이다. 물론 원한다면 뭐든 AGI라고 부를 수는 있다.

Pachocki가 OpenAI가 Astra에서 수학 최적화를 의도적으로 건너뛰었다고 지적한 것은 '뾰족한' 가설의 증거다. 최고의 AI 연구소조차 모든 방향에서 동시에 최대한의 발전을 이룰 수는 없다. 능력은 최적화하는 곳에서 성장하며, 수학을 더 강하게 만들면 다른 곳을 줄여야 한다. 다만 RSI 우선의 배후에는 모델이 언젠가 스스로 그런 최적화를 수행해 수학을 포함한 전 영역에서 더 빠르게 확장하리라는 희망이 자리하고 있다.

수학은 AI와, 그리고 자기 자신과 씨름 중이다

Astra가 이전 세대보다 5% 나았든 50% 나았든, 더 깊은 질문은 남는다. 문제를 푸는 데는 도움이 되지만 이해하는 데는 반드시 도움이 되지 않는 새로운 컴퓨팅 파워의 세계와 씨름하는 분야에게 말이다. 수학자 테렌스 타오(Terence Tao)는 2026 국제수학자대회에서 이 문제를 제기했다.

원문 보기
원문 보기 (영어)
GPT-6 Astra gives mathematicians a breather, and OpenAI says that's by design Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 10, 2026 GPT-Image-2 prompted by THE DECODER OpenAI's GPT-6 Astra tops the ErdosBench for open math problems, even though chief scientist Jakub Pachocki says the company didn't prioritize math. That says something about where "AGI" actually stands. Mathematicians can take a brief ( likely very brief ) breather after Astra's release. The field has been grappling with existential questions lately, and while the new model does push math research forward, it's not moving at the pace some had hoped for and others had feared. On ulam.ai 's ErdosBench , Astra took first place. The benchmark covers 226 open math problems inspired by the famous Erdős problems. Astra scored 3.23 and solved 106 problems, 43 of them completely. It disproved 27 more. It also disproved 27 others. Compared to Sol, which solved 78 problems at maximum reasoning, Astra showed stronger scientific writing and was less prone to overblown claims. In some cases, it actually understated its own results. Overall, benchmark developer Przemek Chojecki called it "a solid 5%-10% gain on various math-research skills tested," but the benchmark is far from saturated. OpenAI chose not to optimize for math OpenAI could have made Astra much stronger in math research but decided against it, even though the company had put math wins front and center in its first announcement. In his essay "An Alien Mind," OpenAI chief scientist Jakub Pachocki writes, "[…] we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later." That means the most capable math model right now isn't the result of targeted optimization. It's a byproduct of other priorities. OpenAI is pouring its resources into recursive self-improvement and securing future AI systems, since "we believe it is the only way to remain at the frontier of AI research moving forward," Pachocki writes. (Note: Since the essay was published, OpenAI has reportedly trained better internal math models , a complicated story, and the topics of RSI and AI safety have exploded .) Pachocki's statement is interesting for another reason, too: it shows that hard optimization trade-offs are being made at the jagged frontier of AI development. A model that dominates math won't necessarily dominate everything else. That reminded me of a visualization by Cambridge researcher Adam Hunt , which contrasts two possible paths for AI. The "mainstream AGI" thesis assumes models improve gradually and broadly until they cover all human tasks. The alternative scenario describes an increasingly "spiky" trajectory, with extreme strength in a few domains like coding and math but stagnation or even regression in areas like language quality, common sense, or social reasoning. That would give us a highly specialized model, not something most people would call Artificial General Intelligence. Of course, you can call anything AGI if you feel like it . Pachocki's point that OpenAI deliberately skipped math optimization for Astra is evidence for the "spiky" thesis. Even the leading AI lab can't push maximum progress in every direction at once. Capabilities grow where you optimize, and making math stronger means cutting back somewhere else. Behind the RSI priority, though, is the hope that the model will eventually make those optimizations itself, scaling faster across the board, including in math. Math is wrestling with AI and with itself Regardless of whether Astra is 5 or 50 percent better than its predecessor, the deeper question remains for a field that's grappling with a new world of compute power that helps solve problems but doesn't necessarily help understand them. Mathematician Terence Tao raised it at the 2026 International Congress of Mathematicians: if AI models keep producing proofs faster than humans can check them, the field risks shifting from proof scarcity to proof overload. The critical task would then no longer be solving problems but deciding which results actually matter. Tao says math faces a crisis of its values and practices, one he compares to the foundational upheaval of the early 20th century. The hardest problems in math remain unsolved for now, which should buy the discipline some time. The direction still seems set, though, even if some mathematicians don't think language models can deliver real breakthroughs because they lack human-like creativity . For similar reasons, there are also doubts about AI's potential for genuine self-improvement . AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now --> Read on for the full picture. Subscribe for hype-free coverage. Full access to every article on THE DECODER No ads Join the comments and community discussions A weekly AI news recap via mail 6x/year: "AI Radar" — deep dives on the AI topics that matter most Daily AI news, always up to date Our full ten-year archive Covered by a team with 10+ years in AI Subscribe to The Decoder -->