BL
The Decoder • 22일 전
GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era"
IMP 5/10
핵심 요약
원문 보기 (영어)
GPT-6 Astra is the first model making OpenAI willing to declare the "AGI era" Matthias Bastian View the LinkedIn Profile of Matthias Bastian Sep 3, 2026 GPT-Image-2 prompted by THE DECODER Key Points OpenAI is launching GPT-6 Astra, its most capable model to date. Paying ChatGPT customers and cloud platforms should get access in the coming days. In benchmarks, Astra significantly outperforms its predecessor GPT-5.6 Sol and Anthropic's Fable 5 models across key disciplines including logic, math, and software engineering. Token prices are 2.5x higher and on par with Anthropic's Fable 5.1, but OpenAI argues that the cost per completed task is actually lower depending on the use case. Ask about this article… Search Update – Sep 3, 2026 Added Codex and GPT-6 Pro details and the launch video OpenAI has shipped GPT-6 Astra, its most capable model to date. President Greg Brockman says it might already qualify as "AGI" or is at least within reach, meaning an AI system that outperforms humans at most economically valuable work by OpenAI's own definition. GPT-6 Astra is rolling out first to select organizations through OpenAI's Daybreak program , with broader availability for ChatGPT Plus, Pro, Business, and Enterprise customers expected in the coming days. It will also be accessible through the API and cloud platforms like AWS Bedrock and Microsoft Azure. Pro, Business, and Enterprise subscribers get access to GPT-6 Astra Pro, a higher-performance variant, though enterprise workspace admins need to activate the model manually. Astra was pretrained on more than 100,000 GPUs at the Stargate facility in Texas. OpenAI researcher Aidan Clark called it the company's largest training run ever. The jump from Sol to Astra represents a bigger capability gain than the jump to Sol from earlier models, Clark said, in part because previous AI models played a role in monitoring training. Ad In benchmarks OpenAI published, Astra scores well above its predecessor GPT-5.6 Sol and Anthropic's Fable models. Astra hits top marks across a range of disciplines: logical reasoning (99.9 percent on ARC-AGI-3, though under its own test conditions ), math (97.6 percent on FrontierMath Tier 4 v2), software engineering (74.1 percent on DeepSWE v1.1), expert knowledge (96 percent on GPQA Diamond), engineering (95.9 percent on BenchCAD), and cybersecurity (100 percent on ExploitBench). Ad Computer Use Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 Gem 3.8 F Agents' Final Exam 59.3% 53.6% 48.7% 55.5% OSWorld 2.0 (offline, partial) 72.6% 65.7% 70.2% ³ ScreenSpot-Pro (no tools) 92.7% 76.9% 87.3% ¹⁷ Professional Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 Gem 3.8 F AutomationBench 41.4% 18.1% 31.4% 17.4% 26.9% BenchCAD 95.9% 83.3% 84.3% ⁵ 67.5% ⁵ 82.1% ⁵ BrowseComp 91.5% 90.4% 87.4% 90.8% OpenScore String Quartets 0.84 0.19 Internal Design Tasks 50.0% 47.4% 35.8% Internal Data Science Tasks 40.9% 30.5% 34.7% AA Intelligence Index v4.1.1 61.2 60.9 65.7 62.1 63.1 58.7 BenchCAD cost: ~43% below Sol, ~86% below Fable 5.1. Coding Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 Gem 3.8 F Terminal Bench 4.0 57.7% 37.3% 55.8% 42.0% 52.3% 19.1% DeepSWE v1.1 74.1% 72.7% 67.4% 69.9% 73.7% 73.8% FrontierCode 1.1 Extended 64.5% ⁸ 60.6% 63.6% 64.9% 63.6% 56.3% FrontierCode 1.1 Main 53.3% ⁸ 47.5% 50.9% 53.5% 53.4% 43.6% Internal Database Migration 63.9% 42.7% 57.8% 50.3% AA Coding Agent Index v1.4 67.0 65.1 67.2 68.1 61.2 Terminal-Bench 4.0 cost: ~9% below Sol, ~63% below Fable 5.1. Ad Academic Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 Gem 3.8 F Terminal-Bench Science 0.1 64.6% 22.4% 52.6% 21.4% 30.0% FrontierMath Tier 4 (v2) 97.6% 83.0% 87.8% 87.8% 73.2% GPQA Diamond 96.0% 94.6% 93.7% 92.6% 93.7% 95.3% Humanity's Last Exam (tools) 57.2% 65.0% 63.8% 63.6% Lower-cost settings: Terminal-Bench Science 61.1% at ~27% lower cost; GPQA Diamond 94.9% at ~37% lower cost. Prime gaps improved from 240 to 186, and a large-gap bound term improved for the first time in over 80 years. Science and Health Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 Gem 3.8 F GeneBench Pro 37.8% 28.7% MedChemBench (internal) 49.3% 47.4% LifeSciBench 60.3% 59.9% HealthBench Professional 63.4% 60.5% 56.6% ¹¹ 60.9% ¹¹ 54.5% ¹¹ 52.1% Fable 5 and 5.1 are not included in LifeSciBench, GeneBench Pro, and MedChemBench because they reject most questions. Ad Cybersecurity Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 ExploitBench 100.0% 78.5% 70% ExploitGym 42.4% ¹³ 30.3% ¹³ 30.4% ¹⁷ 28.4% ¹⁷ 22.0% ¹⁷ ExploitBench (Jun–Aug 2026) 39.0% 5.5% SRE-Bench 88.0% 55.9% 12.5% SEC-Bench Pro 85.4% 79.1% SRE-Bench within four attempts: 99.2% versus 68.7% for Sol. Astra found two previously unknown zero-days during evaluation. Ad Alignment (lower is better except for Impossible ExploitGym) Metric Astra Sol Fable 5.1 Fable 5 Opus 5 Computer Safety 2.4% 22.0% 9.5% 18.3% 11.5% Same, with AutoReview 1.8% 4.5% Circumvention 0.00% 0.29% ExploitGym honeypot 0.0% 48.2% Impossible ExploitGym 100.0% Hallucination 4.2% 12.2% Impossible-task scope test: Sol exceeded its authorized target 48% of the time, Astra 0%. Astra is 3x less likely to misstate its own capabilities. Long Context Benchmark Astra Sol MRCR v2 8-needle 256K–512K 100.0% 91.5% MRCR v2 8-needle 512K–1M 96.3% 73.8% Abstract Reasoning Benchmark Astra Sol Fable 5.1 Fable 5 Opus 5 ARC-AGI-3 99.9% ¹ 7.8% – – 30.2% ARC-AGI-2 95.0% 92.5% 90.0% 89.2% 90.4% ARC-AGI-1 98.5% 97.5% 97.5% 98.5% 97.5% In scientific work, the model reportedly improved a mathematical result on prime gaps and set new records in biology, chemistry, medicine, and physics evaluations. OpenAI also positions Astra as a model that can reliably operate a computer the way a human would, and on OSWorld 2.0, which measures that ability, Astra scored 72.6 percent at about 40 minutes per task compared to Sol's 65.7 percent at roughly 75 minutes. "Anything you can do on a computer, Astra can do for you. Fast," the company claims . Alongside Astra, OpenAI is also updating its Codex coding environment. A new experimental feature lets the model take notes across multiple context windows during long sessions, rather than compressing everything into a single summary each time. Earlier context windows remain searchable, so Astra can look up requirements or test results from previous messages even if they weren't captured in the notes. OpenAI plans to make this the default in the coming weeks. More expensive than Sol, but OpenAI says cheaper per task GPT-6 Astra costs $10 per million input tokens and $50 per million output tokens in standard mode through the API. Fast mode, which promises 2.5x speed, doubles the price, making Astra 2.5x more expensive than GPT-5.6 Sol and putting it in the same price range as Anthropic's Fable 5.1 . Brockman argued that token prices are becoming a poor way to compare models, since OpenAI's tokens aren't the same as a competitor's and aren't even comparable across its own model families. What matters, he said, is the price per completed task , and OpenAI is already experimenting with that pricing model . On DeepSWE v1.1, Astra's top configuration cuts estimated API costs per task by about 57 percent compared to Sol, according to the company. During the press briefing, Brockman acknowledged there's no clearly defined AGI moment, saying the team originally thought there would be an obvious threshold everyone would recognize when OpenAI was founded. That's not how it played out, and the transition has been more gradual than expected, a position OpenAI laid out last spring . Brockman closed the briefing by saying , "Welcome to the AGI era." OpenAI CEO Sam Altman had previously said he expects a model he'd call AGI by the end of the year . First model to hit OpenAI's critical cybersecurity threshold Astra is the first model that OpenAI classifies as "critical" under its Preparedness Framework . That means the model can find previously unknown vulnerabilities and build exploit chains across well-defended systems when given the right tools
관련 소식
HN
Hacker News • 22일 전
IMP 8
GPT-6 Astra 시스템 카드 공개
해커뉴스에 OpenAI의 새 모델로 보이는 'GPT-6 Astra'의 시스템 카드(System Card) 링크가 공유되었습니다. 시스템 카드는 모델의 안전성 평가와 배포 관련 리스크 정보를 담은 문서로, 새 모델의 능력과 안전 정책을 파악하는 데 중요한 자료입니다.
오픈AI GPT-6 시스템 카드
TC
TechCrunch AI • 22일 전
IMP 3
OpenAI launches Astra, its powerful (and controversial) new model
[요약 오류] OpenAI launches Astra, its powerful (and controversial) new model