메뉴
HN
Hacker News • 44일 전

스페이스스페이스AI 그록 4.6, 저비용 고성능으로 최고 수준 모델 반열

IMP
8/10
핵심 요약

스페이스스페이스AI의 '그록 4.6(Grok 4.6)' 모델이 지능 지수 61점을 기록하며 GPT-5.6과 동등한 최고 수준의 성능을 달성했습니다. 특히 에이전트(AI 자율 작업) 성능이 탁월하면서도 가격은 이전 세대와 동일하게 유지하여 경쟁 모델 대비 압도적인 가성비를 자랑합니다.

번역된 본문

인공지능 분석 지수(AAII)에서 스페이스스페이스AI의 그록 4.6이 61점을 기록했습니다. 이는 GPT-5.6 Sol과 동등한 최고 수준(Frontier)으로, 낮은 비용으로 뛰어난 에이전트 성능을 보여줍니다.

그록 4.6은 출시 약 한 달 만에 지능 지수에서 그록 4.5 대비 5점(그록 4.3 대비 +23점) 상승했습니다. 이를 통해 스페이스스페이스AI는 앤스로픽(Anthropic)에 이어 OpenAI와 함께 지능 최전선에 재합류했습니다.

핵심 요약:

➤ 최고 수준의 인공지능 분석 지수 진입: 점수 61점으로 GPT-5.6 Sol(최고 설정)과 동등하며, Claude Opus 5(최고, 63점) 및 Claude Fable 5(대체 모델 포함 최고, 62점) 다음이자 Kimi K3 바로 앞서는 위치입니다.

➤ 강력한 에이전트 성능: 실제 업무 수행 능력을 평가하는 GDPval-AA v2에서 1753 Elo를 기록해 Claude Opus 5 다음으로 높습니다. 통계적으로 Claude Fable 5 및 Qwen3.8 Max와 동등한 수준입니다. 𝜏³-Banking(은행 다중 턴 고객 서비스)에서는 Qwen3.8 Max(51.3%)와 함께 최상위권인 50.7%를 기록했으며, Terminal-Bench v2.1(터미널 기반 소프트웨어 작업)에서는 88.4%로 최고 수준 모델들과 어깨를 나란히 합니다.

➤ 저비용의 최고 수준 지능: 기본 가격은 입력/출력 100만 토큰당 $2/$6로 그록 4.5와 동일합니다. 이는 Claude Opus 5($5/$25)나 GPT-5.6 Sol($5/$30)보다 60% 이상 저렴합니다. 작업당 비용은 $0.84로 Kimi K3와 동일하지만 더 높은 지능을 제공합니다.

➤ 장기 지식 작업(Long-horizon knowledge work): 장기적인 에이전트 작업을 측정하는 자체 비공개 벤치마크인 AA-Briefcase에서 Elo 1577점(클로드 Opus 5 다음)을 기록했습니다. 눈에 띄는 점은 높은 효율성으로, Claude Opus 5(평균 103턴, 약 20억 입력 토큰 소모)과 비교해 평균 약 53턴, 약 5천만(0.5B) 입력 토큰만으로 작업을 완료합니다.

기타 모델 세부 정보:

➤ 컨텍스트 윈도우는 50만 토큰(그록 4.5와 동일) ➤ 가격은 입력/출력 100만 토큰당 $2/$6; 캐시 적중 시 할인 가격은 100만 토큰당 $0.5로, 그록 4.5의 $0.3에서 인상되었습니다.

에이전트 성능 그록 4.6의 가장 강력한 결과는 단순 추론이 아닌 '에이전트 작업'에서 나타납니다. 실제 환경의 지식 작업을 측정하는 GDPval-AA v2에서 Claude Opus 5 다음으로 높은 Elo(1753)를 기록했으며, 통계적 오차 범위 내에서 Claude Fable 5 및 Qwen3.8 Max와 차이가 없습니다. 이러한 패턴은 모든 작업 유형에서 일관됩니다. 지식 작업, 고객 서비스, 터미널 사용 능력에서 동시에 경쟁력을 갖춘 모델은 드물며, 저렴한 가격과 결합되어 그록 4.6을 모든 에이전트 평가에서 '비용 대비 성능(Pareto frontier)' 1순위 모델로 만듭니다.

비용 최고 수준의 모델에서 성능 향상에도 불구하고 세대 간 기본 가격을 동일하게 유지하는 것은 이례적입니다. 그록 4.6은 지능 지수를 5점 높이면서도 $2/$6 가격을 유지했으며, 이는 우수한 토큰 효율성을 반영하여 작업당 $0.84의 비용으로 측정되었습니다. 구매자에게 중요한 비교 대상은 비슷한 점수를 받은 Claude Opus 5($5/$25) 및 GPT-5.6 Sol($5/$30)입니다. 그록 4.6은 GPT-5.6 Sol과 사실상 동일한 지능 점수를 토큰당 훨씬 낮은 가격에 제공하며, 이는 추론 위주의 작업에서 전체 비용을 좌우하는 핵심 요소입니다.

장기 지식 작업 장기적 에이전트 지식 작업을 평가하는 자체 벤치마크 AA-Briefcase에서 그록 4.6은 1577 Elo로 데뷔했습니다. 이는 Claude Opus 5 제품군 다음이자 Fable 5와 같은 등급으로, 특정 영역의 강점이 다른 영역의 약점을 상쇄한 것이 아니라 채점 기준 준수, 프레젠테이션 품질, 분석 품질 등 모든 영역에서 강력한 성능을 보여줍니다. 점수만큼이나 효율성이 돋보입니다. Claude Opus 5가 평균 약 103턴과 약 20억(2.0B) 입력 토큰을 소모하는 반면, 그록 4.6은 평균 약 53턴과 약 5천만(0.5B) 입력 토큰만으로 작업을 완료합니다.

원문 보기
원문 보기 (영어)
Artificial Analysis K All articles August 12, 2026 SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic. Key takeaways: ➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3 ➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models ➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier ➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max) Other model details: ➤ Context window of 500k tokens (unchanged from Grok 4.5) ➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits Agentic performance Grok 4.6's strongest results are on agentic work rather than static reasoning. On GDPval-AA v2, our leading measure of real-world agentic knowledge work, it scores an Elo of 1753 - behind only Claude Opus 5, and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max given overlapping confidence intervals. The pattern holds across task types. 𝜏³-Banking (50.7%) tests multi-turn customer service with tool use and places Grok 4.6 in the top two, while Terminal-Bench v2.1 (88.4%) puts it level with the leaders on terminal-based software tasks. Few models are simultaneously competitive across knowledge work, customer service and terminal use; combined with its pricing, this places Grok 4.6 on the cost vs. performance Pareto frontier for every agentic evaluation in the Intelligence Index. Cost Holding headline pricing flat across a generation is unusual at the frontier, where intelligence gains have typically been accompanied by price increases. Grok 4.6 delivers a 5-point Intelligence Index gain at unchanged $2/$6 pricing, and our measured cost per task of $0.84 reflects both that pricing and reasonable token efficiency. The comparison that matters for buyers is against the models scoring within two points of it: Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. Grok 4.6 offers effectively the same Intelligence Index score as GPT-5.6 Sol at a fraction of the output token price, which is the dimension that dominates cost in reasoning-heavy workloads. Long-horizon knowledge work Grok 4.6 debuts on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577. This places it at Fable 5-tier, behind the Claude Opus 5 family, with consistently strong performance across rubric grading, presentation quality and analytical quality rather than strength in one dimension offsetting weakness in another. The efficiency profile is as notable as the score. Grok 4.6 resolves tasks in ~53 turns and ~0.5B input tokens on average, against ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). Long-horizon agentic work accumulates context rapidly, so a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage well beyond its per-token pricing. Full results Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index: See Artificial Analysis for further details and benchmarks of Grok 4.6: https://artificialanalysis.ai/models/grok-4-6 Read the latest Upstage Solar Pro 4: Benchmarks and analysis Upstage has released Solar Pro 4 August 12, 2026 NVIDIA launches Nemotron 3.5 Lightning Efficient on-device scale intelligence August 11, 2026 Announcing AA-AnalystAgent: an agentic benchmark for quantitative analysis on real-world spreadsheets and documents AA-AnalystAgent tests agents on 80 private quantitative analysis questions across 14 business and scientific domains, each answered from a folder of real source spreadsheets and documents. Every question is run five times and the headline metric is pass^5 — solved on all five attempts — because an analyst agent that is only sometimes right still has to be checked by hand. August 10, 2026