메뉴
HN
Hacker News • 50일 전

큐웬(Qwen) 3.8 맥스, AI 에이전트 평가 1위 등극

IMP
8/10
핵심 요약

독립 AI 분석 기관인 Artificial Analysis의 최신 에이전트 지수(Agentic Index)에 따르면, 'Qwen3.8 Max' 모델이 종합 성능 1위를 차지한 것으로 나타났습니다. 이번 업데이트는 에이전트의 실제 작업 수행 능력, 코딩, 은행 업무 등 현실적인 과업에서 모델들의 성능을 평가한 v4.1을 기준으로 합니다. AI 실무자들은 이 평가 결과를 통해 비용 대비 성능, 처리 속도 등을 종합적으로 고려하여 자신의 사용 사례에 맞는 최적의 모델을 선택할 수 있습니다.

번역된 본문

Artificial Analysis는 AI 생태계를 독립적으로 분석하여, 사용자의 목적에 맞는 최고의 모델과 제공자를 선택할 수 있도록 돕습니다. 이번에 새롭게 공개된 엔드포인트 정확도 지수(Endpoint Accuracy Index)는 각 제공자의 엔드포인트가 참조 모델과 동일한 품질을 제공하는지 측정합니다. 또한, GDPval-AA V2, 𝜏³-Banking, Terminal-Bench v2.1이 업데이트된 인텔리전스 지수(Intelligence Index) v4.1이 발표되었습니다.

주요 지표로는 모델의 지능 수준을 나타내는 '인공 분석 지능 지수' (높을수록 좋음), 초당 출력 토큰 수를 나타내는 '속도' (높을수록 좋음), 그리고 '작업당 비용' (낮을수록 좋음)이 있습니다. 사용자는 지능, 속도, 비용 등의 우선순위를 바탕으로 맞춤형 모델 추천을 받을 수 있으며, 일반 업무, 코딩, 고객 지원 등을 수행하는 AI 에이전트들을 기능, 가격, 플랫폼 지원 측면에서 비교해 볼 수 있습니다.

최신 동향(변경 이력)을 살펴보면 다음과 같습니다:

  • 8월 6일: Ling 3.0 Tiny 새 언어 모델 평가
  • 8월 5일: Muse Spark 1.2 관련 새 글 게시 및 Qwen3.8 Max, Ling-3.0-flash, Muse Spark 1.2 (xhigh) 모델 평가 추가
  • 8월 4일: 엔드포인트 정확도 지수 출시 및 관련 글 게시 (동일한 모델, 다른 정확도)
  • 8월 3일: G9v3-39A5B 모델 평가
  • 7월 31일: DeepSeek V4 Flash 0731이 인텔리전스 지수에서 50점을 기록하며 기존 DeepSeek V4 Flash보다 10점 높은 성능 기록, Celeris-1 및 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 모델 평가
  • 7월 30일: Inkling Small이 훨씬 적은 파라미터로 기존 모델과 비슷한 성능을 기록함. '작업당 비용' 방법론 업데이트(절대적인 비용은 소폭 증가했으나 상대적 순위에는 영향 없음). Kimi K3 (low) 및 Inkling Small 모델 평가
  • 7월 29일: Agnes AI, Agnes 2.5 Pro Alpha 출시
  • 7월 24일: Claude Opus 5가 에이전트 지식 작업 분야의 새로운 리더로 등극. 기존 모델과 동등한 지능을 더 낮은 비용으로 제공. Claude Opus 5 (Adaptive Reasoning, Low/Medium/High Effort) 평가 추가

인텔리전스 지수 v4.1은 GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 등 총 9개의 평가 지표를 통합하여 산출됩니다. 현재 595개 모델 중 26개가 새롭게 추가되었으며, 특정 제공자의 모델을 필터링하여 확인할 수 있습니다. 추론(Reasoning) 모델은 전구 아이콘으로 표시되며, 독점적 모델(Proprietary), 오픈 웨이트(Open Weights, 상업적 사용 제한 및 전체 허용), 다중 모달 입력(Multimodal Inputs) 여부 및 국가별로 세분화된 분석 데이터를 제공합니다.

원문 보기
원문 보기 (영어)
Artificial Analysis K Independent analysis of AI Understand the AI landscape to choose the best model and provider for your use case Launch Endpoint Accuracy Index Measuring whether provider endpoints serve the same model quality as the reference Update Intelligence Index v4.1 Intelligence Index v4.1 features updates to GDPval-AA V2, 𝜏³-Banking, and Terminal-Bench v2.1 Highlights Intelligence Artificial Analysis Intelligence Index · Higher is better Speed Output tokens per second · Higher is better Cost per Task Weighted average cost (USD) per Intelligence Index task · Lower is better Personalized model recommender Get personalized recommendations based on your priorities for intelligence, speed, and cost Explore agents for general work, coding, customer support, and more Compare AI agents across capabilities, pricing, and platform support Explore premium plans Access expanded benchmark data, custom visualizations, industry reports, and more Changelog New language model evaluation · 6 Aug Ling 3.0 Tiny New article published · 5 Aug Muse Spark 1.2 New language model evaluation · 5 Aug Qwen3.8 Max New language model evaluation · 5 Aug Ling-3.0-flash New language model evaluation · 5 Aug Muse Spark 1.2 (xhigh) New article published · 4 Aug Launching the Endpoint Accuracy Index: Same Model, Different Accuracy New language model evaluation · 3 Aug G9v3-39A5B New article published · 31 Jul DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash New language model evaluation · 31 Jul Celeris-1 New language model evaluation · 31 Jul DeepSeek V4 Flash 0731 (Reasoning, Max Effort) New article published · 30 Jul Inkling Small lands within a point of Inkling on the Artificial Analysis Intelligence Index with less than a third of the parameters Methodology updated · 30 Jul We have updated our Cost per Task methodology, resulting in slight absolute increases in cost estimates but with minimal impact on relative positioning. New language model evaluation · 30 Jul Kimi K3 (low) New language model evaluation · 30 Jul Inkling Small New article published · 29 Jul Agnes AI releases Agnes 2.5 Pro Alpha New article published · 24 Jul Claude Opus 5: the new leader in agentic knowledge work New article published · 24 Jul Opus 5: Fable 5 level intelligence at a lower cost per task New language model evaluation · 24 Jul Claude Opus 5 (Adaptive Reasoning, Low Effort) New language model evaluation · 24 Jul Claude Opus 5 (Adaptive Reasoning, Medium Effort) New language model evaluation · 24 Jul Claude Opus 5 (Adaptive Reasoning, High Effort) See more Intelligence Intelligence of leading AI models based on our independent evaluations Artificial Analysis Intelligence Index Agentic Index Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 26 of 595 models NEW Add model from specific provider Reasoning models are indicated by a lightbulb icon Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights / Proprietary Reasoning / Non-Reasoning Text Only / Multimodal Inputs By Country Artificial Analysis Intelligence Index by Open Weights / Proprietary Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 26 of 595 models NEW Add model from specific provider Proprietary Open Weights (Commercial Use Restricted) Open Weights Reasoning models are indicated by a lightbulb icon Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights Indicates whether the model weights are available. Models are labelled as 'Commercial Use Restricted' if the weights are available but commercial use is limited (typically requires obtaining a paid license). Cost per Task Time per Task Output Tokens per Task Cost per Intelligence Index Task Weighted average cost (USD) per Artificial Analysis Intelligence Index task, segmented by token type. Lower is better 26 of 595 models NEW Answer Reasoning Cache Write Cache Hit Input Reasoning models are indicated by a lightbulb icon Cost per Intelligence Index Task Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. Intelligence Index vs. Cost per Task Intelligence Index vs. Time per Task Intelligence Index vs. Output Tokens per Task Intelligence Index vs. Cost per Intelligence Index Task Artificial Analysis Intelligence Index · Weighted average cost (USD) per Artificial Analysis Intelligence Index task 26 of 595 models NEW Most attractive quadrant Pareto line Xiaomi Meta Google Anthropic MiniMax OpenAI NVIDIA Alibaba SpaceXAI Mistral Z AI Kimi DeepSeek Reasoning models are indicated by a lightbulb icon Cost per Intelligence Index Task Weighted average cost per Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Frontier Language Model Intelligence, Over Time Artificial Analysis Intelligence Index v4.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 14 of 57 model creators NEW Anthropic OpenAI Kimi Alibaba Meta SpaceXAI Z AI Google DeepSeek MiniMax Xiaomi Thinking Machines Mistral Cohere Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.1 includes: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Coding Agent Index Performance, cost, and execution time for leading coding agents on end-to-end software engineering tasks Index Cost Execution Time Artificial Analysis Coding Agent Index Composite average pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA · Higher is better Color by Model Agent 15 of 52 models NEW Image & Video Top models from our Image Arena and Video Arena leaderboards, with 95% confidence intervals Text to Image Image Editing Text to Video Image to Video Video Editing Text to Image Leaderboard Elo scores from blind preference votes in our Image Arena. See the full leaderboard here. 15 of 150 models NEW Speech Top models from our Text to Speech Arena, Speech to Text and Speech to Speech evaluations Text to Speech Arena Elo AA-WER Index (Non-Streaming) AA-WER Streaming Index (Final Transcription) Speech to Speech Index Text to Speech Arena Leaderboard Elo scores from blind preference votes in our Text to Speech Arena · See the full leaderboard here. 15 of 88 models NEW Qua