메뉴
HN
Hacker News • 2일 전

Mercury 2.5 LLM, 초당 770 토큰 속도 달성

IMP
6/10
핵심 요약

Inception이 출시한 독점 모델 Mercury 2.5가 초당 770 토큰의 출력 속도로 동급 모델 중 2위를 기록했습니다. 지능 지수는 12점으로 평균(중앙값 13)보다 다소 낮지만, 비슷한 가격대 모델과 비교하면 가격 대비 효율이 좋고 간결한 답변을 생성합니다. 26만 토큰의 컨텍스트 윈도우와 추론(reasoning) 기능을 지원합니다.

번역된 본문

Artificial Analysis K Inception • 독점 모델 • 2026년 9월 출시

Mercury 2.5 지능, 성능 및 가격 분석 비교 체험해보기 API 제공업체 벤치마크

모델 요약 지능: 업데이트됨, 순위 #91/175, Artificial Analysis 지능 지수 12점 (지능 항목 4단위 중 2단위) 속도: 순위 #2/175, 초당 출력 770.4 토큰 (속도 항목 4단위 만점) 비용: 순위 #21/175, 입력 $0.25, 출력 $0.75, 캐시 할인 90%, 지능 지수 과제당 비용 $0.06 (비용 항목 4단위 중 2단위) 간결성: 순위 #14/175, 지능 지수에서 생성된 출력 토큰 3,500만 개 (간결성 항목 4단위 중 2단위)

비교 요약 Mercury 2.5는 지능 면에서 평균 이하지만, 비슷한 가격대의 다른 모델들과 비교하면 가격 대비 가치가 좋습니다. 또한 눈에 띄게 빠르고 상당히 간결한 편입니다. 이 모델은 텍스트 입력을 지원하고 텍스트를 출력하며, 26만 토큰의 컨텍스트 윈도우를 갖추고 있습니다.

Mercury 2.5는 Artificial Analysis 지능 지수에서 12점을 받아 동급 모델(중앙값 13) 중 평균 이하에 위치합니다. 지능 지수 평가 시 3,500만 토큰을 생성했는데, 이는 중앙값 8,500만 토큰에 비해 상당히 간결한 수치입니다.

가격은 100만 입력 토큰당 $0.25(중간 수준, 중앙값 $0.25), 100만 출력 토큰당 $0.75(중간 수준, 중앙값 $0.90)입니다. 지능 지수에서 Mercury 2.5를 평가하는 과제당 평균 비용은 $0.06입니다. 초당 770 토큰의 속도로 Mercury 2.5는 현저히 빠른 편입니다(109).

기술 사양 추론(Reasoning): 예 — 이 페이지는 이 모델의 추론 버전을 보여줍니다. 비추론 변형도 존재할 수 있습니다. 입력 형식: 텍스트 지원 출력 형식: 텍스트 지원 컨텍스트 윈도우: 26만 토큰 (12포인트 Arial 글꼴 A4 페이지 약 390페이지 분량)

이 등급의 모델은 총 175개입니다. 지표는 같은 등급의 모델들과 비교됩니다: 비추론 모델 → 다른 비추론 모델과만 비교 추론 모델 → 추론 및 비추론 모델 전체와 비교 오픈 웨이트 모델 → 같은 규모 등급의 다른 오픈 웨이트 모델과만 비교: 초소형(Tiny): 4B 파라미터 이하 소형(Small): 4B40B 파라미터 중형(Medium): 40B150B 파라미터 대형(Large): 150B 파라미터 초과 독점 모델 → 독점 및 오픈 웨이트 모델 전체와 비교하되, 입력:출력 3:1 혼합 가격 비율 기준으로 같은 가격대 내에서 비교: 100만 토큰당 $0.15 미만 100만 토큰당 $0.15~$1 100만 토큰당 $1 초과

하이라이트 (업데이트됨) 지능: Artificial Analysis 지능 지수 · 높을수록 좋음 속도: 초당 출력 토큰 · 높을수록 좋음 과제당 비용: 지능 지수 과제당 가중 평균 비용(USD) · 낮을수록 좋음 프롬프트 옵션

지능 (업데이트됨) Artificial Analysis 지능 지수 v4.3.2는 10개 평가를 포함합니다: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 674개 모델 중 28개 표시, 특정 제공업체의 모델 추가

지능 지수 방법론에는 각 평가의 세부 내역과 실행 방법 등 자세한 정보가 있습니다. 오픈 웨이트/독점, 추론/비추론, 텍스트 전용/멀티모달 입력별로 분류된 지능 지수 비교도 제공되며, 독점 모델은 오픈 웨이트(상업적 사용 제한 포함) 모델과 함께 가격대별로 비교됩니다.

원문 보기
원문 보기 (영어)
Artificial Analysis K Inception • Proprietary model • Released September 2026 Mercury 2.5 Intelligence, Performance & Price Analysis Compare Try it out API Provider Benchmarks Model summary Intelligence Updated # 91 / 175 12 Artificial Analysis Intelligence Index 2 out of 4 units for Intelligence. Speed # 2 / 175 770.4 Output tokens per second 4 out of 4 units for Speed. Cost # 21 / 175 In $0.25 Out $0.75 Cache Discount 90% $0.06 Cost per Intelligence Index task 2 out of 4 units for Cost. Verbosity # 14 / 175 35M Output tokens from Intelligence Index 2 out of 4 units for Verbosity. Comparison Summary Mercury 2.5 is below average in intelligence, but well priced when comparing to other models of similar price. It&#x27;s also notably fast and fairly concise. The model supports text input, outputs text, and has a 260k tokens context window. Mercury 2.5 scores 12 on the Artificial Analysis Intelligence Index, placing it below average among comparable models (median: 13). When evaluating the Intelligence Index, it generated 35M tokens, which is fairly concise in comparison to the median of 85M. Pricing for Mercury 2.5 is $0.25 per 1M input tokens (moderately priced, median: $0.25) and $0.75 per 1M output tokens (moderately priced, median: $0.90). On average, it costs $0.06 per task to evaluate Mercury 2.5 on the Intelligence Index. At 770 tokens per second, Mercury 2.5 is notably fast (109). Technical specifications Reasoning Yes This page shows the reasoning version of this model. A non-reasoning variant may also exist. Input modality Supports: text Output modality Supports: text Context window 260k ~390 A4 pages of size 12 Arial font 175 models in this class Metrics are compared against models of the same class: Non-reasoning models → compared only with other non-reasoning models Reasoning models → compared across both reasoning and non-reasoning Open weights models → compared only with other open weights models of the same size class: Tiny: ≤4B parameters Small: 4B–40B parameters Medium: 40B–150B parameters Large: >150B parameters Proprietary models → compared across proprietary and open weights models of the same price range, using a blended 3:1 input/output price ratio: <$0.15 per 1M tokens $0.15–$1 per 1M tokens >$1 per 1M tokens Highlights Updated Intelligence Artificial Analysis Intelligence Index · Higher is better Speed Output tokens per second · Higher is better Cost per Task Weighted average cost (USD) per Intelligence Index task · Lower is better Prompt Options Intelligence Updated Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity&#x27;s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 28 of 674 models Add model from specific provider Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity&#x27;s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights / Proprietary Reasoning / Non-Reasoning Text Only / Multimodal Inputs Artificial Analysis Intelligence Index by Open Weights / Proprietary Artificial Analysis Intelligence Index v4.3.2 incorporates 10 evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity&#x27;s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 28 of 674 models Add model from specific provider Proprietary Open Weights Open Weights (Commercial Use Restricted) Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity&#x27;s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. Open Weights Indicates whether the model weights are available. Models are labelled as &#x27;Commercial Use Restricted&#x27; if commercial use is limited by conditions, and as &#x27;Non-commercial&#x27; if the license prohibits commercial use. Capability Indexes Measures the performance of models on specific capabilities and industries Finance & Accounting Strategy & Ops Legal Engineering Economics Artificial Analysis Finance & Accounting Index Incorporates 7 evaluations: AA-Omniscience, GDPval-AA v2.1, AA-Briefcase v1.1, Humanity&#x27;s Last Exam, AutomationBench-AA, AA-LCR v1.1, GDP.pdf · Higher is better 28 of 171 models Add model from specific provider Benchmarks Intelligence Evaluations Intelligence evaluations measured independently by Artificial Analysis · Higher is better Coding Agentic Tool Use Private Dataset User Interaction Finance Medical Legal Intelligence Index Long Context Multimodal Instruction Following Faithfulness Writing Business See more 18 of 26 evaluations 28 of 674 models Add model from specific provider AA-Briefcase v1.1 Updated Agentic knowledge work, (Elo-500)/2000 GDPval-AA v2.1 Updated Agentic real-world work tasks, (Elo-500)/2000 AutomationBench-AA Updated Agentic SaaS workflows Terminal-Bench 4.0 New Agentic coding & terminal use SciCode Coding Humanity&#x27;s Last Exam Reasoning & knowledge GDP.pdf New Professional document reasoning, All-pass CritPt Under review Physics reasoning AA-Omniscience Accuracy Knowledge AA-Omniscience Non-Hallucination Rate 1 - hallucination rate AA-LCR v1.1 Long context reasoning Harvey LAB-AA Legal agentic work, criterion pass rate EnterpriseOps-Gym-AA Agentic business operations AA-AnalystAgent Quantitative analysis on spreadsheets & documents 𝜏³-Banking Agentic tool use ITBench-AA Kubernetes incident root-cause analysis MMMU-Pro Visual reasoning MLCR-AA New Medical long context reasoning Intelligence Evaluation Relevance While model intelligence generally translates across use cases, specific evaluations may be more relevant for certain use cases. Artificial Analysis Intelligence Index Artificial Analysis Intelligence Index v4.3.2 includes: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity&#x27;s Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1 . See Intelligence Index methodology for further details, including a breakdown of each evaluation and how we run them. AA-Briefcase v1.1 Updated AA-Briefcase Elo AA-Briefcase Rubric Score (%) Analytical Quality & Presentation Elo AA-Briefcase Elo AA-Briefcase v1.1 is an agentic knowledge work benchmark developed by Artificial Analysis. AA-Briefcase Elo is a combined metric that aggregates rubric pass rate, analytical quality Elo and presentation Elo · Higher is better 28 of 192 models Add model from specific provider AA-Briefcase Elo AA-Briefcase Elo is a combined metric that aggregates analytical quality Elo, presentation Elo, and rubric pass rate, with rubric performance converted into Elo via synthetic head-to-head matches. Elo and 95% confidence interval bounds are clamped at 0. AA-Omniscience AA-Omniscience Index AA-Omniscience Accuracy AA-Omniscience Hallucination Rate AA-Omniscience Index AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more incorrect than correct. 28 of 549 models Add model from specific provider AA-Omniscience Index AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. Scores range from -100 to 100, where 0 means as many correct as incorrect answers, and negative scores mean more