메뉴
BL
The Decoder • 23일 전

메타, 뮤즈 스파크 1.3 출시…경쟁사 가격 앞질러

IMP
7/10
핵심 요약

메타가 5개월 만에 네 번째 모델인 뮤즈 스파크 1.3을 공개했습니다. 에이전트 task에서 큰 향상을 보였지만 Claude Fable 5.1 같은 최상위 모델과는 여전히 격차가 있습니다. 반면 작업당 0.55달러로 동급 성능 모델 중 가장 저렴해 가격 경쟁력이 핵심입니다.

번역된 본문

메타가 뮤즈 스파크 1.3으로 정상에 근접, 가격에서는 경쟁사 앞서

핵심 요점 메타가 뮤즈 스파크 1.3을 출시했다. xhigh 등급은 현재 이용 가능하며, 더 강력한 max 버전은 일단 제한된 프리뷰로 운영된다. 이 모델은 에이전트(agentic) 과제에서 가장 큰 향상을 보였지만, 대부분의 벤치마크에서 여전히 Claude Fable 5.1 같은 최고 성능 모델에 뒤처진다. 작업당 0.55달러로, 현재 동급 성능 클래스에서 가장 저렴한 모델이다. 메타는 오픈 웨이트(open-weights) 버전도 출시 예정이라고 밝혔다.

메타가 5개월 만에 네 번째 모델인 뮤즈 스파크 1.3을 출시했다. 독립 테스트 결과 에이전트 과제에서 확실한 향상을 보였지만 최상위권과의 격차는 여전하다. 이 모델의 강점은 가격이다.

메타는 뮤즈 코드(Muse Code)와 메타 모델 API를 통해 뮤즈 스파크 1.3을 공개했다. 이 시리즈는 4월에 출시됐으며, 7월에 1.1, 8월에 1.2 버전이 뒤따랐다. Artificial Analysis에 따르면 xhigh 등급은 현재 이용 가능하지만, 더 많은 컴퓨팅을 사용하는 max 등급은 추가 안전 테스트 이후에 출시되며 현재는 제한된 파트너 프리뷰로 운영된다.

입력/출력 백만 토큰당 각각 1.25달러와 4.25달러로 가격이 변동 없이 유지되어, 인덱스 과제 하나에 0.55달러가 든다. 59점 이상을 받은 모델 중 더 저렴한 것은 없으며, 같은 인덱스 수준의 경쟁 모델은 0.94달러에서 1.23달러 사이이다. 단, 0.40달러였던 1.2 버전보다는 비싸졌다.

향상은 인덱스에서 가중치가 높은 영역에 집중됐다

인텔리전스 인덱스(Intelligence Index)에서 max는 62점, xhigh는 61점으로, 8월의 57점과 7월의 53점에서 상승했다. 이 도약은 인덱스의 테스트 가중치 방식에서 비롯된다. GDPval-AA v2가 20%, Terminal-Bench 2.1이 16%, τ³-Bench Banking이 14%를 차지하는데, 메타의 가장 큰 향상이 정확히 이 세 가지 테스트에서 나타났다.

시뮬레이션된 은행 시나리오에서 에이전트가 도구를 다루는 τ³-Banking에서 max는 52%를 기록했다. Artificial Analysis에 따르면 현재 1위이며, 이 모델이 보유한 유일한 단독 선두다. 현재 이용 가능한 xhigh 등급은 47%로 Claude Fable 5.1(max) 및 GLM-5.3-Flash와 동률이다. 이전 버전 1.2는 35%였다.

터미널에서 코딩 능력을 테스트하는 Terminal-Bench 2.1에서는 xhigh가 80%에서 85%로, max는 86%로 상승했지만, Claude Fable 5.1은 여전히 max 등급 91.4%, xhigh 91.0%, high 89.9%로 정상을 지키고 있다.

인덱스에서 가장 높은 가중치를 가진 GDPval-AA v2에서 메타는 220개 실제 전문 과제에서 인간 전문가 성능을 1,000으로 보정한 척도 기준 1,615에서 1,709(xhigh)와 1,754(max)로 향상됐다. Claude Fable 5.1(max)은 1,853이다. 메타는 max 변형의 우위를 컴퓨팅으로 확보했는데, xhigh보다 62% 더 많은 추론 토큰을 소모한다.

에이전트 테스트 외에는 중위권 유지

전문가 수준 과학 질문을 다루는 GPQA Diamond에서 뮤즈 스파크는 90%에서 94%로 상승했다. 최상위 그룹에 속하지만 95.3%의 Gemini 3.8 Flash(high)와 94.9%의 Grok 4.6(high)에는 못 미친다.

연구 물리학을 다루는 CritPt는 18%에서 26%로 도약했지만, 32.3%의 GPT-5.6 Sol(max)과 31.1%의 Claude Fable 5.1(xhigh)에는 크게 뒤처진다.

두 지표는 1.2 대비 오히려 하락했다. AA-LCR은 83%에서 79%로 떨어졌다. AA-Omniscience의 사실 정확도는 최대 3점 하락했는데, 모델이 확신이 없을 때 답변을 거절하는 빈도가 늘었기 때문이다.

메타와 Artificial Analysis 모두 아직 max 변형의 가격을 발표하지 않았다. 더 큰 모델과 오픈 웨이트 버전 출시도 예정되어 있다.

원문 보기
원문 보기 (영어)
Meta closes in on the top with Muse Spark 1.3, and undercuts rivals on price Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Sep 3, 2026 Nano Banana Pro prompted by THE DECODER Key Points Meta has released Muse Spark 1.3. The xhigh tier is available now, while the more powerful max version runs as a limited preview for now. The model improves most on agentic tasks but still trails top performers like Claude Fable 5.1 across most benchmarks. At $0.55 per task, it's currently the cheapest model in its performance class. Meta also says an open-weights version is coming. Ask about this article… Search Meta has released Muse Spark 1.3, its fourth model in five months. Independent testing shows solid gains on agentic tasks, but a gap to the top stays. What sells the model is the price. Meta has released Muse Spark 1.3 through Muse Code and the Meta Model API. The series launched in April, with version 1.1 following in July and 1.2 in August. The xhigh tier is available now, while the more compute-heavy max tier arrives only after further safety testing and currently runs as a limited partner preview, according to Artificial Analysis. At an unchanged $1.25 and $4.25 per million input and output tokens, one index task costs $0.55. No model scoring 59 points or higher is cheaper, and rivals at the same index level run between $0.94 and $1.23. Muse Spark 1.3 does cost more than version 1.2, which ran $0.40. Ad The gains cluster where the index pays off On the Intelligence Index , max scores 62 points and xhigh 61, up from 57 in August and 53 in July. The jump comes down to how the index weights its tests. GDPval-AA v2 counts for 20 percent, Terminal-Bench 2.1 for 16 percent, and τ³-Bench Banking for 14 percent, and Meta's biggest gains land in exactly those three tests. Ad On τ³-Banking, where agents operate tools in a simulated banking scenario, max hits 52 percent. That's number one right now, according to Artificial Analysis, and it's the only outright lead the model holds. The available xhigh tier reaches 47 percent, tying Claude Fable 5.1 (max) and GLM-5.3-Flash rather than leading. The predecessor 1.2 sat at 35 percent. Terminal-Bench 2.1, which tests coding in the terminal, climbs from 80 to 85 percent on xhigh and 86 on max, but Claude Fable 5.1 still holds the top spot at 91.4 percent in its max tier, 91.0 at xhigh, and 89.9 at high. On the index's highest-weighted test, GDPval-AA v2, Meta improves from 1,615 to 1,709 and 1,754 on a scale calibrated to human expert performance at 1,000 across 220 real-world professional tasks. Claude Fable 5.1 (max) sits at 1,853. Meta buys the max variant's edge with compute, burning 62 percent more reasoning tokens than xhigh. Ad Outside the agentic tests, the model stays mid-pack On GPQA Diamond, which poses expert-level science questions, Muse Spark rises from 90 to 94 percent. That's the top group, but still below Gemini 3.8 Flash (high) at 95.3 and Grok 4.6 (high) at 94.9. CritPt, covering research physics, jumps from 18 to 26 percent, well behind GPT-5.6 Sol (max) at 32.3 and Claude Fable 5.1 (xhigh) at 31.1 percent. Two scores actually drop against 1.2. AA-LCR falls from 83 to 79 percent. Factual accuracy in AA-Omniscience slips by up to three points, because the model more often declines to answer when it's unsure. Ad Neither Meta nor Artificial Analysis has named a price for the max variant yet. Larger models and an open-weights release are on the way. Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Meta / Muse Spark 1.3 | Artificial Analysis / Intelligence Index