메뉴
BL
The Decoder • 50일 전

알리바바 큐웬3.8 맥스, 클로드와 동급...하지만 비용·환각 문제는 여전

IMP
7/10
핵심 요약

알리바바의 최신 모델인 Qwen3.8 Max가 지능 지수에서 Claude Opus 4.8과 맞먹는 수준으로 성능이 크게 향상되었습니다. 하지만 과도한 토큰 처리량 증가로 인해 오히려 작업당 비용은 크게 상승했으며, 모르는 질문을 억지로 추측하면서 환각(Hallucination) 현상도 급증한 한계를 보입니다.

번역된 본문

알리바바의 Qwen3.8 Max가 'Artificial Analysis 지능 지수(Intelligence Index)'에서 56점을 기록하며, 기존 Qwen3.7 Max(46점) 대비 10점이나 끌어올렸습니다. Artificial Analysis에 따르면, 이 점수로 인해 이 모델은 Claude Opus 4.8과 동등한 수준에 올랐고 GLM-5.2(51점)를 앞섰지만, 동시에 25% 더 저렴한 비용으로 운영되는 Kimi K3(57점)에는 뒤처졌습니다.

업무 관련 작업 벤치마크인 GDPval-AA에서 Qwen은 468 Elo 포인트가 상승한 1,739점을 기록하며 Kimi K3(1,685점)를 제쳤습니다. 오직 Claude Opus 5(1,852점)만이 이보다 높은 점수를 기록했습니다.

하지만 문제는 이러한 높은 성과를 어떻게 달성했는지에 있습니다. Qwen3.8 Max는 기존 14단계였던 작업을 처리하는 데 64단계나 필요로 하며, 매 단계마다 전체 대화 내역을 모델에 재전송하기 때문에 입력 토큰(Input Token)이 15배나 증가했습니다. 모델이 훨씬 더 꼼꼼하게 작동하긴 하지만, 속도는 느려지고 비용은 더 많이 발생합니다.

입력 토큰 비용(백만 토큰당 2.50달러에서 2.00달러로 인하), 출력 비용(7.50달러에서 6.00달러로 인하), 캐시 적중 비용(0.50달러에서 0.25달러로 인하)이 모두 하락했음에도 불구하고, 알리바바의 가성비는 큰 타격을 받았습니다. 현재 지능 지수에서 단일 작업을 수행하는 데 드는 비용은 1.14달러로, Qwen3.7 Max(0.53달러)의 두 배가 넘습니다. 반면, Kimi K3는 1점 더 높은 점수를 0.86달러의 비용으로 기록했으며, GLM-5.2 역시 0.57달러의 비용만을 요구합니다.

또한 이전 버전과 비교했을 때 성능이 퇴보한 부분도 존재합니다. 모델이 매우 긴 텍스트에서 정보를 올바르게 수집할 수 있는지 측정하는 'AA-LCR' 테스트에서는 2점이 하락했습니다. 모델이 지식 질문에 정확하게 답하거나 모를 경우 솔직하게 모른다고 인정하는지를 측정하는 'AA-Omniscience' 점수는 10점이나 떨어졌습니다. 정확도는 약 31% 수준을 유지했지만, 환각(Hallucination) 비율은 23%에서 40%로 급증했습니다. 즉, Qwen3.8 Max는 모른다고 말하는 대신 훨씬 더 자주 임의로 추측하여 답변을 내놓고 있는 것입니다.

원문 보기
원문 보기 (영어)
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 6, 2026 Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper. On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step. The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57. Ad There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know. Ad DEC_D_Incontent-1 AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: Artificial Analysis Ask about this article… Search