메뉴
HN
Hacker News • 24일 전

쿼사르 438B: 유럽 최고 성능 AI 모델 등장

IMP
6/10
핵심 요약

멀티버스 컴퓨팅이 기업용 에이전트와 코딩을 위한 438B 규모 추론 모델 '쿼사르 438B'를 공개했습니다. Artificial Analysis 지능 지수에서 43점을 기록하며 유럽 모델 중 최고 성능을 달성했고, 500토큰 응답에 15.3초 만에 완료하는 프론티어급 속도도 갖췄습니다. CompactifAI API를 통해 인프라 구축 없이 테스트할 수 있습니다.

번역된 본문

쿼사르 438B(Quasar 438B)는 기업 규모의 에이전트와 코딩을 위해 설계된 우리의 대표 추론 모델입니다. 이는 멀티버스 컴퓨팅(Multiverse Computing)이 출시한 첫 대형 모델로, 영어와 스페인어로 작동하며, Artificial Analysis 지능 지수에서 43점을 기록해 해당 분야 유럽 모델 중 최고 성적을 달성했습니다. 지능 지수 v4.1.1은 GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR 등 9개 평가를 종합한 것입니다. 쿼사르의 43점은 미스트랄 미디움 3.5(Mistral Medium 3.5)의 30점, 엔비디아 네모트론 3 울트라(NVIDIA Nemotron 3 Ultra)의 38점, 잉클링(Inkling)의 42점을 앞서며, 클로드 오퍼스 5(Claude Opus 5)가 63점으로 선두를 달리는 분야입니다. 쿼사르는 지능적일 뿐만 아니라 빠릅니다. 추론 시간을 포함해 500토큰을 15.3초 만에 반환합니다. 비교 대상 중 단 3개 모델만이 더 빠르며, 그중 지수에서 더 높은 점수를 받은 모델은 제미나이 3.7 플래시(Gemini 3.7 Flash) 하나뿐입니다. 쿼사르보다 점수가 높은 모델 중 25초 이내에 답변하는 모델은 2개뿐이며, 나머지는 38~156초가 걸립니다. 멀티버스 컴퓨팅은 AI를 더 효율적이고 배포 가능하게 만드는 데 입지를 쌓아왔습니다. 쿼사르는 그 성과를 400B 이상 파라미터급으로 확장한 것으로, 계획 수립, 도구 사용, 코드 실행, 대형 컨텍스트가 필요한 다단계 작업을 위한 추론 모델이면서도 해당 급 모델이 보통 수반하는 지연 시간이 없습니다. 이 모델은 CompactifAI API를 통해 제공되어, 인프라를 구축하지 않고도 팀이 테스트할 수 있습니다. 최고 점수를 기록한 유럽 모델. 그림 1. 9개 평가를 종합한 Artificial Analysis 지능 지수 v4.1.1. 높을수록 좋습니다. 쿼사르는 종합 지수에서 43점을 받았습니다. 이는 미스트랄 미디움 3.5보다 13점, 1120억 개 더 많은 파라미터를 갖춘 네모트론 3 울트라보다 5점 앞선 성적입니다. 프론티어급 속도. 그림 2. 종단 간 응답 시간: 추론 시간 포함 500토큰 출력까지의 초. 낮을수록 좋습니다. 쿼사르는 500토큰 응답을 15.3초에 완료합니다. 비교 대상 중 더 빠른 3개 모델은 지수 24점의 네모트론 3.5 라이트닝(9.4초), 37점의 제미나이 3.5 플래시-라이트(10.8초), 56점의 제미나이 3.7 플래시(11.5초)이며, 마지막 모델만이 속도와 성능 모두 앞섭니다. 반대로 보면 격차가 더 큽니다. 쿼사르는 43점에 15.3초, 미스트랄 미디움 3.5는 30점에 18.8초, 네모트론 3 울트라는 36점에 25.7초, 잉클링은 42점에 48.3초로 쿼사르 438B보다 두 배 이상 걸립니다. 미스트랄 미디움 3.5와의 비교는 두 축 모두에서 일방적입니다: 쿼사르가 점수(43 대 30)와 응답 속도(15.3초 대 18.8초) 모두 앞섭니다. 롱컨텍스트 추론. 그림 3. AA-LCR 점수. 높을수록 좋습니다. 쿼사르는 긴 문서에 분산된 정보를 추출·연결·추론하는 능력을 평가하는 AA-LCR에서 75.0점을 받았습니다. 이는 Grok 4.6(하이)(high)의 75.0점과 동률이며, 클로드 오퍼스 5의 75.7점, Qwen3.8 2.4T A95B의 75.3점과 1점 이내 차이입니다. 네모트론 3 울트라보다 4.0점, 미스트랄 미디움 3.5보다 9.7점 앞섭니다. 롱컨텍스트 처리는 기업 리서치, 문서 분석, 에이전트 워크플로의 기반이며, 쿼사르가 프론티어 그룹에 가장 근접한 분야입니다. 에이전트 코딩 및 터미널 작업. 그림 4. Terminal-Bench v2.1 점수. 높을수록 좋습니다. 쿼사르는 실제 터미널 환경에서 에이전트를 평가하는 Terminal-Bench v2.1에서 69.3점을 받았습니다. 미스트랄 미디움 3.5보다 18.7점, 네모트론 3 울트라보다 15.4점 앞서며, 클로드 오퍼스 5가 89.1점으로 이끄는 프론티어 그룹에는 뒤집니다. 이는 쿼사르에게 가장 개선 여지가 큰 평가로, 다음 작업의 목표가 되는 분야입니다. 한눈에 보는 4가지 결과. 표 1. Artificial Analysis가 제공한 4개 차트에서 본 쿼사르 438B. 출처: Artificial Analysis. 기업 규모 에이전트와 코딩을 위해 설계. 쿼사르는 소프트웨어 개발 에이전트, 기술 코파일럿, 리서치 시스템, 워크플로 자동화를 운영하는 조직을 위해 설계되었습니다. Terminal-Bench와 롱컨텍스트 결과는 컨텍스트를 유지하고, 작업을 조율하며, 다단계 작업을 수행해야 하는 에이전트를 뒷받침합니다.

원문 보기
원문 보기 (영어)
Quasar 438B is our flagship reasoning model, built for enterprise-scale agents and coding. It is the first large model Multiverse Computing has released, it runs in English and Spanish, and it scores 43 on the Artificial Analysis Intelligence Index, the highest result of any European model in the field. The Intelligence Index v4.1.1 combines nine evaluations: GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience and AA-LCR. Quasar's 43 puts it ahead of Mistral Medium 3.5 at 30, NVIDIA Nemotron 3 Ultra at 38 and Inkling at 42, in a field led by Claude Opus 5 at 63. Quasar is not only intelligent, it is also fast. It returns 500 tokens, thinking time included, in 15.3 seconds. Only three models in the comparison are faster, and only one of those, Gemini 3.7 Flash, scores higher on the index. Of the models that do outscore Quasar, only two answer in under 25 seconds. The rest take between 38 and 156. Multiverse Computing has built its position on making AI more efficient and deployable. Quasar brings that work into the 400B-plus parameter class: a reasoning model for multi-step tasks that need planning, tool use, code execution and large context, without the latency that class normally carries. The model is available through the CompactifAI API, so teams can test it without standing up infrastructure. The highest-scoring European model Figure 1. Artificial Analysis Intelligence Index v4.1.1, a composite of nine evaluations. Higher is better. Quasar scores 43 on the composite index. That is 13 points ahead of Mistral Medium 3.5 and 5 ahead of Nemotron 3 Ultra, which carries 112 billion more parameters. Frontier-class speed Figure 2. End-to-end response time: seconds to output 500 tokens, including reasoning time. Lower is better. Quasar completes a 500-token response in 15.3 seconds. The three faster models in the comparison are Nemotron 3.5 Lightning at 9.4s, which scores 24 on the index, Gemini 3.5 Flash-Lite at 10.8s, which scores 37, and Gemini 3.7 Flash at 11.5s, which scores 56. Only the last of those is both faster and more capable. Read it the other way and the gap is wider. Quasar scores 43 and takes 15.3 seconds. Mistral Medium 3.5 scores 30 and takes 18.8 seconds. Nemotron 3 Ultra scores 36 and takes 25.7 seconds, Inkling scores 42 and needs 48.3 secods, more than double than Quasar 438B. Against Mistral Medium 3.5 the comparison runs one way on both axes: Quasar scores higher (43 against 30) and answers faster (15.3s against 18.8s). Long-context reasoning Figure 3. AA-LCR score. Higher is better. Quasar scores 75.0 on AA-LCR, which tests the ability to extract, connect and reason over information spread across long documents. That is level with Grok 4.6 (high) at 75.0, and within a point of Claude Opus 5 at 75.7 and Qwen3.8 2.4T A95B at 75.3. It leads Nemotron 3 Ultra by 4.0 points and Mistral Medium 3.5 by 9.7. Long-context handling is what enterprise research, document analysis and agentic workflows are built on, and it is where Quasar comes closest to the frontier group. Agentic coding and terminal work Figure 4. Terminal-Bench v2.1 score. Higher is better. Quasar scores 69.3 on Terminal-Bench v2.1, which puts agents to work in real terminal environments. It leads Mistral Medium 3.5 by 18.7 points and Nemotron 3 Ultra by 15.4, and trails the frontier group led by Claude Opus 5 at 89.1. This is the evaluation with the most headroom for Quasar, and it is where the next round of work is aimed. The four results in one view Table 1. Quasar 438B on the four supplied Artificial Analysis charts. Source: Artificial Analysis. Built for enterprise-scale agents and coding Quasar is built for organizations running software development agents, technical copilots, research systems and workflow automation. The Terminal-Bench and long-context results support agents that have to hold context, coordinate actions and work through multi-step tasks. The response time keeps those loops fast enough to sit inside an interactive product rather than a batch job. Support for English and Spanish makes Quasar a practical foundation for European and international enterprises that need reasoning without narrowing deployment to a single language. Teams can put it to work across engineering, operations and knowledge work, wherever accuracy, context and reliable task completion decide the outcome. More updates are coming for Quasar. Stay tuned. Get started CompactifAI API. Sign up and start sending requests: dashboard.compactif.ai Artificial Analysis. Check out the benchmarks Contact. Reach out our team too at business@multiversecomputing.com