BL
MarkTechPost • 30일 전
알리바바 Qwen3.8-Flash-Next 공개, Qwen4 아키텍처 예고
IMP 8/10
핵심 요약
알리바바 Qwen 팀이 오픈 웨이트 멀티모달 MoE 모델 'Qwen3.8-Flash-Next'(총 180B 파라미터, 토큰당 6B 활성)를 공개하며 Qwen4 아키텍처를 미리 선보였다. Gated DeltaNet과 희소 어텐션 하이브리드, N-gram 임베딩, Muon 옵티마이저 등 4가지 구조적 변화가 적용됐으며, Qwen3.7-Plus 대비 학습 비용 1/9을 주장한다. 단, FP8 체크포인트만 172.78 GiB로 자체 호스팅 요구 사양이 상당하다.
번역된 본문
알리바바의 오픈 웨이트 멀티모달 MoE(Mixture-of-Experts) 모델이자 Qwen4 아키텍처의 초기 프리뷰인 Qwen3.8-Flash-Next를 살펴본다. 180B 파라미터가 실제로 어디에 위치하는지 분석한다: 125B 백본, 51B N-gram 임베딩 테이블, 4B 멀티 토큰 예측 모듈로 구성되며, 토큰당 활성 파라미터는 6B에 불과하다. 네 가지 아키텍처 변화 — Gated DeltaNet과 Qwen 희소 어텐션(Qwen Sparse Attention) 하이브리드, Gated Residual, N-gram 임베딩, Muon 옵티마이저 — 를 차례로 살펴본다. 또한 벤치마크 결과, Qwen3.7-Plus 대비 1/9 수준의 학습 비용 주장, 그리고 172.78 GiB 크기의 FP8 체크포인트를 자체 호스팅하는 데 실제로 요구되는 사양까지 다룬다.
원문 보기 (영어)
We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands.
The post Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture appeared first on MarkTechPost.
관련 소식