메뉴
BL
MarkTechPost • 30일 전

Z.ai, 1M 토큰 컨텍스트 멀티모달 GLM-5.3-Flash 공개

IMP
7/10
핵심 요약

Z.ai가 GLM-5 시리즈 최초의 네이티브 멀티모달 모델인 GLM-5.3-Flash를 공개했습니다. 총 320B/활성 18B 파라미터의 MoE 구조에 100만 토큰 컨텍스트 윈도우를 지원하며, MIT 라이선스로 허깅페이스에 공개되었습니다. 하이브리드 KDA 선형 어텐션과 NoPE 희소 MLA 어텐션을 통해 기존 대비 어텐션 연산량을 약 3배, KV 캐시를 4.4배 줄여 비용 효율성이 높은 것이 특징입니다.

번역된 본문

Z.ai가 GLM-5 시리즈 최초의 네이티브 멀티모달 모델인 GLM-5.3-Flash를 공개했습니다. 이 모델은 총 320B/활성 18B 파라미터의 MoE(전문가 혼합) 구조와 1,048,576 토큰(약 100만 토큰)의 컨텍스트 윈도우를 갖추고 있으며, 허깅페이스에 MIT 라이선스로 가중치가 공개되었습니다. API 가격은 입력 100만 토큰당 0.15달러, 출력 100만 토큰당 0.50달러입니다. Terminal-Bench 2.1에서 84.3점, DeepSWE v1.1에서 63.4점을 기록했으며, 하이브리드 KDA 선형 어텐션과 NoPE 희소 MLA 어텐션을 채택해 GLM-5.3 대비 어텐션 연산량을 약 3배, KV 캐시를 4.4배 절감했습니다.

이 기사는 MarkTechPost에 처음 게재되었습니다.

원문 보기
원문 보기 (영어)
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using hybrid KDA linear plus NoPE sparse MLA attention to cut attention compute ~3× and KV cache 4.4× versus GLM-5.3. The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost.
관련 소식