BL
MarkTechPost • 42일 전
미니맥스, 연산량 28배 줄인 희소 어텐션(MSA) 공개
IMP 7/10
핵심 요약
미니맥스(MiniMax)는 그룹드 쿼리 어텐션(GQA) 기반의 새로운 희소 어텐션 기술인 MSA를 발표했습니다. 이 기술은 가벼운 인덱스 브랜치를 통해 핵심 데이터 블록만 처리하여 연산량을 크게 줄이면서도 기존 모델과 동등한 성능을 유지합니다. 이를 통해 100만 토큰의 긴 컨텍스트를 처리할 때 토큰당 어텐션 연산량을 28.4배나 획기적으로 감소시킬 수 있습니다.
번역된 본문
미니맥스(MiniMax)는 그룹드 쿼리 어텐션(Grouped Query Attention, GQA)을 기반으로 구축된 희소 어텐션(Sparse Attention)인 MSA를 발표했습니다. 가벼운 인덱스 브랜치(Index Branch)는 각 쿼리 및 GQA 그룹당 상위 k개의 키-값(Key-Value) 블록을 선택하며, 메인 브랜치(Main Branch)는 오직 선택된 해당 블록들에만 어텐션을 수행합니다. 이 방식은 다운스트림 벤치마크에서 GQA와 동등한 성능을 보여주면서도, 100만 컨텍스트(1M context)에서 토큰당 어텐션 연산량을 28.4배 감소시킵니다.
이 글 '미니맥스 희소 어텐션(MSA): 3조 토큰 예산으로 109B 매개변수 MoE 모델을 학습한 2-브랜치 블록 희소 어텐션'은 MarkTechPost에 처음 게재되었습니다.
원문 보기 (영어)
MiniMax released MSA, a sparse attention built on Grouped Query Attention. A lightweight Index Branch selects Top-k key-value blocks per query and GQA group; the Main Branch attends only to those blocks. It matches GQA on downstream benchmarks while reducing per-token attention compute 28.4× at 1M context.
The post MiniMax Sparse Attention (MSA): a Two-Branch Block-Sparse Attention Trained on a 109B-Parameter MoE With a 3T-Token Budget appeared first on MarkTechPost.