메뉴
BL
MarkTechPost 20일 전

엔비디아, 압축형 하이브리드 MoE LLM 출시

IMP
7/10
핵심 요약

엔비디아가 기존 모델 대비 서버 처리량을 2배 이상 향상한 압축형 하이브리드 전문가 혼합(MoE) 대형 언어 모델을 공개했습니다. 이 모델은 하드웨어에 최적화된 구조적 압축과 지식 증류 기술을 활용하여 총 파라미터와 활성 파라미터 수를 줄이면서도 사용자별 응답 속도는 유지합니다. 결과적으로 동일한 하드웨어 환경에서 처리할 수 있는 동시 사용자 수와 전체 서버 효율성을 극대화할 수 있어, AI 인프라 운영 비용 절감 및 확장성 측면에서 매우 중요합니다.

번역된 본문

엔비디아(NVIDIA)가 Nemotron-3-Super의 압축된 변형인 'Nemotron-Labs-3-Puzzle-75B-A9B'를 출시했습니다. 이 모델은 하드웨어 특성을 고려한 구조적 압축(hardware-aware structural compression)과 짧은 지식 증류(knowledge distillation) 기반의 회복 단계를 교대로 진행하는 'Iterative Puzzle' 기법을 적용했습니다.

이를 통해 모델의 크기는 총 파라미터 1,207억 개(120.7B) 및 활성 파라미터 128억 개(12.8B)에서 총 753억 개(75.3B) 및 활성 93억 개(9.3B)로 감소했습니다. 단일 8x B200 노드 환경에서 이 모델은 사용자당 초당 100 토큰(100 tok/s) 속도를 유지하면서도, 기존 Super 모델 대비 총 처리량(throughput)을 2.03배 향상시켰습니다. 또한 단일 H100 GPU 기준으로 100만 토큰(1M-token)을 동시에 처리할 수 있는 요청 수가 1개에서 8개로 증가했습니다.

해당 포스트 '엔비디아, 사용자 처리량을 유지하면서 서버 처리량을 2.03배 높인 압축형 하이브리드 MoE LLM인 Nemotron-Labs-3-Puzzle-75B-A9B 출시'는 MarkTechPost를 통해 처음 공개되었습니다.

원문 보기
원문 보기 (영어)
NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.3B / 9.3B. On a single 8xB200 node it delivers 2.03x Super's total throughput at 100 tok/s per user. On one H100, 1M-token concurrency rises from 1 request to 8. The post NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput appeared first on MarkTechPost.