메뉴
BL
The Decoder • 53일 전

중국 미니맥스 H3, 오픈 AI 비디오 모델 최초로 성능 1위 달성

IMP
8/10
핵심 요약

중국의 AI 기업 미니맥스(MiniMax)가 공개한 비디오 생성 모델 'H3'가 오픈소스 모델 최초로 글로벌 AI 비디오 평가 순위 정상을 차지했습니다. 이 모델은 텍스트, 이미지, 오디오를 동시에 처리하여 입체 음향이 포함된 짧은 영상을 생성하며, 자체 데이터로 파인튜닝(미세조정)이 가능해 실무 적용 가치가 높습니다. 단, 고해상도 모듈 등 일부 핵심 기능은 비공개이며 연 매출 2,000만 달러 이하 기업에만 상업적 이용이 허용되는 라이선스 제약이 있습니다.

번역된 본문

중국의 미니맥스(MiniMax) H3, AI 비디오 순위 정상을 차지한 최초의 오픈 모델 작성자: 막시밀리안 슈라이너 (Maximilian Schreiner) 2026년 8월 3일

미니맥스가 H3 비디오 모델의 가중치(weights)를 공개하면서, 오픈 모델이 비디오 평가 순위에서 최초로 1위를 차지하는 역사적인 성과를 냈습니다.

평가 기관인 '인공 지능 분석(Artificial Analysis)'에 따르면, H3는 '비디오 편집(Video Editing)' 부문에서 1위, '텍스트-비디오(Text-to-Video)' 부문에서 2위, '이미지-비디오(Image-to-Video)' 부문에서 3위를 기록했습니다.

이 330억 매개변수(parameters) 모델은 텍스트, 이미지, 비디오, 오디오를 함께 처리하며, 입체 음향(stereo sound)이 포함된 4~15초 분량의 영상 클립을 생성합니다. 모델 카드에 따르면, 단일 프롬프트에 최대 9장의 참조 이미지, 3개의 비디오 클립, 3개의 오디오 클립을 포함할 수 있습니다.

다만, 두 가지 핵심 요소는 여전히 비공개 상태로 남아있습니다. 2K 해상도 모듈과 프롬프트 및 참조 자료를 구조화된 중간 형식으로 변환하는 'H3-Context-IR'은 포함되지 않았습니다. 따라서 'ComfyUI'와 같은 환경에서 로컬로 H3를 구동할 때는 최대 768p 해상도까지만 지원되며, 사용자가 미니맥스가 공개한 프롬프트 가이드를 활용해 직접 컨텍스트를 준비해야 합니다.

그럼에도 불구하고 오픈 가중치를 통해 사용자는 자체 영상 소스, 캐릭터 또는 특정 시각적 스타일에 맞춰 모델을 파인튜닝(fine-tuning)할 수 있습니다. 단, 라이선스와 관련된 주의할 점이 있습니다. 상업적 사용은 연 매출 2,000만 달러 미만의 기업에게만 허용됩니다.

[광고] 같은 날 바이트댄스(ByteDance)는 내장 오디오로 30초 분량의 클립을 생성하는 자체 폐쇄형(Closed) 비디오 모델 '시드댄스(Seedance) 2.5'를 공개했습니다.

[광고]

과장 없는 AI 뉴스 – 전문가가 직접 엄선합니다. 과장된 마케팅 없이 핵심만 전하는 THE DECODER를 구독하세요. 광고 없는 읽기 환경, 주간 AI 뉴스레터, 연 6회 발행되는 독점 프론티어 보고서 'AI Radar', 전체 아카이브 접근 및 댓글 기능을 제공합니다. 지금 구독하세요.

출처: HuggingFace

원문 보기
원문 보기 (영어)
China's MiniMax H3 is the first open model to top an AI video ranking Maximilian Schreiner View the LinkedIn Profile of Maximilian Schreiner Aug 3, 2026 MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video. The 33-billion-parameter model processes text, images, video, and audio together, generating four- to 15-second clips with stereo sound. According to the model card , a single prompt can include up to nine reference images, three video clips, and three audio clips. Video by MiniMax H3 Two pieces remain closed, though. The 2K resolution module and H3-Context-IR, which translates prompts and reference material into a structured intermediate format, aren't included. Running H3 locally in ComfyUI tops out at 768p, and users will need to handle context prep themselves using MiniMax's published prompting guides. The open weights do allow fine-tuning on custom footage, characters, or a specific visual style. One catch on the license side: commercial use is only permitted for companies making under $20 million in revenue. Ad ByteDance released its closed Seedance 2.5 the same day, which generates 30-second clips with built-in audio. Ad DEC_D_Incontent-1 AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: HuggingFace Ask about this article… Search