메뉴
BL
The Decoder • 32일 전

알리바바 Wan3.0, 텍스트·이미지·문서로 최대 30초 AI 영상 생성

IMP
7/10
핵심 요약

알리바바가 공개한 영상 생성 모델 Wan3.0 베타는 최대 30초 영상을 만들 수 있으며, 텍스트뿐 아니라 이미지, PDF, 웹페이지, 파워포인트까지 입력으로 활용할 수 있습니다. 얼굴이나 UI 왜곡 등 기존 AI 영상의 문제를 개선해 참조 자료의 디테일을 일관되게 유지하는 것이 강점이며, 영화 제작부터 마케팅, 자율주행·로봇 학습용 시뮬레이션까지 폭넓은 활용이 기대됩니다.

번역된 본문

알리바바의 영상 생성 모델 Wan3.0이 베타로 공개되었습니다. 최대 30초 길이의 영상을 생성할 수 있으며, 텍스트, PDF, 웹페이지, 파워포인트 파일을 입력으로 받아들입니다. 이는 전작 Wan2.5에 비해 영상 길이가 두 배로 늘어난 것입니다. 또한 사용자의 프롬프트에 맞춰 최적의 영상 길이를 추천하고, 기존 영상을 더 길게 만드는 확장 도구도 포함되어 있습니다.

Wan3.0은 텍스트, 이미지, 영상, 오디오를 동시에 처리합니다. 하나의 프롬프트에 최대 10개의 이미지, 5개의 영상, 5개의 오디오 클립을 포함할 수 있습니다. PDF나 파워포인트 발표자료 같은 웹페이지와 문서도 입력 소스로 사용할 수 있어, 정적인 데이터를 동적인 영상으로 변환할 수 있습니다.

AI 생성 영상은 특히 얼굴이나 사용자 인터페이스에서 시각적 드리프트와 왜곡이 발생하는 경우가 많습니다. 알리바바에 따르면 Wan3.0은 캐릭터, 소품, 공간 배치 같은 참조 자료의 디테일을 더 일관되게 유지함으로써 이러한 문제를 해결하는 것을 목표로 합니다.

Wan3.0은 wan.video 웹사이트, 알리바바 클라우드 모델 스튜디오(Alibaba Cloud Model Studio), 또는 Qwen Cloud의 API를 통해 이용할 수 있습니다. 현재 30% 할인 중인 Standard 버전과 더 빠른 Prime 버전의 두 가지 요금제가 제공됩니다.

해상도 / Standard(초당) / Prime(초당) / 30초 클립(Standard) / 30초 클립(Prime) 480p / $0.05 / $0.068 / $1.50 / $2.04 720p / $0.10 / $0.14 / $3.00 / $4.20 1080p / $0.20 / $0.28 / $6.00 / $8.40

알리바바는 영화 제작부터 로봇 학습까지 폭넓은 활용을 겨냥하고 있습니다. 영화 제작 속도를 높이는 것에서부터 콘텐츠 크리에이터를 위한 숏드라마와 소셜 미디어 클립 제작까지 Wan3.0을 다양한 용도로 내세우고 있습니다. 기업은 텍스트와 이미지를 마케팅·교육 영상으로 변환할 수 있고, 기술 개발자는 자율주행 차량과 로봇 시스템 학습용 현실적인 시뮬레이션 영상 제작에 활용할 수 있습니다.

이번 출시는 알리바바가 AI 투자를 대폭 확대하는 가운데 이루어졌습니다. 알리바바는 최근 AI 확충 자금을 마련하기 위해 홍콩 상장 기업 사상 최대 규모의 신주 매각을 발표했으며, 지난주에는 AI 투자 급증으로 분기 순이익이 전년 동기 대비 75% 감소했다고 보고했습니다.

원문 보기
원문 보기 (영어)
Alibaba's Wan3.0 generates AI videos up to 30 seconds long from text, images, and documents Matthias Bastian View the LinkedIn Profile of Matthias Bastian Aug 24, 2026 Alibaba Ask about this article… Search Alibaba's video generation model Wan3.0 is now available in beta. It produces videos up to 30 seconds long and accepts text, PDFs, web pages, and PowerPoint files as input. That doubles the video length compared to its predecessor, Wan2.5 . The model also recommends the best video length based on the user's prompt and includes an extension tool to make existing videos longer. Wan3.0 processes text, images, video, and audio at the same time. A single prompt can include up to ten images, five videos, and five audio clips. Web pages and documents like PDFs or PowerPoint presentations also work as input sources, turning static data into dynamic video. Ad AI-generated videos often suffer from visual drift and distortion, especially in faces and user interfaces. Wan3.0 aims to fix that by keeping details from reference material like characters, props, and spatial layouts more consistent, according to Alibaba. Ad Wan3.0 is available through the wan.video website , Alibaba Cloud Model Studio , or via API on Qwen Cloud . There are two tiers: a Standard version, currently at a 30 percent discount, and a faster Prime version. Resolution Standard (per sec.) Prime (per sec.) 30-second clip (Standard) 30-second clip (Prime) 480p $0.05 $0.068 $1.50 $2.04 720p $0.10 $0.14 $3.00 $4.20 1080p $0.20 $0.28 $6.00 $8.40 Alibaba targets everything from film production to robotics training Alibaba is pitching Wan3.0 for a wide range of uses, from speeding up film production to generating short dramas and social media clips for content creators. Businesses could turn text and images into marketing and training videos, while tech developers could use it to produce realistic simulation footage for training autonomous vehicles and robotics systems. Ad The launch comes as Alibaba ramps up AI spending. The company just announced the largest share sale by a Hong Kong-listed company to fund its AI push, and last week reported a 75 percent year-over-year drop in quarterly profit driven by sharply higher AI investments. Ad AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: via X