AI video generator from Alibaba that turns text, images, and audio into cinematic 1080p clips with synced sound.
Wan AI is Alibaba's open-source video generation platform that creates high-quality videos from text prompts, images, or audio inputs. Built on a diffusion transformer architecture, it supports text-to-video, image-to-video, and speech-to-video generation up to 1080p with synchronized audio and lip-sync. Wan works for marketers, content creators, filmmakers, and developers who want cinematic output without expensive proprietary tools. You can use the hosted web platform or self-host the open-weight models under Apache 2.0.
Wan's open models can be downloaded and run locally without paying a model access fee. However, local use still requires suitable hardware or cloud GPU resources, which can add compute costs.
Wan is developed by Alibaba's Wan team and is released as an open video generation model family, with model weights and inference code available through platforms such as GitHub, Hugging Face, and ModelScope.
Wan2.2 supports 480p and 720p generation, depending on the model. The TI2V-5B model supports 720p video at 24fps. Generation length varies by model and workflow.
Yes. Wan2.2 provides downloadable model weights and inference code. The TI2V-5B model can run on a GPU with at least 24GB VRAM, while larger 14B models require more powerful hardware.
Wan2.2 includes an S2V-14B speech-to-video model that can generate video based on speech and other inputs. It also supports speech synthesis through CosyVoice.
0 out of 5 stars
Based on 0 reviews
5 star reviews
4 star reviews
3 star reviews
2 star reviews
1 star reviews
If you've used this tool, share your thoughts with other users
Alibaba's open AI video generator with audio sync
All-in-one AI video production with 20+ models
AI video, image, and music generator for creators
Your all-in-one creative assistant.
Create cinematic AI‑generated videos & images in seconds.
Transform your ideas into cinematic-quality videos and images.
Director-grade AI video generation from any input
Turn text and photos into AI videos
Create AI-powered music videos with scene-based visuals and optional lip-sync.