Alibaba Wan3.0 expands AI video generation with support for clips lasting up to 30 seconds and multimodal reference inputs including text, images, videos, audio, web pages, PDFs, and PowerPoint presentations. Alibaba has officially unveiled Wan3.0, the latest version of its advanced AI video generation model, bringing longer video outputs of up to 30 seconds along with expanded multimodal input support for text, images, videos, audio, web pages, PDFs, and PowerPoint presentations. Currently available in public beta, Wan3.0 is designed to combine high-fidelity video generation, visual consistency, and multimodal reference capabilities within a single model. Mainstream AI video generators have traditionally focused on relatively short clips lasting anywhere from just a few seconds to around 15 seconds. The AI model supports text, images, video, and audio inputs simultaneously.