On December 16, Alibaba’s Tongyi Wanxiang team released a new generation of Wanxiang 2.6 series models. This version is defined as the first video generation model in China that supports role-playing functions. It also integrates audio and video synchronization, multi-shot generation and sound driver capabilities.


It is reported that Wanxiang 2.6 can learn the timing information, subject characteristics and acoustic elements of the input video through multi-modal joint modeling at the technical level, aiming to achieve the overall consistency of the generated video in terms of picture and sound. Its storyboard control function can construct original materials into professional narrative paragraphs including multi-shot switching based on semantic understanding.

This upgrade focuses on improving the image quality, sound effects and command following capabilities, and supports a maximum of 15 seconds of video generation at a time. The new role-playing function allows users to upload personal videos and combine them with prompt words. The model can automatically complete storyboard design, role interpretation and dubbing to generate a short film with a cinematic feel. This capability is mainly oriented to professional scenarios such as advertising design and short play production.


At present, the Wanxiang Model family has more than ten kinds of visual creation capabilities, such as Vincent pictures, image editing, and Vincent videos. From now on, users can experience Wanxiang 2.6 through the official website, and enterprise users can also call the model API through the Alibaba Cloud Bailian platform.