AB
AiBoss
project

Wan2.2 - Alibaba's open-source AI video generation model

Wan2.2 is an advanced AI video generation model open-sourced by Alibaba. It includes three open-source versions: text-based video (Wan2.2-T2V-A14B), image-based video (Wan2.2-I2V-A14B), and unified video generation (Wan2.2-IT2V-...).

What is Wan2.2 in Tongyi Wanxiang?

Wan2.2 is an advanced AI video generation model open-sourced by Alibaba. It comprises three open-source models: text-based video (Wan2.2-T2V-A14B), image-based video (Wan2.2-I2V-A14B), and unified video generation (Wan2.2-IT2V-5B), with a total of 27 billion parameters. The model is the first to introduce a hybrid expert (MoE) architecture, effectively improving generation quality and computational efficiency. It also features a pioneering cinematic-level aesthetic control system, precisely controlling aesthetic effects such as lighting, color, and composition. This newly open-sourced 5B parameter compact video generation model supports text and image-based video generation, can run on consumer-grade graphics cards, and is based on a high-efficiency 3D VAE architecture, achieving high compression rates and rapid high-definition video generation capabilities. Currently, developers can obtain the model and code through platforms such as GitHub and HuggingFace, enterprises can use the Alibaba Cloud API for application development, and users can directly experience it on the Wan2.2 official website and the Wan2.2 app.

The main functions of Wan2.2 (通义万相)

  • Text-to-VideoIt generates corresponding video content based on the input text description. For example, if the input is "a cat is running on the grass," the model can generate a video that matches the description.
  • Image-to-VideoThe model generates videos based on the input images, creating dynamic scenes that bring the images to life.
  • Unified video generation (Text-Image-to-Video)It combines text and images to generate videos, using text descriptions and image information to create more accurate video content.
  • Cinematic Aesthetic ControlIt generates videos with a professional cinematic quality by controlling lighting, color, composition, and micro-expressions. Users can customize the aesthetic style of the video by inputting relevant keywords (such as "warm color tone" or "central composition").
  • Complex motion generationIt can generate complex motion scenes and character interactions, enhancing the dynamic expressiveness and realism of videos.

Technical Principles of Tongyi Wanxiang Wan2.2

  • Hybrid Expert (MoE) ArchitectureThe MoE architecture is introduced, dividing the model into high-noise experts and low-noise experts. High-noise experts are responsible for the overall video layout, while low-noise experts handle detailed refinement. This significantly improves the number of model parameters and the quality of generated content while maintaining the same computational cost.
  • Diffusion ModelBased on a diffusion model as its foundation, high-quality video content is generated by progressively removing noise. Combining the MoE architecture with the diffusion model can further optimize the generation results.
  • High compression ratio 3D VAETo improve model efficiency, Tongyi Wanxiang 2.2 is based on a high-compression-ratio 3D variational autoencoder (VAE). The architecture achieves high compression ratios in both time and space, enabling the model to quickly generate high-definition videos on consumer-grade graphics cards.
  • Large-scale data trainingThe model is trained on a large-scale dataset, including more image and video data, to improve the model's generalization ability and generation quality in a variety of scenarios.
  • Aesthetic data annotationBased on meticulously labeled aesthetic data (such as lighting, color, and composition), the model can generate video content with a professional cinematic quality, meeting users' customized needs for video aesthetics.

The project address for Tongyi Wanxiang Wan2.2

  • GitHub repository: https://github.com/Wan-Video/Wan2.2
  • HuggingFace model libraryhttps://huggingface.co/Wan-AI/models

How to use Wan2.2 (通义万相)

  • Visit the official websiteVisit the official website of Tongyi Wanxiang or download the Tongyi APP to experience it..
  • Select ModelIn the model selection drop-down box, select General Universal Phase 2.2.
  • Select experience mode:
    • Text-to-VideoEnter a text description, such as "a cat is running on the grass", and click the generate button to see the generated video.
    • Image-to-VideoUpload an image, and the model will generate a dynamic video based on the image content.
    • Unified video generation (Text-Image-to-Video)By combining text descriptions and uploaded images, more accurate video content can be generated.
  • Adjust parameters (optional)Users can adjust video parameters such as resolution and frame rate as needed. A cinematic aesthetic control system allows users to customize the video's aesthetic style by inputting keywords (such as "warm tones" or "central composition").
  • View the generated resultsThe generated video is displayed directly on the webpage, and users can download or share the generated video.

Application Scenarios of Tongyi Wanxiang Wan2.2

  • Short video creationCreators can quickly generate engaging short video content for social media platforms, saving time and costs.
  • Advertising and MarketingAdvertising agencies and brands generate high-quality advertising videos to enhance advertising effectiveness and brand influence.
  • Education and TrainingEducational institutions and businesses can generate engaging educational videos and training materials to improve learning outcomes and training quality.
  • Film and television productionFilm and television production teams can quickly generate scene designs and animation clips, improving creative efficiency and reducing production costs.
  • News and MediaNews organizations and media outlets can generate animations and visual effects to enhance the visual appeal and audience engagement of news reports.