AB
AiBoss
project

AnyCharV - A role-controlled video generation framework jointly launched by the Chinese University of Hong Kong, Tsinghua University, and other institutions.

AnyCharV is a character-controlled video generation framework jointly developed by the Chinese University of Hong Kong, Tsinghua University Shenzhen Graduate School, and the University of Hong Kong. It can combine arbitrary reference character images with target-driven video to generate high-quality character-controlled videos...

What is AnyCharV?

AnyCharV is a character-controlled video generation framework jointly developed by the Chinese University of Hong Kong, Tsinghua University Shenzhen International Graduate School, and the University of Hong Kong. It combines arbitrary reference character images with target-driven videos to generate high-quality character videos. AnyCharV employs a two-stage training strategy to achieve fine-to-coarse guidance: the first stage uses fine-grained segmentation masks and pose information for self-supervised synthesis; the second stage uses self-reinforcement training and coarse-grained masks to optimize character detail preservation. AnyCharV demonstrates superior performance in experiments, naturally preserving character appearance details and supporting complex human-object interactions and background blending. AnyCharV can be combined with content generated by text-to-image (T2I) and text-to-video (T2V) models, exhibiting strong generalization capabilities.

Main functions of AnyCharV

  • Combining arbitrary characters with target scenesCombine any given character image with target-driven video to generate natural, high-quality video.
  • High-fidelity character details preservedBased on self-reinforcement training and coarse-grained masking guidance, the appearance and details of the character are preserved, avoiding distortion.
  • Complex Scenes and Human-Object InteractionSupports natural interaction of characters in complex backgrounds, such as movement and object manipulation.
  • Flexible input supportThe content generated by combining text-to-image (T2I) and text-to-video (T2V) models has strong generalization ability.

AnyCharV's technical principles

  • Phase 1Self-supervised synthesis and fine-grained guidance: Using the segmentation mask and pose information of the target character as conditional signals, the reference character is accurately synthesized into the target scene. CLIP features from the reference image and character appearance features extracted from ReferenceNet are introduced to preserve the character's identity and appearance. The segmentation mask is strongly enhanced to reduce detail loss due to shape differences.
  • Phase TwoSelf-reinforcement training and coarse-grained guidance: Self-reinforcement training is performed on generated video pairs, using coarse bounding box masks instead of fine segmentation masks to reduce constraints on character shapes. Based on this approach, the model can better preserve the details of the reference character, generating more natural videos during the inference phase.

AnyCharV's project address

AnyCharV Application Scenarios

  • Film and television productionIt allows you to composite any character into a target scene, supports complex interactions, and helps with special effects production.
  • Artistic CreationCombine text-based content to quickly generate high-quality character videos and inspire creativity.
  • Virtual RealityIt generates real-time interactive videos between characters and virtual scenes, enhancing the sense of immersion.
  • Advertising and MarketingQuickly synthesize personalized advertising videos to meet diverse needs.
  • Education and TrainingGenerate videos featuring specific characters and scenarios to aid in teaching and training.