DynamicFace - A video face-swapping technology launched by Xiaohongshu in collaboration with Shanghai Jiao Tong University and others.
DynamicFace is a new video face-swapping technology launched by the Xiaohongshu team. By combining a diffusion model and a plug-and-play temporal layer, based on 3D facial prior knowledge, the technology achieves high-quality and consistent video face-swapping effects.
What is DynamicFace?
DynamicFace is a new video face-swapping technology launched by the Xiaohongshu team. This technology combines a diffusion model with a plug-and-play temporal layer, leveraging 3D facial prior knowledge to achieve high-quality and consistent video face-swapping effects. The core of DynamicFace lies in introducing four refined facial conditions: background, shape-aware normal map, expression-related landmarks, and UV texture map with identity information removed. These conditions are independent of each other, providing accurate motion and identity information. It also employs Face Former and ReferenceNet for identity injection, ensuring consistency across different expressions and poses.
Main functions of DynamicFace
- Detailed facial condition analysisDynamicFace, based on 3D facial prior knowledge, decomposes the face into four fine conditions: background, shape-aware normal map, expression-related landmarks, and UV texture map with identity information removed. This provides precise guidance for face swapping.
- Identity Injection and ConsistencyThrough the Face Former and ReferenceNet modules, DynamicFace can maintain identity consistency under different expressions and poses, ensuring that the face identity after face swapping is highly consistent with the source image.
- Temporal consistency and video face swappingThe introduction of a temporal attention layer effectively solves the temporal consistency problem in video face-swapping, ensuring that the face-swapped video remains coherent across different frames.
- High-quality image generationDynamicFace is based on a diffusion model and can generate high-resolution and high-quality face-swapping images while preserving details such as the expression, pose, and background of the target image.
- Wide applicabilityDynamicFace is suitable for face swapping of static images and can be extended to the video field, applicable to various application scenarios such as portrait reenactment, film and television production, and virtual reality.
The technical principles of DynamicFace
- Diffusion Model and Latent Space GenerationDynamicFace generates high-quality images based on a diffusion model. The diffusion model generates images by progressively reversing a noise-adding process.
- 3D facial priors and decoupling conditionsFour fine-grained conditions based on 3D facial priors are introduced: background, shape-aware normal map, expression-related landmark map, and identity-removed UV texture map.
- Identity Injection ModuleDynamicFace uses Face Former and ReferenceNet for identity injection. Face Former provides high-level identity features, while ReferenceNet injects detailed texture information. The two modules ensure identity consistency across different expressions and poses.
- Time Consistency ModuleTo achieve temporal consistency in video face-swapping, DynamicFace introduces a temporal attention layer. This ensures that the generated video remains coherent across different frames, avoiding abrupt or unnatural transitions.
- Multi-condition guidance mechanismDynamicFace uses a mixture-of-guides mechanism to precisely control facial movement and appearance. It better preserves non-identity attributes of the target face, such as expression, posture, and lighting.
DynamicFace project address
- Project official website:https://dynamic-face.github.io
- arXiv technical paper:https://arxiv.org/pdf/2501.08553v1
Application scenarios of DynamicFace
- Film and television productionDynamicFace can be used in film and television post-production to quickly replace actors' facial expressions or identities, saving reshoot costs and improving production efficiency.
- Human Reenactment and Virtual RealityIn the field of facial reconstruction, DynamicFace can transfer one person's facial expressions and postures to another person's face, achieving highly realistic results.
- Social media and content creationDynamicFace helps creators produce fun and personalized short videos and images for social media. Users can replace their own facial features with images of celebrities or famous people to generate interesting and creative videos.
- Virtual meetings and live streamingUsers can use virtual cameras to replace faces in real time during live streams or virtual meetings, bringing a brand-new visual experience to the audience.
- Personal entertainment and creativityUsers can replace their own faces in various interesting scenarios to generate personalized emojis or creative videos.