One Shot, One Talk - A dynamic image generation technology jointly developed by the University of Science and Technology of China and Hong Kong Polytechnic University.
One Shot, One Talk is an advanced image generation technology that can generate fully animated, talking avatars with personalized details from a single image, supporting realistic animation effects, including natural facial expressions and vivid body movements...
What is One Shot, One Talk?
One Shot, One Talk is an advanced image generation technology that can generate fully animated, talking avatars with personalized details from a single image. It supports realistic animation effects, including natural facial expressions and vivid body movements. Developed by researchers at the University of Science and Technology of China and the Hong Kong Polytechnic University, One Shot, One Talk combines a pose-guided image-to-video diffusion model with a 3DGS-mesh hybrid avatar representation to generalize to new poses and expressions, creating realistic, precisely animated, and expressive full-body talking avatars from a single image.
The main features of One Shot, One Talk
- Single image reconstructionReconstruct a full-body, dynamic speaking avatar from a single image.
- Realistic animationSupports realistic animation effects, including body movements and facial expressions.
- Personalized details: To capture and reproduce the individual characteristics and details of a person.
- Precise controlIt provides precise control over avatar poses and expressions.
- Generalization abilityIt can generalize to new postures and expressions, even those not seen during training.
The technical principles of One Shot, One Talk
- Pose-guided image-to-video diffusion modelThe model generates imperfect video frames as pseudo-labels to generalize to new poses and expressions.
- 3DGS-mesh hybrid avatar representationCombining 3D Gaussian models (3DGS) and parametric mesh models (such as SMPL-X) enhances the expressiveness and realism of avatars.
- Key regularization techniquesRegularization techniques are applied to mitigate inconsistencies caused by pseudo-tags, ensuring the accuracy of avatar structure and dynamic modeling.
- Pseudo-tag generationUsing datasets such as the TED Gesture Dataset to drive a pre-trained model, we can generate video sequences of target individuals performing different poses and expressions.
- Loss function and constraintsThe design incorporates multiple loss functions and constraints, including perceptual loss (such as LPIPS) and pixel-level loss, to effectively extract information from the input image and pseudo-labels and stabilize the head reconstruction process.
- Optimization and trainingThe Adam optimizer is used for training, and different loss functions are balanced based on carefully designed loss weights to achieve the best head reconstruction effect.
One Shot, One Talk project address
- Project official website:xiangjun-xj.github.io/OneShotOneTalk
- arXiv technical paper:https://arxiv.org/pdf/2412.01106
Application scenarios of One Sho, One Talk
- Augmented Reality (AR) and Virtual Reality (VR)In AR/VR applications, creating realistic virtual characters enhances user immersion and interactive experience.
- Remote conferencing and remote presentationBased on generating realistic full-body dynamic avatars, it can be used in remote meetings to make remote communication more natural and efficient.
- Games and entertainmentIn game and film production, it enables the rapid generation or customization of characters, reducing the time and cost of traditional motion capture and modeling.
- Social media and content creationUsers create personalized virtual avatars to use on social media platforms or as virtual anchors for content creation.
- Education and trainingIn a virtual teaching environment, teachers have realistic virtual avatars, which enhances the effectiveness of remote teaching.