AB
AiBoss
project

MDT-A2G - An AI model developed by Fudan University, Tencent YouTu, capable of generating gestures synchronously based on voice.

MDT-A2G is an AI model jointly developed by Fudan University and Tencent YouTu, specifically designed to synchronously generate corresponding gestures based on speech content. MDT-A2G mimics the natural gestures humans make during communication, enabling computers to create more vivid and...

What is MDT-A2G?

MDT-A2G is an AI model jointly developed by Fudan University and Tencent YouTu, specifically designed to generate corresponding hand gestures synchronously based on speech content. MDT-A2G mimics the natural gestures humans make during communication, allowing computers to "perform" more vividly and naturally. MDT-A2G comprehensively analyzes various information sources, including speech, text, and emotion, and uses techniques such as noise reduction and accelerated sampling to generate coherent and realistic gesture sequences.

Main functions of MDT-A2G

  • Multimodal information fusionIt combines multiple information sources such as voice, text, and emotion for comprehensive analysis to generate gestures synchronized with voice.
  • Noise reduction processingBy using noise reduction technology, we correct and optimize gesture movements to ensure that the generated gesture movements are accurate and natural.
  • Accelerated samplingIt employs an efficient inference strategy, utilizing the results of previous calculations to reduce the computational load for denoising, thereby achieving rapid generation.
  • Time-aligned contextual reasoningIt enhances the learning of the temporal relationships between gesture sequences, producing coherent and realistic movements.

MDT-A2G Technical Principles

  • Multimodal feature extractionThe model extracts features from multiple information sources, including speech, text, and emotion. It involves speech recognition technology to convert speech into text, and emotion analysis to identify the speaker's emotional state.
  • Masked diffusion converterMDT-A2G uses a novel masked diffusion transformer architecture. It generates the target output by introducing randomness into the data and then progressively removing this randomness, similar to a denoising process.
  • Time alignment and contextual reasoningThe model needs to understand the temporal relationship between speech and gestures, ensuring that gestures are synchronized with speech. This involves sequence models, which are capable of processing time-series data and learning temporal dependencies.
  • Accelerate the sampling processTo improve generation efficiency, MDT-A2G employs a scale-aware accelerated sampling process. The model uses the results of previous calculations to reduce subsequent computations, thereby speeding up gesture generation.
  • Feature fusion strategyThe model employs an innovative feature fusion strategy, combining temporal embeddings with emotional and identity features, and integrating them with text, audio, and gesture features to generate a comprehensive feature representation.
  • Denoising processDuring the gesture generation process, the model gradually removes noise and optimizes gesture movements to ensure that the generated gestures are both accurate and natural.

MDT-A2G project address

Application scenarios of MDT-A2G

  • Enhance interactive experienceVirtual assistants can enhance non-verbal communication with users through gestures generated by the MDT-A2G model, making the dialogue more natural and human.
  • Education and trainingVirtual teachers or training assistants can use gestures to assist teaching, improving learning efficiency and engagement.
  • Customer ServiceIn customer service scenarios, virtual customer service assistants can use gestures to express information more clearly, thereby improving service quality and user satisfaction.
  • Assisting people with disabilitiesFor people with hearing or speech impairments, virtual assistants can provide a more easily understood way of communicating through gestures.