AB
AiBoss
project

FLOAT - An audio-driven speaker avatar generation model based on stream matching

FLOAT is an audio-driven speaker avatar generation model developed by DeepBrain AI and the Korea Advanced Institute of Science and Technology (KAIST). Based on a flow matching generation model, it learns a motion latent space to achieve efficient temporally consistent motion design. The model is based on...

What is FLOAT?

FLOAT is an audio-driven speaker avatar generation model developed by DeepBrain AI and the Korea Advanced Institute of Science and Technology (KAIST). Based on a flow matching generation model, it learns a motion latent space to achieve efficient temporally consistent motion design. The model utilizes a Transformer-based vector field predictor to achieve inter-frame temporal consistency and supports voice-driven emotion enhancement, making the generated speech movements more natural and expressive. FLOAT surpasses existing diffusion-based and non-diffusion-based methods in visual quality, motion fidelity, and generation efficiency, achieving industry-leading levels.

FLOAT's main functions

  • Audio-driven speaker image generationIt generates a speaking portrait video based on a single source image and driving audio, achieving audio-synchronized head movements, including verbal and non-verbal movements.
  • Time-consistent video generationBy modeling within the motion potential space, FLOAT generates videos with high temporal consistency, solving the temporal coherence problem in traditional diffusion-based video generation.
  • Emotional enhancementVoice-driven emotion tags enhance emotional expression in videos, making generated speech and actions more natural and expressive.
  • High-efficiency samplingBased on stream matching technology, improve the sampling speed and efficiency of video generation.

FLOAT's technical principles

  • potential space for movement: Shifting generative modeling from pixel latent space to learned motion latent space, more effectively capturing and generating temporally coherent motion.
  • Stream matchingBased on flow matching, efficient sampling is performed in the motion latent space to generate time-consistent motion sequences.
  • Transformer-based vector field predictorThe Transformer-based architecture predicts the vector field of the generated stream. The predictor can handle frame conditions and generate time-consistent motion.
  • Frame condition mechanismBased on a simple frame condition mechanism, driving audio and other conditions (such as emotion tags) are integrated into the generation process to achieve effective control over the motion potential space.
  • Emotional controlEmotional labels are generated using a pre-trained speech emotion predictor and then used as conditional inputs to a vector field predictor. Emotional control is introduced during the generation process.
  • Fast sampling and efficient generationBased on stream matching technology, the number of iterations in the generation process is reduced, enabling fast sampling and maintaining high quality of the generated video.

FLOAT's project address

Application scenarios of FLOAT

  • Virtual anchors and virtual assistantsIn fields such as news broadcasting, weather forecasting, and online education, it generates lifelike virtual anchors and provides 24/7 program production.
  • Video conferencing and remote communicationIn video conferencing, create virtual avatars for users so they can communicate via video even without a webcam.
  • Social media and entertainmentOn social media platforms, users generate their own virtual avatars for use in live streaming, interactive entertainment, or virtual social interactions.
  • Games and Virtual RealityIn games and virtual reality applications, it can be used to create or customize the facial expressions and movements of game characters to enhance immersion.
  • Film and animation productionIn film post-production, it generates or enhances a character's facial expressions and lip movements, reducing the need for traditional motion capture.