AB
AiBoss
project

FlowAct-R1 - ByteDance's real-time interactive digital human video generation framework

FlowAct-R1 is a real-time interactive digital human video generation framework launched by ByteDance. It only requires a single reference image and audio to support the streaming generation of full-body dynamic videos of unlimited length.

What is FlowAct-R1?

FlowAct-R1 is a real-time interactive digital human video generation framework launched by ByteDance. It requires only a single reference image and audio, and supports streaming generation of full-body dynamic videos of unlimited duration. The framework achieves low latency (1.5-second first frame) and stable real-time response at 25fps through a block diffusion forced strategy and a multimodal large language model. It can precisely control the facial expressions and body movements of digital humans, making it suitable for scenarios such as video conferencing, virtual companionship, and live interactive streaming. It has strong generalization capabilities and can drive various styles of characters.

Main functions of FlowAct-R1

  • Real-time interaction and unlimited duration generationThe framework requires only a single reference image and audio input, and can stream unlimited full-body dynamic videos, supporting stable operation over long periods without common issues such as unnatural facial expressions.
  • Low latency and high frame rateThe framework can achieve a low latency of 1.5 seconds for the first frame and a stable real-time response of 25fps, ensuring a smooth and natural interaction process, and is suitable for scenarios such as video conferencing and live streaming.
  • Full body movement and facial expression control: By using multimodal commands to precisely control the facial expressions and body movements of digital humans, such as listening, thinking, and gestures, the interaction becomes more vivid and realistic.
  • Strong generalization abilityThe framework is not limited to specific characters; it can drive various styles of characters from a single reference image, including realistic photos, anime, and artistic styles.

Technical principles of FlowAct-R1

  • Streaming generation and infinite duration:frameA block-based diffusion forced strategy is adopted to divide the video into small blocks and generate them one by one. A structured memory library is used to ensure the continuity of the images, so as to achieve theoretically unlimited duration generation.
  • Real-time performance optimization:Frame LoveBy combining multi-stage distillation technology, the number of denoising steps in the diffusion model is reduced to 3 steps.By combining FP8 quantization and operator fusion, the overhead of video memory read and write is greatly reduced, ultimately achieving real-time generation capability of 25fps and 480p.
  • Whole body control and behavior planning:Frame LoveBy introducing a multimodal large language model as the "brain," the digital human can determine the actions it should perform based on speech and context, thereby achieving fine-grained natural motion planning and eliminating the mechanical feel.
  • High-fidelity visual effects:frameMaintaining high-fidelity visuals during the generation process, the system employs optimized model architecture and training strategies to ensure high-quality video performance across different styles and scenarios.

FlowAct-R1 project address

  • Project official website: https://grisoon.github.io/FlowAct-R1/
  • arXiv technical paper: https://arxiv.org/pdf/2601.10103

Application scenarios of FlowAct-R1

  • AI Live StreamingThe framework enables 24/7 uninterrupted, real-time interactive live streaming, supports multiple languages and style switching, and enhances audience engagement.
  • videoconferenceAs a virtual participant, it provides natural body language and interaction to enhance the realism of the meeting and supports multilingual translation.
  • Virtual companionshipIt generates personalized virtual companions, provides emotional support and interactive entertainment, and meets users' needs for companionship.
  • Online EducationAs a virtual teacher, it provides engaging teaching and personalized tutoring, supporting multilingual instruction.
  • Customer ServiceAs a virtual customer service representative, it answers customer questions in real time, provides multilingual support, and improves customer satisfaction.