AB
AiBoss
project

Huya VAM 1.0 - Huya's real-time multimodal digital human basic model

Huya VAM 1.0 (Vivid Avatar Model) is a real-time multimodal digital human basic model launched by Huya based on the DiT architecture. It can generate an AI digital human that can talk, sing and dance from a single photo.

What is Huya VAM 1.0?

Huya VAM 1.0 (Vivid Avatar Model) is a real-time multimodal digital human basic model launched by Huya based on the DiT architecture. It can generate an AI digital human that can talk, sing, and dance from a single photo. The model enables 24/7 real-time live streaming interaction with a resolution of 480×832 and a frame rate of 28. It supports full-duplex dialogue, instant interruption, bullet screen replies, and multi-role strategy games. It leads the way in realism, identity preservation, and reasoning speed, and is suitable for scenarios such as live e-commerce, news broadcasting, and virtual concerts.

Main functions of Huya VAM 1.0

  • One-click digital human creation from photosUpload a photo and it will generate a real-time AI digital human avatar that can talk, sing, and dance.
  • Full-duplex real-time dialogueIt supports both text and voice input, allowing users to interrupt and respond instantly, achieving a smooth, human-like interaction.
  • Multi-talent live performanceIt can generate singing, dancing, and other content in real time, with lip movements synchronized with lyrics and natural, smooth body movements.
  • Multi-role strategy gameSupports complex multiplayer interactive games such as Werewolf and Tarot, with AI characters possessing independent stances and speaking styles.
  • 24/7 live streaming480×832 resolution, 28 frames per second streaming output, can run continuously for more than 24 hours without crashing or distortion.
  • Real-time interaction with bullet commentsIt supports reading and replying to live stream comments in real time, adapting to scenarios such as live e-commerce and news broadcasts.

The technical principles of Huya VAM 1.0

  • DiT Multimodal ArchitectureBuilt on Diffusion Transformer, it integrates VAE image encoding, text encoding and audio encoding, and generates them uniformly by concatenating channels into the DiT Block.
  • Triple cross-attention mechanismDiT Block embeds Self-Attention, Text & Image Cross-Attention, and Adaptive Audio Cross-Attention to handle self-attention, image-text alignment, and audio-driven lip-syncing, respectively.
  • Motion-ControllerThe introduction of a motion latent variable control module enriches the diversity of facial expressions and movements, enabling the head and limbs to slow down synchronously when speech is paused and to nod in time with the beat when music is heard.
  • Three-stage progressive trainingThe first stage uses multiple reference images and motion frames to anchor the character and feeds them into a degraded scene to train stability; the second stage uses DPO preference to optimize and balance multiple objectives such as mouth shape, expression, and movement; the third stage compresses the number of inference steps from 20 steps to 4 steps through model distillation.
  • Self-correction mechanismDuring inference, the already generated frames are used as input to continue generating. During the training phase, it learns to self-correct and prevents accumulated errors from causing facial drift and screen tearing.

How to use Huya VAM 1.0

The model is currently in the internal testing/invitation-only phase and has not yet been made publicly available.

VAM 1.0's core advantages

  • stableMulti-reference image anchoring + motion frame strategy + self-correction mechanism ensures no crashes, distortions, or tearing for 24 consecutive hours.
  • allowIt natively covers three states: silence, listening, and speaking, and the precision of micro-expression and body movement control is close to that of a real person.
  • quickThe first frame latency is approximately 1.3 seconds, the segment generation latency is only 0.77 seconds, and the 8×H200 GPU achieves 36.4 FPS, the fastest in the industry.
  • ProvinceModel distillation reduces the number of inference steps from 20 to 4, with significantly lower computational overhead than similar solutions.
  • realDPO preference optimization balances lip shape, facial expressions, and movements across multiple objectives, maintaining a leading edge in realism and identity.

VAM 1.0 Comparison with Similar Products

Comparison Dimensions Huya VAM 1.0 OmniHuman 1.5
Architecture DiT (Diffusion Transformer) Diffusion model + audio driver
Real-time Real-time streaming output, 28 FPS Not real-time, video needs to be pre-generated.
Interactive capabilities Full-duplex conversation, supports interruption/response. One-way broadcast, no real-time interaction
Continuous operation 24/7 stable live streaming Unable to run continuously for a long time
Input method Photos + Text/Voice/Bullet Comments Photos + Audio
Application scenarios Live-streaming e-commerce, interactive games, and virtual companionship Short video generation, voice-over video
Delay 0.77 seconds/segment Minute-level generation
Multiple roles Supports 10-player simultaneous Werewolf game Single-role drive

Application scenarios of VAM 1.0

  • AI-powered live-streaming e-commerceDigital human anchors are online 24/7, reading real-time comments and interactions, recommending products, and answering questions.
  • Virtual news broadcastNews anchors broadcast around the clock, maintaining a consistent and professional image, with fluent speech and natural body language.
  • Virtual concertAI singer performs in real time, with lip movements synchronized with the music beat, and supports continuous performances of multiple music styles.
  • Game companion interactionIn strategy games such as Tarot readings and Werewolf, AI characters possess independent personalities and game-playing abilities.
  • Emotional companionship chatA personalized AI assistant that supports dialect dialogue, remembers user preferences, and provides immersive companionship.