AB
AiBoss
project

AudioGen-Omni - A multimodal audio generation framework launched by Kuaishou

AudioGen-Omni is a multimodal audio generation framework launched by Kuaishou. The framework can generate high-quality audio, speech, and songs based on inputs such as video and text. The framework utilizes a unified lyrics-text encoder and phase-aligned anisotropic...

What is AudioGen-Omni?

AudioGen-Omni is a multimodal audio generation framework launched by Kuaishou. The framework can generate high-quality audio, speech, and songs based on inputs such as video and text. Through a unified lyrics-text encoder and Phase Aligned Anisotropic Position Injection (PAAPI) technology, the framework achieves accurate audio-visual alignment and cross-modal synchronization. It supports multilingual input, boasts fast inference speed (generating 8 seconds of audio in 1.91 seconds), and performs exceptionally well on various audio generation tasks, making it suitable for scenarios such as video dubbing, speech synthesis, and song creation.

Main functions of AudioGen-Omni

  • Multimodal audio generationGenerate high-quality audio, speech, and songs from video, text, or a combination of both.
  • Precise audiovisual alignmentBased on Phase Aligned Anisotropic Position Injection (PAAPI) technology, lip-sync and rhythm alignment of audio and video are achieved.
  • Multilingual supportIt supports multiple language inputs and generates corresponding language voice and songs.
  • Efficient ReasoningIt has a fast inference speed, generating 8 seconds of audio in 1.91 seconds, which is significantly better than similar models.
  • Flexible input conditionsIt can handle cases with missing modalities, and can generate stable audio output even with only video or text input.
  • High-quality audio generationThe generated audio closely matches the input in both semantics and acoustics, supporting high-fidelity audio generation.

AudioGen-Omni's technical principles

  • Multimodal Diffusion Transformer (MMDiT)It integrates video, audio, and text modalities into a shared semantic space, supporting various audio generation tasks. Based on a joint training paradigm, it enhances cross-modal associations using large-scale video-text-audio data.
  • Lyrics-Text Unified EncoderEncodes text (grapheme) and phonemes into frame-level dense representations, adapted for speech and singing tasks. Uses multilingual unified word segmentation and ConvNeXt refinement to generate frame-aligned representations.
  • Phase-aligned anisotropic injection (PAAPI)Selectively apply Rotation Position Encoding (RoPE) to temporal modalities (such as video and audio) to improve cross-modal temporal alignment accuracy.
  • Dynamic condition mechanismBased on unfreezing all modalities and masking missing inputs, it avoids the semantic limitations of the text freeze paradigm and supports flexible multimodal conditional generation.
  • Joint attention mechanismA method based on AdaLN (Adaptive Layer Normalization) enhances cross-modal feature fusion and promotes cross-modal information exchange through a joint attention mechanism.

AudioGen-Omni's project address

  • Project official websitehttps://ciyou2.github.io/AudioGen-Omni/
  • arXiv technical paper: https://arxiv.org/pdf/2508.00733

Application scenarios of AudioGen-Omni

  • Video dubbingIt automatically generates precisely matched voice, songs, or sound effects for videos, improving video creation efficiency and content richness.
  • Speech SynthesisIt can quickly convert text into natural and fluent speech, and is suitable for audiobooks, voice assistants, intelligent customer service and other fields.
  • SongwritingGenerate matching songs based on video content or lyrics to assist in music creation and enrich video background music.
  • Sound effect generationGenerate natural environmental sound effects and motion sound effects based on text descriptions or video content to enhance the immersive experience of the content.