AB
AiBoss
project

Kling-Foley - A multimodal video audio generation model launched by Kling AI

Kling-Foley is a multimodal video audio generation model developed by Kling AI. The model takes video and text prompts as input and generates high-quality stereo audio that is semantically relevant to the video content and time-synchronized, covering sound effects...

What is Kling-Foley?

Kling-Foley is a multimodal video audio generation model developed by Kling AI. Taking video and text prompts as input, the model generates high-quality stereo audio that is semantically relevant to the video content and time-synchronized. It covers various types of sound content, including sound effects and background music, and supports audio generation of any length. Based on a multimodal controlled stream matching architecture, the model uses multimodal feature fusion and specific module processing to accurately achieve audio-video alignment. Trained on a large-scale self-built multimodal dataset, the model demonstrates excellent audio generation performance, placing it at the forefront of the audio generation field and providing a more efficient and high-quality audio solution for video content creation.

Kling-Foley's main functions

  • High-quality sound generationBased on the input video content and optional text prompts, it generates high-quality stereo audio that is semantically related to the video and time-synchronized, covering various types of sound content such as sound effects and background music to meet the audio needs of different scenarios.
  • Generate audio of any durationIt supports generating audio content of any length and can dynamically adapt to the length of the input video.
  • Stereo renderingIt has the capability of stereo rendering, supports spatially oriented sound source modeling and rendering, and makes the generated audio have a stronger sense of space and immersion.

Kling-Foley's technical principles

  • Flow matching model for multimodal controlKling-Foley is a multimodal controlled streaming matching model. Its core principle is to use text, video, and time-extracted video frames as conditional inputs, fuse them using a multimodal joint conditional module, and then input them into the MMDit module for processing. This multimodal control approach allows the model to better understand and generate audio that matches the video content.
  • Modular processing flowThe model's processing flow includes several key modules. Multimodal features are fused based on the multimodal joint conditional module and input into the MMDit module to predict VAE latent features. The latent features are reconstructed into a mono-channel Mel spectrogram by a pre-trained Mel decoder. The mono-channel spectrogram is rendered into a stereo spectrogram based on the Mono2Stereo module, and the output waveform is generated using a vocoder.
  • Visual semantic representation and audio/video synchronization moduleThe Kling-Foley architecture introduces a visual semantic representation module and an audio-video synchronization module, which supports the alignment of video conditions and audio latent elements at the frame level, improving the effect of video semantic alignment and audio-video synchronization, and ensuring that the generated audio is highly matched with the video in terms of time and content.
  • Discrete Duration EmbeddingKling-Foley introduces discrete-time embeddings as part of its global conditional mechanism. This allows the model to better handle video inputs of varying lengths and generate audio content that adapts to the video length.
  • Universal Subsurface Audio CodecAt the audio latent representation level, Kling-Foley applies a universal latent audio codec, enabling high-quality modeling across diverse scenarios such as sound effects, speech, singing, and music. The core component is Mel-VAE, which jointly trains the Mel encoder, Mel decoder, and discriminator, allowing the model to learn a continuous and complete latent spatial distribution, significantly enhancing its audio representation capabilities.

Kling-Foley's project address

  • Project official websitehttps://klingfoley.github.io/Kling-Foley/
  • GitHub repositoryhttps://github.com/klingfoley/Kling-Foley
  • arXiv technical paper: https://www.arxiv.org/pdf/2506.19774

Application scenarios of Kling-Foley

  • Video content creationIt provides precisely matched sound effects and background music for the production of videos such as animations, short videos, and advertisements, enhancing the appeal and professionalism of the videos and improving creative efficiency.
  • Game developmentGenerate realistic scene sound effects and background music, such as weapon firing, character movements, and environmental sound effects, to enhance the game's immersion and player experience.
  • Education and TrainingAdding appropriate sound effects and background music to teaching videos and virtual training environments enhances the realism and appeal of teaching and training, thereby improving learning outcomes.
  • Film and television productionGenerate high-quality sound effects and background music for movies, TV series, and other film and television works, enhancing the sound quality and the emotional impact of the story.
  • social mediaUsers can quickly add matching sound effects and background music to their shared videos, enhancing the appeal of their content.