AB
AiBoss
project

Draw an Audio - A video-to-audio system jointly launched by the Chinese Academy of Sciences and Meituan

Draw an Audio is a video-to-audio generation system developed by researchers from the Institute of Automation, Chinese Academy of Sciences, and Meituan-Dianping. It automatically generates matching sound effects based on video content, similar to Foley art direction in filmmaking...

What is Draw an Audio?

Draw an Audio is a video-to-audio generation system developed by researchers from the Institute of Automation, Chinese Academy of Sciences, and Meituan-Dianping. It automatically generates matching sound effects based on video content, similar to Foley art in filmmaking. The system analyzes the video and combines various input instructions, such as text, video masks, and loudness signals, to generate audio that matches the video content, time, and loudness. Its core architecture includes a Latent Diffusion Model (LDM), a Text Conditional Model, a Mask Attention Module (MAM), and a Time-Loudness Module (TLM), which together ensure high-quality and accurate audio generation. It provides video content creators with a powerful tool, making the sound design process more efficient and flexible.

Draw an Audio's main functions

  • Content consistencyThe system analyzes video content and generates sounds that match the semantics of the video scene, such as generating corresponding animal sounds when animals appear in the video.
  • Time ConsistencyThe generated audio is precisely synchronized with the actions in the video, ensuring that sound effects appear at the correct time, such as the sound of an object colliding and the collision action occurring simultaneously in the video.
  • Loudness consistencyThe system adjusts the volume of the sound based on the intensity of the action in the video. For example, the sound of distant objects in the video is relatively quiet, while the sound of nearby objects is louder.
  • Multiple command inputThe system supports a variety of input commands, including the video itself, related text descriptions, video masks, and loudness signals, making audio generation more flexible and controllable.
  • High-quality synchronized audioBy utilizing multiple commands, Draw an Audio can generate high-quality audio that is naturally synchronized with the video content, enhancing the viewing experience.

The technical principles of Draw an Audio

  • Latent Diffusion Model (LDM)As a basic model, it is responsible for the basic generation and processing of audio data.
  • Text Conditional ModelProcess text instructions to ensure that the generated audio matches the text description, improving the semantic consistency of the content.
  • Masked-Attention Module (MAM)By using video masking, we can focus on key areas of the video and enhance the consistency between the video content and the generated audio.
  • Time-Loudness Module (TLM)Process signal instructions, such as loudness signals, to ensure that the generated sound is synchronized with the video in terms of time and loudness.

Draw an Audio project address

Application scenarios of Draw an Audio

  • Film and video productionIn film and television post-production, Draw an Audio automatically adds matching sound effects, such as footsteps and vehicle sounds, to silent videos, improving production efficiency and reducing costs.
  • Game developmentIt generates realistic sound effects for animations and scenes in the game, enhancing the player's immersion and gaming experience.
  • Virtual Reality (VR) and Augmented Reality (AR)Generate sounds that match the scene in a virtual environment to enhance the user's interactive experience and perceived realism.
  • Education and trainingIt automatically generates explanatory audio for educational videos, helping students better understand and absorb knowledge.
  • Animation ProductionIt automatically generates dialogue and ambient sound effects for animated characters, making animation production more efficient.
  • Advertising productionGenerate engaging audio effects for ad videos to enhance their appeal and memorability.