AB
AiBoss
project

MultiFoley - A sound generation system developed by Adobe in collaboration with the University of Michigan

MultiFoley is a sound generation system jointly developed by Adobe Research and the University of Michigan. It can generate Foley sound effects based on multimodal control of text, audio, and video. The system allows users to generate Foley sound effects based on text prompts, reference audio, etc.

What is MultiFoley?

MultiFoley, a sound generation system jointly developed by Adobe Research and the University of Michigan, can generate Foley sound effects based on multimodal control of text, audio, and video. The system allows users to customize and generate sound synchronized with videos based on text prompts, reference audio, or portions of video, enhancing the video viewing experience. MultiFoley is trained on an internet video dataset and professional sound recordings to achieve high-quality, full-bandwidth (48kHz) audio generation. MultiFoley provides flexible sound design control for video production, helping users create clean and creative sound effects.

MultiFoley's main functions

  • Text-controlled Foley generationUse text prompts to guide and generate sound effects synchronized with the video, whether realistic or creative.
  • Foley generation with audio controlIt allows users to select reference audio from the sound effects library, apply the sound to silent videos, and synchronize it with the video.
  • Foley Audio Extensions: Expand a portion of the audio track to produce the complete Foley sound.
  • Quality controlBased on adding quality tags to the text, high-quality full-band (48kHz) audio is generated.
  • Multimodal controlIt combines conditional signals from text, audio, and video to provide fine-grained control over sound design.

The technical principles of MultiFoley

  • Joint trainingTrain on internet video datasets (low-quality audio) and professional sound effects (SFX) recordings to generate high-quality full-band audio.
  • Diffusion Transformer: Generate new samples from random noise based on a diffusion model, use them for video-guided Foley sound generation, and combine them with multimodal control.
  • High-quality audio autoencoder (DAC-VAE)Based on the variational autoencoder (VAE), a 48kHz audio waveform is encoded into a 40Hz latent feature for use in audio-video synchronization.
  • Frozen Video EncoderUsed in audio-video synchronization, it encodes video into features and uses them in conjunction with the underlying audio encoding.
  • Multi-condition training strategyThis allows the model to flexibly support downstream tasks, such as audio extensions and text-driven sound design.
  • Multi-head attention mechanismEnhance the model's expressive power and learn different types of features or dependencies in parallel.

MultiFoley project address

MultiFoley application scenarios

  • Film and video productionIn film production, generating sound effects synchronized with on-screen actions, such as footsteps and door closing sounds, enhances the audience's immersion.
  • Game developmentIn the game, realistic sounds are generated for different game environments and actions, enhancing the gaming experience.
  • Animation ProductionFor animation, corresponding sounds are generated based on the movements of the animated characters, making the animation more vivid.
  • Advertising productionIn the advertising industry, eye-catching sound effects are generated based on advertising creatives to increase the appeal of advertisements.
  • Virtual Reality (VR)In VR experiences, generating sounds synchronized with the virtual environment enhances the user's immersion and experience quality.