AB
AiBoss
project

SeedFoley - ByteDance's end-to-end video audio generation model

SeedFoley is an end-to-end video audio generation model developed by ByteDance's Doubao Big Model Speech Team, providing intelligent audio generation services for video creation. By integrating spatiotemporal video features with a diffusion generation model, it achieves audio effects and...

What is SeedFoley?

SeedFoley is an end-to-end video audio generation model developed by ByteDance's Doubao Big Model Speech Team, providing intelligent audio generation services for video creation. By fusing spatiotemporal video features with a diffusion generation model, it achieves high synchronization between audio and video. The model employs a video encoder combining fast and slow features to extract spatiotemporal features from the video, while an audio representation model based on the original waveform as input retains high-frequency information, enhancing the subtlety of the audio effects. The diffusion model reduces the number of inference steps and lowers inference costs by optimizing the continuous mapping relationship on the probability path. SeedFoley can accurately extract frame-level visual information from videos, intelligently distinguish between action audio effects and environmental audio effects, supports various video lengths, and demonstrates excellent performance in audio accuracy, synchronization, and matching.

SeedFoley's main functions

  • Intelligent sound effect generationSeedFoley can accurately extract visual information at the frame level from videos. By analyzing information from multiple frames, it can accurately identify the main speaker and action scenes in the video, such as moments of music with a strong rhythm or tense plots in movies. It can precisely time the moments to create an immersive and realistic experience.
  • Distinguish sound effect typesSeedFoley can intelligently distinguish between action sound effects and ambient sound effects, significantly improving the narrative tension and emotional delivery efficiency of videos.
  • Supports multiple video lengthsSeedFoley supports variable-length video input and has achieved leading levels in metrics such as audio accuracy, audio synchronization, and audio matching.

SeedFoley's technical principles

  • Video encoderSeedFoley's video encoder employs a combination of fast and slow features. It extracts local motion information between frames at high frame rates and semantic information from the video at low frame rates. This approach allows the model to achieve frame-level video feature extraction at 8fps with low computational resources, enabling fine-grained motion localization. Finally, it fuses the fast and slow features using a Transformer structure to extract spatiotemporal features from the video.
  • Audio representation modelUnlike traditional Mel-spectrum-based VAE models, SeedFoley uses the original waveform as input, which is then encoded to obtain a 1D representation. The audio is sampled at 32kHz to ensure high-frequency information is preserved. 32 latent audio representations are extracted per second of audio, effectively improving the temporal resolution of the audio and enhancing the subtlety of the sound effects.
  • diffusion modelSeedFoley employs the Diffusion Transformer framework, optimizing continuous mapping relationships along probabilistic paths to achieve probabilistic matching from Gaussian noise distributions to the target audio representation space. Compared to traditional diffusion models that rely on Markov chain sampling, SeedFoley effectively reduces the number of inference steps and lowers inference costs by constructing continuous transformation paths. During the training phase, video features and audio semantic labels are encoded into latent space vectors, which are then concatenated along the channel dimension and mixed with temporal encoding and noise signals to form a joint conditional input. This improves the temporal consistency between sound effects and video visuals.

How to use SeedFoley

  • Visit the JiMeng platformVisit the official website of JiMeng or use the JiMeng App to register and log in.
  • Generate videoSelect the video generation function on Jimeng to generate video content according to your needs.
  • Select the "AI Sound Effects" functionAfter generating the video, select the "AI Audio Effects" feature. The system will automatically generate three professional-grade audio effect schemes for your video.
  • Preview and select sound schemesPreview the generated sound effect scheme and choose the one that best suits your video content.
  • Application sound effectsApply the selected sound effects scheme to your video.
  • Precautions:
    • Video lengthSeedFoley supports variable-length video input, but it is recommended that the video length not be too long to ensure the generated effect.
    • Sound effects typeSeedFoley can intelligently distinguish between action sound effects and ambient sound effects, enhancing the narrative tension and emotional delivery efficiency of videos.
    • PreviewWhen choosing a sound effect scheme, it is recommended to carefully preview the effect of each scheme and choose the sound effect that best suits your video content.

Application scenarios of SeedFoley

  • Life VlogAdd realistic ambient sound effects to your personal vlog, such as street noise or background music from a coffee shop.
  • Short film productionAdd action sound effects and ambient sound effects to the short film to match the plot and enhance the audience's immersion.
  • Game DevelopmentAdd realistic sound effects to game videos, such as combat sound effects and environmental sound effects, to enhance the gaming experience.
  • Video post-productionIn video post-production, SeedFoley can quickly generate sound effects that closely match the video content, saving time and costs in post-production.
  • Ad videoAdding engaging sound effects to advertising videos enhances their appeal and spread.
  • Educational VideosAdding appropriate sound effects to educational videos can enhance viewers' learning interest and attention.