AB
AiBoss
project

AudioFly - iFlytek's open-source Vinyl sound effect model

AudioFly is an open-source AI model from iFlytek for generating audio effects from text. The model uses a latent diffusion model architecture, has 1 billion parameters, and utilizes a large number of open datasets (such as AudioSet, AudioCaps, and TUT) as well as internal proprietary data...

What is AudioFly?

AudioFly is an open-source AI model from iFlytek for generating audio effects from text. The model uses a latent diffusion model architecture with 1 billion parameters, trained on a large number of open datasets (such as AudioSet, AudioCaps, and TUT) and internal proprietary data. AudioFly can generate high-quality audio from text descriptions, with a sampling rate of up to 44.1kHz, and the generated sound effects closely match the text descriptions. The model performs excellently in both single-event and multi-event scenarios, and its performance on the AudioCaps dataset is outstanding, surpassing previous audio generation models. AudioFly is suitable for short video dubbing, audio story generation, and other fields, bringing unlimited possibilities to sound creation.

AudioFly's main functions

  • Text-to-sound generationIt generates corresponding sound effects based on the text description input by the user. For example, if the user inputs "thunder is booming in the distance," the model can generate the corresponding thunder sound effect.
  • High-quality audio outputThe generated audio has a sampling rate of 44.1kHz, with clear sound quality, making it suitable for various application scenarios.
  • Multi-scenario supportIt supports sound effect generation for single-event (such as "dog barking") and multi-event (such as "dog barking and wind sound") scenarios, and can accurately reflect the described content.
  • High-efficiency generationBased on an advanced diffusion model architecture, the generation process is efficient and can quickly respond to user needs.

AudioFly's technical principles

  • Potential Diffusion Model (LDM) ArchitectureAudioFly uses a latent diffusion model architecture, a generative model based on deep learning. The model generates target audio by progressively removing noise, similar to the diffusion process in image generation.
  • Large-scale data trainingThe model is trained on a large number of open datasets (such as AudioSet, AudioCaps, TUT) and internal proprietary data, which cover a variety of sound effects and scenes, enabling the model to generate diverse sound effects.
  • Feature AlignmentBy optimizing the training objectives of the model, we ensure that the generated audio is highly consistent with the real audio in terms of features, and closely aligned with the text description in terms of content.

AudioFly's project address

  • Magic Dash Communityhttps://modelscope.cn/models/iflytek/AudioFly

Application scenarios of AudioFly

  • Short video dubbingIt can quickly generate matching sound effects for short videos, enhancing their appeal and immersion.
  • Audio story creation: Generate sound effects based on text content to enhance the atmosphere and emotional expression of the story.
  • Film and television sound effects productionIt assists film and television production teams in quickly generating the sound effects they need, thereby improving production efficiency.
  • Game sound designGenerate real-time sound effects for game scenes to enhance player immersion and experience.
  • Advertising and MarketingGenerate custom sound effects for advertising video or audio content to enhance the appeal and memorability of the ads.