AB
AiBoss
project

AudioGenie - A multimodal audio generation tool launched by Tencent AI Lab

AudioGenie is a multimodal audio generation tool developed by Tencent AI Lab. It can generate various audio outputs, including sound effects, speech, and music, from multiple modal inputs such as video, text, and images. The tool employs a training-free multi-agent approach...

What is AudioGenie?

AudioGenie is a multimodal audio generation tool developed by Tencent AI Lab. It can generate various audio outputs, including sound effects, speech, and music, from multiple modal inputs such as video, text, and images. The tool employs a training-free multi-agent framework, achieving efficient collaboration through a two-layer architecture of a generation team and a supervision team. The generation team is responsible for decomposing complex inputs into specific audio sub-events and dynamically selecting the most suitable model for generation through an adaptive hybrid expert (MoE) collaboration mechanism. The supervision team is responsible for spatiotemporal consistency verification, using a feedback loop for self-correction to ensure highly reliable generated audio.

AudioGenie has established the world's first benchmark set, MA-Bench, for the multimodal to multi-audio generation (MM2MA) task, which includes 198 videos with various audio annotations. In the tests, AudioGenie achieved or approached state-of-the-art performance in 8 out of 9 metrics, with particularly outstanding performance in audio quality, accuracy, content alignment, and aesthetic experience.

AudioGenie's main functions

  • Multimodal input and multi-audio outputIt supports input from multiple modalities such as video, text, and images, and generates various audio types such as sound effects, speech, and music.
  • Untrained multi-agent frameworkThe system employs a two-tier architecture: the generation team is responsible for task decomposition and dynamic model selection, while the oversight team is responsible for verification and self-correction to ensure the reliability of the output.
  • Refined task breakdownIt decomposes complex multimodal inputs into specific audio sub-events, accurately labels audio type, start and end time, and content description, forming a structured generation blueprint.
  • Trial and error and iterative optimizationThe system employs an iterative optimization process based on a "mind tree". It generates candidate audio files, which are then evaluated by the supervisory team based on dimensions such as quality, alignment, and aesthetics. If any flaws are found, the system automatically triggers a correction or retry process until the output meets the requirements.

AudioGenie's technical principles

  • Two-layer multi-agent architectureThe system employs a two-tiered architecture consisting of a generation team and a supervision team. The generation team is responsible for decomposing and executing the audio generation task, while the supervision team is responsible for verifying the spatiotemporal consistency of the output and providing feedback to optimize the generation results.
  • Adaptive Hybrid Expert (MoE) CollaborationBased on different audio subtasks, the most suitable model is dynamically selected for generation, and the generation scheme is optimized through a collaborative correction mechanism among experts to improve generation quality and efficiency.
  • No training frameworkThe use of a training-free multi-agent system avoids the problems of data scarcity and overfitting in traditional training methods, thereby improving the system's generalization ability and adaptability.
  • Spatiotemporal consistency verificationThe oversight team verifies the spatiotemporal consistency of the generated audio through feedback loops, ensuring that the generated audio is consistent with the input content in both time and space.

AudioGenie's project address

  • Project official websitehttps://audiogenie.github.io/

Application scenarios of AudioGenie

  • Film and television productionIt can quickly generate background music, ambient sound effects, and character voice-overs that closely match the video content, improving production efficiency and enhancing audience immersion.
  • Virtual character voice actingIt generates natural and fluent voices for virtual anchors, virtual customer service representatives, and other virtual characters, making them more expressive and realistic.
  • Game developmentIt automatically generates realistic environmental sound effects, background music, and character voices based on the game scene, enhancing the player's immersion and gaming experience.
  • Podcast ProductionAutomatically generate background music that follows the plot of the podcast, enhancing its appeal and professionalism.
  • Commercial editing: Quickly match brand tone with sound effects and music, saving production time and costs, and enhancing the appeal and impact of advertisements.