SongBloom - A full-length song generation model launched by Tencent AI Lab
SongBloom is a full-length song generation framework developed by Tencent AI Lab. It combines autoregressive sketching and diffusion-based refinement techniques, alternately generating semantic content through an interleaved generation paradigm...
What is SongBloom?
SongBloom is a full-length song generation framework developed by Tencent AI Lab. It combines autoregressive sketching and diffusion-based refinement techniques, using an interleaved generation paradigm to alternately generate semantic and acoustic context, resulting in high-quality, complete songs. The model requires only a 10-second audio sample and corresponding lyrics to generate a 2-minute-30-second dual-channel, 48kHz audio file. SongBloom excels in audio quality and lyric accuracy, approaching state-of-the-art (SOTA) performance, and has been successfully open-sourced.
SongBloom's main functions
-
High-efficiency song generationIt can generate a complete song of up to 2 minutes and 30 seconds with only 10 seconds of audio sample and corresponding lyrics, and supports dual-channel, 48kHz high-quality audio output.
-
Innovative Generative ParadigmIt employs an interleaved generation paradigm, combining autoregressive sketching and diffusion-based refinement techniques to alternately generate semantic and acoustic contexts, thereby optimizing the overall structure and sound quality of the song.
-
Superior sound quality and accuracyIt performs exceptionally well in terms of audio quality and lyric accuracy, approaching state-of-the-art (SOTA) levels and surpassing existing open-source models.
-
Open source and ease of useThe project is open source, providing detailed usage guides and multiple model versions, supporting operation on devices with low video memory, and making it easy for users to get started quickly.
-
Broad application prospectsIt provides powerful tools for music composition, audio production and other fields, which can significantly improve creative efficiency and inspire new inspiration for music creation.
SongBloom's technical principles
-
Interleaved generation paradigmBy alternately generating semantic and acoustic contexts and dynamically switching the generation process, the overall structure and sound quality of the song can be optimized.
-
Autoregressive sketch drawing: Generate music sketches using an autoregressive model to ensure structural coherence and phoneme alignment.
-
Diffusion model refinementThe generated sketches are refined using a diffusion model to improve audio quality with high fidelity.
-
Combination of discrete and continuous outputThe final result is output using discrete sketch tokens and VAE latents, balancing structure and sound quality.
-
Multimodal input fusionThe input includes lyrics and audio samples, and the model achieves accurate generation through multimodal fusion.
SongBloom's project address
- Github repositoryhttps://github.com/tencent-ailab/SongBloom
- HuggingFace model libraryhttps://huggingface.co/CypressYang/SongBloom
- arXiv technical paper: https://arxiv.org/pdf/2506.07634
- Experience the demo onlinehttps://cypress-yang.github.io/SongBloom_demo/
SongBloom's application scenarios
-
Music compositionIt provides inspiration for musicians and creators, quickly generating high-quality song frameworks to help them explore new musical styles and creative directions.
-
Audio productionIn audio production for industries such as film, games, and advertising, it is used to quickly generate background music or theme songs, improving production efficiency.
-
EducationAs a music education tool, it helps students understand music structure and the creative process, and stimulates their interest in learning.
-
Entertainment industryOn social media, short video and other platforms, personalized music content is generated for users to enhance interactivity and fun.
-
Business applicationsGenerate customized music for businesses and brands for product promotion, event advertising, and other purposes to enhance brand influence.