Qinyue Big Model - Tencent's AI Music Creation Big Model
The QinYue Big Model is an AI-powered music creation model jointly developed by Tencent AI Lab and Tencent TME Tianqin Lab. This model can directly generate stereo audio or... by inputting Chinese or English keywords, descriptive phrases, or audio.
What is the Qinle Big Model?
Qinyue Big Model is an AI-powered music creation model jointly developed by Tencent AI Lab and Tencent TME Tianqin Lab. This model can directly generate stereo audio or multi-track scores by inputting Chinese and English keywords, descriptive phrases, or audio. Qinyue Big Model supports automatic editing, such as continuing or regenerating specified tracks or measures, as well as modifying instrument types and rhythms. Currently, the Qinyue Big Model technology is available on Tencent Music's Qimingxing platform, where users can register for free to experience it. In the future, the research team plans to add the ability to generate vocals, lyrics, and other elements to the model to better serve the needs of music creation.
Features of the Qinle Large Model
- Music generationThe model can intelligently generate music based on user-provided keywords, descriptive phrases, or audio input. This generation is not only based on text descriptions but also understands audio content, enabling automatic music creation.
- Music score generationIn addition to generating audio, "Qinyue Big Model" can also generate detailed scores, which include multiple tracks such as melody, chords, accompaniment and percussion, providing users with rich musical structures.
- Automatic editingThe model supports a series of automatic editing operations on the generated score, including but not limited to continuing the score, regenerating specific tracks or measures, adjusting the orchestration, and modifying the instrument type and rhythm, which greatly improves the flexibility and efficiency of creation.
- Audio text alignmentBy using contrastive learning techniques, the model constructs a shared feature space, aligning audio tags or text descriptions with the audio itself, providing conditional control signals for the generative model, and enhancing the relevance and accuracy of music generation.
- Music score/audio representation extractionThe model can transform musical scores or audio into a series of discrete feature (token) sequences, which provide the basis for predictions by large language models.
- Large language model predictionUsing a decoder-only structure, the model is trained through feature prediction (next token prediction), and the predicted sequence can be converted back into sheet music or audio, realizing the conversion from text to music.
- Audio recoveryBy using stream matching and vocoder techniques, the model can recover audible audio from the predicted audio representation sequence, enhancing the realism and quality of the audio.
- Music theory followsIn the process of generating music, the "Qinyue Big Model" follows music theory to ensure that elements such as melody, chords, and rhythm conform to musical logic and human aesthetics.
How to experience and use the Qinyue large model
- Registration and LoginAccess Tencent Music Rising Star Platform (https://y.qq.com/venus/#/venus/aigc/ai_compose( ), and register an account or log in using an existing account.
- Input creation conditionsOn the experience page, enter music keywords, phrases, or descriptions, which will serve as the basis for the model to generate music.
- Select music modelCurrently, only the Qin music generation large model v1.0 is available for selection.
- Select music durationYou can choose the music duration from 10 seconds to 30 seconds.
- Generate musicClick "Start Generation," and wait approximately one minute for the music to be generated. The generated music can then be played and downloaded.
Technical principles of the large-scale piano model
- Audio text alignment modelThis module uses contrastive learning to construct a shared feature space between audio labels or text descriptions and the audio. In this way, the model can understand the semantic relationships between the text and audio and use this information as conditional control signals during the generation process.
- Music score/audio representation extractionThe model converts musical scores or audio into discrete sequences of features, which can be representations of MIDI attributes or encoded and compressed representations of pre-trained audio spectra.
- Large Language ModelThis method uses a large language model with a decoder-only structure for training in next token prediction. This model can predict the next feature based on the input feature sequence, thus generating continuous musical elements.
- Stream matching and vocoder technologyDuring the audio generation process, the model uses stream matching technology and a vocoder module to convert the predicted audio representation sequence into audible audio, enhancing the realism of the audio.
- Multi-module collaborative workThe "Music Generation Model" comprises multiple modules that work together to achieve the desired music generation effect. For example, the audio-text alignment model provides conditional control signals during training, while using text representations as control signals during inference.
- Music theory followsIn the process of generating music, the model needs to follow music theory, including the rationality of elements such as melody, chords, and rhythm, to ensure that the generated music conforms to human auditory habits and aesthetic standards.
- Automatic editing and adjustmentThe model supports automatic editing of the generated score, such as continuing the song, regenerating a specified track or measure, and modifying the instrument type and rhythm, which makes the music creation process more flexible.
- End-to-end generation processFrom text input to audio output, the "Qinyue Big Model" realizes an end-to-end generation process, reducing manual intervention and improving the efficiency of music creation.
- Large-scale double-blind audiometryThe model's generation quality was validated through large-scale double-blind listening tests, and its multi-dimensional subjective scores surpassed industry standards.