GTSinger - A large-scale, multilingual, high-quality singing dataset open-sourced by Zhejiang University.
GTSinger is a large-scale, open-source, high-quality singing dataset developed by a research team at Zhejiang University, designed to support diverse singing tasks. GTSinger contains 80.59 hours of professionally recorded vocals, covering nine different languages...
What is GTSinger?
GTSinger is a large-scale, open-source, high-quality singing dataset developed by a research team at Zhejiang University, designed to support diverse singing tasks. GTSinger contains 80.59 hours of professionally recorded studio vocals in nine different languages (Chinese, English, Japanese, Korean, Russian, Spanish, French, German, and Italian), sung by 20 professional singers, offering rich timbre and stylistic diversity. GTSinger emphasizes the control and modeling of singing techniques, providing control groups and phoneme-level annotations for six commonly used singing techniques. GTSinger provides authentic musical scores, aiding in actual music composition. The dataset includes artificial phoneme alignment, global style labels, and paired reading data, adapting to various singing tasks.
GTSinger's main functions
- Multilingual singing datasetGTSinger includes vocals in nine different languages, offering a diverse range of timbres and styles, and supports cross-language vocal synthesis and analysis.
- Singing technique controlThe dataset provides control groups and phoneme-level annotations for six commonly used singing techniques, allowing researchers to better model and control techniques in singing.
- Real sheet music supportProviding accurate sheet music that matches the vocals is very helpful in applying vocal synthesis technology to actual music creation.
- Multi-tasking adaptationGTSinger is designed to support a variety of singing tasks, including singing synthesis, skill recognition, style transfer, and speech-to-singing conversion.
- BenchmarkingProvides benchmark tests to evaluate the performance and applicability of the dataset on different singing tasks.
GTSinger's technical principles
- High-quality audio recordingThe GTSinger dataset is constructed from recordings of professional singers in professional recording studios, ensuring high-quality audio data.
- Phoneme alignment and annotationBased on music information retrieval technologies, such as MFA and Praat, phoneme alignment and annotation are performed to achieve precise control at the phoneme level.
- Singing TechniquesBased on expert listening and audio analysis technology, the singing techniques in the singing are labeled to facilitate model learning and control.
- Music score generationCombining audio signal processing technology and music theory, pitch information is extracted from singing, converted into MIDI sheet music, and then adjusted by experts into actual sheet music.
- Dataset construction and validationBased on manual review and subsequent processing, the quality and applicability of the dataset are ensured, including semantic segmentation of audio segments and processing of silent regions.
GTSinger's project address
- Project official websitegtsinger.github.io
- GitHub repository:https://github.com/GTSinger/GTSinger
- HuggingFace model library:https://huggingface.co/datasets/GTSinger/GTSinger
- arXiv technical paper:https://arxiv.org/pdf/2409.13832
Application scenarios of GTSinger
- Singing synthesisBased on vocal samples and technique annotations in the dataset, a system was developed to synthesize high-quality vocals with specific techniques and styles.
- Vocal Technique Recognition: Analyze the phoneme-level technique annotations in singing, and train a model to recognize and classify different singing techniques.
- Singing style transferTransforming a song into a different style, such as converting a pop song into a classical style.
- Voice-to-singing conversion(Speech-to-Singing, STS): Converts ordinary speech into melodic singing, used in speech synthesis and music composition.
- Music EducationBased on real musical scores and singing samples in the dataset, we develop music teaching tools to help students learn and practice singing skills.