Chinese-LiPS - A Chinese multimodal speech recognition dataset jointly released by the Beijing Academy of Artificial Intelligence and Nanjing University.
Chinese-LiPS is a high-quality Chinese multimodal speech recognition dataset jointly developed by the Beijing Academy of Artificial Intelligence and Nankai University. It contains 100 hours of audio, video, and manually transcribed text, innovatively integrating lip-reading videos and speeches...
What is Chinese-LiPS?
Chinese-LiPS is a high-quality Chinese multimodal speech recognition dataset jointly developed by the Beijing Academy of Artificial Intelligence (BAAI) and Nankai University. It contains 100 hours of audio, video, and manually transcribed text, innovatively integrating lip-reading videos and speaker slides. The slides were meticulously designed by domain experts to ensure high-quality and rich visual images. By combining lip-reading and slide information, the dataset improves speech recognition performance. Experiments show that lip-reading information and slide information can improve ASR performance by approximately 8% and 25%, respectively, with the combination achieving an improvement of approximately 35%. It is designed for complex contexts such as Chinese explanations, popular science, teaching, and knowledge dissemination.
Main functions of Chinese-LiPS
- Improve speech recognition performanceThe dataset significantly improved the performance of the speech recognition system by fusing lip-reading information and slide semantic information. Experimental results show that lip-reading information can reduce the character error rate by about 8%, slide information by about 25%, and the combination of the two can reduce it by about 35%.
- Reduce error typesLip reading information plays a crucial role in reducing deletion errors, capturing pronunciation-related details and effectively supplementing easily missing parts in speech recognition, such as filler words and speech fragments not fully expressed due to hesitation. Slide information significantly reduces substitution errors; its rich semantic and contextual information provides key recognition clues for the model when recognizing domain-specific terms such as professional vocabulary and place names.
- Provide high-quality multimodal dataAs a high-quality multimodal Chinese speech recognition dataset, it contains 100 hours of audio, video, and corresponding manual transcriptions, covering lip-reading videos and speaker slides, enabling a more comprehensive exploration of audio-visual speech recognition tasks.
The technical principle of Chinese-LiPS
- Multimodal data fusionThe dataset integrates speech, lip-reading information, text extracted from slides using OCR technology, and semantic information from images and graphics. This combination of multimodal information provides the speech recognition model with richer context and cues, significantly improving the accuracy and robustness of recognition.
- The role of lip reading informationLip reading can capture details related to pronunciation, such as filler words and speech fragments that are not fully expressed due to hesitation, which are easily lost in speech recognition. With the help of lip reading information, these parts can be effectively supplemented, reducing deletion errors.
- The role of slide informationThe slides contain rich semantic and contextual information, which can provide the model with key recognition clues when facing the recognition of words with specific domain attributes such as professional terms and place names, and greatly reduce replacement errors.
Chinese-LiPS project address
- Project official website:https://data.baai.ac.cn/datadetail/Chinese-LiPS
- Github repository:https://github.com/flageval-baai/Chinese-LiPS
- HuggingFace model library:https://huggingface.co/datasets/BAAI/Chinese-LiPS
- arXiv technical paper:https://arxiv.org/pdf/2504.15066
Application scenarios of Chinese-LiPS
- Virtual TeacherData sets can help create interactive language learning materials, making virtual teachers' explanations more engaging. By integrating lip-reading information and semantic information from slides, virtual teachers can present teaching content more naturally, improving teaching effectiveness.
- Intelligent tutoringIn intelligent tutoring systems, multimodal speech recognition technology can more accurately understand students' questions and needs, and provide more personalized tutoring solutions.
- Museum and exhibition hall guided toursIn museums, exhibition halls, and other venues, virtual guides can use multimodal information provided by datasets to introduce exhibits and exhibition content more vividly and accurately, enhancing the visitor experience.
- Company Product IntroductionBusinesses can use datasets to create virtual presenters for product introductions, training sessions, and other scenarios, improving the efficiency and accuracy of information delivery.