Tonggu Large-Scale Model - A Large-Scale Language Model for Ancient Books Developed by South China University of Technology
Tonggu Model is an AI language model developed by the Deep Learning and Visual Computing Laboratory (SCUT-DLVCLab) at South China University of Technology, focusing on the processing of classical Chinese texts. It is incrementally pre-trained based on Baichuan 2-7B-Base and uses...
What is the Tonggu Grand Model?
Tonggu Model is an AI language model developed by the Deep Learning and Visual Computing Laboratory (SCUT-DLVCLab) at South China University of Technology, focusing on the processing of classical Chinese texts. It is incrementally pre-trained on the Baichuan 2-7B-Base platform, unsupervised training using a corpus of 2.41 billion classical Chinese texts, and fine-tuned using 4 million dialogue data points from classical Chinese texts. The model employs Redundancy Aware Tuning (RAT) technology, effectively improving performance on classical Chinese text tasks. It helps users understand and translate classical Chinese documents more easily. Through Retrieval-Enhanced Generation (CCU-RAG) technology, it reduces the illusion problem in knowledge-intensive tasks, improving the accuracy and reliability of generated content.
The main functions of the Tonggu Grand Model
- Classical Chinese punctuationThe Tonggu model can automatically add punctuation marks to classical Chinese texts, solving common punctuation problems in ancient books and helping users better understand the content.
- Classical and vernacular translationThe model supports bidirectional translation between classical Chinese and modern Chinese, translating obscure ancient texts into modern Chinese, and vice versa, making it convenient for users to read and study ancient books.
- Poetry CreationThe Tonggu model can generate poems that conform to the rules and styles of classical Chinese poetry. Users can provide themes or keywords according to their needs, and the model will generate corresponding poems.
- Appreciation of Ancient BooksThe model can appreciate classic chapters in ancient books, interpret their literary value, historical background and cultural connotations, and help users to study ancient books in depth.
- Ancient Books Search and Q&ABy combining search enhancement technology, the Tonggu Big Data Model can quickly retrieve ancient book content and provide accurate answers to user questions, helping users efficiently obtain ancient book information.
- Assisting in the collation of ancient booksThe model can identify textual errors, omissions, and other problems in ancient books, and provide repair suggestions to assist in the collation and digitization of ancient books.
Technical principles of the Tonggu large model
- Basic model architectureThe Tonggu Big Model is incrementally pre-trained based on Baichuan 2-7B-Base. Baichuan 2-7B-Base is a powerful pre-trained language model that provides Tonggu Big Model with basic language understanding and generation capabilities.
- Unsupervised incremental pre-trainingThe model was unsupervised incremental pre-training on a corpus of 2.41 billion ancient Chinese texts. This enabled the model to learn the language style and structure of the ancient texts, laying the foundation for subsequent ancient text processing tasks.
- Multi-stage instruction fine-tuningThe Tonggu model employs a multi-stage instruction fine-tuning technique and proposes a redundancy-aware fine-tuning (RAT) method. This improves the performance of downstream tasks while preserving the capabilities of the base model. Through instruction fine-tuning, the model can better adapt to specific tasks in ancient text processing, such as classical Chinese translation and punctuation.
- Search Enhancement Generation (RAG) technologyTonggu's large model combines retrieval-enhanced generation (RAG) technology to reduce the illusion problem in knowledge-intensive tasks. The core idea is to combine information retrieval with text generation, retrieving relevant information from external knowledge bases and feeding it as contextual input to the language model to generate more accurate and context-appropriate answers.
Project address of Tonggu Large Model
- Github repository:https://github.com/SCUT-DLVCLab/TongGu-LLM
- HuggingFace model library:https://huggingface.co/SCUT-DLVCLab/TongGu-7B-Instruct
Application scenarios of Tonggu large model
- Ancient Book Processing and DigitizationThe Tonggu Big Data Model can efficiently process ancient texts and documents, supporting functions such as classical and vernacular Chinese translation, punctuation correction, and ancient text retrieval. It assists in the organization of ancient texts by intelligently identifying and correcting textual errors, thereby improving the efficiency of ancient text digitization.
- Educational supportTeachers can use it to generate lesson plans, teaching PPTs, and design interactive classroom activities. For students, the model provides functions such as classical Chinese translation, idiom explanation, and poetry creation to help them better understand classical Chinese texts.
- Cultural inheritance and popularizationThe Tonggu Big Model reduces the difficulty of reading ancient books, allowing more people to access and understand traditional Chinese culture.
- academic researchThe Tonggu Big Data Model provides powerful technical support for the study of ancient books, enabling scholars to quickly retrieve and analyze the content of ancient books.