Jiaojiao - A comprehensive oral dialogue and emotional model launched by Shanghai Jiao Tong University
Jiaojiao is the world's first purely academic, self-developed spoken dialogue emotion model, launched by the Auditory Cognition and Computational Acoustics Laboratory at Shanghai Jiao Tong University. Jiaojiao features multi-person dialogue, multilingual communication, dialect understanding, role-playing, and emotional interaction...
What is Jiaojiao?
Jiaojiao, developed by the Auditory Cognition and Computational Acoustics Laboratory at Shanghai Jiao Tong University, is the world's first purely academic, self-developed spoken dialogue emotion model. Jiaojiao boasts powerful functions including multi-person dialogue, multilingual communication, dialect understanding, role-playing, emotional interaction, and knowledge-based question answering. It supports multiple languages, including Chinese, English, Japanese, and French, and can accurately recognize Chinese dialects. Based on innovative technology, Jiaojiao achieves end-to-end voice dialogue, multilingual understanding, multi-person interaction, and real-time voice cloning. Jiaojiao demonstrates powerful voice interaction capabilities, bringing a new breakthrough to the field of intelligent voice assistants.
The main functions of Jiaojiao
- Multi-person dialogueIt can simultaneously engage in natural and fluent conversations with multiple users, accurately identify each person's identity and content, and provide personalized responses.
- Multilingual communicationIt supports four major languages: Chinese, English, Japanese, and French, and has cross-language response capabilities.
- Role-playing and emotional interactionIt understands user emotions based on dialogue content and context, and generates emotionally resonant responses.
- Knowledge Q&AIt covers a wide range of knowledge areas, such as reciting ancient poems, explaining scientific principles, and interpreting literary masterpieces.
- Real-time timbre cloningIt provides high-fidelity voice imitation technology, supports multiple voice acting styles and real-time seamless switching between the voice and the user's own voice.
The technical principles of intersection
- End-to-end voice dialogueBased on a robust audio encoder, the discrete sequence obtained from the audio input streaming encoder is aligned to the text sequence space. Without the need for large-scale fine-tuning with high-quality data, it can maintain and utilize the basic generalization ability of the large text model to achieve real-time knowledge question answering.
- Multilingual understanding and generationBased on an innovative cross-modal alignment mechanism, it accurately maps multilingual speech signals to corresponding text in the feature space, preserves language-specific information using implicit representation learning, and combines the contextual modeling capabilities of deep language models to achieve seamless switching and efficient semantic understanding in cross-language scenarios.
- Multi-person dialogue modelingConstruct multi-person dialogue data to simulate real-world scenarios and enhance the model's dialogue processing capabilities. Use an end-to-end model to fuse contextual information, generate personalized responses and summaries, and achieve natural and coherent multi-party interactions.
- Emotional understanding and expressionBased on contextual information, the system uses thought chain technology to generate a global emotional representation that fits the dialogue scenario, which is then used to generate vivid emotional voice responses and enhance the realism of the dialogue.
- Real-time timbre cloning and switchingIt provides high-fidelity voice imitation technology, uses thought chain technology for control signal inference, and supports multi-role voice acting styles and real-time seamless switching between the voice and the user's own voice.
- Flexible expansionA powerful alignment strategy supports arbitrary splicing and fusion of text and audio modalities, providing a unified and scalable interface for integrating various enhancement mechanisms (such as web search, RAG retrieval enhancement generation, etc.) in large-scale text models.
Project address of the project
- Application for trial address:https://wj.sjtu.edu.cn/q/4FiP8hsB
Application scenarios of Jiaojiao
- Educational guidanceIt provides students with personalized learning guidance, answers their questions, and assists teachers in their teaching.
- Family Interaction: To entertain and add to the fun at family gatherings, and to keep family members company and chat with them to relieve boredom on a daily basis.
- Business Communication: Assist in meeting minutes and summaries, and support cross-language business communication.
- Customer SupportWe respond quickly to customer inquiries, provide professional answers, and improve service efficiency.
- Entertainment and companionshipEngaging in role-playing provides emotional support and adds fun to life.