AB
AiBoss
project

Jiaojiao - A comprehensive oral dialogue and emotional model launched by Shanghai Jiao Tong University

Jiaojiao is the world's first purely academic, self-developed spoken dialogue emotion model, launched by the Auditory Cognition and Computational Acoustics Laboratory at Shanghai Jiao Tong University. Jiaojiao features multi-person dialogue, multilingual communication, dialect understanding, role-playing, and emotional interaction...

What is Jiaojiao?

Jiaojiao, developed by the Auditory Cognition and Computational Acoustics Laboratory at Shanghai Jiao Tong University, is the world's first purely academic, self-developed spoken dialogue emotion model. Jiaojiao boasts powerful functions including multi-person dialogue, multilingual communication, dialect understanding, role-playing, emotional interaction, and knowledge-based question answering. It supports multiple languages, including Chinese, English, Japanese, and French, and can accurately recognize Chinese dialects. Based on innovative technology, Jiaojiao achieves end-to-end voice dialogue, multilingual understanding, multi-person interaction, and real-time voice cloning. Jiaojiao demonstrates powerful voice interaction capabilities, bringing a new breakthrough to the field of intelligent voice assistants.

The main functions of Jiaojiao

  • Multi-person dialogueIt can simultaneously engage in natural and fluent conversations with multiple users, accurately identify each person's identity and content, and provide personalized responses.
  • Multilingual communicationIt supports four major languages: Chinese, English, Japanese, and French, and has cross-language response capabilities.
  • Role-playing and emotional interactionIt understands user emotions based on dialogue content and context, and generates emotionally resonant responses.
  • Knowledge Q&AIt covers a wide range of knowledge areas, such as reciting ancient poems, explaining scientific principles, and interpreting literary masterpieces.
  • Real-time timbre cloningIt provides high-fidelity voice imitation technology, supports multiple voice acting styles and real-time seamless switching between the voice and the user's own voice.

The technical principles of intersection

  • End-to-end voice dialogueBased on a robust audio encoder, the discrete sequence obtained from the audio input streaming encoder is aligned to the text sequence space. Without the need for large-scale fine-tuning with high-quality data, it can maintain and utilize the basic generalization ability of the large text model to achieve real-time knowledge question answering.
  • Multilingual understanding and generationBased on an innovative cross-modal alignment mechanism, it accurately maps multilingual speech signals to corresponding text in the feature space, preserves language-specific information using implicit representation learning, and combines the contextual modeling capabilities of deep language models to achieve seamless switching and efficient semantic understanding in cross-language scenarios.
  • Multi-person dialogue modelingConstruct multi-person dialogue data to simulate real-world scenarios and enhance the model's dialogue processing capabilities. Use an end-to-end model to fuse contextual information, generate personalized responses and summaries, and achieve natural and coherent multi-party interactions.
  • Emotional understanding and expressionBased on contextual information, the system uses thought chain technology to generate a global emotional representation that fits the dialogue scenario, which is then used to generate vivid emotional voice responses and enhance the realism of the dialogue.
  • Real-time timbre cloning and switchingIt provides high-fidelity voice imitation technology, uses thought chain technology for control signal inference, and supports multi-role voice acting styles and real-time seamless switching between the voice and the user's own voice.
  • Flexible expansionA powerful alignment strategy supports arbitrary splicing and fusion of text and audio modalities, providing a unified and scalable interface for integrating various enhancement mechanisms (such as web search, RAG retrieval enhancement generation, etc.) in large-scale text models.

Project address of the project

Application scenarios of Jiaojiao

  • Educational guidanceIt provides students with personalized learning guidance, answers their questions, and assists teachers in their teaching.
  • Family Interaction: To entertain and add to the fun at family gatherings, and to keep family members company and chat with them to relieve boredom on a daily basis.
  • Business Communication: Assist in meeting minutes and summaries, and support cross-language business communication.
  • Customer SupportWe respond quickly to customer inquiries, provide professional answers, and improve service efficiency.
  • Entertainment and companionshipEngaging in role-playing provides emotional support and adds fun to life.