EMOVA - A multimodal all-around processing model jointly developed by Huawei Noah's Ark and several universities.
EMOVA (Emotionally Omni-present Voice Assistant) is a multimodal, all-around model jointly developed by the Hong Kong University of Science and Technology, the University of Hong Kong, and Huawei Noah's Ark Lab, among other institutions. EMOVA can process images,...
What is EMOVA?
EMOVA (Emotionally Omni-present Voice Assistant) is a multimodal, all-around model jointly developed by the Hong Kong University of Science and Technology, the University of Hong Kong, and Huawei Noah's Ark Lab. EMOVA can process image, text, and speech modalities, enabling full-modal interaction that allows it to see, hear, and speak. Based on semantic acoustic separation technology and a lightweight emotion control module, EMOVA supports emotionally rich voice dialogue, making human-computer interaction more natural and human-like. EMOVA demonstrates superior performance in both visual language and speech tasks, providing new implementation ideas for the AI field and driving the development of emotional interaction.
EMOVA's main functions
- Multimodal processing capabilityIt can process data in three modalities simultaneously: image, text, and voice, enabling multimodal interaction.
- Emotionally rich dialogueBased on semantic acoustic separation technology and emotion control module, it can generate voice output with emotional color, such as happiness and sadness.
- End-to-end voice dialogueThe model supports a complete dialogue process from voice input to voice output without relying on external voice processing tools.
- Visual Language UnderstandingIt understands and generates text related to image content, maintaining leading performance in visual language understanding.
- Speech understanding and generationThe model can understand and generate speech, achieving speech recognition and speech synthesis.
- Personalized voice generationIt supports control over the style, emotion, speed, and tone of the voice to adapt to different communication scenarios and user needs.
EMOVA's technical principles
- Continuous vision encoder: Capture the fine visual features of an image using a continuous visual encoder and encode them into a vector representation that can be aligned with the text embedding space.
- Semantic-acoustic separation speech segmenterThe input speech is decomposed into two parts: semantic content and acoustic style. The semantic content is quantized into discrete units and aligned with the language model, while the acoustic style controls emotion and tone.
- Lightweight style moduleIntroducing a lightweight style module to control the emotion and tone of voice output, making voice dialogue more natural and expressive.
- Full-modal alignmentUsing text as a bridge, we perform full-modal training based on publicly available image-text and speech-text data to achieve effective alignment between different modalities.
- End-to-end architectureIt adopts an end-to-end architecture to directly generate text and speech output from multimodal input, realizing a direct mapping from input to output.
- A data-efficient full-modal alignment methodThe goal is to enhance full-modal capabilities based on bimodal data, avoid dependence on scarce trimodal data, and enhance cross-modal capabilities through joint optimization.
EMOVA's project address
- Project official website:emova-ollm.github.io
- arXiv technical paper:https://arxiv.org/pdf/2409.18042
Application scenarios of EMOVA
- Customer ServiceIn the field of customer service, chatbots interact with customers using voice, text, and images to provide emotional service and support.
- Educational SupportIn the field of education, virtual teachers provide personalized teaching and learning experiences through multimodal interaction of images, text, and voice.
- Smart Home ControlIn a smart home system, it acts as a central control system, using voice commands to control devices in the home and providing visual feedback.
- Health ConsultationIn the healthcare field, we provide voice-interactive health consultation services, offering corresponding health advice based on analysis of users' questions and needs.
- Emergency RescueIn emergency situations, voice recognition and image analysis can be used to quickly understand the situation on-site and provide rescue guidance.