NEXUS-O - A multimodal AI model that enables comprehensive perception and interaction across language, audio, and vision.
NEXUS-O is a multimodal AI model developed by HiThink Research Institute, Imperial College London, Zhejiang University, Fudan University, Microsoft, Meta AI, and other institutions. It enables comprehensive perception and interaction of language, audio, and visual information...
What is NEXUS-O?
NEXUS-O is a multimodal AI model developed by HiThink Research Institute, Imperial College London, Zhejiang University, Fudan University, Microsoft, Meta AI, and other institutions. It enables comprehensive perception and interaction of language, audio, and visual information. NEXUS-O can process any combination of audio, image, video, and text inputs, outputting results in audio or text form. Based on a visual language model pre-trained, NEXUS-O uses high-quality synthetic audio data to improve its trimodal alignment capabilities. NEXUS-O introduces a new audio testing platform, Nexus-O-audio, covering various real-world scenarios (such as meetings and live streaming) to evaluate the model's robustness in practical applications. NEXUS-O performs exceptionally well in tasks such as visual understanding, audio question answering, speech recognition, and speech translation, demonstrating efficiency and effectiveness based on trimodal alignment analysis.
Main functions of NEXUS-O
- Voice processing capabilitiesIt supports tasks such as automatic speech recognition (ASR), speech-to-text translation (S2TT), speech synthesis, and voice command interaction, and is suitable for a variety of voice application scenarios.
- Visual understanding and interactionIt processes image and video inputs, performs tasks such as visual question answering (VQA), image description generation, and video analysis, and has strong visual understanding capabilities.
- Language Interaction and ReasoningIt can understand natural language instructions, perform tasks such as dialogue interaction, text generation, and multimodal reasoning, and support complex language interaction scenarios.
- Cross-modal alignment and understandingBased on multimodal alignment technology, it achieves collaborative understanding between audio, visual and language modalities, improving the overall performance of the model in complex scenarios.
NEXUS-O's technical principles
- Multimodal architecture:
- Visual encoderBased on the improved Vision Transformer (ViT) architecture, it supports high-resolution image input and improves computational efficiency with a window attention mechanism.
- Audio encoders and decodersThe audio encoder maps speech features to semantic space based on a pre-trained Whisper-large-v3 model; the audio decoder generates discrete speech codes using autoregression and synthesizes the final speech waveform from the pre-trained generator.
- Language ModelBased on Qwen2.5-VL-7B, it contains 28 layers of causal Transformers and is responsible for handling language modalities.
- Multimodal alignment and pre-trainingBased on the pre-training phase, features from audio, visual, and linguistic modalities are aligned into a unified semantic space, enabling the model to understand and generate cross-modal information. A phased pre-training approach, including audio alignment, audio instruction following (SFT), and audio output tuning, progressively improves the model's multimodal interaction capabilities.
- Data Composition and AugmentationText-to-speech (TTS) technology is used to convert text data into natural speech, enhancing data diversity. Synthetic data is filtered for length, non-text elements, and pattern matching to ensure data quality.
- Joint training of multimodal tasksNexus-O supports a variety of multimodal tasks during the pre-training phase, such as automatic speech recognition, speech-to-text translation, voice command interaction, and visual question answering, and joint training improves the model's generalization ability.
- Spatial alignment analysisMethods such as kernel alignment are used to evaluate the degree of alignment of different modalities in the representation space within the model, and to optimize the multimodal feature fusion effect.
NEXUS-O project address
- arXiv technical paper:https://arxiv.org/pdf/2503.01879
Application scenarios of NEXUS-O
- Intelligent voice interactionAs the core of a voice assistant, it supports multilingual dialogue, voice control of devices, and real-time translation, and is widely used in smart homes, in-vehicle systems, and intelligent customer service.
- Video conferencing and collaborationIt offers real-time voice translation, intelligent meeting recording, and virtual assistant features to facilitate efficient remote work and multilingual meetings.
- Education and Content CreationIt assists in language learning, intelligent tutoring, and educational game development, supports video subtitle generation, audio content creation, and multimodal content recommendation, and enhances the learning and creation experience.
- Intelligent driving and securityBased on voice control of vehicle functions, environmental perception assistance, smart home control, and security monitoring, it enhances driving safety and convenience of life.
- Public services and healthcareIt supports intelligent navigation, emergency response assistance, voice diagnosis assistance, and rehabilitation training guidance, contributing to the intelligentization of public services and personalized services in the field of healthcare.