AB
AiBoss
project

HumanOmni - Alitongyi and others launch a multimodal large model focusing on human-centric scenarios

HumanOmni is a large, multimodal model focused on human-centric scenarios, fusing visual and auditory modalities. By processing video, audio, or a combination of both inputs, it comprehensively understands human behavior, emotions, and interactions. The model is based on...

What is HumanOmni?

HumanOmni is a large, multimodal model focused on human-centric scenarios, fusing visual and auditory modalities. By processing video, audio, or a combination of both inputs, it comprehensively understands human behavior, emotions, and interactions. The model is pre-trained on over 2.4 million video clips and 14 million instructions, employing a dynamic weight adjustment mechanism to flexibly fuse visual and auditory information according to different scenarios. HumanOmni excels in emotion recognition, facial description, and speech recognition, and is suitable for various scenarios such as film analysis, close-up video interpretation, and live-action video understanding.

HumanOmni's main functions

  • Multimodal fusionHumanOmni can process visual (video), auditory (audio), and text information simultaneously. Through a command-driven dynamic weight adjustment mechanism, it fuses features from different modalities to achieve a comprehensive understanding of complex scenes.
  • Human-centered scene understandingThe model handles face-related, body-related, and interaction-related scenarios through three specialized branches, and adaptively adjusts the weights of each branch according to user instructions to adapt to different task requirements.
  • Emotion recognition and facial expression descriptionHumanOmni excels in dynamic facial emotion recognition and facial expression description tasks, surpassing existing video-language multimodal models.
  • Action understandingThrough body-related branches, the model can effectively understand human movements and is suitable for action recognition and analysis tasks.
  • Speech Recognition and UnderstandingIn speech recognition tasks, HumanOmni achieves efficient speech understanding through audio processing modules (such as Whisper-large-v3) and supports speech recognition for specific speakers.
  • Cross-modal interactionThe model combines visual and auditory information to gain a more comprehensive understanding of the scene, making it suitable for tasks such as movie clip analysis, close-up video interpretation, and live-action video understanding.
  • Flexible fine-tuning supportDevelopers can fine-tune the pre-trained parameters of HumanOmni to suit specific datasets or task requirements.

HumanOmni's technical principles

  • Multimodal fusion architectureHumanOmni achieves a comprehensive understanding of complex scenes by fusing visual, auditory, and textual modalities. In the visual part, the model is designed with three branches: a face-related branch, a body-related branch, and an interaction-related branch, used to capture features of facial expressions, body movements, and environmental interactions, respectively. A command-driven fusion module dynamically adjusts the weights, adaptively selecting the most suitable visual features for the task based on user commands.
  • Dynamic weight adjustment mechanismHumanOmni introduces an instruction-driven feature fusion mechanism. It encodes user commands using BERT to generate weights, dynamically adjusting the feature weights for different branches. In emotion recognition tasks, the model focuses more on features related to facial features; in interactive scenarios, it prioritizes interaction-related branches.
  • Coordinated processing of hearing and visionIn terms of auditory processing, HumanOmni uses the Whisper-large-v3 audio preprocessor and encoder to process audio data, mapping it to the text domain via MLP2xGeLU. Visual and auditory features are combined in a unified representation space and further fed into the decoder of the large language model for processing.
  • Multi-stage training strategyHumanOmni training is divided into three stages:
    • The first phase involves building visual capabilities and updating the parameters of the visual mapper and instruction fusion module.
    • The second stage develops auditory abilities by updating only the parameters of the audio mapper.
    • The third stage involves cross-modal interaction integration to enhance the model's ability to process multimodal information.
  • Data-driven optimizationHumanOmni is pre-trained on over 2.4 million human-centric video clips and 14 million instruction data. The data covers multiple tasks, including emotion recognition, facial description, and speaker-specific speech recognition, and the model performs excellently in various scenarios.

HumanOmni's project address

HumanOmni Application Scenarios

  • Film and EntertainmentHumanOmni can be used in film and television production, such as virtual character animation generation, virtual anchors, and music video creation.
  • Education and TrainingIn the education field, HumanOmni can create virtual teachers or simulated training videos to assist language learning and vocational skills training.
  • Advertising and MarketingHumanOmni can generate personalized ads and brand promotion videos by analyzing people's emotions and actions to provide more engaging content and increase user engagement.
  • Social media and content creationHumanOmni helps creators quickly generate high-quality short videos, supports interactive video creation, and increases the fun and appeal of content.