AB
AiBoss
project

Daily New Fusion Model - SenseTime's Native Fusion Modal Model

SenseNova, a multimodal large-scale model, was officially launched by SenseTime on January 10, 2025. The model achieves native fusion modality, significantly improving both deep reasoning and multimodal information processing capabilities...

What is the RiRiXin Fusion Grand Model?

SenseNova, a multimodal large-scale model, was officially launched by SenseTime on January 10, 2025. The model achieves native modality fusion, significantly improving both deep reasoning and multimodal information processing capabilities. It can handle various types of information, including text, images, and videos, breaking through the limitations between modalities. It achieved first place on both the authoritative SuperCLUE and OpenCompass benchmark lists, becoming a "double champion."

The main functions of the Rixin Fusion Model

  • Image recognition and analysisIt can accurately identify and analyze the content in images, including blurry text and complex scenes.
  • Video processingIt can process video content, extract key information, perform video editing and generation, and improve the video interactive experience.
  • Speech recognition and synthesisCombining voice and natural language processing capabilities enhances the interactive experience, such as in scenarios like voice customer service and online education.
  • Text processingIt possesses powerful text understanding and generation capabilities, and can handle complex rich-modal documents, such as documents that combine tables, text, images, and videos.
  • Mathematical calculation and logical reasoningIt can solve complex mathematical problems, such as calculating whether 2 to the power of 31 or 3 to the power of 21 is larger, using the logarithmic function method to solve the problem.
  • Data Analysis and Decision SupportIt can analyze information in data charts, extract key elements, draw conclusions, and provide decision support for users.

The technical principles of the daily fusion model

  • Native fusion modeThe model can process multiple types of information such as text, images, and videos simultaneously, breaking through the limitation of traditional large language models that only support single text input.
  • Fusion modal data synthesis:
    • Reverse rendering technologyBy using inverse rendering technology, image and text data are fused to generate a large amount of synthetic data. This synthetic data establishes numerous interactive bridges between the image and text modalities, enabling the model to more solidly grasp the rich relationships between the modalities.
    • Image generation based on hybrid semanticsBy utilizing hybrid semantic generation technology, the fused modal data was further enriched, and the model's ability to understand multimodal information was improved.
  • Integrated Task Enhancement TrainingA rich set of cross-modal tasks has been constructed, providing a solid foundation for model training. These tasks not only include traditional text processing tasks, but also cover multimodal tasks such as image recognition and video analysis, enabling the model to effectively respond to user needs in various business scenarios.
  • Deep reasoning ability:
    • A balanced education in both arts and sciencesIn the SuperCLUE annual assessment, the humanities score ranked first in the world with 81.8 points, and the science score won the gold medal, with the calculation dimension ranking first in China with 78.2 points.
    • Complex Problem SolvingIt can handle complex, rich-modal documents, such as documents that combine tables, text, images, and videos, and provides in-depth reasoning support.

The project address of the RiRiXin Fusion Large Model

Application scenarios of the daily fusion model

  • autonomous drivingIt can process complex multimodal information and improve decision-making capabilities.
  • Video InteractionImprove the efficiency of video content generation, editing, and analysis.
  • Office EducationIt efficiently processes complex, rich-modal documents, improving office and educational efficiency.
  • finance: Analyze and process heterogeneous data from multiple sources to provide accurate risk assessments and investment advice.
  • Park ManagementTo improve the management efficiency and security of the park.
  • Industrial manufacturingOptimize production processes and quality control.