SenseNova V6 - A multimodal fusion model series launched by SenseTime
SenseNova V6 is SenseTime's sixth-generation multimodal fusion model series, based on a 600 billion parameter multimodal MoE architecture, enabling native fusion of text, images, and videos. SenseNova V6...
What is the Ririshin SenseNova V6?
SenseNova V6 is the sixth generation of the SenseTime multimodal fusion model series. Based on a 600 billion parameter multimodal MoE architecture, it achieves native fusion of text, images, and videos. SenseNova V6 performs exceptionally well in both pure text and multimodal tasks, outperforming models such as GPT-4.5 and Gemini 2.0 Pro in multiple metrics.
Ririsun's SenseNova V6 includes four versions: SenseNova V6 Pro, a hybrid expert architecture model with 620 billion parameters, supports native fusion of text, images, and video, and is comparable to mainstream international models; SenseNova V6 Reasoner Pro possesses reasoning capabilities to assist in solving complex problems; SenseNova V6 Video specializes in video understanding and is suitable for scenarios such as education and cultural tourism; SenseNova V6 Omni is a lightweight, multimodal interactive model that provides a real-time interactive experience. Ririsun's SenseNova V6 features strong reasoning, strong interactivity, and long memory, performing reasoning and analysis on medium-to-long videos, accurately answering questions in real-time audio and video interactions, and providing emotional expression. The model is applied in fields such as education and embodied intelligence, providing robots with a brain, eyes, ears, and a mouth.
Main functions of Ririshin SenseNova V6
- Video processing and analysisSupports reasoning and parsing of medium- to long-form videos.
- Real-time audio and video interaction: Provide accurate answers to questions about the video content, such as character relationships and plot development.
- Educational guidanceIt recognizes handwriting and provides one-on-one guided explanations to help children with math problems.
- Emotional understanding and expressionIt possesses highly human-like perception, expression, and emotional understanding abilities, and can switch tone, emotion, and pitch according to different dialogue content and scene requirements.
- Embodied IntelligenceTo enable robots to have stronger perception and interaction capabilities.
The technical principles of RiRiXin SenseNova V6
- Native multimodal fusion training technologyIt deeply integrates multiple modal information such as text, images, videos, and audio during model architecture and training, avoiding the problem in traditional methods where enhancing one modality's capabilities leads to a decline in another, and better handles complex scenes and captures detailed cross-modal relationships.
- Multimodal long thought chain synthesis technologyBased on multi-agent collaboration, it enables the generation and verification of ultra-long thought chains, allowing the model to have the ability to think deeply for a long time and in multiple steps, and is suitable for scenarios such as mathematical derivation, scientific analysis, and long document comprehension.
- Multimodal Hybrid Reinforcement LearningThe model balances logical reasoning and emotional expression capabilities by combining human preference-based RLHF and deterministic answer-based RFT, ensuring that the model can naturally express emotions while improving its reasoning ability.
- Unified representation and dynamic compression of long videosIt achieves efficient alignment and compression of cross-modal information, unifies the encoding of images, audio, subtitles, and time logic to form a coherent temporal representation, and significantly improves processing efficiency.
Project address for RiRiXin SenseNova V6
- Project official website:https://platform.sensenova.cn
Application scenarios of Ririshin SenseNova V6
- Video Creation and AnalysisQuickly generate video highlights, edit specific scenes, and add narration and sound effects.
- Educational guidanceTutoring in math problems, providing one-on-one explanations to help students understand problem-solving strategies.
- Intelligent Customer ServiceIt provides accurate answers to user questions, offers personalized suggestions, and enhances the user experience.
- Embodied IntelligenceIt provides robots with perception and interaction capabilities, and can be applied in scenarios such as home, industry, and healthcare.
- Content RecommendationRecommend personalized videos, articles, music, and other content based on user preferences.