AB
AiBoss
project

Daily Update V6.5 - A large-scale multimodal inference model launched by SenseTime.

RiRiXin V6.5 is a new multimodal reasoning model launched by SenseTime. The model features an original image-text interwoven thought chain, in which images participate in reasoning in the form of ontologies, significantly improving cross-modal reasoning accuracy and surpassing Gemini 2.5 Pro.

What is Daily Fresh V6.5?

RiRiXin V6.5 is a new multimodal inference model launched by SenseTime. The model features a unique image-text interleaved thought chain, where images participate in inference in the form of ontologies, significantly improving cross-modal inference accuracy and surpassing Gemini 2.5 Pro. Compared to RiRiXin 6.0, inference capability is improved by 6.99%, while inference cost is only 30%, resulting in a 5x improvement in cost-effectiveness. Based on the lightweight Vision Encoder+ and a deep LLM architecture, the model possesses efficient inference capabilities and can be widely applied in embodied intelligence scenarios such as autonomous driving and robotics.

Main features of Daily New V6.5

  • Multimodal reasoningIt supports processing mixed input of images and text, enabling complex reasoning tasks such as understanding image content and combining it with text information to generate accurate descriptions or answers to related questions.
  • High-efficiency reasoning abilityIt performs exceptionally well on multiple datasets, significantly improving inference accuracy, greatly reducing inference costs, and improving cost-effectiveness by 5 times.

The technical principles of Rixin V6.5

  • Interwoven text and image thought chainImages participate in the reasoning process in the form of ontology, and the mixed image and text thinking mode enables the model to more accurately understand and process multimodal information.
  • Lightweight Vision Encoder+Based on the optimized visual encoder, image processing efficiency is improved while reducing computational resource consumption.
  • In-depth LLM architectureIt combines the powerful language understanding and generation capabilities of deep language models (LLM) to achieve efficient cross-modal reasoning.
  • Multimodal collaborative trainingBy processing both image and text data simultaneously, the model can learn richer semantic information and improve inference accuracy.

The project address for RiRiXin V6.5

  • Project official websitehttps://platform.sensenova.cn/

Application scenarios of Daily New V6.5

  • autonomous drivingIt analyzes the road environment in real time, accurately identifies traffic signs, pedestrians, and vehicles, and provides efficient and safe decision support for autonomous driving systems, thereby improving the intelligence level of autonomous vehicles.
  • robotIn the fields of industrial, service, and logistics robots, it helps robots achieve precise object grasping, flexible navigation and obstacle avoidance, and natural human-robot interaction, significantly improving the robot's work efficiency and adaptability.
  • Smart HomeIt provides real-time monitoring of the home environment, intelligent security alarms, and personalized home management services, creating a more convenient and intelligent home living experience for users.
  • Smart EducationIt provides students with personalized learning guidance, quickly solves math problems and corrects homework through image recognition and natural language processing technology, and generates multimedia teaching materials to improve teaching effectiveness and learning experience.
  • HealthcareIn the medical field, it assists doctors in analyzing medical images to quickly and accurately identify lesions, while also providing patients with intelligent triage services, optimizing the medical process, and improving the level of intelligence in medical services.