AB
AiBoss
project

Skywork R1V - Kunlun Tech's open-source multimodal thinking chain inference model

Skywork R1V is Kunlun Tech's first open-source industrial multimodal thought chain reasoning model, possessing powerful visual chain reasoning capabilities. Skywork R1V can perform multi-step logical reasoning on visual input, solving complex visual tasks...

What is Skywork R1V?

Skywork R1V is Kunlun Tech's first open-source industrial multimodal reasoning model, possessing powerful visual chain reasoning capabilities. Skywork R1V can perform multi-step logical reasoning on visual input, solving complex visual tasks such as visual logic reasoning, visual mathematics problems, scientific phenomenon analysis, and medical image diagnosis. The model performs exceptionally well in multiple authoritative benchmark tests, achieving high scores of 94.0 and 72.0 in the MATH-500 and AIME tests respectively, significantly outperforming other mainstream models. The open-source nature of Skywork R1V promotes the development of multimodal reasoning models, contributing to academic research and industrial application exploration.

Main functions of Skywork R1V

  • Visual chain reasoningIt performs multi-step logical reasoning on visual input (such as images or videos) to gradually analyze and deduce the answer to complex problems.
  • Mathematical and scientific problem solvingIt identifies and analyzes mathematical problems or scientific phenomena in images, and provides step-by-step solutions based on reasoning ability.
  • Cross-modal understandingDeeply integrate visual and textual information to achieve richer semantic understanding.
  • Complex visual task processingIt can handle complex visual tasks, such as medical image diagnostic reasoning and art analysis.

The technical principles of Skywork R1V

  • Multimodal transfer of text reasoning abilityBased on a visual projector, it efficiently transfers text reasoning capabilities to visual tasks without retraining the language model and visual encoder. It retains the model's powerful capabilities in text reasoning tasks while handling visual input.
  • Multimodal hybrid training (Iterative SFT + GRPO)This approach combines Iterative Supervised Finite Fibre (SFT) and Group Relative Policy Optimization (GRPO) reinforcement learning to align visual and textual representations in stages. Through iterative training using a combination of high-quality and challenging data, the model's performance on cross-modal tasks is improved, achieving or surpassing existing leading models in visual reasoning benchmarks.
  • Adaptive Length Mind Chain DistillationAn adaptive inference chain length control mechanism based on visual-text complexity is introduced to dynamically optimize the model's inference process. Combined with a multi-stage self-distillation strategy, this avoids the model "overthinking" and improves inference efficiency and quality.
  • Three-stage training method:
    • Initial alignmentA lightweight vision adapter (MLP) is used to connect the visual encoder and the language model, and the model is trained on regular multimodal data to initially align visual and language representations.
    • Reasoning ability transferThe trained adapter is connected to the strong inference language model to form a visual inference model, giving the model initial visual inference capabilities.
    • Precise alignmentBased on the hybrid optimization framework (Iterative SFT + GRPO), the visual and linguistic modalities are further aligned more accurately, improving the model's multimodal reasoning capabilities.

Skywork R1V performance

  • Logical reasoning ability:
    • In the MATH-500 benchmark test, Skywork R1V achieved a high score of 94.0, significantly higher than other open-source models of similar or larger scale.
    • In the AIME 2024 benchmark test, the Skywork R1V achieved a pass rate of 72.0%.
    • In the GPQA (General Physics Question Answering) benchmark test, Skywork R1V achieved a pass rate of 61.6%.
  • Visual comprehension ability:
    • In the MathVista (Visual Mathematical Reasoning) benchmark test, Skywork R1V scored 67.5 points.
    • In the MMMU (Multimodal Medical Understanding) benchmark test, Skywork R1V achieved a score of 69.0.

Skywork R1V project address

Application scenarios of Skywork R1V

  • Educational guidanceIt helps students solve problems in subjects such as mathematics and physics, providing solution steps and analysis.
  • Medical image analysisIt assists doctors in analyzing medical images, inferring lesion characteristics, and providing diagnostic suggestions.
  • Scientific research assistanceIt analyzes experimental images and literature to deduce scientific phenomena and help researchers verify results.
  • Content creation and review: Analyze artworks, detect illegal content, and assist in art appreciation and content review.
  • Industrial Quality Inspection and Market AnalysisIt can detect product defects, analyze advertising and market data, and assist in quality control and business decisions.