AB
AiBoss
project

Intern-S1 - A scientific multimodal large-scale model launched by Shanghai AI Lab

Intern-S1 is a scientific multimodal large-scale model officially open-sourced and released by the Shanghai Artificial Intelligence Laboratory at the World Artificial Intelligence Conference. It integrates language and multimodal performance, possesses a high level of balanced development capability, and incorporates multidisciplinary expertise...

What is Intern-S1?

Intern-S1, officially open-sourced at the World Artificial Intelligence Conference by the Shanghai Artificial Intelligence Laboratory, is a large-scale scientific multimodal model that integrates language and multimodal performance. It boasts high-level balanced development capabilities and is rich in multidisciplinary expertise, demonstrating outstanding performance in the scientific field. Intern-S1 pioneered a "cross-modal scientific analysis engine," capable of accurately interpreting complex scientific modal data such as chemical formulas, protein structures, and seismic wave signals. It can predict compound synthesis pathways and assess the feasibility of chemical reactions. It surpasses top-tier closed-source models on multidisciplinary professional task benchmarks, showcasing superior scientific reasoning and understanding capabilities. Intern-S1 achieves deep integration of multiple scientific modalities through a dynamic tokenizer and a temporal signal encoder, employing a general-specific integrated scientific data synthesis method, possessing powerful general-purpose reasoning capabilities and multiple top-tier professional capabilities.

Main functions of Intern-S1

  • Cross-modal scientific analysis
    • ChemistryIt can accurately interpret chemical molecular formulas, predict the synthetic pathways of compounds, and determine the feasibility of chemical reactions.
    • Biomedical fieldIt can analyze protein sequences, aiding in drug target discovery and clinical translation value assessment.
    • Earth SciencesIt can identify seismic wave signals, analyze seismic wave events, and provide support for earthquake research.
  • Language and visual integrationIt combines language and visual information to perform complex multimodal tasks, such as text-based question answering and explanation of scientific phenomena.
  • Scientific Data ProcessingIt supports the input of various complex scientific modal data, including light curves in materials science and gravitational wave signals in astronomy.
  • Answers to scientific questionsIt can provide accurate answers based on the input scientific questions, combined with its powerful knowledge base and reasoning capabilities.
  • Experimental Design and OptimizationIt assists researchers in designing experimental plans, optimizing experimental procedures, and improving research efficiency.
  • Multi-agent collaborationIt supports multi-agent systems and can work collaboratively with other agents to complete complex scientific research tasks.
  • Autonomous learning and evolutionIt possesses a certain degree of self-learning ability and can continuously optimize its own performance through interaction with the environment.
  • Data processing and analysisIt provides data processing and analysis tools to help researchers quickly process and analyze scientific data.
  • Model Deployment and ApplicationIt supports multiple deployment methods, including local deployment and cloud services, making it convenient for researchers to use in different scenarios.

Intern-S1 Technical Principles

  • Innovative multimodal architectureThe Intern-S1, through the addition of a dynamic tokenizer and a time-series signal encoder, supports a variety of complex scientific modal data, including chemical formulas, protein sequences, light curves, gravitational wave signals, and seismic waveforms. This innovation enables a deeper understanding and more efficient processing of scientific modal data; for example, its compression rate for chemical formulas is more than 70% higher than that of DeepSeek-R1.
  • Large-scale scientific pre-trainingThe model is built upon a MoE language model with 235 billion parameters and a visual encoder with 6 billion parameters, and pre-trained on multimodal data containing 5 trillion tokens, of which over 2.5 trillion tokens are from scientific fields. This enables the model to excel in both general capabilities and specialized scientific domains, demonstrating outstanding performance in professional tasks such as chemical structure interpretation and protein sequence understanding.
  • Joint Optimization Systems and AlgorithmsThe Intern-S1 R&D team has achieved efficient and stable reinforcement learning training of large-scale multimodal MoE models with FP8 accuracy, reducing training costs by 10 times compared to recently released MoE models. At the system level, a training-inference separation RL scheme is adopted, using a self-developed inference engine for efficient large-scale asynchronous inference with FP8. At the algorithm level, a Mixture of Rewards learning algorithm is proposed, integrating multiple reward and feedback signals to improve training efficiency and stability.
  • Scientific data synthesis that integrates general and specialized knowledgeTo address the specialized needs of high-value tasks in the scientific field, Intern-S1 employs a hybrid approach to scientific data synthesis that integrates general and specialized data. On one hand, it leverages massive amounts of general scientific data to broaden the model's knowledge base; on the other hand, it generates highly readable scientific data through specialized models, with quality control performed by domain-specific verification agents.

Intern-S1 Project Address

  • Project official websiteThe Scholar Model
  • Github repositoryhttps://github.com/InternLM/Intern-S1
  • HuggingFace model libraryhttps://huggingface.co/internlm/Intern-S1-FP8

Application scenarios of Intern-S1

  • Image and text fusionIntern-S1 can handle image and text fusion tasks, such as describing the content in an image and explaining scientific phenomena in an image.
  • Complex scientific modal data processingIt supports the input of various complex scientific modal data, including light curves in materials science and gravitational wave signals in astronomy, and enables deep fusion and efficient processing of these data.
  • Scientific research tool integrationIntern-S1 can be integrated into research tools to help researchers quickly process and analyze scientific data.
  • Answers to scientific questionsAs an intelligent assistant, Intern-S1 can answer various scientific questions based on its powerful knowledge base and reasoning capabilities.