AB
AiBoss
project

360 GPT2-O1 - 360 launches domestically developed AI large model, outperforming GPT-4O in multiple tests.

360gpt2-o1 is a self-developed AI model by 360, demonstrating significant improvements in reasoning capabilities, particularly excelling in mathematical and logical reasoning tasks. The model is optimized through synthetic data, post-training, and the "slow thinking" paradigm...

What is 360gpt2-o1?

360gpt2-o1 is a self-developed AI model by 360, demonstrating significant improvements in reasoning capabilities, particularly excelling in mathematical and logical reasoning tasks. The model achieves technological breakthroughs through synthetic data optimization, post-training, and a "slow thinking" paradigm, resulting in outstanding performance in numerous authoritative evaluations. In basic mathematics assessments (such as MATH and the National College Entrance Examination in Mathematics) and authoritative mathematics competitions (including AIME24 and AMC23), 360gpt2-o1 surpasses its predecessor, 360gpt2-pro, and outperforms the GPT-4o model. In mathematics competition evaluations, 360gpt2-o1 surpasses Alibaba's latest open-source o1 series model, QWQ-32B-preview.

Main functions of 360gpt2-o1

  • Improved reasoning abilityThe 360gpt2-o1 performed exceptionally well on mathematical and logical reasoning tasks, with a particularly significant improvement in reasoning ability.
  • Synthetic data optimizationBy employing methods such as instruction synthesis and quality/diversity screening, the problem of scarcity of high-quality mathematical and logical reasoning data has been solved, effectively expanding the training dataset.
  • Post-training of the modelA two-stage training strategy is adopted: first, a small model is used to generate diverse reasoning paths, and then a large model is used for RFT training and reinforcement learning training to improve the model's reasoning ability and reflection and error correction ability.
  • "Slow Thinking" ParadigmBased on Monte Carlo tree search, we explore diverse solutions, introduce LLM for error verification and correction, simulate the human process of step-by-step reasoning and reflection, and finally form a long thought chain that includes reflection, verification, error correction and backtracking.

Technical principles of 360gpt2-o1

  • Data synthesis and filteringThrough synthetic data optimization, 360gpt2-o1 can generate and filter high-quality training data, which is crucial for model training.
  • Two-stage training strategyThe first stage uses a small model to generate reasoning paths, and the second stage uses a large model for training, so that the model can improve the accuracy and depth of reasoning while maintaining the diversity of reasoning.
  • Monte Carlo Tree Search Combined with LLMMonte Carlo tree search allows the model to explore multiple possible solutions, while the introduction of LLM provides the model with error verification and correction capabilities, enhancing the model's robustness.

How to use 360gpt2-o1

  • Access 360 Smart BrainCurrently, 360gpt2-o1 has been launched on the 360 Smart Brain API Open Platform.
  • Experience addressLink: https://ai.360.com/playground/?model=360gpt2-o1?src=weixinmp

Application scenarios of 360gpt2-o1

  • Mathematical Problem SolvingThe 360gpt2-o1 has achieved remarkable results in basic mathematics assessments (such as MATH and the National College Entrance Examination in Mathematics) and authoritative mathematics competitions (including AIME24 and AMC23), demonstrating its strong ability in solving mathematical problems.
  • Logical reasoningThe model uses "slow thinking" technology to simulate the human process of gradual reasoning and reflection, and has the ability to solve complex logical problems.
  • Programming problemsIn mathematics, programming and other fields, it performs close to or even surpasses O1, and 360GPT2-O1 provides support for solving programming problems.
  • Complex Problem Solving360gpt2-o1 can handle complex problems that require deep logical reasoning abilities, including the ability to self-reflect and correct errors.
  • Education and academicThe application of the model to mathematical and logical problems in the field of education can assist teaching and academic research.
  • Enterprise Decision SupportThrough logical reasoning and data analysis, 360gpt2-o1 can assist enterprises in providing logical support during complex decision-making processes.