TinyR1-Preview - A reasoning model jointly developed by Qihoo 360 and Peking University.
TinyR1-Preview is a 32-byte inference model jointly developed by the School of Computer Science at Peking University and 360 Security Technology Inc. Using only 5% of the parameters, the model approximates the performance of Deepseek-R1-671B. TinyR1-Preview has significant advantages in mathematical fields...
What is TinyR1-Preview?
TinyR1-Preview is a 32-byte inference model jointly developed by the School of Computer Science and Technology at Peking University and 360 Security Technology Inc. Using only 5% of the parameters, the model approaches the performance of Deepseek-R1-671B. In the mathematics domain (AIME score 78.1), TinyR1-Preview is close to the original R1 (79.8 points), far exceeding the 70-byte Deepseek-R1-Distill-Llama (70.0 points). TinyR1-Preview is based on a "divide and conquer-fusion" strategy, training models separately for three vertical domains: mathematics, programming, and science. It then uses the Mergekit tool to achieve intelligent fusion, breaking through performance limits.
Main functions of TinyR1-Preview
- Strong mathematical reasoning abilityIt excels at complex mathematical problems (such as AIME 2024) and solves highly difficult mathematical problems quickly and accurately.
- Highly efficient programming aidsIt supports code generation and debugging, helping developers quickly solve problems and improve programming efficiency.
- Answers to scientific questionsIt supports the handling of complex scientific questions, providing accurate answers and explanations.
- Lightweight deploymentIt requires only 32 bytes of parameters, resulting in lower inference costs compared to large models, making it suitable for resource-constrained scenarios.
The technical principles of TinyR1-Preview
- Divide and conquer strategyBased on massive domain data generated by DeepSeek-R1, sub-models are trained for vertical domains such as mathematics, programming, and science, with each sub-model focusing on tasks in a specific domain.
- Intelligent integrationBased on the Arcee team's Mergekit tool, it intelligently merges sub-models from different domains, breaking through the performance limit of a single model and achieving balanced optimization across multiple tasks.
- Distillation technologyBased on the model distillation method, knowledge from a large model is transferred to a smaller model, achieving more than 95% of the performance of the original R1 model with only 5% of the parameters.
- Optimize trainingBased on domain data training and intelligent fusion, TinyR1-Preview significantly improves inference efficiency and performance while maintaining its lightweight nature, making it suitable for rapid deployment and application.
TinyR1-Preview project address
- HuggingFace model library:https://huggingface.co/qihoo360/TinyR1-32B-Preview
Application scenarios of TinyR1-Preview
- EducationIt assists in math learning and programming education by providing problem-solving strategies and code generation.
- Scientific research and academicIt helps researchers answer scientific questions, design experiments, and analyze data.
- Software developmentGenerate code, optimize algorithms, and improve development efficiency.
- Enterprise ApplicationsIt supports data analysis and process optimization, assisting enterprises in decision-making.
- Personal lifeAs an intelligent assistant, it provides knowledge retrieval and learning support.