Auto Think - Kuaishou's open-source large-scale automatic thinking model
Auto Think is an open-source automatic thinking model developed by the KwaiCoder-AutoThink-preview team from Kuaishou's Kwaipilot platform. This model addresses the "overthinking" problem inherent in deep thinking models by conducting in-depth research and proposing a...
What is Auto Think?
Auto Think is an open-source automatic thinking model from the KwaiPilot team at Kuaishou, developed by KwaiCoder-AutoThink-preview. Addressing the "overthinking" problem inherent in deep thinking models, the model proposes a novel training paradigm for automatic thinking models. Based on the traditional reinforcement learning algorithm (GRPO), it introduces a process-supervised reinforcement learning method, Step-SRPO, to further improve the model's performance on complex tasks. The model integrates "thinking" and "non-thinking" abilities, automatically switching its thinking mode according to the difficulty of the problem. Through this thinking mode training, the model has achieved performance improvements on multiple "thinking" and "non-thinking" benchmarks. In some coding and mathematical tasks, enabling the automatic thinking mode resulted in a score increase of up to 20 points.
Main functions of Auto Think
- Automatically switch thinking modesThe model integrates "thinking" and "non-thinking" abilities, automatically switching between thinking modes based on the difficulty of the problem. For simple problems, the model uses a "fast thinking" mode to provide the answer directly, avoiding unnecessary complex reasoning processes; for complex problems, it switches to a "slow thinking" mode to conduct in-depth reasoning and analysis, solving the problem more accurately.
-
Improve efficiency and performanceThe ability to automatically switch thinking modes improved the model's performance across multiple "thinking" and "non-thinking" benchmarks. In some coding and mathematical tasks, enabling the automatic thinking mode resulted in a score increase of up to 20 points.
The technical principles of Auto Think
- Minimal interventionBy using an ellipsis-emphasized ellipsis prompt, the model's ability to randomly switch thinking modes is activated. This prompt structure is simple yet effective, guiding the model to switch between different thinking modes and providing a foundation for subsequent reinforcement learning training.
- Multi-stage reinforcement learning
-
Phase 1The goal of this stage is to enable the model to consistently employ both fast and slow thinking modes. "Fast thinking" is used to solve simple problems, while "slow thinking" is used for complex problems.
-
Phase TwoThis stage of training optimizes the model's ability to respond correctly under both fast and slow thinking patterns. Through this phase of training, the model can handle problems more accurately under different thinking patterns, thus improving its overall performance.
-
Phase ThreeThis stage of training refines the thought process chain output of fast and slow thinking. After this stage, the model no longer randomly decides whether to think deeply, but can autonomously select the thinking mode according to the difficulty of the problem, achieving a more efficient and accurate reasoning process.
-
Auto Think's project address
- HuggingFace model library:https://huggingface.co/Kwaipilot/KwaiCoder-AutoThink-preview
Auto Think Application Scenarios
-
Video generationAuto Think's automatic thinking capabilities can further optimize the video generation process, making the generated video content more suitable for different levels of difficulty and complexity.
-
CopywritingAuto Think can automatically switch thinking modes based on the difficulty of the question, providing more efficient and accurate ideas and methods for copywriting.
-
Intelligent Customer ServiceAuto Think's automatic thinking capability allows it to quickly and accurately respond to user interactions based on the complexity of the questions, thus improving the user experience.
-
Precise SearchAuto Think's automatic thinking capabilities can further optimize search results, providing more accurate information that better meets user needs.
-
Personalized recommendationsAuto Think can automatically switch thinking modes based on the user's personalized needs, providing more accurate recommendation results.