AB
AiBoss
project

DeepSeek-R1-0528 - The latest open-source R1 model from DeepSeek.

DeepSeek-R1-0528 is the latest AI model released by the DeepSeek team. The model is trained based on DeepSeek-V3-0324 and has 660 bytes of parameters. It is open-source on HuggingFace, allowing developers to freely use and modify it...

What is DeepSeek-R1-0528?

DeepSeek-R1-0528 is the latest AI model released by the DeepSeek team. Trained on DeepSeek-V3-0324, it boasts 660 bytes of parameters. The model is open-source on HuggingFace, allowing developers to freely use and modify it. Key highlights of DeepSeek-R1-0528 include deep inference capabilities, optimized text generation, a unique inference style, and the ability to process single tasks for 30-60 minutes. The model performs exceptionally well on programming tasks, particularly in complex task processing and code generation, surpassing top-tier models such as Claude 4 Sonnet and Gemini 2.5 Pro. Users can experience the latest version by accessing the dialogue interface through the official website, app, or mini-program and enabling the "Deep Thinking" feature. The API has been updated, but the calling method remains unchanged.

Main functions of DeepSeek-R1-0528

  • Deep reasoningIt supports complex logical reasoning and multi-step thinking to solve complex problems.
  • Programming skillsGenerates high-quality code and supports a variety of programming tasks, such as simulating physical phenomena and front-end design.
  • Text generationGenerates natural and fluent text with standardized formatting, suitable for writing tasks.
  • Long-term thinkingSingle-task processing time can reach 30-60 minutes, suitable for complex tasks.
  • Tool callSupports tool calls and extends model functionality.
  • role playSupports multi-role dialogue, suitable for interactive scenarios.

Technical Principles of DeepSeek-R1-0528

  • Model Architecture and Training FundamentalsIt is trained based on the DeepSeek-V3-0324 model, with a parameter count of 660B. It inherits the features of the V3 version in terms of basic architecture and further optimizes it.
  • Text generation optimizationOptimizations have been made to text generation, resulting in more natural-looking and better-formatted text. This is based on fine-tuning of the language model, including improvements to vocabulary selection, sentence structure generation, and contextual understanding.

Performance of DeepSeek-R1-0528

  • Programming skills:In the LiveCodeBench benchmark test, its performance is almost equivalent to OpenAI's o3-high, and even surpasses top-tier large models such as Claude 4 Sonnet and Gemini 2.5 Pro.
  • Mathematical reasoningIn the AIME 2025 test, accuracy improved from 70% to 87.5%. In the AIME 2024 test, DeepSeek-R1-0528-Qwen3-8B performed second only to DeepSeek-R1-0528, surpassing Qwen3-8B (+10.0%), and was comparable to Qwen3-235B.
  • Tool callIn the Tau-Bench benchmark, its performance was comparable to OpenAI o1-high, but it still lagged behind o3-High and Claude 4 Sonnet.

The project address for DeepSeek-R1-0528

Application scenarios of DeepSeek-R1-0528

  • Natural Language ProcessingGenerate news, stories, copy, etc., support multilingual translation, and build an intelligent question-and-answer system.
  • Programming aidsGenerates high-quality code, supports multiple programming languages, optimizes existing code, improves efficiency and readability, and provides debugging suggestions for developers.
  • Educational supportIt provides students with personalized learning suggestions and tutoring to help users better understand and master knowledge.
  • Corporate OfficeAutomatically generates meeting minutes, reports, emails, and other documents to improve office efficiency; generates market research reports to analyze market trends and consumer behavior, providing support for corporate decision-making.