AB
AiBoss
project

Tongyi DeepResearch - Alibaba's open-source deep research intelligent agent

Tongyi DeepResearch is an open-source deep research agent launched by Alibaba, designed specifically for long-term, deep information retrieval tasks. It boasts 30 billion parameters, activating 3 billion parameters per session, and supports ReAct mode and deep learning...

What is DeepResearch in general terms?

Tongyi DeepResearch is an open-source deep research agent launched by Alibaba, designed specifically for long-term, deep information retrieval tasks. It boasts 30 billion parameters, with 3 billion parameters activated each time. It supports ReAct mode and Heavy Mode, the latter enhancing complex reasoning capabilities through IterResearch. The agent employs a fully synthetic data solution, generating high-quality datasets without human intervention, pushing the limits of agent capabilities. The training process encompasses Agentic CPT, Supervised Fine-tuning (SFT), and Reinforcement Learning (RL), forming a complete end-to-end training chain. Tongyi DeepResearch has already empowered multiple applications within Alibaba, such as the AI-native travel agent for Gaode Maps and "Tongyi FaRui" in the legal field.

The main functions of Tongyi DeepResearch

  • Long-term deep information retrievalDesigned specifically for complex, long-cycle information retrieval tasks, it can handle multi-step reasoning and planning, and is suitable for scenarios such as academic research, market analysis, and policy making.
  • Multimodal reasoning supportSupports ReAct mode and Heavy Mode. ReAct mode strictly follows the "think-act-observe" cycle and is suitable for evaluating the core capabilities of the model; Heavy Mode improves complex reasoning capabilities through an iterative research paradigm.
  • Full-process synthetic data generationIt adopts a self-developed end-to-end synthetic data solution, which can generate high-quality datasets without human intervention, breaking through the upper limit of intelligent agent capabilities and supporting the complete training chain from pre-training to fine-tuning to reinforcement learning.
  • End-to-end reinforcement learningBy using customized reinforcement learning algorithms (such as Group Relative Policy Optimization, GRPO), we can ensure that the behavior of the agent is consistent with the higher-order objective, thereby improving the adaptability and stability of the model in dynamic environments.
  • Empowering Practical ApplicationsIt has been successfully applied to multiple scenarios within Alibaba, such as the AI-native travel agent of Gaode Maps and "Tongyi Farui" in the legal field, demonstrating strong practicality and value.
  • Open source collaborationThe project is completely open source, providing complete code, models, and data, encouraging developers to participate in co-construction, and promoting the development and innovation of deep research intelligent agents.

The technical principles of Tongyi DeepResearch

  • End-to-end synthetic data solutionIt automatically generates high-quality datasets without human intervention, supports the complete training chain from pre-training to fine-tuning to reinforcement learning, and breaks through the upper limit of intelligent agent capabilities.
  • Iterative Research ParadigmComplex tasks are broken down into multiple research rounds, with each round dynamically reconstructing and streamlining the workspace. Through the "think-synthesis-action" process, complex reasoning ability and decision-making quality are improved.
  • End-to-end reinforcement learning: Employ customized reinforcement learning algorithms, such as Group Relative Policy Optimization (GRPO), to ensure that the learning signals are accurately matched with the model's current capabilities, thereby improving the model's adaptability and stability in dynamic environments.
  • Large-scale continuous pre-training: Build an open-world knowledge memory by utilizing continuously updated knowledge documents, crawled data, knowledge graphs, etc., generate multi-style (question, answer) pairs, and continuously expand the model's capabilities.
  • Automated data management: Optimize data in real time under the guidance of training dynamics, and dynamically adjust the training set through fully automatic data synthesis and data funnel to ensure training stability and performance improvement.
  • Stable and efficient tool sandboxDevelop a unified sandbox environment to handle concurrency and failures, ensure the stability and reliability of tool calls, and provide a fast and robust interaction environment for intelligent agents.

Tongyi DeepResearch project address

  • Project official website: https://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/
  • Github repositoryhttps://github.com/Alibaba-NLP/DeepResearch
  • HuggingFace model libraryhttps://huggingface.co/Alibaba-NLP/Tongyi-DeepResearch-30B-A3B

Tongyi DeepResearch family members

  • WebWalker: Focuses on webpage traversal tasks and is used to evaluate the performance of language models in webpage navigation.
  • WebDancerIt is committed to achieving autonomous information seeking capabilities and promoting the autonomy of intelligent agents in information retrieval.
  • WebSailorUsed for navigating complex web environments and enhancing the superhuman reasoning ability of intelligent agents.
  • WebShaperBy formalizing information seeking, we can synthesize agent data, thereby improving data quality and model performance.
  • WebWatcher: To explore new frontiers of visual-language intelligent agents and conduct in-depth research by combining visual and language capabilities.
  • WebResearcherUnleash the unbounded reasoning capabilities of long-cycle intelligent agents and improve their performance in complex tasks.
  • ReSumBy summarizing context, we can unlock long-term search intelligence and optimize the information management capabilities of intelligent agents.
  • WebWeaver: Utilizing evidence of the scale of dynamically structured network outlines to support open-ended in-depth research.
  • WebSailor-V2By using synthetic data and scalable reinforcement learning, we can narrow the gap with proprietary agents.

Application scenarios of generalized DeepResearch

  • academic researchIt can quickly organize literature reviews, helping scholars to efficiently complete complex academic research tasks and improve research efficiency.
  • Market AnalysisIt provides companies with competitor analysis, industry trend reports, and other information to help them develop precise market strategies.
  • Legal ResearchIn the legal field, applications such as "Tongyi Farui" automatically retrieve legal provisions, similar cases, and judgments, and conduct in-depth analysis, providing legal professionals with powerful productivity tools.
  • Travel planningIn partnership with Amap, we launched an AI-native travel agent that combines real-time data to provide users with accurate travel suggestions and plans.
  • Complex Information RetrievalIt is suitable for complex information retrieval tasks that require multi-step reasoning and planning, such as cross-disciplinary research and policy making, helping users to quickly acquire and integrate information.