AB
AiBoss
project

WebThinker - A deep research intelligent agent launched by Renmin University of China in conjunction with the Institute of Artificial Intelligence and other institutions.

WebThinker is a deep research agent proposed by Renmin University of China, Beijing Academy of Artificial Intelligence, and Huawei Poisson Lab, among other institutions. WebThinker empowers large-scale inference models (LRMs) to autonomously perform inference processes...

What is WebThinker?

WebThinker is a deep research agent proposed by institutions such as Renmin University of China, Beijing Academy of Artificial Intelligence, and Huawei Poisson Lab. WebThinker empowers Large Reasoning Models (LRMs) to autonomously perform web searches, web navigation, and report writing during the reasoning process. Based on a deep web explorer and autonomous thinking, searching, and writing strategies, WebThinker enables LRMs to dynamically acquire information and generate high-quality research reports in real time. WebThinker further optimizes tool efficiency through a reinforcement learning training strategy. WebThinker performs exceptionally well in complex reasoning and report generation tasks, significantly improving the reliability and practicality of LRMs in knowledge-intensive tasks.

WebThinker's main functions

  • Autonomous decision-makingLRM autonomously determines when external knowledge is needed and when reports need to be updated during the reasoning process.
  • In-depth explorationIt supports multi-step searches and page navigation, allowing for in-depth information exploration.
  • Dynamic writingThe model can write and modify report content in real time, and is equipped with a dedicated toolset (such as writing, checking, and editing) to ensure the consistency and completeness of the report.
  • Tool optimization: Optimize the efficiency of LRM in using research tools.

WebThinker's technical principles

  • Deep Web ExplorerThis empowers LRM to go beyond traditional simple search, enabling it to navigate between web pages based on interactive elements such as clicked links and buttons, and to delve deeper into information. The model autonomously determines search queries, continuously exploring until sufficient information is collected, and then returns a refined summary.
  • Reinforcement learning-based training strategiesThis approach improves the efficiency of LRM in utilizing research tools (including search, navigation, and report writing tools) by using iterative online direct preference optimization (DPO) training. It constructs a preference dataset and prioritizes inference paths that yield correct answers, high-quality reports, and more efficient tool usage.
  • Operating modeThe problem-solving mode equips LRM with a deep web explorer, enabling in-depth exploration of the web to solve complex problems. The report generation mode further empowers LRM with writing, reviewing, and editing capabilities, allowing for iterative creation of comprehensive research reports while simultaneously thinking and searching.

WebThinker project address

WebThinker Application Scenarios

  • Solutions to complex problemsIt provides quick and accurate answers to doctoral-level scientific questions or interdisciplinary challenges.
  • Research report generatedIndependently search and write scientific research reports to ensure comprehensive, accurate, and coherent content, thereby improving report generation efficiency.
  • Deep information miningBased on multi-step search and page navigation, it retrieves in-depth information and supports complex analysis and research.
  • Educational SupportIn the field of education, it helps students find learning materials and answer academic questions, and generates teaching outlines for teachers to improve learning and teaching efficiency.
  • Enterprise Decision SupportIt provides decision support to businesses, including market analysis and competitor analysis, helping management quickly obtain key information and make more informed decisions.