Alibaba-NLP/DeepResearch
41.4
Weak · 19 September 2026
22.7k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a comprehensive framework for building and evaluating LLM-based web agents capable of complex, multi-turn information seeking and deep research. It provides various agent architectures that orchestrate tools for web search, page navigation, file parsing, and media analysis to synthesize answers from unstructured data. The repository also includes specialized components for context summarization, multimodal reasoning, and dual-agent planning, alongside a suite of evaluation benchmarks and inference pipelines for measuring performance on factual and exploratory tasks.
Features
Added AgentFold inference script and multi-GPU serving setup
This change introduces the \WebAgent/AgentFold\ component, providing a new inference entry point (\infer.py\) and a shell script (\serve.sh\) to deploy the underlying model. The inference script implements an agent loop that utilizes a Qwen3-30b tokenizer and connects to a pool of local API endpoints (ports 8000–8006) to execute tools such as web search and page visiting. The accompanying \serve.sh\ script configures vLLM to serve the model across multiple GPUs (ports 8000–8005) and includes a separate instance for a GPT-OSS-120b model on port 8006, enabling the agent to distribute inference load.
WebAgent/AgentFold · high confidence
Added official evaluation scripts and prompts for Tongyi DeepResearch
New evaluation tools have been added to the \evaluation\ directory to support benchmarking the Tongyi DeepResearch model. This includes \evaluate\_deepsearch\_official.py\ and \evaluate\_hle\_official.py\ for running benchmarks, along with \prompt.py\ containing the specific system, extraction, and judge prompts used for these evaluations. A \README.md\ has also been added to document the required environment variables and execution commands for both HLE and other benchmark datasets.
evaluation · high confidence
Initial release of WebAgent repository
The WebAgent directory has been initialized with core project files, including a .gitignore configuration for Python development environments, an MIT License, and a comprehensive README. The documentation introduces the WebAgent framework for information-seeking tasks, detailing the release of the WebLeaper model and providing quick-start guides and demos for related models such as WebDancer, WebSailor, and WebWatcher.
WebAgent · high confidence
Introduce NestBrowse and ParallelMuse web agent implementations
Added the NestBrowse and ParallelMuse web agent components, which enable automated, multi-turn web browsing and information synthesis. NestBrowse provides an asynchronous agent loop that interacts with a browser via an MCP server (supporting visit, click, and fill actions) and uses LLMs to summarize webpage content and extract evidence. ParallelMuse introduces a compressed reasoning aggregation module that generates detailed problem-solving reports from agent trajectories and consolidates multiple independent team reports to derive a final answer.
WebAgent/NestBrowse, WebAgent/ParallelMuse · high confidence
Introduce ReSum context summarization for long-horizon web agent exploration
Adds the ReSum inference paradigm to the WebAgent, enabling periodic compression of conversation history into compact summaries to support unlimited exploration without hitting context limits. This includes the core agent logic in src/react\_agent.py that triggers summarization, the summary generation utilities in src/summary\_utils.py, and dedicated prompts in src/prompt.py. The change also provides a complete evaluation setup with src/evaluate.py and judge prompts in src/judge\_prompt.py, along with shell scripts (run\_react.sh, run\_resum.sh) to launch inference and a README documenting the workflow.
WebAgent/WebResummer · high confidence
Introduce Tongyi DeepResearch inference pipeline
The \inference\ directory now contains a complete deep research agent system. This includes a \MultiTurnReactAgent\ that orchestrates a ReAct loop using tools for web search (\search\), academic search (\google\_scholar\), file parsing (\parse\_file\), Python code execution (\PythonInterpreter\), and webpage visiting (\visit\). The agent uses a structured prompt (\prompt.py\) to guide the model, handles context limits and token counting, and supports multi-branch rollouts via \run\_multi\_react.py\ and \run\_react\_infer.sh\ for batch evaluation on datasets like GAIA.
inference · high confidence
Introduce WebWalker: LLM-based web traversal benchmark and agent
Adds the WebWalker project, a benchmark and multi-agent framework for evaluating LLMs on complex web traversal tasks. The release includes the WebWalkerQA dataset (680 queries across 1,373 pages), a Streamlit-based interactive demo (\src/app.py\) for navigating websites, and a RAG-system evaluation pipeline (\src/rag\_system.py\, \src/evaluate.py\) supporting multiple API providers (OpenAI, Gemini, Ark, Moonshot, Baidu). The core agent (\src/agent.py\) implements a ReAct-style loop with memory management and information extraction to answer queries by exploring web pages.
WebAgent/WebWalker · high confidence
Introducing WebSailor: an open-source web agent for complex information-seeking tasks
This release introduces WebSailor, a new open-source web agent designed to handle high-uncertainty, multi-step reasoning tasks on the web. The package includes the \MultiTurnReactAgent\ implementation, which utilizes \search\ and \visit\ tools to navigate information landscapes, along with a dedicated evaluation framework for benchmarks like BrowseComp and GAIA. It also provides the \SailorFog-QA\ dataset, a collection of complex questions synthesized to test an agent's ability to navigate extreme uncertainty.
_WebAgent/WebSailor, WebAgent/WebShaper, inference/eval\data · high confidence
Introducing WebWeaver, a dual-agent deep research system
WebAgent/WebWeaver introduces a new deep research agent that uses a dual-agent framework (planner and writer) to perform open-ended research. The planner dynamically optimizes a living outline by interleaving web searches with refinement, while the writer synthesizes reports using memory-grounded hierarchical retrieval. The package includes a DashScope API client for model inference, shell scripts for running the planner and writer, and sample evaluation data, requiring specific API keys (Serper, ScraperAPI, DashScope) and a summary model.
WebAgent/WebWeaver · high confidence
New evaluation datasets added for WebAgent/WebDancer
Added two new JSONL dataset files, \sample\_qa.jsonl\ and \sample\_traj.jsonl\, to the \WebAgent/WebDancer/datasets\ directory. \sample\_qa.jsonl\ contains 200 question-answer pairs tagged with \e2hqa\ and \crawlqa\ for evaluating factual and crawl-based queries. \sample\_traj.jsonl\ contains 200 multi-turn agent interaction trajectories tagged with \agent/multiturn\_search\, demonstrating tool usage (search, visit) and reasoning steps for evaluating web information seeking capabilities.
WebAgent/WebDancer · high confidence
New file parsing and video/audio analysis tools
Added new capabilities for processing documents and media files. The \file\_parser.py\ module introduces support for parsing PDF, Word, PowerPoint, Excel, text, HTML, CSV, and ZIP files, including integration with an Alibaba Cloud Intelligent Document Processing (IDP) service for advanced layout extraction. Additionally, \video\_analysis.py\ and \video\_agent.py\ provide tools to analyze video and audio files (MP4, MOV, MP3, WAV, etc.) by extracting transcripts, key frames, and performing AI-based content analysis.
_inference/file\tools · high confidence
WebWatcher deep research agent and BrowseComp-VL benchmark released
This change introduces WebWatcher, a multimodal vision-language agent for deep research, along with the BrowseComp-VL benchmark dataset. The release includes the WebWatcher README with setup instructions and performance highlights, as well as the BrowseComp-VL Level 1 and Level 2 evaluation data files (JSONL), enabling users to evaluate multimodal reasoning capabilities on complex visual search tasks.
WebAgent/WebWatcher · high confidence
Dependencies
Initial dependency manifests for WebAgent components
This change introduces the initial Python dependency manifests (requirements.txt) for the WebAgent project, establishing the runtime environment for its sub-components. WebDancer and WebSailor are configured with sglang and qwen-agent (including GUI, RAG, code interpreter, and MCP extras), while WebWalker depends on crawl4ai, langchain, and streamlit. The WebWatcher component and the root project include extensive pinned dependencies for AI inference and data processing, featuring torch, transformers, vllm, openai, and various Alibaba Cloud SDKs.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 41.
Lenses
- Code Health 79
- Architecture 100
- Maturity 75
- Readiness 12
- Security 56
Changes since last survey
- 300 commits — 277 feature/other, 23 fixes
By area
- (root) — 137 commits
- (repo) — 30 commits
- WebDancer/readme.md — 16 commits
- assets/wechat.jpg — 8 commits
- WebAgent/WebWatcher — 6 commits
- WebDancer/assets — 6 commits
- WebSailor/README.md — 6 commits
- WebWatcher/README.md — 5 commits
- assets/roadmap.png — 5 commits
- inference/react_agent.py — 5 commits
- WebDancer/demos — 4 commits
- WebSailor/assets — 4 commits
- WebSailor/src — 4 commits
- WebShaper/readme.md — 4 commits
- WebWatcher/assets — 4 commits
- Agent/AgentFounder — 3 commits
- Agent/AgentScaler — 3 commits
- WebAgent/README.md — 3 commits
- WebAgent/WebResummer — 3 commits
- assets/wechat_new.jpg — 3 commits
Notable commits
- fix: Fix OpenAI API key problem in Issue (#130)
- fix: Fix inconsistent parameter passing for Jina API key in WebSailor
- fix: Fix: Inconsistent parameter (#60)
- fix: Fix: Update Crawl4AI name (#8)
- fix: Fix: Update Crawl4AI name and branding. Thanks for the acknowledgment.
- fix: Merge pull request #6 from Alibaba-NLP/fix
- fix: fix Citation
- fix: fix Citation (#79)
- fix: fix WebWalker env variable and root url bug (#14)
- fix: fix arxiv type
- fix: fix bug
- fix: fix bug
- fix: fix bugs #13: update env
- fix: fix count_tokens bug (#165)
- fix: fix crawl4ai bug
- fix: fix env bug (#104)
- fix: fix logic bug
- fix: fix parse_retry_times bug (#233)
- fix: fix search bug
- fix: fix the 130 issue
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Alibaba-NLP/DeepResearch was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f72f75d8c3eb842f2bbbab096a12206ff66e270f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.