assafelovic/gpt-researcher
52.1
Adequate · 18 September 2026
28.4k
lines of production code
Python
with TypeScript, JavaScript
1
measurement over time
What this system is
This system is an automated AI research assistant that generates comprehensive reports by orchestrating web searches, document retrieval, and large language model reasoning. It supports multiple research depths, from quick summaries to deep, recursive exploration, and offers various output formats including PDF, DOCX, and Markdown. The platform features a modular backend with multi-agent workflows and a modern frontend that enables real-time monitoring, interactive chat, and embeddable usage.
How it got here
2023–2024 — modular architecture and multi-agent expansion
60 changes.
The project underwent a comprehensive structural overhaul, refactoring the monolithic GPT Researcher agent into modular skill classes and establishing a LangGraph-based multi-agent workflow with specialized roles. This period also introduced extensive support for diverse search retrievers, embedding providers, and scraping backends, while rebuilding the frontend as a Next.js application with real-time WebSocket monitoring and PWA capabilities.
2025–2026 — extensive feature expansion and infrastructure
25 changes.
This period focused on significantly expanding the platform's capabilities by introducing deep research workflows, Model Context Protocol integration, and numerous new search and image generation providers. Concurrently, the team established a robust evaluation framework for factuality and hallucination detection while delivering a distributable React component and initial AWS deployment infrastructure via Terraform.
Features
Add Bing Search Retriever
Introduces a new Bing Search retriever component that allows users to perform web searches via the Bing API. The implementation handles API key configuration via the BING\_API\_KEY environment variable, parses search results, and normalizes them into a standard format (title, href, body) while filtering out YouTube links and handling malformed responses gracefully.
_gpt\_researcher/retrievers/arxiv, gpt\researcher/retrievers/bing · high confidence
Add BoCha Search Retriever with robust error handling
Users can now utilize the BoCha search provider for web research. The new BoChaSearch retriever integrates with the BoCha AI API, normalizing results into the standard title, href, and body format. It includes defensive coding to handle missing keys in provider payloads and gracefully returns an empty result list on network or parsing errors, preventing research runs from crashing.
_gpt\researcher/retrievers/bocha · high confidence
Add Brave Search as a new retrieval source
Users can now use Brave Search as a data source for research queries. This change introduces a new BraveSearch retriever that integrates with the Brave Search API, requiring the BRAVE\_API\_KEY environment variable to be set. The retriever normalizes search results into the standard format (title, href, body) used by other search providers in the system.
_gpt\researcher/retrievers/brave · high confidence
Add FireCrawl scraper with concurrency controls and SDK v4.6+ compatibility
Introduces a new FireCrawl scraper implementation that supports the FireCrawl Python SDK v4.6.0+ API (using \scrape()\ instead of \scrape\_url()\ and accessing metadata as attributes rather than dict keys). To prevent rate-limit errors on the FireCrawl Free Tier, the scraper includes a module-level async semaphore that caps concurrent API calls (defaulting to 2, configurable via \FIRECRAWL\_CONCURRENCY\). It also safely handles optional sessions for image extraction and guards against null responses.
_gpt\researcher/scraper/firecrawl · high confidence
Add GroundRoute multi-engine search retriever
Users can now use the GroundRoute retriever to route search queries across multiple web-search engines (Serper, Brave, Exa, Tavily, Firecrawl, Perplexity), automatically selecting the cheapest option that meets a quality bar. This new component caches repeated queries and includes failover logic, exposing a unified search API that normalizes results into a consistent format with href and body fields.
_gpt\researcher/retrievers/groundroute · high confidence
Add OpenAlex scholarly works retriever
Users can now retrieve information from the OpenAlex open catalog of scholarly works. This new retriever supports searching by query with configurable sorting (relevance, citation count, or publication date) and automatically handles API authentication via environment variables (OPENALEX\_EMAIL, OPENALEX\_API\_KEY). It robustly parses API responses, reconstructing abstracts from inverted indices and selecting the best available link (PDF, landing page, or work ID) for each result.
_gpt\researcher/retrievers/openalex · high confidence
Add PubMed Central retriever for full-text medical research
Introduces a new PubMed Central retriever that searches the NCBI database and retrieves full-text content (title, abstract, and body) from PMC articles. The implementation handles API key configuration via environment variables, supports custom search parameters, and ensures robust error handling by returning empty lists instead of None to prevent downstream type errors.
_gpt\_researcher/retrievers/pubmed\central · high confidence
Add SearchApi retriever
A new SearchApi retriever has been added to the system, allowing users to perform internet searches via the SearchApi service. This component requires the SEARCHAPI\_API\_KEY environment variable and filters out YouTube results from the search output, ensuring that only relevant web results are returned to the user.
_gpt\researcher/retrievers/searchapi · high confidence
Add SearxNG API Retriever with snippet length guard
A new SearxNG API retriever has been added to the system, allowing users to search via a SearxNG instance configured via the SEARX\_URL environment variable. The implementation includes a safeguard that truncates result snippets to 100 characters to prevent the system from incorrectly treating search snippets as full article text, ensuring that actual page content is scraped when needed. The retriever also declares that results require scraping and supports passing query domains to the search API.
_gpt\researcher/retrievers/searx · high confidence
Add Semantic Scholar retriever with camelCase sort support
A new Semantic Scholar retriever has been added to the system, allowing users to search academic papers via the Semantic Scholar API. The implementation supports sorting by relevance, citation count, or publication date, with a specific fix to preserve the API-required camelCase format for sort criteria (e.g., 'citationCount') to ensure valid requests. It also includes robust handling for open-access PDF links, filtering results to only include papers where the open access PDF data is present and valid.
_gpt\_researcher/retrievers/semantic\scholar · high confidence
Add Xquik X/Twitter search retriever
Users can now search for real-time perspectives and discussions on X (Twitter) using the Xquik API. This new retriever, located in gpt\_researcher/retrievers/xquik, requires the XQUIK\_API\_KEY environment variable and returns results in the standard title, href, and body format. The implementation includes robust handling for malformed API responses, such as null or non-list tweet data, to prevent silent failures.
_gpt\_researcher/retrievers/getxapi, gpt\researcher/retrievers/xquik · high confidence
Added Next.js API action handlers and Nginx configuration
The frontend now includes server-side action handlers in \frontend/nextjs/actions/apiActions.ts\ that communicate with backend endpoints (\/api/getSources\, \/api/getAnswer\, \/api/getSimilarQuestions\, \/api/generateLanggraph\) to retrieve sources, generate answers (including streaming responses via Server-Sent Events), and fetch similar questions. Additionally, \frontend/nextjs/nginx/default.conf\ provides an Nginx configuration to serve the Next.js static assets on port 3000, ensuring proper routing for client-side navigation.
frontend/nextjs/actions, frontend/nextjs/nginx · high confidence
Added SerpApi web search retriever
Users can now use SerpApi as a search provider for retrieving web results. This new retriever supports filtering search results to specific domains via the \query\_domains\ parameter and includes robust handling for malformed API responses or missing fields to prevent crashes during search operations.
_gpt\researcher/retrievers/serpapi · high confidence
Added TypeScript type definitions for data models and analytics
Introduced new TypeScript declaration files to provide type safety for frontend data structures and third-party libraries. The \data.ts\ file defines interfaces for various content types (such as basic text, chat, and Langgraph buttons), configuration settings for chat boxes and MCP (Model Context Protocol) support, and structures for chat messages and research history. Additionally, \react-ga4.d.ts\ adds type definitions for the Google Analytics 4 library, enabling proper typing for initialization and event tracking methods.
frontend/nextjs/types · high confidence
Added custom retriever for external API integration
Users can now integrate an external search API by setting the RETRIEVER\_ENDPOINT environment variable. The new CustomRetriever fetches results from this endpoint, automatically mapping query parameters from environment variables prefixed with RETRIEVER\ARG\, and normalizes the response into a consistent list of objects containing 'url' and 'raw\_content'. It includes robust error handling to tolerate malformed payloads or non-list responses, ensuring the research pipeline does not crash on unexpected API output.
_gpt\researcher/retrievers/custom · high confidence
Added fastCRW search retriever with robust error handling
A new CRWRetriever has been added to support web searches via the fastCRW API (Firecrawl-compatible). The implementation allows configuration via the CRW\_API\_KEY and CRW\_API\_URL environment variables (or headers) to support both managed cloud and self-hosted deployments. The retriever includes specific error handling to skip individual search results that lack a URL or contain malformed data, ensuring that partial failures do not discard the entire result set.
_gpt\researcher/retrievers/crw · high confidence
Added memory state definitions for draft and research workflows
New Python modules have been introduced in the backend memory package to define the data structures for multi-agent workflows. Specifically, \DraftState\ and \ResearchState\ TypedDicts are now available to structure the context for drafting and research tasks, including fields for tasks, topics, drafts, review notes, research data, and report layout details.
backend/memory · high confidence
Backend server startup and report export utilities
The backend now includes a Dockerfile and Procfile to containerize and run the application via uvicorn on port 8000. Additionally, the backend provides utilities to export research reports as PDF and DOCX files, handling markdown-to-PDF conversion with CSS styling and markdown-to-DOCX conversion using python-docx, while ensuring filenames are sanitized to prevent empty or invalid paths.
backend · high confidence
Initial deployment infrastructure for the GPT Researcher service
This change introduces the complete Terraform configuration to deploy the GPT Researcher application on AWS. It establishes an immutable ECR repository for storing Docker images, configures an ECS service to run the application container (defaulting to image tag v1.0.5 on port 3535), and sets up GitHub Actions IAM roles with OIDC authentication for CI/CD pipelines. The infrastructure includes necessary security groups for EFS NFS access, CloudMap service discovery for inter-service communication, and integration with AWS Secrets Manager for managing sensitive credentials like API keys.
terraform · high confidence
Initial issue triage data added to .triage
The \.triage\ directory now contains \all\_results.json\ and \batch\_1.json\, which store structured metadata and comments for a batch of GitHub issues (e.g., HuggingFace integration, code interpreter support, local data reference bugs). This data serves as the input for the automated triage workflow, allowing the system to analyze issue history and suggest actions like closing, requesting info, or assigning PRs.
.triage · high confidence
Initial project scaffolding and configuration
Established the foundational project structure by adding essential configuration files including \.cursorrules\ and \.cursorignore\ for AI-assisted development, \.env.example\ with comprehensive environment variable documentation (API keys, scraper settings, token limits), \.gitignore\ and \.dockerignore\ for build hygiene, \.mcp.json\ for Model Context Protocol integration, and \.python-version\ for environment consistency. Added standard governance documents (CODE\_OF\_CONDUCT.md, CONTRIBUTING.md) and a Dockerfile for containerized deployment.
(repo-wide) · high confidence
Introduce LangGraph-based multi-agent research workflow
Users can now run an in-depth research process orchestrated by a team of specialized AI agents (Chief Editor, Researcher, Editor, Reviewer, Revisor, Writer, Publisher) using LangGraph. This new capability, located in the \multi\_agents\ directory, automates planning, parallel data collection, review, and report generation (PDF, Docx, Markdown). The workflow supports optional human-in-the-loop feedback, configurable research guidelines, local document sources, and optional observability via Monocle or LangSmith.
_multi\agents · high confidence
Introduce MCP-based research retriever with two-stage tool selection
Users can now leverage Model Context Protocol (MCP) servers for research via a new \MCPRetriever\. This component implements a two-stage intelligent approach: first, an LLM selects the 2-3 most relevant tools from available MCP capabilities, and second, it executes research using only those selected tools. This modular implementation, located in \gpt\_researcher/retrievers/mcp\, relies on the \langchain-mcp-adapters\ library and integrates with existing researcher configurations to provide more targeted and efficient research results.
_gpt\researcher/retrievers/mcp · high confidence
Introduce Model Context Protocol (MCP) integration for external tool usage
GPT Researcher now supports connecting to external tools and data sources via the Model Context Protocol (MCP). This update adds a new \gpt\_researcher/mcp\ module that manages MCP server connections (stdio, WebSocket, and HTTP), intelligently selects relevant tools using an LLM, and executes research tasks through these tools. Users can configure MCP servers via \mcp\_configs\ to extend GPT Researcher's capabilities with local databases, third-party APIs, or custom services, with results seamlessly integrated into the final research report.
_gpt\researcher/mcp · high confidence
Introduce SerperSearch retriever with robust error handling and filtering
Added a new SerperSearch retriever that supports country, language, and time-range filtering, as well as domain inclusion and site exclusion. The implementation includes defensive coding practices to prevent crashes from malformed API responses or missing keys, ensuring the search method always returns a list of results rather than None.
_gpt\researcher/retrievers/serper · high confidence
Introduce context compression and retrieval modules
The \gpt\researcher/context\ package now provides new classes for managing and compressing research context. \ContextCompressor\ filters raw documents using embedding similarity and includes a fast-path optimization that skips expensive compression for small document sets. \VectorstoreCompressor\ retrieves relevant context from an existing vector store. \SearchAPIRetriever\ and \SectionRetriever\ handle document retrieval, with the former truncating content to prevent embedding token limit errors and guarding against \None\ raw content. These components are exported via the new \\\init\\_.py\.
_gpt\researcher/context · high confidence
Introduce vector store wrapper for document ingestion
A new \VectorStoreWrapper\ class has been added to the \gpt\_researcher\ package to handle the conversion, chunking, and storage of documents into a LangChain-compatible vector store. This component translates internal document structures into LangChain \Document\ objects, safely skipping non-dict inputs or entries missing content, splits the text into configurable chunks, and provides an asynchronous similarity search interface, enabling the application to persist and retrieve research data via vector embeddings.
_gpt\_researcher/vector\store · high confidence
Introduces a split-pane research interface with an interactive copilot chat panel
The research view now supports a side-by-side layout where the main research results are displayed in a resizable ResearchPanel alongside a new CopilotPanel. This copilot panel allows users to ask follow-up questions about the research findings via a dedicated chat interface, with automatic scrolling and status indicators. The layout is responsive, collapsing to a single column on mobile devices, and includes a 'New Research' button and share functionality within the panel headers.
frontend/nextjs/components/research · high confidence
Introduction of BasicReport class with extended configuration support
The backend now exposes a new BasicReport class that initializes the GPTResearcher with support for query domains, tone selection, and MCP (Model Context Protocol) configurations. This change allows users to filter search results by specific domains, customize the report's writing style, and integrate external tools via MCP, while also introducing unique research IDs for better tracking.
_backend/report\_type/basic\report · high confidence
Introduction of SimpleQA evaluation framework with structured logging and metrics
The \evals/simple\_evals\ module now provides a dedicated framework for running SimpleQA evaluations, featuring structured JSON logging of results and detailed aggregate metrics including accuracy, F1 score, latency percentiles, and cost tracking. The directory includes a multi-domain test set (Simple QA Test Set.csv), example output demonstrating the new logging format, and a logs directory configured to persist historical evaluation runs for performance tracking.
_evals/simple\evals · high confidence
Introduction of generic LLM and image generation provider interfaces
The \gpt\_researcher/llm\provider\ package now exposes a unified entry point via \\\init\\_.py\, making \GenericLLMProvider\ and \ImageGeneratorProvider\ directly importable from the package root. This change consolidates the provider abstractions, allowing users to access the core LLM and image generation interfaces without needing to know the specific internal module paths (e.g., \.generic\ or \.image\).
_gpt\_researcher/llm\provider · high confidence
Introduction of generic LLM provider abstraction
The GPT Researcher now uses a unified \GenericLLMProvider\ to manage interactions with a wide variety of Large Language Model providers. This change introduces support for numerous new providers including DashScope, DeepSeek, OpenRouter, GigaChat, MiniMax, xAI, vLLM, Atlas Cloud, Nebius, Forge, and Avian, alongside existing ones like OpenAI, Anthropic, and Azure. The provider abstraction also adds support for custom base URLs (e.g., for OpenAI-compatible APIs), tracks real-time usage metadata for cost estimation, and implements logic to handle models that do not support temperature settings or support reasoning effort configurations.
_gpt\_researcher/llm\provider/generic · high confidence
New AG2-based multi-agent research workflow
A new multi-agent orchestration example using the AG2 framework has been added under the \multi\_agents/ag2\ directory. This implementation mirrors the existing LangGraph flow, utilizing a team of eight agents (Chief Editor, Editor, Researcher, Reviewer, Reviser, Writer, Publisher, and Human) to conduct in-depth research, plan outlines, review drafts, and publish final reports in various formats. Users can run this workflow by installing the specific requirements and executing \python -m multi\_agents.ag2.main\, with behavior customizable via \task.json\ settings such as the LLM model, section limits, and human feedback inclusion.
_multi\agents/ag2 · high confidence
New BeautifulSoup scraper with robust text and metadata extraction
A new BeautifulSoup-based scraper has been introduced to handle web page fetching and content parsing. This component improves research reliability by implementing retry logic for transient HTTP errors (429, 500, 502, 503, 504), enforcing a 10MB content size limit to prevent resource exhaustion, and correctly handling character encodings by respecting server-declared charsets while falling back to document detection. It extracts cleaned text, relevant images, and page titles, returning empty values gracefully when pages are unreachable or parsing fails.
_gpt\_researcher/scraper/beautiful\soup · high confidence
New Deep Agents example: GPT Researcher as the research engine
This location introduces a new example that integrates GPT Researcher into LangChain's Deep Agents harness. The example features a Chief Editor agent that plans and reviews, delegating parallel research tasks to subagents powered by GPT Researcher's \quick\_search\ and \deep\_research\ tools. It includes a comprehensive benchmark comparing this setup against a raw Tavily search baseline, demonstrating a significant increase in verified citations and improved coverage of private documents when using GPT Researcher's hybrid mode.
_deep\agents · high confidence
New Deep Research capability with recursive exploration
A new 'Deep Research' report type is introduced, enabling recursive, tree-like exploration of topics with configurable breadth, depth, and concurrency. This feature utilizes reasoning models (e.g., o3-mini) for generating follow-up questions and analyzing results, while standard models handle search query generation. Users can now trigger this advanced research mode via \report\_type='deep'\ and monitor progress through real-time callbacks that report on current depth, breadth, and query status.
_backend/report\_type/deep\research · high confidence
New GPTResearcher React component for embedding
The frontend now includes a distributable React component (\GPTResearcher\) that allows users to embed the research interface into their own React applications. This component accepts configuration props such as \apiUrl\, \apiKey\, and \defaultPrompt\, and manages the research workflow via WebSocket connections and REST API calls. It exposes an \onResultsChange\ callback to notify parent components of updates and includes built-in styling via Tailwind CSS and custom CSS imports.
frontend/nextjs/src · high confidence
New PWA and embed capabilities for the frontend
The frontend now supports Progressive Web App (PWA) functionality and an embeddable widget. A new \manifest.json\ configures the app as a standalone PWA with icons, shortcuts, and theme colors, while a new service worker (\sw.js\) using Workbox precaches static assets and implements caching strategies for offline availability. Additionally, a new \embed.js\ script allows the application to be embedded in other websites via an iframe, supporting custom API URL configuration through local storage and dynamic height adjustments.
frontend/nextjs/public · high confidence
New Task UI components with domain filtering and deep research options
The frontend now includes a new set of components for the Task interface, introducing a DomainFilter that allows users to restrict web research to specific domains (persisted in local storage), and a ResearchForm that exposes a 'Deep Research Report' option alongside existing summary and detailed modes. Additionally, an Accordion component is provided to display agent logs and source differences in a collapsible format, while separate components handle Agent Logs and Report rendering.
frontend/nextjs/components/Task · high confidence
New Tavily search retriever with domain filtering and robust error handling
A new TavilySearch retriever has been added to the system, enabling web searches via the Tavily API. This implementation supports filtering results by specific domains through the query\_domains parameter and automatically translates Google-style 'site:' operators in search queries into compatible domain filters. The retriever is configured to always require content scraping after the initial search and includes defensive coding to handle malformed API responses, missing fields, and empty result sets gracefully, ensuring that a single bad response does not crash the entire search process.
_gpt\researcher/retrievers/tavily · high confidence
New browser-based web scraper with Selenium and async NoDriver support
A new browser scraping module has been added to the researcher, introducing two distinct scraping backends: a synchronous \BrowserScraper\ powered by Selenium (supporting Chrome, Firefox, and Safari) and an asynchronous \NoDriverScraper\ built on zendriver. The Selenium scraper handles cookie management and headless execution, while the zendriver-based scraper introduces advanced concurrency features including browser load balancing, domain-specific rate limiting, and a tab-based browsing mode to improve performance. Both scrapers support PDF extraction (via PyMuPDF and Arxiv), image extraction, and title extraction, providing a more robust and flexible foundation for web research tasks.
_gpt\researcher/scraper/browser · high confidence
New chat agent with optional web search and RAG fallback
The backend now includes a new chat agent (backend/chat) that processes chat messages using the configured LLM provider and supports tool-based web search via Tavily. Web search is optional and disabled if the TAVILY\_API\_KEY is not set or the Tavily package is missing. The agent also implements Retrieval-Augmented Generation (RAG) by chunking report documents into a vector store; if embedding setup fails, it gracefully falls back to using the full report without RAG to prevent chat breakage.
backend/chat · high confidence
New detailed report generation with STORM-inspired architecture
Introduces a new DetailedReport class that generates long, structured reports using an architecture inspired by the STORM paper. The process involves an initial research phase, automatic subtopic generation, and the creation of individual subtopic reports that avoid duplicating previously written content. The final output combines an introduction, a table of contents, and the unique subtopic sections. This implementation supports passing query domains, tone settings, and MCP configurations to ensure consistent behavior across the research workflow.
_backend/report\_type/detailed\report · high confidence
New document loading infrastructure with Azure and LangChain support
The document processing module has been restructured to introduce three new loader classes: \DocumentLoader\ for local files, \OnlineDocumentLoader\ for remote URLs, and \LangChainDocumentLoader\ for pre-loaded LangChain documents. \OnlineDocumentLoader\ now normalizes URL extension casing to ensure correct file-type detection and includes a User-Agent header during downloads. A new \AzureDocumentLoader\ has been added to fetch and process documents from Azure Blob Storage, featuring path validation to prevent directory traversal. The local \DocumentLoader\ has been expanded to support additional formats including EPUB, HTML/HTM (via BSHTMLLoader), and various Unstructured loaders for CSV, Excel, and Markdown.
_gpt\researcher/document · high confidence
New evaluation framework for factuality and hallucination detection
Added a new \evals\ directory containing tools to assess GPT-Researcher's performance. This includes a \simple\_evals\ module for measuring short-form factual accuracy using a zero-shot, chain-of-thought approach adapted from OpenAI's SimpleQA, and a \hallucination\_eval\ module for detecting non-factual content by comparing reports against source materials. The framework tracks metrics such as accuracy, F1 score, cost, and latency, and outputs structured JSON logs for trend analysis.
evals · high confidence
New hallucination evaluation tool for model outputs
A new evaluation suite has been added to assess model responses for hallucinations. The \evaluate.py\ script introduces a \HallucinationEvaluator\ class that utilizes the \HaluEvalDocumentSummaryNonFactual\ judge to compare model-generated summaries against source texts, returning a hallucination score and reasoning. This is supported by a multi-domain dataset of search queries in \inputs/search\_queries.jsonl\ covering areas such as healthcare, legal, history, climate, finance, education, and science, along with sample aggregate results in \results/aggregate\_results.json\.
_evals/hallucination\eval · high confidence
New helper utilities for object diffing, host resolution, and secure markdown rendering
The frontend now includes three new helper modules in the \frontend/nextjs/helpers\ directory. \findDifferences.ts\ provides a utility to recursively compare two objects and return their differences, supporting nested objects, arrays, and primitive values. \getHost.ts\ centralizes backend URL resolution, prioritizing \localStorage\, URL parameters, and environment variables (\NEXT\_PUBLIC\_GPTR\_API\_URL\, \REACT\_APP\_GPTR\_API\_URL\) before falling back to the current window host or a default localhost port. \markdownHelper.ts\ introduces secure markdown-to-HTML conversion using \remark\ and \DOMPurify\, fixing list item formatting issues and ensuring all links open in new tabs with proper security attributes, which is critical for rendering untrusted content like LLM outputs safely.
frontend/nextjs/helpers · high confidence
New image gallery and modal viewer components
Added ImagesAlbum and ImageModal components to the frontend, enabling users to view images in a responsive grid layout with a full-screen modal viewer. The modal supports keyboard navigation (arrow keys, Escape) and mobile swipe gestures, while the album automatically filters out broken images to prevent display errors.
frontend/nextjs/components/Images · high confidence
New image generation providers for Gemini and ModelsLab
Users can now generate images within research workflows using two new providers: Google's Gemini/Imagen models (via the \google-genai\ SDK) and ModelsLab's API (supporting Flux, SDXL, and other community models). The \ImageGeneratorProvider\ handles Google models with automatic landscape cropping and style-aware prompt enhancement, while \ModelsLabImageGeneratorProvider\ supports asynchronous generation with polling. Both providers save images to the \outputs/images/\ directory and require their respective API keys (\GOOGLE\_API\_KEY\/\GEMINI\_API\_KEY\ or \MODELSLAB\_API\_KEY\) to be set.
_gpt\_researcher/llm\provider/image · high confidence
New layout components and PDF styling for research interface
This change introduces new layout structures for the frontend and styling for backend PDF generation. In the frontend, three new layout components are added: \CopilotLayout\ for the copilot mode, \MobileLayout\ which includes a dedicated mobile header with history and settings toggles, and \ResearchPageLayout\ which adds a scroll-to-bottom button for long research results. On the backend, a new \pdf\_styles.css\ file defines the visual appearance of generated PDFs, applying the 'Libre Baskerville' serif font and academic formatting to headers, paragraphs, and tables.
backend/styles, frontend/nextjs/components/layouts · high confidence
New memory state schemas for draft and research workflows
The multi-agent system now exposes explicit state definitions for the drafting and research phases via the \multi\_agents.memory\ module. \DraftState\ tracks the current draft, review feedback, and revision counts, while \ResearchState\ manages research data, human feedback, report structure (including headers, table of contents, and sections), and fact-checking details. These schemas provide the structured data backbone for the new FactChecker and Visualizer agents and the human-in-the-loop workflow.
_multi\agents/memory · high confidence
New modular backend server with agent discovery and local report persistence
The backend/server directory has been restructured into a new FastAPI application (app.py) that serves the frontend static files and exposes REST and WebSocket endpoints for research, chat, and report management. A new agent discovery endpoint (/.well-known/agent-discovery.json) advertises available services (research, reports, chat, research\_stream) to external agents. Report history is now persisted locally via a JSON-based ReportStore instead of MongoDB, and structured JSON logging is introduced for research events. Multi-agent execution is supported via a dedicated runner that resolves the implementation from either the standard multi\_agents module or the nested ag2 variant.
backend/server · high confidence
New research history, analytics, and WebSocket management hooks
This change introduces several new React hooks in the frontend Next.js application to enhance user experience and observability. The \useResearchHistory\ hook and its associated context (\ResearchHistoryContext\) provide a unified way to manage, save, and retrieve research reports, syncing local storage with the server backend. A new \useAnalytics\ hook integrates Google Analytics (react-ga4) to track research queries and report submissions. Additionally, \useWebSocket\ establishes and maintains real-time connections for live report generation, including a heartbeat mechanism to keep connections alive, while \useScrollHandler\ manages automatic scrolling and a 'scroll to bottom' button for long content.
frontend/nextjs/hooks · high confidence
New static frontend with real-time WebSocket monitoring and conversation history
The frontend now includes a new static HTML/JS interface (served via FastAPI) that replaces or supplements previous UI implementations. This interface features a dedicated WebSocket status panel that displays connection state, research activity, duration, and message counts, allowing users to monitor backend connectivity in real-time. It also introduces a conversation history panel with search, sort, and import/export capabilities, persisting research sessions across page reloads. The UI has been restyled with a dark theme, animated gradients, and Font Awesome icons, and supports new report types such as 'Deep Research' and 'Resource Report'.
frontend · high confidence
New utility module for multi-agent file handling, LLM calls, and output formatting
A new \multi\_agents/agents/utils\ package has been introduced to centralize shared functionality for the multi-agent workflow. This includes robust file format conversion utilities (\file\_formats.py\) that asynchronously write reports to Markdown, PDF (with a new academic-style CSS theme), and DOCX, ensuring UTF-8 compliance and fixing previous extension bugs. The module also provides a standardized interface for calling LLMs (\llms.py\) with JSON parsing support, strict sentinel parsers (\none\_sentinels.py\) to accurately interpret 'None' or 'no' responses from reviewers and humans without false positives, and colored console output helpers (\views.py\) for distinct agent identification. Additionally, a filename sanitization utility (\utils.py\) ensures cross-platform compatibility for generated files.
_multi\agents/agents/utils · high confidence
New utility modules for cost tracking, structured logging, and URL security
The \gpt\_researcher/utils\ package has been populated with several new modules that enhance operational visibility and safety. \costs.py\ introduces LLM cost estimation utilities, including specific pricing tables for OpenAI and Anthropic models and logic to extract usage metadata (such as cache tokens) for accurate billing. \logger.py\ and \logging\_config.py\ provide a structured logging system with colorized console output and JSON-based research event tracking to \logs/\. \url\_security.py\ adds SSRF and local-file-read protections by validating that URLs resolve to public IP addresses, blocking private or loopback targets unless explicitly allowed. Additionally, \rate\_limiter.py\ implements a global singleton rate limiter to coordinate request frequency across all worker pools, and \enum.py\ defines configuration enumerations for report types, sources, tones, and prompt families.
_gpt\researcher/utils · high confidence
Next.js frontend initialization with PWA and library build support
The Next.js application in the frontend/nextjs directory has been initialized with a production-ready configuration, including a multi-stage Dockerfile for containerization and a Rollup-based build pipeline to package the UI as a reusable React component library (gpt-researcher-ui). The frontend now supports Progressive Web App (PWA) capabilities via the next-pwa plugin and integrates Google Analytics (react-ga4) for usage tracking. Styling is standardized using Tailwind CSS with custom theme extensions, and the build process is configured to handle TypeScript, React, and CSS injection for the library distribution.
frontend/nextjs · high confidence
Next.js frontend with API proxying and research history management
The application has been migrated to a Next.js frontend structure located in frontend/nextjs/app. This introduces a new API layer (app/api) that proxies chat and report requests to the backend service, handling report retrieval, updates, and deletion. The main page and research detail pages now utilize a ResearchHistoryContext to manage local storage of research data and chat messages, enabling offline access and persistence. The UI includes a new sidebar for research history, mobile-responsive layouts, and PWA capabilities with Google Analytics integration.
frontend/nextjs/app · high confidence
Redesigned Settings modal with new configuration options
The Settings interface has been completely redesigned into a modal experience, now labeled 'Preferences' in the UI. This update introduces several new configuration capabilities: users can now select a 'Layout Type' (switching between 'Research' and 'Copilot' modes), choose from an expanded list of content 'Tone' options (such as Objective, Formal, Analytical, etc.), and manage 'MCP' (Model Context Protocol) server configurations via a JSON editor with validation. Additionally, a new 'FileUpload' component allows users to drag-and-drop, view, and delete files directly within the settings, while the underlying chat and report logic has been consolidated into a new ChatBox component.
frontend/nextjs/components/Settings · high confidence
Unified embedding provider management via new Memory class
The \gpt\_researcher/memory\ module now provides a centralized \Memory\ class that manages embedding generation across a wide range of providers, including OpenAI, Azure OpenAI, Cohere, Google Vertex AI, Google Generative AI, Ollama, HuggingFace, AWS Bedrock, DashScope, and several others (GigaChat, Netmind, MiniMax, Nebius, OpenRouter, etc.). This change introduces a unified interface for selecting and configuring embedding models via environment variables and constructor arguments, replacing previous fragmented or hardcoded provider logic with a single, extensible entry point for document similarity and retrieval.
_gpt\researcher/memory · high confidence
Web base scraper now extracts page title alongside content and images
The web base loader scraper has been updated to return the page title in addition to the extracted content and image URLs. When scraping a webpage, the system now uses a utility function to parse the HTML and extract the title, ensuring this metadata is available to downstream processes. This change enhances the richness of the scraped data by providing context about the source page.
_gpt\_researcher/scraper/web\_base\loader · high confidence
Architecture
Research workflow restructured into modular skill classes
The monolithic research logic has been refactored into a set of dedicated, modular skill classes within the \gpt\_researcher/skills\ package. This change introduces \ResearchConductor\ to manage the overall research orchestration and query planning, \BrowserManager\ to handle web scraping and image selection, \ContextManager\ for content retrieval and compression, \SourceCurator\ for ranking and filtering sources, \ReportGenerator\ for writing the final report, and \ImageGenerator\ for creating contextually relevant visuals. This modularization improves code clarity and maintainability while preserving the existing research capabilities.
_gpt\researcher/skills · high confidence
Behavioural changes
Automatic image URL rewriting for npm package compatibility
A new image transformation utility has been added to the frontend utilities to ensure images render correctly when the code is consumed as an npm package. This plugin automatically rewrites local image paths (such as \/img/...\ or \img/...\) to absolute URLs pointing to \https://gptr.app/img/...\, resolving issues where relative paths would otherwise break in external package contexts.
frontend/nextjs/src/utils · high confidence
Exa Search retriever now supports domain filtering and handles malformed responses
The Exa retriever has been updated to accept a \query\_domains\ parameter, which is passed to the Exa API via \include\_domains\ to restrict search results to specific domains. Additionally, the implementation now robustly handles malformed or incomplete API responses by safely extracting URLs and text/summary fields using \getattr\, skipping any result objects that lack a valid URL or ID, and returning an empty list on API errors instead of crashing.
_gpt\researcher/retrievers/exa · high confidence
GPT Researcher agent restructured with new capabilities and configuration
The GPT Researcher module has been reorganized into a new package structure (gpt\_researcher) with a refactored main agent class. This update introduces support for Model Context Protocol (MCP) integrations via configurable server connections, enables deep research workflows, and adds image generation capabilities for reports. Users can now restrict searches to specific domains, complement source URLs with web searches, and utilize a new prompt family system for better model-specific formatting. The agent also tracks per-step costs and supports UTF-8 encoding for output files.
_gpt\researcher · high confidence
Langgraph research component initialization with empty authentication token
The Langgraph research component in the frontend Next.js application has been introduced to facilitate AI-driven research tasks. This component initializes a Langchain LangGraph client to search for assistants and stream run responses based on user queries. Notably, the component currently includes a hardcoded, empty string for the API authentication token ('X-Api-Key'), meaning the client is configured without valid credentials for the specified host URL.
frontend/nextjs/components/Langgraph · medium confidence
Multi-agent research workflow with bounded revision loops and new agent roles
The multi-agent system now uses a LangGraph-based workflow orchestrated by ChiefEditorAgent, introducing FactCheckerAgent and VisualizerAgent alongside existing roles. To prevent infinite loops, the system enforces configurable ceilings for plan, draft, and fact-check revisions (defaulting to 3 rounds); once a ceiling is exceeded, the workflow force-accepts and proceeds. Human-in-the-loop feedback is supported via WebSocket or console, and the PublisherAgent can output final reports as PDF, DOCX, or Markdown.
_multi\agents/agents · high confidence
New dedicated styling for rendered markdown content
A new CSS file (markdown.css) has been introduced to define the visual appearance of markdown content within the Next.js frontend. This stylesheet establishes a consistent theme for reports and other text-based outputs, specifying typography (Georgia serif, 18px base), spacing, and formatting for elements like headings, paragraphs, links, code blocks, and blockquotes. It also includes specific styling for tables (with dark theme matching, alternating row colors, and hover effects), images (including special handling for generated illustrations/diagrams), and task lists, ensuring that rendered markdown aligns with the application's dark UI design.
frontend/nextjs/styles · high confidence
New modular configuration system with deep research and image generation support
The configuration system has been restructured into a modular setup (base, default, and local overrides) to support new capabilities and finer control. Users can now configure Deep Research parameters (breadth, depth, concurrency) and enable optional inline image generation via Google or ModelsLab providers. The default models have been updated to the GPT-5.4 family, and new settings allow for MCP server integration, reasoning effort tuning, and configurable rate limiting for scrapers.
_gpt\researcher/config/variables · high confidence
New modular configuration system with multi-retriever and LLM separation
The configuration module has been refactored into a new structure (\gpt\_researcher/config\) that introduces a \Config\ class for centralized management of settings. This change allows users to configure multiple retrievers simultaneously via the \RETRIEVER\ environment variable (defaulting to 'tavily') instead of a single provider. It also separates LLM settings into distinct \fast\_llm\ and \smart\_llm\ configurations, enabling different models for different research steps. The system now supports loading custom JSON configuration files via \CONFIG\_PATH\, merges them with defaults, and handles deprecated environment variables (like \LLM\_PROVIDER\) with warnings.
_gpt\researcher/config · high confidence
New research interface components with premium visual styling
The research block UI now includes new components for the input experience and result display. InputArea and ChatInput provide text entry with auto-resizing textareas, debounced input handling, and a consistent 'premium' aesthetic featuring teal-to-cyan gradient glows and backdrop blurs. SubQuestions renders interactive suggestion chips for follow-up angles. LogMessage processes and displays structured logs, including markdown rendering for 'differences' and accordion views. SourceCard displays source attribution with domain extraction and favicon fallbacks.
frontend/nextjs/components/ResearchBlocks/elements · high confidence
New retriever architecture with explicit scraping requirements
The retriever module has been restructured to introduce a base contract that explicitly declares whether a search provider returns URLs needing further scraping or already-fetched content. This replaces a previous heuristic that incorrectly inferred scraping needs based on snippet length, which caused citation failures when snippets were long. Users benefit from more reliable report generation and verifiable citations, as the system now correctly handles full-text providers like PubMed Central and Semantic Scholar without manual workarounds. Additionally, the module now supports a significantly expanded set of search engines, including Arxiv, Bing, Brave, Exa, OpenAlex, and several others, all registered in a centralized init file.
_gpt\researcher/retrievers · high confidence
New scraping backends and improved content filtering
The scraper module now supports Tavily Extract and FireCrawl as additional scraping backends, with automatic installation of their required packages if missing. It also introduces URL deduplication to prevent redundant scraping, enhanced image relevance scoring based on size and CSS classes, and robust detection of anti-bot block pages and unextracted PDF content to ensure only valid article text is ingested.
_gpt\researcher/scraper · high confidence
Normalize DuckDuckGo search results and fix snippet length handling
The DuckDuckGo retriever now normalizes search results to the shared \{href, body, title}\ contract, ensuring compatibility with downstream research steps that expect \href\ keys regardless of whether the underlying \ddgs\ library returns \link\, \url\, \body\, \snippet\, or \description\. Additionally, the retriever explicitly declares that results require scraping (\requires\_scraping = True\) and caps the prefetched \body\ snippet at 100 characters to prevent the system from incorrectly treating standard search snippets as full article text, which previously caused citation failures.
_gpt\researcher/retrievers/duckduckgo · high confidence
Redesigned Next.js frontend with mobile-first UI and new component structure
The frontend has been completely rebuilt using Next.js components, introducing a modern, premium visual design with animated particle backgrounds, gradient glows, and responsive layouts. New components include a sticky Header with stop/new research controls, a Hero section for input, a collapsible Research Sidebar for history, and a Footer with social links. A dedicated mobile experience is provided via MobileHomeScreen, MobileResearchContent, and MobileChatPanel, ensuring full functionality on smaller screens. The ResearchResults component now intelligently filters data to show only the final report block, while HumanFeedback and LoadingDots enhance the interaction flow.
frontend/nextjs/components · high confidence
Redesigned research report interface with new interactive components
The research report view has been completely redesigned to offer a more premium, interactive experience. Users can now access reports in PDF, DocX, and JSON formats via the new AccessReport component, which handles file path resolution. The interface includes a ChatInterface for asking questions about the report content, with real-time markdown rendering and loading states. Chat responses (ChatResponse) now display web search sources when available, and the main report (Report) features a copy-to-clipboard function. Additional components like ImageSection, LogsSection, and Sources provide structured views for related images, agent work logs, and source citations, all styled with a consistent dark theme and backdrop blur effects.
frontend/nextjs/components/ResearchBlocks · high confidence
Restructured frontend layout and data processing logic
The frontend now uses a new \getAppropriateLayout\ utility to dynamically select between Mobile, Copilot, and Research page layouts based on screen size and user settings, replacing the previous static layout approach. Additionally, data processing has been modularized into \dataProcessing.ts\ and \consolidateBlocks.ts\, which now handle the aggregation of report content, deduplication of source URLs, and consolidation of image blocks before rendering, ensuring a cleaner and more consistent presentation of research results.
frontend/nextjs/utils · high confidence
Restructured research actions into a modular, robust action package
The \gpt\researcher/actions\ module has been reorganized into a dedicated package with an explicit \\\init\\_.py\ that centralizes exports for retrievers, query processing, agent creation, web scraping, report generation, and markdown processing. This restructuring introduces several behavioral improvements: agent selection and sub-query generation now use \json\_repair\ and a greedy regex fallback to reliably extract JSON from LLM responses, preventing crashes on malformed output; the retriever factory now strips whitespace from header-supplied retriever lists to ensure correct resolution; and report references are now emitted in a deterministic, sorted order. Additionally, the report introduction and conclusion generation now respect the configured language setting, and the web scraper integrates a \WorkerPool\ to manage concurrency and properly closes HTTP sessions.
_gpt\researcher/actions · high confidence
Fixes
ArXiv scraper now uses the official Client API to fix compatibility and improve data extraction
The ArXiv scraper has been refactored to use the maintained \arxiv\ Client API instead of the deprecated \Search.results()\ method, resolving breakage with arxiv library versions 2.2 and later. This change also adds robustness by gracefully handling empty or malformed results without raising errors, and enhances the scraped context by including the paper's publication date and author list alongside the abstract.
_gpt\researcher/scraper/arxiv · high confidence
Fix Google Search retriever robustness and query encoding
The Google Search retriever now properly URL-encodes query parameters to prevent malformed requests when search terms contain special characters (e.g., '&', '\#'). It also guards against malformed API responses by safely parsing JSON and filtering out non-dict items or missing fields, preventing uncaught exceptions. Additionally, the retriever supports domain filtering via the \query\_domains\ parameter and respects the \max\_results\ limit when returning results.
_gpt\researcher/retrievers/google · high confidence
PyMuPDF scraper now extracts content from all pages and handles SSL/temp-file edge cases
The PyMuPDF scraper has been refactored to read and return content from all pages in a PDF document, rather than just the first page, ensuring that documents with cover pages or multi-page structures are fully processed. Additionally, the scraper now includes an SSL verification fallback (retrying without verification on SSL errors) and guarantees that temporary downloaded files are cleaned up even if PDF loading fails, improving reliability and preventing disk clutter.
_gpt\researcher/scraper/pymupdf · high confidence
Test coverage
Added backend tests for report chat routing and PDF export filename handling; Added tests for deep research termination and Tavily MCP deduplication; Expanded test coverage for core research and agent capabilities.
Dependencies
Initial dependency manifests for GPT Researcher ecosystem
This change introduces the foundational dependency configuration files for the project, establishing the required libraries for the backend, frontend, and various agent modules. The root \requirements.txt\ and \pyproject.toml\ define the core Python environment, mandating Python 3.11+ and upgrading to LangChain v1 (including \langchain\, \langchain-openai\, and \langgraph\), while also addressing security by pinning \brotli\ to version 1.2.0 or higher to mitigate [CVE redacted]. The \frontend/nextjs/package.json\ sets up the React-based UI with Next.js 14 and Tailwind CSS, and separate manifests are added for the Discord bot, npm client, multi-agent systems, and evaluation suites to isolate their specific tooling needs.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 52.
Lenses
- Code Health 73
- Architecture 98
- Maturity 89
- Readiness 44
- Security 61
- Accessibility 45
Changes since last survey
- 300 commits — 158 feature/other, 142 fixes
By area
- (repo) — 107 commits
- (root) — 30 commits
- gpt_researcher/retrievers — 30 commits
- docs/docs — 21 commits
- gpt_researcher/scraper — 16 commits
- gpt_researcher/skills — 12 commits
- backend/server — 9 commits
- multi_agents/agents — 9 commits
- gpt_researcher/utils — 7 commits
- gpt_researcher/actions — 6 commits
- backend/report_type — 5 commits
- gpt_researcher/config — 4 commits
- gpt_researcher/context — 4 commits
- .github/workflows — 3 commits
- gpt_researcher/agent.py — 3 commits
- gpt_researcher/llm_provider — 3 commits
- .claude/references — 2 commits
- .claude/skills — 2 commits
- backend/utils.py — 2 commits
- deep_agents/BENCHMARK.md — 2 commits
Notable commits
- fix: "Fix: README paths -> gpt_researcher/config/variables/default.py"
- fix: #1673: Fixed the reference error in code, added 'strategic_llm' variable instead of 'model_name'.
- fix: Fix Azure blob loader path handling
- fix: Fix badge label in README.md
- fix: Fix config path from env var not loaded
- fix: Fix for issue 1718
- fix: Fix issue 1712
- fix: Fix missing os import in websocket manager
- fix: Fix silently dropped token limits, real usage-based cost tracking, and research pipeline robustness (#1861)
- fix: Fix silently dropped token limits, real usage-based costs, and research pipeline robustness
- fix: Fix: resolve unhashable dict error in detailed report context deduplication
- fix: Merge branch 'assafelovic:main' into fix/hash-mcp-context
- fix: Merge branch 'main' into fix-dict-unhashable-bug
- fix: Merge branch 'main' into fix/hash-mcp-context
- fix: Merge pull request #1589 from kga245/claude/fix-automated-checks-V08GV
- fix: Merge pull request #1594 from PriscaAmajuoyi/fix/reports-persistence
- fix: Merge pull request #1595 from PriscaAmajuoyi/fix/cors-origins
- fix: Merge pull request #1607 from Carton/fixes
- fix: Merge pull request #1623 from MattBenesch/fix/pymupdf-read-all-pages
- fix: Merge pull request #1637 from technot80/fix/deep-research-context-type-error
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
assafelovic/gpt-researcher was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 6f998577d547b1e54ec662dac63583aa11e3b84b — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.