ScrapeGraphAI/Scrapegraph-ai
59.5
Adequate · 18 September 2026
16.4k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a Python-based web scraping framework that constructs and executes data extraction pipelines using directed acyclic graphs. It leverages large language models to dynamically generate scraping logic from natural language prompts and supports a wide variety of data sources, including HTML, CSV, JSON, and images via OCR. The architecture provides modular components for fetching content, processing data, and integrating with external services like Burr for stateful execution.
Features
Add telemetry module for tracking graph executions
A new telemetry module has been introduced to scrapegraphai, enabling the collection and sending of usage data. The module allows users to disable tracking via the SCRAPEGRAPHAI\_TELEMETRY\_ENABLED environment variable or a local config file, generates a persistent anonymous ID, and sends execution details (such as prompts, schemas, and LLM models) to a custom tracing endpoint in the background to minimize performance impact.
scrapegraphai/telemetry · high confidence
Added dedicated tokenizers for Mistral, Ollama, and OpenAI models
The tokenization utilities in \scrapegraphai/utils/tokenizers\ now include specific implementations for Mistral, Ollama, and OpenAI models. Users can now accurately estimate token counts for text processed by these providers: Mistral tokenization uses the \mistral\_common\ library, Ollama leverages the model's built-in \get\_num\_tokens\ method, and OpenAI tokenization utilizes \tiktoken\ with the \gpt-4o\ encoding.
scrapegraphai/utils/tokenizers · high confidence
Initialize scrapegraphai package with logging setup
The scrapegraphai package now initializes with a dedicated \_\init\\_.py file that imports and configures the logging utilities. This ensures that the application logger is available and set to info verbosity by default upon import, providing immediate visibility into application events without requiring manual configuration.
scrapegraphai · high confidence
Introduction of new graph nodes and restructured node module
The \scrapegraphai/nodes\ module has been restructured with a new \\_\init\\_.py\ that explicitly exports a comprehensive set of node classes, including \BatchGenerateAnswerNode\ for cost-efficient OpenAI batch processing, \FetchNodeLevelK\ for recursive multi-level link fetching, \FetchScreenNode\ for screenshot capture, and specialized answer generators like \GenerateAnswerCSVNode\ and \GenerateAnswerFromImageNode\. The \BaseNode\ class now enforces a stricter node type system ('node' or 'conditional\_node') and includes a centralized logger, while the \ConditionalNode\ has been updated to support custom boolean condition evaluation via \simple\_eval\ for more flexible graph branching logic.
scrapegraphai/nodes · high confidence
New GraphBuilder for natural language scraping graph generation
A new GraphBuilder class has been introduced in the builders module, allowing users to generate web scraping graph configurations from natural language prompts. This component leverages an LLM (supporting OpenAI, Gemini, and Ernie models) to parse user intent and output a structured graph definition, which can then be visualized using Graphviz. This provides a programmatic way to construct scraping pipelines dynamically based on user requirements.
scrapegraphai/builders · high confidence
New LLM and AI model integrations added to the models module
The models module now exposes wrappers for several new AI providers and capabilities, allowing users to connect to Atlas Cloud, CLōD, DeepSeek, MiniMax, OneAPI, xAI (Grok), and NVIDIA (via langchain\_nvidia\_ai\_endpoints). Additionally, new classes for OpenAI Image-to-Text and OpenAI Text-to-Speech have been added, expanding the types of content the application can process and generate.
scrapegraphai/models · high confidence
New document loaders and browser integration options
Users can now scrape web pages using three new backend options: PlasmateLoader for lightweight, low-memory extraction without Chrome; ChromiumLoader with support for undetected\_chromedriver and Firefox backends; and direct integrations with BrowserBase and Scrape.do APIs. The module also introduces lazy loading for Chromium and Plasmate to prevent DLL crashes on Windows and adds configurable timeouts and retry limits.
scrapegraphai/docloaders · high confidence
New example scripts for Code, CSV, Custom, Depth Search, and Document Scrapers
Added new example directories and scripts for Code Generator, CSV Scraper, Custom Graph, Depth Search, and Document Scraper graphs, each including OpenAI and Ollama implementations, configuration files, and documentation.
examples · high confidence
New helper module for model tokens, node metadata, and graph schemas
A new \scrapegraphai.helpers\ package has been introduced to centralize configuration data. It exports \models\_tokens\ (mapping providers like OpenAI, Azure, Google, and Ollama to their specific model names and token limits), \nodes\_metadata\ (descriptions and I/O signatures for graph nodes such as Search, Fetch, and RAG), \robots\_dictionary\ (mapping models to associated web crawlers), and \graph\_schema\ (a JSON schema defining the structure of graph configurations). This change provides a single source of truth for model capabilities and graph structure definitions.
scrapegraphai/helpers · high confidence
New integration modules for Burr, Indexify, and scrapegraph-py SDK compatibility
The \scrapegraphai/integrations\ package now exposes three new capabilities: a \BurrBridge\ class that wraps ScrapeGraphAI graphs into Burr applications for stateful execution and tracking; an \IndexifyNode\ for indexing content within the graph state; and a compatibility layer (\scrapegraph\_py\_compat\) that abstracts differences between scrapegraph-py SDK v2 and v3, allowing users to call extract, scrape, and search endpoints without managing version-specific imports.
scrapegraphai/integrations · high confidence
New screenshot scraping and OCR utilities
Added a new \screenshot\_scraping\ module providing utilities to capture webpage screenshots via Playwright and extract text from images using the Surya OCR library. The module exposes functions to take screenshots (\take\_screenshot\), manually select image areas for processing via OpenCV or ipywidgets (\select\_area\_with\_opencv\, \select\_area\_with\_ipywidget\), and detect text (\detect\_text\). All optional dependencies (PIL, OpenCV, ipywidgets, Surya) are imported dynamically to avoid hard requirements, with clear error messages guiding users to install the \\[ocr\]\ extra.
_scrapegraphai/utils/screenshot\scraping · high confidence
New utility modules for code error handling, HTML cleanup, and LLM cost tracking
The \scrapegraphai/utils\ package has been restructured with a new \\_\init\\_.py\ that exposes a comprehensive set of utility functions. This includes new modules for analyzing and correcting code errors (syntax, execution, validation, and semantic), cleaning and minifying HTML content, and exporting data to CSV, JSON, and XML formats. Additionally, the package now provides robust LLM cost tracking via \CustomLLMCallbackManager\ and \CustomCallbackHandler\, supports OpenAI Batch API operations for cost savings, and includes utilities for proxy rotation, screenshot preparation, and text tokenization.
scrapegraphai/utils · high confidence
Repository initialization with development tooling and documentation
The repository is initialized with essential configuration files including a \.gitignore\ for Python artifacts, a \.pre-commit-config.yaml\ enforcing code style via Black, Ruff, and isort, and a \.releaserc.yml\ for semantic-release automation. It also includes \AGENTS.md\ providing contribution guidelines for AI coding agents, a \Makefile\ for standardizing linting, type-checking, and testing workflows, and a \Dockerfile\ for containerized deployment.
(repo-wide) · high confidence
Behavioural changes
Centralized prompt template management
The \scrapegraphai/prompts\ package has been restructured to centralize all LLM prompt templates. A new \\_\init\\_.py\ file now exports a comprehensive set of template constants (e.g., \TEMPLATE\_CHUNKS\, \TEMPLATE\_REASONING\, \TEMPLATE\_ROBOT\) from dedicated modules for specific graph nodes such as answer generation, code generation, search, and reasoning. This change organizes the prompt definitions into a single, importable namespace, making it easier to manage and access the instructions used by the scraping pipeline.
scrapegraphai/prompts · high confidence
ScrapeGraphAI v2 API surface and graph architecture overhaul
The \scrapegraphai/graphs\ module has been completely rewritten to align with the scrapegraph-py v2 API. This change introduces a new \AbstractGraph\ and \BaseGraph\ foundation that standardizes how scraping pipelines are constructed and executed. Users will now interact with a unified set of graph classes (such as \SmartScraperGraph\, \CodeGeneratorGraph\, and \DocumentScraperGraph\) that support Pydantic output schemas, configurable LLM providers via a standardized config dict, and new capabilities like multi-URL iteration (\GraphIteratorNode\) and Burr integration. The previous graph implementations have been replaced by this new modular structure, which centralizes common parameters (like \verbose\, \headless\, and \loader\_kwargs\) and enforces a consistent execution flow across all scraper types.
scrapegraphai/graphs · high confidence
Test coverage
Added comprehensive test suite for graph components and LLM integrations; Added integration tests and test fixtures for ScrapeGraphAI; Added sample test input files for XML, JSON, HTML, and CSV formats; Added test coverage for utility modules; Added unit tests for Fetch, Robots, SearchInternet, and SearchLink nodes; Comprehensive test infrastructure and coverage for ScrapeGraphAI.
Dependencies
Initial project setup with Python 3.12 and LangChain v1 dependencies
The project now uses a pyproject.toml manifest, establishing a minimum Python version of 3.12 and defining core dependencies including LangChain 1.2.0, LangChain Classic 1.0.0, and scrapegraph-py 2.0.0. Optional features are configured for Burr, NVIDIA AI endpoints, and OCR, while the build system is set to hatchling 1.26.3.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 60.
Lenses
- Code Health 94
- Architecture 95
- Maturity 55
- Readiness 54
- Security 56
Changes since last survey
- 300 commits — 251 feature/other, 49 fixes
By area
- (root) — 131 commits
- (repo) — 89 commits
- scrapegraphai/helpers — 12 commits
- scrapegraphai/graphs — 11 commits
- scrapegraphai/nodes — 9 commits
- scrapegraphai/utils — 8 commits
- scrapegraphai/models — 5 commits
- scrapegraphai/telemetry — 5 commits
- .github/workflows — 4 commits
- docs/assets — 4 commits
- scrapegraphai/docloaders — 4 commits
- examples/markdownify — 3 commits
- tests/graphs — 3 commits
- docs/timeout_configuration.md — 2 commits
- tests/test_chromium.py — 2 commits
- PullRequests/PR_1027_reviews.md — 1 commit
- docs/chinese.md — 1 commit
- docs/korean.md — 1 commit
- docs/source — 1 commit
- examples/smart_scraper_graph — 1 commit
Notable commits
- fix: Add Italian README translation and fix outdated links (#1070)
- fix: Add Italian README translation and fix outdated links (#1070) (#1071)
- fix: Add format key to LLM configuration, solve bug.
- fix: Fix E402 errors in smart_scraper_graph.py by moving imports to top
- fix: Fix critical schema transformation bugs and improve logging
- fix: Fix issue: Burr integration by updating fetch_node.py
- fix: Fix langchain import issues blocking tests
- fix: Fix linting issues - remove unused imports and whitespace
- fix: Fixed Issue: Burr integration ParseNode by updating parse_node.py
- fix: Merge pull request #1001 from Mirza-Samad-Ahmed-Baig/fix-schema-transform-bugs
- fix: Merge pull request #1028 from ScrapeGraphAI/copilot/fix-whitespace-formatting-errors
- fix: Merge pull request #1029 from ScrapeGraphAI/copilot/fix-e402-import-issues
- fix: Merge pull request #1033 from majiayu000/fix/langchain-v1-compatibility
- fix: Merge pull request #1053 from Vikrant-Khedkar/fix/replace-print-with-logging
- fix: Merge pull request #1126 from aayushbaluni/fix/1121-expose-model-tokens-fallback
- fix: Merge pull request #1131 from primorLee/codex/fix-relative-markdown-links
- fix: Merge pull request #1132 from Excelius-Wang/docs/fix-timeout-links
- fix: Merge pull request #1135 from ScrapeGraphAI/fix/1102-silent-error-page-detection
- fix: Merge pull request #1136 from ScrapeGraphAI/fix/1102-silent-error-page-detection-main
- fix: Merge pull request #1140 from amirshahzadhashmi7145/fix/1121-add-gemini-2.5-tokens
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ScrapeGraphAI/Scrapegraph-ai was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit c75c8084fae2d4f5ba01a8c218bc1168b67e3569 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.