D4Vinci/Scrapling
67.8
Adequate · 26 September 2026
28.5k
lines of production code
Python
primary language
4
measurements over time
What this system is
Scrapling is a Python web scraping library that provides a unified framework for fetching, parsing, and crawling web content. It supports multiple fetching strategies, including static HTTP requests and browser-based automation with stealth capabilities to bypass anti-bot protections. The system includes a built-in spider engine for managing crawl lifecycles, pausing, and data export, along with templates for common tasks like feed parsing and e-commerce extraction. Additionally, it offers integrations with Scrapy and an MCP server for AI-driven interaction.
How it got here
2024 — Core architecture and engine implementation
11 changes.
This period focused on establishing the foundational architecture of the Scrapling library, introducing a new core module with text handling and storage, and implementing the static HTTP fetching engine. The work also included restructuring dependencies into a modern pyproject.toml format and adding comprehensive test suites for fetchers, parsers, and sessions to ensure robustness.
2025 — Browser engine modernization and test expansion
6 changes.
The project replaced its browser backend with a new stealth-focused engine based on patchright, introducing advanced features like Cloudflare bypass and dynamic session management. This architectural shift was accompanied by a comprehensive restructuring of the fetcher API for better performance and a significant expansion of test coverage across CLI, AI, and core storage modules.
2026 — Spider framework and integrations
6 changes.
This period focused on introducing a new Spider framework featuring pause/resume capabilities, auto-throttling, and structured data export, alongside a comprehensive suite of ready-to-use templates for feeds, e-commerce, and RAG pipelines. The work also established robust Scrapy integration, allowing Scrapling parsers to be used within existing Scrapy projects, and was supported by extensive test coverage for both the new crawling engine and the integration layer.
Features
Add Scrapy integration for using Scrapling parsers in Scrapy spiders
Users can now integrate Scrapling with Scrapy by decorating spider callbacks with \scrapling\_response\ or using \convert\_response\ directly. This allows existing Scrapy projects to receive Scrapling \Response\ objects in their callbacks, enabling the use of Scrapling's full parsing API without altering the spider's crawling logic. The integration is optional and requires Scrapy to be installed separately.
scrapling/integrations · high confidence
Added Scrapling agent skill examples
The agent-skill/Scrapling-Skill/examples directory now includes a complete set of usage examples for the Scrapling library, accompanied by a README. These examples demonstrate four distinct scraping approaches: FetcherSession for fast, persistent HTTP requests; DynamicSession for JavaScript-heavy pages using Playwright; StealthySession for bypassing anti-bot protections like Cloudflare; and Spider for automated multi-page crawling with structured data export. The examples target quotes.toscrape.com and serve as a practical guide for users to integrate Scrapling into their workflows.
agent-skill/Scrapling-Skill/examples · high confidence
Introduce new Spider framework with pause/resume, auto-throttling, and export capabilities
The \scrapling/spiders\ package introduces a new crawling system that supports pausing and resuming crawls via disk checkpoints, automatically throttling request rates based on server response times, and exporting collected items to JSON, CSV, or XML formats. The engine now respects \robots.txt\ directives (with configurable compliance), caches HTTP responses for development replay, and includes specialized spider templates for sitemaps, Shopify, and feed parsing.
scrapling/spiders · high confidence
Introduction of the Scrapling engines module with static HTTP fetching capabilities
This change introduces the new \scrapling/engines\ package, establishing the foundational structure for the library's fetching logic. It adds a \static.py\ module implementing a \\_SyncSessionLogic\ class that leverages \curl\cffi\ for HTTP requests, supporting features like browser impersonation, stealth headers, proxy rotation, and HTTP/3. The module includes \constants.py\ with optimized Chromium launch arguments for stealth and performance, and an empty \\\init\\_.py\ to define the package namespace. This represents the initial implementation of the static engine component within the broader engine architecture.
scrapling/engines · high confidence
New spider templates for feeds, e-commerce, sitemaps, and RAG pipelines
The \scrapling/spiders/templates\ module now provides several ready-to-use spider classes to simplify common scraping tasks. Users can parse XML and CSV feeds with \XMLFeedSpider\ and \CSVFeedSpider\ (which handle gzip decompression and namespace stripping), automatically follow link rules with \CrawlSpider\ and \CrawlRule\, or seed crawls from sitemaps/robots.txt using \SitemapSpider\. A dedicated \ShopifySpider\ extracts product data via Shopify's JSON API, while \SiteToMarkdownSpider\ crawls sites and converts pages to Markdown for LLM/RAG ingestion, including optional file output. These templates are exported from the \scrapling.spiders.templates\ package.
scrapling/spiders/templates · high confidence
New toolbelt utilities for proxy rotation, ad blocking, and response handling
The \scrapling/engines/toolbelt\ package now provides core utilities for the library's fetching engine. Users can leverage the new \ProxyRotator\ class for thread-safe, cyclic proxy rotation across fetchers. The library introduces built-in ad and tracker blocking capabilities via the \ad\_domains.py\ list and navigation interceptors, allowing requests to specific domains or resource types to be dropped. Response handling is standardized through a new \Response\ class and \ResponseFactory\, which unify data from Playwright and other sources, support XHR capture, and expose a \follow()\ helper for spider workflows.
scrapling/engines/toolbelt · high confidence
Scrapling 0.4.15: Lazy imports, CLI shell, and MCP server with auth
Scrapling 0.4.15 introduces lazy loading for top-level imports (Fetcher, Selector, etc.) to reduce startup memory, adds a \scrapling install\ CLI command to manage Playwright dependencies, and includes a new \scrapling mcp\ command that runs an MCP server with optional HTTP transport, bearer-token authentication, and DNS-rebinding protection.
scrapling · high confidence
Behavioural changes
Introduce flexible, context-aware logging and new utility modules
The \scrapling/core/utils\ package has been restructured to support more flexible logging and provide new helper utilities. Logging is now context-aware via \ContextVar\, allowing spider classes to temporarily override the logger using \set\_logger\ and \reset\_logger\ without affecting other concurrent contexts, while retaining a default INFO-level console output. Additionally, new modules \\_shell.py\ and \\_utils.py\ introduce utilities for parsing HTTP headers and cookies, flattening iterables, cleaning whitespace, and converting HTML elements to dictionaries, supporting the core scraping logic with improved memory efficiency and modularity.
scrapling/core/utils · high confidence
New browser-based fetcher engine with stealth and Cloudflare bypass
The browser engine layer has been replaced with a new implementation built on \patchright\ (replacing the previous \rebrowser\ and \Camoufox\ backends). This introduces \DynamicSession\ and \StealthySession\ classes that manage browser tabs via a page pool, offering features such as Cloudflare Turnstile/Interstitial solving, ad blocking, DNS-over-HTTPS, and the ability to block specific domains or disable resources for speed. Users can now configure stealth options (canvas noise, WebRTC blocking, WebGL control), set browser timezones and locales, use persistent contexts with user data directories, and inject custom initialization scripts or page setup actions.
_scrapling/engines/\browsers · high confidence
New core module with shell signatures, text handling, and storage
The library introduces a new \scrapling.core\ package that centralizes core infrastructure. It provides \\_shell\_signatures.py\ to define parameter types for the interactive shell, enabling better autocompletion and introspection. A new \TextHandler\ class extends Python strings to preserve its type through operations like slicing and regex matching, improving developer experience when chaining text manipulations. Additionally, the package includes \SQLiteStorageSystem\ for persistent element storage and a \Translator\ that enhances CSS-to-XPath conversion with support for \::text\ and \::attr()\ pseudo-elements.
scrapling/core · high confidence
Reworked fetcher API with lazy imports and backward compatibility
The fetcher module has been restructured to use lazy imports, improving startup performance by deferring the loading of heavy browser engines until they are actually used. The public API now exposes \Fetcher\, \AsyncFetcher\, \DynamicFetcher\, and \StealthyFetcher\ directly from \scrapling.fetchers\. To ensure backward compatibility, the \custom\_config\ argument is still supported alongside the new \selector\_config\, and \PlayWrightFetcher\ remains an alias for \DynamicFetcher\. Additionally, the \ProxyRotator\ class is now explicitly exposed in the public API.
scrapling/fetchers · high confidence
Test coverage
Added CLI test suite for MCP, shell, and extraction commands; Added async test suite for fetchers and sessions; Added comprehensive test suite for sync fetchers and sessions; Added comprehensive test suite for the spiders subsystem; Added integration tests for Scrapy compatibility; Added test coverage for AI MCP server functionality; Added test coverage for core shell and storage modules; Added test package initialization; Expanded test coverage for fetcher components; Expanded test coverage for parser and selector features.
Dependencies
Project migration to pyproject.toml and dependency restructuring
The project has migrated its build configuration from legacy setup files to a modern pyproject.toml format, establishing version 0.4.15 and raising the minimum Python requirement to 3.10. Core dependencies have been updated, notably replacing tldextract with the tld library. The dependency structure has been reorganized into optional extras to reduce bloat: fetchers (including playwright, patchright, and curl\_cffi), rag, ai (adding MCP v2 support), and shell. Additionally, dedicated requirements files have been introduced for documentation (docs/requirements.txt) and testing (tests/requirements.txt) to isolate tooling dependencies.
(dependencies) · high confidence
Housekeeping
Initial project scaffolding and development tooling
The repository is initialized with essential configuration files for the Scrapling project. This includes a Dockerfile for containerized deployment, a .dockerignore file, and a MANIFEST.in for package distribution. Development workflows are established via a .pre-commit-config.yaml integrating Ruff and Bandit, alongside ruff.toml and .bandit.yml for linting and security checks. Documentation is configured for ReadTheDocs using Zensical, and the project structure includes a CHANGELOG, AI\_POLICY, CODE\_OF\_CONDUCT, and CONTRIBUTING guides to standardize contributions.
(repo-wide) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 48 → 68 (+20.1)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 95 → 87 (-7.4)
- Architecture 96 → 97 (+1.4)
- Maturity 65 → 75 (+9.8)
- Readiness 30 → 58 (+28.6)
- Security 49 → 71 (+21.9)
Resolved (33)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (11 lines × 2) (tests/spiders/test_force_stop_checkpoint.py)
- Duplicated block (17 lines × 2) (scrapling/core/ai.py)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- …and 13 more
New (119)
- AsyncDynamicSession.fetch (cognitive 37) (scrapling/engines/_browsers/_controllers.py)
- AsyncDynamicSession.fetch (cyclomatic 18) (scrapling/engines/_browsers/_controllers.py)
- AsyncStealthySession._cloudflare_solver (cognitive 64) (scrapling/engines/_browsers/_stealth.py)
- AsyncStealthySession._cloudflare_solver (cyclomatic 21) (scrapling/engines/_browsers/_stealth.py)
- AsyncStealthySession.fetch (cognitive 39) (scrapling/engines/_browsers/_stealth.py)
- AsyncStealthySession.fetch (cyclomatic 19) (scrapling/engines/_browsers/_stealth.py)
- BaseSessionMixin.__generate_options__ (cognitive 31) (scrapling/engines/_browsers/_base.py)
- BaseSessionMixin.__generate_options__ (cyclomatic 19) (scrapling/engines/_browsers/_base.py)
- Change coupling: _controllers.py ↔ _stealth.py (scrapling/engines/_browsers/_controllers.py)
- Change coupling: chrome.py ↔ stealth_chrome.py (scrapling/fetchers/chrome.py)
- Convertor._extract_content (cognitive 18) (scrapling/core/shell.py)
- CrawlerEngine._process_request (cognitive 25) (scrapling/spiders/engine.py)
- CrawlerEngine._process_request (cyclomatic 20) (scrapling/spiders/engine.py)
- CrawlerEngine._run_callbacks (cognitive 18) (scrapling/spiders/engine.py)
- CrawlerEngine.crawl (cognitive 36) (scrapling/spiders/engine.py)
- CrawlerEngine.crawl (cyclomatic 20) (scrapling/spiders/engine.py)
- CurlParser.parse (cognitive 38) (scrapling/core/shell.py)
- CurlParser.parse (cyclomatic 31) (scrapling/core/shell.py)
- Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
- Documentation: written for insiders
- …and 99 more
Changes since last survey
- 64 commits — 50 feature/other, 14 fixes
By area
- (root) — 31 commits
- (repo) — 7 commits
- agent-skill/Scrapling-Skill — 6 commits
- scrapling/core — 6 commits
- scrapling/spiders — 3 commits
- docs/assets — 2 commits
- docs/index.md — 2 commits
- scrapling/parser.py — 2 commits
- tests/fetchers — 2 commits
- docs/ai — 1 commit
- images/novada.jpg — 1 commit
- scrapling/engines — 1 commit
Notable commits
- fix: fix(ai)!: keep session settings on MCP session fetches (#418)
- fix: fix(ai): size the MCP bulk page pools within the validator bounds (#393)
- fix: fix(browser version): extract current chrome version from playwright package data
- fix: fix(mcp): Set the session proxy to be browser-level argument
- fix: fix(mcp): point the server card website_url at the MCP docs page
- fix: fix(parser): find/find_all with class_ silently miss multi-class elements (#410)
- fix: fix(parser): keep filtering on blank class_ and escape CSS string values (#417)
- fix: fix(parser): match multi-class elements in find/find_all class_ filter
- fix: fix(spiders): keep the request meta on cached responses (#419)
- fix: fix(static): send the request once when retries is below 1 (#420)
- fix: fix(stealth): solve Cloudflare challenges regardless of locale + retry cap + fix crashing mid-solve
- fix: fix(xml template): Be specific about the node type hint
- fix: fix: require authentication by default for the HTTP transport (#414)
- fix: fix:(mcp): upgrade mcp server to v2
- change: Merge branch 'main' into dev
- change: Merge branch 'main' into dev
- change: Merge branch 'main' into dev
- change: Updating contribution rules
- change: build: correcting dep version
- change: build: pump version and deps
- …and 44 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
D4Vinci/Scrapling was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit e0d4d7563207b70c2cb38487c4e0dcbe0e2f04ca — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-a15879f6f801.