unclecode/crawl4ai
40.7
Weak · 26 September 2026
111.4k
lines of production code
Python
with JavaScript
4
measurements over time
What this system is
Crawl4AI is a web crawling and data extraction library that provides adaptive, deep, and multi-source crawling capabilities for websites, PDFs, and search engines. It features a secure-by-default Docker deployment with strict egress controls, SSRF protection, and API authentication, alongside a modular architecture supporting custom scripts and real-time monitoring. The system includes robust content processing tools for HTML-to-Markdown conversion and integrates with LLMs for intelligent data extraction and adaptive depth management.
How it got here
2024 — Security hardening and crawler modernization
10 changes.
This period focused on significantly enhancing security by removing the legacy web crawler module, implementing strict egress policies to prevent SSRF, and patching critical RCE vulnerabilities. Concurrently, the project modernized its architecture by adopting a pyproject.toml build system, integrating robust HTML-to-Markdown conversion, and introducing adaptive crawling strategies with advanced anti-bot detection capabilities.
2025–2026 — Security hardening and deep crawling infrastructure
30 changes.
This period focused on significantly hardening the Docker deployment with mandatory authentication, SSRF protection, and secure artifact storage, while introducing a comprehensive deep crawling framework with BFS, DFS, and Best-First strategies. The team also expanded core capabilities by adding dedicated PDF processing, Amazon and Google Search crawlers, and a new DSL for web automation. Extensive testing efforts covered these new features alongside critical security regressions, browser management stability, and high-volume stress scenarios.
Features
Added real-time terminal dashboard for crawler monitoring
A new \CrawlerMonitor\ component with a \TerminalUI\ has been introduced, providing a live, interactive dashboard for tracking crawler performance. This UI displays real-time statistics including memory usage, completion progress, and task details, and supports keyboard input (press 'q' to quit) on both Unix/macOS and Windows platforms.
crawl4ai/components · high confidence
Added scripts for SBOM generation and stats dashboard updates
Added a new shell script (scripts/gen-sbom.sh) to generate a CycloneDX Software Bill of Materials (SBOM) using Syft, and a Python script (scripts/update\_stats.py) to fetch live data from GitHub, PyPI, and Docker Hub to generate a stats dashboard page (docs/md\_v2/stats.md) with embedded Chart.js charts.
scripts · high confidence
Crawl4AI v0.9.4 release with adaptive crawling and security hardening
This release introduces the Adaptive Crawler, a new capability that uses statistical or embedding-based strategies to intelligently determine when sufficient information has been gathered for a query, automatically managing crawl depth and page selection. It also includes significant security hardening, specifically patching a critical Remote Code Execution (RCE) vulnerability caused by unsafe deserialization and eval() usage in the crawl endpoint and config deserializer. Additionally, the library adds support for C4A Script language, enhances image processing with srcset validation, and introduces a configurable cache directory for constrained environments.
crawl4ai · high confidence
Integrate html2text library for HTML-to-Markdown conversion
The crawl4ai package now includes the html2text library to provide robust HTML-to-Markdown conversion capabilities. This addition introduces a full suite of components including the core HTML2Text parser, configuration options, CLI utilities, and helper utilities, enabling users to transform crawled HTML content into structured Markdown text directly within the crawl4ai ecosystem.
crawl4ai/html2text · high confidence
Introduce C4A-Script, a domain-specific language for web automation
Users can now write automation workflows using C4A-Script, a new DSL that is transpiled into JavaScript for execution within Crawl4AI. This location provides the compiler infrastructure, including the \C4ACompiler\ class and result types, which parse the script syntax, validate it, and generate the corresponding JS code without modifying the core crawler logic.
crawl4ai/script · high confidence
Introduce dedicated PDF processing strategy with security and performance controls
The \crawl4ai/processors\ area now includes a new \pdf\ subpackage that provides \PDFCrawlerStrategy\ and \PDFContentScrapingStrategy\ to handle PDF documents. This change adds the capability to download, parse, and extract text, images, and links from PDFs using the \pypdf\ library. To address security and stability concerns, the implementation enforces configurable limits on download size (default 100 MiB), page count (default 2000), and redirect hops (default 5), and includes hooks for validating egress destinations and peer IPs to prevent SSRF attacks. Additionally, it fixes a potential DOM XSS vulnerability by escaping paragraph text in the generated HTML output, and improves performance by processing PDF pages in parallel batches.
crawl4ai/processors · high confidence
Introduce modular JavaScript snippets for shadow DOM, overlays, and navigator spoofing
The crawler now uses a dedicated \js\_snippet\ module to load and execute JavaScript files, replacing inline scripts. This update adds support for flattening Shadow DOM trees into the light DOM for accurate serialization, removes consent popups and overlay elements using extensive, CMP-specific selectors, and spoofs the \navigator\ object (permissions, webdriver, plugins, languages) to bypass bot detection. It also includes a script to update image dimensions based on natural size and filters out small or placeholder images.
_crawl4ai/js\snippet · high confidence
New Amazon and Google Search crawlers added
Added new crawler implementations for Amazon product pages and Google Search results (including text and image search). The Amazon crawler provides a basic structure for extracting product name and price, while the Google Search crawler supports querying for organic results, top stories, and suggested queries, as well as extracting image metadata via a dedicated JavaScript execution script.
crawl4ai/crawlers · high confidence
New Crawl4AI Playground UI for Docker deployments
A new interactive playground interface is now available at the /playground route within the Docker deployment. This single-page application provides a user-friendly environment for testing API requests, featuring a token-based authentication bar, tabbed navigation between the Playground and Stress Test modes, and a JSON configuration editor powered by CodeMirror. The UI includes syntax highlighting for responses and integrates with the application's monitoring dashboard.
deploy/docker/static/playground · high confidence
New deep crawling strategies and filtering infrastructure
The \crawl4ai/deep\_crawling\ module introduces a structured framework for advanced web traversal, providing three distinct strategies: Breadth-First Search (\BFSDeepCrawlStrategy\), Best-First (\BestFirstCrawlingStrategy\) which prioritizes URLs using configurable scorers, and Depth-First Search (\DFSDeepCrawlStrategy\). This update adds a comprehensive filtering system via \FilterChain\ and specific filters like \URLPatternFilter\ and \ContentTypeFilter\, alongside a suite of URL scorers (e.g., \KeywordRelevanceScorer\, \FreshnessScorer\) to control crawl depth and relevance. The implementation includes support for crash recovery, cancellation, and strict page limits to manage resource usage during deep crawling operations.
_crawl4ai/deep\crawling · high confidence
Removals
Removal of Python example script
The example script \examples/test.py\, which demonstrated how to initialize the WebCrawler and fetch pages using OpenAI providers, has been removed from the repository.
examples · high confidence
Removal of crawler module and associated components
The entire crawler package has been removed from the codebase. This deletion eliminates the web crawling functionality, including the \WebCrawler\ class that utilized Selenium for page fetching, the SQLite-based caching and database logic, the configuration for AI provider models, the Pydantic data models, the prompt templates for content extraction, and the utility functions for HTML processing and JSON parsing.
crawler · high confidence
Removal of legacy web-based crawling interfaces
The \index.html\ and \index\_pooling.html\ landing pages have been removed from the \pages\ directory. These files previously provided browser-based UIs for crawling websites, including input fields for URLs, API tokens, and model selection, as well as tabs for viewing results in JSON, Cleaned HTML, or Markdown. Their removal indicates that the web-based interactive crawling experience is no longer part of this component.
pages · high confidence
Security
Introduce egress policy to enforce network egress controls
The library now supports a process-wide egress policy that restricts outbound network requests to a configured proxy. This mechanism is used by the Docker server to enforce egress controls for internal HTTP clients, ensuring that requests made during crawling (such as fetching robots.txt or link previews) are routed through a pinned proxy that validates destinations. This change closes SSRF vulnerabilities by preventing untrusted requests from reaching internal or cloud-metadata endpoints.
(repo-wide) · high confidence
Behavioural changes
Secure-by-default Docker deployment with mandatory API authentication and hardened artifact storage
The Docker server now enforces a secure-by-default posture: it requires a \CRAWL4AI\_API\_TOKEN\ to bind to external interfaces (defaulting to loopback-only without one), implements JWT authentication for API access, and replaces the insecure \output\_path\ parameter for screenshots and PDFs with a server-managed, opaque-ID artifact store that prevents arbitrary file writes. Additionally, the deployment introduces a smart 3-tier browser pool for memory efficiency, declarative-only hook actions (replacing arbitrary code execution), and comprehensive security hardening including SSRF protection, Redis password requirements, and read-only root filesystem support.
deploy/docker · high confidence
Test coverage
Add Wikipedia sample HTML for async tests; Added CLI test suite for crawling, configuration, and deep-crawl output; Added Docker server behavioral security and performance test suites; Added adversarial and integration tests for DomainMapper; Added browser module test suite; Added cache validation test suite; Added demo script for browser pooling and page pre-warming; Added high-volume stress testing and benchmarking framework; Added integration tests for MCP server protocols; Added regression tests for release 0.6.4 and 0.7.0; Added test coverage for deep crawling stability, cancellation, and crash recovery; Added test coverage for proxy configuration, anti-bot detection, and sticky sessions; Added test coverage for webhook, CDP, logging, and anti-bot fixes; Added test suite for adaptive crawler performance and embedding strategies; Added tests for BrowserProfiler profile creation and keyboard input handling; Added tests for ContentTypeFilter PHP MIME type handling; Added tests for CrawlerHub integration examples; Added tests for Docker-based browser automation strategy; Added tests for the AI Assistant extract pipeline; Added tests for the new file logger implementation; Added unit tests for browser flags, config provenance, domain mapping, egress proxy, PDF security, and robots.txt parsing; Comprehensive regression test suite for Crawl4AI; Expanded Docker integration test coverage; Expanded test coverage for crawler strategies, deep crawling, and markdown generation.
Dependencies
Modernize build system and pin dependencies for security and compatibility
The project has migrated from legacy setup files to a modern pyproject.toml configuration, establishing a strict Python 3.10+ requirement and defining core dependencies with specific version constraints. Key updates include pinning litellm to the secure unclecode-litellm fork (v1.81.13) to address a PyPI supply chain compromise, upgrading pyOpenSSL to \>=25.3.0 to fix a security vulnerability, and replacing chromedriver\_autoinstaller with playwright-stealth and patchright for browser management. The dependency list has also been expanded to include chardet, brotli, fake-useragent, and pdf2image, while optional dependencies for PDF, torch, and transformers are now explicitly categorized.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 35 → 41 (+5.6)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 32 → 25 (-7.1)
- Architecture 94 → 98 (+4.2)
- Maturity 73 → 71 (-2.4)
- Readiness 19 → 47 (+28.5)
- Security 55 → 70 (+15.3)
- Accessibility 46 (new)
Resolved (134)
- (anonymous) (cognitive 21) (docs/md_v2/apps/crawl4ai-assistant/content/scriptBuilder.js)
- (anonymous) (cognitive 21) (docs/md_v2/assets/github_stats.js)
- (anonymous) (cognitive 22) (docs/md_v2/apps/crawl4ai-assistant/content/content.js)
- (anonymous) (cognitive 86) (crawl4ai/js_snippet/remove_consent_popups.js)
- (anonymous) (cyclomatic 16) (docs/md_v2/apps/crawl4ai-assistant/content/content.js)
- (anonymous) (cyclomatic 17) (docs/md_v2/assets/github_stats.js)
- (anonymous) (cyclomatic 19) (docs/md_v2/apps/crawl4ai-assistant/content/scriptBuilder.js)
- (anonymous) (cyclomatic 42) (crawl4ai/js_snippet/remove_consent_popups.js)
- Change coupling: admin.js ↔ app-detail.js (docs/md_v2/marketplace/admin/admin.js)
- Change coupling: admin.js ↔ app-detail.js (docs/md_v2/marketplace/admin/admin.js)
- Change coupling: admin.js ↔ marketplace.js (docs/md_v2/marketplace/admin/admin.js)
- Change coupling: app-detail.js ↔ app-detail.js (docs/md_v2/marketplace/app-detail.js)
- Change coupling: app-detail.js ↔ marketplace.js (docs/md_v2/marketplace/app-detail.js)
- Change coupling: app-detail.js ↔ marketplace.js (docs/md_v2/marketplace/frontend/app-detail.js)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (10 lines × 2) (crawl4ai/async_configs.py)
- Duplicated block (10 lines × 2) (crawl4ai/async_crawler_strategy.py)
- Duplicated block (10 lines × 2) (crawl4ai/html2text/init.py)
- Duplicated block (11 lines × 2) (crawl4ai/deep_crawling/bff_strategy.py)
- …and 114 more
New (669)
- (anonymous) (cognitive 102) (crawl4ai/js_snippet/remove_consent_popups.js)
- (anonymous) (cognitive 16) (crawl4ai/js_snippet/update_image_dimensions.js)
- (anonymous) (cognitive 32) (crawl4ai/js_snippet/remove_overlay_elements.js)
- (anonymous) (cyclomatic 18) (crawl4ai/crawlers/google_search/script.js)
- (anonymous) (cyclomatic 21) (crawl4ai/js_snippet/flatten_shadow_dom.js)
- (anonymous) (cyclomatic 28) (crawl4ai/js_snippet/remove_overlay_elements.js)
- (anonymous) (cyclomatic 56) (crawl4ai/js_snippet/remove_consent_popups.js)
- (anonymous)::extractImageData (cognitive 17) (crawl4ai/crawlers/google_search/script.js)
- (anonymous)::serializeShadowChild (cognitive 17) (crawl4ai/js_snippet/flatten_shadow_dom.js)
- 0.print_summary (cognitive 20) (docs/releases_review/demo_v0.8.0.py)
- 5.print_summary (cognitive 20) (docs/releases_review/demo_v0.8.5.py)
- 8.print_summary (cognitive 20) (docs/releases_review/demo_v0.7.8.py)
- 8.test_adaptive_crawler_embedding (cognitive 22) (docs/releases_review/demo_v0.7.8.py)
- 8.test_import_formatting (cognitive 19) (docs/releases_review/demo_v0.7.8.py)
- AdaptiveCrawler.digest (cognitive 73) (crawl4ai/adaptive_crawler.py)
- AdaptiveCrawler.digest (cyclomatic 32) (crawl4ai/adaptive_crawler.py)
- AdaptiveCrawler.print_stats (cognitive 111) (crawl4ai/adaptive_crawler.py)
- AdaptiveCrawler.print_stats (cyclomatic 35) (crawl4ai/adaptive_crawler.py)
- AdminDashboard.getAppForm (cognitive 19) (docs/md_v2/marketplace/admin/admin.js)
- AdminDashboard.getAppForm (cyclomatic 20) (docs/md_v2/marketplace/admin/admin.js)
- …and 649 more
Changes since last survey
- 101 commits — 25 feature/other, 76 fixes
By area
- (repo) — 40 commits
- deploy/docker — 23 commits
- (root) — 8 commits
- crawl4ai/async_configs.py — 5 commits
- crawl4ai/processors — 4 commits
- crawl4ai/js_snippet — 3 commits
- .github/workflows — 2 commits
- crawl4ai/async_crawler_strategy.py — 2 commits
- crawl4ai/browser_manager.py — 2 commits
- crawl4ai/deep_crawling — 2 commits
- crawl4ai/table_extraction.py — 2 commits
- crawl4ai/utils.py — 2 commits
- docs/md_v2 — 2 commits
- crawl4ai/async_webcrawler.py — 1 commit
- docs/assets — 1 commit
- tests/general — 1 commit
- tests/test_docker_pdf_crawler_pairing.py — 1 commit
Notable commits
- fix: Fix: Make the untrusted timeout ceiling configurable
- fix: Merge pull request #2094 from unclecode/fix/docker-deploy-0.9-gaps
- fix: Merge pull request #2117 from nightcityblade/fix/issue-2116
- fix: Merge pull request #2130 from nightcityblade/fix/issue-2127
- fix: Merge pull request #2131 from nightcityblade/fix/issue-2129
- fix: Merge pull request #2134 from nightcityblade/fix/issue-2133
- fix: Merge pull request #2138 from unclecode/fix/pdf-antibot-false-positive
- fix: Merge pull request #2139 from unclecode/fix/csp-sandbox-overlay-hang
- fix: Merge pull request #2142 from unclecode/fix/egress-proxy-upstream-chaining
- fix: Merge pull request #2145 from unclecode/fix/issue-2144
- fix: Merge pull request #2148 from weike-zhang/fix/mcp2-docker-requirements
- fix: Merge pull request #2150 from unclecode/fix/docker-pdf-crawler-pairing
- fix: Merge pull request #2156 from unclecode/fix/compose-comment-indent
- fix: Merge pull request #2157 from unclecode/fix/monitor-dashboard-redirect
- fix: Merge pull request #2158 from unclecode/test/per-url-pdf-ssrf-regression
- fix: Merge pull request #2159 from unclecode/fix/pdf-redirect-cookie-session
- fix: Merge pull request #2160 from unclecode/fix/browser-startup-driver-leak
- fix: Merge pull request #2212 from damusix/fix/configurable-max-timeout
- fix: Merge pull request #2224 from unclecode/fix/playground-2222
- fix: Merge pull request #2229 from Nalhin/fix-robots-parsing
- …and 81 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
unclecode/crawl4ai was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit e5d2e786d1a101225f3f6a3e6fd344d76eeb13af — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-a15879f6f801.