elixir-crawly/crawly
55.7
Adequate · 23 September 2026
3.7k
lines of production code
Elixir
primary language
5
measurements over time
What this system is
Crawly is an Elixir-based web scraping framework that manages the full lifecycle of headless browser and HTTP-based crawling. It provides a pluggable architecture for fetching content, processing requests through configurable middlewares, and storing scraped data. The system includes a built-in HTTP API and management UI for scheduling, monitoring, and controlling spider jobs, alongside tools for generating boilerplate code and publishing packages.
How it got here
2019 — Initial framework scaffolding and core architecture
14 changes.
This period established the foundational structure of the Crawly web scraping framework, introducing the core crawling engine, pluggable fetcher architecture, and data storage systems. It also added essential middlewares, pipelines, and a management API to support spider execution and monitoring.
2020–2024 — management UI and developer tooling
9 changes.
This period focused on expanding the project's usability and infrastructure by introducing a standalone management UI and corresponding EEx templates. Developers were provided with new Mix tasks and a modern Elixir quickstart example to streamline spider creation and configuration. The release environment and test coverage were also enhanced to support these new features.
Features
Add Crawly quickstart example
Added a new quickstart example in the \examples/quickstart\ directory, demonstrating how to use the Crawly library to scrape the Books to Scrape website. The example includes a basic Elixir module, an OTP application setup, and a spider implementation that parses book titles, prices, and URLs from the target site.
examples/quickstart · high confidence
Add HTTP API and management UI for controlling spiders
Crawly now includes a built-in HTTP API and a simple management UI, allowing users to schedule, stop, and monitor all running spiders. The API is enabled by default and can be disabled via the \:start\_http\_api?\ application environment. The UI provides a list of known and running spiders, their status, and job history. The engine now supports starting and stopping spiders via the API, and the manager handles worker processes for each spider. The job model tracks crawl history, and the YML spider model allows storing and loading spider configurations. The API router handles GET requests for the UI and POST requests for creating new spiders from YAML definitions.
lib/crawly · high confidence
Add SendToUiBackend logger for UI communication
A new logger module, SendToUiBackend, has been introduced to handle log messages and forward them to a specified backend via RPC calls, enabling the UI to receive real-time log updates.
lib/crawly/loggers · high confidence
Add modern Elixir quickstart example for Crawly
A new quickstart example has been added to the examples directory, providing a modern Elixir implementation of the Crawly web scraper. The example includes a full project structure with configuration for Crawly's spider behavior, middlewares, and pipelines, along with standard Elixir tooling files (.formatter.exs, .gitignore) and a basic test suite to demonstrate usage.
examples · high confidence
Added Mix task generators for Crawly configuration and spider templates
Users can now generate boilerplate code for Crawly projects directly from the command line. New Mix tasks, \crawly.gen.config\ and \crawly.gen.spider\, allow developers to quickly scaffold a Crawly configuration file and a spider template, reducing the manual setup required to start a new crawling project.
lib/mix · high confidence
Added management UI templates and spider templates
Added new EEx templates in the priv directory to support a standalone management tool interface. This includes the main index page, a spider list view with job history, a spider creation/editing form with YAML preview, and a scheduled requests list. Additionally, new template files were added to generate spiders from YAML definitions and a default Elixir spider template, enabling users to create and manage spiders directly through the UI.
priv · high confidence
Added new request processing middlewares for crawling
Added several new middlewares to the Crawly library to enhance request handling and filtering. These include AutoCookiesManager for automatic cookie management, DomainFilter and SameDomainFilter for restricting crawls to specific domains, RequestOptions for configuring HTTP request settings like timeouts and proxies, RobotsTxt for respecting robots.txt directives, UniqueRequest for deduplicating requests, and UserAgent for rotating user-agent headers.
lib/crawly/middlewares · high confidence
Automated Hex package and documentation publishing
A new 'scripts/hex.sh' script has been added to automate the publishing of the Hex package and its generated documentation. The script configures Hex authentication, retrieves dependencies, and executes the build, documentation generation, and publishing steps in a single workflow.
scripts · high confidence
Initial configuration for Crawly web scraping framework
Added configuration files for the Crawly application, establishing default settings for logging, HTTP API, and request handling. The main config/config.exs and environment-specific files (dev, test, standalone) define Crawly's behavior, including the HTTPoison fetcher, retry policies (e.g., max\_retries: 3), and middleware pipelines (DomainFilter, UniqueRequest, RobotsTxt, UserAgent). This sets up the core structure for crawling, including item validation, deduplication, JSON encoding, and file output.
config · high confidence
Initial project scaffolding and configuration
The repository was initialized with essential Elixir project configuration files, including \.credo.exs\ for code linting, \.formatter.exs\ for code formatting, and \.gitignore\ to exclude build artifacts and IDE files. A \.tool-versions\ file was added to specify Elixir 1.14.0, and a \Dockerfile\ was introduced to support running Crawly as a standalone application in a Docker container. Additionally, a \coveralls.json\ file was added to configure code coverage reporting, and the \LICENSE\ file was updated.
(repo-wide) · high confidence
Introduce RequestsStorage module for managing spider request queues
Added the Crawly.RequestsStorage GenServer and its associated Worker to handle storing, popping, and retrieving requests for each spider. The new module manages per-spider worker processes, supports batch storage, and exposes stats and request retrieval endpoints, enabling the crawler to manage its queue of URLs to visit.
_lib/crawly/requests\storage · high confidence
Introduce data storage architecture for crawling
Added a new data storage system that manages item persistence through a central GenServer (DataStorage) that routes requests to per-spider worker processes (DataStorage.Worker). This architecture allows each spider to maintain its own state and pipeline processing, with automatic cleanup when workers terminate.
_lib/crawly/data\storage · high confidence
Introduce pluggable fetcher architecture with new implementations
The library now supports pluggable fetchers via a new \Crawly.Fetchers.Fetcher\ behavior, allowing users to swap out the default HTTP client. This change introduces three new fetcher implementations: \CrawlyRenderServer\ for headless browser rendering via a local server, \Splash\ for the Splash JavaScript rendering engine, and \HTTPoisonFetcher\ for standard HTTP requests. Users can now configure their crawlers to use any of these fetchers by specifying the module and options in their configuration.
lib/crawly/fetchers · high confidence
New Crawly pipelines for data processing and storage
Added five new pipeline implementations for the Crawly framework: CSVEncoder, DuplicatesFilter, JSONEncoder, Validate, and WriteToFile. These pipelines provide built-in support for encoding scraped items into CSV or JSON formats, filtering out duplicate items based on a unique identifier, validating that required fields are present, and writing processed items to the filesystem with configurable folder paths, file extensions, and timestamped filenames.
lib/crawly/pipelines · high confidence
Behavioural changes
Added release environment and VM configuration templates
Added new environment and VM configuration templates for the release, including scripts for Windows (env.bat.eex) and Unix (env.sh.eex) to support interactive code loading and distributed node communication, alongside VM argument templates (vm.args.eex and remote.vm.args.eex) that configure Erlang/OTP settings such as port limits and garbage collection tuning.
rel · medium confidence
Introduce Crawly module with fetch, spider integration, and spider management
The public API for web crawling is now exposed via the Crawly module. Users can fetch URLs using Crawly.fetch/2, which supports custom headers, request options, and pluggable HTTP clients. The module also provides Crawly.fetch\_with\_spider/3 to automatically process responses with a spider, and a parse/2 function for direct spider invocation. Additionally, the module exposes utility functions to list and load spiders from directories or YAML storage.
lib · medium confidence
Test coverage
Added comprehensive test coverage for Crawly core components; Added test coverage for Crawly middlewares; Added tests for Crawly pipeline components; Added tests for Job model; Added tests for the CrawlyRenderServer fetcher; Added tests for the spider generator task.
Dependencies
Updated Elixir project dependencies and added a quickstart example
The Crawly Elixir project and its quickstart example have been updated with new dependency versions. The main \mix.exs\ now specifies \httpoison \~\> 2.2\, \elixir\_uuid \~\> 1.2\, \poison \~\> 3.1\, \gollum \~\> 0.5.0\, \plug\_cowboy \~\> 2.0\, \yaml\_elixir \~\> 2.9\, and \ex\_json\_schema \~\> 0.9.2\. The quickstart example (\examples/quickstart/mix.exs\) has been added, configuring the project to use \crawly\ from the local path and \floki \~\> 0.33.0\. The lock files (\mix.lock\ and \examples/quickstart/mix.lock\) have been updated to reflect the resolved versions of all transitive dependencies, including \httpoison 2.2.1\, \gollum 0.5.0\, and \floki 0.33.1\.
(dependencies) · medium confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 62 → 56 (-5.8)
- Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 99 (-0.9)
- Architecture 100 → 93 (-7.5)
- Maturity 57 → 53 (-4.1)
- Readiness 50 → 49 (-1.0)
- Security 77 → 85 (+8.1)
- Accessibility 51 (new)
Resolved (19)
- Coverage not included — suite not readable by the collector
- Dependency hygiene not measured — no supported dependency manifest was read
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High IaC: DS-0025 (Dockerfile)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- Medium CVE: EEF-[CVE redacted] (mix.lock)
- Medium CVE: EEF-[CVE redacted] (mix.lock)
- Medium CVE: EEF-[CVE redacted] (mix.lock)
- Medium IaC: CKV_DOCKER_3 (Dockerfile)
- Medium IaC: CKV_DOCKER_7 (Dockerfile)
- No exposed public API
- Off-boarding risk: anonymized user #1
- Test reliability not included
New (31)
- Coverage not measured — no coverage collector is wired up
- Duplicated block (10 lines × 2) (lib/crawly/data_storage/data_storage.ex)
- Duplicated block (12 lines × 2) (lib/crawly/data_storage/data_storage.ex)
- Duplicated block (6 lines × 2) (lib/crawly/data_storage/data_storage.ex)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High CVE: [GHSA redacted] (mix.lock)
- High IaC: DS-0025 (Dockerfile)
- High IaC: DS-0025 (Dockerfile)
- High IaC: DS-0025 (Dockerfile)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- Medium CVE: EEF-[CVE redacted] (mix.lock)
- Medium CVE: EEF-[CVE redacted] (mix.lock)
- Medium IaC: WD-DOCKER-0003 (Dockerfile)
- …and 11 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
elixir-crawly/crawly was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 23 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 89db0d6d12d7c0e97961692ce0d07989c7b988c1 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-955b9cee9818.