Skip to content
CAI
Software that uses CAICheck a score

ruippeixotog/scala-scraper

65.4

Adequate · 28 September 2026

1.4k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Scala Scraper is a library for web scraping that provides a unified abstraction over Jsoup and HtmlUnit browsers, supporting features like cookie management, proxy configuration, and JavaScript execution. It offers a type-safe, composable DSL for extracting and validating HTML content, including support for CSS queries, table parsing, and declarative configuration via Typesafe Config. The system is built on modern Scala versions and includes robust utilities for resource management and generic type operations.

How it got here

2014 — Scala Scraper 3.2.0 initial setup and modernization

4 changes.

This period established the initial project structure and documentation for Scala Scraper 3.2.0, including comprehensive guides and configuration files. The work involved modernizing the build system by upgrading SBT and dependencies to support Scala 2.13 and 3, while simultaneously removing legacy scraping DSL components to streamline the core API.

2017 — Core API and DSL Refactoring

12 changes.

This period focused on a comprehensive refactoring of the library's core architecture, introducing a unified browser abstraction with cookie and proxy support alongside a new structured DOM model. The scraping DSL was redesigned to use explicit extractors and type-safe combinators, replacing implicit conversions with a more composable and readable pipeline. These changes were accompanied by the addition of configuration-driven extraction capabilities and extensive test coverage to validate the new APIs.

Features

Add DeepFunctor utility and resource management helper

The library introduces a new \DeepFunctor\ type class in the \util\ package, enabling the deconstruction of nested type constructors into a composite functor for advanced generic operations. Additionally, a \using\ helper method is added to the \util\ package to simplify safe resource management by automatically closing \Closeable\ objects after execution.

core/src/main/scala/net/ruippeixotog/scalascraper/util · high confidence

Configuration-driven HTML extraction and validation

Users can now define HTML extractors and validators via Typesafe Config files instead of writing code. This change introduces \ConfigHtmlExtractor\ and \ConfigHtmlValidator\ objects that parse configuration paths (such as CSS queries, attribute selectors, date formats, and match conditions) to build scraper components. The \ConfigLoaders\ DSL provides convenient methods like \extractorAt\ and \validatorAt\ to load these components from configuration, enabling declarative scraper definitions.

modules/config/src/main · high confidence

Initial project setup and documentation for Scala Scraper 3.2.0

This change introduces the initial project structure and documentation for Scala Scraper version 3.2.0. It adds configuration files for code formatting (scalafmt) and import organization (scalafix), sets the project version to 3.2.1-SNAPSHOT, and includes a comprehensive CHANGELOG detailing the history from 0.1 to 3.2.0. The README provides a quick start guide, core model documentation, and usage examples for the DSL, browsers (Jsoup and HtmlUnit), and content extraction features. It also updates the .gitignore to include standard build and IDE artifacts.

(repo-wide) · high confidence

Introduction of core DOM model abstractions

The library now exposes a structured model layer for HTML parsing, introducing traits for Document, Element, Node, and ElementQuery. Users can now interact with a Document's location, title, and body, and traverse the DOM via Element methods such as text, ownText, and CSS selection. The Node trait distinguishes between ElementNode and TextNode, allowing access to raw child and sibling nodes alongside structured element queries.

core/src/main/scala/net/ruippeixotog/scalascraper/model · high confidence

The library introduces a common \Browser\ trait that standardizes the interface for web scraping implementations, including \HtmlUnitBrowser\ (which executes JavaScript) and \JsoupBrowser\ (which parses static HTML). This change adds first-class support for managing cookies via \cookies\, \setCookie\, and \clearCookies\ methods, and allows users to configure HTTP or SOCKS proxies using a new \Proxy\ case class and the \withProxy\ method on browser instances.

core/src/main/scala/net/ruippeixotog/scalascraper/browser · high confidence

Removals

Removal of legacy scraping DSL and configuration support

The library has removed the legacy HTML scraping DSL and its associated configuration loading infrastructure. This includes the deletion of the \application.conf\ example file, the \Examples.scala\ demonstration code, and core DSL components such as \ScrapingOps\, \ImplicitConversions\, and \ConfigLoadingHelpers\. Additionally, foundational scraper types like \HtmlExtractor\, \HtmlStatusMatcher\, and utility traits like \Mappable\ and \Validated\ have been removed from the \src/main\ package, indicating a significant restructuring or deprecation of the previous extraction and validation API.

src/main · high confidence

Behavioural changes

Redesigned DSL with explicit extractors and new validation operators

The scraping DSL has been refactored to replace previous implicit conversions with explicit \HtmlExtractor\ combinators and a new \ToQuery\ type class for type-safe element queries. Users now interact with the DSL through the \DSL\ object, which provides factory methods for extractors and aliases for content parsers. The \ScrapingOps\ trait introduces new operator-based syntax for extraction (\\>\>\, \\>?\>\) and validation (\\>/\~\), allowing for more composable and readable scraping pipelines while deprecating older, less explicit patterns.

core/src/main/scala/net/ruippeixotog/scalascraper/dsl · high confidence

Refactored extraction API with new HtmlExtractor and PolyHtmlExtractor types

The scraper's extraction logic has been restructured to use a new \HtmlExtractor\ trait and a \PolyHtmlExtractor\ for polymorphic type-safe extraction, replacing the previous \SimpleExtractor\ approach. This change introduces \ContentExtractors\ and \ContentParsers\ as dedicated objects for primitive data extraction (text, attributes, tables) and parsing (dates, regex), providing a more composable and type-safe DSL for users. The \HtmlValidator\ has also been updated to work with these new extractor types, allowing validation logic to be built on top of the new extraction primitives.

core/src/main/scala/net/ruippeixotog/scalascraper/scraper · high confidence

Test coverage

Added HTML test fixtures for encoding, structure, and JavaScript behavior; Added comprehensive browser test suite and test infrastructure; Added test infrastructure and example applications for proxying and DSL usage; Added tests for ElementQuery CSS selection and equality; Added tests for configuration-based extractors and validators; Added tests for the scraping DSL extraction, validation, and typing features.

Dependencies

Upgrade SBT build tool to 1.13.0 and update plugins

The SBT build tool version has been upgraded from 0.13.6 to 1.13.0. Additionally, several SBT plugins have been updated to newer versions: sbt-scalafix to 0.14.9, sbt-release to 1.5.0, sbt-pgp to 2.3.2, sbt-mdoc to 2.9.2, sbt-scalafmt to 2.6.2, sbt-coveralls to 1.3.15, and sbt-scoverage to 2.4.4.

project · high confidence

Upgrade to Scala 2.13/3.9 and modernize dependencies

The build system has been upgraded to support Scala 2.13.18 and Scala 3.9.0, dropping older Scala versions. Core dependencies have been updated to their latest versions, including HtmlUnit 5.5.0, Jsoup 1.23.2, nscala-time 3.2.0, and scalaz-core 7.3.9. The project structure has been converted to a multi-project build with a separate \config\ module, and test dependencies like specs2-core (4.23.0) and akka-http (10.2.10) have also been updated.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 68 → 65 (-2.6)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 99 → 99 (+0.0)
  • Architecture 97 → 69 (-27.8)
  • Maturity 56 → 56 (+0.0)
  • Readiness 73 → 75 (+2.8)
  • Security 73 → 73 (+0.0)

Resolved (3)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • Off-boarding risk: anonymized user #1

New (1)

  • Off-boarding risk: anonymized user #1

Changes since last survey

  • 3 commits — 3 feature/other, 0 fixes

By area

  • (root) — 2 commits
  • project/plugins.sbt — 1 commit

Notable commits

  • change: Update nscala-time to 3.2.0 (#694)
  • change: Update sbt-scalafix to 0.14.9 (#693)
  • change: Update slf4j-nop to 2.0.20 (#696)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

ruippeixotog/scala-scraper was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 694d99ef0b20ebc9dfae9240b1998be370d78726 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.