ruippeixotog/scala-scraper
65.4
Adequate · 28 September 2026
1.4k
lines of production code
Scala
primary language
2
measurements over time
What this system is
Scala Scraper is a library for web scraping that provides a unified abstraction over Jsoup and HtmlUnit browsers, supporting features like cookie management, proxy configuration, and JavaScript execution. It offers a type-safe, composable DSL for extracting and validating HTML content, including support for CSS queries, table parsing, and declarative configuration via Typesafe Config. The system is built on modern Scala versions and includes robust utilities for resource management and generic type operations.
How it got here
2014 — Scala Scraper 3.2.0 initial setup and modernization
4 changes.
This period established the initial project structure and documentation for Scala Scraper 3.2.0, including comprehensive guides and configuration files. The work involved modernizing the build system by upgrading SBT and dependencies to support Scala 2.13 and 3, while simultaneously removing legacy scraping DSL components to streamline the core API.
2017 — Core API and DSL Refactoring
12 changes.
This period focused on a comprehensive refactoring of the library's core architecture, introducing a unified browser abstraction with cookie and proxy support alongside a new structured DOM model. The scraping DSL was redesigned to use explicit extractors and type-safe combinators, replacing implicit conversions with a more composable and readable pipeline. These changes were accompanied by the addition of configuration-driven extraction capabilities and extensive test coverage to validate the new APIs.
Features
Add DeepFunctor utility and resource management helper
The library introduces a new \DeepFunctor\ type class in the \util\ package, enabling the deconstruction of nested type constructors into a composite functor for advanced generic operations. Additionally, a \using\ helper method is added to the \util\ package to simplify safe resource management by automatically closing \Closeable\ objects after execution.
core/src/main/scala/net/ruippeixotog/scalascraper/util · high confidence
Configuration-driven HTML extraction and validation
Users can now define HTML extractors and validators via Typesafe Config files instead of writing code. This change introduces \ConfigHtmlExtractor\ and \ConfigHtmlValidator\ objects that parse configuration paths (such as CSS queries, attribute selectors, date formats, and match conditions) to build scraper components. The \ConfigLoaders\ DSL provides convenient methods like \extractorAt\ and \validatorAt\ to load these components from configuration, enabling declarative scraper definitions.
modules/config/src/main · high confidence
Initial project setup and documentation for Scala Scraper 3.2.0
This change introduces the initial project structure and documentation for Scala Scraper version 3.2.0. It adds configuration files for code formatting (scalafmt) and import organization (scalafix), sets the project version to 3.2.1-SNAPSHOT, and includes a comprehensive CHANGELOG detailing the history from 0.1 to 3.2.0. The README provides a quick start guide, core model documentation, and usage examples for the DSL, browsers (Jsoup and HtmlUnit), and content extraction features. It also updates the .gitignore to include standard build and IDE artifacts.
(repo-wide) · high confidence
Introduction of core DOM model abstractions
The library now exposes a structured model layer for HTML parsing, introducing traits for Document, Element, Node, and ElementQuery. Users can now interact with a Document's location, title, and body, and traverse the DOM via Element methods such as text, ownText, and CSS selection. The Node trait distinguishes between ElementNode and TextNode, allowing access to raw child and sibling nodes alongside structured element queries.
core/src/main/scala/net/ruippeixotog/scalascraper/model · high confidence
Unified browser abstraction with cookie and proxy support
The library introduces a common \Browser\ trait that standardizes the interface for web scraping implementations, including \HtmlUnitBrowser\ (which executes JavaScript) and \JsoupBrowser\ (which parses static HTML). This change adds first-class support for managing cookies via \cookies\, \setCookie\, and \clearCookies\ methods, and allows users to configure HTTP or SOCKS proxies using a new \Proxy\ case class and the \withProxy\ method on browser instances.
core/src/main/scala/net/ruippeixotog/scalascraper/browser · high confidence
Removals
Removal of legacy scraping DSL and configuration support
The library has removed the legacy HTML scraping DSL and its associated configuration loading infrastructure. This includes the deletion of the \application.conf\ example file, the \Examples.scala\ demonstration code, and core DSL components such as \ScrapingOps\, \ImplicitConversions\, and \ConfigLoadingHelpers\. Additionally, foundational scraper types like \HtmlExtractor\, \HtmlStatusMatcher\, and utility traits like \Mappable\ and \Validated\ have been removed from the \src/main\ package, indicating a significant restructuring or deprecation of the previous extraction and validation API.
src/main · high confidence
Behavioural changes
Redesigned DSL with explicit extractors and new validation operators
The scraping DSL has been refactored to replace previous implicit conversions with explicit \HtmlExtractor\ combinators and a new \ToQuery\ type class for type-safe element queries. Users now interact with the DSL through the \DSL\ object, which provides factory methods for extractors and aliases for content parsers. The \ScrapingOps\ trait introduces new operator-based syntax for extraction (\\>\>\, \\>?\>\) and validation (\\>/\~\), allowing for more composable and readable scraping pipelines while deprecating older, less explicit patterns.
core/src/main/scala/net/ruippeixotog/scalascraper/dsl · high confidence
Refactored extraction API with new HtmlExtractor and PolyHtmlExtractor types
The scraper's extraction logic has been restructured to use a new \HtmlExtractor\ trait and a \PolyHtmlExtractor\ for polymorphic type-safe extraction, replacing the previous \SimpleExtractor\ approach. This change introduces \ContentExtractors\ and \ContentParsers\ as dedicated objects for primitive data extraction (text, attributes, tables) and parsing (dates, regex), providing a more composable and type-safe DSL for users. The \HtmlValidator\ has also been updated to work with these new extractor types, allowing validation logic to be built on top of the new extraction primitives.
core/src/main/scala/net/ruippeixotog/scalascraper/scraper · high confidence
Test coverage
Added HTML test fixtures for encoding, structure, and JavaScript behavior; Added comprehensive browser test suite and test infrastructure; Added test infrastructure and example applications for proxying and DSL usage; Added tests for ElementQuery CSS selection and equality; Added tests for configuration-based extractors and validators; Added tests for the scraping DSL extraction, validation, and typing features.
Dependencies
Upgrade SBT build tool to 1.13.0 and update plugins
The SBT build tool version has been upgraded from 0.13.6 to 1.13.0. Additionally, several SBT plugins have been updated to newer versions: sbt-scalafix to 0.14.9, sbt-release to 1.5.0, sbt-pgp to 2.3.2, sbt-mdoc to 2.9.2, sbt-scalafmt to 2.6.2, sbt-coveralls to 1.3.15, and sbt-scoverage to 2.4.4.
project · high confidence
Upgrade to Scala 2.13/3.9 and modernize dependencies
The build system has been upgraded to support Scala 2.13.18 and Scala 3.9.0, dropping older Scala versions. Core dependencies have been updated to their latest versions, including HtmlUnit 5.5.0, Jsoup 1.23.2, nscala-time 3.2.0, and scalaz-core 7.3.9. The project structure has been converted to a multi-project build with a separate \config\ module, and test dependencies like specs2-core (4.23.0) and akka-http (10.2.10) have also been updated.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 68 → 65 (-2.6)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 99 → 99 (+0.0)
- Architecture 97 → 69 (-27.8)
- Maturity 56 → 56 (+0.0)
- Readiness 73 → 75 (+2.8)
- Security 73 → 73 (+0.0)
Resolved (3)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- Off-boarding risk: anonymized user #1
New (1)
- Off-boarding risk: anonymized user #1
Changes since last survey
- 3 commits — 3 feature/other, 0 fixes
By area
- (root) — 2 commits
- project/plugins.sbt — 1 commit
Notable commits
- change: Update nscala-time to 3.2.0 (#694)
- change: Update sbt-scalafix to 0.14.9 (#693)
- change: Update slf4j-nop to 2.0.20 (#696)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ruippeixotog/scala-scraper was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 694d99ef0b20ebc9dfae9240b1998be370d78726 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.