Skip to content
CAI
Software that uses CAICheck a score

helgeho/ArchiveSpark

51.2

Adequate · 20 September 2026

3.5k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

ArchiveSpark is a Scala-based library built on Apache Spark for downloading, loading, and analyzing web archive data. It provides core functionality for processing WARC and CDX files, including extracting hyperlinks, identifying embedded resources, and building text corpora. The system leverages the Sparkling library for configuration management and supports integration with Stanford CoreNLP for natural language processing tasks.

Features

New ArchiveSpark notebooks for web archive analysis and data preparation

Added a suite of new Jupyter notebooks that demonstrate how to use ArchiveSpark to download, load, and analyze web archive data. These recipes cover downloading WARC/CDX files from the Wayback Machine, generating CDX metadata for efficient processing, and performing common analysis tasks such as extracting hyperlinks, identifying embedded resources (like stylesheets), counting term distributions, and building text corpora from selected URLs. The notebooks also include a demo for extracting page titles and an example adapted from the IEEE BigData 2017 conference, illustrating both legacy and modern loading patterns for HDFS-stored archives.

notebooks · high confidence

Behavioural changes

ArchiveSpark library refactored to depend on Sparkling

The ArchiveSpark codebase has been restructured to remove internal code duplication by adopting the Sparkling library as a dependency. This change introduces a new initialization flow in ArchiveSpark that registers Kryo classes and delegates configuration and partition management to Sparkling, while also introducing a new DistributedConfig class to manage runtime settings such as exception handling and WARC decompression limits.

src/main · high confidence

Dependencies

Initial build configuration for ArchiveSpark 3.3.8

The project now uses a build.sbt file to define the ArchiveSpark library (version 3.3.8-SNAPSHOT) built against Scala 2.12.8. It declares provided dependencies on Apache Spark 2.4.5 (core and SQL), the internal sparkling library (0.3.8-SNAPSHOT), and Stanford CoreNLP (4.3.1), establishing the baseline environment for users running the tool.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 51.

Lenses

  • Code Health 98
  • Architecture 99
  • Maturity 49
  • Readiness 26
  • Security 100

Changes since last survey

  • 154 commits — 109 feature/other, 45 fixes

By area

  • src/main — 113 commits
  • (root) — 27 commits
  • subprojects/benchmarking — 4 commits
  • notebooks/ContentLengthExample.ipynb — 3 commits
  • notebooks/Demo1.ipynb — 2 commits
  • notebooks/IEEE_BigData_2017.ipynb — 2 commits
  • (repo) — 1 commit
  • docs/Use_Library.md — 1 commit
  • notebooks/WebSciHackathonHandsOn.ipynb — 1 commit

Notable commits

  • fix: Add benchmarks, fix JSON output for maps, update README
  • fix: Fix README
  • fix: Fixed datetime comparison syntax;picking only one location per (W)ARC
  • fix: Update Sparkling and recent fixes
  • fix: bug fixes
  • fix: bug fixes
  • fix: bug fixes
  • fix: fix
  • fix: fix
  • fix: fix Demo1 notebook
  • fix: fix SURT util
  • fix: fix a couple of things
  • fix: fix bound function inheritance and result sorting
  • fix: fix build
  • fix: fix dependency issues with HBase's guava version and fix loading issues with WarcBase
  • fix: fix enrich dependency issues for multi-value enrich functions and http client being serialized
  • fix: fix error caused by (malformed) CDX pointing to a missing WARC file
  • fix: fix error caused by empty ARC headers
  • fix: fix exludeFromOutput for multi-value fields
  • fix: fix filterExists helper method
  • …and 134 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

helgeho/ArchiveSpark was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9ac4ac710803cb682fde8e0e832e3f1072994c01 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.