Skip to content
CAI
Software that uses CAICheck a score

archivesunleashed/aut

55.8

Adequate · 20 September 2026

4.2k

lines of production code

Scala

with Python

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Scala-based toolkit for processing and analyzing web archives using Apache Spark. It provides utilities to load WARC and ARC files, extract metadata and content from diverse file types, and compute hashes or detect MIME types. The toolkit exposes these capabilities through both a Scala API and a Python interface, supporting data export in formats like CSV, Parquet, GEXF, and GraphML.

Features

Added Spark log4j configuration template

A new log4j.properties file has been added to the resources directory to configure logging for Spark applications. This configuration sets the default log level to WARN for the console and applies specific levels to quiet verbose third-party logs (such as Jetty and Parquet) and suppresses annoying messages related to Hive metastore and UDF lookups, improving the clarity of application output.

src/main/resources · high confidence

New DataFrame-based extractors and graph export formats

The command-line application now supports a comprehensive suite of new extractors that process web archives using Spark DataFrames, including Audio, CSS, HTML, JavaScript, JSON, PDF, Plain Text, Presentation Programs, Spreadsheets, Videos, Word Processors, and XML. It also introduces Domain and Image Graph extractors, a Plain Text extractor that filters for boilerpipe text, and an Extract Popular Images component. For graph analysis, the DomainGraphExtractor now supports exporting results in GEXF and GraphML formats in addition to the existing CSV and Parquet options.

src/main/scala/io/archivesunleashed/app · high confidence

New DataFrameLoader API and SaveBytes utility for PySpark integration

This change introduces the \DataFrameLoader\ class in \src/main/scala/io/archivesunleashed/df\, providing a wrapper around \RecordLoader\ to expose archive data as Spark DataFrames with specific schemas for various file types (e.g., \all\, \audio\, \html\, \images\, \pdfs\, \webpages\). It also adds a \SaveBytes\ implicit class to the \df\ package, enabling users to save binary content from a DataFrame to disk by specifying column names for bytes, filename, and extension, using MD5 hashing for unique file naming.

src/main/scala/io/archivesunleashed/df · high confidence

New Python API for Archive Unleashed Spark operations

The \src/main/python/aut\ package now provides a Python interface for Archive Unleashed, exposing PySpark DataFrames and UDFs. Users can load various web archive file types (HTML, images, PDFs, etc.) via the \WebArchive\ class, which includes methods like \webpages()\, \images()\, and \imagegraph()\. The package also introduces new application functions \ExtractPopularImages\, \SaveBytes\, \WriteGEXF\, and \WriteGraphML\ for processing and exporting data, alongside a suite of UDFs in \aut.udfs\ for tasks such as computing hashes, detecting MIME types, and extracting links.

src/main/python · high confidence

New utility objects for web archive data processing

This release introduces a suite of new utility objects in the \io.archivesunleashed.matchbox\ package to support web archive analysis. Key additions include \ComputeImageSize\ and \ExtractImageDetails\ for determining image dimensions and metadata, \ComputeMD5\ and \ComputeSHA1\ for generating content checksums, and \DetectLanguage\ and \DetectMimeTypeTika\ for identifying content language and MIME types using Apache Tika. The update also adds \ExtractBoilerpipeText\ for cleaning HTML content, \ExtractDate\ and \CovertLastModifiedDate\ for parsing and formatting timestamps, and \ExtractDomain\ for resolving hostnames using the public suffix list. Additional utilities include \ExtractImageLinks\ and \ExtractLinks\ for parsing HTML with Jsoup, \ExtractTextFromPDFs\ for PDF text extraction, \GetExtensionMIME\ for file extension resolution, and \RemoveHTML\/\RemoveHTTPHeader\ for content sanitization.

src/main/scala/io/archivesunleashed/matchbox · high confidence

Behavioural changes

ArchiveRecord interface and Sparkling-based implementation introduced

The core archive record abstraction has been refactored: \ArchiveRecord\ is now a \Serializable\ trait defining the record interface, and \SparklingArchiveRecord\ provides the concrete implementation backed by the Sparkling library. This change replaces the previous Java/ARC-based processing with a Sparkling-based WARC loader, ensuring records are serializable and consistent across the codebase.

src/main/scala/io/archivesunleashed · high confidence

Centralized UDF registration for Spark DataFrames

Users can now access Archives Unleashed's extraction, detection, and filtering functions through a single \udfs\ package object. This change consolidates previously scattered definitions, allowing direct access to Matchbox utilities (such as \computeMD5\, \extractDomain\, and \detectLanguage\) and filter functions (like \hasDate\ and \hasMIMETypes\) via \import io.archivesunleashed.udfs.\_\, simplifying DataFrame transformations in both Scala and PySpark.

src/main/scala/io/archivesunleashed/udfs · high confidence

Test coverage

Added Scala test suite for Archives Unleashed core components; Added WARC test fixture for wget redirect handling; Added comprehensive test coverage for DataFrame extractors and UDFs; Added test coverage for app extractors and command-line interface; Added unit tests for Matchbox utility functions.

Dependencies

Archives Unleashed Toolkit 1.2.1-SNAPSHOT dependency baseline

The project's Maven build (pom.xml) establishes a dependency baseline for the Archives Unleashed Toolkit, targeting Java 11 and Scala 2.12.10. It integrates Apache Spark 3.0.1, Apache Tika 1.23, and Guava 32.0.0-jre, while configuring the Maven Shade plugin to produce a fat JAR with relocated Guava classes and excluded Hadoop/Spark artifacts to prevent runtime conflicts.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 56.

Lenses

  • Code Health 97
  • Architecture 100
  • Maturity 42
  • Readiness 63
  • Security 57

Changes since last survey

  • 300 commits — 285 feature/other, 15 fixes

By area

  • (root) — 137 commits
  • src/main — 115 commits
  • src/test — 42 commits
  • .github/ISSUE_TEMPLATE — 3 commits
  • .github/workflows — 2 commits
  • .github/PULL_REQUEST_TEMPLATE.md — 1 commit

Notable commits

  • fix: Add ExtractGraphTest; lint fixes on RemoveHttpHeaderTest. (#92)
  • fix: Add and update tests, resolve textFiles bug. (#388)
  • fix: [CVE redacted] fix. (#281)
  • fix: Fix README badges.
  • fix: Fix Scaladocs build. (#523)
  • fix: Fix TravisCI build issues (#244)
  • fix: Fix bug -- label type should be "string" not "label". (#166)
  • fix: Fix bug and unit test for ExtractDomain; resolves #277 (#278)
  • fix: Fix codecov GitHub action. (#536)
  • fix: Fix exception error when processing corrupted ARC files, and empty files. (#272)
  • fix: Fix relative links extraction (#504)
  • fix: Gexf Fixes & StringUtil Functions #172 (#173)
  • fix: Minor fix to improve coverage. #55 (#98)
  • fix: Revert "make ArchiveRecord a trait (#175)" (#181)
  • fix: Update Bug report template. (#268)
  • change: Add graphml output to CommandLineApp and DomainGraphExtractor. (#438)
  • change: Add office document binary extraction. (#346)
  • change: Add option to save to Parquet for app. (#454)
  • change: - Add tests for RecordLoader (#149)
  • change: - Include test for flatten.default (#145)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

archivesunleashed/aut was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 57c9b5ed167f97c9ea5e08ebd1a4a958c6da1818 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.