archivesunleashed/aut
55.8
Adequate · 20 September 2026
4.2k
lines of production code
Scala
with Python
1
measurement over time
What this system is
This system is a Scala-based toolkit for processing and analyzing web archives using Apache Spark. It provides utilities to load WARC and ARC files, extract metadata and content from diverse file types, and compute hashes or detect MIME types. The toolkit exposes these capabilities through both a Scala API and a Python interface, supporting data export in formats like CSV, Parquet, GEXF, and GraphML.
Features
Added Spark log4j configuration template
A new log4j.properties file has been added to the resources directory to configure logging for Spark applications. This configuration sets the default log level to WARN for the console and applies specific levels to quiet verbose third-party logs (such as Jetty and Parquet) and suppresses annoying messages related to Hive metastore and UDF lookups, improving the clarity of application output.
src/main/resources · high confidence
New DataFrame-based extractors and graph export formats
The command-line application now supports a comprehensive suite of new extractors that process web archives using Spark DataFrames, including Audio, CSS, HTML, JavaScript, JSON, PDF, Plain Text, Presentation Programs, Spreadsheets, Videos, Word Processors, and XML. It also introduces Domain and Image Graph extractors, a Plain Text extractor that filters for boilerpipe text, and an Extract Popular Images component. For graph analysis, the DomainGraphExtractor now supports exporting results in GEXF and GraphML formats in addition to the existing CSV and Parquet options.
src/main/scala/io/archivesunleashed/app · high confidence
New DataFrameLoader API and SaveBytes utility for PySpark integration
This change introduces the \DataFrameLoader\ class in \src/main/scala/io/archivesunleashed/df\, providing a wrapper around \RecordLoader\ to expose archive data as Spark DataFrames with specific schemas for various file types (e.g., \all\, \audio\, \html\, \images\, \pdfs\, \webpages\). It also adds a \SaveBytes\ implicit class to the \df\ package, enabling users to save binary content from a DataFrame to disk by specifying column names for bytes, filename, and extension, using MD5 hashing for unique file naming.
src/main/scala/io/archivesunleashed/df · high confidence
New Python API for Archive Unleashed Spark operations
The \src/main/python/aut\ package now provides a Python interface for Archive Unleashed, exposing PySpark DataFrames and UDFs. Users can load various web archive file types (HTML, images, PDFs, etc.) via the \WebArchive\ class, which includes methods like \webpages()\, \images()\, and \imagegraph()\. The package also introduces new application functions \ExtractPopularImages\, \SaveBytes\, \WriteGEXF\, and \WriteGraphML\ for processing and exporting data, alongside a suite of UDFs in \aut.udfs\ for tasks such as computing hashes, detecting MIME types, and extracting links.
src/main/python · high confidence
New utility objects for web archive data processing
This release introduces a suite of new utility objects in the \io.archivesunleashed.matchbox\ package to support web archive analysis. Key additions include \ComputeImageSize\ and \ExtractImageDetails\ for determining image dimensions and metadata, \ComputeMD5\ and \ComputeSHA1\ for generating content checksums, and \DetectLanguage\ and \DetectMimeTypeTika\ for identifying content language and MIME types using Apache Tika. The update also adds \ExtractBoilerpipeText\ for cleaning HTML content, \ExtractDate\ and \CovertLastModifiedDate\ for parsing and formatting timestamps, and \ExtractDomain\ for resolving hostnames using the public suffix list. Additional utilities include \ExtractImageLinks\ and \ExtractLinks\ for parsing HTML with Jsoup, \ExtractTextFromPDFs\ for PDF text extraction, \GetExtensionMIME\ for file extension resolution, and \RemoveHTML\/\RemoveHTTPHeader\ for content sanitization.
src/main/scala/io/archivesunleashed/matchbox · high confidence
Behavioural changes
ArchiveRecord interface and Sparkling-based implementation introduced
The core archive record abstraction has been refactored: \ArchiveRecord\ is now a \Serializable\ trait defining the record interface, and \SparklingArchiveRecord\ provides the concrete implementation backed by the Sparkling library. This change replaces the previous Java/ARC-based processing with a Sparkling-based WARC loader, ensuring records are serializable and consistent across the codebase.
src/main/scala/io/archivesunleashed · high confidence
Centralized UDF registration for Spark DataFrames
Users can now access Archives Unleashed's extraction, detection, and filtering functions through a single \udfs\ package object. This change consolidates previously scattered definitions, allowing direct access to Matchbox utilities (such as \computeMD5\, \extractDomain\, and \detectLanguage\) and filter functions (like \hasDate\ and \hasMIMETypes\) via \import io.archivesunleashed.udfs.\_\, simplifying DataFrame transformations in both Scala and PySpark.
src/main/scala/io/archivesunleashed/udfs · high confidence
Test coverage
Added Scala test suite for Archives Unleashed core components; Added WARC test fixture for wget redirect handling; Added comprehensive test coverage for DataFrame extractors and UDFs; Added test coverage for app extractors and command-line interface; Added unit tests for Matchbox utility functions.
Dependencies
Archives Unleashed Toolkit 1.2.1-SNAPSHOT dependency baseline
The project's Maven build (pom.xml) establishes a dependency baseline for the Archives Unleashed Toolkit, targeting Java 11 and Scala 2.12.10. It integrates Apache Spark 3.0.1, Apache Tika 1.23, and Guava 32.0.0-jre, while configuring the Maven Shade plugin to produce a fat JAR with relocated Guava classes and excluded Hadoop/Spark artifacts to prevent runtime conflicts.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 56.
Lenses
- Code Health 97
- Architecture 100
- Maturity 42
- Readiness 63
- Security 57
Changes since last survey
- 300 commits — 285 feature/other, 15 fixes
By area
- (root) — 137 commits
- src/main — 115 commits
- src/test — 42 commits
- .github/ISSUE_TEMPLATE — 3 commits
- .github/workflows — 2 commits
- .github/PULL_REQUEST_TEMPLATE.md — 1 commit
Notable commits
- fix: Add ExtractGraphTest; lint fixes on RemoveHttpHeaderTest. (#92)
- fix: Add and update tests, resolve textFiles bug. (#388)
- fix: [CVE redacted] fix. (#281)
- fix: Fix README badges.
- fix: Fix Scaladocs build. (#523)
- fix: Fix TravisCI build issues (#244)
- fix: Fix bug -- label type should be "string" not "label". (#166)
- fix: Fix bug and unit test for ExtractDomain; resolves #277 (#278)
- fix: Fix codecov GitHub action. (#536)
- fix: Fix exception error when processing corrupted ARC files, and empty files. (#272)
- fix: Fix relative links extraction (#504)
- fix: Gexf Fixes & StringUtil Functions #172 (#173)
- fix: Minor fix to improve coverage. #55 (#98)
- fix: Revert "make ArchiveRecord a trait (#175)" (#181)
- fix: Update Bug report template. (#268)
- change: Add graphml output to CommandLineApp and DomainGraphExtractor. (#438)
- change: Add office document binary extraction. (#346)
- change: Add option to save to Parquet for app. (#454)
- change: - Add tests for RecordLoader (#149)
- change: - Include test for flatten.default (#145)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
archivesunleashed/aut was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 57c9b5ed167f97c9ea5e08ebd1a4a958c6da1818 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.