G-Research/spark-extension
73.1
Strong · 20 September 2026
7.1k
lines of production code
Scala
with Python
1
measurement over time
What this system is
This system is a Spark extension library that provides utilities for data comparison, metadata inspection, and transformation. It enables users to diff DataFrames with customizable comparators, compute histograms, generate global row numbers, and read Parquet file metadata. The project supports Scala, Java, and Python, with the Python package bundling the necessary Scala artifacts for seamless integration.
How it got here
2020 — Spark 3.5 upgrade and Python API launch
7 changes.
The project upgraded its core dependencies to Spark 3.5 and Scala 2.13 while introducing a new Python extension API for DataFrame diffing and Parquet metadata inspection. This period also established comprehensive documentation, formatting standards, and test suites for both Scala and Python components to support the new features and build systems.
2022–2026 — Python packaging and multi-version testing
5 changes.
This period focused on releasing the official Python package with integrated build processes and documentation, alongside adding Java integration tests for the Spark Diff library. It also involved establishing test compatibility helpers and directory structures for upcoming Spark 4 and 5 versions, as well as providing examples for installing Python dependencies in Spark environments.
Features
Add example for installing Python dependencies in Spark
A new example located in the \examples/python-deps\ directory demonstrates how to install Python packages (specifically \pandas\ and \pyarrow\) into a Spark environment using the \install\_pip\_package\ API. The example includes a Dockerfile based on Apache Spark 3.5.0, a docker-compose configuration for a local Spark master and worker setup, and a Python script (\example.py\) that initializes a Spark session, installs the dependencies, and executes a pandas UDF.
examples · high confidence
Initial Python package release with integrated build and documentation
The \python/\ directory now contains the official \pyspark-extension\ package, including \setup.py\ and \README.md\. The \setup.py\ script integrates the Scala JAR build process, automatically compiling the Spark extension JAR and bundling it into the Python distribution, while declaring \typing\_extensions\ as a runtime dependency and supporting Python 3.7 through 3.13. The accompanying \README.md\ provides comprehensive installation instructions for PyPI, PySpark API, REPL, and \spark-submit\, and documents available features such as dataset diffing, histogram computation, global row numbering, Parquet metadata inspection, and .Net DateTime conversion.
python · high confidence
Initial release of Python Spark extension API
This change introduces the initial Python API for the G-Research Spark extension, providing new capabilities for data comparison and metadata inspection. Users can now compare DataFrames using the \diff\ module, which supports various modes (ColumnByColumn, SideBySide, etc.), sparse output, and custom comparators for strings, maps, and numeric values with epsilon tolerance. Additionally, the \parquet\ module exposes functions to read detailed metadata from Parquet files, including file-level statistics, schema details, block information, and partition data, with support for controlling parallelism during metadata reads.
python/gresearch · high confidence
New utility functions and diffing capabilities
This release introduces several new capabilities for data processing and comparison. Users can now generate global row numbers with the new \RowNumbers\ transformation, which supports custom ordering and storage levels. A \Histogram\ transformation is added to compute column histograms aggregated by other columns. The diffing API is expanded with new \DiffComparators\ for epsilon, duration, whitespace-agnostic strings, and maps, alongside a new \App\ for command-line diffing of datasets. Additionally, the \Backticks\ utility simplifies quoting column names with special characters, and a \ConditionalCall\ API enables fluent conditional execution of transformations.
src/main · high confidence
Project documentation and configuration files added
The repository now includes a \.scalafmt.conf\ file to enforce code formatting rules (Scala 2.13, max column 120), a \LICENSE\ file adopting the Apache License 2.0, and a \CHANGELOG.md\ documenting the project's history and upcoming changes. Additionally, several documentation files have been added to explain library features: \DIFF.md\ for the dataset diffing API, \CONDITIONAL.md\ for fluent conditional transformations, \GROUPS.md\ for sorted group operations, \HISTOGRAM.md\ for histogram computations, \PARQUET.md\ for reading Parquet metadata, \PARTITIONING.md\ for optimized partitioned writing, \PYSPARK-DEPS.md\ for installing Python dependencies in PySpark jobs, \MAINTAINERS.md\ listing project maintainers, and \RELEASE.md\ detailing the release process.
(repo-wide) · high confidence
Test coverage
Added Java integration tests for Spark Diff and utility functions; Added Python test suite for Spark extension features; Added Scala test suites for new group-by, histogram, partitioned write, and diff features; Added Spark 4 test compatibility helpers; Added Spark 5 test symlink; Added test logging configuration files.
Dependencies
Upgrade to Spark 3.5 and Scala 2.13 with new Python build system
The project has upgraded its core dependencies to Apache Spark 3.5.1 and Scala 2.13.8, replacing the previous Spark 2.4/Scala 2.12 baseline. This change includes adding explicit support for Spark Hive and managing Parquet Hadoop dependencies (version 1.16.0) to ensure compatibility. Additionally, a new Python build configuration (pyproject.toml) using setuptools has been introduced alongside the existing Maven build, and test dependencies have been updated to ScalaTest 3.3.0-SNAP4 and JUnit 4.13.2.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 73.
Lenses
- Code Health 99
- Architecture 100
- Maturity 66
- Readiness 72
- Security 78
Changes since last survey
- 300 commits — 256 feature/other, 44 fixes
By area
- (root) — 118 commits
- .github/workflows — 86 commits
- .github/actions — 33 commits
- src/main — 33 commits
- src/test — 13 commits
- python/gresearch — 5 commits
- python/README.md — 3 commits
- python/test — 3 commits
- .github/show-spark-versions.sh — 2 commits
- python/setup.py — 2 commits
- (repo) — 1 commit
- examples/python-deps — 1 commit
Notable commits
- fix: Add fix versions for partitioned write issue (#130)
- fix: Add more Python code to DIFF.md, fix broken examples
- fix: CI: Fix warnings & deprecation messages in the workflows (#241)
- fix: Fix 2.12 scala version for 3.5.8-SNAPSHOT test (#296)
- fix: Fix FILES deprecation warning
- fix: Fix PyPi release workflow (#321)
- fix: Fix Python API calling into Scala code (#132)
- fix: Fix Python comparator option methods in DIFF.md
- fix: Fix Spark 3.5 patch version in test-jvm (#347)
- fix: Fix bump-version.sh
- fix: Fix cache condition in test-jvm action
- fix: Fix cache restore-keys, fix Python test version (#305)
- fix: Fix cache save keys
- fix: Fix changelog (#338)
- fix: Fix clear caches pagination (#262)
- fix: Fix condition for PyPi release (#319)
- fix: Fix consolidating success job (#303)
- fix: Fix detection of python test failure (#207)
- fix: Fix expected error message for 4.0.0 SNAPSHOT (#203)
- fix: Fix for DataFrame._sort_cols changes in PySpark4 (#272)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
G-Research/spark-extension was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 24fc7308d289e53e0cff2c3793ed4307fbf33b2f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.