holdenk/spark-testing-base
62.3
Adequate · 27 September 2026
4.2k
lines of production code
Scala
with Python
4
measurements over time
What this system is
This system is a testing utility library for Apache Spark, providing infrastructure to validate Spark applications across multiple versions including 2.4, 3.0, 4.0, and Spark Connect. It offers base test classes and assertion helpers for DataFrames, Datasets, RDDs, and streaming operations in both Scala and Python. The library facilitates consistent testing by managing SparkContext lifecycles, handling embedded Kafka integration, and standardizing logging configurations.
How it got here
2015 — Python testing library and Spark 4 support
5 changes.
This period focused on releasing the initial version of the spark-testing-base Python library, providing base test classes and utilities for validating Spark RDDs, DataFrames, and streaming applications. Concurrently, the project updated its build infrastructure to support Spark 4.x and JDK 17, introducing compatibility fixes and a shaded sub-project to manage classpath conflicts.
2016–2021 — Testing infrastructure and build standardization
4 changes.
This period focused on standardizing the build and testing environments by introducing local sbt launch scripts for CI consistency and consolidating Spark 2.4 testing implementations. It also enhanced integration testing capabilities by adding Kafka 0.8 utilities and standardized logging configurations to reduce noise in Spark applications.
2022–2026 — Spark version test coverage expansion
5 changes.
This period focused on expanding test infrastructure and coverage across multiple Spark versions, including 2.4, 3.0, 3.5, and 4.0. The work involved adding specific test suites for DataFrame, Dataset, RDD, and streaming operations, as well as validating compatibility with Spark Connect and expression codegen.
Features
Added local sbt launch scripts for CI usage
The build directory now includes the \sbt\ launcher script and its supporting library \sbt-launch-lib.bash\. This allows the project to download and run the specific sbt version defined in \project/build.properties\ without requiring sbt to be pre-installed on the system, facilitating consistent builds in CI environments.
build · high confidence
Initial release of spark-testing-base Python module
This change introduces the initial version (0.11.1) of the \spark-testing-base\ Python library, providing a framework for testing Spark applications. The package includes a \setup.py\ configuration that declares dependencies on \findspark\ and \pytest\, along with test requirements for \nose\ and \coverage\. It also adds a \run-tests\ script to facilitate test execution by automatically locating the Spark home directory and running tests with coverage reporting, alongside standard project files like the Apache 2.0 license and manifest template.
python · high confidence
Initial release of the SparkTestingBase Python testing library
This change introduces the \sparktestingbase\ Python package, providing a suite of base test classes to simplify writing tests for Apache Spark applications. It includes \SparkTestingBaseTestCase\ and \SparkTestingBaseReuse\ for managing SparkContext lifecycles, \SQLTestCase\ for testing Spark SQL with DataFrame comparison helpers, and \StreamingTestCase\ for testing Spark Streaming with utilities to queue input and collect results. The library also handles PySpark path configuration via \findspark\ or \SPARK\_HOME\ and reduces log noise from py4j.
python/sparktestingbase · high confidence
Behavioural changes
Added default log4j configuration for Spark environments
A new log4j.properties file has been introduced to standardize logging behavior for Spark applications. By default, the root logging level is set to WARN, and output is directed to the console error stream. The configuration also explicitly sets specific log levels for third-party libraries (such as Jetty and Parquet) and Spark REPL components to reduce noise, and addresses SPARK-9183 by setting the Hive metastore handler to FATAL to suppress messages regarding nonexistent UDFs.
log4j · high confidence
Project documentation and configuration updates
The repository now includes standard governance and contribution files (CODE\_OF\_CONDUCT.md, CONTRIBUTING.md, SECURITY.md) and a dedicated release notes file (RELEASE\_NOTES.md). The README has been expanded with installation instructions for Scala and Python, memory configuration guidance for JDK 17+, and notes on disabling parallel test execution. Additionally, a \.codecov.yml\ file has been added to exclude specific source directories and tests from coverage reports, and a \.scala-steward.conf\ file has been created to ignore updates for the \jetty-util\ dependency.
(repo-wide) · high confidence
Unified Spark 2.4 testing implementation in core
The core testing library now consolidates its Spark 2.x support into a single 2.4 implementation, replacing previous fragmented 2.2 and 2.3 versions. This change introduces a unified set of Scala and Java source files under the 2.4 directory structure, including new generators for DataFrames, Datasets, and RDDs, as well as updated base traits for streaming and dataset assertions, ensuring consistent testing capabilities across the supported 2.4.x Spark versions.
core/src/main · high confidence
Test coverage
Added Java test fixtures and dataset/RDD comparison tests for Spark 3.0; Added Kafka 0.8 test utilities; Added Spark 2.4+ testing infrastructure and Java streaming test coverage; Added Spark 2.4-specific test suite; Added Spark 4.0 preview test suite for expression codegen; Added tests for Spark Connect client integration; Added unit tests for RDD, SQL, and streaming assertions.
Dependencies
Add Spark 4 and Spark Connect support with JDK 17 compatibility
The build configuration now supports Spark 4.x (including preview versions) and integrates the Spark Connect client for distributed DataFrame testing. To ensure compatibility with modern Java runtimes, the build explicitly configures JVM arguments (via --add-opens) for JDK 17 and later, and updates specific dependencies like Netty and zstd-jni for Spark 4. Additionally, a shaded sub-project is introduced to relocate Spark SQL classes from the Connect client, preventing classpath conflicts when used alongside standard Spark SQL components.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 47 → 62 (+15.4)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 99 → 95 (-3.4)
- Architecture 69 → 100 (+31.0)
- Maturity 46 → 41 (-5.6)
- Readiness 30 → 77 (+47.4)
- Security 83 → 77 (-6.1)
Resolved (14)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (12 lines × 3) (core/src/test/3.0/java/com/holdenkarau/spark/testing/SampleJavaDatasetTest.java)
- Duplicated block (33 lines × 2) (core/src/test/3.0/java/com/holdenkarau/spark/testing/SampleJavaRDDTest.java)
- Duplicated block (35 lines × 2) (core/src/test/3.0/java/com/holdenkarau/spark/testing/SampleJavaDatasetTest.java)
- Duplicated block (35 lines × 2) (core/src/test/3.0/java/com/holdenkarau/spark/testing/SampleJavaDatasetTest.java)
- Duplicated block (7 lines × 2) (core/src/test/3.0/java/com/holdenkarau/spark/testing/SampleJavaStreamingTest.java)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- No exposed public API
- No tests found
- Test reliability not included
New (20)
- Concentrated knowledge decay
- DataFrameGenerator.getGenerator (cyclomatic 22) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/DataFrameGenerator.scala)
- DataFrameSuiteBase.approxEquals (cognitive 72) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/DataFrameSuiteBase.scala)
- DataFrameSuiteBase.approxEquals (cyclomatic 31) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/DataFrameSuiteBase.scala)
- Duplicated block (15 lines × 2) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/StreamingActionBase.scala)
- Duplicated block (8–9 lines × 2) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/StreamingSuiteBase.scala)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- MLUserDefinedType.unapply (cognitive 16) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/MLUserDefinedType.scala)
- No SBOM
- No build provenance
- No dependency advisory monitoring
- TodoComment (core/src/main/2.4/scala/com/holdenkarau/spark/testing/PerfListener.scala)
- TodoComment (core/src/main/2.4/scala/org/apache/spark/streaming/StreamingContextWithExtraInputStreamGenerators.scala)
- Unpinned build actions
- Utils.deleteRecursively (cognitive 19) (core/src/main/2.4/scala/com/holdenkarau/spark/testing/Utils.scala)
- Workflow token permissions not restricted
Changes since last survey
- 2 commits — 2 feature/other, 0 fixes
By area
- core/src — 1 commit
- project/build.properties — 1 commit
Notable commits
- change: Add support for use assertDataFrameDataEquals with map types (#473)
- change: Update sbt, sbt-dependency-tree, ... to 1.12.15 (#503)
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- repository
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
holdenk/spark-testing-base was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit cb4b10b3be9a4bb46e8fcd60c2b1e76b01b83961 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.