LucaCanali/sparkMeasure
69.5
Adequate · 28 September 2026
2.8k
lines of production code
Scala
with Python
2
measurements over time
What this system is
SparkMeasure is a library and tool for measuring and troubleshooting Apache Spark job performance, providing both Scala and Python APIs to collect stage and task metrics. It enables users to generate reports, export aggregated data as DataFrames, and integrate with external observability systems like Prometheus, InfluxDB, and Kafka through various sink listeners. The system supports continuous, out-of-band monitoring via a flight-recorder mode and includes comprehensive testing and examples for containerized environments.
Features
Added Scala example for instrumenting Spark workloads with sparkMeasure
A new example project, testSparkMeasureScala, has been added to demonstrate how to use the sparkMeasure library to instrument Scala code running on Apache Spark. The example includes source code and documentation showing how to measure stage and task metrics, print reports, and save metric data to disk using Spark 4 and Scala 2.13.
examples/testSparkMeasureScala · high confidence
Initial release of Python API for Spark performance metrics
This change introduces the Python wrapper for the sparkMeasure library, enabling users to collect and analyze Apache Spark performance data directly from Python/PySpark applications. The new API includes \StageMetrics\ and \TaskMetrics\ classes, which allow users to begin and end metric collection, generate reports, and export aggregated metrics as Spark DataFrames. Additionally, a new \jmxexport\ function is provided to publish these metrics to Dropwizard via JMX, and the package is configured as a universal wheel for broader Python compatibility.
python · high confidence
Initial release of SparkMeasure performance tool
Introduces SparkMeasure, a library and tool for measuring and troubleshooting Apache Spark job performance. The release includes Scala and Python implementations for collecting stage and task metrics, a Dockerfile for containerized testing, and documentation with examples for Jupyter notebooks and CLI usage.
(repo-wide) · high confidence
New Python examples and documentation for sparkMeasure
The examples directory now includes a comprehensive set of Python-focused resources to help users instrument and analyze Spark workloads. This includes a new README.md that serves as a central index, a dedicated Databricks example page linking to Scala and Python notebooks, and a new Colab-compatible Jupyter notebook for quick start. Additionally, a standard Python script example (test\_sparkmeasure\_python.py) and a task metrics analysis notebook are provided, all demonstrating the use of sparkMeasure version 0.28 with Scala 2.13.
examples · high confidence
New end-to-end test suite for SparkMeasure on Kubernetes
Added a comprehensive end-to-end testing workflow for SparkMeasure on Kubernetes, including scripts to build test images, provision a Kind cluster with ArgoCD, deploy Spark jobs, and verify metrics availability via a Prometheus exporter. The suite includes example PySpark applications (spark-pi, spark-sql) that demonstrate StageMetrics instrumentation and JMX-based metric export, enabling validation of the library's functionality in a realistic containerized environment.
e2e · high confidence
New flight-recorder and external-metric-sink listeners for Spark monitoring
This change introduces a suite of new Spark listeners that enable continuous, out-of-band monitoring of Spark applications. The Flight Recorder mode (FlightRecorderStageMetrics and FlightRecorderTaskMetrics) automatically captures stage and task metrics at application end, writing them to local files, Hadoop-compatible filesystems, or stdout in JSON or Java-serialized formats. Additionally, new sink listeners allow real-time metric export to external systems: InfluxDBSink writes to InfluxDB v1.x, KafkaSink and KafkaSinkV2 ship metrics to Kafka topics (with V2 adding application-level aggregates and custom labels), PushGatewaySink sends metrics to a Prometheus Push Gateway, and DropwizardMetrics exposes metrics via JMX. These components provide users with flexible, low-overhead ways to integrate Spark performance data into existing observability stacks without modifying application code.
src/main · high confidence
Test coverage
Added test coverage for KafkaSinkV2, StageMetrics, TaskMetrics, Utils, and PushGateway
New test suites have been added to verify the correctness of the KafkaSinkV2 implementation (including application start/end events, custom labels, and Spark configuration extraction), the aggregation of Stage and Task metrics, the serialization and JSON writing capabilities in Utils, and the PushGateway integration with label/metric validation and HTTP posting.
src/test · high confidence
Dependencies
Updated build configuration to Spark 3.5.8 and Scala 2.12/2.13
The main Scala build (build.sbt) now targets Apache Spark 3.5.8 and supports Scala 2.12.18 and 2.13.16, replacing the previous Spark 2.1.0 and Scala 2.11.8 setup. Jackson dependencies are pinned to version 2.15.2 to ensure compatibility with Spark 3.5.x, and the project publishes to Central Sonatype. A new example project (examples/testSparkMeasureScala) demonstrates usage with Spark 4.0.1 and spark-measure 0.28, while the Python package (python/pyproject.toml) is updated to version 0.28.0 with optional PySpark support (\>=3.0.0) and requirements.txt specifies PySpark 3.5.8.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 65 → 70 (+4.5)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 98 → 98 (-0.2)
- Architecture 100 → 100 (+0.0)
- Maturity 80 → 80 (+0.0)
- Readiness 46 → 52 (+6.4)
- Security 76 → 83 (+6.8)
Resolved (2)
- Documentation: no architecture or design documentation (docs/Flight_recorder_mode_PrometheusPushgatewaySink.md)
- Documentation: no installation or build instructions (docs/Instrument_Scala_code.md)
New (2)
- Duplicated block (9 lines × 2) (python/sparkmeasure/stagemetrics.py)
- Outdated: org.apache.kafka:kafka-clients
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
LucaCanali/sparkMeasure was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0ab0ec7079d6d9310cff9dcf58ce22ef18c2ae1e — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.