AbsaOSS/spline-spark-agent
56.1
Adequate · 20 September 2026
12.5k
lines of production code
Scala
primary language
1
measurement over time
What this system is
This system is the Spline Agent for Apache Spark, a library designed to capture and record data lineage for Spark-based data processing workloads. It provides a modular agent that integrates with various Spark versions (2.2 through 3.4) to track execution plans and data dependencies, supporting both codeless configuration and programmatic initialization in Java and Scala. The system exposes lineage data via a standardized Producer API and includes comprehensive tooling for local testing, Docker-based examples, and integration verification across diverse data sources like Delta Lake and BigQuery.
Features
Added BigQuery example application
A new example application has been added to demonstrate lineage tracking with Google BigQuery. This standalone Scala application reads data from a public BigQuery table, performs a transformation to select and limit specific columns, and writes the result back to a BigQuery table, illustrating how to integrate the Spline lineage harvester with BigQuery data sources.
examples/src/main/scala/za/co/absa/spline/example/bigquery · high confidence
Added Java batch example with programmatic Spline initialization
A new Java example job (JavaExampleJob.java) has been added to the examples directory, demonstrating how to programmatically initialize Spline lineage tracking within a Spark application. The example shows the creation of a SparkSession and the explicit invocation of SparkLineageInitializer.enableLineageTracking with an AgentConfig, providing a reference implementation for users integrating Spline into Java-based batch workloads.
examples/src/main/java · high confidence
Added batch dependency example jobs
New example jobs (JansBeerJob, MareksJob, OtherJob) were added to the batchWithDependencies example directory. These jobs demonstrate chained data processing workflows where one job's output serves as input for subsequent jobs, illustrating how lineage tracking handles dependencies between separate Spark applications.
examples/src/main/scala/za/co/absa/spline/example/batchWithDependencies · high confidence
Added local mock servers for Spline Gateway for development and testing
Developers can now test the agent without a running Spline server by using two new local mock environments: a Mockoon-based setup and a Node.js mockserver-based setup. These tools simulate the Spline Producer API endpoints (execution-events, execution-plans, and status) on localhost:8080, allowing for offline development and easier local testing of the agent's HTTP interactions.
dev · high confidence
Added pseudo DCE example jobs
New example jobs have been added to demonstrate a Data Conformance and Enrichment (DCE) workflow. The set includes J1StandardizationJob for initial data standardization, J2ConformanceJob for applying reference data mappings (country, currency, rates), and three additional jobs (MyOtherAJob, MyOtherBJob, MyOtherCJob) that process the conformed output, all with lineage tracking enabled.
examples/src/main/scala/za/co/absa/spline/example/dce · high confidence
Added sample datasets for batch processing examples
New input data files have been added to the examples directory to support batch processing demonstrations. These include domain mappings (domain.csv), a NASA astronomy catalog in XML format (nasa.xml), Wikipedia page view statistics (wikidata.csv), and economic datasets for European countries (beerConsum.csv) and global development indicators (devIndicators.csv).
examples/data · high confidence
Added shell scripts for non-Spark lineage examples
New shell scripts have been added to the examples directory to demonstrate how to record lineage for non-Spark workloads using the Spline Producer API. These examples, including \cur-rate-file-replacement-example.sh\, \non-spark-example-1.sh\, and \non-spark-example-2.sh\, show how to manually construct and POST execution plans and events to the Spline gateway, covering scenarios such as file replacement, simple job chaining, and complex multi-source join operations.
examples/src/main/shell · high confidence
Expanded batch example jobs for lineage tracking scenarios
The batch examples directory now includes a comprehensive set of new demonstration jobs showcasing various Apache Spark data processing patterns with Spline lineage tracking. These additions cover codeless initialization via Spark configuration, handling of failed executions, multi-stage data pipelines with unions, XML and Excel file I/O, RDD conversions, and complex window function operations, providing users with diverse reference implementations for integrating lineage capture into different types of batch workloads.
examples/src/main/scala/za/co/absa/spline/example/batch · high confidence
New Docker-based example suite with configurable runtime settings
The examples directory now includes a Dockerfile, entrypoint script, and run scripts that allow users to execute Spline example jobs via a containerized environment. This setup exposes environment variables such as SPLINE\_PRODUCER\_URL, SPLINE\_MODE, and DISABLE\_SSL\_VALIDATION, enabling users to configure the Spline agent's behavior (e.g., lineage mode and SSL validation) without modifying code. The entrypoint script maps these variables to JVM system properties, ensuring the examples respect the specified configuration when run via Docker.
examples · high confidence
New SparkApp base class for example applications
A new abstract SparkApp base class has been added to the examples module to standardize how example applications initialize the SparkSession. This class allows example developers to configure the application name, master URL, custom Spark configurations, and custom tags, while automatically disabling the Spark UI and setting the driver host to localhost for consistent execution environments.
examples/src/main/scala/za/co/absa/spline · high confidence
New commons library with utility extensions and configuration helpers
The commons module now includes a comprehensive set of Scala utility classes and implicit extensions to support the Spline codebase. This adds configuration helpers for safely retrieving required and optional properties (including enum support), graph traversal utilities for topological sorting of DAGs, and various language extensions for collections, options, iterators, and arrays. It also introduces resource management (ARM), immutable properties, build info retrieval, and temporary file/directory handling.
commons · high confidence
New example jobs for Delta Lake DSV2 and complex lineage scenarios
Added four new example applications to demonstrate specific lineage tracking capabilities: DeltaDSV2Job and DeltaMergeDSV2Job showcase lineage for Delta Lake operations (including CREATE, INSERT, OVERWRITE, UPDATE, MERGE, and DELETE) on Spark 3+, while GitHub718Job and GitHub738Job provide test scenarios for lineage graph alignment and complex end-to-end lineage graphs involving multiple CSV sources and iterative processing.
examples/src/main/scala/za/co/absa/spline/issue · high confidence
Project initialization with development tooling and documentation
The repository has been initialized with standard development configuration files, including an \.editorconfig\ for consistent code formatting, a \.gitignore\ for build artifacts and IDE files, and a \.sdkmanrc\ pinning the Java version to 8.0.362-amzn. A \build-all.sh\ script is provided to facilitate cross-building for Scala 2.11 and 2.12, and a \README.md\ documents the Spline Agent for Apache Spark, its usage, and compatibility matrix. Additionally, a \.sonarcloud.properties\ file configures SonarCloud to ignore duplication in test files.
(repo-wide) · high confidence
Behavioural changes
Spline Agent configuration and API definitions updated
The Spline Spark Agent now supports YAML-based configuration via the new \spline.default.yaml\ file, replacing or supplementing previous property files. This configuration introduces a \scanClasspath\ property to control plugin discovery, allows capturing failed executions via \sql.failure.capture\, and defaults to UUID v5 for execution plan IDs. Additionally, the resource directory now includes \spline-build.properties\ to expose build metadata (version, timestamp, revision) and updated OpenAPI definitions for the Spline Producer API versions 1.0, 1.1, and 1.2, reflecting changes in schema structures and endpoint paths.
core/src/main/resources · high confidence
Spline Agent initialization and configuration restructured
The Spline agent's initialization flow and configuration model have been restructured to support a more modular and programmatic setup. A new \AgentBOM\ (Bill of Materials) and \AgentConfig\ builder now centralize the creation of core components like the \LineageDispatcher\, \PostProcessingFilter\, and \IgnoredWriteDetectionStrategy\ from configuration. This is driven by a new \HierarchicalObjectFactory\ that handles component instantiation via reflection. Additionally, new configuration enums (\SplineMode\, \SQLFailureCaptureMode\, \AuthenticationType\) provide explicit control over agent behavior, such as enabling/disabling tracking and handling SQL failures, while the \SparkLineageInitializer\ has been updated to use this new configuration model for both programmatic and codeless initialization.
core/src/main/scala · high confidence
Test coverage
Added test coverage for core configuration and agent components; New integration test infrastructure and fixtures for Spline lineage verification.
Dependencies
Added Maven bundles for Spark 2.2 through 3.4
The project now provides dedicated Spline Spark Agent bundles for Spark versions 2.2, 2.3, 2.4, 3.0, 3.1, 3.2, 3.3, and 3.4. Each bundle is configured with a \dependencyManagement\ section that pins the specific transitive dependencies required by its target Spark version, ensuring compatibility and preventing classpath collisions when the agent is deployed.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 56.
Lenses
- Code Health 100
- Architecture 99
- Maturity 59
- Readiness 30
- Security 100
Changes since last survey
- 300 commits — 226 feature/other, 74 fixes
By area
- (root) — 133 commits
- core/src — 83 commits
- (repo) — 31 commits
- core/pom.xml — 24 commits
- integration-tests/src — 13 commits
- .github/workflows — 3 commits
- bundle-2.2/pom.xml — 3 commits
- examples/src — 3 commits
- integration-tests/pom.xml — 3 commits
- bundle-2.4/pom.xml — 1 commit
- bundle-3.4/pom.xml — 1 commit
- commons/src — 1 commit
- examples/pom.xml — 1 commit
Notable commits
- fix: Bugfix/agent 426 iceberg support (#434)
- fix: Bugfix/agent 450 view and attribute lineage (#454)
- fix: Bugfix/agent 574 merge into using fix (#575)
- fix: Bugfix/agent 652 excel hdfs issue (#704)
- fix: Bugfix/spline spark agent 714 classpath collision (#715)
- fix: Feature/agent 443 fix compiler warnings (#444)
- fix: Fix MergeIntoNodeBuilder (#844)
- fix: Merge pull request #582 from AbsaOSS/bugfix/spline-spark-agent-581-integration-tests
- fix: POM: fix missing reference to the bundle-3.2 POM, as a result of incomplete cherrypick in commit 001d2d54c1340e023d8a1b07728d8ae7cb59d542
- fix: fix "example" dependencies
- fix: fix merge errors
- fix: fix post merge issue
- fix: fix post-merge leftovers
- fix: fix: pom.xml to reduce vulnerabilities
- fix: fix: upgrade com.fasterxml.jackson.core:jackson-annotations from 2.12.3 to 2.15.0
- fix: fix: upgrade com.fasterxml.uuid:java-uuid-generator from 4.0.1 to 4.1.0
- fix: fix: upgrade com.fasterxml.uuid:java-uuid-generator from 4.1.0 to 4.1.1
- fix: fix: upgrade com.fasterxml.uuid:java-uuid-generator from 4.1.1 to 4.2.0
- fix: fix: upgrade com.fasterxml.uuid:java-uuid-generator from 4.2.0 to 4.3.0
- fix: fix: upgrade commons-io:commons-io from 2.14.0 to 2.15.0
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
AbsaOSS/spline-spark-agent was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit d347fa0b3b1f47fef7c56dc3594e45a90130b25c — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.