mrpowers-io/spark-daria
54.4
Adequate · 28 September 2026
4k
lines of production code
Scala
primary language
2
measurements over time
What this system is
This system is a Scala library that extends Apache Spark with custom SQL functions, DataFrame utilities, and schema manipulation tools. It provides a compatibility layer for Spark versions 3.2 through 3.5, enabling features like random number generation, date truncation, and nullability alignment via an 'unsafe' module. The codebase also includes core helpers for ETL pipelines, file I/O, and data validation, supported by a modernized build and testing infrastructure.
Features
New core utilities for DataFrame validation, schema manipulation, and file operations
This release introduces a comprehensive set of new helper objects and extensions in the core module to streamline Spark DataFrame workflows. Users can now validate DataFrames against required or prohibited columns and schemas using the new \DariaValidator\ and related checker classes, which throw descriptive exceptions on mismatch. Schema manipulation is enhanced with \StructTypeHelpers\ for building and flattening nested schemas, while \DataFrameExt\ and \ColumnExt\ provide convenient methods for casting, chaining transformations, and handling null-aware comparisons. The module also adds practical utilities for file I/O, including \FsHelpers\ for merging HDFS directories and \DariaWriters\ for writing single-file outputs, alongside new ETL framework components like \EtlDefinition\ and \Parser\ for structured data processing pipelines.
core · high confidence
New unsafe utilities for schema alignment, column conversions, and random number generation
This change introduces several new components in the \unsafe\ project to support Spark compatibility and data manipulation. \BebeFunctions\ adds a suite of SQL functions (such as \beginningOfDay\, \bebe\_approx\_percentile\, and \bebe\_count\_if\) that wrap Catalyst expressions, providing capabilities missing from the standard Scala API. \DataFrameExt\ introduces an extension method \toSchemaWithNullabilityAligned\ that allows DataFrames to be converted to a target schema while explicitly aligning column nullability. \ShimColumnConversions\ provides a shim layer to handle \Column\ to \Expression\ conversion, preparing the codebase for Spark 4.0 where \Column.expr\ is removed. Finally, \XORShiftRandomAdapted\ adds a random number generator class adapted from Spark's internal \Rand\ to use Apache Commons Math 3.
unsafe/src/main/scala · high confidence
Removals
Removal of legacy date and range helper functions
The \yeardiff\ and \between\ helper functions have been removed from the \com.github.mrpowers.spark.daria.sql.functions\ package. Users relying on these specific utilities for calculating year differences or checking column value ranges will need to replace them with standard Spark SQL functions or custom implementations.
src/main · high confidence
Behavioural changes
Add Spark 3.2/3.3 shim implementations for catalyst expressions and column conversions
This change introduces new shim-layer files for the Spark 3.2 and 3.3 compatibility layer, adding implementations for \BeginningOfMonth\, \KnownNotNull\, \KnownNullable\, and \RandGamma\ catalyst expressions, as well as a \ColumnConversions\ helper. These additions enable the library to support specific date truncation, nullability metadata handling, gamma random number generation, and column-to-expression conversion APIs required by these Spark versions.
_unsafe/src/main/spark\_3.2\3.3 · high confidence
Add Spark 3.4/3.5 shim implementations for native functions
This update introduces specific Scala implementations for the unsafe module targeting Spark 3.4 and 3.5. It adds support for the \rand\_gamma\ random distribution function, the \beginning\_of\_month\ date truncation function, and the \left\/\right\ string manipulation functions (via \BebeLeft\ and \BebeRight\ wrappers). Additionally, it provides a \ColumnConversions\ shim to expose the internal \expr\ field of a \Column\ object, ensuring compatibility with the Catalyst expression engine in these Spark versions.
_unsafe/src/main/spark\_3.4\3.5 · high confidence
Major documentation overhaul and CI migration to GitHub Actions
The project has significantly expanded its README with comprehensive setup instructions, usage examples for core extensions, column functions, and transformations, and detailed publishing guidelines. Documentation for specific methods has been moved from the README into Scaladoc comments. The CI pipeline has migrated from Travis CI to GitHub Actions, reflected by the addition of a GitHub Actions badge and the removal of Travis references. Additionally, the code formatting tool has been switched from scalariform to scalafmt, introducing a new .scalafmt.conf configuration file.
(repo-wide) · high confidence
Fixes
Added Spark 3.2 and 3.3 specific implementations for BebeLeft and BebeRight expressions
Introduced version-specific Scala implementations for the BebeLeft and BebeRight SQL expressions to ensure compatibility across Spark 3.2 and 3.3. For Spark 3.2, these objects provide custom logic for substring operations, including null and length checks, while the Spark 3.3 versions delegate directly to the underlying expression constructors.
_unsafe/src/main/spark\_3.2, unsafe/src/main/spark\3.3 · high confidence
Test coverage
Added test suite for Bebe functions and DataFrame extensions; Removed legacy FunctionsSpec test file.
Dependencies
Modernized build configuration with Spark 3.3.4 and Scala 2.12/2.13 support
The build system has been significantly refactored to support Apache Spark 3.3.4 and cross-compilation for Scala 2.12.20 and 2.13.15, replacing the previous Spark 2.1.0 and Scala 2.11.7 setup. The project structure now includes a 'core' module and an 'unsafe' (bebe) module, with the latter dynamically including source directories based on the Spark major/minor version. Test dependencies have been updated to use spark-fast-tests 3.0.1, utest 0.8.2, and os-lib 0.10.3, replacing the older spark-testing-base and scalatest dependencies. Build settings now include automatic source formatting on compile, forked test execution, and improved publishing configurations.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 52 → 54 (+2.3)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 97 → 97 (+0.0)
- Architecture 100 → 100 (+0.0)
- Maturity 35 → 40 (+4.9)
- Readiness 50 → 50 (+0.0)
- Security 73 → 73 (+0.0)
Resolved (1)
- Off-boarding risk: anonymized user #1
New (1)
- Off-boarding risk: anonymized user #1
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
mrpowers-io/spark-daria was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit eb0cc1d91337493b0aa87f7fb8c441eea84bae46 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.