Skip to content
CAI
Software that uses CAICheck a score

nightscape/spark-excel

56.4

Adequate · 20 September 2026

3.3k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Spark Excel data source connector that enables reading and writing Excel files within Spark applications. It provides core functionality for handling multiple worksheets, schema inference, and various data formatting options, while also addressing edge cases like malformed XML and empty rows. The implementation includes comprehensive test coverage for encrypted files, error handling, and compliance with Spark's DataFrame APIs.

Behavioural changes

Excel connector refactored to new package and rewritten core logic

The Excel data source has been migrated from the \com.crealytics\ to the \dev.mauch\ package, introducing a complete rewrite of the core reading and writing components. This change replaces the previous implementation with new classes such as \DataLocator\ (handling cell range and table addressing), \DataColumn\ (managing cell value extraction and type inference), and \ExcelRelation\ (orchestrating the Spark relation). The rewrite introduces support for reading multiple worksheets, fixes handling of faulty dimension tags and sheet names consisting of digits, and adds a \usePlainNumberFormat\ option to remove decimal points from integer values. It also improves memory management via \rowCacheSize\ for streaming reads and fixes bugs related to empty rows and undefined row handling.

src/main/scala · high confidence

Updated Spark DataSource registration to new package path

The Spark DataSource registration file has been updated to point to the new package location dev.mauch.spark.excel.v2.ExcelDataSource, reflecting the internal package rename from com.crealytics to dev.mauch and the restructuring of the Excel v2 module.

src/main/resources · high confidence

Test coverage

Added Log4j2 test configuration for Spark 3 compatibility; Added comprehensive test suite for spark-excel v2 data source; Added test resources for handling faulty dimension tags and blank rows.

Dependencies

Build system migrated to Mill 1.0 with Scala 2.12/2.13 support

The project has switched its build tool from SBT to Mill, updating the build definition to version 1.0 (build.mill). This change enables cross-compilation for both Scala 2.12 and 2.13, with specific source folders and compiler dialects configured for each version. The build also integrates the \ifdef\ plugin to manage Spark version-specific code paths and updates the Java target version to 17 for Spark 4.1.3+ and 11 for earlier versions.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 56.

Lenses

  • Code Health 94
  • Architecture 100
  • Maturity 54
  • Readiness 65
  • Security 45

Changes since last survey

  • 300 commits — 289 feature/other, 11 fixes

By area

  • (root) — 234 commits
  • .github/workflows — 30 commits
  • src/main — 18 commits
  • src/test — 10 commits
  • project/plugins.sbt — 5 commits
  • .github/ISSUE_TEMPLATE — 1 commit
  • .github/dependabot.yml — 1 commit
  • project/build.properties — 1 commit

Notable commits

  • fix: Fix #747: Remove decimal points from integer values when usePlainNumberFormat=true (#948)
  • fix: Read mulitple worksheets, bug fixes for faulty dimension tag, sheet names consisting of digits (#949)
  • fix: Revert commit(s) 1e19a1d
  • fix: Revert commit(s) 4800a16
  • fix: build: Fix Sonatype Central publishing
  • fix: build: Fix deprecation regarding SbtModuleTests
  • fix: build: Fix source paths messed up after Mill upgrade
  • fix: fix exception when reading xlsx file by sheet index (#994)
  • fix: fix(ci): use Java 17 in publish job for Spark 4.x compatibility (#1029)
  • fix: fix: Fix 3.5.0 compile issues
  • fix: fixed handling of keepUndefinedRows if excel contains empty rows (issue #965) (#966)
  • change: Add 'Reformat with scalafmt 3.7.11' to .git-blame-ignore-revs
  • change: Add 'Reformat with scalafmt 3.7.15' to .git-blame-ignore-revs
  • change: Add 'Reformat with scalafmt 3.7.17' to .git-blame-ignore-revs
  • change: Add 'Reformat with scalafmt 3.8.5' to .git-blame-ignore-revs
  • change: Add 'Reformat with scalafmt 3.9.10' to .git-blame-ignore-revs
  • change: Add 'Reformat with scalafmt 3.9.5' to .git-blame-ignore-revs
  • change: Add Codium PR Agent
  • change: Added Spark 4 Support (#976)
  • change: Apply Spark Session Properties for Writing Files (#728)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

nightscape/spark-excel was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 5a6d63541d619fc7405b8b38b68d7c28513a599f — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.