swoop-inc/spark-alchemy
46.5
Weak · 20 September 2026
1.2k
lines of production code
Scala
with JavaScript
1
measurement over time
What this system is
Spark-Alchemy is a Scala library for Apache Spark that provides HyperLogLog++ functions for approximate distinct counting and cardinality estimation. It enables users to create, merge, and intersect composable sketches with configurable error rates, supporting pluggable backends like StreamLib and Aggregate Knowledge. The system includes comprehensive testing infrastructure to validate accuracy and interoperability with PostgreSQL.
Features
Initial project scaffolding and documentation
This change introduces the foundational structure for the spark-alchemy project, including the README, development guidelines, and versioning strategy. It establishes the supported SDK versions via .sdkmanrc (Java 8, SBT 1.6.2, Scala 2.12.17) and sets up the local development environment with a docker-compose configuration for Postgres interoperability testing. The entry also includes standard project files like .gitignore, NOTICE, and code style settings, marking the initial public release of the library's core documentation and build infrastructure.
(repo-wide) · high confidence
New HyperLogLog++ functions for cardinality estimation
Alchemy now provides a suite of HyperLogLog++ functions for approximate distinct counting, including \hll\_init\, \hll\_init\_collection\, \hll\_init\_agg\, \hll\_init\_collection\_agg\, \hll\_merge\, \hll\_row\_merge\, \hll\_cardinality\, \hll\_intersect\_cardinality\, and \hll\_convert\. These functions allow users to create composable sketches, merge them across partitions, and estimate cardinalities with configurable error rates. The implementation supports pluggable backends (StreamLib and Aggregate Knowledge) and uses a specialized hash function to better differentiate data types and handle nulls compared to Spark's built-in approximations.
alchemy/src/main · high confidence
Test coverage
Added test infrastructure and unit tests for HyperLogLog functions
Added comprehensive test coverage for the new HyperLogLog (HLL) functions, including \hll\_init\, \hll\_merge\, \hll\_cardinality\, \hll\_row\_merge\, and \hll\_intersect\_cardinality\. The changes introduce \CardinalityHashFunctionTest\ and \HLLFunctionsTest\ to verify cardinality estimation accuracy for simple types and collections, validate null handling, and ensure correct behavior for both the Aggregate Knowledge (AGKN) and StreamLib (STRM) implementations. A new \PostgresInteropTest\ validates that HLL aggregates produced in Spark match results from PostgreSQL when using the AGKN implementation. Additionally, the test suite includes new test utilities (\SparkSessionSpec\, \DebugFilesystem\, \SparkFunSuite\, \PlanTest\, \SQLHelper\, \SQLTestData\, \SQLTestUtils\, \SharedSparkSessionBase\, \TestSparkSession\) to support isolated Spark session management, resource leak detection, and SQL plan comparison.
alchemy/src/test · high confidence
Dependencies
Initial project setup with Spark 3.5.2 and Scala 2.12.15
The build configuration establishes the project's core dependencies, targeting Apache Spark 3.5.2 and Scala 2.12.15. It includes the HyperLogLog++ library (net.agkn:hll:1.6.0) for cardinality estimation, PostgreSQL and ScalaTest for testing, and configures the microsite documentation generation using GitHub4s.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 47.
Lenses
- Code Health 97
- Architecture 69
- Maturity 44
- Readiness 32
- Security 75
Changes since last survey
- 49 commits — 48 feature/other, 1 fixes
By area
- (root) — 29 commits
- alchemy/src — 8 commits
- .github/workflows — 3 commits
- (repo) — 2 commits
- .circleci/Dockerfile — 2 commits
- .circleci/config.yml — 2 commits
- docs/main — 1 commit
- docs/src — 1 commit
- project/plugins.sbt — 1 commit
Notable commits
- fix: GitHub Actions config fix for bintray credential handling
- change: Add Aggregate Knowledge HLL implementation
- change: Add hll_row_merge and hll_intersect_cardinality
- change: Add an SDKMAN file while we're at it
- change: Added release create steps to CircleCI config
- change: Addressing comments
- change: Adds HyperLogLog release information
- change: Adds HyperLogLog++ functions for Spark
- change: Adds doc describing the use of HLL functions through PySpark
- change: Adds sbt plugin to enable the easy assembly of a fat JAR of spark-alchemy
- change: Bump to Scalatest 3.2.2
- change: Bump to Spark 3.0.1
- change: Bump to v1.0.1
- change: Bump version to 0.3.0 for release
- change: Bumped VERSION to 0.4.0-SNAPSHOT
- change: Bumped base version in VERSION file
- change: Bumps version because of changed dependencies
- change: Bumps version for 1.0.0 release
- change: CircleCI custom image update, microsite reconfigured with github4s
- change: CircleCI custom image updated with grep
- …and 29 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
swoop-inc/spark-alchemy was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 58293e4cceabdd6a1c963733eddf46100d80b6a9 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.