derrickburns/generalized-kmeans-clustering
72.1
Strong · 20 September 2026
25.2k
lines of production code
Scala
with Python
1
measurement over time
What this system is
This system is a Scala-based library for distributed machine learning clustering, built on Apache Spark. It provides a comprehensive suite of clustering algorithms—including various K-Means variants, spectral clustering, and co-clustering—accessible via a modern DataFrame API with support for multiple Bregman divergences. The library also includes tools for model evaluation, persistence, and PySpark integration, having recently migrated from a legacy RDD interface to a more modular and optimized architecture.
How it got here
2014 — Project modernization and API overhaul
5 changes.
The project underwent a comprehensive modernization, migrating the build system from Maven to sbt and upgrading to Scala 2.13 and Spark 3.5.1. This period involved removing legacy RDD-based components and custom K-Means implementations in favor of a new DataFrame-only API, while introducing several new clustering estimators and extensive test coverage.
2015 — Spark ML clustering feature expansion
5 changes.
This period focused on expanding the Spark ML library with a comprehensive suite of new clustering algorithms, including hierarchical, multi-view, and constrained K-Means variants, alongside distributed evaluation metrics. The work was supported by the introduction of low-level linear algebra utilities and RDD helper tools to efficiently manage data processing and centroid tracking. Concurrently, the build infrastructure was modernized with updated SBT plugins for code quality, coverage, and publishing, while test logging was refined to reduce noise.
2024–2025 — DataFrame API and Bregman divergence support
8 changes.
The project introduced a new DataFrame-based clustering API that replaces the previous RDD implementation, enabling support for a wide range of Bregman divergences and modular kernel strategies. This period saw the addition of numerous new clustering algorithms, such as Balanced K-Means and CLARA, alongside PySpark bindings and comprehensive example suites to demonstrate the new capabilities.
Features
Added Bloop configuration for the sbt build project
The project now includes Bloop configuration files in \project/.bloop\ to enable IDE integration and faster compilation for the sbt build itself. This adds \bloop.settings.json\ defining supported Scala versions (2.11.12, 2.12.17–2.12.20, 2.13.14–2.13.17) and SemanticDB versions, alongside \generalized-kmeans-clustering-build.json\ which configures the build project's classpath, source directories, and dependencies (including sbt 1.9.8 and various Scala libraries).
project/.bloop · high confidence
Introduce PySpark wrapper for Generalized K-Means clustering
This release adds a new Python package (\massivedatascience-clusterer\) that provides PySpark bindings for the generalized K-Means clustering algorithm. The wrapper exposes \GeneralizedKMeans\ as a native Spark ML Estimator, allowing users to perform clustering using multiple Bregman divergences (Squared Euclidean, KL Divergence, Itakura-Saito, Generalized I-divergence, and Logistic Loss). It supports standard Spark ML features such as Pipeline integration, model persistence (save/load), weighted clustering, and detailed quality metrics (WCSS, Calinski-Harabasz, Davies-Bouldin, etc.) via the model summary. The package is versioned at 0.7.0 and includes example scripts and a smoke test to verify functionality.
python · high confidence
Modular clustering strategy architecture with optimized assignment paths
The clustering engine in the strategies package has been restructured into a modular strategy pattern, introducing traits for assignment, center updates, convergence checking, empty cluster handling, and input validation. For point assignment, the system now supports multiple optimized paths: a fast Squared Euclidean cross-join strategy, a memory-efficient chunked broadcast approach, and an adaptive strategy that automatically selects the best method based on executor memory configuration. Additionally, an accelerated assignment strategy using triangle inequality pruning is available for Squared Euclidean kernels to reduce distance computations. Center updates now support both gradient-based means for Bregman kernels and component-wise weighted medians for K-Medians clustering.
src/main/scala/com/massivedatascience/clusterer/ml/df/strategies · high confidence
New DataFrame-based clustering API with Bregman divergence support
The library introduces a new DataFrame API for clustering algorithms, replacing the previous RDD-based implementation. This new API supports a variety of Bregman divergences (such as KL, Itakura-Saito, and L1) through a modular kernel system, enabling more flexible distance metrics for clustering. It includes new algorithms like Spectral Clustering and Information Bottleneck, along with performance optimizations like Elkan-accelerated Lloyd's iterations for squared Euclidean distance. The API also provides robust numeric guards to detect and prevent common issues like NaN/Inf propagation and domain violations, ensuring more stable and reliable clustering results.
src/main/scala/com/massivedatascience/clusterer/ml/df · high confidence
New Spark ML clustering algorithms and evaluation metrics
This release introduces a suite of new clustering estimators and diagnostic tools integrated with the Spark ML pipeline. Users can now perform hierarchical clustering with Bregman divergences via AgglomerativeBregman, handle multi-modal data with MultiViewKMeans, and leverage specialized K-Means variants including SparseKMeans for high-dimensional sparse data, RobustKMeans for outlier handling, ConstrainedKMeans for must-link/cannot-link constraints, and CoClustering for simultaneous row/column matrix clustering. Additionally, the new ClusteringMetrics object provides distributed evaluation capabilities, computing silhouette scores, inertia, and cluster balance directly from prediction DataFrames, while TrainingSummary offers detailed convergence diagnostics for all trained models.
repository · high confidence
New clustering algorithms and model infrastructure added to the DataFrame API
This release introduces several new clustering algorithms to the Spark ML DataFrame API, including Balanced K-Means (with soft/hard size constraints), Bisecting K-Means (hierarchical divisive clustering), CLARA (sampling-based K-Medoids for large datasets), Coreset K-Means (approximation for large-scale speedup), DP-Means (automatic cluster count determination), and Bregman Mixture Models (probabilistic EM-based clustering). These new estimators share a common infrastructure, including a unified \ClusteringKernel\ for pluggable Bregman divergences (e.g., KL, Itakura-Saito, Cosine) and a \HasTrainingSummary\ trait for consistent model metadata handling.
src/main/scala/com/massivedatascience/clusterer/ml · high confidence
New executable examples for clustering algorithms and persistence
Added a suite of executable Scala examples in the \examples\ package that demonstrate the usage and persistence round-trips for various clustering algorithms, including GeneralizedKMeans, CoresetKMeans, KMedoids, SoftKMeans, StreamingKMeans, XMeans, and Spherical K-Means. These examples show how to train models, inspect training summaries, and save/load models to disk to verify that parameters and predictions are preserved correctly across sessions.
src/main/scala/examples · high confidence
New low-level linear algebra and vector iteration utilities
The library introduces a new set of internal components in the \com.massivedatascience.linalg\ package to support clustering algorithms. This includes a \BLAS\ object providing vector operations like \axpy\ and \merge\ (leveraging Netlib), a \BregmanFunction\ trait with a \SquaredEuclideanFunction\ implementation for divergence calculations, and mutable/immutable vector types (\EagerCentroid\, \MutableWeightedVector\, \WeightedVector\) to efficiently track cluster centroids. Additionally, \VectorIterator\ classes and implicit extensions are added to allow efficient, inlined iteration over dense and sparse vectors.
src/main/scala/com/massivedatascience/linalg · high confidence
New modular kernel hierarchy and factory for DataFrame-based clustering
The clustering module now uses a structured kernel hierarchy with a base \ClusteringKernel\ trait and a \BregmanKernel\ extension for gradient-based center updates, alongside a \KernelFactory\ for unified creation of dense and sparse-optimized kernels. This change introduces support for multiple divergence types including Squared Euclidean, KL, Itakura-Saito, Generalized I, Logistic, L1, and Spherical (cosine), with specific implementations for sparse data (e.g., \SparseKLKernel\) and a factory that can auto-select sparse optimizations based on data sparsity ratios.
src/main/scala/com/massivedatascience/clusterer/ml/df/kernels · high confidence
New utility helpers and validation logic for Spark RDDs
The \com.massivedatascience.util\ package now includes \SparkHelper\, a trait providing methods to manage RDD lifecycle (caching, unpersisting, and synchronous computation) and broadcast variables, alongside \ValidationUtils\, which adds strict input validation for weights, probabilities, vectors, and indices. Additionally, \XORShiftRandom\ has been refactored and moved to this package, improving its seed hashing and exposing a global random instance for default seeding in clustering methods.
src/main/scala/com/massivedatascience/util · high confidence
Project modernization and structural overhaul
This release introduces a comprehensive modernization of the project, including a major architectural refactor that removes the legacy RDD API in favor of a DataFrame-only interface, and migrates the default Scala version to 2.13.14. It adds several new clustering estimators: RobustKMeans for outlier-resistant clustering, SparseKMeans for high-dimensional sparse data, MultiViewKMeans for multi-feature clustering, TimeSeriesKMeans with DTW-based kernels, Information Bottleneck clustering, and Spectral Clustering. The release also includes a new ClusteringMetrics component for model selection and diagnostics, and restructures documentation using the Diátaxis framework.
(repo-wide) · high confidence
Removals
Removal of Euclidean distance metric implementations
The Euclidean distance metric implementations, specifically the standard \EuOps\ and the optimized \FastEuOps\ classes, have been removed from the clusterer metrics module. This deletion eliminates the ability to perform clustering using Euclidean distance calculations within this component.
src/main/scala/com/rincaro/clusterer/metrics · high confidence
Removal of custom K-Means clustering implementation
The custom K-Means clustering components located in the \com.rincaro.clusterer.base\ package have been removed. This includes the core \GeneralizedKMeans\ algorithm, the \GeneralizedKMeansModel\, the \KMeans\ entry point, and all initialization strategies (\KMeansInitializer\, \KMeansParallel\, \KMeansPlusPlus\, \KMeansRandom\). Additionally, the supporting type definitions and operations (\FP\, \FPoint\, \Centroid\, \PointOps\) defined in the \package.scala\ file are no longer present, eliminating this local clustering capability from the application.
src/main/scala/com/rincaro/clusterer/base · high confidence
Fixes
Add Scala 2.13 parallel collections compatibility
Users can now use parallel collections in Scala 2.13 through the new \compat\ package, which provides a \.par\ extension method that delegates to \CollectionConverters\. This change ensures cross-version compatibility by maintaining an empty compat package for Scala 2.12 (where parallel collections are built-in) while implementing the necessary adapter for Scala 2.13.
src/main/scala-2.13 · high confidence
Test coverage
Added comprehensive test suites for clustering algorithms and metrics; Added test logging configuration files.
Dependencies
Build system modernization and dependency updates
The project has migrated its build infrastructure from Maven to sbt, replacing the legacy pom.xml with a new build.sbt that targets Scala 2.13.14 (with 2.12.18 support) and enforces Java 17 compatibility. Core dependencies have been updated to Apache Spark 3.5.1 and Scala 2.13, including the addition of the scala-parallel-collections module required for Scala 2.13. The Python package (pyproject.toml) now specifies PySpark \>=3.4.0 and NumPy \>=1.20.0, while publishing is configured for Maven Central via sbt-ci-release.
(dependencies) · high confidence
Modernized build configuration with SBT 1.11 and updated plugins
The project build has been updated to use SBT version 1.11.0 and replaced legacy plugin definitions with modern equivalents for code quality, coverage, and publishing. Key additions include sbt-scoverage 2.2.2 for code coverage, sbt-ci-release 1.11.2 for Maven Central publishing via the Central Portal, sbt-scalafmt 2.5.2 for code formatting, and sbt-dependency-check 5.1.0 for vulnerability scanning and SBOM generation. The configuration also resolves a scala-xml version conflict and includes sbt-unidoc for unified documentation.
project · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 72.
Lenses
- Code Health 93
- Architecture 100
- Maturity 61
- Readiness 78
- Security 80
Changes since last survey
- 300 commits — 221 feature/other, 79 fixes
By area
- (root) — 111 commits
- src/main — 109 commits
- .github/workflows — 30 commits
- src/test — 18 commits
- project/plugins.sbt — 4 commits
- docs/CNAME — 3 commits
- release-notes/releases.md — 3 commits
- (repo) — 2 commits
- .bloop/generalized-kmeans-clustering — 2 commits
- docs/_config.yml — 2 commits
- docs/_howto — 2 commits
- docs/api — 2 commits
- project/project — 2 commits
- .metals/.reports — 1 commit
- docs/_explanation — 1 commit
- docs/guides — 1 commit
- docs/howto — 1 commit
- docs/index.md — 1 commit
- project/target — 1 commit
- python/examples — 1 commit
Notable commits
- fix: Avoid compiler bug?
- fix: Fix .scalafmt
- fix: Fix .scalafmt.conf
- fix: Fix Bregman divergence negative distance bug
- fix: Fix CI
- fix: Fix CI
- fix: Fix CI
- fix: Fix CI: format code and update CodeQL workflow
- fix: Fix CodeQL: use build-mode manual for Scala
- fix: Fix PySpark smoke test JAR discovery
- fix: Fix PySpark wrapper: add missing setter methods
- fix: Fix RobustKMeans divergence case mismatch and SparseKMeans L1 ClassCastException
- fix: Fix Scala 2.13 compilation errors
- fix: Fix broken test.
- fix: Fix bugs in K-means implementation: 1) Fix potential empty sequence return in KMeansPlusPlus.pickWeighted, 2) Improve handling of negative distances in BregmanPointOps, 3) Enhance robustness of empty cluster handling in ColumnTrackingKMeans
- fix: Fix codeql ci
- fix: Fix doc gen
- fix: Fix permissions on docs.yml
- fix: Fix publish workflow: use PGP_SECRET/PGP_PASSPHRASE secret names for sbt-ci-release
- fix: Fix publish: use +ci-release to cross-publish both Scala versions in single deployment
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
derrickburns/generalized-kmeans-clustering was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f0668a833099db2be002abaa3ed972eb8e3318f3 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.