Skip to content
CAI
Software that uses CAICheck a score

derrickburns/generalized-kmeans-clustering

72.1

Strong · 20 September 2026

25.2k

lines of production code

Scala

with Python

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Scala-based library for distributed machine learning clustering, built on Apache Spark. It provides a comprehensive suite of clustering algorithms—including various K-Means variants, spectral clustering, and co-clustering—accessible via a modern DataFrame API with support for multiple Bregman divergences. The library also includes tools for model evaluation, persistence, and PySpark integration, having recently migrated from a legacy RDD interface to a more modular and optimized architecture.

How it got here

2014 — Project modernization and API overhaul

5 changes.

The project underwent a comprehensive modernization, migrating the build system from Maven to sbt and upgrading to Scala 2.13 and Spark 3.5.1. This period involved removing legacy RDD-based components and custom K-Means implementations in favor of a new DataFrame-only API, while introducing several new clustering estimators and extensive test coverage.

2015 — Spark ML clustering feature expansion

5 changes.

This period focused on expanding the Spark ML library with a comprehensive suite of new clustering algorithms, including hierarchical, multi-view, and constrained K-Means variants, alongside distributed evaluation metrics. The work was supported by the introduction of low-level linear algebra utilities and RDD helper tools to efficiently manage data processing and centroid tracking. Concurrently, the build infrastructure was modernized with updated SBT plugins for code quality, coverage, and publishing, while test logging was refined to reduce noise.

2024–2025 — DataFrame API and Bregman divergence support

8 changes.

The project introduced a new DataFrame-based clustering API that replaces the previous RDD implementation, enabling support for a wide range of Bregman divergences and modular kernel strategies. This period saw the addition of numerous new clustering algorithms, such as Balanced K-Means and CLARA, alongside PySpark bindings and comprehensive example suites to demonstrate the new capabilities.

Features

Added Bloop configuration for the sbt build project

The project now includes Bloop configuration files in \project/.bloop\ to enable IDE integration and faster compilation for the sbt build itself. This adds \bloop.settings.json\ defining supported Scala versions (2.11.12, 2.12.17–2.12.20, 2.13.14–2.13.17) and SemanticDB versions, alongside \generalized-kmeans-clustering-build.json\ which configures the build project's classpath, source directories, and dependencies (including sbt 1.9.8 and various Scala libraries).

project/.bloop · high confidence

Introduce PySpark wrapper for Generalized K-Means clustering

This release adds a new Python package (\massivedatascience-clusterer\) that provides PySpark bindings for the generalized K-Means clustering algorithm. The wrapper exposes \GeneralizedKMeans\ as a native Spark ML Estimator, allowing users to perform clustering using multiple Bregman divergences (Squared Euclidean, KL Divergence, Itakura-Saito, Generalized I-divergence, and Logistic Loss). It supports standard Spark ML features such as Pipeline integration, model persistence (save/load), weighted clustering, and detailed quality metrics (WCSS, Calinski-Harabasz, Davies-Bouldin, etc.) via the model summary. The package is versioned at 0.7.0 and includes example scripts and a smoke test to verify functionality.

python · high confidence

Modular clustering strategy architecture with optimized assignment paths

The clustering engine in the strategies package has been restructured into a modular strategy pattern, introducing traits for assignment, center updates, convergence checking, empty cluster handling, and input validation. For point assignment, the system now supports multiple optimized paths: a fast Squared Euclidean cross-join strategy, a memory-efficient chunked broadcast approach, and an adaptive strategy that automatically selects the best method based on executor memory configuration. Additionally, an accelerated assignment strategy using triangle inequality pruning is available for Squared Euclidean kernels to reduce distance computations. Center updates now support both gradient-based means for Bregman kernels and component-wise weighted medians for K-Medians clustering.

src/main/scala/com/massivedatascience/clusterer/ml/df/strategies · high confidence

New DataFrame-based clustering API with Bregman divergence support

The library introduces a new DataFrame API for clustering algorithms, replacing the previous RDD-based implementation. This new API supports a variety of Bregman divergences (such as KL, Itakura-Saito, and L1) through a modular kernel system, enabling more flexible distance metrics for clustering. It includes new algorithms like Spectral Clustering and Information Bottleneck, along with performance optimizations like Elkan-accelerated Lloyd's iterations for squared Euclidean distance. The API also provides robust numeric guards to detect and prevent common issues like NaN/Inf propagation and domain violations, ensuring more stable and reliable clustering results.

src/main/scala/com/massivedatascience/clusterer/ml/df · high confidence

New Spark ML clustering algorithms and evaluation metrics

This release introduces a suite of new clustering estimators and diagnostic tools integrated with the Spark ML pipeline. Users can now perform hierarchical clustering with Bregman divergences via AgglomerativeBregman, handle multi-modal data with MultiViewKMeans, and leverage specialized K-Means variants including SparseKMeans for high-dimensional sparse data, RobustKMeans for outlier handling, ConstrainedKMeans for must-link/cannot-link constraints, and CoClustering for simultaneous row/column matrix clustering. Additionally, the new ClusteringMetrics object provides distributed evaluation capabilities, computing silhouette scores, inertia, and cluster balance directly from prediction DataFrames, while TrainingSummary offers detailed convergence diagnostics for all trained models.

repository · high confidence

New clustering algorithms and model infrastructure added to the DataFrame API

This release introduces several new clustering algorithms to the Spark ML DataFrame API, including Balanced K-Means (with soft/hard size constraints), Bisecting K-Means (hierarchical divisive clustering), CLARA (sampling-based K-Medoids for large datasets), Coreset K-Means (approximation for large-scale speedup), DP-Means (automatic cluster count determination), and Bregman Mixture Models (probabilistic EM-based clustering). These new estimators share a common infrastructure, including a unified \ClusteringKernel\ for pluggable Bregman divergences (e.g., KL, Itakura-Saito, Cosine) and a \HasTrainingSummary\ trait for consistent model metadata handling.

src/main/scala/com/massivedatascience/clusterer/ml · high confidence

New executable examples for clustering algorithms and persistence

Added a suite of executable Scala examples in the \examples\ package that demonstrate the usage and persistence round-trips for various clustering algorithms, including GeneralizedKMeans, CoresetKMeans, KMedoids, SoftKMeans, StreamingKMeans, XMeans, and Spherical K-Means. These examples show how to train models, inspect training summaries, and save/load models to disk to verify that parameters and predictions are preserved correctly across sessions.

src/main/scala/examples · high confidence

New low-level linear algebra and vector iteration utilities

The library introduces a new set of internal components in the \com.massivedatascience.linalg\ package to support clustering algorithms. This includes a \BLAS\ object providing vector operations like \axpy\ and \merge\ (leveraging Netlib), a \BregmanFunction\ trait with a \SquaredEuclideanFunction\ implementation for divergence calculations, and mutable/immutable vector types (\EagerCentroid\, \MutableWeightedVector\, \WeightedVector\) to efficiently track cluster centroids. Additionally, \VectorIterator\ classes and implicit extensions are added to allow efficient, inlined iteration over dense and sparse vectors.

src/main/scala/com/massivedatascience/linalg · high confidence

New modular kernel hierarchy and factory for DataFrame-based clustering

The clustering module now uses a structured kernel hierarchy with a base \ClusteringKernel\ trait and a \BregmanKernel\ extension for gradient-based center updates, alongside a \KernelFactory\ for unified creation of dense and sparse-optimized kernels. This change introduces support for multiple divergence types including Squared Euclidean, KL, Itakura-Saito, Generalized I, Logistic, L1, and Spherical (cosine), with specific implementations for sparse data (e.g., \SparseKLKernel\) and a factory that can auto-select sparse optimizations based on data sparsity ratios.

src/main/scala/com/massivedatascience/clusterer/ml/df/kernels · high confidence

New utility helpers and validation logic for Spark RDDs

The \com.massivedatascience.util\ package now includes \SparkHelper\, a trait providing methods to manage RDD lifecycle (caching, unpersisting, and synchronous computation) and broadcast variables, alongside \ValidationUtils\, which adds strict input validation for weights, probabilities, vectors, and indices. Additionally, \XORShiftRandom\ has been refactored and moved to this package, improving its seed hashing and exposing a global random instance for default seeding in clustering methods.

src/main/scala/com/massivedatascience/util · high confidence

Project modernization and structural overhaul

This release introduces a comprehensive modernization of the project, including a major architectural refactor that removes the legacy RDD API in favor of a DataFrame-only interface, and migrates the default Scala version to 2.13.14. It adds several new clustering estimators: RobustKMeans for outlier-resistant clustering, SparseKMeans for high-dimensional sparse data, MultiViewKMeans for multi-feature clustering, TimeSeriesKMeans with DTW-based kernels, Information Bottleneck clustering, and Spectral Clustering. The release also includes a new ClusteringMetrics component for model selection and diagnostics, and restructures documentation using the Diátaxis framework.

(repo-wide) · high confidence

Removals

Removal of Euclidean distance metric implementations

The Euclidean distance metric implementations, specifically the standard \EuOps\ and the optimized \FastEuOps\ classes, have been removed from the clusterer metrics module. This deletion eliminates the ability to perform clustering using Euclidean distance calculations within this component.

src/main/scala/com/rincaro/clusterer/metrics · high confidence

Removal of custom K-Means clustering implementation

The custom K-Means clustering components located in the \com.rincaro.clusterer.base\ package have been removed. This includes the core \GeneralizedKMeans\ algorithm, the \GeneralizedKMeansModel\, the \KMeans\ entry point, and all initialization strategies (\KMeansInitializer\, \KMeansParallel\, \KMeansPlusPlus\, \KMeansRandom\). Additionally, the supporting type definitions and operations (\FP\, \FPoint\, \Centroid\, \PointOps\) defined in the \package.scala\ file are no longer present, eliminating this local clustering capability from the application.

src/main/scala/com/rincaro/clusterer/base · high confidence

Fixes

Add Scala 2.13 parallel collections compatibility

Users can now use parallel collections in Scala 2.13 through the new \compat\ package, which provides a \.par\ extension method that delegates to \CollectionConverters\. This change ensures cross-version compatibility by maintaining an empty compat package for Scala 2.12 (where parallel collections are built-in) while implementing the necessary adapter for Scala 2.13.

src/main/scala-2.13 · high confidence

Test coverage

Added comprehensive test suites for clustering algorithms and metrics; Added test logging configuration files.

Dependencies

Build system modernization and dependency updates

The project has migrated its build infrastructure from Maven to sbt, replacing the legacy pom.xml with a new build.sbt that targets Scala 2.13.14 (with 2.12.18 support) and enforces Java 17 compatibility. Core dependencies have been updated to Apache Spark 3.5.1 and Scala 2.13, including the addition of the scala-parallel-collections module required for Scala 2.13. The Python package (pyproject.toml) now specifies PySpark \>=3.4.0 and NumPy \>=1.20.0, while publishing is configured for Maven Central via sbt-ci-release.

(dependencies) · high confidence

Modernized build configuration with SBT 1.11 and updated plugins

The project build has been updated to use SBT version 1.11.0 and replaced legacy plugin definitions with modern equivalents for code quality, coverage, and publishing. Key additions include sbt-scoverage 2.2.2 for code coverage, sbt-ci-release 1.11.2 for Maven Central publishing via the Central Portal, sbt-scalafmt 2.5.2 for code formatting, and sbt-dependency-check 5.1.0 for vulnerability scanning and SBOM generation. The configuration also resolves a scala-xml version conflict and includes sbt-unidoc for unified documentation.

project · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 72.

Lenses

  • Code Health 93
  • Architecture 100
  • Maturity 61
  • Readiness 78
  • Security 80

Changes since last survey

  • 300 commits — 221 feature/other, 79 fixes

By area

  • (root) — 111 commits
  • src/main — 109 commits
  • .github/workflows — 30 commits
  • src/test — 18 commits
  • project/plugins.sbt — 4 commits
  • docs/CNAME — 3 commits
  • release-notes/releases.md — 3 commits
  • (repo) — 2 commits
  • .bloop/generalized-kmeans-clustering — 2 commits
  • docs/_config.yml — 2 commits
  • docs/_howto — 2 commits
  • docs/api — 2 commits
  • project/project — 2 commits
  • .metals/.reports — 1 commit
  • docs/_explanation — 1 commit
  • docs/guides — 1 commit
  • docs/howto — 1 commit
  • docs/index.md — 1 commit
  • project/target — 1 commit
  • python/examples — 1 commit

Notable commits

  • fix: Avoid compiler bug?
  • fix: Fix .scalafmt
  • fix: Fix .scalafmt.conf
  • fix: Fix Bregman divergence negative distance bug
  • fix: Fix CI
  • fix: Fix CI
  • fix: Fix CI
  • fix: Fix CI: format code and update CodeQL workflow
  • fix: Fix CodeQL: use build-mode manual for Scala
  • fix: Fix PySpark smoke test JAR discovery
  • fix: Fix PySpark wrapper: add missing setter methods
  • fix: Fix RobustKMeans divergence case mismatch and SparseKMeans L1 ClassCastException
  • fix: Fix Scala 2.13 compilation errors
  • fix: Fix broken test.
  • fix: Fix bugs in K-means implementation: 1) Fix potential empty sequence return in KMeansPlusPlus.pickWeighted, 2) Improve handling of negative distances in BregmanPointOps, 3) Enhance robustness of empty cluster handling in ColumnTrackingKMeans
  • fix: Fix codeql ci
  • fix: Fix doc gen
  • fix: Fix permissions on docs.yml
  • fix: Fix publish workflow: use PGP_SECRET/PGP_PASSPHRASE secret names for sbt-ci-release
  • fix: Fix publish: use +ci-release to cross-publish both Scala versions in single deployment
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

derrickburns/generalized-kmeans-clustering was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit f0668a833099db2be002abaa3ed972eb8e3318f3 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.