Skip to content
CAI
Software that uses CAICheck a score

h2oai/sparkling-water

61.6

Adequate · 28 September 2026

35.7k

lines of production code

Scala

with Python

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Sparkling Water is a bridge that integrates the H2O machine learning platform with Apache Spark, enabling users to run H2O algorithms and models within Spark workflows. It provides multi-language APIs for Scala, Python, and R, allowing seamless conversion between Spark DataFrames and H2O Frames while supporting standard Spark ML pipelines. The system also includes infrastructure for deploying on AWS EMR and Kubernetes, along with automated tools for benchmarking, documentation, and continuous integration.

How it got here

2014–2019 — Project scaffolding and multi-language API expansion

49 changes.

This period established the foundational structure of the Sparkling Water project, introducing multi-version support for Spark and a modular Gradle build system. It focused on expanding the project's reach by releasing initial Python (PySparkling) and R (RSparkling) APIs, while simultaneously enhancing the ML infrastructure with new algorithm wrappers, MOJO scoring capabilities, and comprehensive test coverage.

2020 — API generation and CI migration

48 changes.

This period focused on automating the generation of Scala, Python, and R APIs through a new code-generation infrastructure, while simultaneously migrating the continuous integration pipeline to Jenkins on AWS. Significant efforts were also directed toward modernizing the RSparkling client interface, expanding test coverage across multiple backends, and adding comprehensive documentation generation tools.

2021–2023 — Python scoring package and Spark compatibility

15 changes.

This period focused on releasing the standalone Python scoring package for H2O MOJO models, enabling seamless integration into Python pipelines with features like SHAP contributions and prediction intervals. Concurrently, the codebase expanded its compatibility layer to support Spark versions 3.1 through 3.5 by implementing version-specific shims for SQL encoders and user-defined functions. Extensive test coverage was added to validate the new Python package, serialization across Scala versions, and the correctness of model metrics and SHAP calculations.

Features

Add AWS EMR Terraform configuration for benchmarks

New Terraform files (main.tf, variables.tf, outputs.tf) have been added to the benchmarks infrastructure to provision Amazon EMR clusters for running benchmarks. This configuration allows users to deploy benchmark workloads on AWS using EMR, with customizable parameters for instance types, cluster size, H2O version, and memory settings.

benchmarks/src/main/terraform · high confidence

Add AWS EMR bootstrap script for Sparkling Water installation

A new bootstrap script (install\_sparkling\_water.sh) has been added to the AWS templates to automate the installation of Sparkling Water on AWS EMR clusters. This script configures the Python environment by upgrading pip, installing specific dependencies (requests, tabulate, six, scikit-learn), and downloading the Sparkling Water archive, enabling PySparkling functionality on worker nodes.

templates/src/aws · high confidence

Add Terraform module for deploying Sparkling Water on AWS EMR

A new Terraform module has been added to provision a Sparkling Water cluster on AWS EMR. This module manages the underlying infrastructure, including an S3 bucket for deployment artifacts, an EMR cluster configured with Spark, Hadoop, and JupyterHub, and the necessary IAM roles and security groups. It includes bootstrap scripts to install Sparkling Water dependencies and configure Jupyter access, and exposes outputs for the Jupyter notebook URL, master DNS, and S3 bucket name.

_templates/src/terraform/aws/modules/emr\deployment · high confidence

Add conda recipe for version-specific PySparkling package

A new conda build recipe has been added for the PySparkling package, allowing it to be built and installed via conda. The recipe defines build and runtime dependencies including Python, pip, setuptools, tabulate, requests, and future, and uses a placeholder for the Spark major version to support multiple Spark versions.

py/conda · high confidence

Added Chicago weather datasets for external testing

Two new CSV files, \Chicago\_Ohare\_International\_Airport.csv\ and \chicagoAllWeather.csv\, have been added to the \examples/smalldata/chicago\ directory. These files provide historical weather data (including temperature, precipitation, and snow metrics) for Chicago, making them available for use in external tests and examples.

examples/smalldata/chicago · high confidence

Added ECG anomaly detection datasets

The examples/smalldata/anomaly directory now includes \ecg\_discord\_train.csv\ and \ecg\_discord\_test.csv\, providing new training and test data for ECG-based anomaly detection examples.

examples/smalldata/anomaly · high confidence

Added NYC City Bike and weather sample datasets

The examples/smalldata/citybike-nyc directory now includes two new CSV files: citybike\_2013.csv, containing 2013 NYC Citi Bike trip data (durations, stations, user types), and New\_York\_City\_Hourly\_Weather\_2013.csv, providing hourly weather observations for New York City in 2013. These files serve as sample data for the citybike-nyc example.

examples/smalldata/citybike-nyc · high confidence

Added enriched airline dataset with graph features

The \examples/smalldata/airlines\ directory now includes a new CSV file, \airlines\_big\_data\_100.csv\, which provides an expanded version of the airline dataset. This file contains the original flight data augmented with numerous graph-theory features, including various centrality measures (Degree, Closeness, Harmonic, SubGraph, Load, Pagerank, EigenVector) and node embeddings (node2vec) for both origin and destination airports, enabling more complex network-based analysis.

examples/smalldata/airlines · high confidence

Added heart transplant dataset for CoxPH testing

The \examples/smalldata/coxph\_test\ directory now includes \heart.csv\ and \heart\_test.csv\, providing training and test data for the Cox Proportional Hazards model. These files contain columns for start/stop times, event status, age, year, surgery, and transplant status, enabling users to run and validate CoxPH algorithm examples.

_examples/smalldata/coxph\test · high confidence

Added prostate cancer dataset and anomaly validation files to examples

The \examples/smalldata/prostate\ directory now includes \prostate.csv\ and \prostate\_anomaly\_validation.csv\, providing a new dataset for external testing and anomaly detection validation. This change ensures the data is available for external tests as well, supporting the addition of Isolation Forest to GridSearch.

examples/smalldata/prostate · high confidence

Added sample datasets for examples

New CSV datasets have been added to the \examples/smalldata\ directory to support example workflows. These include \USArrests.csv\ (US crime statistics), \birds.csv\ (ecological landscape data), \cars\_20mpg.csv\ (vehicle fuel economy and specs), and \craigslistJobTitles.csv\ (job listing text).

examples/smalldata · high confidence

Automated Nexus staging and release scripts

Added shell scripts to automate the publishing workflow to the Sonatype Nexus repository. The new \createStagingId.sh\ initiates a staging repository, \uploadToNexus.sh\ uploads artifacts to that staging area, and \closeBucket.sh\ finalizes the release by closing the bucket. These scripts handle the necessary HTTP interactions with the OSS Sonatype service to streamline the release process.

gradle/publish · high confidence

Automated benchmark execution on AWS EMR

A new shell script, run\_benchmarks.sh, has been added to the benchmarks directory to automate the lifecycle of benchmark runs on AWS EMR. The script handles provisioning the infrastructure via Terraform, executing the benchmarks, polling for completion, downloading the results, and finally destroying the cluster. It supports configuration of AWS credentials, instance types, memory settings, and dataset specifications, providing a streamlined way to run benchmarks in a cloud environment.

benchmarks · high confidence

Automated generation of configuration, parameter, and metrics documentation

The documentation build process now automatically generates RST documentation for Sparkling Water configuration properties, algorithm parameters, model details, and metrics classes. New Scala tools in the doc generation package scan backend configuration objects and ML algorithm/model classes via reflection to produce structured tables and API references, ensuring that configuration tables, parameter lists (including default values and MOJO availability), and metric class details are kept in sync with the codebase without manual maintenance.

doc/src/main · high confidence

Documentation now supports interactive tabs and collapsible sections

A new Sphinx extension has been added to the documentation build process, enabling authors to create interactive UI elements such as tabbed content panels and collapsible toggle sections. This change introduces new RST directives (\content-tabs\, \tab-container\, \toggle-header\) along with the necessary CSS and JavaScript assets, allowing users to navigate between different content variants (e.g., language-specific examples) or hide/show detailed information directly within the documentation pages.

doc/src/site/sphinx/extensions · high confidence

Expanded R API for H2O MOJO models with new prediction options and metrics

The R API for H2O MOJO models now exposes additional prediction capabilities and detailed model information. Users can enable feature contributions, leaf node assignments, stage results, and prediction intervals via H2OMOJOSettings. The H2OMOJOModel class provides direct access to training, validation, and cross-validation metrics (as both DataFrames and objects), scoring history, feature importances, and model metadata such as start/end time and default threshold. A new H2OBinaryModel class allows reading and writing binary models, and H2OMOJOPipelineModel supports internal contributions and prediction intervals.

r/src/R/ai/h2o/sparkling/ml · high confidence

Expose Spark internal utilities and logging via new wrapper classes

Added \Logging.scala\ and \Utils.scala\ to the \org.apache.spark.expose\ package to provide public access to previously internal Spark components. The new \Logging\ trait extends \org.apache.spark.internal.Logging\ and is serializable, allowing external code to use Spark's logging infrastructure. The \Utils\ object wraps several \org.apache.spark.util.Utils\ and \ShutdownHookManager\ methods—including temporary directory creation, local directory retrieval, listener bus posting, and shutdown hook management—making these capabilities available outside the \org.apache\ namespace.

utils/src/main/scala/org/apache/spark/expose · high confidence

Expose binary model handling and target encoder model implementation

Users can now load and save H2O binary models directly via the Sparkling Water Scala API using the new H2OBinaryModel class, which validates H2O version compatibility and distributes model files via SparkFiles. Additionally, the H2OTargetEncoderModel is now fully exposed, allowing users to transform datasets using the target encoder's MOJO model or perform training-time transformations via the H2O REST API, with internal H2O frames properly managed and cleaned up.

ml/src/main/scala/ai/h2o/sparkling/ml/models · high confidence

H2OFrame support as a native Spark SQL data source

Users can now read and write H2O Frames directly using standard Spark SQL syntax (e.g., \spark.read.format("h2o")\). This change introduces the \DefaultSource\ relation provider and registers it via the \DataSourceRegister\ service file, enabling seamless integration of H2O data into Spark SQL workflows without requiring explicit API calls for data loading or saving.

core/src/main · high confidence

Initial AWS EMR Terraform template with modular network and EMR deployment

This change introduces the initial set of Terraform templates for deploying Sparkling Water on AWS EMR. The root configuration in \templates/src/terraform/aws/\ orchestrates two main modules: \network\, which provisions a VPC, subnet, internet gateway, and DHCP options, and \emr\, which handles the EMR cluster setup. The EMR module further delegates to \emr\_security\ for IAM roles and security groups, and \emr\_deployment\ for the actual cluster resources. Users can configure the deployment via variables such as \aws\_emr\_version\ (defaulting to a placeholder \SUBST\_EMR\_VERSION\), \aws\_instance\_type\ (defaulting to \m5.xlarge\), and \aws\_core\_instance\_count\ (defaulting to 2). The template also supports SSH access via \aws\_ssh\_public\_key\ and Jupyter notebook configuration via \jupyter\_name\. Outputs include the Jupyter notebook URL, master public DNS, and the S3 bucket name.

templates/src/terraform/aws/modules/emr · high confidence

Initial AWS infrastructure for Kubernetes testing

Added Terraform configuration to provision an AWS EKS cluster, VPC, and ECR repository, enabling the deployment and testing of the application on Kubernetes. The setup includes a VPC with public and private subnets, an EKS cluster using the terraform-aws-modules/eks module (version 13.2.1) with worker groups, and an immutable ECR repository for storing Sparkling Water images. Security groups allow all ingress and egress traffic for the worker group, and outputs expose the cluster endpoint, kubeconfig, and registry details for integration.

kubernetes/src/terraform · high confidence

Initial project scaffolding and multi-version support configuration

This change establishes the foundational structure for the Sparkling Water project. It introduces a comprehensive README.rst detailing usage, installation, and documentation links for Spark versions 2.3 through 3.5, alongside a new Code of Conduct. The build system is configured via Gradle wrapper scripts and specific property files (gradle-spark2.3.properties through gradle-spark3.5.properties) that define supported Spark, Scala, Python, and EMR versions for each release track. Additionally, development standards are enforced through .scalafmt.conf and greclipse.properties, while .gitignore is expanded to exclude generated resources, build artifacts, and Terraform state files.

(repo-wide) · high confidence

Initial release of the PySparkling Python API

This change introduces the complete Python API for Sparkling Water under the \ai.h2o.sparkling\ package. It provides the core \H2OContext\ and \H2OConf\ classes for managing the H2O cluster connection and configuration, along with a comprehensive set of PySpark ML-compatible estimators and transformers. Users can now access H2O algorithms (such as GLM, GBM, Deep Learning, AutoML, and Isolation Forest) and feature transformers (including Target Encoder, Word2Vec, PCA, and AutoEncoder) directly within the PySpark ML pipeline framework.

py/src/ai · high confidence

Initial release of the Python scoring package

The \py-scoring/src\ directory now contains the build infrastructure for the \h2o\_pysparkling\_scoring\ Python package. This includes the \setup.py\ configuration, \MANIFEST.in\ to ensure version files are included in source distributions, and Conda recipe files (\meta.yaml\, \build.sh\, \bld.bat\) to facilitate packaging and distribution. This change establishes the foundation for distributing the scoring component as a standalone Python package.

py-scoring/src · high confidence

Initial release of the RSparkling package

This change introduces the RSparkling package, providing an R interface to the H2O Sparkling Water machine learning library. The package extends sparklyr, allowing users to interact with H2O contexts, create H2O frames, and utilize H2O MOJO models for prediction directly within R. It includes core classes like H2OContext and H2OMOJOModel, along with necessary dependencies on sparklyr and the h2o package.

r/src · high confidence

Initial release of the Sparkling Water Booklet documentation

The Sparkling Water user guide (Booklet) is now available as a standalone LaTeX project within the repository. This includes the main document structure, configuration templates for code listings, and content sections covering the introduction, design, data manipulation, algorithm usage, deployment, and FAQ. A new Scala-based generation tool has been added to automatically produce the configuration properties section from backend code, ensuring the documentation stays synchronized with the current configuration options.

booklet · high confidence

Introduce Python Scoring Package for H2O MOJO Models

A new Python scoring package is added under \py-scoring/src/ai\, providing a self-contained set of PySparkling classes for loading, configuring, and scoring H2O MOJO models without requiring manual JAR configuration. This includes an \Initializer\ that automatically injects the \sparkling\_water\_scoring\_assembly.jar\ into the Spark environment, \H2ODataFrameConverters\ for seamless Scala-to-Python DataFrame translation, and a comprehensive suite of model wrappers (e.g., \H2OMOJOPipelineModel\, \H2OTargetEncoderModel\, \H2OAutoEncoderMOJOModel\) along with their corresponding parameter definitions. Users can now import and use these models directly in Python pipelines, with support for features like prediction intervals, SHAP contributions, and leaf node assignments exposed through the new \H2OMOJOSettings\ and model base classes.

py-scoring/src/ai · high confidence

New AWS infrastructure modules for test environment networking and container registry

Added Terraform modules to provision the underlying AWS infrastructure for the test environment. The network module creates a VPC, subnet, internet gateway, and DHCP options to provide network connectivity, while the ECR module provisions an immutable container registry for storing test artifacts. These changes establish the foundational cloud resources required to support the migration of the test infrastructure to AWS.

ci/aws/terraform/modules/ecr, ci/aws/terraform/modules/network · high confidence

New AWS-hosted Jenkins CI infrastructure module

The CI pipeline now provisions a dedicated Jenkins master on AWS (t2.medium) via a new Terraform module. This setup includes an S3 bucket for storing initialization scripts and credentials, Route53 DNS records for public access, and an Apache reverse-proxy with SSL termination managed by Certbot. The Jenkins instance is pre-configured with essential plugins (Blue Ocean, GitHub, AWS, Docker, etc.) and bootstrapped with global credentials for GitHub, AWS, Nexus, Docker Hub, and Databricks, enabling automated testing and artifact publishing workflows.

ci/aws/terraform/modules/jenkins · high confidence

New EMR security module for AWS infrastructure

A new Terraform module has been added to manage the security infrastructure for Amazon EMR clusters. This module provisions the necessary IAM roles and instance profiles for both the EMR service and EC2 nodes, configures master and slave security groups with open ingress/egress rules, and handles the VPC and subnet data lookups required for network placement.

_templates/src/terraform/aws/modules/emr\security · high confidence

New ExposeUtils helper for Spark type and class introspection

A new ExposeUtils object has been added to the Spark utilities package, providing static helper methods to check for specific data types (ML and MLlib vectors, general UserDefinedTypes) and to verify the presence of Hive classes. It also exposes a wrapper for the internal Utils.classForName method, allowing other components to perform class loading and type checking without directly accessing Spark's internal utility classes.

utils/src/main/scala/org/apache/spark · high confidence

New Jenkins pipeline for building and publishing test Docker images to AWS

A new Jenkinsfile (ci/docker/Jenkinsfile-build-docker) has been added to automate the build, tagging, and publishing of the Sparkling Water test Docker image. The pipeline authenticates with the internal Harbor registry, builds the image using Gradle, and pushes the final artifact to an AWS Docker repository. It also includes logic to automatically increment the image version in gradle.properties after a successful build. Additionally, a YARN configuration file (ci/docker/conf/yarn-site.xml) is included to define resource allocation and logging settings for the test environment.

ci/docker · high confidence

New Python API code-generation templates for algorithms, models, and metrics

The Python API generation system in api-generation now includes a comprehensive set of new Scala templates (AlgorithmTemplate, ConfigurationTemplate, MOJOModelTemplate, MetricsFactoryTemplate, ModelMetricsTemplate, ParametersTemplate, and others) that automatically produce Python classes for H2O algorithms, configuration objects, MOJO models, and model metrics. This enables the automatic generation of Python wrappers that expose algorithm parameters, handle MOJO model creation and cross-validation model retrieval, and provide access to model metrics, ensuring the Python API stays synchronized with the underlying Scala implementation.

api-generation/src/main/scala/ai/h2o/sparkling/api/generation/python · high confidence

New Python examples for Chicago Crime prediction and multi-algorithm spam classification

Added three new demonstration scripts to the Python examples directory: ChicagoCrimeDemo.py, which illustrates a full workflow for predicting crime arrests using H2O GBM and Deep Learning models on Chicago crime, weather, and census data; H2OContextInitDemo.py, a minimal example for initializing an H2OContext within a Spark session; and HamOrSpamMultiAlgorithmDemo.py, which demonstrates building and evaluating text classification pipelines using Spark ML features combined with H2O GBM, Deep Learning, AutoML, and XGBoost estimators.

py/examples · high confidence

New REST API endpoints for frame import and chunk management

The extensions module now registers a suite of new REST API endpoints to support external frame import and management. This includes endpoints for initializing and finalizing frames, retrieving upload plans, and reading or writing individual data chunks (with support for categorical domain handling and compression). Additionally, new handlers allow users to get and set the cluster log level, verify Sparkling Water availability, and check web openness and version across nodes. These capabilities are exposed via new servlets and API handlers registered through the standard extension mechanism.

extensions · high confidence

New Scala API code-generation templates for algorithms, parameters, and MOJO models

The Scala API generation engine now uses a new set of templates in the api-generation module to produce algorithm wrappers, parameter classes, MOJO model classes, and model metrics. This means the generated Scala API will have updated structure and behavior: algorithm classes are generated with specific MOJO model return types, parameter traits include explicit getters/setters and H2O parameter mapping, MOJO models expose algorithm-specific metrics and outputs with robust JSON parsing, and model metrics are generated as distinct objects with typed getters. Users relying on the generated Scala API will see these structural changes in the exposed classes and methods.

api-generation/src/main/scala/ai/h2o/sparkling/api/generation/scala · high confidence

New Spark SQL utility classes for data type handling and dataset manipulation

This change introduces three new utility files in the Spark SQL utils package to support internal data processing needs. DataTypeExtensions provides an implicit wrapper to convert Spark DataTypes to JSON values and parse JSON back to DataTypes. DatasetExtensions adds implicit methods to DataFrames for adding multiple columns at once and adding columns with specific metadata. SWGenericRow provides a custom subclass of Spark's GenericRow, likely to support specific serialization or schema requirements for the scoring package preparation.

utils/src/main/scala/org/apache/spark/sql · high confidence

New Sphinx-based documentation site structure

The documentation has been restructured into a new Sphinx site located at doc/src/site/sphinx. This change introduces a new build configuration (conf.py) using the Read the Docs theme and includes a comprehensive set of new documentation pages: an overview (index.rst), an about page, installation and requirements guides, PySparkling and RSparkling usage instructions, a migration guide covering breaking changes from version 3.30 through 3.44, a FAQ section addressing common configuration and error issues, and a changelog entry that includes the main CHANGELOG.rst file. This provides a centralized, structured location for user guides and migration instructions.

doc/src/site/sphinx · high confidence

New Terraform module for EMR-based Sparkling Water benchmarks

A new Terraform module has been added to provision an Amazon EMR cluster specifically for running Sparkling Water benchmarks. This infrastructure deploys the necessary S3 buckets for storing benchmark artifacts and results, configures the EMR cluster with Spark and Hadoop, and executes benchmark scripts via bootstrap actions. The module supports configurable execution modes (YARN internal/external and local) and includes an automatic cluster shutdown feature after a specified timeout to manage costs.

_benchmarks/src/main/terraform/aws/modules/emr\_benchmarks\deployment · high confidence

New and updated example applications for Sparkling Water

The examples directory now includes a comprehensive set of runnable Scala applications demonstrating Sparkling Water capabilities. New additions include structured streaming examples (CraigslistJobTitlesStructuredStreamingApp), a client-less H2OContext usage pattern (AirlinesWithWeatherDemo, ChicagoCrimeApp, CityBikeSharingDemo), and various ML pipelines (DeepLearningDemo, HamOrSpamDemo, ProstateDemo). Existing examples have been refactored to use the new Sparkling Water API (SW API) and H2OContext.getOrCreate(), ensuring they run without requiring a separate client process.

examples/src/main/scala · high confidence

New benchmark suite for DataFrame/H2OFrame conversion and model training

The benchmarks module now includes a comprehensive set of new benchmarks to measure the performance of Spark DataFrame to H2OFrame conversions (including direct conversion, via CSV files, and with S3 load times) and the training of supervised algorithms (GBM and GLM) from DataFrames and H2OFrames. This change introduces a new \BenchmarkBase\ framework and a \Runner\ that automatically discovers and executes these benchmarks, allowing users to evaluate the efficiency of data movement and model training workflows.

benchmarks/src/main/scala · high confidence

New common API generation infrastructure for Sparkling Water

This change introduces a new set of configuration and context classes in the \api-generation\ module (including \AlgorithmConfigurations\, \FeatureEstimatorConfigurations\, \AutoMLConfiguration\, and \GridSearchConfiguration\) that define how H2O algorithm parameters, metrics, and model outputs are mapped to the Sparkling Water Scala API. It establishes the structural basis for generating algorithm-specific parameter classes and metrics for algorithms such as GLM, GBM, Deep Learning, AutoML, and feature estimators like PCA and Word2Vec, effectively replacing or supplementing previous manual or less structured generation approaches.

api-generation/src/main/scala/ai/h2o/sparkling/api/generation/common · high confidence

New download page and build metadata for Sparkling Water

The distribution now includes a dedicated download page (index.html) and a structured build metadata file (buildinfo.json). The download page provides a user-friendly interface to access the Sparkling Water ZIP archive, along with links to developer documentation, Scala Scaladoc, and example templates. The buildinfo.json file exposes key versioning details—including the Sparkling Water, H2O, and Spark versions, build timestamp, and commit hash—enabling users and scripts to verify the exact composition of the distribution.

dist · high confidence

New feature estimators and utilities for Sparkling ML features

This change introduces several new components to the Sparkling ML features package. It adds a ColumnPruner transformer to selectively keep or drop columns from datasets. It also introduces base classes for new H2O-powered feature estimators: H2OAutoEncoderBase for autoencoder models, H2ODimReductionEstimator as a base for dimensionality reduction, H2OGLRMBase for Generalized Low Rank Models, and H2OWord2VecBase for word embeddings. Additionally, the H2OTargetEncoder is updated to support multinomial classification problems and regression tasks, and exposes parameters for setting output columns and handling interactions.

ml/src/main/scala/ai/h2o/sparkling/ml/features · high confidence

New macro for deprecating API methods with warnings

The macros module now includes a new \@DeprecatedMethod\ annotation and its corresponding macro implementation. When applied to methods, this annotation automatically injects a standard Scala \@deprecated\ annotation and logs a warning message at runtime, informing users of the deprecation, the recommended replacement, and the version in which the method will be removed. A test suite has been added to verify the macro's behavior with various annotation usages and method signatures.

macros · high confidence

New micro-benchmark suite for DataFrame-to-H2OFrame conversion performance

A new set of micro-benchmarks has been added to the core benchmarking infrastructure to measure the performance of converting Spark DataFrames to H2OFrames. The suite includes specific tests for flat DataFrames, DataFrames with nested structs, DataFrames with flat arrays, and DataFrames containing wide sparse or dense vectors. It also benchmarks the performance of flattening DataFrames and schemas, as well as converting rows to row schemas, providing detailed statistical results (mean, standard deviation, min, max) to help track regression or improvement in these data conversion paths.

core/src/bench · high confidence

New scoring package with MOJO model classes and metrics

The scoring module now includes a complete set of new Scala classes for handling H2O MOJO models, including H2OMOJOModel, H2OMOJOPipelineModel, and algorithm-specific variants like H2OAlgorithmMOJOModel. This introduces support for detailed predictions, feature contributions, leaf node assignments, and stage results. Additionally, a new H2OMetric enum and H2OMetrics trait system are added to expose training, validation, and cross-validation metrics (such as AUC, RMSE, and Logloss) as structured objects, replacing the previous map-based approach.

scoring · high confidence

New utilities for analyzing and flattening nested Spark DataFrames

Added \DatasetShape\ to classify Spark schemas as Flat, StructsOnly, or Nested, and \SchemaUtils\ to flatten nested DataFrames (containing structs, arrays, maps, or binary data) into a single-level schema. This enables downstream components to handle complex nested data structures uniformly by converting them into a flat representation suitable for standard processing.

utils/src/main/scala/ai/h2o/sparkling/ml · high confidence

New utility classes for serialization, compression, and Spark session management

Added a suite of new utility classes in the utils package to support data handling and interoperability. This includes Base64Encoding for converting byte, int, and long arrays to and from strings; Compression utilities supporting NONE, DEFLATE, GZIP, and SNAPPY formats; and a comprehensive DataFrame serialization system (DataFrameJsonSerialization, JSONDataFrameSerializer, and wrappers) that allows DataFrames to be serialized to JSON and deserialized back, enabling compatibility with Java serialization wrappers. Additionally, CompatibilityObjectInputStream is introduced to handle serialVersionUID mismatches during deserialization, FinalizingOutputStream ensures cleanup actions upon stream closure, and SparkSessionUtils provides helpers for managing active Spark sessions and reading HDFS files.

utils/src/main/scala/ai/h2o/sparkling/utils · high confidence

PySparkling Python API restructured with expanded algorithm and MOJO support

The PySparkling package has been reorganized into a single source directory, exposing a comprehensive set of H2O algorithms and model types in the Python API. This update adds support for new algorithms including PCA, GLRM, AutoEncoder, CoxPH, GAM, RuleFit, StackedEnsemble, and Extended Isolation Forest, alongside existing ones like DRF and Word2Vec. It also introduces algorithm-specific MOJO model classes (e.g., H2OAlgorithmMOJOModel, H2OSupervisedMOJOModel) and separates classification and regression variants for easier access. Users can now import these new estimators and models directly from pysparkling.ml.algos and pysparkling.ml.models.

py/src/pysparkling · high confidence

Architecture

Refactored ML parameter definitions into modular traits

The parameter definitions in the Sparkling ML params package have been reorganized into a modular trait-based architecture. Common parameters are now shared via base traits like H2OCommonParams and H2OAlgorithmCommonParams, while algorithm-specific options (such as monotone constraints, beta constraints, and calibration frames) are isolated in dedicated traits (e.g., HasMonotoneConstraints, HasBetaConstraints). This change improves code maintainability and reduces duplication across different model algorithms.

ml/src/main/scala/ai/h2o/sparkling/ml/params · high confidence

Refactored algorithm training infrastructure with new base classes and preparation traits

The algorithm wrappers in the \ml/src/main/scala/ai/h2o/sparkling/ml/algos\ package have been restructured to use a new inheritance hierarchy and preparation logic. New base classes such as \H2OAlgorithm\, \H2OSupervisedAlgorithm\, and \H2OUnsupervisedAlgorithm\ now serve as the foundation for specific models, while \H2OAlgoCommonUtils\ centralizes dataset preparation, column handling, and frame lifecycle management. New traits like \H2OTrainFramePreparation\, \DistributionBasedH2OTrainFramePreparation\, and \FamilyBasedH2OTrainFramePreparation\ allow algorithms to automatically convert label columns to categorical types based on their distribution or family settings. This change also introduces specific implementations for \H2OAutoML\, \H2OGridSearch\, and \H2OStackedEnsemble\ that leverage this new common infrastructure for training and model management.

ml/src/main/scala/ai/h2o/sparkling/ml/algos · high confidence

Behavioural changes

Add Spark 3.2 compatibility layer for SQL encoders and UDFs

This change introduces a new Scala source directory for Spark 3.2 that provides version-specific adapters for SQL functionality. It adds a RowEncoder wrapper to handle schema-based row encoding, a functions object to expose UserDefinedFunction creation using the Spark 3.2-specific SparkUserDefinedFunction constructor, and a SparkUserDefinedFunction facade that bridges the internal expressions API. These components ensure that Sparkling Water's SQL features, particularly custom UDFs and row encoding, function correctly on Spark 3.2 by abstracting away version-specific API differences.

_utils/src/main/scala\_spark\3.2 · high confidence

Added Spark 3.1 and 3.5 compatibility shims

New adapter files were added for Spark 3.1 and 3.5 to handle API differences in row encoding and user-defined functions. For Spark 3.1, the RowEncoder shim delegates directly to the Spark API, while for 3.5 it uses the new encoderFor method to accommodate the removal of the apply method. Additionally, a SparkUserDefinedFunction shim was introduced to abstract away constructor signature changes, ensuring consistent UDF creation across both versions.

_utils/src/main/scala\_spark\3.1 · high confidence

Added Spark 3.3-specific SQL compatibility shims

This change introduces version-specific adapter files for Spark 3.3 to maintain compatibility with the SQL API. It adds a \RowEncoder\ wrapper that delegates to the Spark 3.3 catalyst encoder, a \functions\ object that provides a \udf\ helper using the \SparkUserDefinedFunction\ signature required by this version, and a \SparkUserDefinedFunction\ companion object that bridges the internal \expressions\ package with the public API. These shims ensure that user-defined functions and row encoding behave consistently across supported Spark versions.

_utils/src/main/scala\_spark\3.3 · high confidence

Added Spark 3.4 compatibility layer for SQL encoders and UDFs

This change introduces a new source directory for Spark 3.4 that provides compatibility shims for SQL functionality. Specifically, it adds a \RowEncoder\ wrapper to handle schema encoding, a \functions\ object to expose User Defined Functions (UDFs) using the \SparkUserDefinedFunction\ API, and a \SparkUserDefinedFunction\ companion object to bridge internal Spark expression types. These additions allow the application to maintain a consistent API surface across different Spark versions by abstracting away version-specific implementation details in the SQL module.

_utils/src/main/scala\_spark\3.4 · high confidence

Added default logging configuration for examples

A new log4j.properties file has been added to the examples resources, configuring the default logging level to WARN for the console output. This setup suppresses verbose logs from third-party libraries like Jetty and sets specific INFO levels for Spark REPL components, providing a cleaner log experience when running demos from IDEs like IntelliJ.

examples/src/main/resources · high confidence

CRAN rsparkling package now redirects users to the custom repository

The rsparkling package available on CRAN has been updated to a dummy version that no longer provides functional machine learning capabilities. Instead, attempting to use the package now triggers an error message directing users to install the maintained version from the H2O custom repository, with specific installation commands provided for Spark versions 2.1 through 2.4. This change ensures users are guided toward the actively supported release rather than the deprecated CRAN distribution.

r-cran · high confidence

Compatibility layer for Spark UDFs and RowEncoders

Added version-specific compatibility shims in the Sparkling Water utilities to handle API differences between Spark versions. The new \RowEncoder\ and \functions\ objects (along with \SparkUserDefinedFunction\ for Spark 3.0) abstract away changes in how User-Defined Functions and Row Encoders are constructed, ensuring that code using these features works consistently across supported Spark versions without requiring conditional logic in the application layer.

_utils/src/main/scala\_spark\_3.0, utils/src/main/scala\_spark\others · high confidence

Enforced H2O version verification and bundled Sparkling Water JAR in RSparkling

The RSparkling package now includes a new \package.R\ file that enforces strict compatibility by verifying the installed H2O R package version against the specific Sparkling Water version at load time. If a mismatch is detected, users receive clear instructions to uninstall the current H2O package and install the correct version from the H2O release repository. Additionally, the package now bundles the Sparkling Water assembly JAR directly, removing the need for users to manually specify the JAR path or version, thereby simplifying setup and ensuring the R package version is tightly coupled with the correct Sparkling Water backend.

r/src/R · high confidence

GBM pipeline example now includes detailed predictions by default

The H2OGBM pipeline example in the build artifacts has been updated to enable the 'detailed\_prediction' column in MOJO outputs by default. This change is reflected in the stage metadata, which now sets the 'detailedPredictionCol' parameter to 'detailed\_prediction' and enables 'namedMojoOutputColumns', ensuring that users running this specific pipeline example receive granular prediction details without needing to manually configure these parameters.

py/examples/build · high confidence

Internal H2OModel class now exposes cross-validation MOJO models

The new H2OModel internal class in the Sparkling ML internals package now supports converting models to MOJO format while including cross-validation models. When the withCVModels flag is true, the toMOJOModel method retrieves cross-validation model details from the H2O cluster API and sets them on the resulting H2OMOJOModel instance, making cross-validation MOJO models accessible to users who request them.

ml/src/main/scala/ai/h2o/sparkling/ml/internals · high confidence

Migrate CI infrastructure to Jenkins and AWS

The continuous integration system has been migrated from the previous platform to Jenkins, with the testing infrastructure now hosted on AWS. This change introduces a suite of new Jenkins pipelines (including main, benchmarks, Databricks, Kubernetes, nightly, and release workflows) and replaces the old infrastructure provisioning with Terraform and Packer scripts located in ci/aws/terraform. The new setup manages AWS resources, Docker registries, and credential handling through a shared Groovy library, enabling automated testing on Kubernetes and Databricks alongside standard Spark builds.

ci · high confidence

Migrate test infrastructure to AWS

The test infrastructure has been moved to AWS, introducing Terraform modules to provision the network, an ECR repository for Docker images, and a Jenkins instance. This change updates the CI environment to use AWS resources, including a specific Docker registry and Jenkins URL, replacing the previous setup.

ci/aws/terraform · high confidence

New API generation runners for algorithms, configurations, and MOJO models

The api-generation module now includes three new entry-point runners: AlgorithmAPIRunner generates algorithm-specific classes (including Word2Vec and problem-specific classifiers/regressors) for Scala and Python; ConfigurationRunner generates configuration bindings for Scala, Python, and R by reflecting on backend configuration classes; and MOJOModelAPIRunner generates MOJO model classes, factories, and metrics objects for Scala, Python, and R. These runners replace or supplement previous manual generation steps, ensuring that algorithm, configuration, and MOJO APIs are consistently generated across languages.

api-generation/src/main/scala/ai/h2o/sparkling/api/generation · medium confidence

New AWS Packer configuration for Java 11 Jenkins slaves

Added a new Packer template and initialization script to provision AWS EC2 instances for Jenkins slaves. The configuration builds an Amazon Linux AMI with Java 11 (OpenJDK 11), Docker, and Git, creating a dedicated 'jenkins' user with Docker group access to support the CI infrastructure migration.

ci/aws/packer · high confidence

New Scala editor styling in Flow CSS

A new CSS file (scala-editor.css) has been added to the core/flow-css directory, defining syntax highlighting and layout styles for the Scala code editor. The styles adopt an IntelliJ IDEA default theme, applying specific colors to keywords, numbers, strings, and comments, while also setting the font family to a monospace stack and configuring the editor container's dimensions, padding, and a distinctive left border.

core/flow-css · high confidence

New utility traits for REST-based model training and parameter reading

The \ml/src/main/scala/ai/h2o/sparkling/ml/utils\ package now includes \EstimatorCommonUtils\ and \H2OParamsReadable\. \EstimatorCommonUtils\ enables training models via the REST API by sending parameters to a cluster endpoint, waiting for job completion with progress printing, and handling warnings from model builders; it also provides utilities for downloading binary models and resolving model ID conflicts. \H2OParamsReadable\ introduces an \MLReadable\ implementation for reading H2O parameters, supporting the standard Spark ML parameter serialization pattern.

ml/src/main/scala/ai/h2o/sparkling/ml/utils · high confidence

R API generation now supports overloaded configuration setters and exposes model metrics as R objects

The R API generation templates have been updated to automatically generate R configuration classes from Scala code, including proper handling of overloaded setter methods in configuration objects. Additionally, the generation logic now exposes model metrics as distinct R objects via a new factory and template system, allowing users to access detailed metric data from H2OMOJOModel instances in R.

api-generation/src/main/scala/ai/h2o/sparkling/api/generation/r · high confidence

RSparkling API modernization and client-less default

RSparkling now uses a client-less approach by default, simplifying initialization via a parameterless H2OContext.getOrCreate() and removing the need to pass SparkContext explicitly. The API exposes H2OConf getters and setters, adds H2OContext.asSparkFrame() (deprecating asDataFrame), and enables REST API conversions with compression support.

r/src/R/ai/h2o/sparkling · high confidence

Repl interpreter refactored to support parallel sessions and Spark 2.12

The Scala REPL interpreter in the \repl/src/main\ module has been restructured to support multiple concurrent interpreter sessions and compatibility with Spark 2.12. This change introduces a new \H2OIMain\ class that wraps the Scala \IMain\ and uses session-specific package naming (e.g., \intp\id\\<sessionId\>\) to isolate classes defined in each session. A custom \InterpreterClassLoader\ routes class loading to the correct session-specific interpreter, allowing parallel execution. The refactoring also includes a runtime patch (\PatchUtils\) to fix \OuterScopes\ regex matching for these session-specific classes, and consolidates interpreter logic into \BaseH2OInterpreter\ and \H2OInterpreter\ to remove version-specific duplication.

repl/src/main · high confidence

Restructure Python package build configuration

The Python package build system has been reorganized to ensure correct packaging and version handling. A new MANIFEST.in file explicitly includes version and build info files (version.txt, buildinfo.txt) for both h2o and sparkling water components, ensuring they are included in source distributions. The setup.py and setup.cfg files have been updated to define the package metadata, dependencies (requests, tabulate), and package data, facilitating proper installation via pip.

py/src · high confidence

Unify and modernize Sparkling Water launcher scripts and environment configuration

The bin directory has been refactored to provide a consistent, unified set of launcher scripts for Sparkling Water across Linux/macOS and Windows. The new scripts (such as run-sparkling.sh, run-sparkling.cmd, sparkling-shell, and pysparkling) centralize environment validation by sourcing sparkling-env.sh (or sparkling-env.cmd), which now reads version and path details directly from gradle.properties. This change introduces a new default MASTER value of 'local\[\*\]' and sets a default driver memory of 2G. Additionally, the PySparkling launcher now explicitly checks for required Python packages (requests, tabulate) before starting, and the build-kubernetes-images.sh script has been updated to support building images for the external backend mode.

bin · high confidence

Updated RSparkling examples to target Spark 2.4.6

The included R demo scripts (NYC flights analysis and simple start) have been updated to explicitly install and connect to Spark version 2.4.6, replacing the previous 2.4.5 references. This ensures that users running these examples will provision the correct Spark runtime version compatible with the current release.

r/src/inst · high confidence

Updated documentation theme with Font Awesome 4.2.0 and new CSS assets

The Sphinx documentation theme has been updated to include Font Awesome 4.2.0, providing a broader set of icons for use in the documentation interface. This change introduces new CSS files (\badge\_only.css\ and \theme.css\) that define the styling for the theme, including the icon definitions and layout adjustments for the version switcher and navigation elements.

_doc/src/site/sphinx/sphinx\_rtd\theme · high confidence

Fixes

Added Iris dataset with headers for external testing

A new CSV file, \iris\_wheader.csv\, has been added to the \examples/smalldata/iris\ directory. This file contains the standard Iris dataset including a header row (sepal\_len, sepal\_wid, petal\_len, petal\_wid, class) and 150 data points across the three species (Iris-setosa, Iris-versicolor, Iris-virginica). This change ensures the data is available as a committed resource for external tests, addressing the issue where the data location was previously inaccessible.

examples/smalldata/iris · high confidence

Test coverage

1 commit adding/updating tests in ml/src/test/resources/target\_encoder; Add R Kubernetes integration test script; Add test data for prediction interval validation; Added API test suites for DataFrames, H2OFrames, RDDs, and Scala Interpreter; Added Kubernetes smoke tests for cluster initialization and data workflows; Added Python integration test infrastructure and baseline examples; Added R tests for H2O configuration, frame conversions, and MOJO model features; Added Scala 2.11 and 2.12 test resources for H2OMOJOModel serialization; Added comprehensive test coverage for Spark-to-H2O data conversion; Added integration tests for SSL, LDAP, and PAM authentication modes; Added integration tests for complex schema flattening; Added integration tests for example applications; Added integration tests for external H2O backend scenarios; Added test data for H2O MOJO SHAP contributions; Added test data for Mojo pipeline with multiple output columns; Added test data for SHAPley contributions on transformed features; Added test resources for H2O MOJO pipeline validation; Added test resources for binary model predictions and test logging configuration; Added test suites for H2O feature estimators; Added test suites for ML algorithm prediction outputs and metrics; Added test suites for MOJO model behavior and binary model lifecycle; Added test suites for backend utility components; Added test suites for configuration, data sources, authentication, and frame operations; Added tests for EnumParamValidator; Added tests for H2O model metrics handling; Added tests for H2OContext lifecycle and import scenarios; Added tests for JSON serialization and schema flattening utilities; Added tests for PartitionStatsGenerator; Added tests for Sparkling Water ML pipeline prediction and streaming; Added unit tests for REPL interpreter settings and patch utilities; Initial Python scoring test suite for PySparkling; New unit test suite for PySparkling algorithms and data conversions; Test logging configuration added; Updated test resource model to use H2O MOJO pipeline.

Dependencies

Introduce modular Gradle build structure with dedicated assembly and utility subprojects

The build system has been restructured into a multi-module Gradle project, introducing new subprojects for API generation, extensions assembly, scoring assembly, and utilities. This change centralizes dependency management and artifact assembly, ensuring that the main Sparkling Water assembly jar correctly excludes transitive dependencies provided by Apache Spark and Hadoop, while explicitly relocating conflicting packages (such as Jetty and Guava) to the \ai.h2o\ namespace to prevent classpath collisions. It also adds dedicated build configurations for benchmarks and documentation generation.

(dependencies) · high confidence

Upgrade Gradle wrapper to version 7.6

The Gradle wrapper has been updated to use Gradle 7.6 (all distribution). This ensures that builds are executed with a consistent, modern version of the Gradle build tool, leveraging improvements and bug fixes present in the 7.x release line.

gradle/wrapper · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 58 → 62 (+3.4)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 94 → 94 (-0.0)
  • Architecture 99 → 92 (-7.4)
  • Maturity 55 → 55 (+0.0)
  • Readiness 59 → 59 (-0.3)
  • Security 50 → 62 (+12.5)

Resolved (2)

  • Documentation: no installation or build instructions
  • Documentation: no usage examples

New (5)

  • Documentation: no project overview (README.rst)
  • Duplicated block (6 lines × 2) (py-scoring/src/ai/h2o/sparkling/H2ODataFrameConverters.py)
  • Duplicated block (8 lines × 2) (extensions/src/main/scala/ai/h2o/sparkling/extensions/internals/ConvertCategoricalToStringColumnsTask.java)
  • No ADRs found
  • Split core

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

h2oai/sparkling-water was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 1fc2265322f5ff1575890c75df384119ff55aa98 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.