h2oai/sparkling-water
61.6
Adequate · 28 September 2026
35.7k
lines of production code
Scala
with Python
2
measurements over time
What this system is
Sparkling Water is a bridge that integrates the H2O machine learning platform with Apache Spark, enabling users to run H2O algorithms and models within Spark workflows. It provides multi-language APIs for Scala, Python, and R, allowing seamless conversion between Spark DataFrames and H2O Frames while supporting standard Spark ML pipelines. The system also includes infrastructure for deploying on AWS EMR and Kubernetes, along with automated tools for benchmarking, documentation, and continuous integration.
How it got here
2014–2019 — Project scaffolding and multi-language API expansion
49 changes.
This period established the foundational structure of the Sparkling Water project, introducing multi-version support for Spark and a modular Gradle build system. It focused on expanding the project's reach by releasing initial Python (PySparkling) and R (RSparkling) APIs, while simultaneously enhancing the ML infrastructure with new algorithm wrappers, MOJO scoring capabilities, and comprehensive test coverage.
2020 — API generation and CI migration
48 changes.
This period focused on automating the generation of Scala, Python, and R APIs through a new code-generation infrastructure, while simultaneously migrating the continuous integration pipeline to Jenkins on AWS. Significant efforts were also directed toward modernizing the RSparkling client interface, expanding test coverage across multiple backends, and adding comprehensive documentation generation tools.
2021–2023 — Python scoring package and Spark compatibility
15 changes.
This period focused on releasing the standalone Python scoring package for H2O MOJO models, enabling seamless integration into Python pipelines with features like SHAP contributions and prediction intervals. Concurrently, the codebase expanded its compatibility layer to support Spark versions 3.1 through 3.5 by implementing version-specific shims for SQL encoders and user-defined functions. Extensive test coverage was added to validate the new Python package, serialization across Scala versions, and the correctness of model metrics and SHAP calculations.
Features
Add AWS EMR Terraform configuration for benchmarks
New Terraform files (main.tf, variables.tf, outputs.tf) have been added to the benchmarks infrastructure to provision Amazon EMR clusters for running benchmarks. This configuration allows users to deploy benchmark workloads on AWS using EMR, with customizable parameters for instance types, cluster size, H2O version, and memory settings.
benchmarks/src/main/terraform · high confidence
Add AWS EMR bootstrap script for Sparkling Water installation
A new bootstrap script (install\_sparkling\_water.sh) has been added to the AWS templates to automate the installation of Sparkling Water on AWS EMR clusters. This script configures the Python environment by upgrading pip, installing specific dependencies (requests, tabulate, six, scikit-learn), and downloading the Sparkling Water archive, enabling PySparkling functionality on worker nodes.
templates/src/aws · high confidence
Add Terraform module for deploying Sparkling Water on AWS EMR
A new Terraform module has been added to provision a Sparkling Water cluster on AWS EMR. This module manages the underlying infrastructure, including an S3 bucket for deployment artifacts, an EMR cluster configured with Spark, Hadoop, and JupyterHub, and the necessary IAM roles and security groups. It includes bootstrap scripts to install Sparkling Water dependencies and configure Jupyter access, and exposes outputs for the Jupyter notebook URL, master DNS, and S3 bucket name.
_templates/src/terraform/aws/modules/emr\deployment · high confidence
Add conda recipe for version-specific PySparkling package
A new conda build recipe has been added for the PySparkling package, allowing it to be built and installed via conda. The recipe defines build and runtime dependencies including Python, pip, setuptools, tabulate, requests, and future, and uses a placeholder for the Spark major version to support multiple Spark versions.
py/conda · high confidence
Added Chicago weather datasets for external testing
Two new CSV files, \Chicago\_Ohare\_International\_Airport.csv\ and \chicagoAllWeather.csv\, have been added to the \examples/smalldata/chicago\ directory. These files provide historical weather data (including temperature, precipitation, and snow metrics) for Chicago, making them available for use in external tests and examples.
examples/smalldata/chicago · high confidence
Added ECG anomaly detection datasets
The examples/smalldata/anomaly directory now includes \ecg\_discord\_train.csv\ and \ecg\_discord\_test.csv\, providing new training and test data for ECG-based anomaly detection examples.
examples/smalldata/anomaly · high confidence
Added NYC City Bike and weather sample datasets
The examples/smalldata/citybike-nyc directory now includes two new CSV files: citybike\_2013.csv, containing 2013 NYC Citi Bike trip data (durations, stations, user types), and New\_York\_City\_Hourly\_Weather\_2013.csv, providing hourly weather observations for New York City in 2013. These files serve as sample data for the citybike-nyc example.
examples/smalldata/citybike-nyc · high confidence
Added enriched airline dataset with graph features
The \examples/smalldata/airlines\ directory now includes a new CSV file, \airlines\_big\_data\_100.csv\, which provides an expanded version of the airline dataset. This file contains the original flight data augmented with numerous graph-theory features, including various centrality measures (Degree, Closeness, Harmonic, SubGraph, Load, Pagerank, EigenVector) and node embeddings (node2vec) for both origin and destination airports, enabling more complex network-based analysis.
examples/smalldata/airlines · high confidence
Added heart transplant dataset for CoxPH testing
The \examples/smalldata/coxph\_test\ directory now includes \heart.csv\ and \heart\_test.csv\, providing training and test data for the Cox Proportional Hazards model. These files contain columns for start/stop times, event status, age, year, surgery, and transplant status, enabling users to run and validate CoxPH algorithm examples.
_examples/smalldata/coxph\test · high confidence
Added prostate cancer dataset and anomaly validation files to examples
The \examples/smalldata/prostate\ directory now includes \prostate.csv\ and \prostate\_anomaly\_validation.csv\, providing a new dataset for external testing and anomaly detection validation. This change ensures the data is available for external tests as well, supporting the addition of Isolation Forest to GridSearch.
examples/smalldata/prostate · high confidence
Added sample datasets for examples
New CSV datasets have been added to the \examples/smalldata\ directory to support example workflows. These include \USArrests.csv\ (US crime statistics), \birds.csv\ (ecological landscape data), \cars\_20mpg.csv\ (vehicle fuel economy and specs), and \craigslistJobTitles.csv\ (job listing text).
examples/smalldata · high confidence
Automated Nexus staging and release scripts
Added shell scripts to automate the publishing workflow to the Sonatype Nexus repository. The new \createStagingId.sh\ initiates a staging repository, \uploadToNexus.sh\ uploads artifacts to that staging area, and \closeBucket.sh\ finalizes the release by closing the bucket. These scripts handle the necessary HTTP interactions with the OSS Sonatype service to streamline the release process.
gradle/publish · high confidence
Automated benchmark execution on AWS EMR
A new shell script, run\_benchmarks.sh, has been added to the benchmarks directory to automate the lifecycle of benchmark runs on AWS EMR. The script handles provisioning the infrastructure via Terraform, executing the benchmarks, polling for completion, downloading the results, and finally destroying the cluster. It supports configuration of AWS credentials, instance types, memory settings, and dataset specifications, providing a streamlined way to run benchmarks in a cloud environment.
benchmarks · high confidence
Automated generation of configuration, parameter, and metrics documentation
The documentation build process now automatically generates RST documentation for Sparkling Water configuration properties, algorithm parameters, model details, and metrics classes. New Scala tools in the doc generation package scan backend configuration objects and ML algorithm/model classes via reflection to produce structured tables and API references, ensuring that configuration tables, parameter lists (including default values and MOJO availability), and metric class details are kept in sync with the codebase without manual maintenance.
doc/src/main · high confidence
Documentation now supports interactive tabs and collapsible sections
A new Sphinx extension has been added to the documentation build process, enabling authors to create interactive UI elements such as tabbed content panels and collapsible toggle sections. This change introduces new RST directives (\content-tabs\, \tab-container\, \toggle-header\) along with the necessary CSS and JavaScript assets, allowing users to navigate between different content variants (e.g., language-specific examples) or hide/show detailed information directly within the documentation pages.
doc/src/site/sphinx/extensions · high confidence
Expanded R API for H2O MOJO models with new prediction options and metrics
The R API for H2O MOJO models now exposes additional prediction capabilities and detailed model information. Users can enable feature contributions, leaf node assignments, stage results, and prediction intervals via H2OMOJOSettings. The H2OMOJOModel class provides direct access to training, validation, and cross-validation metrics (as both DataFrames and objects), scoring history, feature importances, and model metadata such as start/end time and default threshold. A new H2OBinaryModel class allows reading and writing binary models, and H2OMOJOPipelineModel supports internal contributions and prediction intervals.
r/src/R/ai/h2o/sparkling/ml · high confidence
Expose Spark internal utilities and logging via new wrapper classes
Added \Logging.scala\ and \Utils.scala\ to the \org.apache.spark.expose\ package to provide public access to previously internal Spark components. The new \Logging\ trait extends \org.apache.spark.internal.Logging\ and is serializable, allowing external code to use Spark's logging infrastructure. The \Utils\ object wraps several \org.apache.spark.util.Utils\ and \ShutdownHookManager\ methods—including temporary directory creation, local directory retrieval, listener bus posting, and shutdown hook management—making these capabilities available outside the \org.apache\ namespace.
utils/src/main/scala/org/apache/spark/expose · high confidence
Expose binary model handling and target encoder model implementation
Users can now load and save H2O binary models directly via the Sparkling Water Scala API using the new H2OBinaryModel class, which validates H2O version compatibility and distributes model files via SparkFiles. Additionally, the H2OTargetEncoderModel is now fully exposed, allowing users to transform datasets using the target encoder's MOJO model or perform training-time transformations via the H2O REST API, with internal H2O frames properly managed and cleaned up.
ml/src/main/scala/ai/h2o/sparkling/ml/models · high confidence
H2OFrame support as a native Spark SQL data source
Users can now read and write H2O Frames directly using standard Spark SQL syntax (e.g., \spark.read.format("h2o")\). This change introduces the \DefaultSource\ relation provider and registers it via the \DataSourceRegister\ service file, enabling seamless integration of H2O data into Spark SQL workflows without requiring explicit API calls for data loading or saving.
core/src/main · high confidence
Initial AWS EMR Terraform template with modular network and EMR deployment
This change introduces the initial set of Terraform templates for deploying Sparkling Water on AWS EMR. The root configuration in \templates/src/terraform/aws/\ orchestrates two main modules: \network\, which provisions a VPC, subnet, internet gateway, and DHCP options, and \emr\, which handles the EMR cluster setup. The EMR module further delegates to \emr\_security\ for IAM roles and security groups, and \emr\_deployment\ for the actual cluster resources. Users can configure the deployment via variables such as \aws\_emr\_version\ (defaulting to a placeholder \SUBST\_EMR\_VERSION\), \aws\_instance\_type\ (defaulting to \m5.xlarge\), and \aws\_core\_instance\_count\ (defaulting to 2). The template also supports SSH access via \aws\_ssh\_public\_key\ and Jupyter notebook configuration via \jupyter\_name\. Outputs include the Jupyter notebook URL, master public DNS, and the S3 bucket name.
templates/src/terraform/aws/modules/emr · high confidence
Initial AWS infrastructure for Kubernetes testing
Added Terraform configuration to provision an AWS EKS cluster, VPC, and ECR repository, enabling the deployment and testing of the application on Kubernetes. The setup includes a VPC with public and private subnets, an EKS cluster using the terraform-aws-modules/eks module (version 13.2.1) with worker groups, and an immutable ECR repository for storing Sparkling Water images. Security groups allow all ingress and egress traffic for the worker group, and outputs expose the cluster endpoint, kubeconfig, and registry details for integration.
kubernetes/src/terraform · high confidence
Initial project scaffolding and multi-version support configuration
This change establishes the foundational structure for the Sparkling Water project. It introduces a comprehensive README.rst detailing usage, installation, and documentation links for Spark versions 2.3 through 3.5, alongside a new Code of Conduct. The build system is configured via Gradle wrapper scripts and specific property files (gradle-spark2.3.properties through gradle-spark3.5.properties) that define supported Spark, Scala, Python, and EMR versions for each release track. Additionally, development standards are enforced through .scalafmt.conf and greclipse.properties, while .gitignore is expanded to exclude generated resources, build artifacts, and Terraform state files.
(repo-wide) · high confidence
Initial release of the PySparkling Python API
This change introduces the complete Python API for Sparkling Water under the \ai.h2o.sparkling\ package. It provides the core \H2OContext\ and \H2OConf\ classes for managing the H2O cluster connection and configuration, along with a comprehensive set of PySpark ML-compatible estimators and transformers. Users can now access H2O algorithms (such as GLM, GBM, Deep Learning, AutoML, and Isolation Forest) and feature transformers (including Target Encoder, Word2Vec, PCA, and AutoEncoder) directly within the PySpark ML pipeline framework.
py/src/ai · high confidence
Initial release of the Python scoring package
The \py-scoring/src\ directory now contains the build infrastructure for the \h2o\_pysparkling\_scoring\ Python package. This includes the \setup.py\ configuration, \MANIFEST.in\ to ensure version files are included in source distributions, and Conda recipe files (\meta.yaml\, \build.sh\, \bld.bat\) to facilitate packaging and distribution. This change establishes the foundation for distributing the scoring component as a standalone Python package.
py-scoring/src · high confidence
Initial release of the RSparkling package
This change introduces the RSparkling package, providing an R interface to the H2O Sparkling Water machine learning library. The package extends sparklyr, allowing users to interact with H2O contexts, create H2O frames, and utilize H2O MOJO models for prediction directly within R. It includes core classes like H2OContext and H2OMOJOModel, along with necessary dependencies on sparklyr and the h2o package.
r/src · high confidence
Initial release of the Sparkling Water Booklet documentation
The Sparkling Water user guide (Booklet) is now available as a standalone LaTeX project within the repository. This includes the main document structure, configuration templates for code listings, and content sections covering the introduction, design, data manipulation, algorithm usage, deployment, and FAQ. A new Scala-based generation tool has been added to automatically produce the configuration properties section from backend code, ensuring the documentation stays synchronized with the current configuration options.
booklet · high confidence
Introduce Python Scoring Package for H2O MOJO Models
A new Python scoring package is added under \py-scoring/src/ai\, providing a self-contained set of PySparkling classes for loading, configuring, and scoring H2O MOJO models without requiring manual JAR configuration. This includes an \Initializer\ that automatically injects the \sparkling\_water\_scoring\_assembly.jar\ into the Spark environment, \H2ODataFrameConverters\ for seamless Scala-to-Python DataFrame translation, and a comprehensive suite of model wrappers (e.g., \H2OMOJOPipelineModel\, \H2OTargetEncoderModel\, \H2OAutoEncoderMOJOModel\) along with their corresponding parameter definitions. Users can now import and use these models directly in Python pipelines, with support for features like prediction intervals, SHAP contributions, and leaf node assignments exposed through the new \H2OMOJOSettings\ and model base classes.
py-scoring/src/ai · high confidence
New AWS infrastructure modules for test environment networking and container registry
Added Terraform modules to provision the underlying AWS infrastructure for the test environment. The network module creates a VPC, subnet, internet gateway, and DHCP options to provide network connectivity, while the ECR module provisions an immutable container registry for storing test artifacts. These changes establish the foundational cloud resources required to support the migration of the test infrastructure to AWS.
ci/aws/terraform/modules/ecr, ci/aws/terraform/modules/network · high confidence
New AWS-hosted Jenkins CI infrastructure module
The CI pipeline now provisions a dedicated Jenkins master on AWS (t2.medium) via a new Terraform module. This setup includes an S3 bucket for storing initialization scripts and credentials, Route53 DNS records for public access, and an Apache reverse-proxy with SSL termination managed by Certbot. The Jenkins instance is pre-configured with essential plugins (Blue Ocean, GitHub, AWS, Docker, etc.) and bootstrapped with global credentials for GitHub, AWS, Nexus, Docker Hub, and Databricks, enabling automated testing and artifact publishing workflows.
ci/aws/terraform/modules/jenkins · high confidence
New EMR security module for AWS infrastructure
A new Terraform module has been added to manage the security infrastructure for Amazon EMR clusters. This module provisions the necessary IAM roles and instance profiles for both the EMR service and EC2 nodes, configures master and slave security groups with open ingress/egress rules, and handles the VPC and subnet data lookups required for network placement.
_templates/src/terraform/aws/modules/emr\security · high confidence
New ExposeUtils helper for Spark type and class introspection
A new ExposeUtils object has been added to the Spark utilities package, providing static helper methods to check for specific data types (ML and MLlib vectors, general UserDefinedTypes) and to verify the presence of Hive classes. It also exposes a wrapper for the internal Utils.classForName method, allowing other components to perform class loading and type checking without directly accessing Spark's internal utility classes.
utils/src/main/scala/org/apache/spark · high confidence
New Jenkins pipeline for building and publishing test Docker images to AWS
A new Jenkinsfile (ci/docker/Jenkinsfile-build-docker) has been added to automate the build, tagging, and publishing of the Sparkling Water test Docker image. The pipeline authenticates with the internal Harbor registry, builds the image using Gradle, and pushes the final artifact to an AWS Docker repository. It also includes logic to automatically increment the image version in gradle.properties after a successful build. Additionally, a YARN configuration file (ci/docker/conf/yarn-site.xml) is included to define resource allocation and logging settings for the test environment.
ci/docker · high confidence
New Python API code-generation templates for algorithms, models, and metrics
The Python API generation system in api-generation now includes a comprehensive set of new Scala templates (AlgorithmTemplate, ConfigurationTemplate, MOJOModelTemplate, MetricsFactoryTemplate, ModelMetricsTemplate, ParametersTemplate, and others) that automatically produce Python classes for H2O algorithms, configuration objects, MOJO models, and model metrics. This enables the automatic generation of Python wrappers that expose algorithm parameters, handle MOJO model creation and cross-validation model retrieval, and provide access to model metrics, ensuring the Python API stays synchronized with the underlying Scala implementation.
api-generation/src/main/scala/ai/h2o/sparkling/api/generation/python · high confidence
New Python examples for Chicago Crime prediction and multi-algorithm spam classification
Added three new demonstration scripts to the Python examples directory: ChicagoCrimeDemo.py, which illustrates a full workflow for predicting crime arrests using H2O GBM and Deep Learning models on Chicago crime, weather, and census data; H2OContextInitDemo.py, a minimal example for initializing an H2OContext within a Spark session; and HamOrSpamMultiAlgorithmDemo.py, which demonstrates building and evaluating text classification pipelines using Spark ML features combined with H2O GBM, Deep Learning, AutoML, and XGBoost estimators.
py/examples · high confidence
New REST API endpoints for frame import and chunk management
The extensions module now registers a suite of new REST API endpoints to support external frame import and management. This includes endpoints for initializing and finalizing frames, retrieving upload plans, and reading or writing individual data chunks (with support for categorical domain handling and compression). Additionally, new handlers allow users to get and set the cluster log level, verify Sparkling Water availability, and check web openness and version across nodes. These capabilities are exposed via new servlets and API handlers registered through the standard extension mechanism.
extensions · high confidence
New Scala API code-generation templates for algorithms, parameters, and MOJO models
The Scala API generation engine now uses a new set of templates in the api-generation module to produce algorithm wrappers, parameter classes, MOJO model classes, and model metrics. This means the generated Scala API will have updated structure and behavior: algorithm classes are generated with specific MOJO model return types, parameter traits include explicit getters/setters and H2O parameter mapping, MOJO models expose algorithm-specific metrics and outputs with robust JSON parsing, and model metrics are generated as distinct objects with typed getters. Users relying on the generated Scala API will see these structural changes in the exposed classes and methods.
api-generation/src/main/scala/ai/h2o/sparkling/api/generation/scala · high confidence
New Spark SQL utility classes for data type handling and dataset manipulation
This change introduces three new utility files in the Spark SQL utils package to support internal data processing needs. DataTypeExtensions provides an implicit wrapper to convert Spark DataTypes to JSON values and parse JSON back to DataTypes. DatasetExtensions adds implicit methods to DataFrames for adding multiple columns at once and adding columns with specific metadata. SWGenericRow provides a custom subclass of Spark's GenericRow, likely to support specific serialization or schema requirements for the scoring package preparation.
utils/src/main/scala/org/apache/spark/sql · high confidence
New Sphinx-based documentation site structure
The documentation has been restructured into a new Sphinx site located at doc/src/site/sphinx. This change introduces a new build configuration (conf.py) using the Read the Docs theme and includes a comprehensive set of new documentation pages: an overview (index.rst), an about page, installation and requirements guides, PySparkling and RSparkling usage instructions, a migration guide covering breaking changes from version 3.30 through 3.44, a FAQ section addressing common configuration and error issues, and a changelog entry that includes the main CHANGELOG.rst file. This provides a centralized, structured location for user guides and migration instructions.
doc/src/site/sphinx · high confidence
New Terraform module for EMR-based Sparkling Water benchmarks
A new Terraform module has been added to provision an Amazon EMR cluster specifically for running Sparkling Water benchmarks. This infrastructure deploys the necessary S3 buckets for storing benchmark artifacts and results, configures the EMR cluster with Spark and Hadoop, and executes benchmark scripts via bootstrap actions. The module supports configurable execution modes (YARN internal/external and local) and includes an automatic cluster shutdown feature after a specified timeout to manage costs.
_benchmarks/src/main/terraform/aws/modules/emr\_benchmarks\deployment · high confidence
New and updated example applications for Sparkling Water
The examples directory now includes a comprehensive set of runnable Scala applications demonstrating Sparkling Water capabilities. New additions include structured streaming examples (CraigslistJobTitlesStructuredStreamingApp), a client-less H2OContext usage pattern (AirlinesWithWeatherDemo, ChicagoCrimeApp, CityBikeSharingDemo), and various ML pipelines (DeepLearningDemo, HamOrSpamDemo, ProstateDemo). Existing examples have been refactored to use the new Sparkling Water API (SW API) and H2OContext.getOrCreate(), ensuring they run without requiring a separate client process.
examples/src/main/scala · high confidence
New benchmark suite for DataFrame/H2OFrame conversion and model training
The benchmarks module now includes a comprehensive set of new benchmarks to measure the performance of Spark DataFrame to H2OFrame conversions (including direct conversion, via CSV files, and with S3 load times) and the training of supervised algorithms (GBM and GLM) from DataFrames and H2OFrames. This change introduces a new \BenchmarkBase\ framework and a \Runner\ that automatically discovers and executes these benchmarks, allowing users to evaluate the efficiency of data movement and model training workflows.
benchmarks/src/main/scala · high confidence
New common API generation infrastructure for Sparkling Water
This change introduces a new set of configuration and context classes in the \api-generation\ module (including \AlgorithmConfigurations\, \FeatureEstimatorConfigurations\, \AutoMLConfiguration\, and \GridSearchConfiguration\) that define how H2O algorithm parameters, metrics, and model outputs are mapped to the Sparkling Water Scala API. It establishes the structural basis for generating algorithm-specific parameter classes and metrics for algorithms such as GLM, GBM, Deep Learning, AutoML, and feature estimators like PCA and Word2Vec, effectively replacing or supplementing previous manual or less structured generation approaches.
api-generation/src/main/scala/ai/h2o/sparkling/api/generation/common · high confidence
New download page and build metadata for Sparkling Water
The distribution now includes a dedicated download page (index.html) and a structured build metadata file (buildinfo.json). The download page provides a user-friendly interface to access the Sparkling Water ZIP archive, along with links to developer documentation, Scala Scaladoc, and example templates. The buildinfo.json file exposes key versioning details—including the Sparkling Water, H2O, and Spark versions, build timestamp, and commit hash—enabling users and scripts to verify the exact composition of the distribution.
dist · high confidence
New feature estimators and utilities for Sparkling ML features
This change introduces several new components to the Sparkling ML features package. It adds a ColumnPruner transformer to selectively keep or drop columns from datasets. It also introduces base classes for new H2O-powered feature estimators: H2OAutoEncoderBase for autoencoder models, H2ODimReductionEstimator as a base for dimensionality reduction, H2OGLRMBase for Generalized Low Rank Models, and H2OWord2VecBase for word embeddings. Additionally, the H2OTargetEncoder is updated to support multinomial classification problems and regression tasks, and exposes parameters for setting output columns and handling interactions.
ml/src/main/scala/ai/h2o/sparkling/ml/features · high confidence
New macro for deprecating API methods with warnings
The macros module now includes a new \@DeprecatedMethod\ annotation and its corresponding macro implementation. When applied to methods, this annotation automatically injects a standard Scala \@deprecated\ annotation and logs a warning message at runtime, informing users of the deprecation, the recommended replacement, and the version in which the method will be removed. A test suite has been added to verify the macro's behavior with various annotation usages and method signatures.
macros · high confidence
New micro-benchmark suite for DataFrame-to-H2OFrame conversion performance
A new set of micro-benchmarks has been added to the core benchmarking infrastructure to measure the performance of converting Spark DataFrames to H2OFrames. The suite includes specific tests for flat DataFrames, DataFrames with nested structs, DataFrames with flat arrays, and DataFrames containing wide sparse or dense vectors. It also benchmarks the performance of flattening DataFrames and schemas, as well as converting rows to row schemas, providing detailed statistical results (mean, standard deviation, min, max) to help track regression or improvement in these data conversion paths.
core/src/bench · high confidence
New scoring package with MOJO model classes and metrics
The scoring module now includes a complete set of new Scala classes for handling H2O MOJO models, including H2OMOJOModel, H2OMOJOPipelineModel, and algorithm-specific variants like H2OAlgorithmMOJOModel. This introduces support for detailed predictions, feature contributions, leaf node assignments, and stage results. Additionally, a new H2OMetric enum and H2OMetrics trait system are added to expose training, validation, and cross-validation metrics (such as AUC, RMSE, and Logloss) as structured objects, replacing the previous map-based approach.
scoring · high confidence
New utilities for analyzing and flattening nested Spark DataFrames
Added \DatasetShape\ to classify Spark schemas as Flat, StructsOnly, or Nested, and \SchemaUtils\ to flatten nested DataFrames (containing structs, arrays, maps, or binary data) into a single-level schema. This enables downstream components to handle complex nested data structures uniformly by converting them into a flat representation suitable for standard processing.
utils/src/main/scala/ai/h2o/sparkling/ml · high confidence
New utility classes for serialization, compression, and Spark session management
Added a suite of new utility classes in the utils package to support data handling and interoperability. This includes Base64Encoding for converting byte, int, and long arrays to and from strings; Compression utilities supporting NONE, DEFLATE, GZIP, and SNAPPY formats; and a comprehensive DataFrame serialization system (DataFrameJsonSerialization, JSONDataFrameSerializer, and wrappers) that allows DataFrames to be serialized to JSON and deserialized back, enabling compatibility with Java serialization wrappers. Additionally, CompatibilityObjectInputStream is introduced to handle serialVersionUID mismatches during deserialization, FinalizingOutputStream ensures cleanup actions upon stream closure, and SparkSessionUtils provides helpers for managing active Spark sessions and reading HDFS files.
utils/src/main/scala/ai/h2o/sparkling/utils · high confidence
PySparkling Python API restructured with expanded algorithm and MOJO support
The PySparkling package has been reorganized into a single source directory, exposing a comprehensive set of H2O algorithms and model types in the Python API. This update adds support for new algorithms including PCA, GLRM, AutoEncoder, CoxPH, GAM, RuleFit, StackedEnsemble, and Extended Isolation Forest, alongside existing ones like DRF and Word2Vec. It also introduces algorithm-specific MOJO model classes (e.g., H2OAlgorithmMOJOModel, H2OSupervisedMOJOModel) and separates classification and regression variants for easier access. Users can now import these new estimators and models directly from pysparkling.ml.algos and pysparkling.ml.models.
py/src/pysparkling · high confidence
Architecture
Refactored ML parameter definitions into modular traits
The parameter definitions in the Sparkling ML params package have been reorganized into a modular trait-based architecture. Common parameters are now shared via base traits like H2OCommonParams and H2OAlgorithmCommonParams, while algorithm-specific options (such as monotone constraints, beta constraints, and calibration frames) are isolated in dedicated traits (e.g., HasMonotoneConstraints, HasBetaConstraints). This change improves code maintainability and reduces duplication across different model algorithms.
ml/src/main/scala/ai/h2o/sparkling/ml/params · high confidence
Refactored algorithm training infrastructure with new base classes and preparation traits
The algorithm wrappers in the \ml/src/main/scala/ai/h2o/sparkling/ml/algos\ package have been restructured to use a new inheritance hierarchy and preparation logic. New base classes such as \H2OAlgorithm\, \H2OSupervisedAlgorithm\, and \H2OUnsupervisedAlgorithm\ now serve as the foundation for specific models, while \H2OAlgoCommonUtils\ centralizes dataset preparation, column handling, and frame lifecycle management. New traits like \H2OTrainFramePreparation\, \DistributionBasedH2OTrainFramePreparation\, and \FamilyBasedH2OTrainFramePreparation\ allow algorithms to automatically convert label columns to categorical types based on their distribution or family settings. This change also introduces specific implementations for \H2OAutoML\, \H2OGridSearch\, and \H2OStackedEnsemble\ that leverage this new common infrastructure for training and model management.
ml/src/main/scala/ai/h2o/sparkling/ml/algos · high confidence
Behavioural changes
Add Spark 3.2 compatibility layer for SQL encoders and UDFs
This change introduces a new Scala source directory for Spark 3.2 that provides version-specific adapters for SQL functionality. It adds a RowEncoder wrapper to handle schema-based row encoding, a functions object to expose UserDefinedFunction creation using the Spark 3.2-specific SparkUserDefinedFunction constructor, and a SparkUserDefinedFunction facade that bridges the internal expressions API. These components ensure that Sparkling Water's SQL features, particularly custom UDFs and row encoding, function correctly on Spark 3.2 by abstracting away version-specific API differences.
_utils/src/main/scala\_spark\3.2 · high confidence
Added Spark 3.1 and 3.5 compatibility shims
New adapter files were added for Spark 3.1 and 3.5 to handle API differences in row encoding and user-defined functions. For Spark 3.1, the RowEncoder shim delegates directly to the Spark API, while for 3.5 it uses the new encoderFor method to accommodate the removal of the apply method. Additionally, a SparkUserDefinedFunction shim was introduced to abstract away constructor signature changes, ensuring consistent UDF creation across both versions.
_utils/src/main/scala\_spark\3.1 · high confidence
Added Spark 3.3-specific SQL compatibility shims
This change introduces version-specific adapter files for Spark 3.3 to maintain compatibility with the SQL API. It adds a \RowEncoder\ wrapper that delegates to the Spark 3.3 catalyst encoder, a \functions\ object that provides a \udf\ helper using the \SparkUserDefinedFunction\ signature required by this version, and a \SparkUserDefinedFunction\ companion object that bridges the internal \expressions\ package with the public API. These shims ensure that user-defined functions and row encoding behave consistently across supported Spark versions.
_utils/src/main/scala\_spark\3.3 · high confidence
Added Spark 3.4 compatibility layer for SQL encoders and UDFs
This change introduces a new source directory for Spark 3.4 that provides compatibility shims for SQL functionality. Specifically, it adds a \RowEncoder\ wrapper to handle schema encoding, a \functions\ object to expose User Defined Functions (UDFs) using the \SparkUserDefinedFunction\ API, and a \SparkUserDefinedFunction\ companion object to bridge internal Spark expression types. These additions allow the application to maintain a consistent API surface across different Spark versions by abstracting away version-specific implementation details in the SQL module.
_utils/src/main/scala\_spark\3.4 · high confidence
Added default logging configuration for examples
A new log4j.properties file has been added to the examples resources, configuring the default logging level to WARN for the console output. This setup suppresses verbose logs from third-party libraries like Jetty and sets specific INFO levels for Spark REPL components, providing a cleaner log experience when running demos from IDEs like IntelliJ.
examples/src/main/resources · high confidence
CRAN rsparkling package now redirects users to the custom repository
The rsparkling package available on CRAN has been updated to a dummy version that no longer provides functional machine learning capabilities. Instead, attempting to use the package now triggers an error message directing users to install the maintained version from the H2O custom repository, with specific installation commands provided for Spark versions 2.1 through 2.4. This change ensures users are guided toward the actively supported release rather than the deprecated CRAN distribution.
r-cran · high confidence
Compatibility layer for Spark UDFs and RowEncoders
Added version-specific compatibility shims in the Sparkling Water utilities to handle API differences between Spark versions. The new \RowEncoder\ and \functions\ objects (along with \SparkUserDefinedFunction\ for Spark 3.0) abstract away changes in how User-Defined Functions and Row Encoders are constructed, ensuring that code using these features works consistently across supported Spark versions without requiring conditional logic in the application layer.
_utils/src/main/scala\_spark\_3.0, utils/src/main/scala\_spark\others · high confidence
Enforced H2O version verification and bundled Sparkling Water JAR in RSparkling
The RSparkling package now includes a new \package.R\ file that enforces strict compatibility by verifying the installed H2O R package version against the specific Sparkling Water version at load time. If a mismatch is detected, users receive clear instructions to uninstall the current H2O package and install the correct version from the H2O release repository. Additionally, the package now bundles the Sparkling Water assembly JAR directly, removing the need for users to manually specify the JAR path or version, thereby simplifying setup and ensuring the R package version is tightly coupled with the correct Sparkling Water backend.
r/src/R · high confidence
GBM pipeline example now includes detailed predictions by default
The H2OGBM pipeline example in the build artifacts has been updated to enable the 'detailed\_prediction' column in MOJO outputs by default. This change is reflected in the stage metadata, which now sets the 'detailedPredictionCol' parameter to 'detailed\_prediction' and enables 'namedMojoOutputColumns', ensuring that users running this specific pipeline example receive granular prediction details without needing to manually configure these parameters.
py/examples/build · high confidence
Internal H2OModel class now exposes cross-validation MOJO models
The new H2OModel internal class in the Sparkling ML internals package now supports converting models to MOJO format while including cross-validation models. When the withCVModels flag is true, the toMOJOModel method retrieves cross-validation model details from the H2O cluster API and sets them on the resulting H2OMOJOModel instance, making cross-validation MOJO models accessible to users who request them.
ml/src/main/scala/ai/h2o/sparkling/ml/internals · high confidence
Migrate CI infrastructure to Jenkins and AWS
The continuous integration system has been migrated from the previous platform to Jenkins, with the testing infrastructure now hosted on AWS. This change introduces a suite of new Jenkins pipelines (including main, benchmarks, Databricks, Kubernetes, nightly, and release workflows) and replaces the old infrastructure provisioning with Terraform and Packer scripts located in ci/aws/terraform. The new setup manages AWS resources, Docker registries, and credential handling through a shared Groovy library, enabling automated testing on Kubernetes and Databricks alongside standard Spark builds.
ci · high confidence
Migrate test infrastructure to AWS
The test infrastructure has been moved to AWS, introducing Terraform modules to provision the network, an ECR repository for Docker images, and a Jenkins instance. This change updates the CI environment to use AWS resources, including a specific Docker registry and Jenkins URL, replacing the previous setup.
ci/aws/terraform · high confidence
New API generation runners for algorithms, configurations, and MOJO models
The api-generation module now includes three new entry-point runners: AlgorithmAPIRunner generates algorithm-specific classes (including Word2Vec and problem-specific classifiers/regressors) for Scala and Python; ConfigurationRunner generates configuration bindings for Scala, Python, and R by reflecting on backend configuration classes; and MOJOModelAPIRunner generates MOJO model classes, factories, and metrics objects for Scala, Python, and R. These runners replace or supplement previous manual generation steps, ensuring that algorithm, configuration, and MOJO APIs are consistently generated across languages.
api-generation/src/main/scala/ai/h2o/sparkling/api/generation · medium confidence
New AWS Packer configuration for Java 11 Jenkins slaves
Added a new Packer template and initialization script to provision AWS EC2 instances for Jenkins slaves. The configuration builds an Amazon Linux AMI with Java 11 (OpenJDK 11), Docker, and Git, creating a dedicated 'jenkins' user with Docker group access to support the CI infrastructure migration.
ci/aws/packer · high confidence
New Scala editor styling in Flow CSS
A new CSS file (scala-editor.css) has been added to the core/flow-css directory, defining syntax highlighting and layout styles for the Scala code editor. The styles adopt an IntelliJ IDEA default theme, applying specific colors to keywords, numbers, strings, and comments, while also setting the font family to a monospace stack and configuring the editor container's dimensions, padding, and a distinctive left border.
core/flow-css · high confidence
New utility traits for REST-based model training and parameter reading
The \ml/src/main/scala/ai/h2o/sparkling/ml/utils\ package now includes \EstimatorCommonUtils\ and \H2OParamsReadable\. \EstimatorCommonUtils\ enables training models via the REST API by sending parameters to a cluster endpoint, waiting for job completion with progress printing, and handling warnings from model builders; it also provides utilities for downloading binary models and resolving model ID conflicts. \H2OParamsReadable\ introduces an \MLReadable\ implementation for reading H2O parameters, supporting the standard Spark ML parameter serialization pattern.
ml/src/main/scala/ai/h2o/sparkling/ml/utils · high confidence
R API generation now supports overloaded configuration setters and exposes model metrics as R objects
The R API generation templates have been updated to automatically generate R configuration classes from Scala code, including proper handling of overloaded setter methods in configuration objects. Additionally, the generation logic now exposes model metrics as distinct R objects via a new factory and template system, allowing users to access detailed metric data from H2OMOJOModel instances in R.
api-generation/src/main/scala/ai/h2o/sparkling/api/generation/r · high confidence
RSparkling API modernization and client-less default
RSparkling now uses a client-less approach by default, simplifying initialization via a parameterless H2OContext.getOrCreate() and removing the need to pass SparkContext explicitly. The API exposes H2OConf getters and setters, adds H2OContext.asSparkFrame() (deprecating asDataFrame), and enables REST API conversions with compression support.
r/src/R/ai/h2o/sparkling · high confidence
Repl interpreter refactored to support parallel sessions and Spark 2.12
The Scala REPL interpreter in the \repl/src/main\ module has been restructured to support multiple concurrent interpreter sessions and compatibility with Spark 2.12. This change introduces a new \H2OIMain\ class that wraps the Scala \IMain\ and uses session-specific package naming (e.g., \intp\id\\<sessionId\>\) to isolate classes defined in each session. A custom \InterpreterClassLoader\ routes class loading to the correct session-specific interpreter, allowing parallel execution. The refactoring also includes a runtime patch (\PatchUtils\) to fix \OuterScopes\ regex matching for these session-specific classes, and consolidates interpreter logic into \BaseH2OInterpreter\ and \H2OInterpreter\ to remove version-specific duplication.
repl/src/main · high confidence
Restructure Python package build configuration
The Python package build system has been reorganized to ensure correct packaging and version handling. A new MANIFEST.in file explicitly includes version and build info files (version.txt, buildinfo.txt) for both h2o and sparkling water components, ensuring they are included in source distributions. The setup.py and setup.cfg files have been updated to define the package metadata, dependencies (requests, tabulate), and package data, facilitating proper installation via pip.
py/src · high confidence
Unify and modernize Sparkling Water launcher scripts and environment configuration
The bin directory has been refactored to provide a consistent, unified set of launcher scripts for Sparkling Water across Linux/macOS and Windows. The new scripts (such as run-sparkling.sh, run-sparkling.cmd, sparkling-shell, and pysparkling) centralize environment validation by sourcing sparkling-env.sh (or sparkling-env.cmd), which now reads version and path details directly from gradle.properties. This change introduces a new default MASTER value of 'local\[\*\]' and sets a default driver memory of 2G. Additionally, the PySparkling launcher now explicitly checks for required Python packages (requests, tabulate) before starting, and the build-kubernetes-images.sh script has been updated to support building images for the external backend mode.
bin · high confidence
Updated RSparkling examples to target Spark 2.4.6
The included R demo scripts (NYC flights analysis and simple start) have been updated to explicitly install and connect to Spark version 2.4.6, replacing the previous 2.4.5 references. This ensures that users running these examples will provision the correct Spark runtime version compatible with the current release.
r/src/inst · high confidence
Updated documentation theme with Font Awesome 4.2.0 and new CSS assets
The Sphinx documentation theme has been updated to include Font Awesome 4.2.0, providing a broader set of icons for use in the documentation interface. This change introduces new CSS files (\badge\_only.css\ and \theme.css\) that define the styling for the theme, including the icon definitions and layout adjustments for the version switcher and navigation elements.
_doc/src/site/sphinx/sphinx\_rtd\theme · high confidence
Fixes
Added Iris dataset with headers for external testing
A new CSV file, \iris\_wheader.csv\, has been added to the \examples/smalldata/iris\ directory. This file contains the standard Iris dataset including a header row (sepal\_len, sepal\_wid, petal\_len, petal\_wid, class) and 150 data points across the three species (Iris-setosa, Iris-versicolor, Iris-virginica). This change ensures the data is available as a committed resource for external tests, addressing the issue where the data location was previously inaccessible.
examples/smalldata/iris · high confidence
Test coverage
1 commit adding/updating tests in ml/src/test/resources/target\_encoder; Add R Kubernetes integration test script; Add test data for prediction interval validation; Added API test suites for DataFrames, H2OFrames, RDDs, and Scala Interpreter; Added Kubernetes smoke tests for cluster initialization and data workflows; Added Python integration test infrastructure and baseline examples; Added R tests for H2O configuration, frame conversions, and MOJO model features; Added Scala 2.11 and 2.12 test resources for H2OMOJOModel serialization; Added comprehensive test coverage for Spark-to-H2O data conversion; Added integration tests for SSL, LDAP, and PAM authentication modes; Added integration tests for complex schema flattening; Added integration tests for example applications; Added integration tests for external H2O backend scenarios; Added test data for H2O MOJO SHAP contributions; Added test data for Mojo pipeline with multiple output columns; Added test data for SHAPley contributions on transformed features; Added test resources for H2O MOJO pipeline validation; Added test resources for binary model predictions and test logging configuration; Added test suites for H2O feature estimators; Added test suites for ML algorithm prediction outputs and metrics; Added test suites for MOJO model behavior and binary model lifecycle; Added test suites for backend utility components; Added test suites for configuration, data sources, authentication, and frame operations; Added tests for EnumParamValidator; Added tests for H2O model metrics handling; Added tests for H2OContext lifecycle and import scenarios; Added tests for JSON serialization and schema flattening utilities; Added tests for PartitionStatsGenerator; Added tests for Sparkling Water ML pipeline prediction and streaming; Added unit tests for REPL interpreter settings and patch utilities; Initial Python scoring test suite for PySparkling; New unit test suite for PySparkling algorithms and data conversions; Test logging configuration added; Updated test resource model to use H2O MOJO pipeline.
Dependencies
Introduce modular Gradle build structure with dedicated assembly and utility subprojects
The build system has been restructured into a multi-module Gradle project, introducing new subprojects for API generation, extensions assembly, scoring assembly, and utilities. This change centralizes dependency management and artifact assembly, ensuring that the main Sparkling Water assembly jar correctly excludes transitive dependencies provided by Apache Spark and Hadoop, while explicitly relocating conflicting packages (such as Jetty and Guava) to the \ai.h2o\ namespace to prevent classpath collisions. It also adds dedicated build configurations for benchmarks and documentation generation.
(dependencies) · high confidence
Upgrade Gradle wrapper to version 7.6
The Gradle wrapper has been updated to use Gradle 7.6 (all distribution). This ensures that builds are executed with a consistent, modern version of the Gradle build tool, leveraging improvements and bug fixes present in the 7.x release line.
gradle/wrapper · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 58 → 62 (+3.4)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 94 → 94 (-0.0)
- Architecture 99 → 92 (-7.4)
- Maturity 55 → 55 (+0.0)
- Readiness 59 → 59 (-0.3)
- Security 50 → 62 (+12.5)
Resolved (2)
- Documentation: no installation or build instructions
- Documentation: no usage examples
New (5)
- Documentation: no project overview (README.rst)
- Duplicated block (6 lines × 2) (py-scoring/src/ai/h2o/sparkling/H2ODataFrameConverters.py)
- Duplicated block (8 lines × 2) (extensions/src/main/scala/ai/h2o/sparkling/extensions/internals/ConvertCategoricalToStringColumnsTask.java)
- No ADRs found
- Split core
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
h2oai/sparkling-water was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 1fc2265322f5ff1575890c75df384119ff55aa98 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.