Skip to content
CAI
Software that uses CAICheck a score

combust/mleap

58.2

Weak · 27 September 2026

28.2k

lines of production code

Scala

with Python

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This release delivers a comprehensive overhaul of the MLeap serialization and runtime architecture, introducing a new Bundle.ML format that supports JSON, Protobuf, and binary serialization for a wide range of machine learning models. The update significantly expands model support, adding transformers for classification, clustering, regression, and feature engineering, alongside new integrations for TensorFlow, XGBoost, and Scikit-learn. Additionally, the release introduces a robust model serving layer with gRPC and HTTP endpoints, while also upgrading the underlying build system to support Scala 2.13, Java 17, and Spark 4.1.2.

Features

Add ALS recommendation transformer

A new ALS (Alternating Least Squares) transformer has been added to the MLeap runtime, enabling recommendation capabilities. The implementation defines an ALS case class that wraps an ALSModel and exposes a user-defined function to compute recommendations for a given user and item.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/recommendation · high confidence

Add Avro serialization support for LeapFrame and Row data

The mleap-avro module now includes new implementations for serializing and deserializing MLeap data structures to and from the Avro format. This introduces DefaultFrameReader and DefaultFrameWriter for full LeapFrame serialization, as well as DefaultRowReader and DefaultRowWriter for individual row-level Avro operations. The change includes a SchemaConverter that maps MLeap data types (including scalars, lists, maps, and various tensor types) to Avro schemas, and a ValueConverter that handles the actual data conversion between MLeap and Avro representations.

mleap-avro/src/main · high confidence

Add CategoricalDrilldown transformer for ensemble models

A new CategoricalDrilldown transformer has been added to the runtime, enabling the application of multiple underlying transformers based on a categorical label. This allows users to route data through specific ensemble members, with output columns automatically prefixed with 'drilldown.' to distinguish the transformed fields.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/ensemble · high confidence

Add Databricks runtime testkit with Spark ML, TensorFlow, and XGBoost integration tests

A new testkit module for Databricks runtime has been added, providing integration tests for loading and saving ML models. The suite includes test cases for Spark ML pipelines (using LogisticRegression), TensorFlow models (add and multiply graph operations), and XGBoost classifiers, each verifying the ability to write and load MLeap bundles in a Databricks environment.

mleap-databricks-runtime-testkit/src · high confidence

Add HDFS Bundle File System support

Users can now store and retrieve ML bundles on HDFS. This change introduces a new Hadoop-based file system implementation (HadoopBundleFileSystem) that handles loading and saving bundles to HDFS, along with the corresponding configuration in reference.conf and unit tests to verify the new functionality.

bundle-hdfs/src · high confidence

Add JSON serialization support for Bundle.ML types

A new \JsonSupport\ trait is introduced in the \bundle-ml\ module, providing \JsonFormat\ implementations for core Bundle.ML types including \BasicType\, \ByteString\, \Tensor\, \Scalar\, and \DataShape\. This enables these types to be serialized to and deserialized from JSON, facilitating the storage and transfer of machine learning model bundles in a human-readable format.

bundle-ml/src/main/scala/ml/combust/bundle/json · high confidence

Add JSON serialization support for MLeap frames and rows

Users can now serialize and deserialize MLeap LeapFrames and Rows to and from JSON format. This change introduces new components in the mleap-runtime module (DefaultFrameReader, DefaultFrameWriter, DefaultRowReader, DefaultRowWriter, JsonSupport, and RowFormat) that enable converting MLeap data structures into human-readable JSON strings and back, supporting various data types including basic types, lists, maps, tensors, and byte strings.

mleap-runtime/src/main/scala/ml/combust/mleap/json · high confidence

Add JSON serialization support for decision and clustering trees

Users can now serialize and deserialize decision trees and clustering nodes to and from JSON format. This change introduces new serialization logic in the bundle-ml module, including JSON format writers and readers for both decision trees (internal/leaf nodes) and clustering nodes, as well as supporting type classes (NodeWrapper) to bridge the gap between internal model representations and the bundle format.

bundle-ml/src/main/scala/ml/combust/bundle/tree · high confidence

Add MLEP serialization for XGBoost classification and regression models

Added new MLEP operators for serializing and deserializing XGBoost classification and regression models. The new \XGBoostClassificationModelOp\ and \XGBoostRegressionModelOp\ classes handle the storage and loading of model parameters such as tree limits, thresholds, and missing values, enabling these Spark ML models to be exported and imported via MLEP bundles.

mleap-xgboost-spark/src/main/scala · high confidence

Add MLeap extensions for scikit-learn transformers

New MLeap extension classes are introduced for scikit-learn's SimpleImputer and a custom DefineEstimator wrapper, enabling these transformers to be used within MLeap pipelines and serialized to bundles. The Imputer class wraps scikit-learn's SimpleImputer to support feature extraction and serialization, while the DefineEstimator class facilitates running transformers on specific columns of input data.

python/mleap/sklearn/extensions · high confidence

Add MLeap serialization support for Gensim Word2Vec models

Users can now serialize Gensim Word2Vec models to the MLeap bundle format. This change introduces a new module at python/mleap/gensim/word2vec.py that patches the gensim.models.Word2Vec class with MLeap-specific methods (serialize\_to\_bundle, sent2vec) and a SimpleSparkSerializer to handle the conversion of word vectors and metadata into MLeap-compatible bundles.

python/mleap/gensim · high confidence

Add MLeap serialization support for scikit-learn preprocessing transformers

The MLeap Python library now supports serializing and deserializing several scikit-learn preprocessing transformers (StandardScaler, MinMaxScaler, OneHotEncoder, SimpleImputer, Binarizer, and PolynomialFeatures) into MLeap bundles. This enables users to export models trained with these scikit-learn components for execution in MLeap or Spark environments, maintaining parity with the MLeap runtime's expected input/output feature structures.

python/mleap/sklearn/preprocessing · high confidence

Add MLeap serving module for model serving

A new mleap-serving module is introduced, providing a unified entry point to start both gRPC and HTTP servers backed by a single MLeap executor. The module exposes configuration for the gRPC port and coordinates the startup of the gRPC server (mleap-grpc-server) and the HTTP server (mleap-spring-boot), ensuring that models loaded through one interface are available through the other.

mleap-serving · high confidence

Add MLeap-TensorFlow converter implementation for tensor type mapping

The MLeap-TensorFlow integration now includes the core conversion logic to map between MLeap tensor types and TensorFlow tensors. The new \MleapConverter\ and \TensorflowConverter\ classes handle bidirectional conversion for numeric types (int, long, float, double), strings, and byte strings, enabling the TensorFlow backend to serialize and deserialize model data using the updated TensorFlow Java API.

mleap-tensorflow/src/main/scala/ml/combust/mleap/tensorflow/converter · medium confidence

Add S3-backed repository support for MLeap executor

Users can now store and download MLeap bundles directly from Amazon S3 using s3:// URIs. The new S3Repository implementation allows the executor to fetch bundles from S3 buckets, with configuration provided via the standard AWS credential provider chain. This adds S3 as a supported repository type alongside existing file and HTTP repositories.

mleap-repository-s3 · high confidence

Add Sklearn PolynomialFeatures transformer

Introduces a new PolynomialFeatures transformer for the Sklearn integration, enabling polynomial feature expansion within the MLeap runtime.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/sklearn · high confidence

Add SparkUtil utility for direct pipeline creation

A new SparkUtil object has been added to the Spark ML MLeap module, providing a convenience method to create a PipelineModel directly from an array of Transformers. This simplifies pipeline construction by allowing users to instantiate a PipelineModel without explicitly creating a Pipeline first.

mleap-spark-base/src/main/scala/org/apache/spark/ml/mleap · high confidence

Add TensorFlow model loading and transformation support

Introduces new Scala classes (TensorflowModel, TensorflowTransformer, and TensorflowTransformerOp) that enable MLeap to load and run TensorFlow models in both 'graph' and 'saved\_model' formats, allowing users to integrate TensorFlow-based transformations into their MLeap pipelines.

mleap-tensorflow/src/main/scala/ml/combust/mleap/tensorflow · high confidence

Add automatic casting of data types in LeapFrame and Spark DataFrame conversion

A new TypeConverters trait has been introduced to handle the conversion between Spark DataFrames and MLeap LeapFrames. This includes mapping Spark data types (such as Boolean, Byte, Short, Int, Long, Float, Double, String, and various Array types) to their corresponding MLeap types, and vice versa. The implementation also adds support for converting Vector and Matrix types to and from MLeap Tensor types, ensuring that data shapes and nullability are correctly preserved during the conversion process.

mleap-spark-base/src/main/scala/org/apache/spark/sql · high confidence

Add built-in Spark ML operation registry

The MLeap Spark integration now includes a comprehensive built-in registry of Spark ML operations, enabling automatic serialization and deserialization of a wide range of Spark ML models and transformers. This configuration file registers support for classification algorithms (e.g., Decision Tree, Naive Bayes, Random Forest), clustering (K-Means, GMM), feature transformers (e.g., Tokenizer, PCA, Vector Assembler), regression models, and tuning components (Cross-Validation, Train-Validation Split). This allows users to seamlessly export and import Spark ML pipelines and individual models using MLeap's format without manual registration.

mleap-spark/src/main/resources · high confidence

Add bundle serialization for new and existing feature transformers

Adds bundle serialization support for several feature transformers, including Imputer, MapEntrySelector, MathBinary, MathUnary, MultinomialLabeler, StringMap, and WordLengthFilter. This enables these Spark ML transformers to be serialized into MLeap bundles for deployment.

mleap-spark-extension/src/main/scala/org/apache/spark/ml/bundle/extension/ops/feature · high confidence

Add high-performance XGBoost Predictor runtime support

Introduces a new, high-performance implementation for XGBoost classification and regression models using the XGBoost Predictor library. This adds new runtime classes (e.g., XGBoostPredictorClassification, XGBoostPredictorRegression) and their corresponding bundle operations, enabling faster inference for MLeap users while maintaining compatibility with existing XGBoost model formats.

mleap-xgboost-runtime/src/main · high confidence

Add serialization support for OneVsRest and SVM classification models

Users can now serialize and deserialize OneVsRest and Support Vector Machine (SVM) classification models. The new OneVsRestOp and SupportVectorMachineOp implementations in the Spark bundle extension allow these models to be saved to and loaded from MLeap bundles, including handling of model parameters such as thresholds and class metadata.

mleap-spark-extension/src/main/scala/org/apache/spark/ml/bundle/extension/ops/classification · high confidence

Add serialization support for Scikit-learn models

The Python integration with Scikit-learn now supports serializing and deserializing several core model types, including LinearRegression, LogisticRegression, LogisticRegressionCV, FeatureUnion, Pipeline, SVC, and LinearSVC. These changes enable users to export trained Scikit-learn models into the MLeap bundle format for deployment.

python/mleap/sklearn · high confidence

Add serialization support for scikit-learn Decision Trees and Random Forests

Users can now serialize scikit-learn DecisionTreeClassifier, DecisionTreeRegressor, RandomForestClassifier, and RandomForestRegressor models to MLeap bundles. This change adds the necessary serialization logic in \python/mleap/sklearn/tree/tree.py\ and \python/mleap/sklearn/ensemble/forest.py\, enabling these specific model types to be exported and stored in the MLeap format.

python/mleap/sklearn/tree · high confidence

Added MLeap and Spark transform benchmarks

A new benchmarking tool has been introduced in the mleap-benchmark module, providing performance measurement for MLeap and Spark model transformations. The addition includes a Boot entry point for CLI execution, a shared Benchmark trait, and specific benchmark classes for MLeap row/transform operations and Spark transforms, enabling users to evaluate execution performance across different runtime environments.

mleap-benchmark/src/main/scala · high confidence

Added MLeap gRPC client module for remote model serving

Users can now interact with a remote MLeap gRPC server to load, unload, transform, and stream data against hosted models. This new \mleap-grpc\ module provides a \GrpcClient\ implementation of the \TransformService\ trait, handling the translation between MLeap types and their Protobuf representations via Akka stream helpers. This enables client-side integration with the \mleap-grpc-server\ for remote inference workflows.

mleap-grpc · high confidence

Added Travis CI deployment and release scripts

Added new shell scripts (docker.sh, extract.sh, publish.sh, release.sh, travis\_publish.sh) to automate the release and publishing process on Travis CI. The release script ensures builds only run on the master branch for pull requests, configures git credentials, and triggers SBT publishing tasks for MLeap and Spring Boot modules.

travis · high confidence

Added XGBoost model operation registrations for Spark

A new configuration file (reference.conf) was added to register the XGBoost classification and regression model operations with the MLeap Spark registry, enabling serialization and deserialization of these models within the Spark integration.

mleap-xgboost-spark/src/main/resources · high confidence

Added benchmark configuration and test data for MLeap

The MLeap benchmark module now includes its own configuration files and test data. A new JSON file (frame.airbnb.json) provides sample data for the LeapFrame, while mleap.conf and spark.conf define benchmark parameters such as iteration ranges and run counts. Additionally, a log4j.properties file is added to control logging verbosity for the benchmarking process.

mleap-benchmark/src/main/resources · medium confidence

Added clustering transformer implementations

New transformer classes have been added to the runtime to support specific clustering algorithms. The diff introduces Scala implementations for K-Means, Gaussian Mixture Models (GMM), Latent Dirichlet Allocation (LDA), and Bisecting K-Means. Each class wraps its corresponding model (e.g., KMeansModel, GaussianMixtureModel) to execute predictions or topic distributions on input tensors, enabling these clustering models to be executed within the MLeap runtime.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/clustering · high confidence

Added sample datasets and parity test infrastructure to the Spark testkit

The mleap-spark-testkit module now includes sample data files (a text file for NLP tasks, a ratings file for recommendation systems, and a libSVM file for classification) and a new Scala base class, SparkParityBase. This base class provides the infrastructure for running Spark/MLeap parity tests, including helper methods to load the new datasets, serialize models to bundles, and compare Spark and MLeap outputs with relative tolerance.

mleap-spark-testkit · high confidence

Added script to compile and export Scala classpath for Python integration

A new shell script, scripts/scala\_classpath\_for\_python.sh, has been added to automate the compilation of MLeap Spark extension classes and export the resulting Scala classpath. This script ensures that the necessary Scala classes are compiled and available in the environment for Python-based tests and integrations, streamlining the setup process for developers working with the Python bindings.

scripts · high confidence

Added serialization support for scikit-learn text vectorizers

Users can now serialize scikit-learn's CountVectorizer and TfidfVectorizer models to MLeap bundles. This change adds the necessary serialization logic in python/mleap/sklearn/feature\_extraction/text.py, enabling these text processing components to be exported and reused within the MLeap ecosystem.

_python/mleap/sklearn/feature\extraction · high confidence

Added support for serializing and deserializing tensors to and from JSON

Users can now serialize and deserialize DenseTensor, SparseTensor, and general Tensor objects directly to and from JSON format. This is enabled by new Scala source files (ByteString, JsonSupport, and Tensor) that define the JSON structure for tensor data, allowing for easier interchange and storage of tensor data in JSON-compatible formats.

mleap-tensor/src/main · high confidence

Adds core model implementations and neural network layer definitions

Introduces new model classes for classification (GBT, LinearSVC, MultiLayerPerceptron, Naive Bayes), clustering (BisectingKMeans, GaussianMixture, KMeans, LDA), and feature engineering (Binarizer, Bucketizer, ChiSqSelector, Coalesce, CountVectorizer). Additionally, adds the \ann\ package containing neural network layer definitions (AffineLayer, LossFunctions) and utility classes (BreezeUtil, Layer, Model) to support deep learning operations outside of a Spark context.

mleap-core/src/main · high confidence

Classification transformers now expose raw predictions and probabilities

The classification transformers in the MLeap runtime (including DecisionTree, LogisticRegression, RandomForest, SupportVectorMachine, and others) have been updated to output raw predictions and probability scores alongside standard predictions. This change allows users to access intermediate model outputs (rawPrediction, probability) in addition to the final prediction, enabling more flexible downstream processing and evaluation.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/classification · high confidence

Expanded MLeap runtime supports a wider range of ML models and transformers

The MLeap runtime now includes a comprehensive set of built-in operations for classification, clustering, regression, and feature transformation. This update enables the serialization and deserialization of a broader range of machine learning models and data processing steps, including classifiers like Decision Tree, GBT, and Logistic Regression; clustering models such as K-Means and Gaussian Mixture; regression models including Linear and Generalized Linear Regression; and various feature transformers like StringIndexer, VectorAssembler, and Tokenizer. Users can now deploy and serve these specific model types directly via MLeap.

mleap-runtime/src/main/resources · high confidence

Introduce BundleContext and BundleFileSystem for extensible bundle storage

The bundle library now supports pluggable file systems for reading and writing Bundle.ML models. A new BundleContext class holds serialization format, registry, and file system information, while the BundleRegistry tracks registered file systems by URI scheme. This allows models to be loaded from and saved to various storage backends (e.g., local file, JAR, or custom URIs) through a unified interface.

bundle-ml/src/main/scala/ml/combust/bundle · high confidence

Introduce MLeap executor with configurable model serving and streaming

The MLeap executor now provides a new local model serving and transformation service, exposing APIs for loading, unloading, and transforming models via both synchronous calls and Akka-based streams. The executor supports configurable timeouts for memory and disk caching, stream and flow parallelism, buffer sizes, and throttling. It also introduces a multi-repository system that can load models from local files or HTTP sources, with configurable thread pools and idle timeouts.

mleap-executor/src · high confidence

Introduce MLeap gRPC server for model inference

A new gRPC server implementation has been added to MLeap, enabling model inference via the gRPC protocol. The server exposes endpoints for loading, unloading, and transforming frames and rows, supporting both unary and streaming requests. It includes a configurable port (defaulting to 65328, overridable via the MLEAP\_GRPC\_PORT environment variable) and an error interceptor that maps internal exceptions like NotFoundException or TimeoutException to appropriate gRPC status codes.

mleap-grpc-server/src/main · high confidence

Introduce MLeap serialization and deserialization for Python transformers

Added the \python/mleap/bundle\ module, which provides \MLeapSerializer\ and \MLeapDeserializer\ classes to convert Python transformers (such as Scikit-learn estimators) into MLeap bundle formats (model.json and node.json). This enables serializing transformer attributes like coefficients, intercepts, and scalers into a structured JSON representation, facilitating interoperability with MLeap-based pipelines.

python/mleap/bundle · high confidence

Introduce Python package structure and tooling

The Python integration is now packaged as a proper Python library, enabling installation via pip and standard build tools. The package includes a \setup.py\ that enforces Python 3.10+ and lists dependencies such as scikit-learn 1.0, pandas, and scipy. Additionally, a \Makefile\ and \tox.ini\ are provided to streamline running tests and managing the development environment.

python · high confidence

Introduce binary serialization for LeapFrame and Row data

A new binary serialization format is added for serializing and deserializing LeapFrame and Row objects. This includes new readers and writers (DefaultFrameReader, DefaultFrameWriter, DefaultRowReader, DefaultRowWriter) and a ValueSerializer that handles primitive types, strings, ByteStrings, lists, and tensors, enabling more compact and efficient storage or transmission of model data compared to previous formats.

mleap-runtime/src/main/scala/ml/combust/mleap/binary · high confidence

Introduces new MLeapOp base classes and serialization ops for ML models

The MLeap runtime now includes a new abstract \MleapOp\ base class and a \MultiInOutMleapOp\ variant to standardize how model operations are serialized and deserialized. This change introduces serialization logic for a wide range of machine learning models, including classification (Decision Tree, GBT, Logistic Regression, Naive Bayes, Random Forest, SVM), clustering (K-Means, Gaussian Mixture, LDA), and feature transformers (Binarizer, Bucketizer, Count Vectorizer, etc.). Users will see these models supported in the MLeap bundle format, enabling consistent storage and loading of these algorithms.

mleap-runtime/src/main/scala/ml/combust/mleap/bundle · high confidence

Major overhaul of feature transformers to use a new NodeShape-based API

The feature transformer implementations in the runtime have been refactored to use a new \NodeShape\-based API, replacing the previous \TransformBuilder\-based \build\ methods with \SimpleTransformer\ or \Transformer\ implementations that define an \exec\ function. This change standardizes how transformers like \StringIndexer\, \OneHotEncoder\, and \StandardScaler\ handle inputs and outputs, often simplifying the code and aligning with a more consistent internal interface. Additionally, several new feature transformers have been added, including \Binarizer\, \Coalesce\, \CountVectorizer\, \DCT\, \ElementwiseProduct\, \FeatureHasher\, \IDF\, \Imputer\, \Interaction\, \MapEntrySelector\, \MathBinary\, \MathUnary\, \MaxAbsScaler\, \MinHashLSH\, \MinMaxScaler\, \MultinomialLabeler\, \NGram\, \Normalizer\, \Pca\, \PolynomialExpansion\, \RegexIndexer\, \RegexTokenizer\, \StopWordsRemover\, \StringMap\, \VectorIndexer\, \VectorSlicer\, \WordLengthFilter\, and \WordToVector\. The \HashingTermFrequency\ transformer was also updated to use the new API.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/feature · high confidence

New Bundle.ML serialization framework with JSON and Protobuf support

The Bundle.ML serialization system has been refactored to support both JSON and Protobuf formats. A new \BundleSerializer\ handles high-level bundle read/write operations, while dedicated serializers (\ModelSerializer\, \NodeSerializer\, \GraphSerializer\) manage the underlying data. The \SerializationFormat\ trait and its \Json\/\Protobuf\ implementations allow the system to switch between serialization backends. Additionally, a deprecated \FileUtil\ class is provided for backward compatibility, delegating to the new \ml.combust.bundle.util.FileUtil\.

bundle-ml/src/main/scala/ml/combust/bundle/serializer · high confidence

New Java DSL for MLeap runtime operations

A new Java DSL is introduced in the mleap-runtime module, providing convenient builder and support classes for Java users. This includes BundleBuilder and BundleBuilderSupport for loading and saving MLeap bundles, ContextBuilder for creating MleapContext and loading registries, LeapFrameBuilder and LeapFrameBuilderSupport for constructing LeapFrames, schemas, and rows, as well as support classes for RowTransformer, Tensor, and LeapFrame operations. These additions streamline the creation and manipulation of MLeap runtime objects from Java code.

mleap-runtime/src/main/java · high confidence

New PySpark feature transformers: MathBinary, MathUnary, and StringMap

Added new PySpark wrappers for MLeap feature transformers: MathBinary (supporting Add, Subtract, Multiply, Divide, Remainder, LogN, Pow, Min, Max), MathUnary (supporting Sin, Cos, Tan, Log, Exp, Abs, Sqrt, Logit), and StringMap (with configurable invalid handling). These allow users to perform mathematical operations on columns and map string labels to numeric values directly within PySpark pipelines.

python/mleap/pyspark/feature · high confidence

New Spark bundle serialization infrastructure

The MLeap Spark integration now includes a new set of core classes for handling Spark ML model serialization. This includes the SparkBundleContext to manage datasets and registries, SimpleSparkOp and MultiInOutSparkOp to handle single and multi-column transformer serialization, and SparkShape/ParamSpec to map Spark DataFrame schemas to bundle node shapes. This infrastructure enables the framework to serialize and deserialize Spark ML pipelines and transformers into MLeap bundles.

mleap-spark-base/src/main/scala/org/apache/spark/ml/bundle · high confidence

New Spark integration layer for MLeap serialization

Added three new Scala files (SimpleSparkSerializer, SparkLeapFrame, SparkSupport) that provide a bridge between Spark DataFrames and MLeap's internal LeapFrame representation. This enables users to serialize Spark ML transformers to MLeap bundles and deserialize them back, with automatic type conversion between Spark and MLeap schemas.

mleap-spark-base/src/main/scala/ml · high confidence

New feature transformers: Imputer, MapEntrySelector, MathBinary, MathUnary, MultinomialLabeler, StringMap, and WordLengthFilter

Added several new Spark ML transformers to the mleap-spark-extension module. These include Imputer for filling missing values using mean or median, MapEntrySelector for extracting map entries, MathBinary and MathUnary for performing binary and unary mathematical operations, MultinomialLabeler for converting probability vectors to labels, StringMap for mapping string labels to numeric values, and WordLengthFilter for filtering words by length. Each transformer implements MLWritable/MLReadable for serialization and provides standard Spark ML API interfaces.

mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/feature · high confidence

Serialization support for Spark ML transformers and classifiers

Adds serialization and deserialization support for a broad set of Spark MLlib transformers and classifiers, including GBTClassifier, MultiLayerPerceptronClassifier, NaiveBayes, BisectingKMeans, GaussianMixture, KMeans, LDA, Binarizer, BucketedRandomProjectionLSH, Bucketizer, ChiSqSelector, CountVectorizer, DCT, ElementwiseProduct, FeatureHasher, IDF, Interaction, MaxAbsScaler, MinHashLSH, MinMaxScaler, NGram, and Normalizer. This enables users to save and load these models in the MLeap bundle format.

mleap-spark/src/main/scala · high confidence

Removals

Removed Bundle.ML DSL and serialization support files

The \bundle-ml\ module's domain-specific language (DSL) and associated serialization infrastructure have been removed. This includes the deletion of core DSL classes such as \Attribute\, \AttributeList\, \Bundle\, \Model\, \Node\, \Shape\, and \Value\, as well as the \JsonSupport\ for serializing Bundle.ML objects to JSON. This change eliminates the ability to construct, serialize, and deserialize Bundle.ML pipelines and graphs using the previous Scala-based DSL and JSON-based serialization format.

bundle-ml/src/main/scala/ml/bundle/dsl · high confidence

Removed legacy MLeap serialization and registry classes

The \MleapBundle\ and \MleapRegistry\ objects in the \mleap-runtime\ module have been removed. These classes previously handled the serialization and deserialization of MLeap bundle files and maintained a static registry of machine learning operators. Their removal indicates a shift in how the runtime manages model serialization and operator registration, likely consolidating these responsibilities into a newer, unified registry or bundle format.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization/bundle · high confidence

Behavioural changes

Add default Spark extension operations to MLeap configuration

A new reference configuration file registers default Spark extension operations, including OneVsRest, SupportVectorMachine, Imputer, MapEntrySelector, MathBinary/Unary, MultinomialLabeler, WordLengthFilter, and StringMap. This change ensures these Spark ML components are automatically available for serialization and deserialization without requiring manual configuration.

mleap-spark-extension/src/main/resources · high confidence

Added file utility for safe ZIP extraction and directory removal

A new FileUtil utility class was added to provide safe file system operations, including a recursive directory removal method and a ZIP extraction method that validates entry paths to prevent Zip Slip vulnerabilities.

bundle-ml/src/main/scala/ml/combust/bundle/util · medium confidence

Adds PySpark integration utilities and module registration

The PySpark wrapper now includes dedicated modules for serializing and deserializing MLeap bundles directly from Spark DataFrames, exposing \serializeToBundle\ and \deserializeFromBundle\ methods on Spark Transformers. Additionally, the package registers MLeap features (StringMap, MathBinary, MathUnary) under the \pyspark.ml.mleap\ namespace to simplify imports and improve classpath injection for Python tests.

python/mleap/pyspark · high confidence

MLeap Python package version updated to 0.25.2

The MLeap Python package now exposes its version number via the public API, allowing users to programmatically check the installed version. The package version has been updated to 0.25.2.

python/mleap · high confidence

MLeap Spring Boot overhaul: new REST API and model loading

The MLeap Spring Boot module has been overhauled to provide a complete REST API for model management and scoring. New controllers (JsonScoringController, ProtobufScoringController, LeapFrameScoringController) expose endpoints for loading, unloading, and transforming data using both JSON and Protobuf formats. A new GlobalExceptionHandler standardizes error responses across the application, and a ModelLoader component enables automatic loading of models from a configured directory at startup.

mleap-spring-boot/src/main/scala · medium confidence

Migrated type system to support Bundle protocol conversions

The internal type system was refactored to support conversion between MLeap's internal data types and the external Bundle protocol. The previous \DataType\ definitions (including \TensorType\ and \ListType\) were removed and replaced with a new \BundleTypeConverters\ trait that maps basic types, shapes, and schemas between the two systems. This enables interoperability with the Bundle format, supporting scalar, list, and tensor data shapes with nullable support.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/types · high confidence

New binary serialization for tensors and array types

Added new \ArraySerializer\ implementations and a \TensorSerializer\ that convert tensors into a compact, binary format for storage and transmission. This replaces previous serialization methods, ensuring that tensor data (including booleans, bytes, shorts, ints, longs, floats, doubles, strings, and byte strings) is serialized efficiently using \ByteBuffer\ and \DataOutputStream\/\DataInputStream\ mechanisms.

bundle-ml/src/main/scala/ml/combust/bundle/tensor · high confidence

New param traits for labels and probabilities columns

The Spark ML extension now includes new parameter traits, HasLabelsCol and HasProbabilitiesCol, which allow transformers to specify output columns for labels and probabilities respectively. Additionally, the existing HasDropLast trait was moved from the core mleap-spark module to the extension module.

mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/param · high confidence

Pipeline transformer refactored to support async execution and schema introspection

The Pipeline transformer has been refactored to implement the new FrameTransformer interface, enabling asynchronous transformation via a new transformAsync method. The implementation now exposes input, output, and intermediate schemas through dedicated properties, allowing users to inspect the data structure at each stage of the pipeline. Additionally, the Pipeline class now manages the lifecycle of its constituent transformers by implementing a close method to ensure proper resource cleanup.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer · medium confidence

Refactor MLeap runtime API to use new frame and context abstractions

The MLeap runtime has been refactored to use a new \frame\ package for core data structures, replacing the previous \Dataset\, \LeapFrame\, and \Row\ implementations with \DefaultLeapFrame\, \Row\, and \ArrayRow\ in the \ml.combust.mleap.runtime.frame\ package. \MleapContext\ now manages the \BundleRegistry\ and class loading, while \MleapSupport\ provides updated implicit conversions for serializing/deserializing transformers and converting case classes to/from LeapFrames. The old \TransformBuilder\ and \LeapFrameBuilder\ are removed in favor of the new \FrameWriter\ and \RowReader\/\RowWriter\ abstractions.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime · high confidence

Refactor OneVsRest and SVM classification components

The OneVsRest and SVM classification components have been moved from the mleap-spark module to the mleap-spark-extension module. The OneVsRest implementation now supports probabilistic classification models by selecting the appropriate output column (probability vs. raw prediction), and the SVM implementation updates the threshold handling to use the new clearThreshold method, while also correcting the column selection order in the training pipeline.

mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/classification · high confidence

Refactored Bundle.ML DSL and serialization interfaces to support context-aware operations

The Bundle.ML DSL has been refactored to introduce a new \HasAttributes\ trait and \Attributes\ class for managing model metadata, alongside new \Bundle\, \Model\, \Node\, and \NodeShape\ classes that define the structure of serialized pipelines. Additionally, the \OpModel\ and \OpNode\ type classes have been updated to include a \Context\ type parameter, allowing serialization and deserialization logic to access the \BundleContext\ for handling custom types and deferred model names.

bundle-ml/src/main/scala/ml/combust/bundle/dsl · high confidence

Refactored LeapFrame and Row abstractions for consistent interface

The runtime frame package has been refactored to provide a consistent interface for working with data frames and rows. A new \ArrayRow\ class and \Row\ trait define the row abstraction, while \DefaultLeapFrame\ and \LeapFrame\ traits standardize frame operations like \select\, \withColumn\, \withColumns\, \drop\, and \filter\. The \FrameBuilder\ trait unifies the API for frame construction and modification, and \RowTransformer\ enables chaining of transformations. This change simplifies how users interact with and manipulate data frames and rows within the MLeap runtime.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame · high confidence

Refactored serialization architecture to separate readers and writers

The serialization module has been restructured to decouple reading and writing operations into distinct components. New \FrameReader\, \FrameWriter\, \RowReader\, and \RowWriter\ traits and objects have been introduced to handle data serialization and deserialization, replacing the previous monolithic \FrameSerializer\ trait. This change provides a more modular approach to handling different data formats (such as JSON, binary, and Avro) by allowing users to interact with specific read or write operations rather than a single serializer interface.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization · high confidence

Regression transformers migrated to the new typing system

The regression transformer implementations (AFTSurvivalRegression, GBTRegression, GeneralizedLinearRegression, IsotonicRegression, DecisionTreeRegression, LinearRegression, and RandomForestRegression) have been refactored to use the new typing system. This replaces the previous \TransformBuilder\-based approach with a direct \UserDefinedFunction\ execution model, simplifying how these models process input features and produce predictions.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/regression · high confidence

Relocated BuildInfo and added ClassLoaderUtil utility

The BuildInfo class has been moved to a new location within the mleap-base module, updating the package structure for build metadata. Additionally, a new ClassLoaderUtil object has been introduced to provide a robust method for finding the appropriate classloader by inspecting the call stack, supporting consistent class loading behavior across the library.

mleap-base · medium confidence

Removal of MLeap serialization ops for classification, regression, and feature transformers

The serialization operations for a range of MLeap models have been removed, including Pipeline, DecisionTree, LogisticRegression, OneVsRest, RandomForest, SupportVectorMachine, ReverseStringIndexer, StandardScaler, StringIndexer, VectorAssembler, and LinearRegression. This eliminates the ability to serialize and deserialize these specific model types in the current runtime.

mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization/bundle/ops · high confidence

Removed obsolete Bundle.ML serialization classes

The serializer package in bundle-ml has been cleaned up by removing the legacy serialization infrastructure. Specifically, the following files and classes have been deleted: BundleContext, BundleSerializer, ModelSerializer, NodeSerializer, SerializationContext, SerializationFormat, and the entire custom serialization hierarchy (CustomSerializer, CustomType, and their JSON/Protobuf implementations). This removes the ability to serialize/deserialize Bundle.ML models using the old mixed JSON/Protobuf approach, likely in preparation for a new serialization strategy.

bundle-ml/src/main/scala/ml/bundle/serializer · high confidence

TensorFlow integration configuration added

A new reference.conf file was added to configure the TensorFlow transformer operations, specifically registering the TensorflowTransformerOp and the associated operations package in the MLeap registry.

mleap-tensorflow/src/main/resources · high confidence

Updated Spark example data

The Spark example in the repository has been updated with a new CSV file (spark-demo.csv) containing sample data for testing or demonstration purposes. This change modifies the example files to reflect current usage patterns.

examples · medium confidence

Test coverage

Add serialization test for CountVectorizer null vocabulary; Add test utilities and specs for the MLeap executor; Added ALS parity test; Added GrpcSpec and test utilities for the MLeap gRPC server; Added LinearSVC parity test; Added PySpark feature tests for MathUnary, MathBinary, and StringMap; Added Python unit tests for sklearn feature extraction and tree models; Added Python unit tests for sklearn model serialization and deserialization; Added agaricus dataset in CSV format to work around libsvm loading bug; Added and standardized regression model tests; Added comprehensive test coverage for MLeap runtime components; Added comprehensive test coverage for TensorFlow integration; Added comprehensive unit tests for MLeap Tensor equality and legacy indexing; Added comprehensive unit tests for MLeap core feature models; Added integration tests for MLeap Spring Boot scoring endpoints; Added integration tests for MLeap executor components; Added parity test for Support Vector Machine (SVM) classification; Added parity tests for CrossValidator and TrainValidationSplit; Added parity tests for Spark ML classification algorithms; Added parity tests for Spark ML feature transformers; Added parity tests for XGBoost classification and regression models; Added parity tests for multiple regression models; Added test coverage for clustering model schemas; Added test coverage for sklearn PolynomialFeaturesModel; Added test data for XGBoost integration; Added test infrastructure for PySpark; Added test suite for Bundle.ML serialization and file system operations; Added tests for ALS model prediction and schema; Added tests for Avro serialization of LeapFrames and Rows; Added tests for Gensim Word2Vec integration; Added tests for multi-output transformer support; Added tests for sklearn extension components; Added tests for sklearn preprocessing transformers; Added tests for type casting and struct type operations; Added unit tests for MleapReflection utility; Added unit tests for classification model schemas; Added unit tests for vector conversion utilities.

Dependencies

Build system overhauled for Scala 2.13, Java 17, and Spark 4.1.2

The project's build infrastructure has been significantly upgraded to support Scala 2.13.18, Java 17, and Spark 4.1.2. This includes migrating from sbt 0.13 to 1.12.5, updating core dependencies such as Spring Boot 3.2.0, XGBoost 2.0.3, and TensorFlow Java 1.0.0, and introducing new build configuration files (BuildInfo, DockerConfig, MleapProject) to manage the new module structure and publishing settings.

project · high confidence

MLeap 0.25.2 release and dependency upgrades

This release upgrades the project to Spark 4.1.2, TensorFlow 2.16.2, and XGBoost 2.0.3, while also updating the Scala version to 2.13.18 and Java to 17. The release notes document these changes, and the README is updated to reflect the new dependency compatibility matrix and installation instructions for PySpark integration.

(repo-wide) · high confidence

MLeap 0.25.2: Upgrade to Spark 4.1.2

The project has been upgraded to Spark 4.1.2 (MLeap 0.25.2). This update also includes a reorganization of the build configuration, moving module definitions into a centralized \MleapProject\ object and introducing individual \build.sbt\ files for each module to standardize dependency and settings management.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 45 → 58 (+13.5)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 100 → 92 (-8.2)
  • Architecture 94 → 97 (+2.8)
  • Maturity 52 → 57 (+4.2)
  • Readiness 27 → 44 (+17.3)
  • Security 48 → 76 (+28.2)

Resolved (40)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • High IaC: DS-0002 (.devcontainer/Dockerfile)
  • High IaC: DS-0029 (.devcontainer/Dockerfile)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • …and 20 more

New (114)

  • Casting.cast (cognitive 61) (mleap-core/src/main/scala/ml/combust/mleap/core/types/Casting.scala)
  • Casting.cast (cyclomatic 77) (mleap-core/src/main/scala/ml/combust/mleap/core/types/Casting.scala)
  • Concentrated knowledge decay
  • Documentation: no contributor guidance (README.md)
  • Duplicated block (10 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/classification/DecisionTreeClassifierModel.scala)
  • Duplicated block (10 lines × 2) (mleap-grpc/src/main/scala/ml/combust/mleap/grpc/GrpcClient.scala)
  • Duplicated block (12 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame/IndexedRowUtil.scala)
  • Duplicated block (12 lines × 2) (mleap-spark/src/main/scala/org/apache/spark/ml/bundle/ops/tuning/CrossValidatorOp.scala)
  • Duplicated block (15 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame/IndexedRowUtil.scala)
  • Duplicated block (15 lines × 2) (mleap-spark-base/src/main/scala/org/apache/spark/ml/bundle/SparkShape.scala)
  • Duplicated block (15–16 lines × 2) (mleap-executor/src/main/scala/ml/combust/mleap/executor/service/LocalTransformService.scala)
  • Duplicated block (17–18 lines × 2) (mleap-databricks-runtime-testkit/src/main/scala/ml/combust/mleap/databricks/runtime/testkit/TestSparkMl.scala)
  • Duplicated block (18 lines × 2) (mleap-executor/src/main/scala/ml/combust/mleap/executor/service/LocalTransformService.scala)
  • Duplicated block (21–22 lines × 2) (mleap-databricks-runtime-testkit/src/main/scala/ml/combust/mleap/databricks/runtime/testkit/TestSparkMl.scala)
  • Duplicated block (5 lines × 2) (mleap-xgboost-spark/src/main/scala/ml/dmlc/xgboost4j/scala/spark/mleap/XGBoostClassificationModelOp.scala)
  • Duplicated block (6 lines × 2) (bundle-ml/src/main/scala/ml/combust/bundle/tree/cluster/NodeSerializer.scala)
  • Duplicated block (6 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/feature/FeatureHasherModel.scala)
  • Duplicated block (6 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/feature/InteractionModel.scala)
  • Duplicated block (6 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/bundle/tree/clustering/MleapNodeWrapper.scala)
  • Duplicated block (6 lines × 2) (mleap-spark-base/src/main/scala/org/apache/spark/sql/mleap/TypeConverters.scala)
  • …and 94 more

Architecture

  • Containers 0 added · 0 removed · contexts 12 added · 0 removed · edges 35 added · 0 removed

Added bounded contexts (12)

  • bundle-ml
  • mleap-avro
  • mleap-base
  • mleap-core
  • mleap-executor
  • mleap-runtime
  • mleap-spark-base
  • mleap-spark-extension
  • mleap-spark-testkit
  • mleap-tensor
  • mleap-tensorflow
  • mleap-xgboost-runtime

Added dependency edges (35)

  • bundle-ml → mleap-core (coupling)
  • bundle-ml → mleap-runtime
  • bundle-ml → mleap-tensor
  • mleap-avro → mleap-core (coupling)
  • mleap-avro → mleap-runtime (coupling)
  • mleap-base → mleap-core
  • mleap-core → bundle-ml (coupling)
  • mleap-core → mleap-tensor
  • mleap-executor → bundle-ml (coupling)
  • mleap-executor → mleap-core (coupling)
  • mleap-executor → mleap-runtime (coupling)
  • mleap-runtime → bundle-ml (coupling)
  • mleap-runtime → mleap-core (coupling)
  • mleap-runtime → mleap-tensor
  • mleap-spark-base → bundle-ml (coupling)
  • mleap-spark-base → mleap-core (coupling)
  • mleap-spark-base → mleap-runtime (coupling)
  • mleap-spark-extension → bundle-ml (coupling)
  • mleap-spark-extension → mleap-core (coupling)
  • mleap-spark-extension → mleap-runtime (coupling)
  • …and 15 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

combust/mleap was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 88de54dd4fb7720c7c56ff5c2b51cdec4353a128 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.