combust/mleap
58.2
Weak · 27 September 2026
28.2k
lines of production code
Scala
with Python
4
measurements over time
What this system is
This release delivers a comprehensive overhaul of the MLeap serialization and runtime architecture, introducing a new Bundle.ML format that supports JSON, Protobuf, and binary serialization for a wide range of machine learning models. The update significantly expands model support, adding transformers for classification, clustering, regression, and feature engineering, alongside new integrations for TensorFlow, XGBoost, and Scikit-learn. Additionally, the release introduces a robust model serving layer with gRPC and HTTP endpoints, while also upgrading the underlying build system to support Scala 2.13, Java 17, and Spark 4.1.2.
Features
Add ALS recommendation transformer
A new ALS (Alternating Least Squares) transformer has been added to the MLeap runtime, enabling recommendation capabilities. The implementation defines an ALS case class that wraps an ALSModel and exposes a user-defined function to compute recommendations for a given user and item.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/recommendation · high confidence
Add Avro serialization support for LeapFrame and Row data
The mleap-avro module now includes new implementations for serializing and deserializing MLeap data structures to and from the Avro format. This introduces DefaultFrameReader and DefaultFrameWriter for full LeapFrame serialization, as well as DefaultRowReader and DefaultRowWriter for individual row-level Avro operations. The change includes a SchemaConverter that maps MLeap data types (including scalars, lists, maps, and various tensor types) to Avro schemas, and a ValueConverter that handles the actual data conversion between MLeap and Avro representations.
mleap-avro/src/main · high confidence
Add CategoricalDrilldown transformer for ensemble models
A new CategoricalDrilldown transformer has been added to the runtime, enabling the application of multiple underlying transformers based on a categorical label. This allows users to route data through specific ensemble members, with output columns automatically prefixed with 'drilldown.' to distinguish the transformed fields.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/ensemble · high confidence
Add Databricks runtime testkit with Spark ML, TensorFlow, and XGBoost integration tests
A new testkit module for Databricks runtime has been added, providing integration tests for loading and saving ML models. The suite includes test cases for Spark ML pipelines (using LogisticRegression), TensorFlow models (add and multiply graph operations), and XGBoost classifiers, each verifying the ability to write and load MLeap bundles in a Databricks environment.
mleap-databricks-runtime-testkit/src · high confidence
Add HDFS Bundle File System support
Users can now store and retrieve ML bundles on HDFS. This change introduces a new Hadoop-based file system implementation (HadoopBundleFileSystem) that handles loading and saving bundles to HDFS, along with the corresponding configuration in reference.conf and unit tests to verify the new functionality.
bundle-hdfs/src · high confidence
Add JSON serialization support for Bundle.ML types
A new \JsonSupport\ trait is introduced in the \bundle-ml\ module, providing \JsonFormat\ implementations for core Bundle.ML types including \BasicType\, \ByteString\, \Tensor\, \Scalar\, and \DataShape\. This enables these types to be serialized to and deserialized from JSON, facilitating the storage and transfer of machine learning model bundles in a human-readable format.
bundle-ml/src/main/scala/ml/combust/bundle/json · high confidence
Add JSON serialization support for MLeap frames and rows
Users can now serialize and deserialize MLeap LeapFrames and Rows to and from JSON format. This change introduces new components in the mleap-runtime module (DefaultFrameReader, DefaultFrameWriter, DefaultRowReader, DefaultRowWriter, JsonSupport, and RowFormat) that enable converting MLeap data structures into human-readable JSON strings and back, supporting various data types including basic types, lists, maps, tensors, and byte strings.
mleap-runtime/src/main/scala/ml/combust/mleap/json · high confidence
Add JSON serialization support for decision and clustering trees
Users can now serialize and deserialize decision trees and clustering nodes to and from JSON format. This change introduces new serialization logic in the bundle-ml module, including JSON format writers and readers for both decision trees (internal/leaf nodes) and clustering nodes, as well as supporting type classes (NodeWrapper) to bridge the gap between internal model representations and the bundle format.
bundle-ml/src/main/scala/ml/combust/bundle/tree · high confidence
Add MLEP serialization for XGBoost classification and regression models
Added new MLEP operators for serializing and deserializing XGBoost classification and regression models. The new \XGBoostClassificationModelOp\ and \XGBoostRegressionModelOp\ classes handle the storage and loading of model parameters such as tree limits, thresholds, and missing values, enabling these Spark ML models to be exported and imported via MLEP bundles.
mleap-xgboost-spark/src/main/scala · high confidence
Add MLeap extensions for scikit-learn transformers
New MLeap extension classes are introduced for scikit-learn's SimpleImputer and a custom DefineEstimator wrapper, enabling these transformers to be used within MLeap pipelines and serialized to bundles. The Imputer class wraps scikit-learn's SimpleImputer to support feature extraction and serialization, while the DefineEstimator class facilitates running transformers on specific columns of input data.
python/mleap/sklearn/extensions · high confidence
Add MLeap serialization support for Gensim Word2Vec models
Users can now serialize Gensim Word2Vec models to the MLeap bundle format. This change introduces a new module at python/mleap/gensim/word2vec.py that patches the gensim.models.Word2Vec class with MLeap-specific methods (serialize\_to\_bundle, sent2vec) and a SimpleSparkSerializer to handle the conversion of word vectors and metadata into MLeap-compatible bundles.
python/mleap/gensim · high confidence
Add MLeap serialization support for scikit-learn preprocessing transformers
The MLeap Python library now supports serializing and deserializing several scikit-learn preprocessing transformers (StandardScaler, MinMaxScaler, OneHotEncoder, SimpleImputer, Binarizer, and PolynomialFeatures) into MLeap bundles. This enables users to export models trained with these scikit-learn components for execution in MLeap or Spark environments, maintaining parity with the MLeap runtime's expected input/output feature structures.
python/mleap/sklearn/preprocessing · high confidence
Add MLeap serving module for model serving
A new mleap-serving module is introduced, providing a unified entry point to start both gRPC and HTTP servers backed by a single MLeap executor. The module exposes configuration for the gRPC port and coordinates the startup of the gRPC server (mleap-grpc-server) and the HTTP server (mleap-spring-boot), ensuring that models loaded through one interface are available through the other.
mleap-serving · high confidence
Add MLeap-TensorFlow converter implementation for tensor type mapping
The MLeap-TensorFlow integration now includes the core conversion logic to map between MLeap tensor types and TensorFlow tensors. The new \MleapConverter\ and \TensorflowConverter\ classes handle bidirectional conversion for numeric types (int, long, float, double), strings, and byte strings, enabling the TensorFlow backend to serialize and deserialize model data using the updated TensorFlow Java API.
mleap-tensorflow/src/main/scala/ml/combust/mleap/tensorflow/converter · medium confidence
Add S3-backed repository support for MLeap executor
Users can now store and download MLeap bundles directly from Amazon S3 using s3:// URIs. The new S3Repository implementation allows the executor to fetch bundles from S3 buckets, with configuration provided via the standard AWS credential provider chain. This adds S3 as a supported repository type alongside existing file and HTTP repositories.
mleap-repository-s3 · high confidence
Add Sklearn PolynomialFeatures transformer
Introduces a new PolynomialFeatures transformer for the Sklearn integration, enabling polynomial feature expansion within the MLeap runtime.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/sklearn · high confidence
Add SparkUtil utility for direct pipeline creation
A new SparkUtil object has been added to the Spark ML MLeap module, providing a convenience method to create a PipelineModel directly from an array of Transformers. This simplifies pipeline construction by allowing users to instantiate a PipelineModel without explicitly creating a Pipeline first.
mleap-spark-base/src/main/scala/org/apache/spark/ml/mleap · high confidence
Add TensorFlow model loading and transformation support
Introduces new Scala classes (TensorflowModel, TensorflowTransformer, and TensorflowTransformerOp) that enable MLeap to load and run TensorFlow models in both 'graph' and 'saved\_model' formats, allowing users to integrate TensorFlow-based transformations into their MLeap pipelines.
mleap-tensorflow/src/main/scala/ml/combust/mleap/tensorflow · high confidence
Add automatic casting of data types in LeapFrame and Spark DataFrame conversion
A new TypeConverters trait has been introduced to handle the conversion between Spark DataFrames and MLeap LeapFrames. This includes mapping Spark data types (such as Boolean, Byte, Short, Int, Long, Float, Double, String, and various Array types) to their corresponding MLeap types, and vice versa. The implementation also adds support for converting Vector and Matrix types to and from MLeap Tensor types, ensuring that data shapes and nullability are correctly preserved during the conversion process.
mleap-spark-base/src/main/scala/org/apache/spark/sql · high confidence
Add built-in Spark ML operation registry
The MLeap Spark integration now includes a comprehensive built-in registry of Spark ML operations, enabling automatic serialization and deserialization of a wide range of Spark ML models and transformers. This configuration file registers support for classification algorithms (e.g., Decision Tree, Naive Bayes, Random Forest), clustering (K-Means, GMM), feature transformers (e.g., Tokenizer, PCA, Vector Assembler), regression models, and tuning components (Cross-Validation, Train-Validation Split). This allows users to seamlessly export and import Spark ML pipelines and individual models using MLeap's format without manual registration.
mleap-spark/src/main/resources · high confidence
Add bundle serialization for new and existing feature transformers
Adds bundle serialization support for several feature transformers, including Imputer, MapEntrySelector, MathBinary, MathUnary, MultinomialLabeler, StringMap, and WordLengthFilter. This enables these Spark ML transformers to be serialized into MLeap bundles for deployment.
mleap-spark-extension/src/main/scala/org/apache/spark/ml/bundle/extension/ops/feature · high confidence
Add high-performance XGBoost Predictor runtime support
Introduces a new, high-performance implementation for XGBoost classification and regression models using the XGBoost Predictor library. This adds new runtime classes (e.g., XGBoostPredictorClassification, XGBoostPredictorRegression) and their corresponding bundle operations, enabling faster inference for MLeap users while maintaining compatibility with existing XGBoost model formats.
mleap-xgboost-runtime/src/main · high confidence
Add serialization support for OneVsRest and SVM classification models
Users can now serialize and deserialize OneVsRest and Support Vector Machine (SVM) classification models. The new OneVsRestOp and SupportVectorMachineOp implementations in the Spark bundle extension allow these models to be saved to and loaded from MLeap bundles, including handling of model parameters such as thresholds and class metadata.
mleap-spark-extension/src/main/scala/org/apache/spark/ml/bundle/extension/ops/classification · high confidence
Add serialization support for Scikit-learn models
The Python integration with Scikit-learn now supports serializing and deserializing several core model types, including LinearRegression, LogisticRegression, LogisticRegressionCV, FeatureUnion, Pipeline, SVC, and LinearSVC. These changes enable users to export trained Scikit-learn models into the MLeap bundle format for deployment.
python/mleap/sklearn · high confidence
Add serialization support for scikit-learn Decision Trees and Random Forests
Users can now serialize scikit-learn DecisionTreeClassifier, DecisionTreeRegressor, RandomForestClassifier, and RandomForestRegressor models to MLeap bundles. This change adds the necessary serialization logic in \python/mleap/sklearn/tree/tree.py\ and \python/mleap/sklearn/ensemble/forest.py\, enabling these specific model types to be exported and stored in the MLeap format.
python/mleap/sklearn/tree · high confidence
Added MLeap and Spark transform benchmarks
A new benchmarking tool has been introduced in the mleap-benchmark module, providing performance measurement for MLeap and Spark model transformations. The addition includes a Boot entry point for CLI execution, a shared Benchmark trait, and specific benchmark classes for MLeap row/transform operations and Spark transforms, enabling users to evaluate execution performance across different runtime environments.
mleap-benchmark/src/main/scala · high confidence
Added MLeap gRPC client module for remote model serving
Users can now interact with a remote MLeap gRPC server to load, unload, transform, and stream data against hosted models. This new \mleap-grpc\ module provides a \GrpcClient\ implementation of the \TransformService\ trait, handling the translation between MLeap types and their Protobuf representations via Akka stream helpers. This enables client-side integration with the \mleap-grpc-server\ for remote inference workflows.
mleap-grpc · high confidence
Added Travis CI deployment and release scripts
Added new shell scripts (docker.sh, extract.sh, publish.sh, release.sh, travis\_publish.sh) to automate the release and publishing process on Travis CI. The release script ensures builds only run on the master branch for pull requests, configures git credentials, and triggers SBT publishing tasks for MLeap and Spring Boot modules.
travis · high confidence
Added XGBoost model operation registrations for Spark
A new configuration file (reference.conf) was added to register the XGBoost classification and regression model operations with the MLeap Spark registry, enabling serialization and deserialization of these models within the Spark integration.
mleap-xgboost-spark/src/main/resources · high confidence
Added benchmark configuration and test data for MLeap
The MLeap benchmark module now includes its own configuration files and test data. A new JSON file (frame.airbnb.json) provides sample data for the LeapFrame, while mleap.conf and spark.conf define benchmark parameters such as iteration ranges and run counts. Additionally, a log4j.properties file is added to control logging verbosity for the benchmarking process.
mleap-benchmark/src/main/resources · medium confidence
Added clustering transformer implementations
New transformer classes have been added to the runtime to support specific clustering algorithms. The diff introduces Scala implementations for K-Means, Gaussian Mixture Models (GMM), Latent Dirichlet Allocation (LDA), and Bisecting K-Means. Each class wraps its corresponding model (e.g., KMeansModel, GaussianMixtureModel) to execute predictions or topic distributions on input tensors, enabling these clustering models to be executed within the MLeap runtime.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/clustering · high confidence
Added sample datasets and parity test infrastructure to the Spark testkit
The mleap-spark-testkit module now includes sample data files (a text file for NLP tasks, a ratings file for recommendation systems, and a libSVM file for classification) and a new Scala base class, SparkParityBase. This base class provides the infrastructure for running Spark/MLeap parity tests, including helper methods to load the new datasets, serialize models to bundles, and compare Spark and MLeap outputs with relative tolerance.
mleap-spark-testkit · high confidence
Added script to compile and export Scala classpath for Python integration
A new shell script, scripts/scala\_classpath\_for\_python.sh, has been added to automate the compilation of MLeap Spark extension classes and export the resulting Scala classpath. This script ensures that the necessary Scala classes are compiled and available in the environment for Python-based tests and integrations, streamlining the setup process for developers working with the Python bindings.
scripts · high confidence
Added serialization support for scikit-learn text vectorizers
Users can now serialize scikit-learn's CountVectorizer and TfidfVectorizer models to MLeap bundles. This change adds the necessary serialization logic in python/mleap/sklearn/feature\_extraction/text.py, enabling these text processing components to be exported and reused within the MLeap ecosystem.
_python/mleap/sklearn/feature\extraction · high confidence
Added support for serializing and deserializing tensors to and from JSON
Users can now serialize and deserialize DenseTensor, SparseTensor, and general Tensor objects directly to and from JSON format. This is enabled by new Scala source files (ByteString, JsonSupport, and Tensor) that define the JSON structure for tensor data, allowing for easier interchange and storage of tensor data in JSON-compatible formats.
mleap-tensor/src/main · high confidence
Adds core model implementations and neural network layer definitions
Introduces new model classes for classification (GBT, LinearSVC, MultiLayerPerceptron, Naive Bayes), clustering (BisectingKMeans, GaussianMixture, KMeans, LDA), and feature engineering (Binarizer, Bucketizer, ChiSqSelector, Coalesce, CountVectorizer). Additionally, adds the \ann\ package containing neural network layer definitions (AffineLayer, LossFunctions) and utility classes (BreezeUtil, Layer, Model) to support deep learning operations outside of a Spark context.
mleap-core/src/main · high confidence
Classification transformers now expose raw predictions and probabilities
The classification transformers in the MLeap runtime (including DecisionTree, LogisticRegression, RandomForest, SupportVectorMachine, and others) have been updated to output raw predictions and probability scores alongside standard predictions. This change allows users to access intermediate model outputs (rawPrediction, probability) in addition to the final prediction, enabling more flexible downstream processing and evaluation.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/classification · high confidence
Expanded MLeap runtime supports a wider range of ML models and transformers
The MLeap runtime now includes a comprehensive set of built-in operations for classification, clustering, regression, and feature transformation. This update enables the serialization and deserialization of a broader range of machine learning models and data processing steps, including classifiers like Decision Tree, GBT, and Logistic Regression; clustering models such as K-Means and Gaussian Mixture; regression models including Linear and Generalized Linear Regression; and various feature transformers like StringIndexer, VectorAssembler, and Tokenizer. Users can now deploy and serve these specific model types directly via MLeap.
mleap-runtime/src/main/resources · high confidence
Introduce BundleContext and BundleFileSystem for extensible bundle storage
The bundle library now supports pluggable file systems for reading and writing Bundle.ML models. A new BundleContext class holds serialization format, registry, and file system information, while the BundleRegistry tracks registered file systems by URI scheme. This allows models to be loaded from and saved to various storage backends (e.g., local file, JAR, or custom URIs) through a unified interface.
bundle-ml/src/main/scala/ml/combust/bundle · high confidence
Introduce MLeap executor with configurable model serving and streaming
The MLeap executor now provides a new local model serving and transformation service, exposing APIs for loading, unloading, and transforming models via both synchronous calls and Akka-based streams. The executor supports configurable timeouts for memory and disk caching, stream and flow parallelism, buffer sizes, and throttling. It also introduces a multi-repository system that can load models from local files or HTTP sources, with configurable thread pools and idle timeouts.
mleap-executor/src · high confidence
Introduce MLeap gRPC server for model inference
A new gRPC server implementation has been added to MLeap, enabling model inference via the gRPC protocol. The server exposes endpoints for loading, unloading, and transforming frames and rows, supporting both unary and streaming requests. It includes a configurable port (defaulting to 65328, overridable via the MLEAP\_GRPC\_PORT environment variable) and an error interceptor that maps internal exceptions like NotFoundException or TimeoutException to appropriate gRPC status codes.
mleap-grpc-server/src/main · high confidence
Introduce MLeap serialization and deserialization for Python transformers
Added the \python/mleap/bundle\ module, which provides \MLeapSerializer\ and \MLeapDeserializer\ classes to convert Python transformers (such as Scikit-learn estimators) into MLeap bundle formats (model.json and node.json). This enables serializing transformer attributes like coefficients, intercepts, and scalers into a structured JSON representation, facilitating interoperability with MLeap-based pipelines.
python/mleap/bundle · high confidence
Introduce Python package structure and tooling
The Python integration is now packaged as a proper Python library, enabling installation via pip and standard build tools. The package includes a \setup.py\ that enforces Python 3.10+ and lists dependencies such as scikit-learn 1.0, pandas, and scipy. Additionally, a \Makefile\ and \tox.ini\ are provided to streamline running tests and managing the development environment.
python · high confidence
Introduce binary serialization for LeapFrame and Row data
A new binary serialization format is added for serializing and deserializing LeapFrame and Row objects. This includes new readers and writers (DefaultFrameReader, DefaultFrameWriter, DefaultRowReader, DefaultRowWriter) and a ValueSerializer that handles primitive types, strings, ByteStrings, lists, and tensors, enabling more compact and efficient storage or transmission of model data compared to previous formats.
mleap-runtime/src/main/scala/ml/combust/mleap/binary · high confidence
Introduces new MLeapOp base classes and serialization ops for ML models
The MLeap runtime now includes a new abstract \MleapOp\ base class and a \MultiInOutMleapOp\ variant to standardize how model operations are serialized and deserialized. This change introduces serialization logic for a wide range of machine learning models, including classification (Decision Tree, GBT, Logistic Regression, Naive Bayes, Random Forest, SVM), clustering (K-Means, Gaussian Mixture, LDA), and feature transformers (Binarizer, Bucketizer, Count Vectorizer, etc.). Users will see these models supported in the MLeap bundle format, enabling consistent storage and loading of these algorithms.
mleap-runtime/src/main/scala/ml/combust/mleap/bundle · high confidence
Major overhaul of feature transformers to use a new NodeShape-based API
The feature transformer implementations in the runtime have been refactored to use a new \NodeShape\-based API, replacing the previous \TransformBuilder\-based \build\ methods with \SimpleTransformer\ or \Transformer\ implementations that define an \exec\ function. This change standardizes how transformers like \StringIndexer\, \OneHotEncoder\, and \StandardScaler\ handle inputs and outputs, often simplifying the code and aligning with a more consistent internal interface. Additionally, several new feature transformers have been added, including \Binarizer\, \Coalesce\, \CountVectorizer\, \DCT\, \ElementwiseProduct\, \FeatureHasher\, \IDF\, \Imputer\, \Interaction\, \MapEntrySelector\, \MathBinary\, \MathUnary\, \MaxAbsScaler\, \MinHashLSH\, \MinMaxScaler\, \MultinomialLabeler\, \NGram\, \Normalizer\, \Pca\, \PolynomialExpansion\, \RegexIndexer\, \RegexTokenizer\, \StopWordsRemover\, \StringMap\, \VectorIndexer\, \VectorSlicer\, \WordLengthFilter\, and \WordToVector\. The \HashingTermFrequency\ transformer was also updated to use the new API.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/feature · high confidence
New Bundle.ML serialization framework with JSON and Protobuf support
The Bundle.ML serialization system has been refactored to support both JSON and Protobuf formats. A new \BundleSerializer\ handles high-level bundle read/write operations, while dedicated serializers (\ModelSerializer\, \NodeSerializer\, \GraphSerializer\) manage the underlying data. The \SerializationFormat\ trait and its \Json\/\Protobuf\ implementations allow the system to switch between serialization backends. Additionally, a deprecated \FileUtil\ class is provided for backward compatibility, delegating to the new \ml.combust.bundle.util.FileUtil\.
bundle-ml/src/main/scala/ml/combust/bundle/serializer · high confidence
New Java DSL for MLeap runtime operations
A new Java DSL is introduced in the mleap-runtime module, providing convenient builder and support classes for Java users. This includes BundleBuilder and BundleBuilderSupport for loading and saving MLeap bundles, ContextBuilder for creating MleapContext and loading registries, LeapFrameBuilder and LeapFrameBuilderSupport for constructing LeapFrames, schemas, and rows, as well as support classes for RowTransformer, Tensor, and LeapFrame operations. These additions streamline the creation and manipulation of MLeap runtime objects from Java code.
mleap-runtime/src/main/java · high confidence
New PySpark feature transformers: MathBinary, MathUnary, and StringMap
Added new PySpark wrappers for MLeap feature transformers: MathBinary (supporting Add, Subtract, Multiply, Divide, Remainder, LogN, Pow, Min, Max), MathUnary (supporting Sin, Cos, Tan, Log, Exp, Abs, Sqrt, Logit), and StringMap (with configurable invalid handling). These allow users to perform mathematical operations on columns and map string labels to numeric values directly within PySpark pipelines.
python/mleap/pyspark/feature · high confidence
New Spark bundle serialization infrastructure
The MLeap Spark integration now includes a new set of core classes for handling Spark ML model serialization. This includes the SparkBundleContext to manage datasets and registries, SimpleSparkOp and MultiInOutSparkOp to handle single and multi-column transformer serialization, and SparkShape/ParamSpec to map Spark DataFrame schemas to bundle node shapes. This infrastructure enables the framework to serialize and deserialize Spark ML pipelines and transformers into MLeap bundles.
mleap-spark-base/src/main/scala/org/apache/spark/ml/bundle · high confidence
New Spark integration layer for MLeap serialization
Added three new Scala files (SimpleSparkSerializer, SparkLeapFrame, SparkSupport) that provide a bridge between Spark DataFrames and MLeap's internal LeapFrame representation. This enables users to serialize Spark ML transformers to MLeap bundles and deserialize them back, with automatic type conversion between Spark and MLeap schemas.
mleap-spark-base/src/main/scala/ml · high confidence
New feature transformers: Imputer, MapEntrySelector, MathBinary, MathUnary, MultinomialLabeler, StringMap, and WordLengthFilter
Added several new Spark ML transformers to the mleap-spark-extension module. These include Imputer for filling missing values using mean or median, MapEntrySelector for extracting map entries, MathBinary and MathUnary for performing binary and unary mathematical operations, MultinomialLabeler for converting probability vectors to labels, StringMap for mapping string labels to numeric values, and WordLengthFilter for filtering words by length. Each transformer implements MLWritable/MLReadable for serialization and provides standard Spark ML API interfaces.
mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/feature · high confidence
Serialization support for Spark ML transformers and classifiers
Adds serialization and deserialization support for a broad set of Spark MLlib transformers and classifiers, including GBTClassifier, MultiLayerPerceptronClassifier, NaiveBayes, BisectingKMeans, GaussianMixture, KMeans, LDA, Binarizer, BucketedRandomProjectionLSH, Bucketizer, ChiSqSelector, CountVectorizer, DCT, ElementwiseProduct, FeatureHasher, IDF, Interaction, MaxAbsScaler, MinHashLSH, MinMaxScaler, NGram, and Normalizer. This enables users to save and load these models in the MLeap bundle format.
mleap-spark/src/main/scala · high confidence
Removals
Removed Bundle.ML DSL and serialization support files
The \bundle-ml\ module's domain-specific language (DSL) and associated serialization infrastructure have been removed. This includes the deletion of core DSL classes such as \Attribute\, \AttributeList\, \Bundle\, \Model\, \Node\, \Shape\, and \Value\, as well as the \JsonSupport\ for serializing Bundle.ML objects to JSON. This change eliminates the ability to construct, serialize, and deserialize Bundle.ML pipelines and graphs using the previous Scala-based DSL and JSON-based serialization format.
bundle-ml/src/main/scala/ml/bundle/dsl · high confidence
Removed legacy MLeap serialization and registry classes
The \MleapBundle\ and \MleapRegistry\ objects in the \mleap-runtime\ module have been removed. These classes previously handled the serialization and deserialization of MLeap bundle files and maintained a static registry of machine learning operators. Their removal indicates a shift in how the runtime manages model serialization and operator registration, likely consolidating these responsibilities into a newer, unified registry or bundle format.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization/bundle · high confidence
Behavioural changes
Add default Spark extension operations to MLeap configuration
A new reference configuration file registers default Spark extension operations, including OneVsRest, SupportVectorMachine, Imputer, MapEntrySelector, MathBinary/Unary, MultinomialLabeler, WordLengthFilter, and StringMap. This change ensures these Spark ML components are automatically available for serialization and deserialization without requiring manual configuration.
mleap-spark-extension/src/main/resources · high confidence
Added file utility for safe ZIP extraction and directory removal
A new FileUtil utility class was added to provide safe file system operations, including a recursive directory removal method and a ZIP extraction method that validates entry paths to prevent Zip Slip vulnerabilities.
bundle-ml/src/main/scala/ml/combust/bundle/util · medium confidence
Adds PySpark integration utilities and module registration
The PySpark wrapper now includes dedicated modules for serializing and deserializing MLeap bundles directly from Spark DataFrames, exposing \serializeToBundle\ and \deserializeFromBundle\ methods on Spark Transformers. Additionally, the package registers MLeap features (StringMap, MathBinary, MathUnary) under the \pyspark.ml.mleap\ namespace to simplify imports and improve classpath injection for Python tests.
python/mleap/pyspark · high confidence
MLeap Python package version updated to 0.25.2
The MLeap Python package now exposes its version number via the public API, allowing users to programmatically check the installed version. The package version has been updated to 0.25.2.
python/mleap · high confidence
MLeap Spring Boot overhaul: new REST API and model loading
The MLeap Spring Boot module has been overhauled to provide a complete REST API for model management and scoring. New controllers (JsonScoringController, ProtobufScoringController, LeapFrameScoringController) expose endpoints for loading, unloading, and transforming data using both JSON and Protobuf formats. A new GlobalExceptionHandler standardizes error responses across the application, and a ModelLoader component enables automatic loading of models from a configured directory at startup.
mleap-spring-boot/src/main/scala · medium confidence
Migrated type system to support Bundle protocol conversions
The internal type system was refactored to support conversion between MLeap's internal data types and the external Bundle protocol. The previous \DataType\ definitions (including \TensorType\ and \ListType\) were removed and replaced with a new \BundleTypeConverters\ trait that maps basic types, shapes, and schemas between the two systems. This enables interoperability with the Bundle format, supporting scalar, list, and tensor data shapes with nullable support.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/types · high confidence
New binary serialization for tensors and array types
Added new \ArraySerializer\ implementations and a \TensorSerializer\ that convert tensors into a compact, binary format for storage and transmission. This replaces previous serialization methods, ensuring that tensor data (including booleans, bytes, shorts, ints, longs, floats, doubles, strings, and byte strings) is serialized efficiently using \ByteBuffer\ and \DataOutputStream\/\DataInputStream\ mechanisms.
bundle-ml/src/main/scala/ml/combust/bundle/tensor · high confidence
New param traits for labels and probabilities columns
The Spark ML extension now includes new parameter traits, HasLabelsCol and HasProbabilitiesCol, which allow transformers to specify output columns for labels and probabilities respectively. Additionally, the existing HasDropLast trait was moved from the core mleap-spark module to the extension module.
mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/param · high confidence
Pipeline transformer refactored to support async execution and schema introspection
The Pipeline transformer has been refactored to implement the new FrameTransformer interface, enabling asynchronous transformation via a new transformAsync method. The implementation now exposes input, output, and intermediate schemas through dedicated properties, allowing users to inspect the data structure at each stage of the pipeline. Additionally, the Pipeline class now manages the lifecycle of its constituent transformers by implementing a close method to ensure proper resource cleanup.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer · medium confidence
Refactor MLeap runtime API to use new frame and context abstractions
The MLeap runtime has been refactored to use a new \frame\ package for core data structures, replacing the previous \Dataset\, \LeapFrame\, and \Row\ implementations with \DefaultLeapFrame\, \Row\, and \ArrayRow\ in the \ml.combust.mleap.runtime.frame\ package. \MleapContext\ now manages the \BundleRegistry\ and class loading, while \MleapSupport\ provides updated implicit conversions for serializing/deserializing transformers and converting case classes to/from LeapFrames. The old \TransformBuilder\ and \LeapFrameBuilder\ are removed in favor of the new \FrameWriter\ and \RowReader\/\RowWriter\ abstractions.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime · high confidence
Refactor OneVsRest and SVM classification components
The OneVsRest and SVM classification components have been moved from the mleap-spark module to the mleap-spark-extension module. The OneVsRest implementation now supports probabilistic classification models by selecting the appropriate output column (probability vs. raw prediction), and the SVM implementation updates the threshold handling to use the new clearThreshold method, while also correcting the column selection order in the training pipeline.
mleap-spark-extension/src/main/scala/org/apache/spark/ml/mleap/classification · high confidence
Refactored Bundle.ML DSL and serialization interfaces to support context-aware operations
The Bundle.ML DSL has been refactored to introduce a new \HasAttributes\ trait and \Attributes\ class for managing model metadata, alongside new \Bundle\, \Model\, \Node\, and \NodeShape\ classes that define the structure of serialized pipelines. Additionally, the \OpModel\ and \OpNode\ type classes have been updated to include a \Context\ type parameter, allowing serialization and deserialization logic to access the \BundleContext\ for handling custom types and deferred model names.
bundle-ml/src/main/scala/ml/combust/bundle/dsl · high confidence
Refactored LeapFrame and Row abstractions for consistent interface
The runtime frame package has been refactored to provide a consistent interface for working with data frames and rows. A new \ArrayRow\ class and \Row\ trait define the row abstraction, while \DefaultLeapFrame\ and \LeapFrame\ traits standardize frame operations like \select\, \withColumn\, \withColumns\, \drop\, and \filter\. The \FrameBuilder\ trait unifies the API for frame construction and modification, and \RowTransformer\ enables chaining of transformations. This change simplifies how users interact with and manipulate data frames and rows within the MLeap runtime.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame · high confidence
Refactored serialization architecture to separate readers and writers
The serialization module has been restructured to decouple reading and writing operations into distinct components. New \FrameReader\, \FrameWriter\, \RowReader\, and \RowWriter\ traits and objects have been introduced to handle data serialization and deserialization, replacing the previous monolithic \FrameSerializer\ trait. This change provides a more modular approach to handling different data formats (such as JSON, binary, and Avro) by allowing users to interact with specific read or write operations rather than a single serializer interface.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization · high confidence
Regression transformers migrated to the new typing system
The regression transformer implementations (AFTSurvivalRegression, GBTRegression, GeneralizedLinearRegression, IsotonicRegression, DecisionTreeRegression, LinearRegression, and RandomForestRegression) have been refactored to use the new typing system. This replaces the previous \TransformBuilder\-based approach with a direct \UserDefinedFunction\ execution model, simplifying how these models process input features and produce predictions.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/transformer/regression · high confidence
Relocated BuildInfo and added ClassLoaderUtil utility
The BuildInfo class has been moved to a new location within the mleap-base module, updating the package structure for build metadata. Additionally, a new ClassLoaderUtil object has been introduced to provide a robust method for finding the appropriate classloader by inspecting the call stack, supporting consistent class loading behavior across the library.
mleap-base · medium confidence
Removal of MLeap serialization ops for classification, regression, and feature transformers
The serialization operations for a range of MLeap models have been removed, including Pipeline, DecisionTree, LogisticRegression, OneVsRest, RandomForest, SupportVectorMachine, ReverseStringIndexer, StandardScaler, StringIndexer, VectorAssembler, and LinearRegression. This eliminates the ability to serialize and deserialize these specific model types in the current runtime.
mleap-runtime/src/main/scala/ml/combust/mleap/runtime/serialization/bundle/ops · high confidence
Removed obsolete Bundle.ML serialization classes
The serializer package in bundle-ml has been cleaned up by removing the legacy serialization infrastructure. Specifically, the following files and classes have been deleted: BundleContext, BundleSerializer, ModelSerializer, NodeSerializer, SerializationContext, SerializationFormat, and the entire custom serialization hierarchy (CustomSerializer, CustomType, and their JSON/Protobuf implementations). This removes the ability to serialize/deserialize Bundle.ML models using the old mixed JSON/Protobuf approach, likely in preparation for a new serialization strategy.
bundle-ml/src/main/scala/ml/bundle/serializer · high confidence
TensorFlow integration configuration added
A new reference.conf file was added to configure the TensorFlow transformer operations, specifically registering the TensorflowTransformerOp and the associated operations package in the MLeap registry.
mleap-tensorflow/src/main/resources · high confidence
Updated Spark example data
The Spark example in the repository has been updated with a new CSV file (spark-demo.csv) containing sample data for testing or demonstration purposes. This change modifies the example files to reflect current usage patterns.
examples · medium confidence
Test coverage
Add serialization test for CountVectorizer null vocabulary; Add test utilities and specs for the MLeap executor; Added ALS parity test; Added GrpcSpec and test utilities for the MLeap gRPC server; Added LinearSVC parity test; Added PySpark feature tests for MathUnary, MathBinary, and StringMap; Added Python unit tests for sklearn feature extraction and tree models; Added Python unit tests for sklearn model serialization and deserialization; Added agaricus dataset in CSV format to work around libsvm loading bug; Added and standardized regression model tests; Added comprehensive test coverage for MLeap runtime components; Added comprehensive test coverage for TensorFlow integration; Added comprehensive unit tests for MLeap Tensor equality and legacy indexing; Added comprehensive unit tests for MLeap core feature models; Added integration tests for MLeap Spring Boot scoring endpoints; Added integration tests for MLeap executor components; Added parity test for Support Vector Machine (SVM) classification; Added parity tests for CrossValidator and TrainValidationSplit; Added parity tests for Spark ML classification algorithms; Added parity tests for Spark ML feature transformers; Added parity tests for XGBoost classification and regression models; Added parity tests for multiple regression models; Added test coverage for clustering model schemas; Added test coverage for sklearn PolynomialFeaturesModel; Added test data for XGBoost integration; Added test infrastructure for PySpark; Added test suite for Bundle.ML serialization and file system operations; Added tests for ALS model prediction and schema; Added tests for Avro serialization of LeapFrames and Rows; Added tests for Gensim Word2Vec integration; Added tests for multi-output transformer support; Added tests for sklearn extension components; Added tests for sklearn preprocessing transformers; Added tests for type casting and struct type operations; Added unit tests for MleapReflection utility; Added unit tests for classification model schemas; Added unit tests for vector conversion utilities.
Dependencies
Build system overhauled for Scala 2.13, Java 17, and Spark 4.1.2
The project's build infrastructure has been significantly upgraded to support Scala 2.13.18, Java 17, and Spark 4.1.2. This includes migrating from sbt 0.13 to 1.12.5, updating core dependencies such as Spring Boot 3.2.0, XGBoost 2.0.3, and TensorFlow Java 1.0.0, and introducing new build configuration files (BuildInfo, DockerConfig, MleapProject) to manage the new module structure and publishing settings.
project · high confidence
MLeap 0.25.2 release and dependency upgrades
This release upgrades the project to Spark 4.1.2, TensorFlow 2.16.2, and XGBoost 2.0.3, while also updating the Scala version to 2.13.18 and Java to 17. The release notes document these changes, and the README is updated to reflect the new dependency compatibility matrix and installation instructions for PySpark integration.
(repo-wide) · high confidence
MLeap 0.25.2: Upgrade to Spark 4.1.2
The project has been upgraded to Spark 4.1.2 (MLeap 0.25.2). This update also includes a reorganization of the build configuration, moving module definitions into a centralized \MleapProject\ object and introducing individual \build.sbt\ files for each module to standardize dependency and settings management.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 45 → 58 (+13.5)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 100 → 92 (-8.2)
- Architecture 94 → 97 (+2.8)
- Maturity 52 → 57 (+4.2)
- Readiness 27 → 44 (+17.3)
- Security 48 → 76 (+28.2)
Resolved (40)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- High IaC: DS-0002 (.devcontainer/Dockerfile)
- High IaC: DS-0029 (.devcontainer/Dockerfile)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- …and 20 more
New (114)
- Casting.cast (cognitive 61) (mleap-core/src/main/scala/ml/combust/mleap/core/types/Casting.scala)
- Casting.cast (cyclomatic 77) (mleap-core/src/main/scala/ml/combust/mleap/core/types/Casting.scala)
- Concentrated knowledge decay
- Documentation: no contributor guidance (README.md)
- Duplicated block (10 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/classification/DecisionTreeClassifierModel.scala)
- Duplicated block (10 lines × 2) (mleap-grpc/src/main/scala/ml/combust/mleap/grpc/GrpcClient.scala)
- Duplicated block (12 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame/IndexedRowUtil.scala)
- Duplicated block (12 lines × 2) (mleap-spark/src/main/scala/org/apache/spark/ml/bundle/ops/tuning/CrossValidatorOp.scala)
- Duplicated block (15 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/runtime/frame/IndexedRowUtil.scala)
- Duplicated block (15 lines × 2) (mleap-spark-base/src/main/scala/org/apache/spark/ml/bundle/SparkShape.scala)
- Duplicated block (15–16 lines × 2) (mleap-executor/src/main/scala/ml/combust/mleap/executor/service/LocalTransformService.scala)
- Duplicated block (17–18 lines × 2) (mleap-databricks-runtime-testkit/src/main/scala/ml/combust/mleap/databricks/runtime/testkit/TestSparkMl.scala)
- Duplicated block (18 lines × 2) (mleap-executor/src/main/scala/ml/combust/mleap/executor/service/LocalTransformService.scala)
- Duplicated block (21–22 lines × 2) (mleap-databricks-runtime-testkit/src/main/scala/ml/combust/mleap/databricks/runtime/testkit/TestSparkMl.scala)
- Duplicated block (5 lines × 2) (mleap-xgboost-spark/src/main/scala/ml/dmlc/xgboost4j/scala/spark/mleap/XGBoostClassificationModelOp.scala)
- Duplicated block (6 lines × 2) (bundle-ml/src/main/scala/ml/combust/bundle/tree/cluster/NodeSerializer.scala)
- Duplicated block (6 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/feature/FeatureHasherModel.scala)
- Duplicated block (6 lines × 2) (mleap-core/src/main/scala/ml/combust/mleap/core/feature/InteractionModel.scala)
- Duplicated block (6 lines × 2) (mleap-runtime/src/main/scala/ml/combust/mleap/bundle/tree/clustering/MleapNodeWrapper.scala)
- Duplicated block (6 lines × 2) (mleap-spark-base/src/main/scala/org/apache/spark/sql/mleap/TypeConverters.scala)
- …and 94 more
Architecture
- Containers 0 added · 0 removed · contexts 12 added · 0 removed · edges 35 added · 0 removed
Added bounded contexts (12)
- bundle-ml
- mleap-avro
- mleap-base
- mleap-core
- mleap-executor
- mleap-runtime
- mleap-spark-base
- mleap-spark-extension
- mleap-spark-testkit
- mleap-tensor
- mleap-tensorflow
- mleap-xgboost-runtime
Added dependency edges (35)
- bundle-ml → mleap-core (coupling)
- bundle-ml → mleap-runtime
- bundle-ml → mleap-tensor
- mleap-avro → mleap-core (coupling)
- mleap-avro → mleap-runtime (coupling)
- mleap-base → mleap-core
- mleap-core → bundle-ml (coupling)
- mleap-core → mleap-tensor
- mleap-executor → bundle-ml (coupling)
- mleap-executor → mleap-core (coupling)
- mleap-executor → mleap-runtime (coupling)
- mleap-runtime → bundle-ml (coupling)
- mleap-runtime → mleap-core (coupling)
- mleap-runtime → mleap-tensor
- mleap-spark-base → bundle-ml (coupling)
- mleap-spark-base → mleap-core (coupling)
- mleap-spark-base → mleap-runtime (coupling)
- mleap-spark-extension → bundle-ml (coupling)
- mleap-spark-extension → mleap-core (coupling)
- mleap-spark-extension → mleap-runtime (coupling)
- …and 15 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
combust/mleap was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 88de54dd4fb7720c7c56ff5c2b51cdec4353a128 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.