Skip to content
CAI
Software that uses CAICheck a score

microsoft/SynapseML

75.8

Strong · 27 September 2026

62.2k

lines of production code

Scala

with Python

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

SynapseML is a library that extends Apache Spark with machine learning, deep learning, and cognitive services capabilities. It provides transformers and estimators for integrating Azure AI services, including OpenAI, speech, vision, and search, directly into Spark DataFrames. The system also supports causal inference, recommendation engines, and responsible AI tools for model explainability and fairness analysis.

How it got here

2017–2021 — SynapseML rebranding and feature expansion

103 changes.

This period marked the comprehensive rebranding of the project from MMLSpark to SynapseML, including namespace migrations, website overhauls, and dependency upgrades to Spark 3.5. The team significantly expanded the library's capabilities by introducing new modules for AutoML, cybersecurity, image processing, and responsible AI, while simultaneously hardening the codebase with secure deserialization and robust testing infrastructure.

2022–2023 — Fabric integration and cognitive services expansion

45 changes.

This period focused on integrating Microsoft Fabric support, including platform detection, telemetry, and token management, while significantly expanding the cognitive services module with new features for OpenAI, speech, translation, and form analysis. The work also introduced causal inference estimators, deep learning classifiers, and a global parameter system, accompanied by extensive test coverage and infrastructure improvements.

2024–2026 — Feature expansion and test coverage

28 changes.

This period focused on integrating new AI capabilities, including HuggingFace transformers, Azure AI Foundry support, and GPU-accelerated RAPIDS environments, while deprecating legacy OpenAI APIs. Concurrently, the team significantly broadened test coverage across core, cognitive, and LightGBM modules to ensure robustness and correctness of these new features and existing infrastructure.

Features

Add Azure AI Foundry chat completion support

Users can now use the AIFoundryChatCompletion component to interact with Azure AI Foundry models. This new class extends the existing OpenAI chat completion logic, allowing users to specify a custom service name which is used to construct the endpoint URL (e.g., https://\<service\>.services.ai.azure.com/). It inherits standard text generation parameters like model, maxTokens, temperature, and others, enabling compatibility with reasoning models via maxCompletionTokens.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/aifoundry · high confidence

Add HTTP streaming source and sink for real-time data ingestion and response

Introduces new HTTP streaming components (HTTPSource, HTTPSink, and their V2 implementations) that allow Spark Structured Streaming to ingest data via HTTP requests and send responses back to the originating service. This enables real-time, bidirectional HTTP-based data pipelines within Spark streaming applications.

core/src/main/scala/org/apache/spark/sql/execution · high confidence

Add Isolation Forest anomaly detection estimator and model

Users can now perform anomaly detection using the Isolation Forest algorithm via the new \IsolationForest\ estimator and \IsolationForestModel\ classes in the SynapseML core library. This feature wraps the LinkedIn Relevance Isolation Forest implementation, allowing users to fit models on datasets and transform data to identify anomalies, with support for standard Spark ML parameter handling and prediction column configuration.

core/src/main/scala/com/microsoft/azure/synapse/ml/isolationforest · high confidence

Add Python ModelDownloader API for CNTK models

Introduces a new Python module in the downloader package that provides a \ModelDownloader\ class and \ModelSchema\ data structure. This allows users to browse, download, and manage CNTK pretrained models from the Microsoft server (defaulting to the Azure Blob storage URL) directly within PySpark, wrapping the existing Java implementation for seamless interoperability.

core/src/main/python/synapse/ml/downloader · high confidence

Add Python wrappers and type stubs for UDFTransformer and stage utilities

This change introduces Python wrapper classes and type stubs for the UDFTransformer stage, enabling users to apply custom Python User Defined Functions (UDFs) to DataFrame columns within Spark ML pipelines. It also adds Python wrappers for EnsembleByKey and LumpFeaturesModel, providing methods to retrieve column names and inspect which categorical values a model has retained for auditing purposes.

core/src/main/python/synapse/ml/stages · high confidence

Add Python wrappers for Difference-in-Differences and Double Machine Learning models

New Python model classes, DiffInDiffModel and DoubleMLModel, are now available in the synapse.ml causal module. DiffInDiffModel exposes detailed estimation results including treatment effects, standard errors, intercepts, weights, RMSE, and loss history. DoubleMLModel provides access to average treatment effects, confidence intervals, and p-values, with the latter fix ensuring the getPValue method is correctly exposed on the Python side.

core/src/main/python/synapse/ml/causal · high confidence

Add Spark container images for Kubernetes deployment

This change introduces new Dockerfiles and configuration files for running Apache Spark (versions 2.4.3 and 2.4.5) on Kubernetes. The \Dockerfile\ and \mini.Dockerfile\ set up the Spark and Hadoop environments, install necessary Python libraries (numpy, pandas, etc.), and configure the base image to use OpenJDK 11. Supporting scripts (\start-master\, \start-worker\, \start-common.sh\) and configuration files (\core-site.xml\, \spark-defaults.conf\, \log4j.properties\) are added to handle cluster initialization, GCS integration, and logging, enabling users to deploy Spark workloads via Helm charts.

tools/helm/spark · high confidence

Add automated version bump script with context-aware substitution

A new Python script, \scripts/bump-version.py\, has been added to automate version updates across the SynapseML codebase. The script uses strict context anchoring to ensure version strings are only replaced when they appear in specific, known SynapseML contexts (such as Maven coordinates, Docker tags, or documentation paths), preventing accidental changes to unrelated version numbers. A comprehensive test suite (\scripts/test\_bump\_version.py\) is included to verify the script's regex logic and safety constraints.

scripts · high confidence

Add optimized KNN fitting utilities

New utility classes \MetadataUtilities\ and \OptimizedCKNNFitting\ have been added to the \org.apache.spark.sql.types.injections\ package. \MetadataUtilities\ provides a helper to extract keys from Spark SQL metadata, while \OptimizedCKNNFitting\ introduces traits for optimized fitting of Conditional K-Nearest Neighbors (CKNN) and standard KNN models using Ball Trees, leveraging Breeze for vector operations and SynapseML logging.

core/src/main/scala/org/apache/spark/sql/types · high confidence

Add plotting utilities and CNTK model Python wrappers

This change introduces new Python modules for the library. In the core plotting package, it adds a \plot.py\ module providing functions to visualize model performance, specifically generating confusion matrices and ROC curves from Spark DataFrames using scikit-learn and matplotlib. In the deep-learning CNTK package, it adds a \CNTKModel.py\ wrapper class that bridges Python and Java, enabling users to configure model inputs, outputs, and feed/fetch dictionaries for CNTK models within the Spark ML pipeline.

core/src/main/python/synapse/ml/plot, deep-learning/src/main/python/synapse/ml/cntk · high confidence

Added Certified Event Client for Fabric telemetry

A new CertifiedEventClient has been introduced to handle logging of certified events specifically when running on the Fabric platform. This client constructs a telemetry payload containing the timestamp, feature name, activity name, and optional attributes, and posts it to the Fabric telemetry endpoint via the FabricClient. This enables structured usage logging for certified events within the Fabric environment.

core/src/main/scala/com/microsoft/azure/synapse/ml/logging/fabric · high confidence

Added GPU-accelerated RAPIDS ML initialization script

A new initialization script, init-rapidsml-cuda-11.8.sh, has been added to configure a Databricks environment with NVIDIA RAPIDS libraries. This script installs CUDA 11.8, the Spark RAPIDS ML plugin (version 24.6.0), and Python packages including cuDF, cuML, PyLibRAFT, and RMM (version 24.4.0) to enable GPU-based machine learning workloads.

_tools/init\scripts · high confidence

Added Horovod installation script for distributed deep learning

A new shell script (horovod\_installation.sh) has been added to automate the setup of Horovod for distributed training on Ubuntu 20.04. The script installs specific versions of PyTorch Lightning, Torchvision, Transformers, and Petastorm, configures NVIDIA CUDA and NCCL libraries, and builds Horovod from source with PyTorch support enabled.

deep-learning/src/main/python · high confidence

Added Python wrappers for SAR recommendation models and ranking validation

New Python classes have been introduced in the recommendation module to expose specific model capabilities and tuning workflows. SARModel now includes methods for recommending items to all users, specific user subsets, all items, and specific item subsets. Additionally, RankingTrainValidationSplit and its corresponding model class have been added to support ranking-based model evaluation and tuning, bridging the Python API with the underlying Java implementation.

core/src/main/python/synapse/ml/recommendation · high confidence

Added Sphinx documentation configuration and index for SynapseML

The documentation build system now includes a new Sphinx configuration file (conf.py) and an index page (index.rst) for the Python module. This setup configures the 'sphinx\_rtd\_theme' for HTML output, enables extensions like autodoc and napoleon for API reference generation, and establishes intersphinx links to external libraries such as PyTorch, NumPy, and PyTorch Lightning. The index page serves as the entry point, linking to the generated module documentation and the existing Scala API docs.

core/src/main/python/synapse/doc · high confidence

Added VS Code Dev Container configuration with multi-language support

Developers can now use a pre-configured VS Code Dev Container environment that automatically sets up a Ubuntu-based workspace with Node.js, Azure CLI, Conda (including JupyterLab support), and Scala/SBT tooling via SDKman. The container installs specific VS Code extensions such as GitHub Copilot, Metals, Scala, Java debug/test tools, and Jupyter, and runs a post-creation script to initialize Conda environments and SBT setup, streamlining the local development setup.

.devcontainer · high confidence

Added Vagrantfile for Windows developer setup

A new Vagrantfile has been added to the tools/vagrant directory to support Windows developer environments. It provisions an Ubuntu 18.04 VM with necessary dependencies including Java 8, SBT, and various system libraries, and automatically installs IntelliJ IDEA and the SynapseML project to streamline local development.

tools/vagrant · high confidence

Added local Spark initialization utility for Python development

A new \init\_spark\ function has been added to the \synapse.ml.core\ module to streamline local Python development. This utility automatically configures a Spark session with necessary dependencies (including \synapseml\ and \spark-avro\), sets specific execution parameters like shuffle partitions and cross-join support, and points to the Maven repository, allowing developers to quickly spin up a test environment without manual configuration.

core/src/main/python/synapse/ml/core · high confidence

Added script to generate PyPI MFA QR code

A new Python script, \generate-pypi-mfa-qr.py\, has been added to the \tools/pypi\ directory. This utility retrieves the PyPI MFA URI secret from Azure Key Vault and generates a QR code image, facilitating the setup or rotation of multi-factor authentication for PyPI access.

tools/pypi · high confidence

Added verification scripts for Java-Python mapping and snapshot alignment

New tools have been added to the repository to proactively verify build integrity: \verify\_python\_mapping.py\ checks that Java stage classes can be correctly resolved to their corresponding Python wrapper classes, and \verify\_snapshot\_alignment.py\ ensures that the synapseml-core wheel's version aligns with locally available Maven aggregate POMs and module JARs. These scripts help developers and CI pipelines catch systemic mapping or packaging issues early, rather than discovering them piecemeal during test execution.

tools · high confidence

AnalyzeText now supports asynchronous long-running operations for summarization and other tasks

The AnalyzeText transformer in the language service package now supports asynchronous, long-running operations (LRO) for tasks including Extractive Summarization, Abstractive Summarization, Healthcare, Sentiment Analysis, Key Phrase Extraction, PII Entity Recognition, Entity Linking, Entity Recognition, and custom classification tasks. This change introduces a new \AnalyzeTextLongRunningOperations\ component that submits jobs to the Azure AI Language service and polls for completion, allowing users to handle larger or more complex text analysis workloads without blocking. The implementation includes new schema definitions for LRO job states and results, specific parameter handling for summarization (such as \sentenceCount\ and \summaryLength\), and updated API versioning to support the asynchronous job submission endpoint.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/language · high confidence

Automated JAR file preparation for ESRP release

A new Python script (tools/esrp/prepare\_jar.py) has been added to automate the preparation of JAR artifacts for ESRP release. The script locates local Ivy2 repository directories, flattens nested folder structures by moving files to the top level (handling name collisions), and renames files to append the detected version number, ensuring consistent naming conventions for release artifacts.

tools/esrp · high confidence

Automated gateway port and SSH configuration setup

A new setup script for the gateway tool has been added to automatically configure the firewall and SSH service. Running this script opens TCP ports 8000 through 10000 in the public zone and permanently applies the firewall rules. It also modifies the SSH daemon configuration to enable GatewayPorts and restarts the SSH service to apply the changes.

tools/gateway · high confidence

ICE Transformer adds flexible feature specification

The ICETransformer now accepts both simple strings and dictionaries for categorical and numeric features via setCategoricalFeatures and setNumericFeatures, allowing users to specify feature names directly or provide detailed parameter dictionaries for more granular control over individual feature explanations.

core/src/main/python/synapse/ml/explainers · high confidence

Initial Livy Helm chart Dockerfiles and configuration

This change introduces the foundational Dockerfiles (Dockerfile and mini.Dockerfile) and configuration files for the Livy service within the Helm chart. It establishes the runtime environment by pinning Spark 2.4.5 and Hadoop 3.3.4, installing necessary Python libraries (numpy, pandas, etc.), and building Livy from a specific Git commit. The entry also includes startup scripts for Spark master and worker nodes, configuration for Google Cloud Storage integration, and default settings for Livy impersonation and session timeouts.

tools/helm/livy · high confidence

Introduce Fabric doc generation pipeline

Added a new \docgen\ tool in \tools/docgen\ that automates the generation of documentation for the Fabric channel. The tool processes RST and Jupyter notebook inputs, converting them to Markdown while enforcing Fabric-specific requirements such as adding image alt text, removing locale information from URLs, and injecting required metadata headers. It supports configurable output structures (flat or hierarchy) and includes helpers for processing notebook cells and managing image resources.

tools/docgen · high confidence

Introduce Fabric integration client and token management

Added a new \FabricClient\ component in the \fabric\ package to handle Microsoft Fabric workspace connectivity, including automatic detection of workspace IDs, capacity IDs, and workload endpoints (ML, LLM, Cognitive, OpenAI) based on the environment and Private Endpoint settings. Introduced \FabricTokenParser\ to validate and decode JWT tokens for usage reporting, and \TokenLibrary\ to manage MWC token acquisition and cache invalidation via reflection against the Fabric runtime. Added \OpenAIFabricSetting\ to check tenant-level model allow/deny status before use, ensuring users are notified if default Fabric LLM models are restricted by admin policy.

core/src/main/scala/com/microsoft/azure/synapse/ml/fabric · high confidence

Introduce SynapseML unified logging infrastructure

A new centralized logging module (SynapseMLLogger) has been added to standardize how SynapseML emits telemetry and logs. This component automatically detects the execution environment (Synapse Internal, Fabric Python, or standard Synapse) to route logs appropriately, enriches log entries with environment-specific metadata (such as workspace and artifact IDs), and ensures consistent payload structures including protocol versioning and library identifiers for downstream consumption.

core/src/main/python/synapse/ml/core/logging · high confidence

Introduction of LightGBMBooster for native model management

A new LightGBMBooster class has been added to handle the native LightGBM booster pointer, providing robust lifecycle management for native memory and ensuring the native library is initialized when needed. This component introduces support for loading models from string representations, managing thread-local buffers for predictions and SHAP values, and exposing metadata such as the number of classes, features, and total models.

lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/booster · high confidence

Introduction of SynapseML-based recommendation helper traits

The recommendation module now includes a new \RecommendationHelper.scala\ file that integrates SynapseML components (such as \Wrappable\, \ArrayParamMapParam\, and \EvaluatorParam\) into the ALS recommendation logic. This change introduces new parameter traits (\RecEvaluatorParams\, \RankingTrainValidationSplitParams\) and model base traits (\BaseRecommendationModel\) that enable advanced hyperparameter tuning and evaluation capabilities for recommendation models, leveraging the SynapseML codegen and parameter infrastructure.

core/src/main/scala/org/apache/spark/ml/recommendation · high confidence

Launch of new SynapseML website with multi-version installation support

The SynapseML website has been rebuilt using Docusaurus, introducing a new homepage that highlights key capabilities like Cognitive Services, Deep Learning, and Responsible AI, alongside a dedicated videos page showcasing past conference keynotes. A significant backend change is the introduction of \installArtifacts.js\, which centralizes installation coordinates and explicitly adds support for Spark 4.0 and Spark 4.1 (using Scala 2.13 and Python 3.12/3.13) in addition to the existing Spark 3.5 support, ensuring users can find the correct Maven/PyPI artifacts for their specific runtime environment.

website/src/pages · high confidence

LightGBM training now supports IPv6 worker endpoints

Users can now run distributed LightGBM training on clusters configured with IPv6 addresses. The framework automatically detects when the network topology requires IPv6 and transparently bridges the traffic, allowing the native LightGBM library (which is IPv4-only) to communicate with peers over IPv6 without requiring changes to the native binary or manual network configuration.

lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm · high confidence

Native model serialization and prediction control in LightGBM Python wrappers

The LightGBM Python wrappers for classification, ranking, and regression models now support loading native LightGBM models directly from files or strings via new static methods (loadNativeModelFromFile, loadNativeModelFromString). Additionally, the mixin provides methods to save and retrieve the native model as a string, access booster metadata (such as number of classes, features, and iterations), and control prediction behavior by disabling shape checks via setPredictDisableShapeCheck.

lightgbm/src/main/python · high confidence

New ACR manifest cleanup tool with safe backup and deletion logic

Added a new Python utility in tools/acr/clean\_acr.py that manages Azure Container Registry (ACR) manifests by exporting them to Azure Blob Storage before deletion. The tool ensures safety by verifying blob existence, re-validating digests against metadata to prevent data loss, and handling duplicate entries idempotently. It includes retry logic for Azure CLI commands and enforces strict digest validation. Comprehensive unit tests in tools/acr/test\_clean\_acr.py verify authentication modes, retry behavior, name constraints, and the core backup-then-delete workflow.

tools/acr · high confidence

New AutoML hyperparameter tuning and model selection components

This change introduces the core Scala components for automated hyperparameter tuning and model selection in the AutoML pipeline. It adds \TuneHyperparameters\ to execute randomized grid search across multiple estimators with configurable parallelism, folds, and run counts, and \FindBestModel\ to evaluate trained models against a specified metric (such as accuracy, AUC, or RMSE) and select the best performer. The implementation includes \DefaultHyperparams\ providing predefined search ranges for standard Spark ML learners (Logistic Regression, Decision Trees, GBT, Random Forest, MLP, Naive Bayes), \HyperparamBuilder\ and distribution classes (\IntRangeHyperParam\, \DoubleRangeHyperParam\, etc.) to define the search space, and \EvaluationUtils\ to handle metric evaluation and model type detection.

core/src/main/scala/com/microsoft/azure/synapse/ml/automl · high confidence

New Cybersecurity anomaly detection and feature engineering components

This change introduces a new \synapse.ml.cyber\ package containing core utilities and algorithms for cybersecurity analytics. It adds \AccessAnomalyModel\ and \ComplementAccessTransformer\ for detecting access anomalies using collaborative filtering and complement set sampling, alongside feature engineering tools like \IdIndexer\, \MultiIndexer\, and \PerPartitionScalarScaler\ for transforming user and resource data. A \DataFactory\ is also included to generate synthetic training and test datasets for these models.

core/src/main/python/synapse/ml/cyber · high confidence

New Data Balance Analysis transformers for aggregate, distribution, and feature-level metrics

Added three new Spark ML transformers in the exploratory package to measure data balance and fairness: AggregateBalanceMeasure computes overall inequality using the Atkinson, Theil-L, and Theil-T indices; DistributionBalanceMeasure compares observed sensitive feature distributions against a reference (defaulting to uniform) using KL Divergence, Jensen-Shannon Distance, Wasserstein Distance, Infinity Norm, Total Variation, and Chi-Squared tests, with support for custom reference distributions; and FeatureBalanceMeasure calculates parity metrics (Demographic Parity, Pointwise Mutual Information, Sorensen-Dice, Jaccard Index, Kendall Rank Correlation, Log-Likelihood Ratio, and t-test) between pairs of feature values. These components introduce new capabilities for users to assess and quantify bias and balance in their datasets.

core/src/main/scala/com/microsoft/azure/synapse/ml/exploratory · high confidence

New DeepTextClassifier and DeepVisionClassifier for transfer learning

Users can now perform deep transfer learning and fine-tuning for text and image classification tasks using the new DeepTextClassifier and DeepVisionClassifier components. The text classifier integrates with Hugging Face transformers (specifically version 4.49.0) to support fine-tuning of sequence classification models, while the vision classifier leverages torchvision (version 0.14.1 or higher) to fine-tune a variety of backbone architectures including ResNet, VGG, EfficientNet, and Vision Transformer (ViT). Both classifiers provide Spark ML-compatible APIs that allow users to specify parameters such as the number of classes, learning rate, and the number of additional layers to train, simplifying the process of adapting pre-trained models to specific datasets.

deep-learning/src/main/python/synapse/ml/dl · high confidence

New Docker images for local SynapseML development and minimal environments

Added new Dockerfiles and supporting scripts in tools/docker/demo and tools/docker/minimal to simplify local experimentation with SynapseML. The demo image provides a pre-configured Jupyter Notebook environment with Spark 3.5.4, Python, and SynapseML 1.1.3, including an init script that automatically configures Azure storage connectors and loads the library. The minimal image offers a lighter alternative with just the core Spark and Python dependencies. Both images are based on Ubuntu 22.04 and include patches for Netty security vulnerabilities ([CVE redacted], [CVE redacted], [CVE redacted]) by upgrading from 4.1.96 to 4.1.118.

tools/docker · high confidence

New Document Translator and updated Text Translator API support

Added a new DocumentTranslator component that enables batch translation of documents from source URLs to target storage locations, supporting filtering by prefix/suffix, source/target language specification, and glossary integration. The existing TextTranslator has been updated to support the new Translator API version 2026-06-06 alongside the existing v3.0, with corresponding schema changes to handle the new response formats for text translation and transliteration.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/translate · high confidence

New FastVectorAssembler feature transformer

Added a new FastVectorAssembler transformer that efficiently combines input columns into a single vector column. This implementation is optimized for performance with large numbers of columns by filtering out non-nominal numeric data and enforcing an ordering where categorical columns must precede numeric ones to ensure correct attribute indexing for Spark learners.

core/src/main/scala/org/apache/spark/ml/feature · high confidence

New Form Ontology Learner and Transformer for structured schema inference

Users can now automatically infer a unified schema from form analysis results using the new FormOntologyLearner and FormOntologyTransformer components in the form service package. The Learner analyzes input data to merge field types into a consistent ontology (handling conflicts like String vs Double by preferring String, and merging nested structures), while the Transformer applies this inferred schema to cast raw form recognition outputs into strongly-typed, structured rows, simplifying downstream processing of document fields.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/form · high confidence

New HTTP and Binary file I/O infrastructure

This change introduces a new I/O layer for SynapseML, adding support for reading and writing binary files (including zip inspection) and a comprehensive HTTP client framework. Users can now use \spark.read.binary\ and \sparkSession.readBinaryFiles\ to ingest binary data, and leverage new \DataStreamReaderExtensions\ and \DataFrameReaderExtensions\ for \image\ and \binary\ formats. The HTTP subsystem provides \HTTPTransformer\, \JSONInputParser\, and \JSONOutputParser\ to facilitate sending HTTP requests and parsing responses within Spark pipelines, along with underlying client implementations (\HTTPClients\, \BaseClient\) and schema definitions (\HTTPSchema\) to handle request/response data structures.

core/src/main/scala/com/microsoft/azure/synapse/ml/io · high confidence

New HuggingFace transformers for causal language modeling and sentence embeddings

Users can now apply HuggingFace models directly within PySpark pipelines using two new transformers: HuggingFaceCausalLMTransform for text generation (chat and completion tasks) and HuggingFaceSentenceEmbedder for generating sentence embeddings. The causal language model transformer supports configuration of model parameters, device mapping, and caching, while the sentence embedder allows selection of runtime environments (CPU, CUDA, or TensorRT) and batch sizes, enabling optimized inference for embedding tasks.

deep-learning/src/main/python/synapse/ml/hf · high confidence

New ICE, LIME, and KernelSHAP explainers with image support

This change introduces a new suite of local model explainers to the SynapseML library, including Individual Conditional Expectation (ICE), Partial Dependence Plots (PDP), LIME, and KernelSHAP. The implementation adds support for tabular, vector, text, and image data types, with specific transformers like ImageLIME and ImageSHAP handling superpixel-based preprocessing for images. It also includes the underlying statistical components, such as Lasso and Least Squares regression, and a compatibility workaround for Breeze library versions to ensure stability across different Spark releases.

core/src/main/scala/com/microsoft/azure/synapse/ml/explainers · high confidence

New ImageFeaturizer and ONNXModel Python wrappers in synapse.ml.onnx

The \synapse.ml.onnx\ namespace now includes Python wrappers for ONNX model integration. \ImageFeaturizer\ allows users to load models from OnnxHub or local/HDFS locations and configure mini-batch sizes. \ONNXModel\ provides methods to inspect model inputs and outputs via \NodeInfo\, \TensorInfo\, \MapInfo\, and \SequenceInfo\ classes, enabling better introspection of ONNX model structures within Spark pipelines.

deep-learning/src/main/python/synapse/ml/onnx · high confidence

New ImageTransformer Python API for image processing

Added a new Python module \synapse.ml.opencv.ImageTransformer\ that exposes a PySpark ML transformer for common image processing operations. Users can now chain methods such as \resize\, \crop\, \centerCrop\, \colorFormat\, \blur\, \threshold\, \gaussianKernel\, \flip\, and \normalize\ on image data. The module also includes utility functions \toNDArray\ and \toImage\ to convert between image objects and NumPy arrays, with specific support for handling grayscale images (single channel) in \toNDArray\.

opencv/src/main/python · high confidence

New ONNX-based image featurization and model serving components

This release introduces a new ONNX integration in the deep-learning module, featuring an ImageFeaturizer that automatically resizes and processes image columns through pre-trained ONNX models (including those from the ONNXHub), and a general ONNXModel transformer that supports CPU/CUDA device selection, graph optimization levels, and batched inference. The implementation includes an ONNXHub client for discovering and caching models, shared parameter traits for feed/fetch dictionary mapping, and utilities for handling variable input shapes and boolean input types.

deep-learning/src/main/scala · high confidence

New OpenCV image augmentation and transformation capabilities

Users can now apply image augmentation and transformations directly within Spark pipelines using the new OpenCV integration. The \ImageSetAugmenter\ transformer allows datasets to be supplemented with flipped images (left-right or up-down) to improve model robustness, while the \ImageTransformer\ provides a suite of processing stages including resizing, cropping (standard and center), color format conversion, blurring, and thresholding. These features are implemented in Scala and rely on the OpenCV native library, which is automatically loaded via \OpenCVUtils\ when the components are used.

opencv/src/main/scala · high confidence

New Python API for AutoML hyperparameter tuning and model inspection

This change introduces new Python wrapper classes in the \synapse.ml.automl\ module to expose AutoML capabilities. Users can now define hyperparameter search spaces using \HyperparamBuilder\, \DiscreteHyperParam\, \RangeHyperParam\, \GridSpace\, and \RandomSpace\. Additionally, \BestModel\ and \TuneHyperparametersModel\ classes provide methods to retrieve the best model, scored datasets, evaluation results (including ROC curves), and metrics for all compared models, bridging the Python API to the underlying Java implementation.

core/src/main/python/synapse/ml/automl · high confidence

New Python IO utilities for HTTP serving, binary files, images, and Power BI

This change introduces a new \synapse.ml.io\ package that exposes Python wrappers for several I/O capabilities. It adds convenience methods to Spark DataFrames and readers/writers for HTTP serving (including standard, distributed, and continuous HTTP sources/sinks), binary file reading and streaming, and image loading from paths or byte strings. It also provides a \SimpleHTTPTransformer\ for URL configuration, HTTP request/response schema utilities, and methods to stream or write data directly to Power BI.

core/src/main/python/synapse/ml/io · high confidence

New SWIG utility classes for native array handling

A new \SwigUtils\ object and associated wrapper classes have been added to the LightGBM SWIG layer to facilitate data exchange between Scala and the native C++ library. This includes \FloatChunkedArray\, \DoubleChunkedArray\, and \IntChunkedArray\ for managing large datasets in chunks, as well as \FloatSwigArray\, \DoubleSwigArray\, \IntSwigArray\, and \LongSwigArray\ for direct manipulation of native memory arrays. These utilities provide methods for converting between Java/Scala arrays and native pointers, supporting operations like adding items, retrieving chunks, and coalescing data, which underpins the underlying data transfer mechanisms for LightGBM training and inference.

lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/swig · high confidence

New Zeppelin Docker images and MMLSpark example notebooks

This change introduces new Dockerfiles (Dockerfile and mini.Dockerfile) for the Zeppelin Helm chart, building Zeppelin 0.9.0-SNAPSHOT with Spark 2.4.5, Hadoop 3.2.1, and Scala 2.12. The images include MMLSpark 0.15 and several Python libraries (numpy, pandas, matplotlib, scikit-learn, pyarrow). Additionally, example Zeppelin notebooks for MMLSpark classification, simplification, and serving are added to the image, along with configuration files for Spark and GCS integration.

tools/helm/zeppelin · high confidence

New causal inference estimators for Difference-in-Differences, Double Machine Learning, and Synthetic Control

This update introduces several new causal inference capabilities to the SynapseML causal package. Users can now apply Difference-in-Differences (DiffInDiff) and Synthetic Difference-in-Differences estimators to panel data to estimate treatment effects using time and unit indices. Additionally, a Double Machine Learning (DoubleMLEstimator) is added, which uses cross-fitting with configurable treatment and outcome models (e.g., LogisticRegression, GBTRegressor) to estimate Average Treatment Effects (ATE) with confidence intervals. An Orthogonal Forest DML estimator is also included for heterogeneous treatment effect estimation using random forests. Supporting components include residual transformers and parameter traits to facilitate these new modeling workflows.

core/src/main/scala/com/microsoft/azure/synapse/ml/causal · high confidence

New core contract definitions for parameters and metrics

Added new Scala files in the core contracts package that define standard parameter traits (such as inputCol, outputCol, labelCol, and featuresCol) and metric data structures (including MetricData and TypedMetric). These changes establish a unified interface for model parameters and evaluation metrics, ensuring consistent column naming and metric reporting across the library's components.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/contracts · high confidence

New core environment utilities for file handling, native library loading, and package management

The \core/src/main/scala/com/microsoft/azure/synapse/ml/core/env\ package now includes new utility classes to support runtime operations. \FileUtilities\ provides helpers for joining paths, reading/writing files, and zipping folders. \NativeLoader\ (Java) enables the extraction and loading of native libraries from JAR resources based on the operating system. \PackageUtils\ centralizes Maven coordinates and repository URLs for the SynapseML package and its dependencies (like Avro and ONNX Protobuf), using build-time version information. \StreamUtilities\ adds safe resource management wrappers and a \ZipIterator\ for streaming zip entries.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/env · high confidence

New core utility library for SynapseML

The \core/src/main/scala/com/microsoft/azure/synapse/ml/core/utils\ package now provides a comprehensive set of internal utilities to support the library's functionality. This includes \SafeObjectInputStream\ and \ContextObjectInputStream\ to mitigate unsafe Java deserialization (CWE-502) by enforcing allowlists on deserialized classes, \AsyncUtils\ and \FaultToleranceUtils\ to handle concurrent operations and retry logic with timeouts, and \ClusterUtil\ to assist with Spark cluster topology detection. Additionally, the package introduces \BreezeUtils\ for seamless conversion between Spark and Breeze vector/matrix types, \ParamsStringBuilder\ for constructing command-line arguments for native libraries, \ModelEquality\ for comparing model parameters, and \SlicerFunctions\ to expose vector slicing capabilities as Spark SQL UDFs.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/utils · high confidence

New data transformation and batching stages added to core

This release introduces a suite of new transformation stages in the core library to enhance data manipulation capabilities. Users can now balance imbalanced datasets using the ClassBalancer estimator, which calculates and applies class weights. Data structure operations are expanded with Explode for array expansion, RenameColumn for column renaming, and DropColumns for selective removal. For complex schema handling, MultiColumnAdapter allows applying a unary pipeline stage across multiple columns, while LumpFeatures provides categorical lumping with top-K, min-count, and min-frequency controls. Batching and partitioning are improved with new batchers (DynamicBufferedBatcher, FixedBufferedBatcher, TimeIntervalBatcher) and transformers (MiniBatchTransformer variants, PartitionConsolidator) to optimize data flow. Additionally, EnsembleByKey enables averaging scores by key with robust column resolution, Cacher and Repartition offer control over data persistence and partitioning, and Lambda allows custom DataFrame transformations.

core/src/main/scala/com/microsoft/azure/synapse/ml/stages · high confidence

New dataset construction and streaming infrastructure for LightGBM

The LightGBM dataset layer has been refactored to support streaming execution and reference datasets. New components include DatasetAggregator for chunked, parallel processing of Spark rows into native arrays, SampledData for efficient feature sampling to initialize reference datasets, and ReferenceDatasetUtils for serializing and deserializing these reference datasets to enable streaming training. The refactoring also introduces robust validation for feature sizes and group columns, and adds support for handling initial scores and group columns within the chunked aggregation pipeline.

lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/dataset · high confidence

New featurize components for data cleaning, conversion, and indexing

This change introduces several new featurize modules to the SynapseML library. The \CleanMissingData\ estimator allows users to replace missing values in numeric columns using mean or median imputation, or in any supported type using a custom value. The \DataConversion\ transformer enables explicit type casting of columns (e.g., to boolean, integer, string, or date) and includes utilities to convert between categorical indices and original values via \ValueIndexer\ and \IndexToValue\. Additionally, \CountSelector\ provides a mechanism to drop vector indices with no nonzero data, emitting a zero-width sparse vector if no indices remain. These components collectively enhance the library's capabilities for preprocessing and cleaning tabular data before model training.

core/src/main/scala/com/microsoft/azure/synapse/ml/featurize · high confidence

New fluent API methods and template function for Spark DataFrames

Users can now use the new \mlTransform\ and \mlFit\ methods directly on PySpark DataFrames for a more fluent coding experience, and access a \template\ function in \synapse.ml.core.spark.functions\ to invoke JVM-based template processing from Python.

core/src/main/python/synapse/ml/core/spark · high confidence

New fluent DataFrame API and PEP3101 template support

Users can now use a more fluent syntax for Spark ML pipelines via implicit conversions on DataFrame, allowing direct calls like df.mlTransform(transformer) or df.mlFit(estimator) without explicit pipeline wrapping. Additionally, a new template function supports PEP3101-style string formatting (using {variable} syntax) within Spark SQL expressions, enabling easier dynamic string construction in DataFrame operations.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/spark · high confidence

New global parameter system and expanded complex param support

The library introduces a new GlobalParams mechanism that allows parameters to be defined once and shared across multiple models, automatically filling unset parameters with global defaults. Additionally, the param package now includes a comprehensive set of new complex parameter types—including DataFrame, Estimator, Evaluator, Model, Transformer, and various array/map variants—along with dedicated serialization and interoperability wrappers for Python, R, and .NET, enabling these complex types to be passed and persisted across language boundaries.

core/src/main/scala/com/microsoft/azure/synapse/ml/param · high confidence

New image processing transformers for superpixel segmentation and pixel unrolling

Added \SuperpixelTransformer\ and \UnrollImage\ components to the image processing module. \SuperpixelTransformer\ decomposes images into superpixel clusters based on configurable cell size and modifier parameters, outputting cluster data that can be used for masking operations. \UnrollImage\ converts image data (supporting both Spark image structs and binary formats) into dense vectors of doubles, enabling downstream machine learning tasks that require vectorized image inputs. These new transformers expand the library's capabilities for image feature extraction and transformation within Spark DataFrames.

core/src/main/scala/com/microsoft/azure/synapse/ml/image · high confidence

New nearest-neighbor search models using Ball Tree algorithms

Added new KNN and ConditionalKNN Spark ML estimators and models that use Ball Tree data structures for fast nearest-neighbor lookups. The KNN model finds the top-k nearest neighbors by maximum inner product, returning associated values and distances, while the ConditionalKNN variant adds a conditioner column to filter results by label. These components are implemented in core/src/main/scala/com/microsoft/azure/synapse/ml/nn, including supporting BallTree, ConditionalBallTree, BoundedPriorityQueue, and schema definitions, enabling users to perform efficient similarity searches on large datasets within Spark pipelines.

core/src/main/scala/com/microsoft/azure/synapse/ml/nn · high confidence

New off-policy evaluation metrics and VW hashing parameters

Users can now evaluate contextual bandit policies using new Spark SQL UDAFs: IPS, SNIPS, Cressie-Read, and Cressie-Read with confidence intervals. These are registered as SQL functions (snips, ips, cressieRead, cressieReadInterval, cressieReadIntervalEmpirical) via PolicyEvalUDAFUtil. Additionally, Vowpal Wabbit models now expose parameters to control hashing behavior, including numBits (default 30 in the new trait, 18 in base) and sumCollisions (default true), allowing users to configure how feature collisions are handled during training.

vw/src/main/scala · high confidence

New platform detection and secret management utilities

Added a new \Platform\ module that automatically detects the current execution environment (Synapse, Synapse Internal, Fabric Python, Binder, or Databricks) and provides platform-specific helpers. This includes a \find\_secret\ function that abstracts away the differences in secret retrieval across these platforms (using Azure Key Vault linked services for Synapse, DBUtils for Databricks, etc.) and a \materializing\_display\ function that ensures data is properly rendered in Synapse and Synapse Internal notebooks.

core/src/main/python/synapse/ml/core/platform · high confidence

New recommendation evaluation and ranking components

Added RankingAdapter, RankingEvaluator, RankingTrainValidationSplit, and RecommendationIndexer to the recommendation module. RankingAdapter wraps existing recommenders (ALS, SAR) to produce ranked lists with configurable minimum rating filters. RankingEvaluator introduces advanced ranking metrics including NDCG, MAP, precision/recall at K, diversity, mean reciprocal rank, and fraction concordant pairs. RankingTrainValidationSplit enables stratified train/validation splits with parallel hyperparameter tuning using the new evaluator. RecommendationIndexer provides string-to-integer indexing for user and item identifiers, supporting round-trip recovery of original string values.

core/src/main/scala/com/microsoft/azure/synapse/ml/recommendation · high confidence

New repository statistics collection tools added

Added new scripts in tools/misc to gather and upload GitHub and Docker Hub statistics for the project. The bash script (get-stats) and Node.js module (get-stats.js) both query the GitHub API for repository metrics (stars, forks, issues, pull requests, contributors) and Docker Hub metrics (pulls, stars, tags), then append this data to an Azure Storage blob. An SVG icon (mmlspark.svg) was also added to the directory.

tools/misc · high confidence

New schema utilities for binary files, images, and categorical data

The \core/src/main/scala/com/microsoft/azure/synapse/ml/core/schema\ package now includes new utilities to handle specific data types and metadata. \BinaryFileSchema\ defines a schema for binary file columns (path and bytes) with helper methods to extract data and check types. \ImageSchemaUtils\ provides schema definitions and checks for OpenCV-compatible image data. \Categoricals\ introduces robust support for categorical data, allowing users to set and get levels for string, numeric, and boolean columns while managing metadata in both MLlib and MML formats. Additionally, \DatasetExtensions\ adds helper methods for finding unused column names and retrieving column values as specific types, while \SparkSchema\ and \SchemaConstants\ standardize metadata handling for model predictions and labels.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/schema · high confidence

New speech service components and SDK integration

This change introduces a new set of speech processing components within the cognitive services package. It adds a new \SpeechToTextSDK\ transformer that leverages the Microsoft Cognitive Services Speech SDK for advanced recognition features, including support for FFmpeg-based audio stream ingestion, word-level timestamps, and audio recording. Additionally, it introduces new transformers for \SpeakerEmotionInference\ (annotating text with inferred emotion and style via SSML) and \TextToSpeech\ (synthesizing audio from text/SSML with configurable output formats). Supporting infrastructure includes new audio stream handling classes (\WavStream\, \CompressedStream\), API helpers for speaker profiling, and comprehensive JSON schemas for speech responses and errors.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/speech · high confidence

New text featurization components: MultiNGram, PageSplitter, and TextFeaturizer

Added new text processing transformers and a composite featurizer to the SynapseML library. MultiNGram allows extracting multiple n-gram lengths from text columns in a single step, while PageSplitter breaks long text strings into manageable chunks based on character limits and word boundaries. These are complemented by TextFeaturizer, a pipeline estimator that orchestrates tokenization, stop-word removal, n-gram generation, hashing, and IDF scaling, providing a unified interface for complex text feature engineering workflows.

core/src/main/scala/com/microsoft/azure/synapse/ml/featurize/text · high confidence

New utility classes for Spark internals and UDF handling

Added four new utility objects and classes to the \org.apache.spark.injections\ package to support internal Spark operations: \BlockManagerUtils\ for accessing the block manager, \RegressionUtils\ for identifying regression stages, \SConf\ for serializable Hadoop configurations, and \UDFUtils\ for unpacking and creating User Defined Functions using older API patterns.

core/src/main/scala/org/apache/spark/injections · high confidence

New website includes code syntax highlighting themes and redirect handling

The website now supports code syntax highlighting with two new Prism themes: GitHub and Monokai, which define specific color styles for code elements. Additionally, a new redirect utility has been added to handle URL redirections, including logic to preserve URL hashes or default to an 'about' page when redirecting from specific components.

_website/src/exports, website/src/plugins/prism\themes · high confidence

New website theme components for code snippets and documentation

The website now includes custom React components for rendering code samples and documentation tables. The \SampleSnippet\ component displays code blocks with syntax highlighting (using Prism themes) and a 'Copy' button, while \CodeSnippet\ provides a simpler code display. Additionally, \DocumentationTable\ generates links to Python and Scala API docs, \NotebookExamples\ lists available notebooks, and a custom \NotFound\ page guides users to the new documentation section.

website/src/theme · high confidence

Repository initialization with core configuration and documentation files

The repository has been initialized with essential configuration and documentation files. This includes \.dockerignore\, \.gitattributes\, and \.gitignore\ to manage build artifacts and line endings, along with \AGENTS.md\ and \CONTRIBUTING.md\ to guide contributors and automated agents on the branch model, coding standards, and submission process. A \LICENSE\ file (MIT) and \SECURITY.md\ (Microsoft Security Response Center reporting) are added for legal and security compliance. Development environment setup is supported by \environment.yml\ (Conda dependencies), \apt.txt\ (system packages), and \scalastyle-config.xml\ (Scala style rules). CI/CD is defined in \pipeline.yaml\ (Azure DevOps) and \codecov.yaml\ (coverage reporting), while \postBuild\ configures notebook startup scripts.

(repo-wide) · high confidence

Structured output via JSON Schema and usage tracking for OpenAI models

Users can now request strict structured output from OpenAI models by passing an inner JSON Schema map to the new \setResponseSchema\ method on \OpenAIPrompt\ and \OpenAIChatCompletion\; this automatically configures the \response\_format\ envelope. Additionally, \OpenAIPrompt\ and \OpenAIEmbedding\ now support a \usageCol\ parameter to capture token usage statistics, and \OpenAIPrompt\ exposes \responseIdCol\ to track response IDs when the store feature is enabled.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/openai · high confidence

Support for reading and writing image data sources via PatchedImageFileFormat

A new \PatchedImageFileFormat\ implementation has been added to the Spark ML image source module, enabling users to read image files from storage and write them back out. This component handles schema validation for image data, manages file system interactions for reading binary image content, and includes a retry mechanism to handle JVM flakiness during decoding. It also provides an output writer that converts internal image rows back into image files on the file system, supporting the \dropInvalid\ option to control how malformed images are handled during the read process.

core/src/main/scala/org/apache/spark/ml/source · high confidence

Website static assets and configuration files added

The website's static file directory now includes essential configuration and branding assets. A \.nojekyll\ file is added to prevent Jekyll processing of the site, and a \BingSiteAuth.xml\ file is included to verify the site with Bing. Additionally, new SVG images have been added for the main logo, multilingual support, open-source initiatives, scalability, and various notebook examples (LIME, Spark Serving, CNTK, Cognitive Services on Spark, and a visual word representation).

website/static · high confidence

Security

Secure model deserialization with restricted class loading

Model loading in SynapseML now uses a secure import mechanism that restricts dynamic class loading to trusted modules (pyspark. and synapse.ml.), preventing arbitrary code execution via crafted model metadata. This change patches JavaParams and DefaultParamsReader to use the new secure import logic, ensuring that only allowed classes are instantiated during model deserialization.

core/src/main/python/synapse/ml/core/serialize · high confidence

Secure model persistence with class-filtered deserialization and SparkSession-based I/O

The ML persistence layer in core/src/main/scala/org/apache/spark/ml has been rewritten to address security vulnerabilities in Java deserialization and to support modern Spark execution modes. Model loading and saving now use SparkSession APIs instead of RDD/SparkContext methods, enabling compatibility with Spark Connect and Databricks Unity Catalog shared access modes. Additionally, deserialization is now constrained by default using a class filter (SafeObjectInputStream), blocking arbitrary code execution from untrusted artifacts; a legacy compatibility switch (spark.synapseml.legacy.allowUnsafeJavaDeserialization) is provided for trusted older models, and a new Serializer abstraction manages these safe read/write operations for complex params and datasets.

core/src/main/scala/org/apache/spark/ml · high confidence

Architecture

Refactored build infrastructure with new SBT plugins and centralized dependency management

The build system has been restructured into modular SBT plugins to support multi-language code generation and publishing. New plugins include CodegenPlugin for generating Python, R, and test code; BlobMavenPlugin and PublishPlugin for handling artifact publishing to Azure Blob storage and Maven repositories; CondaPlugin for managing Python environments; and Secrets for secure, cached retrieval of publishing credentials from Azure Key Vault. Additionally, OnnxRuntimeDependency.scala centralizes the ONNX Runtime version to 1.17.3 to fix local inference issues on Spark 3.5 and macOS, and the build now uses SBT 1.10.11 with updated plugins like sbt-sonatype and sbt-pgp.

project · high confidence

Restructured Text Analytics and Computer Vision services into dedicated Scala packages

The Text Analytics and Computer Vision cognitive services have been reorganized into new, dedicated packages (\com.microsoft.azure.synapse.ml.services.text\ and \com.microsoft.azure.synapse.ml.services.vision\). This change introduces new source files for core service logic (e.g., \TextAnalytics.scala\, \ComputerVision.scala\) and their corresponding data schemas (e.g., \TextAnalyticsSchemas.scala\, \ComputerVisionSchemas.scala\, \OCRSchemas.scala\). For users, this represents a structural refactoring of the codebase that consolidates service-specific parameters, request/response models, and JSON serialization formats into their respective domains, without altering the external API surface of the transformers themselves.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/text · high confidence

Behavioural changes

Added Python license and SWIG pointer wrapper support

The Python core module now includes an MIT License file and a MANIFEST configuration to ensure proper packaging distribution. Additionally, the Scala LightGBM module introduces a new SWIG helper class, SwigPtrWrapper, which provides a method to retrieve the underlying native pointer address from SWIG pointer objects, facilitating lower-level memory management interactions.

core/src/main/python, lightgbm/src/main/scala/com/microsoft/lightgbm · high confidence

Azure Maps geospatial services now use query-parameter authentication for async polling

The geospatial components in the cognitive package have been refactored to correctly handle Azure Maps' asynchronous long-running operations. Previously, the polling mechanism attempted to extract authentication credentials from HTTP headers, which failed because Azure Maps supplies the subscription key and API version via query parameters. The new \MapsAsyncReply\ trait in \AzureMapsTraits.scala\ extracts these credentials from the request URI's query string and appends them to the polling endpoint URL, ensuring that async requests (such as those from \AddressGeocoder\ and \ReverseAddressGeocoder\) authenticate correctly during status checks. Additionally, new schema definitions (\AzMapsSearchSchemas.scala\, \AzMapsSpatialSchemas.scala\) and transformers (\CheckPointInPolygon\, \Geocoders\) have been added to support these services, with \CheckPointInPolygon\ now explicitly throwing an error as the underlying Azure Maps Spatial service has been retired.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/geospatial · high confidence

Azure Search integration migrated to 2026-04-01 API with profile-based vector schema and AAD auth

The Azure Search components in the cognitive services package have been refactored to support the latest Azure AI Search API version (2026-04-01) and its profile-based vector search schema. This change updates the default API version, introduces support for Azure Active Directory (AAD) authentication alongside existing subscription keys, and ensures backward compatibility by automatically translating legacy vector schemas to the new profile-based format when using newer API versions. Additionally, the index parsing logic now correctly handles complex objects like analyzers and CORS options as opaque JSON values to prevent deserialization failures, and supports scoring profiles in index definitions.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/search · high confidence

Backwards compatibility for deprecated mmlspark namespace

Importing from the legacy 'mmlspark' namespace now triggers a deprecation warning and transparently redirects to 'synapse.ml', ensuring existing code continues to function while encouraging migration to the new package name.

core/src/main/python/mmlspark · high confidence

CI hardening: resilient sbt bootstrap and selective notebook E2E execution

The CI system now includes a resilient sbt bootstrap wrapper (\sbt\_retry.sh\) that mitigates Maven Central rate limits and recovers from unusable Ivy caches via jittered retries and targeted module eviction, supported by a prewarm job and cache templates. Additionally, a new impact analysis script (\e2e\_impact.py\) selectively skips Databricks CPU/GPU and Fabric notebook E2E jobs for PRs that only modify non-runtime paths (such as governance files, documentation, or test-only sources), while strict validation scripts ensure Python version pins are exact and internal typing compatibility is correctly patched.

tools/ci · high confidence

Centralized metric constants and deterministic model metadata resolution

The metrics module now uses a new \MetricConstants\ object to define all regression and classification metric names and their corresponding output column names, ensuring consistent naming across the library. Additionally, \MetricUtils\ has been updated to resolve scored model metadata deterministically by validating label and prediction column metadata, which prevents ambiguous or conflicting model identification during evaluation.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/metrics · high confidence

Deprecation of legacy OpenAI Completions API and expansion of OpenAI configuration options

The legacy OpenAI Completions API is no longer supported; attempting to use OpenAICompletion now raises a RuntimeError and directs users to switch to OpenAIResponses, OpenAIChatCompletion, or OpenAIPrompt with setApiType('chat\_completions') or setApiType('responses'). Additionally, OpenAIDefaults has been updated to support new configuration parameters including embedding\_deployment\_name, top\_p, seed, verbosity, reasoning\_effort, and api\_type, allowing users to configure these specific model and API behaviors.

cognitive/src/main/python/synapse/ml/services/openai · high confidence

Deprecation of synapse.ml.cognitive package in favor of synapse.ml.services

The \synapse.ml.cognitive\ package has been restructured into \synapse.ml.services\. All imports from the old \cognitive\ namespace (e.g., \synapse.ml.cognitive.face\, \synapse.ml.cognitive.text\) now trigger deprecation warnings and redirect users to the corresponding \synapse.ml.services\ modules. Additionally, the \synapse.ml.cognitive.anomaly\ and \synapse.ml.cognitive.bing\ modules have been removed entirely because the Azure AI Anomaly Detector and Bing Search API v7 services have been retired; users should migrate to \synapse.ml.isolationforest\ for anomaly detection.

cognitive/src/main/python/synapse/ml/cognitive · high confidence

Enhanced platform detection and SAS token scrubbing in logging

The logging subsystem now includes a new \PlatformDetails\ module that accurately identifies the runtime environment (Synapse Internal, Synapse, Databricks, Binder, or Unknown) and reports versioned Fabric runtime metadata (e.g., \fabric\spark\\<version\>\ or \fabric\python\\<version\>\) when running on Fabric. Additionally, a new \SASScrubber\ component has been added to automatically redact Shared Access Signature (SAS) tokens from log messages, preventing sensitive credentials from being exposed in logs.

core/src/main/scala/com/microsoft/azure/synapse/ml/logging/common · high confidence

Introduces structured telemetry logging with certified event support

The logging infrastructure in the core module has been replaced with a new system that emits structured JSON logs containing model metadata, execution timing, and error details. This change introduces a \SynapseMLLogging\ trait and \FeatureNames\ constants to standardize how operations like fit and transform are recorded. A key addition is the integration with certified events, allowing specific feature usage to be logged asynchronously to an external telemetry system. The system also automatically scrubs sensitive information (such as SAS tokens) from log messages and captures environment context like workspace and lakehouse IDs from the Spark session configuration.

core/src/main/scala/com/microsoft/azure/synapse/ml/logging · high confidence

Introduction of ComplexParam for secure serialization of complex types

A new ComplexParam class has been added to the serialization module to handle parameters containing complex objects that cannot be JSON-encoded. This change introduces a deserialization class filter mechanism to constrain Java object deserialization, addressing security concerns related to untrusted legacy data during load operations. Users relying on complex parameter types will now have their data serialized and deserialized through this safer, filtered pathway.

core/src/main/scala/com/microsoft/azure/synapse/ml/core/serialize · high confidence

New Python schema utilities for Java interoperability and service parameter handling

Added new Python modules (TypeConversionUtils.py, Utils.py) in the synapse.ml core schema package to handle type conversion and parameter binding between Python and Java. These utilities introduce support for complex type conversion, secure class importing to prevent RCE vulnerabilities via \_\import\\_, and proper handling of ServiceParam bindings, ensuring that Python wrappers correctly transfer parameters to and from their Java counterparts during ML pipeline operations.

core/src/main/python/synapse/ml/core/schema · high confidence

New base service infrastructure for authentication and header handling

The cognitive services module now includes a new \CognitiveServiceBase\ trait that centralizes service parameter management, introducing support for Azure Active Directory (AAD) token authentication (\AADToken\), custom authorization headers (\CustomAuthHeader\), and global subscription key parameters. This change also adds \CognitiveServiceSchemas\ for common data structures like \Rectangle\ and \CognitiveServiceHeaderValues\ for validating and processing header service parameters, providing a unified foundation for authentication and request configuration across cognitive services.

cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services · high confidence

Python bindings for Vowpal Wabbit and Policy Evaluation moved to SynapseML

The Python source files for Vowpal Wabbit estimators and models (including classification, regression, contextual bandit, and generic variants) as well as the Policy Evaluation utilities have been relocated from the Vowpal Wabbit module into the SynapseML package under the \synapse.ml\ namespace. This change updates the import paths for these components, reflecting the rebranding of the library from mmlspark to synapseml, while preserving the existing functionality for training, prediction, and native model management.

vw/src/main/python · high confidence

Refactored LightGBM parameter system and streaming as default transfer mode

The LightGBM parameter system has been restructured into modular case classes (GeneralParams, DatasetParams, ExecutionParams, etc.) within the new BaseTrainParams hierarchy, improving organization and maintainability. As part of this refactor, the default data transfer mode for LightGBM has changed from bulk to streaming, which may affect performance characteristics for existing users who relied on the previous default. Additionally, the deprecated executionMode parameter has been fully removed in favor of dataTransferMode, and new parameters such as maxStreamingOMPThreads and referenceDataset have been introduced to support streaming execution and custom sampling modes.

lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/params · high confidence

Refactored code generation system to support Python and R bindings

The code generation infrastructure in the core module has been restructured to introduce dedicated generators for Python (PyCodegen) and R (RCodegen), replacing the previous monolithic approach. This change adds support for generating Python type stubs (.pyi) to improve IDE autocomplete and type checking, and introduces a PythonInitMerger to preserve hand-written \_\init\\_.py files during code generation. It also implements specific handling for the deprecated OpenAICompletion class to ensure backward compatibility via import hooks, and adds R package generation capabilities including DESCRIPTION file creation and sparklyr integration.

core/src/main/scala/com/microsoft/azure/synapse/ml/codegen · high confidence

Refactored training and evaluation components in the core train module

The core training module has been refactored to introduce a unified base structure for auto-trained models and trainers, while also enhancing the validation and metric calculation logic for model evaluation. Specifically, \AutoTrainedModel\ and \AutoTrainer\ now provide common inheritance and parameter handling for classification and regression trainers, standardizing how underlying models and feature columns are managed. The \ComputeModelStatistics\ transformer has been updated to include stricter input validation via \ComputeModelStatisticsInputValidator\, ensuring that required columns (like \labelCol\, \scoredLabelsCol\, and \scoresCol\) exist and have correct types before calculating metrics such as AUC, accuracy, or MSE. Additionally, \ComputePerInstanceStatistics\ now explicitly calculates per-instance loss metrics (log\_loss for classification, L1/L2 loss for regression), and the \TrainClassifier\ and \TrainRegressor\ classes have been adjusted to work with this new base structure, ensuring consistent featurization and model fitting pipelines.

core/src/main/scala/com/microsoft/azure/synapse/ml/train · high confidence

Reorganized cognitive services into a dedicated services package

The cognitive module has been refactored to restructure its Python and Scala source files under a new \synapse.ml.services\ hierarchy. This change moves the Azure Search writer implementations (including \streamToAzureSearch\ and \writeToAzureSearch\ DataFrame extensions) into \synapse.ml.services.search\ and relocates the Face API components (such as \DetectFace\, \FindSimilarFace\, and associated data schemas) into \synapse.ml.services.face\. Users should update their imports to reflect this new package structure.

cognitive/src/main/python/synapse/ml/services, cognitive/src/main/python/synapse/ml/services/search, cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/face · high confidence

SynapseML website migrated to Docusaurus 3

The SynapseML documentation site has been rebuilt using Docusaurus 3, replacing the previous static site generator. This migration introduces a modernized user interface with a dark mode theme, Algolia-powered search, and an announcement bar. The site now supports versioned documentation (currently listing versions 0.11.3 through 1.1.3) and includes a new sidebar structure organizing content into categories like 'Explore Algorithms' and 'Responsible AI'. Additionally, the site integrates Google Analytics (Gtag) for telemetry and configures MathJax/KaTeX for rendering mathematical expressions in markdown.

website · high confidence

Fixes

IsolationForestModel now correctly exposes inner model parameters

The IsolationForestModel class now properly delegates calls for the inner model and the outlier score threshold to the underlying Java implementation. Previously, the generated Python wrapper failed to exchange parameters with the inner model, meaning users could not retrieve the outlier score threshold or access the inner model object directly. This fix ensures that these critical model attributes are accessible after fitting.

core/src/main/python/synapse/ml/isolationforest · high confidence

LangChain Transformer serialization and error handling improvements

The LangChain transformer in SynapseML now provides robust serialization support for saving and loading transformer instances, while explicitly preventing the persistence of chains that contain memory objects. The implementation includes enhanced error handling to prevent crashes when used with OpenAI client versions greater than 1.0.0, and introduces security measures to validate secret references and prevent unsafe imports during deserialization.

cognitive/src/main/python/synapse/ml/services/langchain · high confidence

Test coverage

Added Fabric integration test infrastructure and unit tests; Added Python service parameter binding and smoke tests; Added Python test package structure for SynapseML; Added Python tests for OpenAI multimodal, defaults, and structured output features; Added Python unit tests for Vowpal Wabbit components; Added R test runner and tag parsing tests; Added TestBase trait for Spark session management and test utilities; Added automated tests for documentation integrity and installation artifacts; Added benchmark data and test configuration for LightGBM; Added benchmark data for Vowpal Wabbit Regressor verification; Added benchmark test resources and verification data; Added benchmarking infrastructure for performance regression testing; Added comprehensive test coverage for Vowpal Wabbit components; Added fuzzing test infrastructure for generating Python unit tests; Added fuzzing tests to validate module loading, fitting, serialization, and Python interoperability; Added integration tests for HTTP streaming sources and sinks; Added streaming mode validation tests for LightGBM Classifier; Added test coverage for AnalyzeText and long-running operation language services; Added test coverage for Azure Search authentication, persistence, schema parsing, and index retention; Added test coverage for Cyber and Neural Network modules; Added test coverage for Form Recognizer and Face services; Added test coverage for ONNX image featurization, model inference, and runtime dependency configuration; Added test coverage for explainer components; Added test coverage for featurize components; Added test coverage for image processing utilities and superpixel transformers; Added test coverage for model training, evaluation, and statistics modules; Added test coverage guard to detect unexecuted test suites in CI; Added test infrastructure for secret management and token parsing; Added test suite for DeepVisionClassifier and DeepTextClassifier; Added test suites for Data Balance Analysis transformers; Added test suites for OpenCV image processing transformers; Added test suites for Speech services; Added test suites for binary/image I/O, HTTP transformers, and retry logic; Added tests for AIFoundryChatCompletion response format handling; Added tests for ComplexParam serialization and deserialization safety; Added tests for DistributionBalanceMeasure reference distribution handling; Added tests for EnsembleByKey column name handling; Added tests for ImageLIME and ImageSHAP explainers; Added tests for Isolation Forest prediction column configuration; Added tests for LightGBM booster parameter deserialization safety and bulk classifier mode; Added tests for LightGBM classifier streaming and bulk modes; Added tests for LightGBM model serialization, streaming validation, and raw prediction exposure; Added tests for LightGBM network recovery, IPv6 bridging, and task retry error handling; Added tests for LightGBM streaming and bulk data transfer modes; Added tests for Text Analytics service batching and credential handling; Added tests for codegen discovery, generation, and utilities; Added tests for empty featurized text handling, logging infrastructure, and template functions; Added tests for model equality and slicer functions; Added tests for preserving missing numeric rows in Featurize; Added tests for recommendation package exports and SAR string identifier support; Added unit and integration tests for the Translator service; Added unit tests for AutoML hyperparameter tuning and evaluation; Added unit tests for Azure Maps geospatial transformers and HTTP traits; Added unit tests for HTTP client, schema, and helper components; Added unit tests for HuggingFace CausalLM and Sentence Embedder transformers; Added unit tests for LangchainTransformer functionality; Added unit tests for LightGBM dataset utility functions; Added unit tests for causal inference components; Added unit tests for cognitive service base configuration and parameter handling; Added unit tests for core Spark fluent API and template parsing; Added unit tests for core data processing stages; Added unit tests for core environment utilities and explainer parameters; Added unit tests for core metrics, parameters, and I/O utilities; Added unit tests for core parameter types and global params; Added unit tests for core schema and ML feature utilities; Added unit tests for core utility classes; Added unit tests for logging infrastructure components; Added unit tests for metric constants and column mappings; Added unit tests for neural network components and serialization safety; Expanded test coverage for OpenAI services; New Databricks and Fabric end-to-end test infrastructure.

Dependencies

Upgrade to Spark 3.5 and update website dependencies

The SynapseML build system has been upgraded to target Apache Spark 3.5.0, replacing previous versions. Additionally, the website documentation has been updated to use Docusaurus 3.10.2 and requires Node.js 24.0 or higher, with corresponding updates to the npm dependency lock files to ensure compatibility and security.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 42 → 76 (+33.4)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 67 → 88 (+20.4)
  • Architecture 92 → 96 (+3.9)
  • Maturity 56 → 72 (+16.5)
  • Readiness 20 → 79 (+59.1)
  • Security 62 → 73 (+11.1)

Resolved (56)

  • Coverage not measured — test suite did not build
  • Decision and consequences are absent; the body is a marketing/announcement title with no explicit decision (e.g. API contract) or trade-offs (website/blog/2019-06-01-MMLSpark Unifying Machine Learning Ecosystems at Massive Scales.md)
  • Dimension evaluation failed
  • High IaC: DS-0002 (tools/docker/demo/Dockerfile)
  • High IaC: DS-0002 (tools/docker/minimal/Dockerfile)
  • High IaC: DS-0029 (tools/docker/demo/Dockerfile)
  • High IaC: DS-0029 (tools/helm/livy/Dockerfile)
  • High IaC: DS-0029 (tools/helm/livy/Dockerfile)
  • High IaC: DS-0029 (tools/helm/livy/Dockerfile)
  • High IaC: DS-0029 (tools/helm/spark/Dockerfile)
  • High IaC: DS-0029 (tools/helm/spark/Dockerfile)
  • High IaC: DS-0029 (tools/helm/zeppelin/Dockerfile)
  • High IaC: DS-0029 (tools/helm/zeppelin/Dockerfile)
  • High IaC: DS-0029 (tools/helm/zeppelin/Dockerfile)
  • High IaC: DS-0029 (tools/helm/zeppelin/Dockerfile)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • …and 36 more

New (332)

  • AnyJsonFormat.anyFormat (cyclomatic 25) (core/src/main/scala/com/microsoft/azure/synapse/ml/param/UntypedArrayParam.scala)
  • Base-context workflow trigger runs with an unscoped token
  • BulkPartitionTask.mergeChunksIntoAggregatedArrays (cognitive 16) (lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/BulkPartitionTask.scala)
  • CI installs an unverified third-party binary (.github/workflows/pr-validation.yml)
  • Change coupling clique: LightGBMClassifier.scala, LightGBMRanker.scala, LightGBMRegressor.scala (lightgbm/src/main/scala/com/microsoft/azure/synapse/ml/lightgbm/LightGBMClassifier.scala)
  • Change coupling: OpenAIChatCompletion.scala ↔ OpenAIResponses.scala (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/openai/OpenAIChatCompletion.scala)
  • ComputeModelStatistics.createConfusionMatrix (cognitive 17) (core/src/main/scala/com/microsoft/azure/synapse/ml/train/ComputeModelStatistics.scala)
  • ComputeModelStatistics.transform (cognitive 32) (core/src/main/scala/com/microsoft/azure/synapse/ml/train/ComputeModelStatistics.scala)
  • ComputeModelStatistics.transform (cyclomatic 22) (core/src/main/scala/com/microsoft/azure/synapse/ml/train/ComputeModelStatistics.scala)
  • ComputePerInstanceStatistics.transform (cognitive 24) (core/src/main/scala/com/microsoft/azure/synapse/ml/train/ComputePerInstanceStatistics.scala)
  • Concentrated knowledge decay
  • DictionaryExamples.inputFunc (cognitive 18) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/translate/TextTranslator.scala)
  • Duplicated block (10 lines × 2) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/geospatial/Geocoders.scala)
  • Duplicated block (10 lines × 2) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/openai/OpenAIChatCompletion.scala)
  • Duplicated block (10 lines × 2) (core/src/main/scala/org/apache/spark/sql/execution/streaming/DistributedHTTPSource.scala)
  • Duplicated block (10–11 lines × 2) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/language/AnalyzeText.scala)
  • Duplicated block (10–11 lines × 3) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/translate/TextTranslator.scala)
  • Duplicated block (11 lines × 2) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/openai/OpenAIChatCompletion.scala)
  • Duplicated block (11 lines × 2) (core/src/main/scala/org/apache/spark/sql/execution/streaming/DistributedHTTPSource.scala)
  • Duplicated block (12 lines × 2) (cognitive/src/main/scala/com/microsoft/azure/synapse/ml/services/openai/OpenAIChatCompletion.scala)
  • …and 312 more

Changes since last survey

  • 192 commits — 101 feature/other, 91 fixes

By area

  • lightgbm/src — 41 commits
  • cognitive/src — 37 commits
  • .github/workflows — 28 commits
  • core/src — 25 commits
  • (root) — 21 commits
  • website/package-lock.json — 6 commits
  • .github/skills — 5 commits
  • tools/ci — 5 commits
  • tools/docker — 4 commits
  • .pipelines/release-compat-prerequisites.txt — 3 commits
  • docs/Explore Algorithms — 3 commits
  • reviews/fabric-cleanup-relations-20260921 — 3 commits
  • reviews/pr-2728 — 2 commits
  • reviews/pr-2732 — 2 commits
  • website/versioned_docs — 2 commits
  • deep-learning/src — 1 commit
  • reviews/pr-2666 — 1 commit
  • reviews/pr-2708 — 1 commit
  • tools/helm — 1 commit
  • vw/src — 1 commit

Notable commits

  • fix: Avoid fixed ports in validation integration tests
  • fix: Reduce validation scaling regression cost
  • fix: fix(ci): align Docker Python with environment
  • fix: fix(ci): avoid retaining Pip caches in Docker images
  • fix: fix(ci): enforce exact Docker Python pins
  • fix: fix(ci): repair publishing and Cognitive service test credentials (#2691)
  • fix: fix(ci): retire synced release compatibility prerequisites
  • fix: fix(ci): support Bash 3.2 version extraction
  • fix: fix(core): include reference-only distribution categories (#2630)
  • fix: fix(lightgbm): clean up failed streaming datasets
  • fix: fix(lightgbm): close failed validation datasets
  • fix: fix(lightgbm): support IPv6 worker endpoints (#2637)
  • fix: fix(onnx): upgrade runtime for Spark 3.5 local inference (#2636)
  • fix: fix(openai): accept extensionless typed attachments
  • fix: fix(openai): clarify null Responses content
  • fix: fix(openai): harden malformed multimodal inputs
  • fix: fix(openai): harden multimodal request handling
  • fix: fix(openai): harden multimodal request handling
  • fix: fix(openai): normalize attachment types by root locale
  • fix: fix(openai): preserve Responses validation indices
  • …and 172 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

microsoft/SynapseML was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 6fdc83d3116793d874099233a59c3de4cb6f7166 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.