salesforce/TransmogrifAI
70.4
Strong · 27 September 2026
55.7k
lines of production code
Scala
primary language
4
measurements over time
What this system is
TransmogrifAI is a Scala library for building machine learning pipelines on Apache Spark, providing a high-level DSL to define features, select models, and evaluate results. It supports a wide range of classification and regression algorithms, including XGBoost, and offers utilities for feature engineering, automatic feature selection, and model explainability. The system enables workflow serialization via MLeap for local scoring and includes a CLI to scaffold projects and generate code from data schemas.
How it got here
2017 — Rebranding to TransmogrifAI and feature expansion
46 changes.
The project was rebranded from Optimus Prime to TransmogrifAI, accompanied by a major version release to 0.7.0 and integration with Spark 2.4.5 and MLeap. This period focused on expanding the core library with new ML models, serialization infrastructure, and data readers, while simultaneously enhancing the CLI and adding comprehensive test coverage for the new features.
2018–2019 — Local scoring and test coverage expansion
16 changes.
This period focused on introducing local model scoring capabilities to enable testing without a Spark environment, alongside the addition of custom serialization annotations. It was heavily characterized by extensive expansion of unit test coverage across core, features, and utility modules to ensure robustness and correct behavior.
Features
Added Iris and Boston housing regression examples
New example workflows have been added to the helloworld directory: an Iris multi-class classification example (OpIris) and a Boston housing regression example (OpBoston). These examples demonstrate how to define features, configure readers, and run workflows using TransmogrifAI's classification and regression model selectors.
helloworld/src/main/scala/com/salesforce/hw/iris · high confidence
Added Titanic passenger dataset test fixtures
Added test data files for the Titanic passenger dataset, including Avro schema definitions (PassengerDataAll.avsc, PassengerDataAll\_.avsc) and CSV samples (PassengerDataAll.csv, PassengerDataAllWithHeader.csv). These files provide structured test inputs for validating data reading and processing capabilities.
test-data · high confidence
Added sample datasets for housing, email, and iris examples
The helloworld example now includes bundled resource files for three distinct datasets: the Boston Housing dataset (housing.data, housing.names, and a CSV variant), an Email dataset containing click and send logs, and the Iris dataset (iris.data, iris.csv, and metadata). These resources provide the underlying data required to run the corresponding sample scenarios locally without needing external downloads.
helloworld/src/main/resources · high confidence
Expanded test data generation with random feature support and US/Canada area code validation
The testkit now includes a random test feature generator capable of creating datasets with features of all supported types, allowing users to generate arbitrary test data without manual construction. Additionally, the testkit imports a comprehensive NPA (area code) report CSV and exposes a list of valid US and Canadian area codes, enabling more realistic and geographically accurate test data generation for phone-related features.
testkit/src/main · high confidence
Initial release of TransmogrifAI Hello World examples
This change introduces the initial set of example workflows for the TransmogrifAI library, including a minimal Titanic classifier, a full Titanic classifier, an Iris multiclass classifier, a Boston housing regression, and data preparation examples for joins/aggregates and conditional aggregations. The entry provides the necessary project scaffolding (Gradle wrapper, build scripts, .gitignore), Avro schemas for input data, and comprehensive documentation on how to build and run these workflows via Gradle or spark-submit.
helloworld · high confidence
Introduce RawFeatureFilter for automated feature selection and ModelInsights for model explainability
This change introduces the \RawFeatureFilter\, a new stage that automatically analyzes raw feature distributions and correlations to identify and drop low-value or redundant features before model training, thereby simplifying the feature engineering pipeline. It also adds \ModelInsights\, a comprehensive reporting tool that provides detailed explanations of model performance, including feature importance, correlations, and evaluation metrics, allowing users to better understand and trust their model's predictions.
repository · high confidence
Introduction of local scoring capabilities
A new local scoring package has been added to the library, introducing the \OpWorkflowModelLocal\ trait and a \ScoreFunction\ type alias. This enables users to perform model scoring locally on raw records (represented as maps) without requiring a distributed Spark execution environment, facilitating faster iteration and testing workflows.
local/src/main · high confidence
New Jupyter notebook samples for TransmogrifAI workflows
Added Jupyter notebook examples (OpHousingPrices, OpIris, OpTitanicSimple) demonstrating TransmogrifAI workflows for housing price prediction, Iris classification, and Titanic survival prediction. These notebooks are configured to use TransmogrifAI 0.7.0 and Spark 2.4.5, providing users with ready-to-run code for feature engineering, model building, and evaluation.
helloworld/notebooks · high confidence
New Parquet and Streaming Readers with Refactored Reader Hierarchy
This update introduces new data readers for Parquet files, including \ParquetProductReader\, \AggregateParquetProductReader\, and \ConditionalParquetProductReader\, accessible via the \DataReaders\ factory. It also adds streaming capabilities with the \StreamingReader\ trait and \FileStreamingAvroReader\ for monitoring Hadoop-compatible filesystems. Under the hood, the reader hierarchy has been refactored: the \Reader\ trait now extends \ReaderType\ and \ReaderKey\, and several existing readers (Avro, CSV, Custom) no longer extend the \Reader\ trait directly. Additionally, CSV schema inference has been updated to use Spark's \CSVSchemaUtils\ and respects session configuration for column pruning.
readers/src/main/scala/com/salesforce/op/readers · high confidence
New Titanic classification example application
Added a new Helloworld example demonstrating a Titanic survival prediction workflow using TransmogrifAI. The example includes feature definitions for passenger attributes (such as class, sex, and age), an automated workflow that performs feature engineering, selection, and model selection between Logistic Regression and Random Forest via cross-validation, and the necessary Kryo registration for the Passenger data schema.
helloworld/src/main/scala/com/salesforce/hw/titanic · high confidence
New annotation for custom stage serialization
A new @ReaderWriter annotation has been added to the stages package, allowing users to specify custom reader and writer implementations for stage value classes. This enables customization of how stage arguments are serialized and deserialized by pointing to a class that extends ValueReaderWriter.
features/src/main/java · high confidence
New classification models and date vectorization stages
This update adds several new machine learning stages to the core library. It introduces wrappers for Spark ML classification algorithms, including Decision Tree, GBT, Linear SVC, Multilayer Perceptron, Naive Bayes, and XGBoost classifiers, allowing users to incorporate these models into their workflows. Additionally, it adds a \DateMapToUnitCircleVectorizer\ stage that transforms date or datetime fields into circular cartesian coordinates for specific time periods (e.g., hour, month), and includes new resource files (\GenderDictionary\_USandUK.csv\, \Names\_JRC\_Combined.txt\) to support name-based gender detection features.
core/src/main · high confidence
New code-generation templates for ML pipelines and feature types
The CLI's code-generation templates have been expanded to support a wider range of machine-learning scenarios and data types. New templates now generate code for binary, multi-class, and regression model selection with their corresponding evaluators, as well as handling for binary, integral, real, text, and categorical features. A new \SampleObject\ prototype replaces the previous \\GenericObject\\ to provide typed field accessors for these features, and the categorical template has been updated to use the new \FeatureOps\ helper for pick-list extraction.
cli/src/main/scala/com/salesforce/op/cli/gen/templates · high confidence
New serialization, aggregation, and sequence processing infrastructure
This change introduces foundational components for the features module: a new JSON-based serialization system for pipeline stages (including \DefaultOpPipelineStageReaderWriter\, \OpPipelineStageReader\, and \OpPipelineStageWriter\) that enables saving and loading stage configurations; new aggregation utilities (\CommutativeGroupAggregator\ and \ExtendedMultiset\) supporting commutative group operations and multisets with negative counters; and new base classes for sequence processing (\BinarySequenceEstimator\, \BinarySequenceTransformer\, and \BinarySequenceLambdaTransformer\) that allow transforming single features combined with sequences of other features. Additionally, it adds \SparkStageParam\ for managing Spark/MLeap stage persistence and \SparkWrapperParams\ for wrapping Spark ML stages with MLeap support.
features/src/main/scala · high confidence
Streaming histogram support and utility refactoring
This update introduces a new StreamingHistogram implementation for building histograms from data streams, along with Scala wrappers (RichStreamingHistogram) to expose bin and density estimation capabilities. It also renames the JaccardDistance utility to JaccardSimilarity to better reflect its function, removes legacy DirectOutputCommitter classes in favor of standard MapReduce committers, and adds several test infrastructure utilities including a new TestSparkStreamingContext trait and improved resource loading methods in TestCommon.
utils/src/main · high confidence
Behavioural changes
CLI project generation now supports overwriting existing projects and includes updated Titanic dataset examples
The CLI's \op gen\ command now accepts the \--overwrite\ flag (replacing the previous \--override\) to allow users to replace an existing project directory. Additionally, the included Titanic dataset example has been expanded with more feature definitions (such as Name, Sex, Ticket, Cabin, and Embarked) and explicitly sets the problem type to binary classification (\binclass\) for the 'Survived' response field, ensuring the generated project is correctly configured for this common use case.
cli · high confidence
CLI rebranded to TransmogrifAI with new command-line options
The command-line interface has been renamed from 'op' (Optimus Prime) to 'transmogrifai'. This change updates the CLI entry point in CliExec.scala and modifies the CommandParser to reflect the new name in help text and command descriptions. Additionally, the CLI now supports an experimental '--auto' option for automatic data schema detection, requires the '--answers' flag to be optional rather than required, and renames the '--override' flag to '--overwrite'. The underlying project generation logic has also been refactored, renaming the ProjectTemplate trait and related classes to ProjectGenerator.
cli/src/main/scala/com/salesforce/op/cli · high confidence
Enhanced feature type API and stricter null handling
The feature type system in \features/src/main/scala/com/salesforce/op/features/types\ has been updated with several behavioral changes. \OPVector\ now supports direct arithmetic operations (addition, subtraction, dot product, and combination) and is no longer strictly non-nullable. \RealNN\ now throws a \NonNullableEmptyException\ if constructed with an empty or null value, replacing the previous silent fallback to a default value. The \URL\ type's validation is now stricter, requiring both URL format validity and successful Java URL parsing. Additionally, new implicit conversion methods (e.g., \toBinary\, \toRealNN\) have been added to various numeric and option types, and the internal type mapping now uses \ReflectionUtils.dealisedTypeName\ for more robust type resolution.
features/src/main/scala/com/salesforce/op/features/types · high confidence
Project renamed to TransmogrifAI and updated to version 0.7.0
The project has been renamed from 'Octopus Prime' to 'TransmogrifAI' and released as version 0.7.0. This update includes a comprehensive changelog detailing new features such as support for Spark 2.4.5, XGBoost 0.90, and MLeap 0.14, alongside bug fixes and dependency updates. The repository now includes standard open-source governance files like a Code of Conduct, Contributing guidelines, and a Security policy, and the build system has been configured with Travis CI for continuous integration.
(repo-wide) · high confidence
Refactors simple template to use DataClass schema and simplified feature extraction
The simple template now uses a generic \DataClass\ schema instead of the specific \Passenger\ schema, requiring users to adapt their data models accordingly. Feature extraction logic has been simplified by replacing verbose \FeatureBuilder\ calls with direct helper methods (e.g., \realFromInt\, \iPickList\), and the application entry point now extends \OpAppWithRunner\ instead of \OpApp\, utilizing a new \OpWorkflowRunner\ for execution. Additionally, a comment warning that this is a fake template prototype has been added to the Avro schema file.
templates/simple · high confidence
Upgrade Gradle wrapper to 5.2 and rebrand template to TransmogrifAI
The simple project template now uses Gradle 5.2 (upgraded from 3.5) with a SHA-256 checksum for the wrapper distribution and updated JVM options for newer Java versions. The template has been renamed from Optimus Prime to TransmogrifAI, updating all references in the README, build configuration, and dependency declarations to use the TransmogrifAI core library and documentation links.
templates · high confidence
Upgrade Gradle wrapper to version 5.2
The project's Gradle wrapper has been updated to use Gradle 5.2 (previously 4.1). This change ensures that builds are executed with the newer Gradle distribution, which may introduce behavioral differences or improvements in build performance and dependency resolution compared to the previous version.
gradle/wrapper · high confidence
Test coverage
Added CLI code generation and full-cycle integration tests; Added and refactored test infrastructure for OpWorkflow and ModelInsights; Added test coverage for IO, JSON, numeric, and Spark utility classes; Added test coverage for feature distribution and raw feature filter components; Added test coverage for feature transformation stages; Added test logging configuration; Added test resource for alternate reader and updated test logging configuration; Added test resources and updated test configuration; Added test resources for CSV-to-Avro conversion and JSON/YAML utilities; Added test resources for OldModelVersion 0.5.1; Added test resources for old model versions; Added tests for ExtendedMultiset and TimeBasedAggregator; updated MonoidAggregatorDefaultsTest; Added tests for FeatureGeneratorStage serialization and custom extract functions; Added tests for FeatureSparkType mapping and FeatureBuilder DataFrame support; Added tests for OPVectorMetadata, RichEvaluator, RichVector, and RichDataset utilities; Added tests for OpParams copy, swap, and validation logic; Added tests for RecordInsights correlation and LOCO stages; Added tests for RichNumericFeature math operations; Added tests for Spark job grouping and OpenNLP text processing utilities; Added tests for TestFeatureBuilder random dataset generation; Added tests for feature serialization, metadata preservation, and feature graph traversal; Added tests for local model scoring and MLeap conversion; Added tests for model selector metadata serialization and random parameter generation; Added tests for pipeline stage serialization and lambda transformer variants; Added unit tests for FeatureHistory, SensitiveFeatureInformation, and UID; Added unit tests for base feature transformation stages; Added unit tests for regression estimator wrappers and model selectors; Added unit tests for tuning data preparation stages; Expanded test coverage for data readers and metadata handling; Expanded test coverage for feature preparation stages; Expanded test coverage for feature types and conversions; Expanded test coverage for model evaluators and metrics; Updated CLI generator tests to reflect API and configuration changes; Updated test infrastructure for passenger feature extraction and data reading; Updated test logging configuration comment; Updated test logging configuration for Hadoop, Avro, and TransmogrifAI; Updated tests for SparkStageParam serialization using Bucketizer; Updated tests for random data generators in the testkit.
Dependencies
Upgrade to TransmogrifAI 0.7.0 with Spark 2.4.5 and MLeap integration
The project has been renamed to TransmogrifAI and upgraded to version 0.7.0, with the build system now targeting Spark 2.4.5. This release introduces MLeap serialization support for Spark models, allowing models to be saved and loaded locally in a format independent of Spark classes. Additionally, the build configuration updates the Gradle wrapper to version 5.2, upgrades the Shadow JAR plugin to 5.0.0, and refines Spark submission defaults, such as changing the local master configuration to use all available cores.
(dependencies) · high confidence
Housekeeping
Added documentation and version metadata for OpenNLP models
The models module now includes a README explaining that it contains pretrained OpenNLP models (such as POS and NER) and can be included as a runtime dependency. Additionally, a version file has been added to record that the included OpenNLP models are version 1.5.
models · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 47 → 70 (+23.6)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 99 → 88 (-11.0)
- Architecture 69 → 99 (+30.3)
- Maturity 46 → 57 (+11.8)
- Readiness 27 → 78 (+50.9)
- Security 100 → 81 (-19.2)
Resolved (9)
- Dimension evaluation failed
- LLM evaluation failed
- No automated tests
- No exposed public API
- No tests found
- Test reliability not included
- early-stage repository — too little history to judge knowledge freshness
- single-commit history — no usable git history window to measure hotspots
- single-maintainer — knowledge-concentration (bus factor) risk
New (185)
- DateListVectorizer.onGetMetadata (cognitive 17) (core/src/main/scala/com/salesforce/op/stages/impl/feature/DateListVectorizer.scala)
- Dependency hygiene PARTLY measured — Maven/Gradle declarations read, no dependency graph resolved
- Dormant codebase
- Duplicated block (10 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/feature/SmartTextMapVectorizer.scala)
- Duplicated block (10 lines × 2) (helloworld/src/main/scala/com/salesforce/hw/OpBostonSimple.scala)
- Duplicated block (11 lines × 2) (core/src/main/scala/com/salesforce/op/dsl/RichNumericFeature.scala)
- Duplicated block (12 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/classification/BinaryClassificationModelSelector.scala)
- Duplicated block (12 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/tuning/DataCutter.scala)
- Duplicated block (12–13 lines × 2) (core/src/main/scala/com/salesforce/op/evaluators/Evaluators.scala)
- Duplicated block (14 lines × 2) (core/src/main/scala/com/salesforce/op/dsl/RichMapFeature.scala)
- Duplicated block (15 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/classification/BinaryClassificationModelSelector.scala)
- Duplicated block (15 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/feature/TextMapLenEstimator.scala)
- Duplicated block (15 lines × 2) (features/src/main/scala/com/salesforce/op/OpParams.scala)
- Duplicated block (16 lines × 2) (core/src/main/scala/com/salesforce/op/dsl/RichMapFeature.scala)
- Duplicated block (16–17 lines × 3) (helloworld/src/main/scala/com/salesforce/hw/OpBostonSimple.scala)
- Duplicated block (17–18 lines × 2) (core/src/main/scala/com/salesforce/op/dsl/RichMapFeature.scala)
- Duplicated block (18 lines × 3) (core/src/main/scala/com/salesforce/op/dsl/RichMapFeature.scala)
- Duplicated block (25 lines × 2) (core/src/main/scala/com/salesforce/op/dsl/RichMapFeature.scala)
- Duplicated block (39 lines × 4) (core/src/main/scala/com/salesforce/op/stages/impl/classification/OpDecisionTreeClassifier.scala)
- Duplicated block (5 lines × 2) (core/src/main/scala/com/salesforce/op/stages/impl/classification/BinaryClassificationModelSelector.scala)
- …and 165 more
Architecture
- Containers 0 added · 0 removed · contexts 2 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (2)
- repository
- utils
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
salesforce/TransmogrifAI was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 8cec508d0348371d80f993586ffe7c619885c674 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.