recommenders-team/recommenders
61.7
Adequate · 19 September 2026
27.3k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a comprehensive library for building, training, and evaluating recommendation algorithms, supporting both collaborative and content-based filtering approaches. It provides implementations for a wide range of models, including matrix factorization, deep learning architectures, and sequential recommenders, with backends in PyTorch, TensorFlow, and Spark. The library also includes utilities for data preparation, hyperparameter tuning, and production deployment via Azure ML and Databricks.
How it got here
2018–2020 — Project foundation and SARplus integration
21 changes.
This period established the core repository structure, testing infrastructure, and documentation for the Recommenders library. It focused heavily on introducing the SARplus library for Spark-based recommendations, including its Python and Scala implementations, build systems, and comprehensive test suites. The work also expanded the examples directory with extensive notebooks covering data preparation, various recommendation algorithms, evaluation metrics, hyperparameter tuning, and production operationalization.
2021 — PyTorch migration and package restructuring
40 changes.
The library was renamed to 'recommenders' and underwent a significant migration of deep learning models from TensorFlow to PyTorch, including NCF, VAE, Wide & Deep, and SASRec. This period also introduced new model implementations, expanded dataset and evaluation utilities, and established comprehensive unit and smoke test coverage across the codebase.
2023–2026 — PyTorch migration and test expansion
12 changes.
The project focused on migrating key recommendation models, such as EmbeddingDotBias and SLi-Rec, from TensorFlow to PyTorch to modernize the stack. This period also involved a significant expansion of test coverage, including functional, integration, and security tests, to ensure the stability and correctness of the new implementations and existing infrastructure.
Features
Add ImplicitCF data processing class for GCN models
Introduces the ImplicitCF class in the deeprec DataModel module to handle data preprocessing for Graph Convolutional Network (GCN) models using implicit feedback. This new component manages the creation and caching of normalized adjacency matrices, reindexes user and item IDs, and filters training data to include only records with ratings greater than zero, thereby enabling efficient training data preparation for implicit recommendation scenarios.
recommenders/models/deeprec/DataModel · high confidence
Add NNI integration scripts for NCF and SVD model tuning
New files have been added to the recommenders/tuning/nni module to enable hyperparameter tuning of Neural Collaborative Filtering (NCF) and Surprise SVD models using the Neural Network Intelligence (NNI) framework. The addition includes training scripts (ncf\_training.py, svd\_training.py) that handle data loading, model fitting, metric evaluation, and result reporting via NNI, along with utility modules (nni\_utils.py, ncf\_utils.py) for managing NNI experiment status, trial retrieval, and metric computation.
recommenders/tuning/nni · high confidence
Add PyTorch-compatible data loaders for sequential recommendation models
The \recommenders/models/deeprec/io\ package now includes \SequentialDataset\, a PyTorch-native data loader that mirrors the parsing, batching, and time-feature logic of the existing TensorFlow \SequentialIterator\. This allows users to train sequential recommendation models (like SLi-Rec) using PyTorch without changing data formats or preprocessing steps, while the original TensorFlow iterators remain available for TF-based workflows.
recommenders/models/deeprec/io · high confidence
Add pysarplus\_dummy package to trigger Spark packaging
A new dummy Python package (pysarplus\_dummy) has been added to the contrib/sarplus directory. This package serves as a placeholder to trigger the Spark packaging process for the SARplus library, reading its version from a shared VERSION file.
contrib/sarplus/scala/python · high confidence
Add sequential recommendation models (A2SVD, Caser, GRU, NextItNet, SUM)
New sequential recommendation models are now available in the deeprec library, including A2SVD (Attentive Asynchronous Singular Value Decomposition), Caser (Convolutional Sequence Embedding), GRU, NextItNet (Dilated CNN with residual blocks), and SUM (Sequential User Matrix). These models extend the SequentialBaseModel to support item and category history embeddings, enabling personalized recommendations based on user behavior sequences.
recommenders/models/deeprec/models/sequential · high confidence
Added Cornac model wrappers and utility functions for recommendation and ranking
This change introduces a new \recommenders/models/cornac\ module that provides Python wrappers and utilities for integrating the Cornac library. It includes a custom \BPR\ class that extends Cornac's base BPR model with a \recommend\_k\_items\ method, enabling top-k item recommendations per user with an optional \remove\_seen\ flag to exclude previously interacted items. Additionally, \cornac\_utils.py\ provides \predict\ and \predict\_ranking\ functions to compute rating predictions and full user-item ranking scores, respectively, facilitating the calculation of metrics like RMSE and NDCG.
recommenders/models/cornac · high confidence
Added GeoIMC model for matrix completion with side information
Introduced the GeoIMC (Geometric Matrix Completion) model, which leverages row and column side features to improve recommendation accuracy. The implementation includes a core optimization engine using Pymanopt on Stiefel and Symmetric Positive Definite manifolds, data handling utilities for loading and preprocessing datasets like MovieLens-100K, and an inference module supporting dot-product similarity with mean or top-k transformations.
recommenders/models/geoimc · high confidence
Added Scala macro for Spark 3.2+ compatibility
A new Scala macro annotation \since3p2defvisible\ has been added to the \sarplus/scala/compat\ module. This macro enables conditional compilation of methods based on the presence of the \path\ member in Spark's \OutputWriter\ class, allowing the library to support Spark 3.2.x features while maintaining backward compatibility with earlier versions.
contrib/sarplus/scala/compat · high confidence
Added configuration validation for NewsRec models
The NewsRec module now includes a utility script that validates configuration parameters for news recommendation models (NRMS, NAML, LSTUR, NPA). This ensures that required settings such as embedding dimensions, batch sizes, and data formats are present and of the correct type (integer, float, string, list, or boolean) before training begins, helping users catch configuration errors early.
recommenders/models/newsrec · high confidence
Added utility functions for Surprise model integration
New helper functions have been added to the Surprise model module to facilitate interaction with the Surprise library. The \surprise\_trainset\_to\_df\ function allows users to convert a Surprise Trainset object into a pandas DataFrame, mapping internal IDs back to raw user and item identifiers. Additionally, \predict\ and \compute\_ranking\_predictions\ provide standardized ways to generate rating predictions and ranking-based predictions (with optional removal of seen items) from Surprise algorithms, returning results in a consistent DataFrame format suitable for metric calculation.
recommenders/models/surprise · high confidence
Added utility scripts for KDD 2020 tutorial data processing
New utility modules have been added to the KDD 2020 tutorial to support data handling and preprocessing. These include \PandasMagClass\ for reading Microsoft Academic Graph (MAG) streams into Pandas DataFrames, \data\_helper\ for loading paper references, dates, and author relationships, \general\ for common utility functions like directory creation and dictionary dumping, and \task\_helper\ for generating paper content, knowledge relations, and indexed sentence collections used in the tutorial notebooks.
_examples/07\tutorials/KDD2020-tutorial/utils · high confidence
Initial project foundation and documentation
Establishes the core repository structure with essential documentation files including the README, contributing guidelines, code of conduct, glossary, and author credits, alongside the MIT license and package manifest configuration.
(repo-wide) · high confidence
Introduce LightGBM feature engineering utilities
Added \lightgbm\_utils.py\ to the \recommenders/models/lightgbm\ package, providing a \NumEncoder\ class that handles categorical and numerical feature preprocessing for LightGBM models. This utility performs sequential label encoding, count encoding, binary encoding, and target encoding (mean and count) while filtering low-frequency categories and handling missing values, returning processed numpy arrays for training and inference.
recommenders/models/lightgbm · high confidence
Introduce RLRMC matrix completion model using Pymanopt
Added the RLRMC (Regularized Low-Rank Matrix Completion) model to the recommenders library. This new feature implements a Riemannian optimization algorithm for matrix completion, leveraging the Pymanopt library for manifold optimization (specifically using Stiefel and Symmetric Positive Definite manifolds) and a custom Conjugate Gradient solver. The implementation includes dataset preprocessing utilities to handle user/item indexing and sparse matrix construction, enabling users to perform low-rank matrix factorization with support for SVD-based initialization and configurable regularization.
recommenders/models/rlrmc · high confidence
Introduce SARplus library for Spark-based recommendations
Adds the SARplus library, providing an efficient implementation of the Simple Algorithm for Recommendation (SAR) for personalized recommendations in Spark. This includes a Python package (pysarplus) with a C++ extension for prediction, a Scala DataSource implementation for reading and writing SAR cache formats, and associated build and test infrastructure.
contrib/sarplus/python · high confidence
Introduce pysarplus library for Spark-compatible SAR recommendations
Adds the pysarplus package, providing a Python API for the Similarity of Activated Ratings (SAR) algorithm optimized for Apache Spark. The library includes the SARPlus class for training models on Spark DataFrames with support for configurable similarity metrics (cooccurrence, jaccard, lift), time-decay formulas, and caching, alongside the SARModel class which interfaces with a C++ backend for fast prediction inference.
contrib/sarplus/python/pysarplus · high confidence
Introduce recommenders/utils package with shared utilities
The new recommenders/utils package provides a centralized set of helper modules for the Recommenders library. It includes constants for default column names and filtering parameters, general utilities for system resource detection (CPU, memory), and GPU management functions that rely exclusively on PyTorch for device counting and memory information. The package also adds K8s utilities for estimating replica counts based on QPS and node cores, notebook utilities for programmatic execution with parameter injection and memory profiling, and plotting helpers for generating line graphs. Additionally, it introduces Python utilities for calculating item similarity metrics (Jaccard, Lift, Cosine, etc.) from co-occurrence matrices, Spark configuration helpers for starting sessions with MMLSpark, TensorFlow utilities for input functions and model export, and a Timer class for performance measurement.
recommenders/utils · high confidence
Introduction of DeepRec model utilities and configuration validation
The \recommenders/models/deeprec\ module now includes \deeprec\_utils.py\, which provides helper functions for flattening YAML configuration files and validating parameter types (int, float, str, list) for various neural network models. It also includes specific validation logic for model types such as FM, LR, DKN, xDeepFM, and GRU, ensuring that required configuration parameters are present and correctly typed before model initialization.
recommenders/models/deeprec · high confidence
Introduction of SAR Single-Node Model with Multiple Similarity Metrics
The SAR (Simple Algorithm for Recommendations) model is now available as a single-node implementation in the \recommenders\ package. This change introduces the \SARSingleNode\ class, which allows users to generate recommendations based on user transaction history and item descriptions. A key feature is the support for multiple item-item similarity metrics, including cooccurrence, cosine, inclusion index, jaccard, lexicographers mutual information, lift, and mutual information. The model also supports time-decay formulas and normalization of predictions, providing flexible configuration for personalized recommendation scenarios.
recommenders/models/sar · high confidence
New AzureML Designer modules for SAR model training, scoring, and evaluation metrics
This change introduces a set of new entry-point scripts in the AzureML Designer modules that enable users to train, score, and evaluate recommender systems directly within the visual interface. Specifically, it adds modules for training a Single Active Recommendation (SAR) model (\train\_sar\_entry.py\) and scoring it for either item recommendations or rating predictions (\score\_sar\_entry.py\). Additionally, it provides evaluation modules to compute key ranking metrics—MAP at K, nDCG at K, Precision at K, and Recall at K—using the \recommenders\ library, along with a stratified data splitter (\stratified\_splitter\_entry.py\) to prepare datasets for these workflows.
_contrib/azureml\_designer\modules/entries · high confidence
New AzureML notebook submission example and notebook template
Added a new \run\_notebook\_on\_azureml.ipynb\ example that allows users to submit existing Jupyter notebooks directly to Azure Machine Learning compute targets without rewriting training scripts, and introduced a \template.ipynb\ to standardize notebook structure and metadata handling for the examples directory.
examples · high confidence
New Databricks installation and operationalization tool
A new script, tools/databricks\_install.py, has been added to help users install the Recommenders library onto a Databricks workspace. The tool automates the setup of a cluster, installs required PyPI dependencies (such as specific versions of pip, setuptools, and numpy), and optionally configures operationalization libraries like Azure Cosmos DB Spark connectors and mmlspark based on the target Spark version.
tools · high confidence
New LightFM utility functions for evaluation and similarity
Added a new \lightfm\_utils.py\ module providing helper functions for LightFM models. This includes \track\_model\_metrics\ for monitoring precision and recall during training, \model\_perf\_plots\ for visualizing performance, and \similar\_users\/\similar\_items\ for finding nearest neighbors based on cosine similarity. These utilities simplify the evaluation and analysis workflow for LightFM-based recommenders.
recommenders/models/lightfm · high confidence
New MIND data iterators for news recommendation models
Added \MINDAllIterator\ and \MINDIterator\ classes in the \recommenders.models.newsrec.io\ module to handle data loading for the MIND news recommendation dataset. These iterators parse news articles (titles, abstracts, categories) and user behavior logs (click history, impressions) into the specific tensor formats required by deep learning models like NAML, supporting features such as negative sampling and efficient batch-wise memory loading.
recommenders/models/newsrec/io · high confidence
New collaborative filtering deep dive documentation and notebooks
The examples/02\_model\_collaborative\_filtering directory now includes a comprehensive README and a suite of Jupyter notebooks providing deep dives into various collaborative filtering algorithms. These resources cover Spark ALS, baseline performance estimation, Cornac-based BiVAE and BPR models, Factorization Machines (FM/FFM), LightFM, LightGCN, Multinomial VAE, Neural Collaborative Filtering (NCF), Restricted Boltzmann Machines (RBM), SAR, Standard VAE, and Surprise SVD. Each notebook details the algorithm's theory, implementation, and usage with utility functions from the recommenders library, serving as educational guides for model training and evaluation.
_examples/02\_model\_collaborative\filtering · high confidence
New content-based filtering example notebooks added
The \examples/02\_model\_content\_based\_filtering\ directory now includes three new Jupyter notebooks demonstrating content-based recommendation algorithms: \dkn\_deep\_dive.ipynb\ (Deep Knowledge-Aware Network for news recommendation using the MIND dataset), \mmlspark\_lightgbm\_criteo.ipynb\ (LightGBM on Spark for CTR prediction using the Criteo dataset), and \vowpal\_wabbit\_deep\_dive.ipynb\ (Vowpal Wabbit for regression and matrix factorization using the MovieLens dataset). A README.md file has also been added to document these examples and their required environments.
_examples/02\_model\_content\_based\filtering · high confidence
New data preparation notebooks for splitting, transformation, and knowledge graphs
The \examples/01\_prepare\_data\ directory now includes a set of documentation notebooks that demonstrate common data preparation tasks for recommendation systems. The \data\_split\ notebook illustrates random, chronological, and stratified splitting strategies for both Spark and pandas DataFrames. The \data\_transform\ notebook covers techniques for handling explicit and implicit feedback in collaborative filtering scenarios. Additionally, \mind\_utils\ provides guidance on generating dictionaries and embeddings for the MIND news dataset, while \wikidata\_knowledge\_graph\ shows how to extract and construct knowledge graphs from Wikidata for algorithms like DKN and RippleNet.
_examples/01\_prepare\data · high confidence
New dataset loaders and splitters for Amazon, Cosmos, COVID, Criteo, MIND, and MovieLens
The \recommenders/datasets\ module now includes dedicated loaders and utilities for several key datasets. Users can load Amazon Reviews, Criteo DAC, MIND, and MovieLens datasets via new \load\_pandas\_df\ and \load\_spark\_df\ functions, with Criteo and MovieLens supporting both pandas and PySpark outputs. A new \covid\_utils\ module provides tools to load and clean the Azure Open Research COVID-19 dataset, while \cosmos\_cli\ offers helper functions for interacting with CosmosDB collections. Additionally, \python\_splitters\ and \spark\_splitters\ provide new functions for random, chronological, and stratified data splitting for both pandas and Spark DataFrames, and \download\_utils\ centralizes file downloading with retry logic and temporary path management.
recommenders/datasets · high confidence
New evaluation example notebooks for accuracy and diversity metrics
The \examples/03\_evaluate\ directory now includes two new notebooks: \evaluation.ipynb\ and \als\_movielens\_diversity\_metrics.ipynb\. The first notebook demonstrates how to calculate standard rating and ranking metrics (such as RMSE, MAE, Precision, Recall, and NDCG) using both Python/CPU and PySpark environments. The second notebook focuses on non-accuracy metrics, illustrating how to evaluate novelty, diversity, serendipity, and coverage by comparing an ALS recommender against a random baseline in a PySpark environment.
_examples/03\evaluate · high confidence
New hyperparameter tuning examples for Spark ALS, AzureML, and NNI
Added a new example directory demonstrating how to tune and optimize hyperparameters for recommender algorithms using three distinct approaches: Spark native tuning and Hyperopt for Spark ALS, Azure Machine Learning Hyperdrive for Wide-and-Deep and Surprise SVD models, and the Neural Network Intelligence (NNI) toolkit for NCF and Surprise SVD models.
_examples/04\_model\_select\_and\optimize · high confidence
New operationalization examples for Azure ML, AKS, and Databricks
Added three new notebooks to the examples/05\_operationalize directory that demonstrate how to build, evaluate, and deploy recommendation systems in production. The als\_movie\_o16n notebook provides an end-to-end example of a Spark ALS movie recommender using Azure Databricks, Cosmos DB, and Azure Kubernetes Service. The aks\_locust\_load\_test notebook shows how to perform load testing on a deployed model using Locust. The lightgbm\_criteo\_o16n notebook demonstrates content-based personalization deployment for ad click prediction using Azure ML and AKS.
_examples/05\operationalize · high confidence
New parameter sweep utility for hyperparameter tuning
A new \parameter\_sweep.py\ module has been added to the \recommenders/tuning\ package, providing a \generate\_param\_grid\ function. This utility allows users to convert a dictionary of parameter options into a list of all possible parameter combinations, which can be directly used for hyperparameter tuning workflows.
recommenders/tuning · high confidence
New training scripts for SVD and Wide-Deep models
Added new entry-point scripts for the model selection and optimization example: \svd\_training.py\ enables training and evaluating Surprise SVD models with configurable hyperparameters and AzureML logging, while \wide\_deep\_training.py\ serves as an AzureML Hyperdrive entry script that executes the Wide-Deep notebook with support for tuning linear and deep network parameters.
_examples/04\_model\_select\_and\_optimize/train\scripts · high confidence
PyTorch port of the SLi-Rec sequential recommender model
A new PyTorch implementation of the SLi-Rec model has been added to the sequential recommenders module, replacing the previous TensorFlow 1.x version. This port includes a custom \Time4LSTMCell\ that modulates LSTM state updates with time-aware gates, a \SequentialBaseModel\ providing shared embedding and attention building blocks, and the main \SLiRecModel\ class that combines long-term ASVD attention with short-term time-aware LSTM features. The implementation is designed to reproduce the behavior of the original TensorFlow model, preserving specific numeric conventions for weight initialization, batch normalization, and loss calculation to ensure comparable performance metrics.
recommenders/models/deeprec/models/sequential/pytorch · high confidence
Python evaluation module introduced alongside Spark evaluator
The \recommenders/evaluation\ package now includes a new \python\_evaluation.py\ module that provides pandas-based implementations of rating metrics (RMSE, MAE, R², explained variance) and ranking metrics (NDCG, MAP, R-Precision, etc.), complementing the existing \spark\_evaluation.py\ which handles Spark DataFrames. This addition allows users to evaluate recommendation models using pure Python/pandas workflows without requiring a Spark environment, with the new module handling column validation, data merging, and metric calculation for both rating and ranking scenarios.
recommenders/evaluation · high confidence
Quick Start examples now include GeoIMC and LightGBM LambdaRank notebooks
The Quick Start directory has been updated with two new algorithm demonstrations: a notebook for Geometry Aware Inductive Matrix Completion (GeoIMC) using the MovieLens dataset, and a notebook for LightGBM Learning-to-Rank with LambdaRank on MovieLens. The GeoIMC example demonstrates inductive matrix completion using user and item features, while the LightGBM example shows how to build a ranking pipeline by binarizing ratings and optimizing NDCG directly. These additions expand the available quick-start guides for users exploring new recommendation algorithms.
_examples/00\_quick\start · high confidence
TF-IDF recommender model implementation and tokenization logic
The TF-IDF recommender module has been introduced in the recommenders/models/tfidf package, providing content-based recommendations via TF-IDF vectorization and cosine similarity. The TfidfRecommender class supports multiple tokenization strategies including NLTK stemming, Hugging Face BERT (bert-base-cased), SciBERT (allenai/scibert\_scivocab\_cased), and raw word analysis. The implementation includes text cleaning utilities that handle HTML tags, unicode normalization, and special character removal, along with dataframe processing to combine and clean text columns prior to vectorization.
recommenders/models/tfidf · high confidence
Behavioural changes
DeepRec models migrated to TensorFlow 2.x compatibility
The deep learning recommendation models (DKN, XDeepFM, and their variants) have been updated to run on TensorFlow 2.x. This change replaces deprecated TensorFlow 1.x APIs with their \tf.compat.v1\ equivalents, disables eager execution to maintain the existing graph-based logic, and removes dependencies on removed libraries such as \tf.contrib.training.HParams\ and \tf.addons\. Users will see no change in model architecture or output, but the code is now compatible with modern TensorFlow versions.
recommenders/models/deeprec/models · high confidence
Introduce PyTorch-based EmbeddingDotBias collaborative filtering model
Added a new PyTorch implementation of the EmbeddingDotBias model for collaborative filtering, replacing the previous fastai dependency. This change introduces a new data loading pipeline (RecoDataset and RecoDataLoader) that handles user/item ID mapping and train/validation splitting, a core model class using embedding layers and bias terms for rating prediction, and a dedicated Trainer class for training and validation loops. Utility functions for scoring recommendations and computing Cartesian products are also included to support the new model's inference workflow.
recommenders/models/embdotbias · high confidence
Introduce structured benchmarking utilities and notebook for collaborative filtering algorithms
Users can now run a standardized comparison of collaborative filtering algorithms (such as Spark ALS, SAR, LightGCN, and EmbDotBias) on the MovieLens dataset using the new \benchmark\_utils.py\ module and the \movielens.ipynb\ notebook. This change replaces previous ad-hoc benchmarking scripts with a unified framework that handles data preparation, model training, and evaluation across CPU, GPU, and Spark environments, while also correcting metric calculations (e.g., using \map\_at\_k\ instead of \map\ and filtering test sets for BPR) to ensure accurate performance reporting.
_examples/06\benchmarks · high confidence
KDD 2020 tutorial migrated to uv and updated data sources
The KDD 2020 tutorial example has been updated to use \uv\ for environment and dependency management instead of conda, with setup instructions now referencing \uv venv\ and \uv pip install\. The tutorial data is now downloaded from a Hugging Face dataset repository (\Recommenders/kdd2020\) rather than the previous storage location. Additionally, the tutorial notebooks have been refactored to use the new \recommenders\ package structure and updated utility modules (e.g., \utils\ instead of \common\), while model configuration files for DKN and LightGCN have been added to support the hands-on experiments.
_examples/07\tutorials/KDD2020-tutorial · high confidence
LightGCN model migrated to PyTorch
The LightGCN recommendation model has been rewritten from TensorFlow to PyTorch, introducing a new \LightGCN\ class that manages user and item embeddings via \nn.Embedding\ and performs graph propagation using sparse matrix multiplication. This change updates the underlying framework for this specific graph-based recommender, affecting how the model is instantiated, trained, and evaluated within the \recommenders\ library.
recommenders/models/deeprec/models/graphrec · high confidence
Migrate VAE models from TensorFlow to PyTorch
The Variational Autoencoder (VAE) implementations in the recommenders library have been rewritten in PyTorch, replacing the previous TensorFlow versions. This change introduces new \StandardVAE\ and \MultVAE\ classes in \recommenders/models/vae/\, which now utilize PyTorch's \nn.Module\ for model architecture, \torch.optim.Adam\ for optimization, and native PyTorch operations for loss calculation (binary cross-entropy for standard VAE, multinomial log-likelihood for MultVAE). Users should expect potential differences in numerical results due to the framework switch and different default initializations, although the API surface remains largely consistent for training and inference.
recommenders/models/vae · high confidence
NCF model ported from TensorFlow to PyTorch
The Neural Collaborative Filtering (NCF) implementation in \recommenders/models/ncf\ has been rewritten to use PyTorch instead of TensorFlow. This change replaces the previous TensorFlow-based model with a native PyTorch \nn.Module\ (in \ncf\_singlenode.py\) that supports GMF, MLP, and NeuMF architectures, along with updated dataset handling in \dataset.py\. Users relying on this model will now interact with a PyTorch backend, which may affect training performance, device placement (CPU/GPU), and model serialization formats.
recommenders/models/ncf, recommenders/models/rbm · high confidence
NewsRec models migrated to TensorFlow 1.x compatibility layer
The NewsRec model implementations (LSTUR, NAML, NPA, NRMS) and their shared base class and layers have been updated to use the \tensorflow.compat.v1\ API. This change switches the underlying framework usage from native TensorFlow 2.x to the TensorFlow 1.x compatibility layer, which alters how sessions, graphs, and Keras layers are initialized and executed within these models.
recommenders/models/newsrec/models · high confidence
Recommenders package renamed and versioned at 1.2.1
The library has been renamed from \reco\_utils\ to \recommenders\, updating the package title and installation instructions accordingly. The package version is now set to 1.2.1, and the \recommenders.models\ submodule has been initialized to support the new structure.
recommenders · high confidence
SASRec and SSEPT models ported to PyTorch
The SASRec and SSEPT recommendation models in the \recommenders/models/sasrec\ package have been rewritten from TensorFlow to PyTorch. This change introduces new model implementations (\model.py\, \ssept.py\) using PyTorch modules for multi-head attention and feed-forward layers, along with a PyTorch-compatible data sampler (\sampler.py\) and dataset utility (\util.py\) for handling user-item interactions and train/validation/test splits.
recommenders/models/sasrec · high confidence
Updated build infrastructure for SAR+ Scala packaging
The build configuration for the SAR+ Scala component has been updated to support Spark 3.x and improve the packaging process. The sbt version is upgraded to 1.6.2, and build plugins for code coverage (sbt-scoverage 1.9.3) and PGP signing (sbt-pgp 2.1.2) are added. A new utility function is introduced to correctly classify and name JAR artifacts for Spark 3.2 and later versions, ensuring compatibility with newer Spark releases.
contrib/sarplus/scala/project · high confidence
Wide & Deep model implementation migrated to PyTorch
The Wide & Deep recommendation model in the recommenders library has been rewritten from TensorFlow to PyTorch. This change introduces a new \WideDeepModel\ class that defines the wide (linear sparse features) and deep (DNN embeddings) components using native PyTorch layers, along with a complete training and inference API (\fit\, \predict\, \recommend\_k\_items\) built on PyTorch's \Dataset\ and \DataLoader\ utilities.
_recommenders/models/wide\deep · high confidence
Test coverage
Add unit tests for example notebooks across Python, GPU, and PySpark environments; Added \_\init\\.py files to test directories; Added \\init\\_.py to enable AzureML test execution; Added data validation tests for recommender datasets; Added functional tests for example notebooks; Added init file for AzureML integration tests; Added integration tests for Kubernetes utility functions; Added performance benchmarks for Python evaluation metrics; Added privacy validation tests for Criteo and Movielens datasets; Added regression tests for TensorFlow compatibility; Added smoke tests for GPU, PySpark, and Python notebook examples; Added smoke tests for deep learning recommender models and utilities; Added test package initialization for AzureML data validation; Added tests for MIND and Wikidata data validation notebooks; Added tests for dependency security versions; Added unit tests for Python and Spark recommender evaluation metrics; Added unit tests for dataset utilities and splitting logic; Added unit tests for recommender tuning utilities; Added unit tests for recommenders.utils modules; Initial test infrastructure and documentation; Initial test suite for PySARPlus; Initial unit test coverage for recommender models.
Dependencies
Add build and dependency manifests for SARplus and documentation
This change introduces build configuration files for the SARplus library and the project documentation. For SARplus, a Python \pyproject.toml\ defines build requirements including \setuptools\>=58.0.0\ and \wheel\, while a Scala \build.sbt\ configures dependencies for Spark (defaulting to 3.2.1) and Hadoop, and implements support for Scala 2.13+ by switching from the \paradise\ plugin to \-Ymacro-annotations\. Additionally, a root \pyproject.toml\ establishes build dependencies (\setuptools\, \wheel\, \numpy\) and registers custom pytest markers (e.g., \experimental\, \gpu\), and \docs/requirements-doc.txt\ specifies \jupyter-book\>=1.0.0\ for documentation generation.
(dependencies) · high confidence
Housekeeping
Initial documentation and versioning for SARplus
Added the DEVELOPMENT.md guide for packaging and testing the SARplus library, updated the README with usage examples for Python, Jupyter, and PySpark, and introduced a VERSION file to centralize the version number at 0.6.6.
contrib/sarplus · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 62.
Lenses
- Code Health 89
- Architecture 100
- Maturity 57
- Readiness 53
- Security 73
Changes since last survey
- 300 commits — 224 feature/other, 76 fixes
By area
- (repo) — 60 commits
- tools/ci — 46 commits
- examples/00_quick_start — 35 commits
- recommenders/models — 22 commits
- .github/workflows — 21 commits
- examples/06_benchmarks — 21 commits
- (root) — 17 commits
- tests/unit — 16 commits
- examples/02_model_collaborative_filtering — 15 commits
- recommenders/datasets — 15 commits
- recommenders/utils — 8 commits
- tests/test_groups.yml — 5 commits
- tests/functional — 4 commits
- .github/PULL_REQUEST_TEMPLATE.md — 3 commits
- tests/README.md — 3 commits
- recommenders/evaluation — 2 commits
- tests/data_validation — 2 commits
- contrib/sarplus — 1 commit
- docs/superpowers — 1 commit
- examples/07_tutorials — 1 commit
Notable commits
- fix: :bug:
- fix: :bug:
- fix: Add by_threshold ranking metrics regression test
- fix: Fix BPR benchmark evaluation: filter test set to match training threshold
- fix: Fix BPR recommend_k_items: apply remove_seen before top-k selection
- fix: Fix MLLib docs link
- fix: Fix NaNs in test assertions
- fix: Fix Wikidata SPARQL 403 errors by adding User-Agent header and skip retries on 4xx
- fix: Fix asset URL in fm_deep_dive.ipynb
- fix: Fix by_threshold relevancy method to filter by score, not count
- fix: Fix tfidf get_top_k_recommendations for NumPy 2.x (#2353)
- fix: Merge branch 'staging' into bug/minor_pr_template
- fix: Merge fix on empty secrets to make it in effect (#2335)
- fix: Merge fix on wrong working directory in testing workflows (#2341)
- fix: Merge pull request #2284 from recommenders-team/bug/wikiendpoint
- fix: Merge pull request #2285 from recommenders-team/bug/minor_pr_template
- fix: Merge pull request #2289 from LincolnBurrows2017/fix-typo-outputing
- fix: Merge pull request #2293 from ds-wook/fix/sarplus-scala-213-paradise-plugin
- fix: Merge pull request #2295 from ds-wook/fix/negative-feedback-sampler-column-order
- fix: Merge pull request #2296 from ds-wook/fix/groupby-apply-pandas-compat
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
recommenders-team/recommenders was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0bb4b3690941ffb668118e31ccaf8a7d19f8212a — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.