Skip to content
CAI
Software that uses CAICheck a score

rust-ml/linfa

60.8

Adequate · 30 September 2026

32.7k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Linfa is a pure Rust machine learning toolkit designed for classical algorithms, providing a unified API for fitting, transforming, and predicting data. It encompasses a comprehensive suite of capabilities including clustering, linear and non-linear regression, classification, dimensionality reduction, and nearest-neighbor search. The system supports end-to-end workflows by integrating preprocessing utilities, dataset management, and model composition tools within a modular crate structure.

How it got here

2018–2021 — Initial release and algorithm expansion

42 changes.

This period marks the initial public release of the Linfa machine learning toolkit for Rust, establishing the core API, dataset abstractions, and foundational traits. The project rapidly expanded its ecosystem by introducing dedicated crates for a wide range of classical algorithms, including clustering, linear models, dimensionality reduction, and nearest neighbor search. Concurrently, the team built out essential supporting infrastructure such as preprocessing utilities, synthetic dataset generators, and comprehensive benchmarking suites to ensure performance and usability.

2022–2025 — algorithm expansion and benchmarking

8 changes.

This period focused on expanding the Linfa toolkit with new machine learning algorithms, including FTRL, LARS, Random Projection, and ensemble methods like AdaBoost and Random Forest. Concurrently, significant effort was dedicated to establishing a robust benchmarking infrastructure, adding performance suites for linear, ICA, and PLS algorithms, and introducing standardized configuration and profiling support.

Features

Add Diffusion Map non-linear dimensionality reduction

Introduces the Diffusion Map algorithm as a new non-linear dimensionality reduction technique. This feature computes a data embedding by applying PCA on the diffusion operator of a kernel matrix, allowing users to project data along the direction of the largest diffusion flow. The implementation includes parameter validation for embedding size and steps, and supports both BLAS-accelerated and standard eigenvalue decomposition paths depending on feature flags.

_algorithms/linfa-reduction/src/diffusion\map · high confidence

Add Least Angle Regression (LARS) algorithm

The \linfa-lars\ crate introduces a pure Rust implementation of the Least Angle Regression (LARS) algorithm, providing a new linear regression capability within the \linfa\ ecosystem. Users can now fit LARS models using \Lars::params()\ with configurable options such as intercept fitting, verbosity, and maximum non-zero coefficients, and generate predictions via the standard \predict\ interface. The implementation supports optional BLAS/LAPACK backends for performance and includes a usage example demonstrating fitting on the Diabetes dataset.

algorithms/linfa-lars · high confidence

Add OPTICS clustering algorithm

Introduces the OPTICS (Ordering Points To Identify Clustering Structure) algorithm to the linfa-clustering library. This new feature allows users to perform density-based clustering analysis that generates an augmented ordering of data points, enabling the identification of clusters with varying densities. The implementation includes configurable hyperparameters such as minimum points, distance metrics, and nearest neighbor algorithms, along with serialization support when the 'serde' feature is enabled.

algorithms/linfa-clustering/src/optics · high confidence

Add PLS regression example

A new example demonstrating how to use the PlsRegression algorithm from the linfa-pls crate has been added. The example generates synthetic data, fits a PLS regression model with scaling enabled, and prints both the true and estimated regression coefficients to illustrate the model's performance.

algorithms/linfa-pls/examples · high confidence

Users can now fit Generalized Linear Models using the Tweedie distribution via the new \TweedieRegressor\. This implementation supports configurable power parameters to model Normal (power=0), Poisson (power=1), Gamma (power=2), and Inverse Gaussian (power=3) distributions, along with other Tweedie variants. The regressor includes built-in validation for target ranges based on the selected power, automatic link function selection (Identity for Normal, Log for others), and uses the L-BFGS solver for optimization. Hyperparameters such as regularization strength (\alpha\), maximum iterations, and tolerance are exposed through a builder-style API.

algorithms/linfa-linear/src/glm · high confidence

Add approximate DBSCAN clustering algorithm

Introduces the \AppxDbscan\ algorithm, an O(N) approximation of the standard DBSCAN clustering method. This implementation partitions the data space into a grid of cells and uses a counting tree structure for efficient range queries, allowing for scalable clustering on larger datasets. The algorithm supports configurable hyperparameters including \min\_points\, \tolerance\, \slack\, and a pluggable nearest-neighbour algorithm via the \linfa-nn\ integration. It returns cluster labels for core points and assigns border points to clusters or labels them as noise, with optional serialization support for the algorithm parameters.

_algorithms/linfa-clustering/src/appx\dbscan · high confidence

Add clustering algorithm examples (DBSCAN, K-Means, OPTICS)

New example scripts have been added to demonstrate how to use the DBSCAN, K-Means, and OPTICS clustering algorithms. These examples show how to generate synthetic datasets, configure and run the algorithms, and save the resulting cluster assignments and data to NumPy-compatible files for further analysis or visualization.

algorithms/linfa-clustering/examples · high confidence

Add diabetes and GLM regression examples

New example programs have been added to the linfa-linear crate to demonstrate usage of linear regression models. The diabetes.rs example shows how to fit a standard LinearRegression model to the diabetes dataset and print the resulting intercept and parameters. The glm.rs example demonstrates fitting a TweedieRegressor (configured as a standard linear regression by setting power and alpha to zero) and includes code to calculate and print the Mean Absolute Error on the training data.

algorithms/linfa-linear/examples · high confidence

Added benchmark configuration module with default settings

A new \src/benchmarks/mod.rs\ file introduces a benchmark configuration module that defines default settings for the Criterion benchmarking framework. This includes a sample size of 200, a measurement time of 10 seconds, a confidence level of 0.97, a warm-up time of 10 seconds, and a noise threshold of 0.05. On non-Windows platforms, it also provides a function to configure PProf profiling for generating flamegraphs.

src/benchmarks · high confidence

Added logistic regression examples for single and multi-class classification

New example files demonstrate how to use the \LogisticRegression\ and \MultiLogisticRegression\ classifiers from the \linfa-logistic\ crate. The \logistic\_cv.rs\ example shows how to perform cross-validation on a single-target (binary) classification task using the wine quality dataset, while \winequality\_logistic.rs\ and \winequality\_multi\_logistic.rs\ provide basic fit/predict workflows for binary and multi-class scenarios respectively, including confusion matrix and accuracy reporting.

algorithms/linfa-logistic/examples · high confidence

Decision tree example now supports entropy-based splitting and LaTeX export

The decision tree example in the linfa-trees package has been updated to demonstrate training models using the entropy split quality criterion, in addition to the existing Gini criterion. The example also now exports the trained Gini tree structure to a LaTeX/TikZ file for visualization, allowing users to generate visual representations of the decision tree directly from the example output.

algorithms/linfa-trees/examples · high confidence

Initial project release and documentation

This entry establishes the initial public release of the Linfa machine learning toolkit for Rust. It introduces the core project documentation, including a comprehensive CHANGELOG.md tracking versions from 0.3.0 through 0.8.1, a CONTRIBUTE.md guide for algorithm implementation, and an updated README.md featuring the project mascot, badges, and a detailed table of supported algorithms (such as Naive Bayes, K-Means, and Random Forest). The release also formalizes the project's dual MIT/Apache-2.0 licensing structure and configures the build system to link against LAPACK/CBLAS libraries when specific backend features are enabled.

(repo-wide) · high confidence

Initial release of linfa-clustering crate with core algorithms

The linfa-clustering crate is introduced to provide pure Rust implementations of popular clustering algorithms, including K-Means, DBSCAN, Gaussian Mixture Models, and OPTICS. As part of the linfa ecosystem, this new module exposes these algorithms for use in classical machine learning tasks. Notably, the approximate DBSCAN implementation is currently provided as a type alias for the standard DBSCAN to leverage its superior performance and avoid outdated dependencies.

algorithms/linfa-clustering/src · high confidence

Initial release of linfa-svm for Support Vector Machines

The \linfa-svm\ crate is introduced, providing a pure Rust implementation of Support Vector Machines for classification and regression tasks. It includes binary classification (C and Nu variants), one-class classification, and support vector regression (Epsilon and Nu variants), utilizing the Sequential Minimal Optimization (SMO) solver. The API supports various kernel methods (Linear, Gaussian, Polynomial) and optional Platt scaling for probability calibration, with example usage provided for wine quality classification and noisy sine regression.

algorithms/linfa-svm · high confidence

Initial release of the Linfa machine learning toolkit

This change introduces the core Linfa library, providing a comprehensive Rust toolkit for classical machine learning tasks. It establishes the foundational API through key traits for fitting models (\Fit\, \FitWith\), transforming data (\Transformer\), and making predictions (\Predict\, \PredictInplace\). The release includes a robust dataset abstraction (\DatasetBase\) and a parameter validation system (\ParamGuard\) to ensure hyperparameter safety. Users gain access to essential evaluation metrics for classification (including confusion matrices, precision, recall, and F1-score), regression (MAE, MSE, R², MAPE, sMAPE), and clustering (Silhouette Score), along with a new Pearson correlation analysis module for feature inspection. The library is structured to support various algorithm categories such as linear models, clustering, and dimensionality reduction via dedicated sub-crates.

src · high confidence

Initial release of the linfa-reduction crate with PCA and data generation utilities

This change introduces the \linfa-reduction\ crate, providing Principal Component Analysis (PCA) for dimensionality reduction via SVD, along with configurable parameters for whitening and embedding size. The PCA implementation exposes key model attributes such as components, mean, singular values, and explained variance ratios. Additionally, the crate includes utility functions for generating synthetic test datasets, specifically swiss rolls and convoluted rings in 2D and 3D spaces, and defines a comprehensive error type (\ReductionError\) to handle validation issues like insufficient samples or invalid dimensions.

algorithms/linfa-reduction/src · high confidence

Users can now perform density-based spatial clustering of applications with noise (DBSCAN) using the \linfa\_clustering\ crate. The implementation supports configurable hyperparameters including minimum points, distance tolerance, custom distance metrics (defaulting to L2/Euclidean), and pluggable nearest-neighbor algorithms (defaulting to KdTree via \linfa\_nn\). The API provides a builder pattern for parameter configuration and returns cluster assignments as \Option\<usize\>\ values, where \None\ indicates noise points.

algorithms/linfa-clustering/src/dbscan · high confidence

Introduce Elastic Net and Multi-Task Elastic Net regression algorithms

The \linfa-elasticnet\ crate now provides pure Rust implementations of Elastic Net linear regression, combining L1 (LASSO) and L2 (Ridge) penalties via coordinate descent. This release adds support for both standard single-target regression and multi-task regression (multiple output variables) using block coordinate descent. Users can configure models with checked hyperparameters (penalty, l1\_ratio, intercept, tolerance, max\_iterations) and access fitted model details such as the hyperplane, intercept, duality gap, and Z-scores.

algorithms/linfa-elasticnet · high confidence

Introduce FastICA algorithm for Independent Component Analysis

The \linfa-ica\ crate now provides a pure Rust implementation of the Fast Independent Component Analysis (FastICA) algorithm, enabling users to separate mixed multivariate signals into their independent subcomponents. The implementation supports configurable hyperparameters such as the number of components, the neg-entropy approximation function (GFunc), maximum iterations, tolerance, and a random seed for reproducibility. It integrates with the \linfa\ ecosystem traits (\Fit\, \Predict\) and includes an example demonstrating signal unmixing.

algorithms/linfa-ica · high confidence

Introduce Follow The Regularized Leader (FTRL) Proximal model

The \linfa-ftrl\ crate is added to the ecosystem, providing a pure Rust implementation of the Follow The Regularized Leader - Proximal linear model, primarily designed for click-through rate (CTR) prediction in online learning settings. The implementation supports L1 and L2 regularization, stores internal \z\ and \n\ values for incremental weight calculation, and exposes a builder API for hyperparameters like \alpha\, \beta\, and regularization ratios. This release includes the core algorithm logic, parameter validation, error handling, a usage example on the winequality dataset, and benchmarks for training and prediction performance.

algorithms/linfa-ftrl · high confidence

Introduce Gaussian Mixture Model clustering algorithm

Adds a new Gaussian Mixture Model (GMM) implementation to the linfa-clustering library, enabling users to cluster data by fitting a mixture of Gaussian distributions using the Expectation-Maximization (EM) algorithm. The feature includes configurable hyperparameters for the number of clusters, convergence tolerance, regularization, and initialization methods (KMeans or random), along with error handling for common fitting issues like empty clusters or non-convergence.

_algorithms/linfa-clustering/src/gaussian\mixture · high confidence

Introduce K-means clustering with multiple algorithm variants and initialization strategies

Users can now perform K-means clustering using the \linfa-clustering\ crate. The implementation supports standard Lloyd's algorithm and Hamerly's accelerated variant for the assignment step, with Hamerly offering performance gains for well-separated clusters by skipping redundant distance computations. Centroid initialization is configurable via K-means++, K-means\|\| (optimized for large cluster counts), random selection, or precomputed centroids. The API allows parametrization of the distance metric, convergence tolerance, maximum iterations, and number of runs, with validation ensuring valid hyperparameters. The model supports serialization via serde when the feature is enabled.

_algorithms/linfa-clustering/src/k\means · high confidence

Introduce Naive Bayes classifiers to linfa-bayes

The \linfa-bayes\ crate now provides three Naive Bayes algorithm implementations: Gaussian, Multinomial, and Bernoulli. This release adds the core model structures, hyper-parameter validation (including smoothing parameters), and example scripts for each variant, enabling users to perform classification tasks using these probabilistic models within the Linfa toolkit.

algorithms/linfa-bayes · high confidence

Introduce Partial Least Squares (PLS) algorithms

The \linfa-pls\ crate now provides a pure Rust implementation of the Partial Least Squares algorithm family, ported from scikit-learn 0.24. Users can now fit and predict using \PlsRegression\, \PlsCanonical\, and \PlsCca\ models. The implementation supports configurable hyperparameters such as the number of components, scaling, and algorithm choice (NIPALS or SVD), and includes traits for fitting, transforming data, and predicting targets.

algorithms/linfa-pls · high confidence

Introduce Random Projection for dimensionality reduction

Added a new Random Projection reduction module that embeds data into a lower-dimensional space using random matrices, offering a computationally efficient alternative to methods like PCA. The implementation supports two projection strategies: Gaussian (dense) and Sparse (using the Achlioptas distribution), allowing users to choose based on their data sparsity and performance needs. Users can configure the target embedding dimension directly or specify a precision parameter (epsilon) which automatically calculates the required dimension via the Johnson-Lindenstrauss lemma. The module includes parameter validation to prevent dimension increases and supports custom random number generators for reproducible results.

_algorithms/linfa-reduction/src/random\projection · high confidence

Introduce agglomerative hierarchical clustering algorithm

Users can now perform agglomerative hierarchical clustering via the new \linfa-hierarchical\ crate. This feature allows clustering data based on a similarity kernel, supporting configurable merging methods (via the \kodama\ crate) and stopping criteria such as a target number of clusters or a maximum distance threshold. An example demonstrating clustering on the Iris dataset is also provided.

algorithms/linfa-hierarchical · high confidence

Introduce decision tree learning algorithm

The \linfa-trees\ crate now provides a decision tree implementation for classification, allowing users to fit models with configurable hyperparameters such as split quality (Gini or Entropy), maximum depth, and minimum sample weights. The API supports exporting the fitted tree structure to TikZ/LaTeX for visualization and includes an efficient level-order node iterator.

algorithms/linfa-trees · high confidence

Introduce kernel methods for dimensionality expansion

The linfa-kernel crate is added to the ecosystem, providing implementations of kernel methods such as radial basis function (RBF) and polynomial kernels for mapping features to higher-dimensional spaces. It supports both dense and sparse matrix representations and includes a k-nearest neighbour approximation to reduce kernel matrix size, enabling users to leverage these techniques in classical machine learning workflows.

algorithms/linfa-kernel · high confidence

The new \linfa-nn\ crate provides a pure Rust implementation of nearest neighbor algorithms, including Linear Search, K-D Tree, and Ball Tree. It offers a trait-based interface for building spatial indices and performing k-nearest neighbor and range queries, supporting configurable distance metrics such as L1 (Manhattan), L2 (Euclidean), L-inf (Chebyshev), L-p (Minkowski), and Wasserstein (Earth Mover's) distances. The implementation is designed to integrate with the \linfa\ ecosystem and supports optional serialization via Serde.

algorithms/linfa-nn · high confidence

Introduce max\_features and custom tokenizer to CountVectorizer

The CountVectorizer in linfa-preprocessing now supports limiting the generated vocabulary to the top N most frequent terms via the new \max\_features\ parameter, and allows users to supply a custom tokenization function instead of relying solely on the default regex splitter. These changes enable more controlled feature sets and flexible text processing pipelines.

algorithms/linfa-preprocessing/src/countgrams · high confidence

Introduce named features and targets to the Dataset structure

The dataset module now supports storing human-readable names for both features and targets. Users can assign these names using the new \with\_feature\_names\ and \with\_target\_names\ methods, and retrieve them via \feature\_names\ and \target\_names\. These names are preserved when transforming the dataset (e.g., mapping targets) and are included in views created by iterating over features or targets, improving the interpretability of analysis results.

src/dataset · high confidence

Introduce offset support and checked parameters for logistic regression

The \linfa-logistic\ crate now supports an optional \offset\ parameter in binary and multinomial logistic regression models, allowing users to incorporate prior knowledge or exposure adjustments directly into the linear predictor. Additionally, parameter validation has been strengthened: the \LogisticRegressionParams\ builder now enforces that \alpha\ and \gradient\_tolerance\ are positive and finite, that initial parameters are finite, and that offset lengths match the number of samples, returning specific errors (e.g., \InvalidOffset\, \InvalidAlpha\) instead of failing silently or panicking. These changes improve model flexibility and robustness for users configuring logistic regression.

algorithms/linfa-logistic · high confidence

Introduce t-SNE dimensionality reduction with Barnes-Hut approximation

The \linfa-tsne\ crate is added to the \algorithms/linfa-tsne\ location, providing a pure Rust implementation of t-SNE that wraps the \bhtsne\ library. It supports both exact t-SNE and the Barnes-Hut approximation for scalable visualization of high-dimensional data. Users can configure parameters such as embedding size, perplexity, and approximation threshold via \TSneParams\, which are validated against constraints (e.g., negative values, embedding size exceeding feature count) before transformation. The implementation integrates with the \linfa\ ecosystem by implementing the \Transformer\ trait for both raw arrays and \Dataset\ objects, and includes examples for the Iris and MNIST datasets.

algorithms/linfa-tsne · high confidence

Introduction of linfa-datasets crate with built-in datasets and generation utilities

This change introduces the \linfa-datasets\ crate, providing a new location for dataset handling. It adds functions to load real-world datasets (Iris, Diabetes, Wine Quality, and Linnerud) directly from embedded gzipped CSV files, including support for feature and target names. Additionally, it introduces a \generate\ module (behind the \generate\ feature flag) containing utilities to create synthetic datasets, such as generating data blobs around centroids or creating random datasets with specified statistical distributions.

datasets/src · high confidence

Introduction of linfa-linear crate with OLS, GLM, and Isotonic Regression algorithms

The \linfa-linear\ crate has been introduced to provide pure Rust implementations of classical linear regression algorithms. This release adds Ordinary Least Squares (OLS) regression, which supports configurable intercept fitting and uses either BLAS/LAPACK or a QR decomposition solver depending on feature flags; Generalized Linear Models (GLM); and Isotonic Regression, implemented via the Pool Adjacent Violators Algorithm (PVA). The crate includes a dedicated error handling module (\LinearError\) and a unified \Float\ trait to support both \f32\ and \f64\ precision across these algorithms.

algorithms/linfa-linear/src · high confidence

New dimensionality reduction algorithms and examples

The linfa-reduction crate now includes implementations for Gaussian and Sparse Random Projections, alongside existing Diffusion Mapping and PCA capabilities. This update adds usage examples for all four algorithms, demonstrating how to apply these reduction techniques to datasets like MNIST for downstream tasks such as decision tree training.

algorithms/linfa-reduction · high confidence

New example code for ElasticNet, cross-validation, and Multi-Task ElasticNet

Added three Rust example programs demonstrating how to use the linfa-elasticnet library. The first example shows training a standard ElasticNet (LASSO) model on the Diabetes dataset. The second demonstrates using the new \cross\_validate\_single\ method to compare multiple L1 ratios via k-fold cross-validation. The third example introduces the new Multi-Task ElasticNet capability, training a model on the Linnerud dataset to predict multiple targets simultaneously.

algorithms/linfa-elasticnet/examples · high confidence

New linfa-ensemble crate with AdaBoost and Random Forest

The new \linfa-ensemble\ crate introduces ensemble learning algorithms to the Linfa toolkit, specifically adding AdaBoost and Random Forest classifiers. AdaBoost trains weak learners sequentially to focus on misclassified samples, while Random Forest uses bootstrap aggregation with random feature selection. The crate also supports serde serialization for these models and includes examples and tests demonstrating their usage on the Iris dataset.

algorithms/linfa-ensemble · high confidence

New model composition utilities

The \src/composing\ module introduces four new composition models: \MultiClassModel\ to combine multiple binary classifiers into a single multi-class predictor, \MultiTargetModel\ to merge several single-target models into a multi-target predictor, \Platt\ to calibrate classifier outputs into posterior probabilities using Platt scaling, and \ResidualChain\ to fit models sequentially on the residuals of previous stages (L2Boosting).

src/composing · high confidence

New preprocessing examples for text vectorization and data scaling

Added four new example programs in the linfa-preprocessing crate demonstrating key preprocessing workflows: count vectorization and TF-IDF vectorization for text classification using the 20 Newsgroups dataset, linear scaling (standardization) on the Wine Quality dataset, and PCA-based whitening on the same dataset. These examples show how to integrate CountVectorizer, TfIdfVectorizer, LinearScaler, and Whitener with Gaussian Naive Bayes classifiers.

algorithms/linfa-preprocessing/examples · high confidence

Behavioural changes

Introduce linfa-preprocessing crate with validation and serialization support

The \linfa-preprocessing\ crate now enforces parameter validation for algorithms like \CountVectorizer\ and \TfIdfVectorizer\, returning specific \PreprocessingError\ variants (e.g., \InvalidNGramBoundaries\, \FlippedDocumentFrequencies\) instead of panicking on invalid inputs. Additionally, all public structs and enums in the crate now support \serde\ serialization and deserialization when the \serde\ feature is enabled.

algorithms/linfa-preprocessing/src · high confidence

Test coverage

Add benchmarks for DBSCAN, Gaussian Mixture, and K-Means algorithms; Added Criterion-based benchmarks for preprocessing algorithms; Added Fast ICA benchmark suite with Pprof profiling support; Added OLS and GLM benchmark suite; Added benchmark for Decision Tree training performance; Added benchmarks for PLS algorithms; Added benchmarks for nearest neighbor algorithms; Added integration tests for nearest neighbour algorithms.

Dependencies

Linfa 0.8.1 release with updated dependencies

The Linfa machine learning framework for Rust has been updated to version 0.8.1. This release updates core dependencies including ndarray to 0.16, ndarray-linalg to 0.17, and argmin to 0.11.0, while also upgrading thiserror to 2.0 and criterion to 0.5. The workspace now uses the Cargo resolver version 2.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 61 → 61 (-0.1)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 92 → 92 (+0.0)
  • Architecture 100 → 100 (-0.1)
  • Maturity 56 → 56 (+0.0)
  • Readiness 79 → 81 (+2.6)
  • Security 47 → 47 (+0.0)
  • Performance 90 (new)

Resolved (4)

  • Documentation: no installation or build instructions (algorithms/linfa-elasticnet/README.md)
  • Documentation: no installation or build instructions (algorithms/linfa-ensemble/README.md)
  • Off-boarding risk: anonymized user #1
  • Off-boarding risk: anonymized user #2

New (10)

  • Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
  • Documentation: no project overview (README.md)
  • Inconsistency in parameter retrieval methods. Ftrl has params() and params_with_rng(). Most other models (e.g., NaiveBayes, GaussianNb, ElasticNet) only have params(). The params_with_rng variant suggests that the parameters might be randomized or that the RNG state is part of the 'params' snapshot, which is unusual for a fitted model. If it's just for reproducibility, it should be a separate method or handled differently. This breaks the pattern of params() being the standard accessor.
  • Inconsistent handling of file I/O. CountVectorizer has transform_files but FittedTfIdfVectorizer also has transform_files. However, CountVectorizer does not have a fit_files method shown (only transform_files), while FittedTfIdfVectorizer has fit_files. This suggests CountVectorizer might be a fitted object that can transform files, but the naming CountVectorizer vs FittedTfIdfVectorizer is inconsistent. One is 'Fitted', the other is not. Also, CountVectorizer has force_tokenizer_function_redefinition while FittedTfIdfVectorizer has force_tokenizer_redefinition. The naming difference (function vs no function) is inconsistent.
  • Inconsistent return type for log-probability prediction. predict_log_proba returns Array2<F> (likely log-probabilities) but the second element is Vec<&L> (labels). predict_proba returns Array2<F> (probabilities) and Vec<&L>. While the intent is clear (log vs linear), the naming convention predict_log_proba is slightly non-standard compared to predict_proba/predict_log_proba parity in other libraries, but more importantly, the return type structure is identical, which is good. However, looking at FittedLogisticRegression, it has predict_probabilities. The inconsistency here is subtle: NaiveBayes uses predict_proba/predict_log_proba while FittedLogisticRegression uses predict_probabilities. This is a naming inconsistency across the library for similar operations.
  • Low cohesion: KernelParams (LCOM4 6) (algorithms/linfa-kernel/src/lib.rs)
  • Methods ridge() and lasso() on fitted models return hyperparameter types (ElasticNetParams), not model coefficients or predictions. This is semantically confusing. These methods likely return the optimal hyperparameters selected during fitting, but the name suggests they return the model state for ridge/lasso variants or perhaps the coefficients. If they return hyperparameters, they should be named optimal_ridge_params() or similar. If they return coefficients, the return type is wrong. Given the return type is Params, the name is misleading.
  • Off-boarding risk: anonymized user #2
  • Off-boarding risk: anonymized user #1
  • The LinearScaler methods return LinearScalerParams instead of a fitted scaler or a transformation function. This is a significant API design inconsistency. Typically, a 'Scaler' object, once fitted, provides a transform method. Here, the methods seem to be factory methods for creating params, not fitting the scaler. However, the type is LinearScaler, implying it's a fitted object. If these are meant to be 'get the params used for this scaling', the names standard, min_max are confusing because they sound like constructors. If they are constructors, the type should be LinearScalerBuilder or similar. The inconsistency is that LinearScaler looks like a fitted model but these methods look like builders.

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

rust-ml/linfa was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 3b411a4c985cdf6b83430d9c44f11e338d300c50 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.