rust-ml/linfa
60.8
Adequate · 30 September 2026
32.7k
lines of production code
Rust
primary language
2
measurements over time
What this system is
Linfa is a pure Rust machine learning toolkit designed for classical algorithms, providing a unified API for fitting, transforming, and predicting data. It encompasses a comprehensive suite of capabilities including clustering, linear and non-linear regression, classification, dimensionality reduction, and nearest-neighbor search. The system supports end-to-end workflows by integrating preprocessing utilities, dataset management, and model composition tools within a modular crate structure.
How it got here
2018–2021 — Initial release and algorithm expansion
42 changes.
This period marks the initial public release of the Linfa machine learning toolkit for Rust, establishing the core API, dataset abstractions, and foundational traits. The project rapidly expanded its ecosystem by introducing dedicated crates for a wide range of classical algorithms, including clustering, linear models, dimensionality reduction, and nearest neighbor search. Concurrently, the team built out essential supporting infrastructure such as preprocessing utilities, synthetic dataset generators, and comprehensive benchmarking suites to ensure performance and usability.
2022–2025 — algorithm expansion and benchmarking
8 changes.
This period focused on expanding the Linfa toolkit with new machine learning algorithms, including FTRL, LARS, Random Projection, and ensemble methods like AdaBoost and Random Forest. Concurrently, significant effort was dedicated to establishing a robust benchmarking infrastructure, adding performance suites for linear, ICA, and PLS algorithms, and introducing standardized configuration and profiling support.
Features
Add Diffusion Map non-linear dimensionality reduction
Introduces the Diffusion Map algorithm as a new non-linear dimensionality reduction technique. This feature computes a data embedding by applying PCA on the diffusion operator of a kernel matrix, allowing users to project data along the direction of the largest diffusion flow. The implementation includes parameter validation for embedding size and steps, and supports both BLAS-accelerated and standard eigenvalue decomposition paths depending on feature flags.
_algorithms/linfa-reduction/src/diffusion\map · high confidence
Add Least Angle Regression (LARS) algorithm
The \linfa-lars\ crate introduces a pure Rust implementation of the Least Angle Regression (LARS) algorithm, providing a new linear regression capability within the \linfa\ ecosystem. Users can now fit LARS models using \Lars::params()\ with configurable options such as intercept fitting, verbosity, and maximum non-zero coefficients, and generate predictions via the standard \predict\ interface. The implementation supports optional BLAS/LAPACK backends for performance and includes a usage example demonstrating fitting on the Diabetes dataset.
algorithms/linfa-lars · high confidence
Add OPTICS clustering algorithm
Introduces the OPTICS (Ordering Points To Identify Clustering Structure) algorithm to the linfa-clustering library. This new feature allows users to perform density-based clustering analysis that generates an augmented ordering of data points, enabling the identification of clusters with varying densities. The implementation includes configurable hyperparameters such as minimum points, distance metrics, and nearest neighbor algorithms, along with serialization support when the 'serde' feature is enabled.
algorithms/linfa-clustering/src/optics · high confidence
Add PLS regression example
A new example demonstrating how to use the PlsRegression algorithm from the linfa-pls crate has been added. The example generates synthetic data, fits a PLS regression model with scaling enabled, and prints both the true and estimated regression coefficients to illustrate the model's performance.
algorithms/linfa-pls/examples · high confidence
Add Tweedie Regressor with configurable distributions and link functions
Users can now fit Generalized Linear Models using the Tweedie distribution via the new \TweedieRegressor\. This implementation supports configurable power parameters to model Normal (power=0), Poisson (power=1), Gamma (power=2), and Inverse Gaussian (power=3) distributions, along with other Tweedie variants. The regressor includes built-in validation for target ranges based on the selected power, automatic link function selection (Identity for Normal, Log for others), and uses the L-BFGS solver for optimization. Hyperparameters such as regularization strength (\alpha\), maximum iterations, and tolerance are exposed through a builder-style API.
algorithms/linfa-linear/src/glm · high confidence
Add approximate DBSCAN clustering algorithm
Introduces the \AppxDbscan\ algorithm, an O(N) approximation of the standard DBSCAN clustering method. This implementation partitions the data space into a grid of cells and uses a counting tree structure for efficient range queries, allowing for scalable clustering on larger datasets. The algorithm supports configurable hyperparameters including \min\_points\, \tolerance\, \slack\, and a pluggable nearest-neighbour algorithm via the \linfa-nn\ integration. It returns cluster labels for core points and assigns border points to clusters or labels them as noise, with optional serialization support for the algorithm parameters.
_algorithms/linfa-clustering/src/appx\dbscan · high confidence
Add clustering algorithm examples (DBSCAN, K-Means, OPTICS)
New example scripts have been added to demonstrate how to use the DBSCAN, K-Means, and OPTICS clustering algorithms. These examples show how to generate synthetic datasets, configure and run the algorithms, and save the resulting cluster assignments and data to NumPy-compatible files for further analysis or visualization.
algorithms/linfa-clustering/examples · high confidence
Add diabetes and GLM regression examples
New example programs have been added to the linfa-linear crate to demonstrate usage of linear regression models. The diabetes.rs example shows how to fit a standard LinearRegression model to the diabetes dataset and print the resulting intercept and parameters. The glm.rs example demonstrates fitting a TweedieRegressor (configured as a standard linear regression by setting power and alpha to zero) and includes code to calculate and print the Mean Absolute Error on the training data.
algorithms/linfa-linear/examples · high confidence
Added benchmark configuration module with default settings
A new \src/benchmarks/mod.rs\ file introduces a benchmark configuration module that defines default settings for the Criterion benchmarking framework. This includes a sample size of 200, a measurement time of 10 seconds, a confidence level of 0.97, a warm-up time of 10 seconds, and a noise threshold of 0.05. On non-Windows platforms, it also provides a function to configure PProf profiling for generating flamegraphs.
src/benchmarks · high confidence
Added logistic regression examples for single and multi-class classification
New example files demonstrate how to use the \LogisticRegression\ and \MultiLogisticRegression\ classifiers from the \linfa-logistic\ crate. The \logistic\_cv.rs\ example shows how to perform cross-validation on a single-target (binary) classification task using the wine quality dataset, while \winequality\_logistic.rs\ and \winequality\_multi\_logistic.rs\ provide basic fit/predict workflows for binary and multi-class scenarios respectively, including confusion matrix and accuracy reporting.
algorithms/linfa-logistic/examples · high confidence
Decision tree example now supports entropy-based splitting and LaTeX export
The decision tree example in the linfa-trees package has been updated to demonstrate training models using the entropy split quality criterion, in addition to the existing Gini criterion. The example also now exports the trained Gini tree structure to a LaTeX/TikZ file for visualization, allowing users to generate visual representations of the decision tree directly from the example output.
algorithms/linfa-trees/examples · high confidence
Initial project release and documentation
This entry establishes the initial public release of the Linfa machine learning toolkit for Rust. It introduces the core project documentation, including a comprehensive CHANGELOG.md tracking versions from 0.3.0 through 0.8.1, a CONTRIBUTE.md guide for algorithm implementation, and an updated README.md featuring the project mascot, badges, and a detailed table of supported algorithms (such as Naive Bayes, K-Means, and Random Forest). The release also formalizes the project's dual MIT/Apache-2.0 licensing structure and configures the build system to link against LAPACK/CBLAS libraries when specific backend features are enabled.
(repo-wide) · high confidence
Initial release of linfa-clustering crate with core algorithms
The linfa-clustering crate is introduced to provide pure Rust implementations of popular clustering algorithms, including K-Means, DBSCAN, Gaussian Mixture Models, and OPTICS. As part of the linfa ecosystem, this new module exposes these algorithms for use in classical machine learning tasks. Notably, the approximate DBSCAN implementation is currently provided as a type alias for the standard DBSCAN to leverage its superior performance and avoid outdated dependencies.
algorithms/linfa-clustering/src · high confidence
Initial release of linfa-svm for Support Vector Machines
The \linfa-svm\ crate is introduced, providing a pure Rust implementation of Support Vector Machines for classification and regression tasks. It includes binary classification (C and Nu variants), one-class classification, and support vector regression (Epsilon and Nu variants), utilizing the Sequential Minimal Optimization (SMO) solver. The API supports various kernel methods (Linear, Gaussian, Polynomial) and optional Platt scaling for probability calibration, with example usage provided for wine quality classification and noisy sine regression.
algorithms/linfa-svm · high confidence
Initial release of the Linfa machine learning toolkit
This change introduces the core Linfa library, providing a comprehensive Rust toolkit for classical machine learning tasks. It establishes the foundational API through key traits for fitting models (\Fit\, \FitWith\), transforming data (\Transformer\), and making predictions (\Predict\, \PredictInplace\). The release includes a robust dataset abstraction (\DatasetBase\) and a parameter validation system (\ParamGuard\) to ensure hyperparameter safety. Users gain access to essential evaluation metrics for classification (including confusion matrices, precision, recall, and F1-score), regression (MAE, MSE, R², MAPE, sMAPE), and clustering (Silhouette Score), along with a new Pearson correlation analysis module for feature inspection. The library is structured to support various algorithm categories such as linear models, clustering, and dimensionality reduction via dedicated sub-crates.
src · high confidence
Initial release of the linfa-reduction crate with PCA and data generation utilities
This change introduces the \linfa-reduction\ crate, providing Principal Component Analysis (PCA) for dimensionality reduction via SVD, along with configurable parameters for whitening and embedding size. The PCA implementation exposes key model attributes such as components, mean, singular values, and explained variance ratios. Additionally, the crate includes utility functions for generating synthetic test datasets, specifically swiss rolls and convoluted rings in 2D and 3D spaces, and defines a comprehensive error type (\ReductionError\) to handle validation issues like insufficient samples or invalid dimensions.
algorithms/linfa-reduction/src · high confidence
Introduce DBSCAN clustering algorithm with configurable nearest-neighbor search
Users can now perform density-based spatial clustering of applications with noise (DBSCAN) using the \linfa\_clustering\ crate. The implementation supports configurable hyperparameters including minimum points, distance tolerance, custom distance metrics (defaulting to L2/Euclidean), and pluggable nearest-neighbor algorithms (defaulting to KdTree via \linfa\_nn\). The API provides a builder pattern for parameter configuration and returns cluster assignments as \Option\<usize\>\ values, where \None\ indicates noise points.
algorithms/linfa-clustering/src/dbscan · high confidence
Introduce Elastic Net and Multi-Task Elastic Net regression algorithms
The \linfa-elasticnet\ crate now provides pure Rust implementations of Elastic Net linear regression, combining L1 (LASSO) and L2 (Ridge) penalties via coordinate descent. This release adds support for both standard single-target regression and multi-task regression (multiple output variables) using block coordinate descent. Users can configure models with checked hyperparameters (penalty, l1\_ratio, intercept, tolerance, max\_iterations) and access fitted model details such as the hyperplane, intercept, duality gap, and Z-scores.
algorithms/linfa-elasticnet · high confidence
Introduce FastICA algorithm for Independent Component Analysis
The \linfa-ica\ crate now provides a pure Rust implementation of the Fast Independent Component Analysis (FastICA) algorithm, enabling users to separate mixed multivariate signals into their independent subcomponents. The implementation supports configurable hyperparameters such as the number of components, the neg-entropy approximation function (GFunc), maximum iterations, tolerance, and a random seed for reproducibility. It integrates with the \linfa\ ecosystem traits (\Fit\, \Predict\) and includes an example demonstrating signal unmixing.
algorithms/linfa-ica · high confidence
Introduce Follow The Regularized Leader (FTRL) Proximal model
The \linfa-ftrl\ crate is added to the ecosystem, providing a pure Rust implementation of the Follow The Regularized Leader - Proximal linear model, primarily designed for click-through rate (CTR) prediction in online learning settings. The implementation supports L1 and L2 regularization, stores internal \z\ and \n\ values for incremental weight calculation, and exposes a builder API for hyperparameters like \alpha\, \beta\, and regularization ratios. This release includes the core algorithm logic, parameter validation, error handling, a usage example on the winequality dataset, and benchmarks for training and prediction performance.
algorithms/linfa-ftrl · high confidence
Introduce Gaussian Mixture Model clustering algorithm
Adds a new Gaussian Mixture Model (GMM) implementation to the linfa-clustering library, enabling users to cluster data by fitting a mixture of Gaussian distributions using the Expectation-Maximization (EM) algorithm. The feature includes configurable hyperparameters for the number of clusters, convergence tolerance, regularization, and initialization methods (KMeans or random), along with error handling for common fitting issues like empty clusters or non-convergence.
_algorithms/linfa-clustering/src/gaussian\mixture · high confidence
Introduce K-means clustering with multiple algorithm variants and initialization strategies
Users can now perform K-means clustering using the \linfa-clustering\ crate. The implementation supports standard Lloyd's algorithm and Hamerly's accelerated variant for the assignment step, with Hamerly offering performance gains for well-separated clusters by skipping redundant distance computations. Centroid initialization is configurable via K-means++, K-means\|\| (optimized for large cluster counts), random selection, or precomputed centroids. The API allows parametrization of the distance metric, convergence tolerance, maximum iterations, and number of runs, with validation ensuring valid hyperparameters. The model supports serialization via serde when the feature is enabled.
_algorithms/linfa-clustering/src/k\means · high confidence
Introduce Naive Bayes classifiers to linfa-bayes
The \linfa-bayes\ crate now provides three Naive Bayes algorithm implementations: Gaussian, Multinomial, and Bernoulli. This release adds the core model structures, hyper-parameter validation (including smoothing parameters), and example scripts for each variant, enabling users to perform classification tasks using these probabilistic models within the Linfa toolkit.
algorithms/linfa-bayes · high confidence
Introduce Partial Least Squares (PLS) algorithms
The \linfa-pls\ crate now provides a pure Rust implementation of the Partial Least Squares algorithm family, ported from scikit-learn 0.24. Users can now fit and predict using \PlsRegression\, \PlsCanonical\, and \PlsCca\ models. The implementation supports configurable hyperparameters such as the number of components, scaling, and algorithm choice (NIPALS or SVD), and includes traits for fitting, transforming data, and predicting targets.
algorithms/linfa-pls · high confidence
Introduce Random Projection for dimensionality reduction
Added a new Random Projection reduction module that embeds data into a lower-dimensional space using random matrices, offering a computationally efficient alternative to methods like PCA. The implementation supports two projection strategies: Gaussian (dense) and Sparse (using the Achlioptas distribution), allowing users to choose based on their data sparsity and performance needs. Users can configure the target embedding dimension directly or specify a precision parameter (epsilon) which automatically calculates the required dimension via the Johnson-Lindenstrauss lemma. The module includes parameter validation to prevent dimension increases and supports custom random number generators for reproducible results.
_algorithms/linfa-reduction/src/random\projection · high confidence
Introduce agglomerative hierarchical clustering algorithm
Users can now perform agglomerative hierarchical clustering via the new \linfa-hierarchical\ crate. This feature allows clustering data based on a similarity kernel, supporting configurable merging methods (via the \kodama\ crate) and stopping criteria such as a target number of clusters or a maximum distance threshold. An example demonstrating clustering on the Iris dataset is also provided.
algorithms/linfa-hierarchical · high confidence
Introduce decision tree learning algorithm
The \linfa-trees\ crate now provides a decision tree implementation for classification, allowing users to fit models with configurable hyperparameters such as split quality (Gini or Entropy), maximum depth, and minimum sample weights. The API supports exporting the fitted tree structure to TikZ/LaTeX for visualization and includes an efficient level-order node iterator.
algorithms/linfa-trees · high confidence
Introduce kernel methods for dimensionality expansion
The linfa-kernel crate is added to the ecosystem, providing implementations of kernel methods such as radial basis function (RBF) and polynomial kernels for mapping features to higher-dimensional spaces. It supports both dense and sparse matrix representations and includes a k-nearest neighbour approximation to reduce kernel matrix size, enabling users to leverage these techniques in classical machine learning workflows.
algorithms/linfa-kernel · high confidence
Introduce linfa-nn for nearest neighbor search
The new \linfa-nn\ crate provides a pure Rust implementation of nearest neighbor algorithms, including Linear Search, K-D Tree, and Ball Tree. It offers a trait-based interface for building spatial indices and performing k-nearest neighbor and range queries, supporting configurable distance metrics such as L1 (Manhattan), L2 (Euclidean), L-inf (Chebyshev), L-p (Minkowski), and Wasserstein (Earth Mover's) distances. The implementation is designed to integrate with the \linfa\ ecosystem and supports optional serialization via Serde.
algorithms/linfa-nn · high confidence
Introduce max\_features and custom tokenizer to CountVectorizer
The CountVectorizer in linfa-preprocessing now supports limiting the generated vocabulary to the top N most frequent terms via the new \max\_features\ parameter, and allows users to supply a custom tokenization function instead of relying solely on the default regex splitter. These changes enable more controlled feature sets and flexible text processing pipelines.
algorithms/linfa-preprocessing/src/countgrams · high confidence
Introduce named features and targets to the Dataset structure
The dataset module now supports storing human-readable names for both features and targets. Users can assign these names using the new \with\_feature\_names\ and \with\_target\_names\ methods, and retrieve them via \feature\_names\ and \target\_names\. These names are preserved when transforming the dataset (e.g., mapping targets) and are included in views created by iterating over features or targets, improving the interpretability of analysis results.
src/dataset · high confidence
Introduce offset support and checked parameters for logistic regression
The \linfa-logistic\ crate now supports an optional \offset\ parameter in binary and multinomial logistic regression models, allowing users to incorporate prior knowledge or exposure adjustments directly into the linear predictor. Additionally, parameter validation has been strengthened: the \LogisticRegressionParams\ builder now enforces that \alpha\ and \gradient\_tolerance\ are positive and finite, that initial parameters are finite, and that offset lengths match the number of samples, returning specific errors (e.g., \InvalidOffset\, \InvalidAlpha\) instead of failing silently or panicking. These changes improve model flexibility and robustness for users configuring logistic regression.
algorithms/linfa-logistic · high confidence
Introduce t-SNE dimensionality reduction with Barnes-Hut approximation
The \linfa-tsne\ crate is added to the \algorithms/linfa-tsne\ location, providing a pure Rust implementation of t-SNE that wraps the \bhtsne\ library. It supports both exact t-SNE and the Barnes-Hut approximation for scalable visualization of high-dimensional data. Users can configure parameters such as embedding size, perplexity, and approximation threshold via \TSneParams\, which are validated against constraints (e.g., negative values, embedding size exceeding feature count) before transformation. The implementation integrates with the \linfa\ ecosystem by implementing the \Transformer\ trait for both raw arrays and \Dataset\ objects, and includes examples for the Iris and MNIST datasets.
algorithms/linfa-tsne · high confidence
Introduction of linfa-datasets crate with built-in datasets and generation utilities
This change introduces the \linfa-datasets\ crate, providing a new location for dataset handling. It adds functions to load real-world datasets (Iris, Diabetes, Wine Quality, and Linnerud) directly from embedded gzipped CSV files, including support for feature and target names. Additionally, it introduces a \generate\ module (behind the \generate\ feature flag) containing utilities to create synthetic datasets, such as generating data blobs around centroids or creating random datasets with specified statistical distributions.
datasets/src · high confidence
Introduction of linfa-linear crate with OLS, GLM, and Isotonic Regression algorithms
The \linfa-linear\ crate has been introduced to provide pure Rust implementations of classical linear regression algorithms. This release adds Ordinary Least Squares (OLS) regression, which supports configurable intercept fitting and uses either BLAS/LAPACK or a QR decomposition solver depending on feature flags; Generalized Linear Models (GLM); and Isotonic Regression, implemented via the Pool Adjacent Violators Algorithm (PVA). The crate includes a dedicated error handling module (\LinearError\) and a unified \Float\ trait to support both \f32\ and \f64\ precision across these algorithms.
algorithms/linfa-linear/src · high confidence
New dimensionality reduction algorithms and examples
The linfa-reduction crate now includes implementations for Gaussian and Sparse Random Projections, alongside existing Diffusion Mapping and PCA capabilities. This update adds usage examples for all four algorithms, demonstrating how to apply these reduction techniques to datasets like MNIST for downstream tasks such as decision tree training.
algorithms/linfa-reduction · high confidence
New example code for ElasticNet, cross-validation, and Multi-Task ElasticNet
Added three Rust example programs demonstrating how to use the linfa-elasticnet library. The first example shows training a standard ElasticNet (LASSO) model on the Diabetes dataset. The second demonstrates using the new \cross\_validate\_single\ method to compare multiple L1 ratios via k-fold cross-validation. The third example introduces the new Multi-Task ElasticNet capability, training a model on the Linnerud dataset to predict multiple targets simultaneously.
algorithms/linfa-elasticnet/examples · high confidence
New linfa-ensemble crate with AdaBoost and Random Forest
The new \linfa-ensemble\ crate introduces ensemble learning algorithms to the Linfa toolkit, specifically adding AdaBoost and Random Forest classifiers. AdaBoost trains weak learners sequentially to focus on misclassified samples, while Random Forest uses bootstrap aggregation with random feature selection. The crate also supports serde serialization for these models and includes examples and tests demonstrating their usage on the Iris dataset.
algorithms/linfa-ensemble · high confidence
New model composition utilities
The \src/composing\ module introduces four new composition models: \MultiClassModel\ to combine multiple binary classifiers into a single multi-class predictor, \MultiTargetModel\ to merge several single-target models into a multi-target predictor, \Platt\ to calibrate classifier outputs into posterior probabilities using Platt scaling, and \ResidualChain\ to fit models sequentially on the residuals of previous stages (L2Boosting).
src/composing · high confidence
New preprocessing examples for text vectorization and data scaling
Added four new example programs in the linfa-preprocessing crate demonstrating key preprocessing workflows: count vectorization and TF-IDF vectorization for text classification using the 20 Newsgroups dataset, linear scaling (standardization) on the Wine Quality dataset, and PCA-based whitening on the same dataset. These examples show how to integrate CountVectorizer, TfIdfVectorizer, LinearScaler, and Whitener with Gaussian Naive Bayes classifiers.
algorithms/linfa-preprocessing/examples · high confidence
Behavioural changes
Introduce linfa-preprocessing crate with validation and serialization support
The \linfa-preprocessing\ crate now enforces parameter validation for algorithms like \CountVectorizer\ and \TfIdfVectorizer\, returning specific \PreprocessingError\ variants (e.g., \InvalidNGramBoundaries\, \FlippedDocumentFrequencies\) instead of panicking on invalid inputs. Additionally, all public structs and enums in the crate now support \serde\ serialization and deserialization when the \serde\ feature is enabled.
algorithms/linfa-preprocessing/src · high confidence
Test coverage
Add benchmarks for DBSCAN, Gaussian Mixture, and K-Means algorithms; Added Criterion-based benchmarks for preprocessing algorithms; Added Fast ICA benchmark suite with Pprof profiling support; Added OLS and GLM benchmark suite; Added benchmark for Decision Tree training performance; Added benchmarks for PLS algorithms; Added benchmarks for nearest neighbor algorithms; Added integration tests for nearest neighbour algorithms.
Dependencies
Linfa 0.8.1 release with updated dependencies
The Linfa machine learning framework for Rust has been updated to version 0.8.1. This release updates core dependencies including ndarray to 0.16, ndarray-linalg to 0.17, and argmin to 0.11.0, while also upgrading thiserror to 2.0 and criterion to 0.5. The workspace now uses the Cargo resolver version 2.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 61 → 61 (-0.1)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 92 → 92 (+0.0)
- Architecture 100 → 100 (-0.1)
- Maturity 56 → 56 (+0.0)
- Readiness 79 → 81 (+2.6)
- Security 47 → 47 (+0.0)
- Performance 90 (new)
Resolved (4)
- Documentation: no installation or build instructions (algorithms/linfa-elasticnet/README.md)
- Documentation: no installation or build instructions (algorithms/linfa-ensemble/README.md)
- Off-boarding risk: anonymized user #1
- Off-boarding risk: anonymized user #2
New (10)
- Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
- Documentation: no project overview (README.md)
- Inconsistency in parameter retrieval methods. Ftrl has params() and params_with_rng(). Most other models (e.g., NaiveBayes, GaussianNb, ElasticNet) only have params(). The params_with_rng variant suggests that the parameters might be randomized or that the RNG state is part of the 'params' snapshot, which is unusual for a fitted model. If it's just for reproducibility, it should be a separate method or handled differently. This breaks the pattern of params() being the standard accessor.
- Inconsistent handling of file I/O. CountVectorizer has transform_files but FittedTfIdfVectorizer also has transform_files. However, CountVectorizer does not have a fit_files method shown (only transform_files), while FittedTfIdfVectorizer has fit_files. This suggests CountVectorizer might be a fitted object that can transform files, but the naming CountVectorizer vs FittedTfIdfVectorizer is inconsistent. One is 'Fitted', the other is not. Also, CountVectorizer has force_tokenizer_function_redefinition while FittedTfIdfVectorizer has force_tokenizer_redefinition. The naming difference (function vs no function) is inconsistent.
- Inconsistent return type for log-probability prediction. predict_log_proba returns Array2<F> (likely log-probabilities) but the second element is Vec<&L> (labels). predict_proba returns Array2<F> (probabilities) and Vec<&L>. While the intent is clear (log vs linear), the naming convention predict_log_proba is slightly non-standard compared to predict_proba/predict_log_proba parity in other libraries, but more importantly, the return type structure is identical, which is good. However, looking at FittedLogisticRegression, it has predict_probabilities. The inconsistency here is subtle: NaiveBayes uses predict_proba/predict_log_proba while FittedLogisticRegression uses predict_probabilities. This is a naming inconsistency across the library for similar operations.
- Low cohesion: KernelParams (LCOM4 6) (algorithms/linfa-kernel/src/lib.rs)
- Methods ridge() and lasso() on fitted models return hyperparameter types (ElasticNetParams), not model coefficients or predictions. This is semantically confusing. These methods likely return the optimal hyperparameters selected during fitting, but the name suggests they return the model state for ridge/lasso variants or perhaps the coefficients. If they return hyperparameters, they should be named optimal_ridge_params() or similar. If they return coefficients, the return type is wrong. Given the return type is Params, the name is misleading.
- Off-boarding risk: anonymized user #2
- Off-boarding risk: anonymized user #1
- The LinearScaler methods return LinearScalerParams instead of a fitted scaler or a transformation function. This is a significant API design inconsistency. Typically, a 'Scaler' object, once fitted, provides a transform method. Here, the methods seem to be factory methods for creating params, not fitting the scaler. However, the type is LinearScaler, implying it's a fitted object. If these are meant to be 'get the params used for this scaling', the names standard, min_max are confusing because they sound like constructors. If they are constructors, the type should be LinearScalerBuilder or similar. The inconsistency is that LinearScaler looks like a fitted model but these methods look like builders.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
rust-ml/linfa was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3b411a4c985cdf6b83430d9c44f11e338d300c50 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.