elixir-nx/scholar
68.2
Adequate · 3 October 2026
20.6k
lines of production code
Elixir
primary language
2
measurements over time
What this system is
This system is a machine learning library for Elixir that provides a comprehensive suite of algorithms for data preprocessing, statistical analysis, and model training. It supports diverse tasks including clustering, dimensionality reduction, classification, regression, and optimization, with a strong emphasis on numerical accuracy and JIT compilation for GPU compatibility. The library includes extensive test coverage and benchmarking tools to ensure reliability across its various modules.
How it got here
2022 — comprehensive feature expansion and JIT compatibility
21 changes.
The project significantly expanded its machine learning capabilities by introducing new modules for model selection, preprocessing, statistics, linear models, clustering, interpolation, decomposition, imputation, and Naive Bayes classifiers. This feature growth was accompanied by a major dependency upgrade to Nx 1.0 and a comprehensive overhaul of documentation and test coverage to ensure JIT and GPU compatibility via EXLA.
2023–2024 — expansion of statistical and ML modules
14 changes.
This period focused on significantly expanding the library's machine learning and statistical capabilities by introducing new modules for manifold learning, numerical integration, k-NN search, preprocessing, covariance estimation, and cross-decomposition. Each new feature was accompanied by comprehensive test suites to ensure numerical accuracy and stability. The work also included performance benchmarking for neighbor search algorithms and configuration adjustments for the EXLA backend.
2025–2026 — expansion of feature extraction and optimization algorithms
6 changes.
This period focused on expanding the Scholar library's capabilities by introducing CountVectorizer for token count matrix generation and adding four optimization algorithms including BFGS and Nelder-Mead. Additionally, Linear and Quadratic Discriminant Analysis classifiers were implemented to support classification tasks with Gaussian density fitting. Comprehensive test suites were developed for all new modules to ensure numerical accuracy and correct behavior across various input scenarios.
Features
Add BFGS, Nelder-Mead, Golden Section, and Brent optimization algorithms
The \lib/scholar/optimize\ module now includes four new optimization algorithms: BFGS for multivariate function minimization using gradient information, Nelder-Mead for derivative-free multivariate optimization, Golden Section search for univariate minimization, and Brent's method which combines golden section search with parabolic interpolation for robust univariate minimization. These additions expand the library's capabilities to handle both single-variable and multi-variable optimization problems with various convergence properties.
lib/scholar/optimize · high confidence
Add Bezier, Cubic, Linear, and Monotonic Cubic Spline interpolation algorithms
The \lib/scholar/interpolation\ module now includes four new interpolation algorithms: \BezierSpline\ (cubic Bezier with local control), \CubicSpline\ (with configurable boundary conditions like \:not\_a\_knot\ or \:natural\), \Linear\ (with configurable left/right extrapolation values), and \MonotonicCubicSpline\ (PCHIP, which preserves monotonicity to avoid overshoot). Each algorithm provides a \fit/2\ function to train on \(x, y)\ data and a \predict/2\ or \predict/3\ function to evaluate new points, exposing these capabilities directly to users for curve fitting and data interpolation tasks.
lib/scholar/interpolation · high confidence
Add KNN and Simple imputers to Scholar
The \lib/scholar/impute\ module now includes two new imputation strategies: \KNNImputer\, which fills missing values using the mean of k-nearest neighbors, and \SimpleImputer\, which supports univariate strategies like mean, median, mode, and constant fill values. These additions provide users with flexible options for handling missing data in their datasets.
lib/scholar/impute · high confidence
Add Ledoit-Wolf and shrunk covariance estimators
Users can now estimate covariance matrices using shrinkage techniques to improve stability in high-dimensional settings. The library introduces \Scholar.Covariance.LedoitWolf\, which automatically computes the optimal shrinkage coefficient based on the Ledoit-Wolf formula, and \Scholar.Covariance.ShrunkCovariance\, which allows users to specify a fixed shrinkage coefficient. Both estimators return the shrunk covariance matrix along with the estimated location (mean), and support an \assume\_centered?\ option to skip data centering when the mean is known to be zero.
lib/scholar/covariance · high confidence
Add MDS, Trimap, and t-SNE manifold learning algorithms
Users can now perform dimensionality reduction using three new algorithms: Multidimensional Scaling (MDS), which preserves pairwise distances in a low-dimensional embedding; Trimap, which uses triplet constraints to maintain global data structure; and t-SNE, a nonlinear technique for visualizing high-dimensional data. These new modules in \Scholar.Manifold\ provide additional options for exploring and visualizing complex datasets beyond existing methods.
lib/scholar/manifold · high confidence
Add Partial Least Squares SVD transformer
A new \PLSSVD\ module has been added to the cross-decomposition library, implementing Partial Least Squares via Singular Value Decomposition. This transformer allows users to fit a model on training data and targets, projecting both onto their respective singular vectors to perform dimensionality reduction. The implementation supports configurable options for the number of components and data scaling, providing a new capability for multivariate statistical analysis within the library.
_lib/scholar/cross\decomposition · high confidence
Add Simpson's rule numerical integration
The \Scholar.Integrate\ module now includes a \simpson/3\ function that performs numerical integration using Simpson's rule. This new capability allows users to integrate data along a specified axis with options to handle even-length intervals (averaging, first, or last) and control axis reduction behavior, providing an alternative to the existing trapezoidal rule for potentially higher accuracy in certain scenarios.
lib/scholar/integrate · high confidence
Added CountVectorizer for converting indexed text to token count matrices
A new \CountVectorizer\ module has been added to the feature extraction library, allowing users to convert a collection of already indexed text documents into a matrix of token counts. This component accepts a 2D tensor where rows represent documents and integer values represent tokens, supporting padding via negative values, and outputs a count matrix where each column corresponds to a unique token in the vocabulary.
_lib/scholar/feature\extraction · high confidence
Added Linear and Quadratic Discriminant Analysis classifiers
The library now includes two new classification models in the \Scholar.DiscriminantAnalysis\ module. The \Linear\ module implements Linear Discriminant Analysis (LDA), which fits a Gaussian density to each class assuming a shared covariance matrix, resulting in linear decision boundaries. The \Quadratic\ module implements Quadratic Discriminant Analysis (QDA), which allows each class to have its own covariance matrix, creating quadratic decision boundaries. Both models support fitting, prediction, and probability estimation, with QDA offering a \reg\_param\ option to handle singular covariance matrices.
_lib/scholar/discriminant\analysis · high confidence
New Naive Bayes classifier variants added
The library now includes four new Naive Bayes implementations: BernoulliNB for binary/boolean features, CategoricalNB for discrete categorical features, ComplementNB for imbalanced datasets, and MultinomialNB for occurrence counts. Additionally, the GaussianNB module has been restructured and renamed from the previous gaussian\_naive\_bayes.ex file.
_lib/scholar/naive\bayes · high confidence
New and updated machine learning notebooks
Added new Livebook examples for cross-validation with gradient boosting trees, portfolio optimization using the efficient frontier, hierarchical clustering with a custom dendrogram visualization, and K-means clustering on the Iris dataset. Updated existing notebooks to use Tucan for plotting, upgraded dependencies (including Nx, EXGBoost, and Tucan), and refined content for clarity and correctness.
notebooks · high confidence
New clustering algorithms added to Scholar.Cluster
This release introduces nine new clustering algorithms to the \Scholar.Cluster\ module: Affinity Propagation, DBSCAN, Gaussian Mixture Models, HDBSCAN, Hierarchical Clustering, K-Means, Mean Shift, OPTICS, and Spectral Clustering. These additions expand the library's capabilities to support density-based, hierarchical, probabilistic, and spectral clustering techniques, allowing users to choose algorithms best suited for their specific data distributions and cluster shapes.
lib/scholar/cluster · high confidence
New decomposition algorithms: Kernel PCA, Truncated SVD, and incremental PCA
The \lib/scholar/decomposition\ module now includes three new dimensionality reduction algorithms alongside an updated PCA implementation. Users can perform non-linear dimensionality reduction using \Scholar.Decomposition.KernelPCA\, which supports linear, polynomial, RBF, sigmoid, and cosine kernels. A new \Scholar.Decomposition.TruncatedSVD\ module provides linear dimensionality reduction via randomized SVD, suitable for sparse data. Additionally, the existing \Scholar.Decomposition.PCA\ has been enhanced with \incremental\_fit/2\ and \partial\_fit/2\ methods, allowing users to train on datasets that are too large to fit in memory by processing data in batches.
lib/scholar/decomposition · high confidence
New k-NN neighbor search algorithms and classification/regression models
The \lib/scholar/neighbors\ area now includes a comprehensive suite of k-Nearest Neighbor (k-NN) capabilities. Users can perform neighbor searches using several new algorithms: \BruteKNN\ for exact brute-force search, \KDTree\ for space-partitioning based search, \RandomProjectionForest\ for approximate search via random projections, \NNDescent\ for approximate nearest neighbor graph construction, and \LargeVis\ for large-scale graph construction. Additionally, the library introduces \KNNClassifier\ and \KNNRegressor\ for classification and regression tasks respectively, which wrap these underlying search algorithms and support configurable weighting (uniform or distance-based). A \RadiusNNClassifier\ and \RadiusNNRegressor\ are also added for radius-based neighborhood queries. These components are supported by a \Utils\ module for metric handling and common search utilities.
lib/scholar/neighbors · high confidence
New linear regression models and utilities
Added new linear modeling capabilities to Scholar.Linear, including BayesianRidgeRegression, IsotonicRegression, LinearRegression, LogisticRegression, PolynomialRegression, RidgeRegression, and SVM. These models provide users with a range of regression and classification algorithms, supporting features such as sample weighting, intercept fitting, and configurable solvers. A new LinearHelpers module was introduced to support these implementations with common utilities for data preprocessing and weight handling.
lib/scholar/linear · high confidence
New model selection, preprocessing, and statistical modules
This release introduces three new modules to the Scholar library. The new \Scholar.ModelSelection\ module provides utilities for model evaluation, including K-fold splitting, cross-validation (with support for weighted samples), and grid search for hyperparameter tuning. The \Scholar.Preprocessing\ module adds data transformation capabilities, offering functions for standard scaling, max-abs scaling, min-max scaling, binarization, ordinal encoding, one-hot encoding, and normalization. Additionally, the \Scholar.Stats\ module introduces statistical analysis functions, including calculations for moments, skewness, kurtosis, and correlation matrices.
lib/scholar · high confidence
New modular metrics library with classification, regression, and distance functions
The \lib/scholar/metrics\ directory has been restructured into distinct modules (\Classification\, \Regression\, \Clustering\, \Distance\, \Neighbors\, \Ranking\, \Similarity\), introducing a comprehensive suite of evaluation metrics. Users can now access classification metrics like F1, F-beta, precision/recall, confusion matrix, and log loss; regression metrics including MAE, MSE, R2, and Mean Tweedie Deviance; distance metrics such as Euclidean, Manhattan, and Minkowski; and clustering/ranking metrics like Silhouette Score and NDCG. All functions are implemented as numerical functions compatible with Nx compilers.
lib/scholar/metrics · high confidence
New preprocessing utilities: encoders, scalers, and normalizers
The \lib/scholar/preprocessing\ module now includes several new data transformation tools. Users can binarize data with \Binarizer\, scale features using \MaxAbsScaler\, \MinMaxScaler\, \RobustScaler\, or \StandardScaler\, and normalize tensors to unit norm with \Normalizer\. Additionally, categorical data can be encoded using \OrdinalEncoder\ and \OneHotEncoder\. These modules provide \fit\, \transform\, and \fit\_transform\ interfaces for consistent usage.
lib/scholar/preprocessing · high confidence
Removals
Removal of empty Scholar module
The empty Scholar module has been removed from the library. This change eliminates an unused code artifact, simplifying the codebase without affecting any existing functionality.
lib · high confidence
Behavioural changes
Conditional EXLA backend configuration for compile-time and test environments
The application now conditionally configures the EXLA backend based on the environment. By default, EXLA is not used as the default backend for Nx operations, but setting the \USE\_EXLA\_AT\_COMPILE\_TIME\ environment variable enables it. Additionally, the \add\_backend\_on\_inspect\ setting for EXLA is disabled during test runs to optimize performance.
config · high confidence
Documentation overhaul and contributor guidelines for JIT/GPU compatibility
The project documentation has been significantly expanded to clarify installation requirements, specifically mandating the use of a JIT compiler (such as EXLA) for algorithms involving loops to ensure memory efficiency and GPU compatibility. A new AGENTS.md file provides detailed best practices for contributors, enforcing patterns like using \deftransform\/\defnp\, branch-free conditionals, and specific tensor handling to maintain JIT/GPU compatibility. The README also updates the recommended Nx version constraint to \~\> 0.3.0 and includes badges and notebook usage instructions.
(repo-wide) · high confidence
Test coverage
Added KNN benchmark comparing KDTree and brute-force approaches; Added test coverage for Kernel PCA, PCA, and Truncated SVD decomposition modules; Added test coverage for MDS, Trimap, and t-SNE manifold learning algorithms; Added test coverage for Naive Bayes variants; Added test coverage for Scholar metrics modules; Added test coverage for Scholar model selection, preprocessing, and statistics modules; Added test coverage for Scholar.Neighbors algorithms; Added test coverage for SimpleImputer and KNNImputer; Added test coverage for multiple clustering algorithms; Added test coverage for preprocessing utilities; Added test suites for Scholar.Linear regression and classification models; Added test support utilities and datasets; Added tests for BFGS, Nelder-Mead, Brent, and Golden Section optimization algorithms; Added tests for CountVectorizer feature extraction; Added tests for Ledoit Wolf and Shrunk covariance estimators; Added tests for Linear and Quadratic Discriminant Analysis classifiers; Added tests for Partial Least Squares SVD; Added tests for Simpson's rule integration functions; Updated test configuration and added Pima dataset.
Dependencies
Scholar v0.5.0: Major dependency upgrades and documentation overhaul
Scholar has been updated to version 0.5.0, requiring Elixir 1.17 and upgrading the core Nx dependency from 0.1.0 to 1.0. This release also introduces several new dependencies including Polaris for styling, Benchee for benchmarking, and Scidata for data loading, while updating ExDoc to 0.34+ for improved documentation generation. The documentation setup now includes KaTeX for math rendering and Vega-Lite for interactive plots, and the package metadata has been expanded to include maintainers, licenses, and a comprehensive module index for all models and utilities.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 68 → 68 (-0.0)
- Rubric changed (rubric-2026.09.15 → rubric-2026.10.1) — scores are not directly comparable.
Lenses
- Code Health 97 → 99 (+1.7)
- Architecture 100 → 88 (-12.4)
- Maturity 51 → 52 (+1.5)
- Readiness 70 → 73 (+2.5)
- Security 98 → 98 (+0.5)
Resolved (4)
- Off-boarding risk: anonymized user #1
- Outdated: benchee
- Outdated: ex_doc
- Outdated: exla
New (2)
- Off-boarding risk: anonymized user #1
- Projects may be oversized for their cohesion
Changes since last survey
- 2 commits — 2 feature/other, 0 fixes
By area
- (root) — 1 commit
- lib/scholar — 1 commit
Notable commits
- change: Release v0.5.0
- change: chore: Move to Nx 1.0 (#370)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
elixir-nx/scholar was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 3 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 001e47308984dcab6144ba9fd56aa3f4fc686d25 — the exact code this score is about.
- Scored under rubric-2026.10.1 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-8fe32cd45d00.