bheisler/criterion.rs
64.0
Adequate · 30 September 2026
13.6k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is Criterion.rs, a high-performance statistical benchmarking library for Rust that provides comprehensive tools for measuring and analyzing code performance. It enables users to define benchmarks using standard test infrastructure or compatibility shims, execute them with support for async and custom measurements, and generate detailed HTML reports with statistical visualizations. The library includes robust internal modules for univariate and bivariate statistical analysis, including bootstrap resampling, outlier detection, and baseline regression testing to detect performance changes.
How it got here
2014 — Criterion-rs migration and API modernization
6 changes.
The project migrated active development to the Criterion-rs organization and modernized its API by introducing benchmark groups, structured reporting, and a plotters backend. This period also established statistical baseline comparison capabilities using t-tests and updated dependencies to target Rust 2021.
2017–2018 — HTML reports and bencher compatibility
9 changes.
This period focused on enhancing the user experience by introducing HTML benchmark reports for visual and statistical analysis of results. It also added a compatibility layer to allow existing benchmarks written for the bencher crate to run within Criterion.rs, alongside comprehensive integration tests and new benchmark examples.
2019–2022 — Statistical analysis and plotting overhaul
9 changes.
This period focused on building a comprehensive internal statistics library for bootstrap analysis, including univariate, bivariate, and outlier detection modules. It also introduced a new procedural macro for benchmarking and replaced the plotting backend with Gnuplot and Plotters 0.3 to generate detailed SVG visualizations for benchmark reports.
Features
Add Tukey's fences outlier classification
Introduces a new outlier detection module in \src/stats/univariate/outliers\ that implements Tukey's method. This feature classifies data points as normal, mild outliers, or severe outliers based on inner (1.5 \* IQR) and outer (3 \* IQR) fences derived from quartiles, providing a \LabeledSample\ struct to access these classifications and counts.
src/stats/univariate/outliers · high confidence
Add bencher compatibility benchmark example
A new benchmark example file has been added to the bencher compatibility layer, demonstrating how to define and run benchmarks using the \criterion\_bencher\_compat\ crate. This allows users to write benchmarks compatible with the bencher API style, including defining benchmark groups and main entry points.
_bencher\compat/benches · high confidence
Add bencher-compat library for Criterion.rs compatibility
A new \bencher\_compat\ library has been introduced to allow existing benchmarks written for the \bencher\ crate to run using Criterion.rs. This library provides compatibility shims, including a \Bencher\ struct and \benchmark\_group!\/\benchmark\_main!\ macros, which translate \bencher\-style benchmark definitions into Criterion.rs execution, enabling users to migrate their benchmarking infrastructure without rewriting their benchmark logic.
_bencher\compat/src · high confidence
Added CI script to verify nextest compatibility for benchmarks
A new shell script (ci/nextest-compat.sh) has been added to the CI pipeline to validate that the project's benchmarks are compatible with the nextest runner. This script executes 'cargo nextest list --benches' and 'cargo nextest run --benches' to ensure benchmark tests can be discovered and run successfully using nextest.
ci · high confidence
Added bivariate statistical analysis with bootstrap and regression
Users can now perform bivariate statistical analysis, including linear regression through the origin and bootstrap resampling. The new \src/stats/bivariate\ module introduces a \Data\ struct for paired X/Y datasets, a \Slope\ type for fitting lines via ordinary least squares, and a \bootstrap\ method that supports multi-threaded computation when the \rayon\ feature is enabled. This allows users to estimate distributions of statistics like means and slopes from resampled data.
src/stats/bivariate · high confidence
Introduce HTML benchmark reports
Criterion.rs now generates navigable HTML reports for benchmark results. This adds a new \src/html\ module that renders an index page linking to individual benchmark reports, summary pages for groups, and detailed per-benchmark pages displaying statistical tables (mean, median, standard deviation, confidence intervals, throughput) and plots (PDF, regression, iteration times, violin plots, line charts). The reports are generated using TinyTemplate and include links to additional plots where applicable, providing a visual and statistical overview of benchmark performance directly in the browser.
src/html · high confidence
Introduce \`\#\[criterion\]\` procedural macro for benchmarking
The \criterion-macro\ crate now provides a \\#\[criterion\]\ attribute that transforms benchmark functions into standard Rust test cases. When applied, the macro wraps the user's function, instantiates a \Criterion\ instance (using default settings or custom arguments provided in the attribute), and configures it from command-line arguments. This allows users to write benchmarks using the familiar \\#\[test\]\ infrastructure while retaining Criterion's configuration capabilities.
macro · high confidence
Introduce benchmark groups and structured reporting
Users can now organize benchmarks into named groups using the new \BenchmarkGroup\ API, allowing shared configuration (such as measurement time and sample size) across multiple benchmarks and generating consolidated reports. This change also introduces a new communication protocol between the runner and benchmarks using CBOR serialization via the \ciborium\ crate, and adds support for exporting raw benchmark data to CSV files.
src · high confidence
Introduce plotters backend alongside gnuplot for plot generation
The plotting module now supports two rendering backends: the existing gnuplot backend and a new plotters backend, which is enabled via the \plotters\ feature flag. This change introduces a \Plotter\ trait that abstracts the rendering logic, allowing the library to generate plots (such as PDFs, regressions, and violin plots) using either gnuplot or plotters, providing users with an alternative to gnuplot for generating SVG output.
src/plot · high confidence
Introduce statistical baseline comparison with t-test and change estimates
The analysis module now supports comparing current benchmark results against a saved baseline. When a baseline directory exists, the system loads previous sample and estimate data, performs a two-sample t-test to calculate a t-statistic and distribution, and bootstraps relative change estimates for mean and median execution times. These change estimates and distributions are saved to a 'change' subdirectory, enabling users to detect performance regressions or improvements relative to a previous run.
src/analysis · high confidence
Introduction of the criterion-plot library for generating Gnuplot scripts
The \plot/src\ directory now contains the \criterion-plot\ library, which provides a Rust API for generating Gnuplot scripts to visualize benchmark data. This new component introduces support for various plot types including curves (lines, points, steps, impulses), error bars, candlesticks, and filled curves. It also includes configuration capabilities for axes (ranges, scales, labels, tick labels), gridlines, and the legend (key), allowing users to customize the appearance and layout of their benchmark visualizations.
plot/src · high confidence
New Gnuplot backend for benchmark reports
The benchmark report generation now uses a new Gnuplot backend (src/plot/gnuplot\_backend) to produce SVG visualizations. This backend implements the Plotter trait to generate specific charts including probability density functions (PDFs) with outlier detection, linear regression plots with confidence intervals, iteration time scatter plots, and Welch's t-test distributions. It also supports line comparisons and violin plots for multi-benchmark analysis, replacing the previous plotting implementation.
_src/plot/gnuplot\backend · high confidence
New benchmark examples for async, custom measurements, and external processes
The benchmark suite now includes new example files demonstrating async measurement overhead (using FuturesExecutor), custom measurement implementations (such as a HalfSeconds timer), and external process benchmarking (via a Python script). It also adds examples for comparing functions with inputs, using different sampling modes, handling special characters in benchmark names, and measuring large setup/drop overheads.
benches/benchmarks · high confidence
New internal statistics library for bootstrap analysis and tuple handling
This change introduces a new internal statistics module (\src/stats\) that provides core data structures for benchmark analysis. It adds a \Distribution\ type to represent bootstrap distributions, enabling users to compute confidence intervals and p-values for benchmark results. The module also includes utilities for handling tuples of distributions (supporting up to four elements) and implements a lightweight random number generator using \oorandom\ to replace the heavier \rand\ crate dependency for internal statistical sampling.
src/stats · high confidence
New univariate statistical analysis module with bootstrap support
Introduces a new \src/stats/univariate\ module providing core statistical operations on data samples, including mean, median, standard deviation, variance, and percentiles. The module adds bootstrap functionality for both single-sample and two-sample scenarios, supporting parallel execution via the \rayon\ feature. It also includes utilities for generating random resamples and calculating metrics like the median absolute deviation and interquartile range.
src/stats/univariate · high confidence
Removals
Removal of legacy Criterion benchmark examples
The legacy benchmark examples (alloc.rs, fib.rs, math.rs) that relied on the old criterion::Bencher API have been removed. These files previously demonstrated basic benchmarking patterns using the deprecated interface, and their deletion indicates a shift away from that specific API usage in the provided examples.
examples · high confidence
Behavioural changes
Consolidated benchmark entry point
The benchmark suite now uses a single entry point (benches/bench\_main.rs) that explicitly registers all benchmark modules, including new additions for async measurement, sampling mode, and custom measurement, ensuring a unified execution path for all defined benchmarks.
benches · high confidence
Migrate plotting backend to Plotters 0.3
The benchmark report generation has switched from the previous plotting library to Plotters 0.3. This change updates the rendering engine for all benchmark visualizations, including probability density functions, regression lines, iteration time charts, and distribution plots, ensuring compatibility with the newer library API and providing updated visual styles for the generated SVG reports.
_src/plot/plotters\backend · high confidence
Repository migration to the Criterion-rs organization
The project has moved active development to the new Criterion-rs GitHub organization. The README now directs users to submit new issues and pull requests to the new repository, and the contributing guidelines have been updated to reflect this change. Additionally, the repository has been cleaned up by removing legacy build scripts (Makefile, check-line-length.sh) and adding configuration files for editor consistency (.editorconfig) and typo checking (.typos.toml).
(repo-wide) · high confidence
Updated benchmark data for the Fibonacci/Iterative example in the user guide
The HTML report for the Fibonacci/Iterative benchmark in the user guide has been regenerated with fresh statistical data. This update includes new raw CSV samples, JSON estimates (mean, median, standard deviation), and Gnuplot-generated SVG visualizations (MAD, SD, and combined PDF plots) reflecting the current benchmark results.
_book/src/user\_guide/html\report · high confidence
Test coverage
Added integration tests for Criterion.rs benchmarking functionality
Added a new comprehensive test suite in \tests/criterion\_tests.rs\ to verify the core behavior of the Criterion benchmarking library. These tests validate that benchmarks correctly create output directories, generate expected statistical files (JSON estimates, samples, Tukey stats, and raw CSV when enabled), and handle baseline comparisons (saving, retaining, and strict/lenient comparison modes). The tests also verify configuration options such as disabling plots and ensure that output files are written to the specified temporary directories.
tests · high confidence
Dependencies
Criterion.rs 0.7.0 dependency and workspace update
This release updates the Criterion.rs benchmarking library to version 0.7.0, targeting Rust 2021 edition and a minimum supported Rust version (MSRV) of 1.80. The dependency tree has been refreshed with significant version bumps, including itertools to 0.13, clap to 4.5, and async-std to 1.13. Optional async runtime support now includes smol 2.0 and tokio 1.0, while the plotting backend relies on plotters 0.3. The project structure is organized as a Cargo workspace, separating the core library, the criterion-plot crate, and compatibility/macro sub-crates.
(dependencies) · high confidence
Housekeeping
Added documentation and license files for bencher\_compat; Establishes plot subdirectory with documentation and build artifacts.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 63 → 64 (+1.4)
- Rubric changed (rubric-2026.09.10 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 88 → 88 (-0.0)
- Architecture 100 → 98 (-2.5)
- Maturity 46 → 51 (+4.9)
- Readiness 82 → 76 (-5.8)
- Security 83 → 90 (+6.8)
- Accessibility 66 → 66 (+0.0)
- Performance 85 (new)
Resolved (1)
- Documentation: no project overview (README.md)
New (17)
- Asymmetric API between sync and async benchers. AsyncBencher has iter_with_large_setup but Bencher does not have a corresponding iter_with_large_setup (it only has iter_with_setup and iter_with_large_drop). This forces users to handle setup/drop logic differently depending on the executor type.
- Duplicated block (6 lines × 2) (src/plot/gnuplot_backend/regression.rs)
- Inconsistent ID type for benchmark identification. The top-level Criterion API expects a string for the benchmark ID, while the BenchmarkGroup API expects a specific BenchmarkId type.
- Inconsistent ID type for parameterized benchmarks. Criterion expects BenchmarkId, while BenchmarkGroup uses a generic ID type (likely String or BenchmarkId depending on context, but the signature differs).
- Low cohesion: Sample (LCOM4 5) (src/stats/univariate/sample.rs)
- Off the main sequence: criterion-plot
- Outdated: async-std
- Outdated: clap
- Outdated: csv
- Outdated: futures
- Outdated: proc-macro2
- Outdated: quote
- Outdated: rayon
- Outdated: regex
- Outdated: serde
- Outdated: serde_json
- Outdated: tokio
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
bheisler/criterion.rs was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3dbc6c618acb48885066422d81d50729aa17b2b7 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-505904ce13c1.