Skip to content
CAI
Software that uses CAICheck a score

rust-lang/regex

73.5

Strong · 30 September 2026

113.5k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a comprehensive Rust-based regular expression engine ecosystem that provides high-performance pattern matching through a modular architecture. It includes a core library for standard regex operations, a dedicated crate for lightweight matching, and a low-level automata engine that powers the search capabilities. The project also offers a C API for external integration, a command-line tool for debugging and benchmarking, and extensive fuzzing infrastructure to ensure stability and correctness.

How it got here

2014–2016 — Automata rewrite and C API expansion

11 changes.

The project underwent a significant architectural shift by rewriting the core regex engine to leverage the regex-automata crate, replacing the previous hand-rolled implementation with optimized, automata-based matching. This period also saw the modularization of the codebase into distinct crates and the introduction of rure, a new C API binding that exposed the regex engine's capabilities to external consumers.

2018–2023 — regex-syntax and regex-lite development

24 changes.

This period focused on establishing the foundational regex-syntax crate with AST and HIR modules, while introducing the lightweight regex-lite engine to reduce binary size. The work also included building comprehensive testing infrastructure, including fuzzing, benchmarking, and a unified test suite, alongside a CLI tool for debugging and engine comparison.

Features

Add C API example for iterating regex matches

The regex-capi/examples directory now includes a C program (iter.c) and build script (compile) demonstrating how to use the rure C API to iterate over regex matches in a file, extract capturing groups, and print match locations. The example uses the Sherlock Holmes text as sample data.

regex-capi/examples · high confidence

Add RegexSet for matching multiple patterns in a single pass

Users can now match multiple, possibly overlapping, regular expressions against a string or byte slice in a single pass using the new \RegexSet\ API. This feature allows applications, such as URL routers or user-agent matchers, to efficiently identify which specific patterns from a large set matched the input without performing separate searches for each pattern. The implementation provides separate modules for string (\&str\) and bytes (\&\[u8\]\) inputs, with the bytes variant supporting invalid UTF-8 data.

src/regexset · high confidence

Add directory for recording and comparing benchmark results

A new \record\ directory has been added to the repository to store committed benchmark and compilation test results. This includes CSV files tracking compilation time and binary size for the \regex\ and \regex-automata\ crates across different configurations, as well as raw benchmark logs for various regex implementations (such as dynamic, native, NFA, and PCRE). These files allow developers to manually compare performance and size metrics over time.

record · high confidence

Add fuzzing infrastructure and build scripts for OSS-Fuzz

The project now includes dedicated fuzzing infrastructure to support integration with OSS-Fuzz. This change adds build scripts (oss-fuzz-build.sh) that compile and package specific fuzz targets, including regex matchers and AST-based fuzzers. It also introduces configuration files, such as ast-fuzzers.options, to enforce sane limits (e.g., max\_len = 65536) for arbitrary-based fuzzers, ensuring stable and controlled fuzzing execution.

fuzz · high confidence

Add regex! macro and TryFrom implementations for Regex

The \regex!\ macro is now available for lazy compilation of regular expressions from string literals, storing the compiled pattern in a static to avoid recompilation overhead. Additionally, \TryFrom\<&str\>\ and \TryFrom\<String\>\ implementations have been added to the \Regex\ type, allowing for more idiomatic conversion from string types into compiled regex patterns.

src/regex · high confidence

Add regex-cli commands for compile testing and debugging

The regex-cli tool now includes new subcommands to help users evaluate and debug regex performance. The \compile-test\ command measures compilation time and binary size overhead for various regex configurations (including regex-automata and regex-lite feature combinations), outputting results as CSV. The \debug\ command provides a way to inspect the internal debug representations of regex-automata structures. These additions are part of the initial import of the regex-automata integration into the CLI.

regex-cli/cmd · high confidence

Initial import of regex-test crate for TOML-based regex engine testing

The \regex-test\ crate has been added to provide a standardized framework for testing regex engine implementations. It allows developers to define test cases using a TOML format, specifying patterns, input text, expected matches, and various engine options (such as case-insensitivity, Unicode support, and match kinds). The crate includes utilities to load these TOML test files and execute them against different regex backends, facilitating consistent validation of regex behavior across implementations.

regex-test · high confidence

Initial release of regex-cli for debugging and testing regex engines

This change introduces the \regex-cli\ tool, a command-line interface for interacting with the \regex\, \regex-automata\, and \regex-syntax\ crates. Users can now install the tool via \cargo install regex-cli\ to perform development tasks such as printing debug representations of regex internal structures (e.g., NFAs), executing searches with detailed timing metrics, and generating serialized DFA code. The CLI supports multiple regex engines (including \meta\, \lite\, and \hybrid\) and provides flags for configuring search behavior, such as handling invalid UTF-8 or minimizing DFAs.

regex-cli · high confidence

Initial release of regex-syntax crate

The regex-syntax crate is introduced as a standalone library providing a robust regular expression parser. It exports \Ast\ and \Hir\ types to represent the abstract syntax and high-level intermediate representation of regex patterns, respectively. The crate includes comprehensive documentation, Apache 2.0 and MIT license files, and a test script that validates various feature combinations (such as \std\, \unicode\, and specific Unicode categories) to ensure correctness across different configuration options.

regex-automata, regex-syntax · high confidence

Initial release of rure, a C API for Rust's regex engine

This change introduces rure, a new C API binding for Rust's regex library, located in the regex-capi directory. The library provides a C interface to Rust's regex engine, which uses finite automata to guarantee linear time searching. It supports capturing groups, lazy matching, Unicode support, and word boundary assertions, while excluding features like backreferences. The release includes the C header file (includes/rure.h), license files (Apache-2.0 and MIT), a README with usage examples and documentation, and a test script to verify the build and examples.

regex-capi · high confidence

Introduce C API for regex matching and capture group iteration

The \regex-capi/include/rure.h\ header introduces a new C API for the \rure\ regex engine, enabling C applications to compile regular expressions, perform matches, and access capture groups. The API provides functions to compile patterns with configurable flags (such as case-insensitivity and Unicode support) and options, check for matches, and find match offsets. It also exposes an iterator (\rure\_iter\_capture\_names\) to retrieve the names of capture groups defined in a compiled regex, along with memory management functions to safely free compiled expressions, captures, and iterators.

regex-capi/include · high confidence

Introduce Rure, a new C API for regex operations

Adds a new C-compatible library (Rure) that exposes regex compilation, matching, and capture functionality to C and other FFI consumers. The API includes functions to compile regex patterns with configurable flags (case-insensitive, multi-line, etc.), perform matches on byte slices, retrieve match positions, and access capture group names via an iterator. It also provides error handling structures and memory management functions for the opaque regex objects.

regex-capi/src · high confidence

Introduce regex-lite crate as a lightweight regex alternative

A new \regex-lite\ crate is introduced to provide a lightweight regex engine for searching strings, serving as a drop-in replacement for the primary \regex\ crate's \Regex\ type. This crate is designed to significantly reduce binary size and compilation times by avoiding complex optimizations and robust Unicode support, making it suitable for users prioritizing performance and size over full feature parity. The crate supports a regex syntax nearly identical to the \regex\ crate, includes a minimum supported Rust version (MSRV) of 1.65.0, and is licensed under either Apache 2.0 or MIT.

regex-lite · high confidence

Introduce regex-lite crate for lightweight regex support

A new \regex-lite\ crate has been added to provide a lightweight regular expression engine optimized for smaller binary sizes and faster compilation times, at the cost of performance and advanced Unicode support compared to the standard \regex\ crate. This location introduces the core implementation files, including the public \Regex\ API with \TryFrom\ support, an NFA-based compiler, a PikeVM search engine, and utilities for UTF-8 decoding and capture group interpolation.

regex-lite/src · high confidence

Introduce regex-lite crate with High-Level Intermediate Representation (HIR)

This change adds the \regex-lite\ crate, providing a lightweight regular expression engine. The diff introduces the High-Level Intermediate Representation (HIR) module, which includes a parser (\parse.rs\) and core data structures (\mod.rs\) for representing regex patterns. Key features include support for standard regex syntax (character classes, repetitions, captures), configuration flags (case insensitivity, multi-line, dot-matches-newline), and specific handling for word boundaries and escape sequences. The implementation includes a nesting limit to prevent stack overflows during parsing and defines error handling for unsupported features like look-arounds and backreferences.

regex-lite/src/hir · high confidence

Introduce regex-syntax crate for regular expression parsing

This change introduces the \regex-syntax\ crate, a standalone library for parsing regular expressions into Abstract Syntax Trees (AST) and High-level Intermediate Representations (HIR). It provides a robust, safe parser with configurable options (such as nesting limits and UTF-8 validation) and detailed error reporting, serving as the foundational syntax analysis component for the broader regex ecosystem.

regex-syntax/src · high confidence

Introduction of the High-Level Intermediate Representation (HIR) module

The \regex-syntax\ crate now exposes a new \hir\ module that provides a high-level intermediate representation for regular expressions. This module includes an AST-to-HIR translator, a visitor pattern for stack-safe traversal, a printer for debugging or round-tripping patterns, and utilities for literal extraction to optimize search performance. It also introduces an internal \IntervalSet\ type to manage character ranges and case folding efficiently.

regex-syntax/src/hir · high confidence

New debug subcommands for DFA, sparse DFA, and literal extraction

The \regex-cli debug\ command now includes new subcommands to inspect the internal state of the regex engine. Users can run \regex-cli debug dense\ and \regex-cli debug sparse\ to print debug representations and performance metrics (parse, compile, memory) for dense and sparse DFAs, including support for full DFA regexes that display both forward and reverse DFAs. Additionally, a new \regex-cli debug literal\ command allows users to extract and inspect literal prefixes or suffixes from a regex pattern, with options to control optimization and extraction limits.

regex-cli/cmd/debug · high confidence

New regex-cli argument parsing module for regex-automata

The \regex-cli/args\ module has been introduced to provide structured command-line argument parsing for the new \regex-automata\ crate. This change adds configuration support for various regex engine components, including the Thompson NFA, dense and sparse DFAs, hybrid lazy DFAs, and the meta-regex engine. Users can now configure specific engine behaviors via CLI flags, such as setting size limits for the DFA or NFA, controlling match kinds (e.g., leftmost-first vs. all), managing capture states, and adjusting search bounds. The module also handles common flags like quiet/verbose modes and input specification (inline or file path).

regex-cli/args · high confidence

New regex-cli find subcommands for match, capture, half, and which searches

The regex-cli tool now includes a comprehensive 'find' command suite that allows users to execute regex searches using multiple underlying engines. Users can search for full matches, capture groups, half matches (end offsets only), and pattern-set matches ('which') using engines such as the dense and hybrid DFAs, the one-pass DFA, the bounded backtracker, PikeVM, the meta engine, regex-lite, and the top-level API. Each subcommand provides detailed timing tables for parsing, translation, compilation, and search phases, enabling users to benchmark and compare engine performance for specific search types.

regex-cli/cmd/find · high confidence

New regex-cli generate subcommands for test conversion, DFA serialization, and Unicode table generation

The \regex-cli generate\ command now supports three new subcommands. The \fowler\ subcommand converts Glenn Fowler's legacy regex test suite (\.dat\ files) into TOML format for use with the \regex-test\ crate. The \serialize\ subcommand allows users to serialize fully compiled dense or sparse DFAs (and regex DFAs) to disk, optionally generating Rust source code for embedding these compiled automata into other programs. The \unicode\ subcommand generates all required Unicode data tables for both \regex-syntax\ and \regex-automata\ by invoking the external \ucd-generate\ tool, including specific tables for Perl-compatible character classes like \\\w\, \\\d\, and \\\s\.

regex-cli/cmd/generate · high confidence

New regex-syntax AST module with visitor-based traversal

The \regex-syntax/src/ast\ module has been introduced, defining an abstract syntax tree (AST) for regular expressions along with a parser, printer, and a visitor trait. This visitor pattern allows consumers to traverse the AST in constant stack space, preventing stack overflows on deeply nested patterns. The module supports advanced regex features including named capture groups (both \(?P\<name\>\ and \(?\<name\>\ syntax), Unicode character classes, and word boundary assertions. Error handling is robust, with specific error kinds for issues like invalid ranges, unclosed groups, and nesting limit exceedances.

regex-syntax/src/ast · high confidence

Removals

Removal of the \`regex\_macros\` crate

The \regex\_macros\ crate, which previously provided the \regex!\ compile-time macro for generating specialized regex code, has been deleted. Users relying on this macro will need to migrate to the standard runtime \Regex::new\ API provided by the main \regex\ crate, as the compile-time expansion capability is no longer available.

_regex\macros · high confidence

Behavioural changes

Repository infrastructure and contributor guidelines update

This change introduces several repository-level updates: a new AI Policy document clarifies that while using AI tools for coding is welcome, contributors must write their own comments and explanations, and autonomous agents are prohibited; a new \.ignore\ file whitelists the \.github\ directory; the \.gitignore\ file is updated to exclude additional artifacts like \bench-log\, \wiki\, \tags\, and \tmp/\; and the \test\ script is updated to run integration tests across a broader set of feature combinations, including \perf-dfa-full\ and \perf-onepass\. Additionally, \Cross.toml\ is added to configure environment passthrough for cross-compilation testing.

(repo-wide) · high confidence

Rewrite regex engine to use regex-automata

The internal regex engine has been completely rewritten to use the \regex-automata\ crate, replacing the previous hand-rolled NFA/VM implementation. This change introduces a new internal builder system (\src/builders.rs\) that configures the \meta\ engine, and adds support for matching on raw bytes (\src/bytes.rs\) alongside standard UTF-8 strings. The public API remains largely consistent, but the underlying search mechanism now leverages optimized automata-based matching, including support for literal optimizations and SIMD acceleration where available.

src · high confidence

Unicode tables updated to version 16.0.0

The Unicode data tables used by the regex syntax engine have been regenerated from the Unicode 16.0.0 standard. This update refreshes character properties including age, case folding, general categories, grapheme cluster breaks, script assignments, and Perl-compatible word/space/decimal definitions to reflect the latest Unicode specifications. A Unicode license file has also been added to the directory to ensure proper attribution for the data files.

_regex-syntax/src/unicode\tables · high confidence

Fixes

11 commits (6 fixes) fixing fuzz/regressions

A fix in fuzz/regressions — 11 commits (6 fixs), 17 files.

fuzz/regressions · medium confidence · unverified

Test coverage

Add structured fuzzing targets for regex and AST validation; Added fuzz regression tests for regex engine edge cases; Added fuzzing regression tests and comprehensive test suite integration for regex-lite; Added regex-syntax parsing benchmarks; Initial C API test suite for rure; Migrate test suite to \regex\_test\ harness and modernize test structure; Standardized regex test suite with engine-independent TOML format.

Dependencies

Regex library major version bump to 1.13.1 with workspace restructuring

The core regex crate has been updated to version 1.13.1, marking a significant release that includes a migration to the Rust 2021 edition and a minimum supported Rust version (MSRV) of 1.65. This change introduces a comprehensive workspace structure that organizes the library into distinct crates: \regex-automata\ (version 0.4.18) for low-level matching, \regex-syntax\ (version 0.8.11) for parsing, \regex-lite\ (version 0.1.9) for lightweight matching, \regex-capi\ (version 0.2.5) for C bindings, and \regex-cli\ (version 0.2.3) for command-line tools. The update also removes the legacy \regex\_macros\ crate and updates key dependencies such as \aho-corasick\ to 1.0.0 and \memchr\ to 2.6.0, while adding fuzzing infrastructure in the \fuzz\ directory.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 75 → 73 (-1.2)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 87 → 87 (+0.0)
  • Architecture 100 → 94 (-5.6)
  • Maturity 69 → 69 (+0.0)
  • Readiness 71 → 68 (-2.9)
  • Security 82 → 88 (+6.0)
  • Performance 85 (new)

Resolved (2)

  • Edited copy of a member (16 corresponding lines) (regex-automata/src/meta/limited.rs)
  • Off-boarding risk: anonymized user #1

New (14)

  • Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
  • Edited copy of a member (16 corresponding lines) (regex-automata/src/dfa/search.rs)
  • Edited copy of a member (20 corresponding lines) (regex-cli/cmd/find/which/mod.rs)
  • Edited copy of a member (20 corresponding lines) (regex-cli/cmd/find/which/mod.rs)
  • Inconsistent cache requirement. Automaton (low-level) methods try_search_fwd/rev do not take a cache. hybrid.dfa.DFA methods try_search_fwd/rev require a Cache. While this reflects the internal architecture (hybrid uses caching, raw automaton might not or manages it differently), the naming try_search is identical but the signature differs significantly in required arguments, confusing users about whether they need to manage a cache.
  • Inconsistent naming for search/match operations across different DFA types. onepass.DFA uses is_match and find. hybrid.dfa.DFA uses try_search_fwd and try_search_rev (and lacks a simple find or is_match on the DFA itself, requiring a cache). meta.regex uses is_match, find, AND search (which returns a Match). search and find appear to do the same thing in meta.regex.
  • Inconsistent naming for sparse construction. The dense variants use build and build_many on the Builder type, but the Regex type itself exposes new and new_many. However, the sparse variants are named new_sparse and new_many_sparse on the Regex type, whereas the dense Builder uses build_sparse and build_many_sparse. More critically, the meta regex uses new/new_many for standard and has no explicit sparse constructor on the Regex type itself (sparse is a builder option). This creates a fragmented API surface where 'sparse' is a suffix on some constructors but not others, and 'new' vs 'build' is inconsistent across types.
  • Inconsistent return types for serialization methods. dense.DFA returns (Vec<u8>, usize) (bytes and size), while sparse.DFA returns u8 (which seems incorrect or truncated, likely a typo in the API definition provided or a very strange API). Assuming the u8 is a placeholder or error in the prompt, if it returns a single byte, it's a massive inconsistency. If it returns Vec<u8>, it's still inconsistent with the tuple return of dense. dense also has write_to_* methods, while sparse has both to_bytes_* and write_to_*.
  • Low cohesion: Config (LCOM4 5) (regex-automata/src/dfa/determinize.rs)
  • Low cohesion: DFA (LCOM4 4) (regex-automata/src/hybrid/dfa.rs)
  • Off the main sequence: regex-lite
  • Off the main sequence: regex-syntax
  • Off-boarding risk: anonymized user #1
  • Projects may be oversized for their cohesion

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

rust-lang/regex was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 72d650cb0a880a01ab6dc2137c0888e8f89740f7 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.