Skip to content
CAI
Software that uses CAICheck a score

pest-parser/pest

53.0

Adequate · 30 September 2026

18k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a Rust-based PEG (Parsing Expression Grammar) parser library that enables declarative definition of parsers via \.pest\ grammar files. It provides a comprehensive toolchain including a derive macro for code generation, a runtime virtual machine for dynamic execution, and a CLI debugger for inspecting parse states. The library includes built-in grammars for common data formats like JSON, SQL, HTTP, and TOML, along with robust infrastructure for fuzzing, benchmarking, and error reporting.

How it got here

2016–2017 — Pest 2.9.2 refactoring and legacy removal

10 changes.

This period focused on upgrading the Pest parser library to version 2.9.2, which involved removing the legacy RDP-based parser implementation and its associated traits. The core engine underwent significant internal refactoring to improve error handling, stack management, and iterator performance, while also introducing compile-time operator precedence support. Repository governance and developer tooling were simultaneously updated to streamline maintenance and community contribution.

2018–2022 — Grammar expansion and tooling infrastructure

20 changes.

This period focused on expanding the library's capabilities by adding built-in grammars for HTTP, JSON, SQL, and TOML, alongside a new meta-grammar AST and optimizer. It also established core tooling infrastructure, including the pest\_derive macro, a runtime virtual machine for dynamic parsing, and a CLI debugger to support development and debugging workflows.

Features

Added HTTP, JSON, SQL, and TOML grammars

New Pest grammar definitions have been added for HTTP, JSON, SQL, and TOML, enabling the parser to recognize and structure these data formats. The HTTP grammar handles request lines, headers, and delimiters; the JSON grammar supports objects, arrays, strings, numbers, booleans, and null values while rejecting unescaped control characters; the SQL grammar covers queries, DDL, ACL commands, and expressions; and the TOML grammar supports tables, arrays, strings, literals, and date-time values.

grammars/src/grammars · high confidence

Added fuzzing dictionaries for HTTP, JSON, SQL, and TOML grammars

New fuzzing dictionaries have been added to the \grammars/fuzz/dict\ directory to improve coverage for specific protocol and data formats. The \http.dict\ file provides common HTTP methods, headers, and status codes; \json.dict\ includes valid and edge-case JSON structures, escape sequences, and numeric literals; \sql.dict\ covers standard SQL keywords, operators, and data types; and \toml.dict\ contains TOML-specific syntax elements like keys, values, and date/time formats. These files are intended to guide fuzzing tools in generating more effective test inputs for parsers handling these formats.

_grammars/fuzz/dict, grammars/fuzz/fuzz\targets · high confidence

Added fuzzing infrastructure for pest\_meta parser

Added a new fuzzing target for the pest\_meta parser that includes a call limit of 5,000 to prevent crashes or timeouts during execution. The change introduces the necessary fuzzing setup files, including a .gitignore for build artifacts and a README linking to the main fuzzing documentation.

meta/fuzz · high confidence

Added parentheses parsing example

A new example demonstrating how to parse nested parentheses structures using the low-level parser state API. The code defines a \ParenParser\ that recursively handles \(\ and \)\ delimiters, and includes conditional support for \miette\ error reporting via the \miette-error\ feature flag.

pest/examples · high confidence

Generator crate restructured with doc comment support and no\_std capability

The generator module has been reorganized into a new crate structure, introducing dedicated source files for documentation parsing (\docs.rs\), code generation (\generator.rs\), and derive input parsing (\parse\_derive.rs\). This change adds support for extracting and propagating \///\ and \//!\ doc comments from grammar files into the generated Rust code, allowing users to see documentation for rules in their IDE. It also enables \no\_std\ support for the generator via feature flags and introduces the \\#\[grammar\_inline\]\ attribute to allow defining grammars directly in source code without external files. Additionally, the generator now correctly handles the \\#\[non\_exhaustive\]\ attribute on the derived struct to apply it to the generated \Rule\ enum.

generator · high confidence

Initial release of the pest\_derive crate

The \derive/src\ location now contains the \pest\_derive\ crate, providing the \\#\[derive(Parser)\]\ macro that allows users to generate parsers from \.pest\ grammar files or inline grammar definitions. This entry establishes the foundational API for declarative parser generation in the project.

derive/src · high confidence

Introduce on-the-fly AST execution via the VM crate

The new \vm\ crate provides a virtual machine that executes optimized grammar ASTs at runtime rather than relying on generated code. This enables dynamic parsing capabilities, powering tools like the fiddle and debugger by allowing rules to be run on-the-fly. The implementation includes core parsing logic, built-in atomic rules (such as \ANY\, \EOI\, and character classes), and helper macros (\parses\_to\, \fails\_with\) for testing the VM's behavior.

vm/src · high confidence

Introduce pest\_debugger crate and CLI debugger

Adds a new \pest\_debugger\ library crate and a CLI-based debugger (\main.rs\) for stepping through pest grammar parsing. The debugger allows users to load grammars and inputs, set breakpoints on specific rules, and step through execution via commands like \run\, \continue\, and \breakpoint\. It exposes a \DebuggerContext\ API for programmatic integration and supports long command names and prefixes for easier CLI usage.

debugger · high confidence

Meta-grammar AST and parser implementation

The \meta/src\ module now defines the Abstract Syntax Tree (AST) for the pest meta-grammar, introducing \Rule\ and \Expr\ types to represent grammar structures and expressions. This includes support for rule modifiers (silent, atomic, compound atomic, non-atomic), expression operators (sequence, choice, predicates, repetitions), and new features like \PUSH\_LITERAL\, \PEEK\ slices, and node tagging. The module also contains the generated parser (\grammar.rs\) and the source grammar definition (\grammar.pest\) that parses these meta-grammar files, enabling the tool to interpret and validate pest grammar definitions.

meta/src · high confidence

New SQL grammar with Pratt parser and HTTP/TOML/JSON parsers added

The library now exposes four distinct grammars: a new SQL parser that uses a Pratt parser for correct operator precedence, an HTTP request parser, and sample parsers for JSON and TOML. The SQL grammar is adapted from the Picodata sbroad project to simulate SQLite-style queries. Additionally, call limits are enforced on the JSON and TOML fuzzing tests to prevent hangs on deeply nested inputs, and the JSON grammar rejects unescaped control characters.

grammars/src · high confidence

New derive examples for calculator and help-menu grammars

Added new example files in the derive/examples directory demonstrating how to use pest\derive with multiple grammar files and Pratt parsers. The calc example shows a calculator grammar (base.pest and calc.pest) supporting infix operators (+, -, \, /, ^), prefix negation, and postfix factorial, parsed using a Pratt parser to handle precedence. The help-menu example (help-menu.pest and help-menu.rs) demonstrates parsing CLI help text structures with optional, required, and choice arguments.

derive/examples · high confidence

Removals

Removal of StringInput implementation

The \StringInput\ struct and its module have been removed from the \src/inputs\ directory. This eliminates the previous in-memory string matching capability that allowed users to create inputs from \&str\ and perform case-sensitive character comparisons via the \matches\ method.

src/inputs · high confidence

Removal of the legacy RDP parser implementation

The \src/parsers\ module, which contained the \Rdp\ trait and the \impl\_rdp!\ macro for recursive-descent parsing, has been removed. This change eliminates the legacy parser interface, meaning users can no longer implement or use the \Rdp\ trait for parsing logic in this location.

src/parsers · high confidence

Removed legacy RDP-based parser implementation

The legacy recursive-descent parser (RDP) implementation has been removed from the library. This change deletes the \grammar!\ macro, the \Input\ and \Parser\ traits, and the \Rdp\ trait, along with their associated modules (\src/grammar.rs\, \src/input.rs\, \src/parser.rs\, and related exports in \src/lib.rs\). Users relying on the \impl\_rdp!\ macro and the manual rule-definition style provided by this legacy system will no longer have access to these APIs.

src · high confidence

Behavioural changes

Bootstrap process now generates grammar.rs from pest source

The bootstrap tool has been updated to explicitly generate the \grammar.rs\ file from the \grammar.pest\ source during the build process. This change ensures the generated parser code is created locally, addressing previous issues with \pest\_derive\ under workspaces on Windows and keeping the generated source file in the repository tree rather than relying solely on compile-time derivation.

bootstrap · high confidence

Include license files in the package

The package now includes Apache and MIT license files via symlinks, ensuring that licensing information is distributed with the package contents.

meta · medium confidence

Major internal refactoring of parser state, error handling, and stack management

The core parsing engine has been significantly restructured. The \ParserState\ now manages parsing logic directly, introducing a new \Stack\ type with snapshot/restore capabilities for backtracking and a \set\_call\_limit\ function to prevent stack overflows. Error reporting has been enhanced with a new \Error\ structure that includes detailed call stacks, parse attempts, and line/column location information. Additionally, a new \ConstPrattParser\ has been added alongside the existing runtime version to allow for compile-time operator precedence configuration, and benchmarks have been added to measure the performance of these new features.

pest/src · high confidence

New AST optimizer pipeline with dedicated optimization passes

The meta/optimizer module now implements a structured optimization pipeline for pest's ASTs, introducing a new OptimizedExpr type and a sequence of dedicated passes: concatenation of adjacent string literals, factorization of common prefixes in choices, list pattern optimization to reduce redundant matches, rotation of nested sequences/choices, inlining of skip patterns, unrolling of repetition quantifiers, and error-state restoration for stack-modifying expressions. This refactors the previous monolithic optimizer into modular components, improving maintainability and enabling targeted performance improvements for grammar parsing.

meta/src/optimizer · high confidence

Refactored iterator internals and added PairsBuilder for testing

The iterator subsystem in \pest/src/iterators\ has been restructured to improve performance and testability. A new \LineIndex\ component replaces the previous cursor-based approach for line/column calculations, using a binary-searchable offset table to speed up \Pair::line\_col\ and \Pairs::next\. The internal token representation has been replaced by a smaller, more efficient \QueueableToken\ enum that stores direct indices to paired tokens, reducing memory usage and lookup time. Additionally, a new \PairsBuilder\ utility has been added, allowing users to manually construct \Pairs\ from rules and byte offsets without running a parser, which simplifies unit testing for code that consumes parser output.

pest/src/iterators · high confidence

Repository restructured with new configuration, documentation, and licensing files

The repository has been reorganized to improve developer experience and project governance. A new .gitpod.yml file enables ready-to-go cloud development workspaces, while CONTRIBUTING.md and SECURITY.md provide clear guidelines for community involvement and vulnerability reporting. The project license has been updated from MPL-2.0 to the dual MIT/Apache-2.0 model, reflected in the new LICENSE-MIT and LICENSE-APACHE files. Additionally, standard tooling configurations (rustfmt.toml, typos.toml, codecov.yml) and automation scripts (release.sh, semvercheck.sh, update\_unicode.sh) have been added to streamline maintenance, testing, and Unicode data updates.

(repo-wide) · high confidence

Symlinked license and readme files in derive and grammars directories

The \derive\ and \grammars\ directories now use symbolic links for their license files (\LICENSE-APACHE\, \LICENSE-MIT\) and the readme file (\\_README.md\). These links point to the corresponding files in the parent directory, ensuring that license and documentation information is consistent across these subprojects without duplicating content.

derive, grammars · high confidence

Unicode data updated to version 18.0.0

The built-in Unicode property rules in pest have been regenerated to align with Unicode Standard version 18.0.0. This update refreshes the underlying trie data for binary properties (such as Emoji and Alphabetic), general categories, and scripts, ensuring that pattern matching against character classes reflects the latest official Unicode definitions.

pest/src/unicode · high confidence

Fixes

The pest crate now includes symlinks to the project's LICENSE-APACHE, LICENSE-MIT, and README files. This ensures that the license information and documentation are accessible within the crate directory, resolving issues with broken or missing references in sub-crates.

pest · high confidence

Test coverage

Added HTTP grammar benchmark and test data fixtures; Added VM test coverage for grammar rules, lists, reporting, and surrounding; Added comprehensive test suite for pest\_derive grammar parsing and built-in rules; Added fuzzing test fixtures for grammar parsing; Added test coverage for HTTP, JSON, SQL, and TOML grammars; Added test coverage for calculator and JSON parsing.

Dependencies

Pest 2.9.2 workspace and dependency updates

The pest parser library has been updated to version 2.9.2 across its workspace crates (pest, pest\_derive, pest\_generator, pest\_meta, pest\_vm, pest\_debugger, pest\_grammars). This release includes a MSRV bump to Rust 1.83, updates to the \syn\ dependency (v3.0 in the generator), and the addition of \miette\ error support via the \miette-error\ feature. The \pest\_debugger\ crate now depends on \rustyline\ 13 for improved terminal interaction and \thiserror\ 2 for error handling. Fuzzing targets for JSON, HTTP, TOML, and SQL grammars have been added to the \pest\_grammars\ crate.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 61 → 53 (-7.7)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 93 → 93 (+0.0)
  • Architecture 100 → 98 (-1.9)
  • Maturity 54 → 54 (-0.0)
  • Readiness 54 → 40 (-14.6)
  • Security 58 → 58 (+0.0)
  • Performance 85 (new)

Resolved (6)

  • Documentation: no usage examples (README.md)
  • Hotspot: debugger/src/main.rs (debugger/src/main.rs)
  • Hotspot: meta/src/parser.rs (meta/src/parser.rs)
  • Hotspot: pest/src/parser_state.rs (pest/src/parser_state.rs)
  • Off-boarding risk: anonymized user #1
  • PR-triggered workflow without a permissions block

New (6)

  • Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
  • Inconsistent naming for extracting string content from different iterator/span types. Pair and Pairs use as_str(), while Span also uses as_str(). However, Pairs also has concat() which returns a String, whereas Pair::as_str() returns &str. The inconsistency is subtle but Pairs::as_str() behavior (likely joining or taking the first) vs Pairs::concat() is ambiguous. More critically, Pair::as_str() vs Span::as_str() is consistent, but Pairs having both as_str and concat suggests unclear intent on whether as_str should return the raw underlying input slice or a joined string.
  • Off the main sequence: pest
  • Off the main sequence: pest_meta
  • Off-boarding risk: anonymized user #1
  • Three distinct types represent the same conceptual entity (a grammar rule) with overlapping but non-identical structures. ParserRule contains parsing metadata (span), Rule is the AST representation, and OptimizedRule is the post-optimization representation. While they serve different pipeline stages, the duplication of name and ty properties across all three without a shared trait or base type creates cognitive load and potential for drift.

Changes since last survey

  • 8 commits — 8 feature/other, 0 fixes

By area

  • .github/workflows — 3 commits
  • debugger/Cargo.toml — 2 commits
  • pest/src — 2 commits
  • generator/Cargo.toml — 1 commit

Notable commits

  • change: Remove reqwest dependency from pest_debugger (#1209)
  • change: Remove workflow-wide actions write permission (#1208)
  • change: Stop Span::lines from yielding a line past the span (#1201)
  • change: bump syn to 3 (#1210)
  • change: bump version to 2.9.2 (#1206)
  • change: ci: add permissions to fuzzing workflow (#1204)
  • change: ci: add workflow permissions (#1203)
  • change: update to unicode 18 (#1202)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

pest-parser/pest was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9abf9e564f51cf676b69fe250f5ac24a362af207 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-505904ce13c1.