crate-ci/typos
73.4
Strong · 30 September 2026
6.4k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a high-performance spell-checking tool for source code and text files, implemented in Rust. It detects typos and misspellings by leveraging multiple static dictionaries, including corrections from Codespell, Misspell, Wikipedia, and Varcon for dialect variants. The tool provides a command-line interface with extensive configuration options, file-type awareness, and structured reporting formats like SARIF, alongside a GitHub Action for CI integration.
How it got here
2019 — Project rebranding and workspace restructuring
8 changes.
The project was renamed from scorrect to typos and restructured into a multi-crate Rust workspace, replacing the original single-threaded implementation with a modular architecture. This period involved removing legacy assets, benchmarks, and the initial spell-checking logic while establishing new CI infrastructure and a comprehensive benchmark suite for performance comparison.
2020 — Dictionary crate extraction and engine refactoring
11 changes.
This period focused on modularizing the project by extracting spell-checking data into dedicated crates such as codespell-dict, misspell-dict, and typos-dict, each providing generated, case-insensitive word dictionaries. Concurrently, the core typos library was refactored to use the winnow parser combinator library, introducing a new tokenization and typo-checking engine with improved performance and robustness. The work also expanded linguistic support by integrating Varcon for dialect variants and updating data models to handle new part-of-speech categories.
2021–2024 — CLI expansion and codegen testing
12 changes.
The project expanded its CLI capabilities by introducing SARIF reporting, flexible configuration sources, and a ripgrep-style file-type system, while also releasing a cross-platform GitHub Action. Concurrently, the codebase enhanced its dictionary generation infrastructure with new Aho-Corasick and naive backends and significantly increased test coverage to ensure the integrity of generated code and data.
Features
Add Varcon dictionary integration
Integrates the Varcon dictionary into the application by adding a new \varcon\ crate. This includes a generated \codegen.rs\ file containing static variant data (e.g., American vs. British spellings) and a \lib.rs\ that exposes core types like \Category\, \Pos\, and \Tag\ from the \varcon\_core\ library, making dictionary-based variant resolution available to consumers.
crates/varcon/src · high confidence
Added benchmark fixture scripts for Linux, ripgrep, and subtitles
The \benchesuite/fixtures\ directory now includes shell scripts to manage benchmark data sets. These scripts handle downloading and preparing fixtures for the Linux kernel (both clean and built states), the ripgrep tool (clean and built states), and English and Russian OpenSubtitles corpora (full and small samples). This provides the underlying data management layer for the end-to-end benchmark suite.
benchsuite/fixtures · high confidence
Added benchmark utility scripts for spell-checking and search tools
The benchsuite/uut directory now includes shell scripts to manage the download, installation, and version reporting for several benchmark utilities: codespell (v2.0.0), misspell\_go (v0.3.4), ripgrep (v11.0.1), scspell (v2.2), and typos. These scripts automate the setup of these tools within a temporary benchmark directory, allowing the benchmark suite to easily provision specific versions of these text-processing tools for performance testing.
benchsuite/uut · high confidence
Added generated Wikipedia spelling dictionary
The \wikipedia-dict\ crate now exposes a generated \WORD\_DICTIONARY\ containing common misspellings and their corrections. This static data is built from a codegen process and exposed via the crate's public API, enabling applications to look up and correct spelling errors using case-insensitive matching.
crates/wikipedia-dict/src · high confidence
Initial release of the typo correction dictionary assets
The \crates/typos-dict/assets\ directory now contains the core data files for the typo correction engine. This includes \words.csv\, which maps misspelled words to their correct suggestions, and \english.csv\, providing a comprehensive list of valid English words. Additionally, \allowed.csv\ defines terms that should not be flagged as typos (such as technical acronyms, HTML tags, and common contractions), and \.gitattributes\ marks these assets as vendored.
crates/typos-dict/assets · high confidence
Initial release of the typos-dict crate with generated word correction map
The \crates/typos-dict\ crate is introduced, providing a static, case-insensitive map of common typos to their suggested corrections. The dictionary data is generated via codegen and stored in a \phf::Map\ (Perfect Hash Function) for efficient lookups, exposing the corrections through the \WORD\ constant.
crates/typos-dict/src · high confidence
Introduce GitHub Action for typos with cross-platform support and formatted output
This change adds a new GitHub Action (entrypoint.sh and format\_gh.sh) that downloads and runs the 'typos' linter (v1.50.3). The action supports Windows, macOS, and Linux (including aarch64/arm64), allowing users to integrate typo checking into their CI pipelines. It accepts inputs for target files, isolated mode, write changes, and custom config files, and formats output as GitHub workflow warnings using relative paths.
action · high confidence
Introduce codespell-dict crate with generated word dictionary
The \crates/codespell-dict\ crate has been added, providing a Rust library that exposes a generated, case-insensitive word dictionary (\WORD\_DICTIONARY\) via \dict\_codegen.rs\. This change introduces the core data structure and public API for accessing the spell-checking dictionary entries, which are compiled directly into the binary from the codespell project data.
crates/codespell-dict/src · high confidence
Introduce misspell-dict crate with generated case-insensitive dictionary
The new \misspell-dict\ crate provides a generated, case-insensitive misspelling dictionary via the \MAIN\_DICTIONARY\ constant. The dictionary is built from a codegen process and exposed through the crate's public API, enabling consumers to access a pre-built list of common English words and their variants for spell-checking or correction purposes.
crates/misspell-dict/src · high confidence
New codegen backends for Aho-Corasick and naive pattern matching
The \dictgen\ crate now supports generating dictionaries using the Aho-Corasick automaton and a naive match-based approach, in addition to the existing map, ordered map, and trie backends. The Aho-Corasick generator (\aho\_corasick\ feature) produces a hybrid structure that uses a DFA for fast ASCII lookups and an ordered map for Unicode keys, while the naive match generator (\match\ feature) creates a simple Rust \match\ statement for direct string comparison. These additions provide alternative trade-offs between code size and lookup performance for generated dictionary structures.
crates/dictgen · high confidence
New configuration and file-type handling architecture
The CLI now supports loading configuration from \Cargo.toml\ (under \\[package.metadata.typos\]\) and \pyproject.toml\ (under \\[tool.typos\]\), in addition to standard \.typos.toml\ files. It introduces a ripgrep-style file-type system with a comprehensive default type list (including support for \dtso\ devicetree files) and type-specific dictionaries that ignore common identifiers for languages like Go, Python, Rust, and Julia. The engine also adds file-type-specific settings to skip checking lockfiles and certificates, and provides a new \file\_type\_specifics\ module to centralize these defaults.
crates/typos-cli/src · high confidence
New spell-checker benchmark suite
A new end-to-end benchmark suite has been added to the \benchsuite\ directory, providing a standardized way to compare the performance of various spell-checking tools. The \benchsuite.sh\ script orchestrates benchmarks using \hyperfine\ across multiple test fixtures, including Linux source trees (\linux\_clean\, \linux\_built\), ripgrep source trees, and English/Russian subtitle files. It measures and reports execution times for tools such as \rg\ (as a baseline), \typos\, \misspell\, \codespell\, and \scspell\, generating Markdown and JSON reports for each run.
benchsuite · high confidence
Support for English dialect variants in spelling corrections
The \typos-vars\ crate now provides spelling correction data that distinguishes between American, British, Canadian, and Australian English variants. This is implemented via a generated trie (\VARS\) and a \corrections\ function that maps specific word variants to the appropriate correction based on the selected \Category\. Users can now receive context-aware suggestions that respect their preferred English dialect.
crates/typos-vars/src · high confidence
Removals
Removal of assets directory and spell-checking logic
The \assets\ directory has been completely removed, deleting the \main.go\ entry point, the \words.go\ generated dictionary, and the \words.csv\ source data. This eliminates the application's built-in spell-checking capability, which previously used these assets to map common misspellings to their correct forms.
assets · high confidence
Removal of initial spell-checking implementation
The initial spell-checking library and CLI entry point have been removed. This deletes the \src/lib.rs\ file, which contained the basic tokenization logic, file processing, and dictionary correction functions, as well as the \src/main.rs\ file, which provided the command-line interface using \structopt\ and \ignore\ for file walking. This change eliminates the original single-threaded, regex-based typo detection capability.
src · high confidence
Removed legacy \`scorrect\` benchmarks
The \benches/corrections.rs\, \benches/file.rs\, and \benches/tokenize.rs\ files have been deleted. These benchmarks, which used the unstable \test\ crate to measure the performance of the \scorrect\ library's correction, file processing, and tokenization logic, are no longer part of the codebase.
benches · high confidence
Behavioural changes
CLI restructured with new argument parsing and SARIF reporting support
The CLI binary has been refactored to use a new argument definition structure (args.rs) and a modular reporting system (report.rs). This change introduces a new --format flag allowing users to select output styles including Silent, Brief, Long, Json, and the newly added Sarif format for structured error reporting. Additionally, new debugging flags have been added to help users inspect the spell-checking process, including --file-list to read paths from a file or stdin, --sort to order results, --highlight-identifiers and --highlight-words to stylize checked tokens, and --force-exclude to ensure excluded files are respected even when passed explicitly.
crates/typos-cli/src/bin · high confidence
Introduce new tokenization and typo-checking engine
The \crates/typos\ library has been refactored to use a new parsing backend built on the \winnow\ parser combinator library, replacing the previous implementation. This change introduces a new \Tokenizer\ and \Dictionary\ API in \check.rs\ and \dict.rs\, enabling more robust handling of various token types (such as UUIDs, emails, and URLs) and improving performance through optimized UTF-8 validation and slice-based processing.
crates/typos · high confidence
Project renamed to typos with comprehensive configuration and CI overhaul
The project has been renamed from 'scorrect' to 'typos', reflected in the repository name, documentation, and configuration files. This change introduces a formal JSON Schema for configuration validation (config.schema.json), establishes a pre-commit integration with specific hooks for YAML, JSON, TOML, and commit message checking, and migrates the CI infrastructure from Travis CI and Appveyor to GitHub Actions using a composite action. The Docker build process has been updated to use a Debian bullseye-slim base image, and the project now enforces stricter coding standards via .clippy.toml, which disallows methods like map\_or and for\_each in favor of more legible alternatives.
(repo-wide) · high confidence
Updated misspelling dictionary assets
The misspelling dictionary assets have been updated and reformatted to align with the project's internal structure. This change introduces new source files for English, Ukrainian, and US/UK variant mappings, ensuring the dictionary reflects the latest corrections and linguistic standards.
crates/misspell-dict/assets · medium confidence
Updated spell-checking dictionary assets
The spell-checking dictionary assets in the codespell crate have been updated to the latest version of the codespell project. This change replaces the previous dictionary files with new \compatible.csv\ and \dictionary.txt\ files, ensuring that the tool recognizes the most recent set of common typos and misspellings for more accurate code review feedback.
crates/codespell-dict/assets · high confidence
Varcon-core parser and data model updated to support new POS types and version 2020.12.07
The varcon-core crate has been refactored to use the winnow parsing library, replacing the previous implementation. This change introduces support for new Part-of-Speech (POS) categories, specifically Interjection and Preposition, which are now available in the Pos enum. The data model has been adjusted to track entry notes separately from comments, and the parser now correctly handles verified status and level metadata in cluster headers. Additionally, the core types have been updated to align with the Varcon version 2020.12.07 specification, including changes to how plural annotations and abbreviations are parsed (swallowed/ignored). The API is now independent of the internal winnow implementation details, exposing a cleaner ParseError and ClusterIter interface for users.
crates/varcon-core · high confidence
Test coverage
Add Aho-Corasick benchmark suite; Added code generation test for Varcon dictionary; Added code generation tests for spelling variations; Added command-line interface tests for typos-cli; Added regression test for file removal during walk; Added tests for dictionary code generation and compatibility; Added tests for dictionary code generation and data integrity; Migrate benchmarks to Divan and restructure bench files.
Dependencies
Project rebranded to Typos and restructured as a Rust workspace
The project has been renamed from 'scorrect' to 'typos' and reorganized into a multi-crate workspace. The main CLI binary is now in the \typos-cli\ crate (v1.50.3), while dictionary sources are split into dedicated crates (\typos-dict\, \misspell-dict\, \codespell-dict\, \wikipedia-dict\, \varcon\, \varcon-core\) and a shared code-generation library (\dictgen\). The workspace enforces Rust edition 2024 and a minimum supported Rust version (MSRV) of 1.95, and the CLI now depends on \clap\ v4, \winnow\ v1.0, and \schemars\ v1.2.1.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 71 → 73 (+2.7)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 94 → 94 (+0.3)
- Architecture 100 → 98 (-2.2)
- Maturity 67 → 67 (+0.0)
- Readiness 76 → 70 (-6.0)
- Security 63 → 80 (+16.8)
- Performance 91 (new)
Resolved (80)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- …and 60 more
New (10)
- Inconsistent naming for factory methods. 'from_dir', 'from_file', 'from_toml' suggest source-based construction, while 'from_defaults' suggests a state-based construction. While distinct, the pattern 'from_*' is used for both source and default state, which can be slightly confusing. More importantly, 'update(source: Self)' is used for merging, but there is no 'merge' or 'combine' alias, which is fine, but 'from_defaults' is an outlier in the 'from_*' naming convention if it doesn't take a source.
- Inconsistent parameter naming for the input buffer. One uses 'buffer' and the other uses 'content' (in Tokenizer methods), but within the check module itself, the distinction is clear (str vs &[u8]). However, looking at Tokenizer, we see 'parse_str(content: str)' and 'parse_bytes(content: &[u8])'. The naming 'content' vs 'buffer' is inconsistent across the API surface for the same conceptual input.
- Inconsistent parameter naming for the same logical argument. The core trait uses 'ident' and 'word', while the implementation 'BuiltIn' uses 'ident_token' and 'word_token'. This creates confusion about whether the argument is the raw token or a processed identifier/word object.
- Off the main sequence: dictgen
- Off the main sequence: typos
- Off the main sequence: varcon-core
- Outdated: clap
- Outdated: encoding_rs
- Outdated: toml
- Outdated: unicode-ident
Changes since last survey
- 16 commits — 12 feature/other, 4 fixes
By area
- (repo) — 6 commits
- (root) — 4 commits
- .github/workflows — 4 commits
- crates/typos-cli — 2 commits
Notable commits
- fix: Fix merge conflict in rust-next workflow
- fix: Merge pull request #1619 from szepeviktor/fix-wf
- fix: Merge pull request #1621 from antonkesy/fix-case-correct-panic
- fix: fix(cli): Don't panic on non-ASCII corrections
- change: Merge pull request #1618 from epage/template
- change: Merge pull request #1620 from epage/test
- change: Merge pull request #1625 from epage/maturin
- change: chore(ci): Clean up test action
- change: chore(ci): Update maturin
- change: chore: Release
- change: chore: Release
- change: chore: Rename master to main
- change: chore: Update from _rust template
- change: docs: Update changelog
- change: docs: Update changelog
- change: style: Make clippy happy
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
crate-ci/typos was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 00f422f3b19c57bc6338715ebfe3316d38768461 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.