huggingface/tokenizers
71.7
Strong · 29 September 2026
34.2k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a high-performance text tokenization library that has been restructured into a modular, inference-focused architecture. It provides a read-only pipeline for encoding and decoding text using optimized models like BPE, Unigram, and WordPiece, with significant performance improvements via SIMD acceleration and efficient memory management. The library exposes bindings for Python and Node.js, enabling integration into machine learning workflows, while explicitly excluding training capabilities which have been moved to a separate crate.
Features
Introduce bitcannon SIMD pre-tokenizer
Adds a new high-performance pre-tokenizer library named \bitcannon\ in the \tokenizers/bitcannon\ directory. This Rust library implements a bitstream-based execution model (inspired by Parabix) to process text in 64-byte blocks, offering significantly lower and more consistent latency across different tokenization grammars compared to traditional finite-state machine approaches. The initial release includes the core engine, a specification document, and implementations for the class-run family of tokenizers (such as Whitespace, Punctuation, and BERT splits) as well as literal character delimiter splitting.
tokenizers/bitcannon · high confidence
Introduction of the new \`tk-encode\` inference crate
The \tokenizers\ library now includes a dedicated \tk-encode\ crate for the inference runtime, separating it from the legacy engine. This new crate provides a read-only \PipelineTokenizer\ that processes text through Normalizers, PreTokenizers, Models, and PostProcessors to produce Encodings. It is designed for performance and small binary size by excluding \serde\ dependencies (using \hifijson\ for parsing instead) and excluding training logic. Users can enable specific tokenizer models (BPE, Unigram, WordPiece, WordLevel) and features like parallelism or HTTP-based model downloading via feature flags.
tokenizers/tk-encode/src · high confidence
New Python binding examples for encoding, batching, and training
Added six new example scripts in the Python bindings directory to demonstrate core tokenizer capabilities. \encode\_decode.py\ and \from\_pretrained.py\ show basic tokenization and loading from the Hugging Face Hub. \batch\_padding.py\ and \truncation.py\ illustrate how to configure and apply padding and truncation strategies both globally and per-call. \multiprocessing\_workers.py\ demonstrates efficient parallel encoding using Python's multiprocessing pool. Finally, \train\_parity\_bpe.py\ provides a reference implementation for training a parity-aware BPE tokenizer across multiple languages, though it is currently commented out as the Python bindings do not yet expose the training functionality.
bindings/python/examples · high confidence
New Unicode Scripts pre-tokenizer for script-aware tokenization
Users can now apply a pre-tokenizer that splits text boundaries based on Unicode script changes (e.g., separating Latin from CJK). This new \UnicodeScripts\ pre-tokenizer, available in \tk-encode\, groups characters by their script (such as Han, Hiragana, Katakana, Latin, etc.) and emits spans where the script changes, while treating spaces and common punctuation as neutral characters that do not trigger splits. It includes specific handling to map Hiragana, Katakana, and the prolonged sound mark (ー) to the Han script for CJK text processing.
_tokenizers/tk-encode/src/pre\_tokenizers/unicode\scripts · high confidence
New Unigram and WordLevel tokenizer models in tk-encode
The \tk-encode\ crate now includes implementations for the Unigram and WordLevel tokenization models. The Unigram model introduces a lattice-based structure to perform Viterbi decoding for optimal tokenization paths and supports sampling from possible encodings, along with features like byte-fallback and configurable unknown token handling. The WordLevel model provides a straightforward lookup-based tokenization strategy. Both models are integrated into the new encoding pipeline, replacing legacy engine components and enabling more flexible and performant text tokenization for users.
tokenizers/tk-encode/src/models/unigram · high confidence
New WordPiece tokenizer model implementation
The \tk-encode\ crate now includes a new \WordPiece\ model implementation in \tokenizers/tk-encode/src/models/wordpiece\. This change introduces a builder-based API for configuring vocabulary, unknown tokens, and subword prefixes, replacing the legacy encode engine and its serialization layer. Users can now construct WordPiece tokenizers programmatically, with the model handling tokenization logic including character limits and subword continuation prefixes.
tokenizers/tk-encode/src/models/wordpiece · high confidence
New dev-time tool to generate bitcannon atom classification tables
A new \bitmap\_gen\ crate has been added to bake Unicode character classification data into static tables for the \bitcannon\ tokenizer. This dev-time tool generates \atom\_tables.rs\ by mapping codepoints to specific atom tags (handling ASCII, 2-byte, 3-byte, BMP RLE, and astral ranges) based on Unicode general categories and script properties (such as Han). This ensures the runtime tokenizer uses pre-computed, consistent classification data rather than performing per-byte range tests.
_tokenizers/bitmap\gen · high confidence
New tokenization decoders and encoding utilities in tk-encode
The \tk-encode\ crate now includes a comprehensive set of decoder implementations (BPE, ByteFallback, ByteLevel, CTC, Fuse, Metaspace, Replace, Strip, and WordPiece) and supporting utilities for padding, truncation, caching, and parallel execution. This adds the capability to decode tokenized sequences back into readable text using various standard algorithms and to manage encoding parameters like padding direction and strategy, significantly expanding the encoding/decoding pipeline's functionality.
tokenizers/tk-encode/src/utils · high confidence
Node bindings scaffolded with build and tooling configuration
The node bindings directory has been initialized with essential build and development infrastructure. This includes a Cargo configuration for cross-compiling the native Rust backend to aarch64 Linux musl targets, a Yarn 3.5.1 release binary for package management, and standard editor tooling (EditorConfig, Prettier, Taplo, Git attributes, and Gitignore) to ensure consistent code formatting and repository hygiene.
bindings/node · high confidence
Behavioural changes
Introduce canonical tokenizer.json 2.0 reader and writer
The \tk-serialize\ crate now implements the canonical \tokenizer.json\ version 2.0 format, replacing the legacy 1.0 handling. It provides a strict reader that only accepts version 2.0 files (refusing legacy 1.0 files and requiring offline conversion via \tk-convert\ for older formats) and a writer that produces version 2.0 output. This change standardizes the serialization schema, enforcing exact key ordering and explicit component types (e.g., \TemplateProcessing\ instead of legacy spellings like \BertProcessing\), and ensures that legacy constructs like \Metaspace\ pre-tokenizers or \ByteLevel\ pre-tokenizers are refused unless converted, thereby guaranteeing that the serialized configuration matches the canonical specification exactly.
tokenizers/tk-serialize · high confidence
Legacy tokenizer.json upgrade pass
The tk-convert crate now provides a JSON-to-JSON upgrade pass that rewrites legacy tokenizer configurations into a canonical format (version 2.0). This ensures that the serialization reader no longer needs to maintain backwards-compatibility branches. The pass handles specific legacy patterns, such as inferring missing model types (e.g., distinguishing BPE from WordPiece based on merge presence), rewriting merge entries from space-separated strings to arrays, and normalizing Metaspace and ByteLevel pre-tokenizers into their canonical decoder and normalizer forms. It also validates ambiguous configurations and refuses them with descriptive errors rather than guessing.
tokenizers/tk-convert · high confidence
New SIMD-accelerated pre-tokenizers for BERT, ByteLevel, and delimiter-based splitting
The \tk-encode\ pre-tokenization pipeline now includes native implementations for \BertPreTokenizer\, \ByteLevel\, \CharDelimiterSplit\, \Digits\, \FixedLength\, \Punctuation\, \Split\, \Whitespace\, and \WhitespaceSplit\. These new components replace legacy scalar logic with \bitcannon\-powered SIMD finite-state machines and optimized byte-scanning, delivering byte-exact behavior while significantly improving performance. Users benefit from faster tokenization for common patterns (such as BERT-style punctuation isolation, GPT-style regex splitting, and fixed-length chunking) without changing their configuration, as the new implementations are designed to be drop-in replacements for existing pre-tokenizer definitions.
_tokenizers/tk-encode/src/pre\tokenizers · high confidence
New encode pipeline with configurable options and parallel processing
The tokenizer's encoding pipeline has been restructured to use a new \EncodeOptions\ struct that exposes \add\_special\_tokens\ and \encode\_special\_tokens\ flags, allowing users to control how special tokens are handled during encoding. The pipeline now supports parallel processing for large inputs (above 8KB) via a thread pool with sharded scratch buffers to reduce contention, and includes a new \PipelinePostProcessor\ that manages template-based wrapping of sequences with type IDs.
tokenizers/tk-encode/src/tokenizer/pipeline · high confidence
New high-performance BPE merge engines and pipeline integration
The BPE model in \tk-encode\ has been replaced with a new pipeline implementation that uses two specialized merge engines: a multipass engine for short words and a hot/cold queue engine for longer words, significantly improving tokenization speed. This change introduces a new \PipelineBPE\ model, a word cache for repeated tokens, and a new serialization format (\BpeConfig\) for loading and saving models, while removing the legacy encode engine.
tokenizers/tk-encode/src/models/bpe · high confidence
New optimized vocabulary storage and added-token handling in tk-encode
The \tk-encode\ tokenizer engine now uses a new vocabulary backend located in \tokenizers/tk-encode/src/vocab\. This introduces \BucketVocabStore\, which replaces standard hash maps with a minimal perfect hash function (MPHF) and packed keys to reduce memory usage and improve lookup speed for token IDs. It also adds \AddedVocabulary\ and \AddedToken\ structures, allowing users to append new tokens to an existing trained model with configurable behaviors such as \single\_word\ matching, whitespace stripping (\lstrip\/\rstrip\), and normalization. The implementation leverages SIMD instructions (via \pshufb\) and multi-bucket prefix matching to accelerate byte-level token matching.
tokenizers/tk-encode/src/vocab · high confidence
New rc0 release with read-only pipeline tokenizer and legacy config upgrade
The tokenizers library has been restructured into separate crates (\tk\_encode\, \tk\_serialize\, \tk\_convert\) with this release providing a read-only \PipelineTokenizer\ that encodes and decodes based on a \tokenizer.json\ configuration. This version does not yet support building tokenizers programmatically, saving, or training; it focuses on inference and includes a legacy config upgrade pass (\tk\_convert\) to rewrite older tokenizer configurations into the canonical form required by the new reader.
tokenizers/src · high confidence
New tokenization engine with redesigned Encoding and Pattern APIs
The \tk-encode\ crate introduces a new tokenization engine, replacing the legacy encode engine and its serialization layer. This change brings a redesigned \Encoding\ struct that explicitly manages token IDs, type IDs, tokens, word IDs, offsets, special token masks, attention masks, and sequence ranges, ensuring compatibility with existing Python pickle formats. It also introduces a new \Pattern\ trait for pre-tokenization, supporting splitting via literals (offloaded to the \atomsplit\ module for performance), system regexes, and character predicates, along with a \SplitDelimiterBehavior\ enum to control how delimiters are handled during splitting.
tokenizers/tk-encode/src/tokenizer · high confidence
New tokenizer normalizer implementations in tk-encode
The \tokenizers/tk-encode/src/normalizers\ module now contains fresh, standalone implementations for the core text normalizers used in the encoding pipeline. This includes \BertNormalizer\ (handling text cleaning, Chinese character spacing, accent stripping, and lowercasing), \ByteLevel\ (converting bytes to their character representations), \MetaspaceNormalizer\ (inserting the SentencePiece delimiter and supporting prepend behaviors), \PrecompiledNormalizer\ (applying precompiled character maps), \Prepend\ (adding a prefix string), \Replace\ (pattern-based literal or regex replacement), \Strip\ (trimming whitespace), \StripAccents\ (removing combining marks), \Lowercase\, and Unicode normalization forms (NFC, NFD, NFKC, NFKD) plus the NMT normalizer. These modules replace the legacy normalizer logic, providing the specific text transformation capabilities required by models like BERT, T5, and Albert within the new encode path.
tokenizers/tk-encode/src/normalizers · high confidence
Optimized BPE encoding with character folding shortcuts
The BPE tokenizer now uses a precomputed 'fold' table to accelerate encoding. For characters that are guaranteed to merge into a single token without interference from neighbors, the encoder skips the standard merge loop and emits the token directly. This optimization significantly reduces processing time for common characters while maintaining exact encoding correctness.
tokenizers/tk-encode/src/models/bpe/fold · high confidence
Python bindings rework: new Tokenizer API with Encoding views and padding/truncation options
The Python bindings have been reworked to expose a modern Tokenizer API. Users can now load tokenizers from local files or directly from the Hugging Face Hub via \Tokenizer.from\_pretrained\. The \Encoding\ object provides \ids\, \type\_ids\, and \attention\mask\ as both Python lists and read-only NumPy array views (\\\_array\), allowing zero-copy access for downstream libraries like PyTorch. Padding and truncation are now controlled via explicit \Padding\ and \Truncation\ objects, supporting options like \pad\_to\_multiple\_of\, \direction\, and \strategy\, which can be set globally on the tokenizer or passed per-call. The bindings also support pickling for multiprocessing and include auto-generated \.pyi\ stubs for type checking.
bindings/python · high confidence
Simplified processor module structure
The processor module in the encode path has been streamlined to re-export only the byte-level pre-tokenizer, while complex post-processing logic (including the PostProcessorWrapper and Sequence processor) has been moved to the tk-convert crate. Users relying on the encode path will now interact with a cleaner API where post-processing is handled via the pipeline configuration rather than direct processor instantiation.
tokenizers/tk-encode/src/processors · medium confidence
Tokenizer training logic moved to a dedicated tk-train crate
The training components for tokenizers (BPE, Unigram, WordLevel, and WordPiece) have been extracted from the inference-focused tk-encode crate into a new tk-train crate. This separation ensures that the inference library no longer carries training-related dependencies or coupling. Users now drive training directly by feeding data to a Trainer instance and calling train, rather than using legacy tokenizer-level entry points. The crate also includes a new TrainerWrapper enum for dispatching across model types and a dedicated serde module for serializing AddedTokens during training.
tokenizers/tk-train · high confidence
Tokenizers library enters v1.0.0 release candidate with major architectural changes
The library has shifted from version 0.23 to a v1.0.0 release candidate, introducing a new pipeline-native architecture that strips the legacy configuration layer. This change removes the \tk-train\ crate from the default build, meaning training is currently unavailable, and replaces the old \Tokenizer\ object with a read-only \PipelineTokenizer\. Consequently, Python bindings can no longer modify tokenizer components (such as setting normalizers) or save modified tokenizers, and several legacy APIs like \Tokenizer::save\ and \from\_pretrained\ are deprecated in favor of new serialization and loading methods. The repository also adds a \CITATION.cff\ file for academic referencing and updates the \README\ to reflect the new installation and usage patterns.
(repo-wide) · high confidence
Tokenizers library restructured into modular crates with rc0 release
The tokenizers library has been restructured into distinct crates: \tk\_encode\ for inference engines, \tk\_serialize\ for reading tokenizer configurations, and \tk\_convert\ for upgrading legacy JSON files. This release (rc0) introduces a read-only \PipelineTokenizer\ that encodes and decodes text based on a \tokenizer.json\ file, but explicitly excludes the full \Tokenizer\ object model (such as \add\_tokens\, \save\, or training capabilities) which are deferred for v1. The change also includes a new Makefile for managing test data, benchmarks, and size profiling, along with a script to verify Python binding integrity.
tokenizers · high confidence
Test coverage
Added Python bindings test suite; Added benchmark suite for tokenizer encode/decode performance; Added regression tests for tokenizer encoding parity.
Dependencies
Node.js bindings restructured with new Rust engine and platform packages
The Node.js bindings have been restructured to use a new Rust-based engine (\tk-encode\, \tk-serialize\, \tk-convert\) instead of the previous umbrella crate, exposing the pipeline encode path via N-API 3. The main \tokenizers\ package is now version 1.0.0-dev.0 and includes platform-specific sub-packages (e.g., \tokenizers-linux-x64-gnu\) for prebuilt binaries, all targeting Node.js \>= 10. The Rust dependencies have been updated to use \napi\ 3 and \pyo3\ 0.29 for the Python bindings, and the Node.js build now uses \napi-build\ 2.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 72 → 72 (-0.1)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.
Lenses
- Code Health 93 → 86 (-7.2)
- Architecture 98 → 95 (-2.7)
- Maturity 60 → 69 (+8.9)
- Readiness 87 → 67 (-20.9)
- Security 72 → 79 (+6.8)
Resolved (225)
- AddedVocabulary::add_tokens (cognitive 26) (tokenizers/src/tokenizer/added_vocabulary.rs)
- AddedVocabulary::find_matches (cognitive 21) (tokenizers/src/tokenizer/added_vocabulary.rs)
- AllEntities.resolve_pendings (cognitive 21) (docs/source/_ext/entities.py)
- BPE::merge_word (cognitive 40) (tokenizers/src/models/bpe/model.rs)
- BPE::merge_word (cyclomatic 16) (tokenizers/src/models/bpe/model.rs)
- BPEVisitor::visit_map (cognitive 31) (tokenizers/src/models/bpe/serialization.rs)
- BPEVisitor::visit_map (cyclomatic 22) (tokenizers/src/models/bpe/serialization.rs)
- BpeTrainer::do_train (cognitive 27) (tokenizers/src/models/bpe/trainer.rs)
- BpeTrainer::tokenize_words (cognitive 30) (tokenizers/src/models/bpe/trainer.rs)
- ByteFallback::decode_chain (cognitive 28) (tokenizers/src/decoders/byte_fallback.rs)
- Change coupling clique: byte_level_bpe.py, char_level_bpe.py, sentencepiece_bpe.py (bindings/python/py_src/tokenizers/implementations/byte_level_bpe.py)
- Change coupling: trainer.rs ↔ trainer.rs (tokenizers/src/models/bpe/trainer.rs)
- ClassTooLong: NormalizedString (tokenizers/src/tokenizer/normalizer.rs)
- ClassTooLong: ParityBpeTrainer (tokenizers/src/models/bpe/parity_trainer.rs)
- ClassTooLong: PyParityBpeTrainer (bindings/python/src/trainers.rs)
- ClassTooLong: PyTokenizer (bindings/python/src/tokenizer.rs)
- ClassTooLong: TokenizerImpl (tokenizers/src/tokenizer/mod.rs)
- Concentrated knowledge decay
- Documentation: no installation or build instructions (docs/source/installation/main.rst)
- Duplicated block (10 lines × 2) (tokenizers/src/models/bpe/parity_trainer.rs)
- …and 205 more
New (229)
- AddedVocabulary::add_tokens (cognitive 19) (tokenizers/tk-encode/src/vocab/bucket_added_vocabulary.rs)
- BpeTables::build (cognitive 23) (tokenizers/tk-encode/src/models/bpe/tables.rs)
- BpeTables::build (cyclomatic 16) (tokenizers/tk-encode/src/models/bpe/tables.rs)
- BpeTrainer::do_train (cognitive 25) (tokenizers/tk-train/src/trainers/bpe/mod.rs)
- BpeTrainer::tokenize_words (cognitive 22) (tokenizers/tk-train/src/trainers/bpe/mod.rs)
- Buckets::build_nibble_table (cognitive 16) (tokenizers/tk-encode/src/vocab/buckets.rs)
- Buckets::from_tokens (cognitive 18) (tokenizers/tk-encode/src/vocab/buckets.rs)
- ByteFallback::decode_chain (cognitive 28) (tokenizers/tk-encode/src/decoders/byte_fallback.rs)
- ByteLevelFold::fold (cognitive 16) (tokenizers/tk-encode/src/models/bpe/fold/byte_level.rs)
- CI installs an unverified third-party binary (.github/workflows/bitcannon.yml)
- Change coupling: model.rs ↔ cache.rs (tokenizers/tk-encode/src/models/bpe/model.rs)
- Change coupling: model.rs ↔ word.rs (tokenizers/tk-encode/src/models/bpe/model.rs)
- Change coupling: unigram.rs ↔ wordlevel.rs (tokenizers/tk-train/src/trainers/unigram.rs)
- Change coupling: wordlevel.rs ↔ wordpiece.rs (tokenizers/tk-train/src/trainers/wordlevel.rs)
- ClassTooLong: ParityBpeTrainer (tokenizers/tk-train/src/trainers/bpe/parity_trainer.rs)
- Duplicate encoding entry points with different signatures and return types. encode is async/handle-based and takes impl Into<Inputs>, while encode_into is synchronous, takes a str, and writes to an output buffer. This splits the API surface unnecessarily and confuses users about which encoding path to use.
- Duplicate pre-tokenization entry points with inconsistent signatures. The module-level bitcannon.pre_tokenize and the class method CharDelimiterSplit.pre_tokenize perform similar pre-tokenization logic but have different parameter lists (e.g., starts vs _tags usage, different mutability requirements). This forces users to choose between a generic module function and a specific class method without a clear hierarchy.
- Duplicated block (10 lines × 2) (tokenizers/tk-encode/src/pre_tokenizers/unicode_scripts/scripts.rs)
- Duplicated block (10 lines × 2) (tokenizers/tk-encode/src/pre_tokenizers/unicode_scripts/scripts.rs)
- Duplicated block (10 lines × 2) (tokenizers/tk-train/src/trainers/bpe/mod.rs)
- …and 209 more
Changes since last survey
- 42 commits — 34 feature/other, 8 fixes
By area
- .github/workflows — 13 commits
- bindings/python — 7 commits
- (root) — 6 commits
- (repo) — 4 commits
- tokenizers/tk-encode — 4 commits
- bindings/node — 2 commits
- .github/dependabot.yml — 1 commit
- docs/source — 1 commit
- tokenizers/Cargo.lock — 1 commit
- tokenizers/Cargo.toml — 1 commit
- tokenizers/bitcannon — 1 commit
- tokenizers/bitcanon — 1 commit
Notable commits
- fix: fix (#2443)
- fix: fix test
- fix: fix(ci): harden workflow files flagged on #2119 (#2421)
- fix: fix(ci): harden workflow files flagged on #2415 (#2419)
- fix: fix(ci): skip doc builds on the v1 PR, drop dead doc-comment workflows
- fix: fix(release): max number of keywords is 5 (#2444)
- fix: fix: crates.io release env name (#2441)
- fix: 🐛 Fix inline vocab key packing on big-endian targets (#2364)
- change: Add installation section to README
- change: Delete for_the_blogpost.md
- change: Drop fmt/clippy from bitcannon CI, rust.yml already runs them (#2440)
- change: Key the doc build concurrency group on the pull request number (#2401)
- change: Merge branch 'main' into feat/train_encode_split
- change: Merge remote-tracking branch 'origin/main' into feat/train_encode_split
- change: Publish to npm via Trusted Publishing (#2439)
- change: Rename bitcanon -> bitcannon (#2434)
- change: Scope GITHUB_TOKEN permissions per job (#2448)
- change: This was not super fair yet
- change: Tokenizers v1 (#2119)
- change: byte-level: index the char->byte table by codepoint instead of hashing (#2178)
- …and 22 more
Architecture
- Containers 0 added · 0 removed · contexts 0 added · 1 removed · edges 0 added · 0 removed
Removed bounded contexts (1)
- python
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
huggingface/tokenizers was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit bbccb0513ff9afda385ca5c85c66eddb1318cfc7 — the exact code this score is about.
- Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fbec9b1e08c2.