marcelroed/gigatoken
61.2
Adequate · 30 September 2026
33.1k
lines of production code
Rust
with Python
2
measurements over time
What this system is
Gigatoken is a high-performance, Rust-based tokenizer library with Python bindings, designed to replace or accelerate existing tokenization workflows. It provides native support for loading HuggingFace, tiktoken, and SentencePiece models while offering drop-in compatibility adapters for HuggingFace Transformers and tiktoken. The system features a parallel batch encoding engine with SIMD-accelerated pretokenization, capable of processing large-scale structured data formats like JSONL and Parquet efficiently.
How it got here
2025 — Initial release and core engine development
11 changes.
This period marks the initial release of the Gigatoken project, establishing a high-performance Rust-based tokenizer library with Python bindings. The work focused on building the core parallel batch encoding engine, implementing native support for various input formats and compression, and creating a comprehensive test and benchmarking infrastructure to validate performance against existing libraries.
2026 — High-performance tokenizer rewrite
8 changes.
This period focused on a comprehensive rewrite of the core BPE tokenizer and Python bindings in Rust to achieve high performance through SIMD acceleration, optimized caching, and parallel processing. The work included implementing cross-platform release automation, establishing rigorous profiling infrastructure for hardware-specific optimization, and providing new Python-based examples and benchmarking suites to demonstrate the improved throughput and memory efficiency.
Features
Experimental notebooks for Unicode classification and data analysis
Added a suite of Marimo notebooks in the \notebooks/\ directory to support data exploration and experimental implementation of Unicode character classification. These include \bit\_patterns.py\ and \simd\_unicode.py\ for developing SIMD-based classification logic, \dfa\_states.py\ for analyzing deterministic finite automaton states, \gather\_free.py\ for building trie-style lookup tables, \data\_view.py\ for analyzing codepoint frequencies in datasets, and \inspect\_tokenizers.py\ for examining tokenizer internals.
notebooks · high confidence
High-performance BPE tokenizer with optimized caching and multi-architecture support
The BPE module has been rewritten to significantly improve encoding performance and memory efficiency. It introduces a custom, prefetch-aware pretoken cache (\ShortPretokenCache\) that uses 2 MiB huge pages and hardware prefetch hints to reduce TLB misses and cache stalls, alongside a flat \PairRankTable\ for faster merge lookups on cache misses. The implementation now supports multiple tokenizer families including tiktoken (GPT-2, DeepSeek, Qwen, Olmo3), SentencePiece (with NFC normalization and SIMD-accelerated character mapping), and fairseq-style vocabs. It also adds support for Python 3.10-3.14, parallel encoding for large documents, and configurable cache budgets to limit memory usage per worker.
src/bpe · high confidence
Initial release of Gigatoken with documentation and dependency lockfile
This change introduces the Gigatoken project, a high-performance tokenizer library. The repository now includes the core \README.md\ detailing installation, usage (including HuggingFace and Tiktoken compatibility modes), and extensive cross-platform benchmark tables. It also adds the \CONTRIBUTING.md\ guidelines, an \MIT\ license, a \.python-version\ file specifying Python 3.13, and a \uv.lock\ file pinning dependencies such as \awkward\ (2.10.0) and \cffi\ (2.1.0).
(repo-wide) · high confidence
Introduce gigatoken library with native Rust tokenizer, HuggingFace/tiktoken compatibility, and CLI benchmarking
The new gigatoken package provides a high-performance tokenizer backend implemented in Rust, exposing a unified Python \Tokenizer\ class that loads HuggingFace \tokenizer.json\ files, raw SentencePiece \.model\ files, and OpenAI \.tiktoken\ vocabulary files without requiring \transformers\, \tokenizers\, or \tiktoken\ as dependencies. It includes built-in adapters (\as\_hf()\ and \as\_tiktoken()\) to drop into existing HuggingFace \transformers\ and \tiktoken\ codebases, while also offering file sources for parallel processing and a CLI tool (\gigatoken bench\) for measuring encode throughput and validating token IDs against the HuggingFace library.
gigatoken · high confidence
Native support for compressed files, JSONL, and Parquet inputs
The input system now natively reads compressed files (Gzip and Zstd) and structured formats (JSONL and Parquet) alongside plain text. Users can load data from \.gz\, \.zst\, \.jsonl\, and \.parquet\ files; the system automatically detects compression from file extensions and content format from the uncompressed stem. JSONL files are parsed by extracting text from a configurable field (defaulting to "text"), while Parquet files extract text from a specified column, with parallel processing available for both formats to improve throughput.
src/input · high confidence
New cross-library tokenizer throughput benchmarking suite
Added a new benchmarking infrastructure in the \benchmarks/compare\ directory to measure and compare tokenizer throughput across Hugging Face tokenizers, tiktoken, and gigatoken. The suite includes \measure.py\ for single-library execution, \sweep.py\ for sequential cross-library sweeps with deduplication of shared tokenizer definitions, \results.py\ for merging measurements and rendering README tables, and \plot\_throughput.py\ for generating bar charts. Initial results are stored in \benchmarks/results.json\ and coverage data in \benchmarks/families.json\, enabling users to see performance comparisons and model family coverage.
benchmarks · high confidence
New cross-platform release builder and CPU profiling script
Added a new Python script, build\_release\_cross\_platform.py, to automate building and validating PyPI releases for macOS, Linux, and Windows across multiple architectures and Python versions, including a specific clang shim to handle Windows cross-compilation. Also added a Fish shell script, profile-cpu.fish, to record CPU performance traces for benchmarks using cargo-instruments.
scripts · high confidence
New high-performance SIMD pretokenizers for cl100k, o200k, and related tokenization schemes
The fast pretokenization module now includes optimized, SIMD-accelerated implementations for the cl100k\_base (GPT-3.5/GPT-4), o200k\_base (GPT-4o), DeepSeek V3, Kimi (K2 family), Nemotron-3, Olmo3, and Qwen (2/3.5) tokenization schemes. These new pretokenizers replace or supplement previous scalar implementations by utilizing a shared mask-scanner infrastructure that processes 64-byte batches using aarch64 NEON and x86\_64 AVX-512/AVX2 instructions. This architecture enables branchless boundary detection and significantly faster token extraction for these specific model families, while maintaining compatibility with scalar fallbacks for non-SIMD environments.
src/pretokenize/fast · high confidence
New high-performance pretokenization module with multi-scheme support
The \src/pretokenize\ module introduces a new, high-performance pretokenization engine that splits documents into tokens using optimized, scheme-specific regex scanners. It supports nine distinct pretokenization schemes—including GPT-2 (r50k), GPT-4 (cl100k), Qwen2/3.5, DeepSeek V3, Olmo3, O200k, Nemotron, and Kimi—selected via a runtime dispatch mechanism. The production implementations in the \fast\ submodule use SIMD-accelerated byte scanning (NEON on aarch64, SSE4.2/CRC32C on x86\_64) and packed Unicode class tables for speed, while legacy designs (state machine, combinator, AVX-512, portable SIMD) are retained in the \reference\ submodule strictly as benchmark baselines and test oracles.
src/pretokenize · high confidence
New profiling infrastructure for cold-encode optimization
Added a dedicated profiling suite in the \profiling/\ directory to support a multi-round, profile-driven optimization campaign for the single-threaded and multi-threaded cold encode paths. This includes \profile.sh\ for recording traces via samply (Firefox-profiler JSON) and xctrace (PMU CPU Counters), \analyze.py\ and \analyze\_mt.py\ for symbolication, inline-frame resolution, and per-thread phase breakdown, \pmu\_summary.py\ for PMU bottleneck aggregation, and detailed campaign reports (\campaign\_report.md\, \report.md\, \mt\_round3\_findings.md\) documenting the methodology, findings, and results of the optimization rounds.
profiling · high confidence
Parallel batch encoding engine with safe oversized-document splitting
The library now includes a parallel batch encoding engine (src/batch.rs) that groups documents into coarse chunks for multi-core processing via Rayon. It introduces safe splitting for oversized documents at pretoken-safe boundaries (for BPE) or scanner-safe unit boundaries (for SentencePiece), ensuring that a single huge document is encoded across all cores with token-identical output. The engine uses pooled workers with persistent pretoken caches, supports descending chunk sizes for better load balancing (controlled by the GIGATOK\_NO\_LPT environment variable), and handles various input formats including JSONL, separator-delimited text, and Parquet. Small batches are encoded serially to avoid thread overhead, and output buffers are pre-sized and optionally backed by huge pages for performance.
src · high confidence
Behavioural changes
Native Rust HuggingFace tokenizer loading replaces Python dependency
The \src/load\_tokenizer\ module now loads HuggingFace tokenizer configurations and downloads model files directly in Rust, eliminating the previous reliance on Python libraries like \huggingface\_hub\ and \tokenizers\. This change introduces a native Hub client (\hub.rs\) that mirrors standard HuggingFace cache resolution and token discovery, alongside a comprehensive parser (\hf.rs\) for \tokenizer.json\ schemas (supporting SentencePiece BPE, ByteLevel BPE, normalizers, and pretokenizers) and a dedicated loader for tiktoken-style rank files (\tiktoken.rs\). Users benefit from faster startup times, reduced binary size, and the ability to load a wider variety of tokenizer architectures without requiring a Python environment.
_src/load\tokenizer · high confidence
Replace Rust examples with Python examples
The repository's example directory has been updated to provide Python-based demonstrations instead of previous Rust implementations. New scripts include quickstart.py for basic tokenization, encode\_files.py for high-performance file encoding using Rust-backed parallel processing, and drop-in compatibility examples (hf\_drop\_in.py and tiktoken\_drop\_in.py) showing how to integrate the gigatoken library with HuggingFace transformers and tiktoken APIs.
examples · high confidence
Rust-native Python bindings with padded encoding, file sources, and cache control
The Python bindings have been rewritten in Rust to improve performance and expand capabilities. Users can now use \encode\_batch\_padded\ to receive rectangular token matrices with configurable padding, truncation, and special token prefixes/suffixes. New \FileSource\ classes (\TextFileSource\, \JsonlFileSource\, \ParquetFileSource\) and \BytesSource\ allow efficient encoding of files and in-memory buffers with optional separator-based splitting. A global cache budget can be set via \set\_max\_cache\_bytes\ to limit memory usage, and HuggingFace Hub integration is now handled directly in Rust for faster model loading.
src/bindings · high confidence
Test coverage
Added Zen 5 single-threaded encode profiling data and reproduction script; Added scripts for training and comparing BPE tokenizers; Added test suite and benchmarking infrastructure for Gigatoken; New gigatoken encoding and pretokenization benchmarks.
Dependencies
Initial release of Gigatoken Rust tokenizer library
This change introduces the core build configuration and dependency manifest for the Gigatoken project, establishing it as a Rust library with Python bindings via PyO3. The Cargo.toml defines the \gigatoken\_rs\ crate (version 0.10.0) and includes dependencies for high-performance tokenization such as \parquet\, \arrow\, \icu\, \simdutf\, and \pyo3\. The \pyproject.toml\ configures the Python package \gigatoken\, requiring Python 3.10+ and dependencies like \awkward\, \numpy\, and \typer\, while \Cargo.lock\ records the resolved versions for reproducible builds.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 65 → 61 (-3.7)
- Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 80 → 79 (-0.8)
- Architecture 100 → 100 (+0.0)
- Maturity 67 → 67 (+0.2)
- Readiness 53 → 45 (-8.2)
- Security 75 → 82 (+7.1)
- Performance 85 (new)
Resolved (14)
- HFCompat.decode (cognitive 25) (gigatoken/_hf_compat.py)
- HFCompat.decode (cyclomatic 16) (gigatoken/_hf_compat.py)
- Hotspot: src/pretokenize/fast/cl100k.rs (src/pretokenize/fast/cl100k.rs)
- Hotspot: src/pretokenize/fast/deepseek_v3.rs (src/pretokenize/fast/deepseek_v3.rs)
- Hotspot: src/pretokenize/fast/o200k_family.rs (src/pretokenize/fast/o200k_family.rs)
- Hotspot: src/pretokenize/fast/olmo3.rs (src/pretokenize/fast/olmo3.rs)
- Hotspot: src/pretokenize/fast/qwen2.rs (src/pretokenize/fast/qwen2.rs)
- Hotspot: src/pretokenize/fast/qwen3_5.rs (src/pretokenize/fast/qwen3_5.rs)
- Hotspot: src/pretokenize/reference/combinator.rs (src/pretokenize/reference/combinator.rs)
- Hotspot: src/pretokenize/reference/simd.rs (src/pretokenize/reference/simd.rs)
- Hotspot: src/pretokenize/reference/state_machine.rs (src/pretokenize/reference/state_machine.rs)
- Members sharing a duplicated core (4 members, 50+ identical tokens) (src/bpe/mod.rs)
- Members sharing a duplicated core (4 members, 50+ identical tokens) (src/pretokenize/fast/cl100k_family.rs)
- Members sharing a duplicated core (6 members, 50+ identical tokens) (src/pretokenize/fast/cl100k.rs)
New (53)
- Duplicated block (10 lines × 5) (notebooks/bit_patterns.py)
- Duplicated block (10–12 lines × 2) (profiling/analyze.py)
- Duplicated block (110–114 lines × 2) (notebooks/bit_patterns.py)
- Duplicated block (11–12 lines × 2) (profiling/analyze.py)
- Duplicated block (11–12 lines × 2) (profiling/analyze.py)
- Duplicated block (11–13 lines × 3) (notebooks/dfa_states.py)
- Duplicated block (12–13 lines × 2) (profiling/analyze.py)
- Duplicated block (12–14 lines × 2) (profiling/analyze.py)
- Duplicated block (12–16 lines × 3) (notebooks/bit_patterns.py)
- Duplicated block (13–14 lines × 3) (notebooks/dfa_states.py)
- Duplicated block (13–15 lines × 2) (notebooks/simd_unicode.py)
- Duplicated block (14–15 lines × 2) (notebooks/dfa_states.py)
- Duplicated block (16–17 lines × 2) (notebooks/bit_patterns.py)
- Duplicated block (17 lines × 2) (notebooks/simd_unicode.py)
- Duplicated block (18–28 lines × 3) (notebooks/bit_patterns.py)
- Duplicated block (19–32 lines × 3) (notebooks/bit_patterns.py)
- Duplicated block (21 lines × 2) (notebooks/bit_patterns.py)
- Duplicated block (25–27 lines × 3) (notebooks/dfa_states.py)
- Duplicated block (27 lines × 2) (scripts/build_release_cross_platform.py)
- Duplicated block (29–30 lines × 2) (profiling/analyze.py)
- …and 33 more
Architecture
- Unchanged — 0 containers · 1 contexts · 0 edges
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
marcelroed/gigatoken was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit fac0114b37120ec8a76362e9ee8e1c742aaafaef — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.