karpathy/nanochat
51.3
Weak · 26 September 2026
5.9k
lines of production code
Python
primary language
4
measurements over time
What this system is
Features
Add inference benchmarking script
A new \scripts/infer\_bench.py\ script is introduced to measure inference latency, throughput, memory and bandwidth utilization of a trained checkpoint. It sweeps over the decode batch size to trace the latency-throughput tradeoff between the compute-bound prefill and memory-bandwidth-bound decode regimes, reporting metrics like TTFT, MFU, and MBU.
scripts · high confidence
Unified Flash Attention 3 interface with SDPA fallback
The codebase now includes a unified Flash Attention 3 interface that automatically falls back to PyTorch's SDPA on incompatible CUDA GPUs, MPS, and CPU. This new module, \nanochat/flash\_attention.py\, provides a drop-in replacement for the previous Flash Attention 3 implementation, supporting both training (no KV cache) and inference (with KV cache) modes. The implementation detects available hardware and selects the appropriate backend, ensuring compatibility across different GPU architectures and non-CUDA devices.
nanochat · high confidence
Removals
Removed embedded rustbpe source code
The embedded rustbpe Rust source code (lib.rs) and its README have been deleted from the repository. This change reflects the decision to treat rustbpe as an external dependency rather than maintaining the code inline, aligning with the project's move to host rustbpe as a separate repository.
rustbpe · high confidence
Behavioural changes
Replace Hugging Face datasets library with a lightweight parquet-based loader
The tasks module no longer depends on the external \datasets\ library. Instead, it uses a new \load\_hub\_dataset\ function in \tasks/common.py\ that downloads and reads Hugging Face datasets directly from their auto-generated Parquet shards using \pyarrow\. This change also updates the MMLU task to use \auxiliary\_train\ as a split name rather than a subset, and corrects the MMLU subset assertion to only allow \all\ while moving \auxiliary\_train\ to the split argument.
tasks · high confidence
Restructure run scripts and update speedrun configuration
The \speedrun.sh\ script has been moved to the \runs/\ directory and updated to train a deeper (d24) model with a lower data-to-parameter ratio (8 instead of 10.5) to improve training efficiency. The script now uses \--fp8\ precision and a device batch size of 16. The previous midtraining phase has been replaced with Supervised Fine-Tuning (SFT). Additionally, new scripts \miniseries.sh\ and \scaling\_laws.sh\ have been added to the \runs/\ directory to support automated training series and scaling law analysis, while the \runcpu.sh\ script has been added for CPU-based educational runs.
runs · high confidence
Switch pretraining dataset from FineWeb-EDU to NVIDIA ClimbMix
The pretraining dataset has been upgraded from FineWeb-EDU to NVIDIA ClimbMix, a 400B-token mixture of high-quality web text, code, and math. This change reduces the time to reach GPT-2 capability from approximately 2 hours 46 minutes to 2 hours 1 minute. The update includes a migration warning in the data loading logic to guide users to the new dataset, and the reference script for data preparation now supports both the legacy FineWeb-EDU and the new ClimbMix configurations.
dev · high confidence
Test coverage
Added comprehensive test coverage for core components
Added new test files to verify the correctness of the Flash Attention 3 and SDPA fallback implementations, the Engine class, the MuonAdamW optimizer, the Task/HubDataset machinery, the RustBPE tokenizer, and the sandboxed code execution environment. These tests ensure that attention mechanisms produce identical results across backends, that multi-sample generation samples tokens independently, that the optimizer matches PyTorch's AdamW behavior and converges, that task slicing and shuffling are deterministic, that the tokenizer handles special tokens and conversation rendering correctly, and that the execution sandbox properly restricts dangerous operations.
tests · high confidence
Dependencies
Streamlined dependencies and added CPU/GPU installation options
The project has removed the inline Rust-based \rustbpe\ crate and its associated \Cargo.lock\ and \Cargo.toml\ files, replacing them with an external \rustbpe\ PyPI dependency. Additionally, the \pyproject.toml\ has been updated to support both CPU and GPU installations via optional dependencies (\cpu\ and \gpu\), allowing users to install PyTorch for either platform. Unnecessary dependencies like \datasets\, \fastapi\, \uvicorn\, and \tokenizers\ have been removed, while \pyarrow\ and \filelock\ have been added.
(dependencies) · medium confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 46 → 51 (+5.3)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 97 → 94 (-3.6)
- Architecture 94 → 100 (+5.9)
- Maturity 57 → 62 (+5.5)
- Readiness 17 → 25 (+8.3)
- Security 81 → 79 (-1.2)
Resolved (10)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (6 lines × 2) (tests/test_attention_fallback.py)
- Low: security finding (details withheld)
- Low: security finding (details withheld)
- Medium: security finding (details withheld)
- Medium: security finding (details withheld)
- No exposed public API
- No tests found
- Test reliability not included
New (49)
- Critical CVE: [GHSA redacted] (uv.lock)
- Duplicated block (6–8 lines × 2) (nanochat/flash_attention.py)
- Duplicated block (7 lines × 2) (scripts/chat_eval.py)
- Engine.generate (cognitive 33) (nanochat/engine.py)
- Engine.generate (cyclomatic 18) (nanochat/engine.py)
- Further sole-owners (lower concentration)
- HackComment (nanochat/checkpoint_manager.py)
- High CVE: [GHSA redacted] (uv.lock)
- High CVE: [GHSA redacted] (uv.lock)
- High CVE: [GHSA redacted] (uv.lock)
- High CVE: [GHSA redacted] (uv.lock)
- High CVE: [GHSA redacted] (uv.lock)
- Hotspot: scripts/chat_sft.py (scripts/chat_sft.py)
- Low: security finding (details withheld)
- Low: security finding (details withheld)
- Medium CVE: [GHSA redacted] (uv.lock)
- Medium CVE: [GHSA redacted] (uv.lock)
- Medium CVE: [GHSA redacted] (uv.lock)
- Medium CVE: [GHSA redacted] (uv.lock)
- Medium CVE: [GHSA redacted] (uv.lock)
- …and 29 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
karpathy/nanochat was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 92d63d4e8bb4df75c3b71618f31ddde2378b2bcd — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.