Skip to content
CAI
Software that uses CAICheck a score

huggingface/open-r1

66.2

Adequate · 18 September 2026

8.3k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system provides the infrastructure for training and evaluating large language models using the Open R1 pipeline, specifically focusing on reproducing the DeepSeek-R1 methodology. It supports supervised fine-tuning and GRPO training with custom reward functions, including specialized scoring for competitive programming problems. The platform integrates external code execution environments like E2B and MorphCloud to validate code submissions and manages data preparation through decontamination and pass-rate filtering utilities.

Features

Add dataset pass-rate filtering script

A new script and supporting files have been added to the \scripts/pass\_rate\_filtering\ directory to enable filtering datasets based on model pass rates. The \compute\_pass\_rate.py\ script loads a dataset, generates completions using a specified LLM, computes rewards, and filters examples based on configurable minimum and maximum pass-rate thresholds. A shell script (\launch\_filtering.sh\) is also provided to orchestrate this process in chunks, and a README documents the usage for merging and pushing filtered datasets to the Hugging Face Hub.

_scripts/pass\_rate\filtering · high confidence

Initial project scaffolding with Open R1 training and evaluation infrastructure

This change introduces the foundational structure for the Open R1 project, including the core Python package setup via setup.py and setup.cfg, a Makefile for managing development tasks (style, quality, tests) and running evaluations, and a comprehensive README.md detailing the project's goal to reproduce the DeepSeek-R1 pipeline. The setup defines dependencies for training (TRL, vLLM, DeepSpeed, Transformers) and evaluation (LightEval, math-verify), while the Makefile provides commands to install the environment and execute SFT and GRPO training scripts as well as evaluation tasks.

(repo-wide) · high confidence

Introduce open\_r1 training scripts and reward functions

Adds the initial \open\_r1\ package containing GRPO and SFT training scripts, configuration dataclasses, and a registry of reward functions (including accuracy, format, length, and reasoning steps rewards) for training language models.

_src/open\r1 · high confidence

New unified Slurm infrastructure for serving, training, and evaluation

This change introduces a comprehensive set of new Slurm scripts and documentation to streamline the deployment and execution of the OpenR1 workflow. Users can now launch DeepSeek-R1 inference servers using either SGLang (via \serve\_r1.slurm\ and \serve\_router.slurm\) or vLLM (via \experimental/serve\_r1\_vllm.slurm\) on multi-node GPU clusters. The suite includes \train.slurm\ to orchestrate SFT and GRPO training jobs with configurable data/tensor parallelism and vLLM integration, \evaluate.slurm\ to run LightEval benchmarks and automatically upload results to the Hugging Face Hub, and \generate.slurm\ for scalable data generation. Additionally, new scripts support code execution via Piston workers (\piston/launch\_piston\_workers.sh\) and an E2B router (\e2b\_router.slurm\), providing a complete operational environment for model development and testing.

slurm · high confidence

New utility module for training infrastructure and code execution

This change introduces the \src/open\_r1/utils\ package, providing core infrastructure for the training workflow. It adds a dataset mixer to combine multiple datasets with configurable weights and splits, and implements code execution providers for E2B and MorphCloud to support reward-based evaluation. The package also includes utilities for asynchronous model pushing to Hub revisions, automated benchmark evaluation via LightEval, and Weights & Biases logging configuration.

_src/open\r1/utils · high confidence

New utility scripts for dataset decontamination, code execution routing, and benchmarking

This update introduces several new scripts to the \scripts/\ directory to support data preparation and evaluation workflows. \decontaminate.py\ allows users to check datasets for n-gram overlap against benchmark datasets (AIME, MATH-500, GPQA, LiveCodeBench) and optionally remove contaminated rows. \e2b\_router.py\ and \morph\_router.py\ provide FastAPI-based batch execution servers for running code in isolated sandboxes using E2B and MorphCloud respectively. \benchmark\_e2b.py\ benchmarks the \code\_reward\ function's performance with varying parallelization levels, while \run\_benchmarks.py\ serves as a CLI entry point to trigger benchmark jobs after training. Additional utilities include \generate\_reasoning.py\ for asynchronous LLM response generation, \get\_tensor\_parallel\_size.py\ for calculating optimal tensor parallelism, and \upload\_details.py\ for pushing LightEval results to the Hub.

scripts · high confidence

Support for Codeforces and IOI competitive programming problem scoring

Added a new utility module for evaluating code submissions against competitive programming problems from Codeforces and the International Olympiad in Informatics (IOI). The module includes scoring logic for Codeforces (supporting pass/fail, partial, and weighted sum modes) and IOI (supporting subtask-based scoring with detailed status reporting). It also provides code patching utilities to fix common Python 3 and C++ compilation issues (such as deprecated imports and missing headers) and integrates with Piston and MorphCloud execution environments to run and score submissions asynchronously.

_src/open\_r1/utils/competitive\programming · high confidence

Test coverage

Added slow integration tests for code reward execution via E2B and Morph Cloud routers; Added test coverage for reward functions; Added tests for dataset loading and mixture configuration.

Housekeeping

Added logs directory placeholder

A new .gitkeep file has been added to the logs directory to ensure the folder is tracked by version control.

logs · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 66.

Lenses

  • Code Health 95
  • Architecture 99
  • Maturity 54
  • Readiness 65
  • Security 84

Changes since last survey

  • 199 commits — 145 feature/other, 54 fixes

By area

  • (root) — 85 commits
  • src/open_r1 — 74 commits
  • slurm/evaluate.slurm — 8 commits
  • recipes/Qwen2.5-1.5B-Instruct — 5 commits
  • .github/workflows — 3 commits
  • recipes/DeepSeek-R1-Distill-Qwen-7B — 2 commits
  • recipes/Mistral-Small-24B-Instruct-2501 — 2 commits
  • scripts/generate_reasoning.py — 2 commits
  • slurm/generate.slurm — 2 commits
  • slurm/train.slurm — 2 commits
  • .github/dependabot.yml — 1 commit
  • recipes/OpenR1-Qwen-7B — 1 commit
  • recipes/SmolLM2-1.7B — 1 commit
  • recipes/accelerate_configs — 1 commit
  • recipes/dataset_filtering — 1 commit
  • recipes/qwen — 1 commit
  • scripts/datasets — 1 commit
  • scripts/decontaminate.py — 1 commit
  • scripts/generate_reason_data.py — 1 commit
  • scripts/upload_details.py — 1 commit

Notable commits

  • fix: Add GPQA Diamond and fix evaluation deps (#196)
  • fix: Async code reward fixes (#546)
  • fix: Bump DeepSpeed to 0.16.8 to fix OOM on Qwen3 (#653)
  • fix: E2B Router bug fixes (#592)
  • fix: Fix AIME25 naming (#270)
  • fix: Fix LightEval commands and dependencies (#386)
  • fix: Fix README: Correct recipes path and missing --config option (#247)
  • fix: Fix SFT for base models (#604)
  • fix: Fix Slurm SFT and gather Slurm scripts (#19)
  • fix: Fix TP once and for all :) (#613)
  • fix: Fix Weka refresh (#666)
  • fix: Fix cosine_scaled_reward compatibility with GRPO (#229)
  • fix: Fix generate.slurm (#10)
  • fix: Fix accuracy reward for math (#566)
  • fix: Fix code quality after adding puzzles (#178)
  • fix: Fix configs
  • fix: Fix eval comamnds (#18)
  • fix: Fix eval system prompt (#591)
  • fix: Fix help text for --retries to match actual default value (#103)
  • fix: Fix len reward (#385)
  • …and 179 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/open-r1 was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 1416fa0cf21595d2083b399a2a0bbddd7f6e9563 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.