Skip to content
CAI
Software that uses CAICheck a score

deepseek-ai/DeepSeek-Coder

39.8

Weak · 18 September 2026

5k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a framework for evaluating and fine-tuning large language models, specifically focusing on code generation capabilities. It provides scripts and utilities to benchmark model performance on standard coding datasets like HumanEval, MBPP, and LeetCode, supporting multiple programming languages. Additionally, it includes tools for distributed fine-tuning using DeepSpeed and an interactive demo interface for testing instruction-tuned models.

Features

Add DeepSeek-6.7B-Chat demo application

A new interactive chat demo for the DeepSeek-Coder 6.7B model has been added to the demo directory. The application uses Gradio to provide a user interface for chatting with the model, featuring adjustable parameters for system prompts, maximum new tokens, top-p, top-k, and repetition penalty. The demo includes specific styling via style.css and lists required Python dependencies in requirement.txt, including Gradio 3.48.0, Transformers \>=4.35.0, and PyTorch 2.0.0.

demo · high confidence

Add DeepSpeed-supported fine-tuning script and configuration for DeepSeek-Coder

Users can now fine-tune the deepseek-ai/deepseek-coder-6.7b-instruct model on custom instruction-output datasets using the new \finetune\_deepseekcoder.py\ script. This entry point integrates with DeepSpeed (via the provided \ds\_config\_zero3.json\ configuration) to enable efficient distributed training with ZeRO-3 optimization, bfloat16 support, and gradient checkpointing, allowing users to adapt the model to specific coding tasks.

finetune · high confidence

Add instruction-tuned MBPP evaluation script

A new evaluation script (eval\_instruct.py) has been added to the MBPP evaluation suite to support instruction-tuned models. This script loads the MBPP dataset, formats prompts with three-shot examples, and generates Python code using Hugging Face Transformers with chat templates and specific stop tokens. It extracts generated code blocks and runs functional correctness checks against the test suite. The README has also been updated to reflect the new evaluation method and remove outdated benchmark results for the 5.7B model.

Evaluation/MBPP · high confidence

Added LeetCode evaluation dataset for Weekly Contest 381 and Biweekly Contest 122

A new evaluation dataset file (\20240121-Jul-zh.jsonl\) has been added to the \Evaluation/LeetCode\ directory, containing problems from LeetCode Weekly Contest 381 and Biweekly Contest 122. This data includes problem statements, prompts, and test cases for four specific algorithmic challenges: 'Minimum Number of Pushes to Type Word I', 'Count the Number of Houses at a Certain Distance I', 'Minimum Number of Pushes to Type Word II', and 'Count the Number of Houses at a Certain Distance II', along with 'Divide an Array Into Subarrays With Minimum Cost I'. This enables the evaluation system to assess model performance on these specific recent contest problems.

Evaluation/LeetCode · high confidence

Added instruction-based model evaluation script and documentation

Users can now evaluate instruction-tuned models on the HumanEval benchmark using the new \eval\_instruct.py\ script, which handles prompt construction, code generation, and functional correctness evaluation. The \README.md\ has been updated with usage instructions for this new evaluation mode and revised experimental results tables to reflect the latest model performance data.

Evaluation/HumanEval · high confidence

Added language-specific code extraction utilities for HumanEval evaluation

The evaluation utility module now includes configuration for multiple programming languages (Python, C++, Java, C\#, PHP, TypeScript, JavaScript, Bash) and new helper functions to extract generated code blocks. The \extract\_generation\_code\ function parses model outputs to isolate the relevant function body, handling language-specific syntax like main methods and indentation, while \get\_function\_name\ identifies function signatures. This enhances the evaluation pipeline's ability to accurately process and compare code generations across different languages.

Evaluation/HumanEval/utils · high confidence

Behavioural changes

2 commits (0 fixes) modifying Evaluation/HumanEval/human\_eval

A change to existing behaviour in Evaluation/HumanEval/human\_eval — 2 commits, 8 files.

Evaluation/HumanEval/\\pycache\\_, Evaluation/HumanEval/human\eval · medium confidence · unverified

Updated evaluation results and fixed output file path in PAL-Math

The PAL-Math evaluation module now reports updated benchmark scores for various models (such as CodeShell, CodeLLama, and DeepSeek-Coder) across datasets like GSM8k and MATH, reflecting changes in performance metrics. Additionally, the run script has been corrected to use the loop variable 'rank' instead of 'args.rank' when constructing the output file path, ensuring the correct distributed evaluation results are loaded.

Evaluation/PAL-Math · high confidence

Dependencies

Add dependency manifests for finetune and base environments

New requirements.txt files have been added for the finetune directory and the project root, establishing pinned or minimum versions for core machine learning libraries. Key dependencies include PyTorch (2.0.1 or \>=2.0), Hugging Face Transformers (4.35.0), and tokenizers (0.14.0 or \>=0.14.0). The finetune environment specifically adds DeepSpeed (0.12.2) for distributed training, along with accelerate, datasets, and tensorboardX, while the base environment includes sympy, pebble, and timeout-decorator.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 40.

Lenses

  • Code Health 81
  • Architecture 99
  • Maturity 37
  • Readiness 17
  • Security 80

Changes since last survey

  • 109 commits — 104 feature/other, 5 fixes

By area

  • (root) — 59 commits
  • (repo) — 17 commits
  • Evaluation/HumanEval — 9 commits
  • Evaluation/MBPP — 4 commits
  • demo/app.py — 3 commits
  • pictures/Math.png — 3 commits
  • Evaluation/LeetCode — 2 commits
  • Evaluation/PAL-Math — 2 commits
  • demo/requirement.txt — 2 commits
  • finetune/requirements.txt — 2 commits
  • demo/style.css — 1 commit
  • finetune/README.md — 1 commit
  • finetune/finetune_deepseekcoder.py — 1 commit
  • pictures/HumanEval.png — 1 commit
  • pictures/logo.png — 1 commit
  • pictures/table.png — 1 commit

Notable commits

  • fix: Fix the position of add_generation_prompt
  • fix: fix add_generation_prompt in latest version
  • fix: fix in-page link for detailed eval results
  • fix: fix test file
  • fix: fix: update discord invitation link
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add supported programming languages into README.md
  • change: Complete missing import
  • change: Create app.py
  • change: Create requirement.txt
  • change: Create requirements.txt
  • change: Create style.css
  • change: Delete pictures/Math.png
  • change: Fix typos and grammatical errors in README.md
  • change: Merge pull request #105 from deepseek-ai/ydj/leetcode
  • change: Merge pull request #109 from eltociear/patch-1
  • …and 89 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

deepseek-ai/DeepSeek-Coder was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 2f9fd85927c669dae3c0fbb2d607274023af243e — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.