Skip to content
CAI
Software that uses CAICheck a score

QwenLM/Qwen

49.9

Weak · 19 September 2026

6k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is an open-source framework for deploying, fine-tuning, and evaluating the Qwen family of large language models. It provides tools for inference across various hardware accelerators, including NVIDIA GPUs, Ascend NPUs, and Hygon DCUs, alongside scripts for full, LoRA, and QLoRA fine-tuning. The codebase also includes evaluation benchmarks, function-calling examples, and containerized deployment options to facilitate model adaptation and integration.

Features

Add chat-mode evaluation scripts and update existing benchmarks for Qwen-7B-Chat

The eval directory now includes new zero-shot evaluation scripts for the chat model (Qwen-7B-Chat) across C-Eval, MMLU, GSM8K, and HumanEval, alongside a new script for tool-use plugin evaluation. Existing non-chat evaluation scripts (C-Eval, GSM8K, HumanEval) have been updated to support batched inference, proper attention masks, and explicit padding token configuration, while the documentation has been updated to reflect the new chat-specific reproduction commands.

eval · high confidence

Add inference support for Qwen models on Ascend 910 and Hygon DCU hardware

Users can now run Qwen models on Ascend 910 and Hygon DCU accelerators. The \ascend-support\ directory provides Docker-based inference using MindSpore and MindFormers, including scripts to convert and run the Qwen-7B-Chat model. The \dcu-support\ directory adds Hygon DCU compatibility via the fastllm framework, offering CLI and web-based demos, along with tools to convert Hugging Face Qwen models into the fastllm format for optimized inference.

ascend-support, dcu-support · high confidence

Added finetuning scripts and DeepSpeed configurations

The finetune directory now includes shell scripts and configuration files to support various training modes. Users can run distributed training with DeepSpeed ZeRO-2 or ZeRO-3 using the new \ds\_config\_zero2.json\ and \ds\_config\_zero3.json\ files, or use single-GPU scripts. The provided scripts cover full finetuning, LoRA finetuning, and QLoRA finetuning (using 4-bit quantized models), with gradient checkpointing enabled to manage memory usage.

finetune · high confidence

Initial Docker support for Qwen-Chat

This change introduces Docker support for running Qwen-Chat, providing three CUDA-specific Dockerfiles (CUDA 11.4, 11.7, and 12.1) and helper shell scripts to easily launch the CLI demo, web demo, and OpenAI-compatible API in containers. The Dockerfiles bundle necessary dependencies like PyTorch, Flash Attention, and fine-tuning libraries, while the scripts handle image pulling, checkpoint mounting, and container startup with appropriate GPU and port configurations.

docker · high confidence

Initial release of Qwen models and documentation

This commit introduces the initial codebase and documentation for the Qwen series of large language models, including Qwen-1.8B, Qwen-7B, Qwen-14B, and Qwen-72B, as well as their chat-aligned counterparts (Qwen-Chat). It provides the necessary Python scripts for inference (\cli\_demo.py\, \web\_demo.py\), fine-tuning (\finetune.py\), and an OpenAI-compatible API server (\openai\_api.py\). The release includes comprehensive documentation in English, Chinese, Japanese, French, and Spanish, covering installation, usage, performance benchmarks, and licensing, along with an FAQ to assist users with common setup issues such as Flash Attention installation and tokenizer configuration.

(repo-wide) · high confidence

Initial release of Qwen recipe notebooks and Ascend NPU fine-tuning guide

This update introduces a comprehensive set of Jupyter notebook recipes for the Qwen model family, providing step-by-step guides for various use cases. The collection includes application demos such as a web-based chatbot using Gradio and a retrieval-augmented Q&A system with LangChain. It also covers fine-tuning workflows, offering tutorials for full-parameter, LoRA, and QLoRA adaptation on both single and multiple GPUs using DeepSpeed, as well as a specific guide for fine-tuning on Ascend NPUs.

recipes · high confidence

New examples for function calling, code auto-commenting, and vocabulary expansion

The examples directory now includes several new scripts and documentation. function\_call\_examples.py and function\_call\_finetune\_examples.py demonstrate how to use the model for tool use (function calling) via both an OpenAI-compatible API and ReAct-style prompting, including guidance for fine-tuning. auto\_comments.py and its documentation show how to automatically generate Chinese comments for Python code files or folders. add\_merges.py provides a utility to expand the tokenizer vocabulary by learning new BPE merges from a frequency list. Additionally, system\_prompt.md explains how to use system prompts for role-playing and behavior customization, and langchain\_tooluse.ipnb shows how to integrate LangChain tools like Google Search and WolframAlpha using ReAct prompting.

examples · high confidence

Dependencies

Initial dependency manifests for core, web demo, DCU support, and DeepSpeed finetuning

Added four new requirements files to define the project's Python dependencies: the main \requirements.txt\ pins \transformers\ (\>=4.32.0,\<4.38.0) and \transformers\_stream\_generator\ (0.0.4) alongside \accelerate\, \tiktoken\, \einops\, and \scipy\; \requirements\_web\_demo.txt\ introduces \gradio\ (\<3.42) and \mdtex2html\ for the web interface; \dcu-support/requirements.txt\ specifies dependencies for DCU hardware support including \streamlit\ (\>=1.24.0) and \sentencepiece\; and \recipes/finetune/deepspeed/requirements.txt\ adds \deepspeed\ and \peft\ for finetuning workflows.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 50.

Lenses

  • Code Health 82
  • Architecture 100
  • Maturity 52
  • Readiness 30
  • Security 81

Changes since last survey

  • 300 commits — 288 feature/other, 12 fixes

By area

  • (root) — 163 commits
  • (repo) — 69 commits
  • assets/wechat.png — 37 commits
  • .github/workflows — 3 commits
  • assets/radar_14b.jpeg — 2 commits
  • docker/Dockerfile — 2 commits
  • eval/EVALUATION.md — 2 commits
  • examples/react_demo.py — 2 commits
  • finetune/finetune_lora_ds.sh — 2 commits
  • finetune/finetune_lora_single_gpu.sh — 2 commits
  • recipes/finetune — 2 commits
  • assets/cli_demo.gif — 1 commit
  • assets/code_interpreter_showcase_001.jpg — 1 commit
  • assets/logo.jpg — 1 commit
  • assets/radar_14b.jpg — 1 commit
  • docker/Dockerfile-cu121 — 1 commit
  • eval/evaluate_ceval.py — 1 commit
  • eval/evaluate_chat_ceval.py — 1 commit
  • eval/evaluate_plugin.py — 1 commit
  • examples/function_call_finetune_examples.py — 1 commit

Notable commits

  • fix: Fix Dockerfile for <https://github.com/QwenLM/Qwen/issues/783>.
  • fix: Fix bug of fschat version in Dockerfile-cu121.
  • fix: Fix bug of low_cpu_mem_usage in finetune.py.
  • fix: Fix peft version in dockerfiles.
  • fix: Merge pull request #881 from QwenLM/fix-dockerfile-peft
  • fix: Merge pull request #934 from QwenLM/fix-docker-cu121
  • fix: Merge pull request #964 from QwenLM/fix-finetune
  • fix: bugfix streaming mode of openai_api.py
  • fix: fix badcase for 14b
  • fix: fix badcase in doc
  • fix: fix badcase in react_demo.py
  • fix: fix single-gpu qlora, and add profiling
  • change: Add Docker image for CUDA-12.1.
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • change: Add files via upload
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

QwenLM/Qwen was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 2df8e8ac450fa185c421a08b0090ef81826caa6e — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.