Skip to content
CAI
Software that uses CAICheck a score

lm-sys/FastChat

60.5

Adequate · 26 September 2026

57.5k

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Features

Add vision example generation scripts and image processing utilities

Added new Python scripts to generate and manage VQA (Visual Question Answering) example datasets, including downloading images from various sources (Memes, Floorplan, Website, IllusionVQA, NewYorker) and creating JSON metadata files. Also introduced an Image utility class that handles image format conversion, base64 encoding, URL fetching, and resizing to support vision-based interactions.

fastchat/serve/vision · high confidence

Added Nginx gateway configuration for HTTPS and load balancing

Added Nginx configuration files (nginx.conf and README.md) that set up an Nginx reverse proxy to protect Gradio servers, enable HTTPS, and provide load balancing and connection limiting.

fastchat/serve/gateway · high confidence

Added tools to analyze and visualize voting latency statistics

A new \vote\_time\_stats\ directory has been added to the FastChat monitor, containing scripts (\analyze\_data.py\ and \plot.py\) and documentation (\README.md\) to process server logs and generate a distribution plot of the time between chat responses and user votes. The \analyze\_data.py\ script parses JSONL logs to extract vote timestamps and corresponding chat finish times, while \plot.py\ calculates the duration and renders a histogram of the 'time to vote' distribution.

_fastchat/serve/monitor/vote\_time\stats · high confidence

Introduce category classification and evaluation tools for the Chatbot Arena monitor

Added a new \fastchat/serve/monitor/classify\ module that provides scripts and configuration files to classify user prompts into specific categories (such as creative writing, math, or instruction-following) and evaluate the performance of category classifiers against ground truth benchmarks. This includes the \category.py\ file defining various category classes, \label.py\ for generating classification labels via LLMs, \display\_score.py\ for reporting accuracy/precision/recall, and associated YAML configs for both text and vision-based categories.

fastchat/serve/monitor/classify · high confidence

New benchmarking and DeepSpeed configuration files for the playground

Added a new Python script, playground/benchmark/benchmark\_api\_provider.py, which provides a tool for benchmarking API providers by measuring time-to-first-token and average token generation time. Additionally, two new DeepSpeed configuration files were introduced: playground/deepspeed\_config\_s2.json (Stage 2) and playground/deepspeed\_config\_s3.json (Stage 3), which define memory optimization and mixed-precision settings for training.

playground · high confidence

New data cleaning and processing scripts

Added a suite of Python scripts in the \fastchat/data\ directory to clean, filter, and prepare conversation datasets for training. These include \clean\_sharegpt.py\ for HTML-to-Markdown conversion and basic cleaning, \filter\_wrong\_format.py\ to remove malformed entries, \optional\_clean.py\ for language filtering and deduplication, \split\_long\_conversation.py\ to truncate long conversations, and \split\_train\_test.py\ to partition data. The \prepare\_all.py\ script orchestrates this pipeline, and \get\_stats.py\ provides dataset statistics.

fastchat/data · high confidence

New data processing and monitoring scripts for the Chatbot Arena

Added a suite of new Python scripts in the \fastchat/serve/monitor\ directory to handle the cleaning, deduplication, and statistical analysis of chat and battle data. This includes \clean\_battle\_data.py\ and \clean\_chat\_data.py\ for processing raw logs, \add\_markdown\_info.py\ for extracting metadata, \deduplication.py\ for handling duplicate prompts, and \criteria\_labeling.py\ for automated evaluation. Additionally, new scripts were added under \dataset\_release\_scripts\ to filter, sample, and upload datasets (such as \lmsys\_chat\_1m\ and \arena\_33k\) to Hugging Face, alongside a new \copilot\_arena.py\ module to fetch and display the Copilot Arena leaderboard.

fastchat/serve/monitor · high confidence

New model utility scripts and adapters for specialized inference

Added new scripts to apply delta weights and LoRA adapters to base models, and introduced model adapters for CodeT5+, ExLlama, xFasterTransformer, and Yuan2.0, alongside a compression module for 8-bit quantization and a utility to convert models to FP16.

fastchat/model · high confidence

New modular serve architecture and OpenAI-compatible API protocols

The \fastchat/serve\ package has been restructured into a modular architecture, introducing a \base\_model\_worker.py\ that defines a common interface for all model workers, alongside dedicated workers like \dashinfer\_worker.py\. The codebase now includes comprehensive Pydantic protocol definitions in \api\_protocol.py\ and \openai\_api\_protocol.py\, enabling full compatibility with the OpenAI API specification (including chat completions, text completions, embeddings, and token checks). Additionally, the directory now contains specialized UI components for the Chatbot Arena (\gradio\_block\_arena\_anony.py\, \gradio\_block\_arena\_named.py\) and a new \api\_provider.py\ to handle streaming responses from various external API providers.

fastchat/serve · high confidence

New scripts for model serving, training, and deployment

Added several new shell scripts to streamline operations: build-api.sh automates spinning up the FastChat controller, workers, and API server in separate screen sessions; train\_lora.sh provides a ready-to-run configuration for LoRA fine-tuning with DeepSpeed; train\_vicuna\_7b.sh and train\_vicuna\_13b.sh offer optimized training setups for 7B and 13B models using FSDP; test\_readme\_train.sh serves as a minimal example for training; and upload\_pypi.sh simplifies building and publishing the package to PyPI.

scripts · high confidence

New training scripts and attention patches for LLaMA, Baichuan, Flan-T5, and Yuan2

The \fastchat/train\ directory now includes dedicated training scripts for LLaMA, Baichuan, Flan-T5, and Yuan2, alongside LoRA fine-tuning scripts for LLaMA and T5. The update also introduces new monkey-patch modules to enable Flash Attention and xformers for LLaMA, which optimize memory usage and training speed. These changes provide users with specialized training workflows and performance optimizations for a wider range of model architectures.

fastchat/train · high confidence

Release MT-bench evaluation suite and LLM judge tools

The \fastchat/llm\_judge\ directory is introduced, providing a complete toolkit for evaluating language models using the MT-bench benchmark. This includes the MT-bench question dataset, a collection of judge prompts for both pairwise and single-answer grading, and Python scripts to generate model answers, produce LLM-based judgments, compute agreement between judges, and clean judgment files. Users can now automate the evaluation of their models against MT-bench questions using strong LLMs as judges.

_fastchat/llm\judge · high confidence

Behavioural changes

FastChat v0.2.36 release with extensive conversation template updates

FastChat has been updated to version 0.2.36, introducing a comprehensive set of conversation prompt templates and supporting a wide variety of new models. The \conversation.py\ module now includes templates for Llama 2, Llama 3, ChatGLM, ChatGLM3, DeepSeek, Yuan 2.0, Gemma, and many others, alongside various bug fixes for existing templates like Mistral and Alpaca. The release also includes improvements to the Gradio web server, better error handling, and fixes for Windows logging issues.

fastchat · high confidence

Rebrand to FastChat with comprehensive documentation and styling

The project is renamed from ChatServer to FastChat, accompanied by a complete overhaul of the README.md to include installation instructions, model weight links, and usage examples. A new theme.json file is introduced to define the visual styling for the web interface, and a .pylintrc configuration file is added to enforce code quality standards.

(repo-wide) · high confidence

Test coverage

Add comprehensive test suite for FastChat services; Added embedding-based test scripts for classification, semantic search, and sentence similarity.

Dependencies

Introduce pyproject.toml and update dependencies

The project now uses a pyproject.toml file to manage dependencies and build configuration. This includes updating the Pydantic dependency to version 2.0.0 or higher (specifically pydantic\>=2.0.0, \<3), and adding new dependencies such as aiohttp, httpx, markdown2, nh3, psutil, and rich. Optional dependencies for model workers, web UI, training, and development are also defined, with specific versions for packages like transformers (\>=4.31.0) and gradio (\>=4.10).

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 46 → 61 (+14.9)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 89 → 79 (-10.2)
  • Architecture 96 → 96 (+0.2)
  • Maturity 55 → 61 (+6.0)
  • Readiness 19 → 48 (+29.1)
  • Security 78 → 82 (+3.5)

Resolved (17)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • LLM evaluation failed
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Low: security finding (details withheld)
  • Medium: security finding (details withheld)
  • Medium: security finding (details withheld)
  • No exposed public API
  • No tests found
  • Test reliability not included

New (246)

  • Change coupling: api_provider.py ↔ gradio_block_arena_anony.py (fastchat/serve/api_provider.py)
  • Change coupling: huggingface_api.py ↔ inference.py (fastchat/serve/huggingface_api.py)
  • Change coupling: model_chatglm.py ↔ inference.py (fastchat/model/model_chatglm.py)
  • Change coupling: model_chatglm.py ↔ model_worker.py (fastchat/model/model_chatglm.py)
  • Controller.get_worker_address (cognitive 28) (fastchat/serve/controller.py)
  • Conversation.get_prompt (cognitive 204) (fastchat/conversation.py)
  • Conversation.get_prompt (cyclomatic 91) (fastchat/conversation.py)
  • Conversation.to_openai_vision_api_messages (cognitive 19) (fastchat/conversation.py)
  • Conversation.to_reka_api_messages (cognitive 16) (fastchat/conversation.py)
  • Conversation.to_vertex_api_messages (cognitive 18) (fastchat/conversation.py)
  • DashInferWorker.generate_stream (cognitive 28) (fastchat/serve/dashinfer_worker.py)
  • DashInferWorker.generate_stream (cyclomatic 22) (fastchat/serve/dashinfer_worker.py)
  • Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
  • Documentation: no architecture or design documentation (docs/model_support.md)
  • Dormant codebase
  • Duplicated block (10 lines × 2) (fastchat/llm_judge/common.py)
  • Duplicated block (10 lines × 2) (fastchat/modules/awq.py)
  • Duplicated block (10 lines × 2) (fastchat/serve/cli.py)
  • Duplicated block (10 lines × 2) (fastchat/serve/gradio_block_arena_named.py)
  • Duplicated block (10 lines × 2) (fastchat/serve/mlx_worker.py)
  • …and 226 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

lm-sys/FastChat was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 587d5cfa1609a43d192cedb8441cac3c17db105d — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d0929f7ac71f.