hpcaitech/ColossalAI
49.5
Weak · 11 October 2026
203.6k
lines of production code
Python
primary language
3
measurements over time
What this system is
This system is a comprehensive distributed training and inference framework for large language models and diffusion models, built on PyTorch. It provides infrastructure for scaling model training through various parallelism strategies, including tensor, pipeline, data, and expert parallelism, alongside memory optimization techniques like ZeRO and automatic offloading. The system also supports post-training alignment via RLHF algorithms such as PPO, DPO, and GRPO, and includes tools for high-performance inference serving and model evaluation.
How it got here
2021–2022 — Auto-parallel infrastructure and diffusion examples
75 changes.
This period focused on establishing the foundational infrastructure for automatic parallelism, including FX graph tracing, cost profiling, and ILP-based solver strategies for tensor and pipeline sharding. Concurrently, the project expanded its example suite to demonstrate these capabilities on large-scale models like OPT and PaLM, while introducing comprehensive support for training and inference with Stable Diffusion and DreamBooth.
2023 — Booster API and Shardformer expansion
89 changes.
This period focused on establishing the high-level Booster API and the Shardformer module to simplify distributed training and model parallelization for mainstream architectures. The team significantly expanded the framework's capabilities by introducing static graph analysis, advanced memory management via Gemini and auto-offloading, and comprehensive support for LLM inference and serving. Additionally, new application modules like ColossalEval and ColossalQA were introduced to support model evaluation and retrieval-augmented generation workflows.
2024–2025 — Inference engine and RLHF expansion
35 changes.
This period focused on establishing a comprehensive inference engine architecture supporting LLMs and diffusion models, alongside significant expansion of the ColossalChat RLHF framework with GRPO and distributed training capabilities. The work introduced new model implementations, quantization modules, and high-performance kernel management to optimize serving and training workflows. Additionally, it added extensive benchmarking tools and examples for large-scale models like Grok-1 and Mixtral to validate performance and parallelism strategies.
Features
Activation checkpoint code generation with offload support
The codegen module now includes a new \ActivationCheckpointCodeGen\ component that generates Python code for activation checkpointing. This implementation supports offloading intermediate activations to CPU via custom \saved\_tensors\_hooks\ and identifies checkpoint regions based on metadata annotations in the FX graph. It handles nested checkpointing and distinguishes between offloading inputs versus intermediate bars, enabling reduced GPU memory usage during training.
colossalai/fx/codegen · high confidence
Add Auto-Parallel and Auto-Checkpoint tutorials for ResNet and GPT-2
This change introduces a new tutorial directory for auto-parallelism and automatic activation checkpointing. It includes a ResNet-50 example demonstrating auto-parallel sharding strategies across multiple GPUs, along with benchmark scripts (\auto\_ckpt\_solver\_test.py\ and \auto\_ckpt\_batchsize\_test.py\) that evaluate memory usage and throughput for GPT-2 and ResNet models using the \CheckpointSolverRotor\. The tutorial provides synthetic data generation utilities and a README explaining how to run these experimental features, which are tested against specific versions of PyTorch and Transformers.
_examples/tutorial/auto\parallel · high confidence
Add BERT and ALBERT finetuning and benchmarking examples using the Booster API
This change introduces a new example directory for BERT and ALBERT models, demonstrating how to finetune and benchmark these models using the ColossalAI Booster API. The example supports multiple parallelization plugins including TorchDDP, Gemini, LowLevelZero, and HybridParallel, allowing users to compare performance metrics such as accuracy, F1-score, memory usage, and throughput across different distributed strategies.
examples/language/bert · high confidence
Add BERT and DeBERTa-v2 model implementations to community examples
The community examples directory now includes PyTorch model implementations for BERT and DeBERTa-v2. These files provide the core model architectures, including embedding layers, attention mechanisms, and pre-trained model archive lists, allowing users to experiment with these specific architectures within the community examples framework.
examples/community/roberta/pretraining/model · high confidence
Add BF16 mixed precision support via new mixin classes
The \colossalai/amp/naive\_amp/mixed\_precision\_mixin\ module now includes a \BF16MixedPrecisionMixin\ alongside the existing \FP16MixedPrecisionMixin\ and base \MixedPrecisionMixin\. This addition enables users to perform mixed precision training using the bfloat16 data type, providing a straightforward implementation that passes through gradients without scaling, while the FP16 mixin retains its dynamic gradient scaling and overflow detection logic.
_colossalai/amp/naive\_amp/mixed\_precision\mixin · high confidence
Add ChatGLM2-6B model support to Shardformer
This change introduces the ChatGLM2-6B model implementation to the Shardformer module, enabling users to optimize and run this specific large language model. The update includes the model configuration class (\ChatGLMConfig\) defining hyperparameters such as layer count, hidden size, and attention settings, as well as the core model architecture (\modeling\_chatglm.py\) which implements the transformer blocks, rotary embeddings, and prefix encoding logic required for inference. This allows the Shardformer to apply its optimization strategies to ChatGLM2-6B models.
_colossalai/shardformer/modeling/chatglm2\6b · high confidence
Add Chinese RoBERTa preprocessing example with C++ acceleration
A new community example for preprocessing Chinese corpus for Whole Word Masked (WWM) pretraining has been added to \examples/community/roberta/preprocessing\. It provides scripts to split sentences, tokenize text, and generate masked language model (MLM) training instances (input\_ids, masks, segment\_ids, masked\_lm\_positions) in HDF5 format. The example supports both a Python backend and a faster C++ backend (via pybind11) for the masking step, which can be built using the included Makefile.
examples/community/roberta/preprocessing · high confidence
Add ColossalAI Stable Diffusion example with training, inference, and Docker support
Introduces a new example for accelerating Stable Diffusion v1 and v2 using ColossalAI, enabling users to train, fine-tune (including DreamBooth), and run inference with significantly reduced GPU memory consumption. The addition includes the core Python entry point (\main.py\), configuration files, and shell scripts for both ColossalAI and DDP training strategies. It also provides a \Dockerfile\ and instructions for containerized usage, along with a \LICENSE\ file (CreativeML Open RAIL-M) and \environment.yaml\ specifying dependencies like PyTorch 1.12.1, ColossalAI 0.2.5, and Lightning 1.9.0.
examples/images/diffusion · high confidence
Add DeepSeek V3 benchmark and smoke test
Users can now benchmark and validate DeepSeek V3 models using the new \benchmark.py\ script, which supports Expert Parallel (EP) and 3D hybrid parallelism via the \MoeHybridParallelPlugin\. A \smoke\_test.py\ has also been added to verify basic training functionality with expert parallelism, ensuring the model integrates correctly with the framework's distributed plugins.
examples/language/deepseek · high confidence
Add FastFold inference tutorial
Added a new tutorial example for FastFold inference, including a README with quick-start instructions for installing FastFold via conda, downloading datasets, and running inference scripts, along with a submodule reference to the FastFold repository.
examples/tutorial/fastfold · high confidence
Add GPT auto-offload training example
Introduces a new example demonstrating GPT model training with automatic memory offloading. The addition includes a training script (\train\_gpt\_offload.py\) that utilizes Colossal-AI's \memory\_optimize\ and \AMPOptimizer\ to manage GPU memory via a configurable budget, alongside a model zoo supporting various GPT-2 sizes (medium to 24b) and a launch script for easy execution.
_examples/language/gpt/experiments/auto\offload · high confidence
Add GPT pipeline parallelism example
The examples/language/gpt/experiments/pipeline\_parallel directory now contains a complete demo for training GPT models using pipeline parallelism. This includes a model zoo supporting various GPT-2 sizes (medium through 24B), a training script utilizing Colossal-AI's FillDrainPipelineEngine and FX-based partitioning, and a run script for launching experiments with configurable GPU counts and microbatch sizes.
_examples/language/gpt/experiments/pipeline\parallel · high confidence
Add GPT training example using Titans models
Introduces a new example in \examples/language/gpt/titans\ that demonstrates how to train GPT-2 and GPT-3 models using the Titans library. The entry includes the training script (\train\_gpt.py\), configuration files for small GPT-2 and GPT-3 models with ZeRO-3 and 1D tensor parallelism, and a README with instructions for running the demo on single or multiple nodes using either real or dummy datasets.
examples/language/gpt/titans · high confidence
Add GPT-1D model components for parallel training
The \examples/language/gpt/titans/model\ directory now contains the core implementation for a 1D-parallel GPT model. This includes \embed.py\ with vocabulary-parallel and hidden-parallel embedding layers, \gpt1d.py\ with 1D-parallel MLP and self-attention layers (including fused variants), and \pipeline\_gpt1d.py\ which assembles these into pipeline-parallel GPT models (e.g., \PipelineGPT1D\, \FusedPipelineGPT1D\). These components enable users to train GPT models with tensor parallelism across the vocabulary and linear layers, and optionally with pipeline parallelism across transformer blocks.
examples/language/gpt/titans/model · high confidence
Add GPT-2 training example with ColossalAI Gemini and ZeRO strategies
The \examples/language/gpt/gemini\ directory now contains a complete demo for training GPT-2 models using ColossalAI. This includes a main training script (\train\_gpt\_demo.py\) that supports multiple distributed strategies: ColossalAI's Gemini plugin (ZeRO Stage 3), ZeRO Stage 1 and 2, and PyTorch DDP/ZeRO. The example provides a model zoo (\commons/model\_zoo.py\) for various GPT-2 scales (medium to 40B) and includes shell scripts (\run\_gemini.sh\, \benchmark\_gemini.sh\, \test\_ci.sh\) to facilitate local runs, performance benchmarking across different GPU counts and tensor parallelism degrees, and CI testing.
examples/language/gpt/gemini · high confidence
Add Grok-1 314B inference example with tensor parallelism support
This change introduces a new example for running inference on the 314-billion parameter Grok-1 model. It provides two execution modes: a standard inference script using Hugging Face's auto device mapping, and a high-performance script leveraging ColossalAI's tensor parallelism via the \HybridParallelPlugin\ and a custom \Grok1ForCausalLMPolicy\ to shard model layers across multiple GPUs. The addition includes the necessary policy definitions for layer replacement, utility functions for tokenization and generation, and shell scripts to launch the inference on 8x A100/H800 GPUs.
examples/language/grok-1 · high confidence
Add KTO dataset preparation support
The data preparation scripts now support the KTO (Kahneman-Tversky Optimization) dataset type. Users can prepare KTO-formatted data by running the new \prepare\_kto\_dataset.sh\ script or by invoking \prepare\_dataset.py\ with \--type kto\, which utilizes the \tokenize\_kto\ function from the \coati.dataset\ module.
_applications/ColossalChat/examples/data\_preparation\scripts · high confidence
Add Locust-based load testing scripts for inference endpoints
Added a Locust load-testing suite (locustfile.py) and helper scripts (run\_locust.sh, test\_ci.sh) to the inference client examples. The suite defines tasks for online generation, online chat, and offline generation, supporting both streaming and non-streaming modes, allowing users to benchmark the inference server's performance under load.
examples/inference/client · high confidence
Add MNIST example with optional FP8 and Transformer Engine support
A new community example has been added at examples/community/fp8/mnist that demonstrates training and inference on the MNIST dataset. The example integrates NVIDIA's Transformer Engine, allowing users to enable FP8 precision for linear layers via command-line flags (--use-te, --use-fp8) to improve performance and reduce memory utilization on compatible hardware.
examples/community/fp8 · high confidence
Add MiDaS monocular depth estimation module
The \examples/images/diffusion/ldm/modules/midas\ directory now includes the full MiDaS implementation, providing a new \MiDaSInference\ module that enables monocular depth estimation within the diffusion pipeline. This addition introduces support for multiple model architectures—including DPT-Large, DPT-Hybrid, MiDaS v2.1, and a smaller EfficientNet-based variant—allowing users to generate depth maps from images using these pre-trained backbones.
examples/images/diffusion/ldm/modules/midas · high confidence
Add Mixtral MoE application with training and inference scripts
Introduces a new application for running the Mixtral-8x7B Mixture-of-Experts model, including \train.py\ and \infer.py\ scripts, shell launchers (\train.sh\, \infer.sh\), and a \README.md\. The training script supports hybrid parallelism (pipeline, data, and expert parallelism) via ColossalAI's \MoeHybridParallelPlugin\, while the inference script uses expert parallelism. The package is structured as a standalone installable module (\setup.py\) with version 1.0.0.
applications/ColossalMoE · high confidence
Add Mixtral benchmark and smoke test
The examples/language/mixtral directory now includes a benchmark script (benchmark.py) for evaluating Mixtral model performance using the MoeHybridParallelPlugin, along with a smoke test (smoke\_test.py) and CI shell script to verify basic functionality. The benchmark supports various configuration options for tensor, expert, pipeline, and sequence parallelism, while the smoke test ensures the model can run successfully with expert parallelism.
examples/language/mixtral · high confidence
Add OPT fine-tuning and benchmarking examples
New examples for the Meta OPT model are now available in the \examples/language/opt\ directory. Users can fine-tune OPT models (such as facebook/opt-350m) on the Netflix shows dataset using the ColossalAI Booster API with various plugins including TorchDDPPlugin, LowLevelZeroPlugin, HybridParallelPlugin, and GeminiPlugin. A benchmarking script is also included to test performance throughput and memory usage across different plugins and GPU configurations.
examples/language/opt · high confidence
Add OPT model fine-tuning tutorial with synthetic dataset support
This change introduces a new tutorial example for fine-tuning Meta's OPT (Open Pretrained Transformer) models using Colossal-AI's Gemini and ZeRO strategies. The example includes scripts to run training on the WikiText-2 dataset or a newly added synthetic dataset, allowing users to experience up to a 40% speedup on a single GPU compared to traditional frameworks. It provides both a quick-start guide for tutorials and a practical launch script supporting various model sizes (125m to 66b) and multi-GPU configurations.
examples/tutorial/opt/opt · high confidence
Add PaLM PyTorch example implementation
Introduces a new PyTorch example for the PaLM model, including the core model definition, an autoregressive wrapper for text generation, and an \_\init\\_ module for easy importing. The implementation features parallel residual attention, rotary positional embeddings, and SwiGLU feed-forward networks, with optimizations such as replacing einsum operations with matmul for efficiency.
_examples/language/palm/palm\pytorch · high confidence
Add PaLM PyTorch example with Booster API support
Introduces a new PyTorch implementation of the PaLM model in the examples/language/palm directory. The example includes a README, training script (train.py), and shell scripts for running and testing. It demonstrates the new Booster API, allowing users to train the model using various plugins such as TorchDDP, Gemini, and LowLevelZero, with support for distributed training configurations and performance metrics like TFLOPS.
examples/language/palm · high confidence
Add Rank Recorder tool for multi-process debugging
Introduces a new \RankRecorder\ utility in \colossalai.utils.rank\_recorder\ that allows users to record and visualize the execution time of specific code blocks across multiple distributed ranks. The tool provides a context manager (\with recorder(...)\) to capture start/end events, automatically dumps per-rank JSON logs, and merges them into a single file with a Gantt-chart-style visualization (PNG/SVG) to help debug performance bottlenecks in multi-process training workflows.
_colossalai/utils/rank\recorder · high confidence
Add Ray Serve and TorchServe inference deployment examples
New example deployments for Colossal Inference are added under \colossalai/legacy/inference/serving\, providing ready-to-use integration with Ray Serve and TorchServe. The Ray Serve example (\ray\_serve/\) demonstrates a scalable serving architecture using a \Driver\ and \Worker\ pattern with tensor parallelism support, including configuration via Pydantic models and batched generation endpoints. The TorchServe example (\torch\_serve/\) provides a custom handler (\Colossal\_Inference\_Handler.py\) for deploying Bloom and Llama models, complete with Dockerfile, configuration files, and documentation for archiving and launching models. These additions enable users to deploy large language models using established serving frameworks.
colossalai/legacy/inference/serving · high confidence
Add ResNet-18 training example for CIFAR-10
A new example in \examples/images/resnet\ provides scripts to train and evaluate a ResNet-18 model on the CIFAR-10 dataset using ColossalAI. The training script (\train.py\) supports multiple distributed plugins, including \torch\_ddp\, \torch\_ddp\_fp16\, \low\_level\_zero\, and \gemini\, allowing users to experiment with different acceleration strategies. It includes features such as checkpointing, resuming training, and target accuracy validation, with an accompanying evaluation script (\eval.py\) to assess model performance on test data.
examples/images/resnet · high confidence
Add Stable Diffusion 3 and PixArt inference examples with benchmarking tools
The \examples/inference/stable\_diffusion\ directory now includes scripts to run inference on Stable Diffusion 3 and PixArt models using the ColossalAI Inference Engine. Users can generate images via \sd3\_generation.py\ (supporting tensor parallelism and patched parallelism), benchmark performance against the standard Diffusers library using \benchmark\_sd3.py\, and evaluate image quality metrics (PSNR, LPIPS, FID) with \compute\_metric.py\. A shell script \run\_benchmark.sh\ is also provided to automate these comparisons across different models, parallelism levels, and resolutions.
_examples/inference/stable\diffusion · high confidence
Add Stable Diffusion inference scripts with 8-bit quantization support
This change introduces a suite of example scripts for Stable Diffusion inference, including text-to-image (txt2img), image-to-image (img2img), inpainting, and retrieval-based generation (knn2img). A key capability added is support for 8-bit integer (int8) inference via the bitsandbytes library, which reduces memory usage and potentially speeds up generation; this is exposed through a --use\_int8 flag in the txt2img and img2img scripts and implemented via a new utils.py module that replaces standard linear layers with 8-bit quantized equivalents. The package also includes helper scripts to download the necessary first-stage and latent diffusion models, and a script to train the nearest-neighbor searcher used by the knn2img example.
examples/images/diffusion/scripts · high confidence
Add Stable Diffusion model components and sampling algorithms
This change introduces the core model architecture and inference utilities for Stable Diffusion within the \ldm/models\ directory. It adds the \AutoencoderKL\ class for latent space encoding and decoding, along with the \DDPM\ and \LatentDiffusion\ base classes for the diffusion process itself. To support generation, it includes multiple sampling implementations: \DDIMSampler\ for deterministic sampling, \PLMSSampler\ for higher-order Langevin dynamics, and \DPMSolverSampler\ for fast differential equation-based solving. Additionally, it provides a \NoisyLatentImageClassifier\ for classifier-guided generation and utility functions for sampling normalization.
examples/images/diffusion/ldm/models · high confidence
Add Tensor Detector utility for monitoring GPU memory usage
Introduces a new \TensorDetector\ utility in \colossalai.utils\ that allows users to track and report tensor states (device, shape, dtype, memory size) on GPU during model execution. The tool distinguishes between module parameters and intermediate activations, supports logging to a file, and provides a summary of total GPU memory allocated, aiding in debugging memory leaks or understanding memory consumption patterns.
_colossalai/utils/tensor\detector · high confidence
Add ViT finetuning and benchmarking example
Introduces a new Vision Transformer (ViT) example in \examples/images/vit\ that demonstrates finetuning a pretrained ViT model on the Beans dataset using ColossalAI's Booster API. The example supports multiple training strategies via plugins, including TorchDDP, LowLevelZero, Gemini, and HybridParallel (tensor/pipeline parallelism), and includes scripts for both interactive demo runs and performance benchmarking.
examples/images/vit · high confidence
Add WebtextDataset for GPT Titan examples
A new WebtextDataset class is introduced in the GPT Titan examples to handle loading and tokenizing the WebText dataset. It supports reading from a JSONL file, caching the encoded data as a PyTorch tensor, and provides input IDs and attention masks compatible with the GPT-2 tokenizer.
examples/language/gpt/titans/dataset · high confidence
Add code evaluation utilities for code generation tasks
The code reward module now includes \testing\_util.py\ and \utils.py\ to support evaluating generated code. These files provide functions to run code snippets (both call-based and standard input styles) with timeout protection, capture output, and check correctness via local execution or a remote API, enabling the system to assess the quality of code generation tasks.
_applications/ColossalChat/coati/distributed/reward/code\reward · high confidence
Add community RoBERTa pretraining example with ColossalAI distributed training support
This change introduces a new community example for pretraining RoBERTa models, located in \examples/community/roberta/pretraining\. The example provides a complete training pipeline using the ColossalAI framework, supporting distributed strategies such as CAI\_Gemini, CAI\_ZeRO1, and CAI\_ZeRO2. It includes scripts for initial training and resuming from checkpoints, along with utilities for dataset handling (via \NvidiaBertDatasetProvider\), evaluation, logging (TensorBoard and Weights & Biases), and performance monitoring. Users can configure hyperparameters like learning rate, batch size, and sequence length via command-line arguments defined in \arguments.py\.
examples/community/roberta/pretraining · high confidence
Add core LDM training utilities and optimizer support
The examples/images/diffusion/ldm directory now includes essential infrastructure for running Latent Diffusion Models, specifically adding learning rate schedulers (LambdaWarmUpCosineScheduler, LambdaLinearScheduler), distribution classes for latent space sampling (DiagonalGaussianDistribution), and utility functions for model instantiation and logging. A custom AdamW optimizer with Exponential Moving Average (EMA) of parameters is also introduced to support stable training dynamics.
examples/images/diffusion/ldm · high confidence
Add core diffusion model architecture and utility modules
This change introduces the foundational components for the Latent Diffusion Model (LDM) implementation within the \ldm/modules/diffusionmodules\ package. It adds \model.py\ and \openaimodel.py\, which define the neural network building blocks such as ResNet blocks, attention mechanisms (including memory-efficient attention via xformers), and up/down-sampling layers. It also includes \upscaling.py\ for handling low-scale image concatenation and noise augmentation, and \util.py\ for essential diffusion utilities like beta schedules, timestep embeddings, and gradient checkpointing. These files provide the structural backbone required for constructing and running diffusion-based image generation models.
examples/images/diffusion/ldm/modules/diffusionmodules · high confidence
Add diffusion data loaders for CIFAR-10, ImageNet, LSUN, and Teyvat
The \examples/images/diffusion/ldm/data\ module now includes dataset implementations for several image sources. \cifar10.py\ adds support for the CIFAR-10 dataset with label-to-text mapping and Hugging Face integration. \imagenet.py\ introduces loaders for the ImageNet training set, including automatic download, extraction, and class label handling. \lsun.py\ provides base and specific classes for LSUN categories (Churches, Bedrooms, Cats) with configurable resizing and flipping. \teyvat.py\ adds a loader for the Teyvat dataset via Hugging Face, and \base.py\ defines a base iterable dataset for text-to-image chains.
examples/images/diffusion/ldm/data · high confidence
Add distributed RLHF training infrastructure with GPTQ quantization support
This change introduces the core components for a detached, distributed Reinforcement Learning from Human Feedback (RLHF) pipeline using Ray. It adds a GPTQ quantization module for LLaMA models, enabling inference quantization to reduce memory usage. The Ray-based architecture separates experience generation (ExperienceMaker) from model training (Trainer) into distinct actors, communicating via a detached replay buffer. This allows for flexible deployment topologies (e.g., 2 makers to 1 trainer) and overlaps inference with training to improve throughput. The implementation includes LoRA support for efficient model updates, performance evaluation callbacks, and distributed training strategies.
(repo-wide) · high confidence
Add document retrieval QA prompt templates and documentation
This change introduces the prompt engineering assets for the new document retrieval QA feature. It adds a new \prompt.py\ module defining LangChain \PromptTemplate\ objects for English and Chinese retrieval QA, disambiguation, and summarization, along with specific trigger keywords and rejection answers for handling unanswerable queries. A corresponding \README.md\ is added to document the prompt design, including examples for normal and overlength conversation contexts.
applications/ColossalQA/colossalqa/prompt · high confidence
Add fused LayerNorm and scaled masked softmax layers
The \colossalai/nn/layer\ package now exposes fused neural network operations designed for performance optimization. This includes \MixedFusedLayerNorm\, which wraps a fused LayerNorm kernel (loaded via \LayerNormLoader\) to normalize inputs with affine parameters, and \FusedScaleMaskSoftmax\, which combines scaling, masking, and softmax into a single fused operation for attention mechanisms. The softmax layer supports both upper-triangular causal masks and general padding masks, automatically falling back to standard PyTorch implementations when hardware or shape constraints prevent kernel usage. A utility function \divide\ is also added to enforce exact integer division.
colossalai/nn/layer · high confidence
Add hybrid parallelism tutorial with synthetic data
A new tutorial example for hybrid parallelism (combining pipeline and tensor parallelism) has been added to the examples directory. It includes a training script using a synthetic dataset and a configuration file that allows users to adjust tensor parallel size and mode, enabling quick experimentation with Colossal-AI's parallelism features.
_examples/tutorial/hybrid\parallel · high confidence
Add inference operation benchmarks and smoke tests
The \examples/inference/benchmark\_ops\ directory now includes a suite of performance benchmarks and a smoke test for key inference kernels. Users can now measure the performance of Triton, CUDA, and vLLM implementations for context attention, decoding attention, flash decoding, fused rotary embeddings, KV cache memory copy, RMSNorm, and cosine-sine cache retrieval. A new \smoke\_test.py\ validates that the Triton RMSNorm and Rotary Embedding kernels produce results consistent with PyTorch references, ensuring correctness alongside performance profiling.
_examples/inference/benchmark\ops · high confidence
Add local and cloud LLM wrappers for ColossalQA
New LangChain-compatible LLM wrapper classes are introduced in the \colossalqa/local\ module to enable integration with various model backends. This includes \ColossalLLM\ and \VllmLLM\ for running models locally (via direct Hugging Face transformers or vLLM serving), \ColossalCloudLLM\ for the ColossalCloud platform, and \Pangu\ for Huawei Cloud's Pangu models. These components allow users to query these specific LLM services through a unified interface within the ColossalQA application.
applications/ColossalQA/colossalqa/local · high confidence
Add model wrappers for ColossalEval evaluation
ColossalEval now supports evaluating models via Hugging Face Transformers and vLLM. Users can run evaluations using \HuggingFaceModel\ (with optional PEFT and tensor parallelism via ShardFormer) or \vLLMModel\ (with support for quantization, tensor parallelism, and CUDA graph capture). Both wrappers provide loss calculation on target tokens and handle prompt truncation and tokenizer configuration.
_applications/ColossalEval/colossal\eval/models · high confidence
Add sequence parallelism tutorial with synthetic data and legacy engine integration
This change introduces a new tutorial example for sequence parallelism, demonstrating how to train a BERT model by splitting input tensors and intermediate activations along the sequence dimension to improve memory efficiency. The implementation includes a complete training script (\train.py\) that supports both standard and pipeline-parallel execution modes, utilizing a synthetic dataloader for immediate usability. It integrates with the legacy ColossalAI engine and AMP settings, featuring custom loss functions (\BertLoss\) that handle sequence-parallel all-reduces, a vocabulary-aware cross-entropy loss, and an annealing learning rate scheduler. Configuration is managed via \config.py\, allowing users to adjust parallel sizes and pipeline stages.
_examples/tutorial/sequence\parallel · high confidence
Add support for multiple evaluation datasets in ColossalEval
The ColossalEval dataset module now includes loaders for a wide range of benchmark datasets, enabling users to evaluate models on diverse tasks. New supported datasets include AGIEval, CEval, CMMLU, GaoKaoBench, GSM8K, LongBench, MMLU, MT-Bench, CValues, SafetyBench (English and Chinese), and a generic ColossalDataset. Each loader handles specific data formats, prompt construction, and inference configurations (such as few-shot settings and language) required for accurate evaluation.
_applications/ColossalEval/colossal\eval/dataset · high confidence
Added BSRGAN-based image degradation utilities
The \ldm/modules/image\_degradation\ module now includes implementations of the BSRGAN degradation pipeline, providing both a full variant (\bsrgan.py\) and a lighter version (\bsrgan\_light.py\) for generating realistic training data. These modules expose functions to simulate complex image degradations—including anisotropic Gaussian blurs, pixel shifts, and noise—allowing users to apply these effects to images during the diffusion model training process.
_examples/images/diffusion/ldm/modules/image\degradation · high confidence
Added Stable Diffusion v2 model components and encoders
The \ldm/modules\ directory now includes new implementation files for the Stable Diffusion v2 architecture, specifically adding \attention.py\ (with support for xformers and GEGLU activations), \ema.py\ (for exponential moving average model weights), and updated encoder modules in \ldm/modules/encoders\. These changes introduce new text embedding capabilities using Frozen T5, CLIP, and OpenCLIP models, enabling users to run the v2 diffusion models with enhanced text conditioning and memory-efficient attention mechanisms.
examples/images/diffusion/ldm/modules, examples/images/diffusion/ldm/modules/encoders · high confidence
Added legacy mixed-precision and checkpointing utilities
This change introduces a new \colossalai/legacy\ package containing utilities for automatic mixed precision (AMP) and distributed checkpointing. For AMP, it provides a unified \convert\_to\_amp\ entry point that supports three modes: NVIDIA Apex, PyTorch native AMP, and a custom 'naive' FP16 implementation, each with dedicated model, optimizer, and loss wrappers. It also adds distributed checkpointing helpers in \legacy/utils/checkpoint\ that correctly handle \ColoTensor\ sharding by gathering and scattering parameters during save and load operations, along with activation checkpointing support that preserves RNG states across forward and backward passes.
colossalai/legacy/utils · high confidence
Added sample datasets for document retrieval QA
The ColossalQA application now includes a new set of sample data files in the \data/data\_sample\ and \data/tests\ directories to support document retrieval and question-answering workflows. These additions include English and Chinese company profiles (\companies.txt\, \companies\_zh.txt\), a CSV of 100 synthetic organizations, JSON datasets for customer service query/response pairs and classification tasks, and a large 64KB test file, providing the necessary content for testing and demonstrating the system's retrieval capabilities.
applications/ColossalQA/data · high confidence
Added synthetic data pipeline for sequence parallel tutorial
The \examples/tutorial/sequence\_parallel/data\ directory now includes a complete data loading infrastructure to support the sequence parallelism tutorial. This adds synthetic dataset builders (BERT, ICT, and blendable datasets), custom PyTorch data samplers for distributed training, and C++ helpers for efficient index mapping. It also introduces sequence-aware batch processing via \get\_batch\_for\_sequence\_parallel\, which shards token sequences across tensor parallel ranks, enabling users to run the tutorial with generated data without requiring external datasets.
_examples/tutorial/sequence\parallel/data · high confidence
Added utility for post-initialization hooks on PyTorch modules
The \colossalai/utils/model/utils.py\ module now provides the \InsertPostInitMethodToModuleSubClasses\ context manager, which allows users to inject custom logic that runs automatically after any \torch.nn.Module\ subclass is initialized. This utility recursively patches existing and future module subclasses to execute a specified post-init method, and optionally manages the default torch dtype during the context scope, facilitating lazy memory allocation or other initialization-time modifications.
colossalai/utils/model · high confidence
Added utility functions for data generation and performance metrics
A new utils.py file has been added to the examples/language/commons directory, providing helper functions for generating random input data (input\_ids and attention masks) and calculating training throughput (TFLOPS). These utilities support example scripts by abstracting away boilerplate tensor creation and performance calculation logic.
examples/language/commons · high confidence
ColossalChat introduces GRPO training and math-specific reward scoring
The ColossalChat application now supports Group Relative Policy Optimization (GRPO) as a reinforcement learning algorithm, accessible via the \rl\_example.py\ script with the \-a GRPO\ flag. This update includes new math-focused reward scoring utilities in \coati/utils/reward\_score/\ that validate response structures (e.g., \\<answer\>\ tags) and verify final answers for GSM8K and math competition datasets. Additionally, a new \NaiveExperienceBuffer\ implementation in \coati/experience\_buffer/\ manages training experiences with optional CPU offloading, and checkpoint I/O utilities in \coati/utils/ckpt\_io.py\ handle model, optimizer, and scheduler state persistence.
applications/ColossalChat · high confidence
Experimental pass for runtime shape consistency
An experimental FX graph pass has been added to handle runtime shape consistency. This new module, \adding\_shape\_consistency\_pass.py\, introduces a \shape\_consistency\_pass\ that modifies the FX graph by inserting calls to a \ShapeConsistencyManager\. The pass identifies nodes with assigned strategies and injects logic to apply sharding spec conversions between predecessor and successor nodes, ensuring that tensor shapes are consistent across the graph at runtime.
colossalai/fx/passes/experimental · high confidence
Experimental profiler adds FLOP/MAC estimation for common PyTorch operations
The experimental profiler now includes function-level registration for estimating floating-point operations (FLOPs) and multiply-accumulate operations (MACs) across a broad set of PyTorch operations. This update adds support for activation functions (e.g., ReLU, GELU, Sigmoid), arithmetic and element-wise operations (e.g., add, matmul, bmm), normalization layers (e.g., BatchNorm, LayerNorm, InstanceNorm), pooling operations, embedding lookups, and various tensor manipulation ops (e.g., permute, flatten, index\_select). Users can now obtain more detailed computational cost estimates for models using these operations within the experimental profiling pipeline.
_colossalai/fx/profiler/experimental/profiler\function · high confidence
Experimental profiler adds FLOPs/MACs support for common PyTorch modules
The experimental profiler module now includes metadata handlers for a wide range of PyTorch operations, enabling approximate FLOPs and MACs estimation. New support covers activation functions (ReLU, GELU, Sigmoid, etc.), convolution layers (1D/2D/3D and Transpose variants), linear layers, normalization layers (BatchNorm, LayerNorm, GroupNorm, and NVIDIA Apex fused variants), pooling layers, RNN variants (LSTM, GRU, RNN and their cells), attention mechanisms, and utility operations like embedding, dropout, and flatten. This allows users to profile model complexity for these specific components within the experimental FX profiler.
_colossalai/fx/profiler/experimental/profiler\module · high confidence
Initial CLI entry point with check and launcher commands
The ColossalAI CLI is now available via the \colossalai.cli\ module, exposing a Click-based interface that registers two subcommands: \check\ (for installation verification) and \run\ (the distributed launcher). Users can invoke these commands through the standard CLI entry point.
colossalai/cli · high confidence
Initial project scaffolding and configuration
The repository has been initialized with the core project structure, including the \setup.py\ build script, \version.txt\ (set to 0.5.0), and \CHANGE\_LOG.md\. Essential configuration files have been added to standardize development workflows: \.pre-commit-config.yaml\ enforces code formatting via Black, isort, and clang-format; \.clang-format\ sets the C++ style to Google; \.isort.cfg\ configures import sorting; and \.coveragerc\ enables parallel multiprocessing for test coverage. Additionally, \.compatibility\ and \.cuda\_ext.json\ define supported PyTorch/CUDA version pairs (e.g., 2.3.0/cu121, 2.5.1/cu124) for extension building, while \.gitignore\ and \MANIFEST.in\ manage source control and package inclusion.
(repo-wide) · high confidence
Initial release of Colossal-AI tutorial examples
This change introduces the \examples/tutorial\ directory, providing a new set of hands-on tutorials for Colossal-AI. The entry includes a README with setup instructions and links to video guides for topics such as multi-dimensional parallelism, sequence parallelism, large batch optimization, and fine-tuning models like OPT and Stable Diffusion. It also adds a script to download the CIFAR-10 dataset and a .gitignore file to exclude local data directories.
examples/tutorial · high confidence
Introduce AutoChunk memory estimation and code generation
The autochunk module now includes core components for estimating inference memory usage and generating chunked execution code. The new \estimate\_memory.py\ provides logic to track active nodes and calculate peak memory consumption, while \autochunk\_codegen.py\ generates the Python code required to execute models in chunks, including loop structures and tensor slicing. Supporting utilities in \utils.py\ and graph manipulation classes in \reorder\_graph.py\ and \trace\_indice.py\ enable the system to analyze node dependencies and optimize execution order for memory efficiency.
colossalai/autochunk · high confidence
Introduce Booster high-level training API
The \colossalai/booster\ package is introduced, providing a unified \Booster\ class to simplify distributed training setup. Users can now configure hardware acceleration via the \Accelerator\ class, apply mixed-precision strategies (fp16, bf16, fp8), and integrate distributed training logic through \Plugin\ implementations. The \Booster.boost()\ method wraps models, optimizers, and data loaders to inject these features, while \execute\_pipeline()\ supports pipeline-parallel training workflows.
colossalai/booster · high confidence
Introduce Coati RLHF model components and loss functions
The \applications/ColossalChat/coati/models\ package now provides the core model architecture and training utilities for RLHF workflows. This includes base, critic, and reward model classes for handling pretrained language models, along with a new \RLVRRewardModel\ for verifiable rewards. The update adds support for Low-Rank Adaptation (LoRA) via \LoraConfig\ and \LoraManager\, and introduces a comprehensive set of loss functions including \DpoLoss\ (supporting both DPO and SimPO objectives), \KTOLoss\, \PolicyLoss\, \ValueLoss\, and pairwise reward losses (\LogSigLoss\, \LogExpLoss\). Additionally, it provides generation utilities for token sampling with temperature, top-k, and top-p controls, as well as helper functions for masking and log-probability calculations.
applications/ColossalChat/coati/models · high confidence
Introduce ColoTensor and ColoParameter as distributed tensor wrappers
The \colossalai/tensor\ module now provides \ColoTensor\ and \ColoParameter\, which are PyTorch tensor subclasses designed to intercept operations via \\_\_torch\function\\_\. This enables automatic integration with the \ColoParamOpHook\ system for distributed memory management and communication. The module also introduces \ShardingSpec\ and \ShapeConsistencyManager\ to define and transform tensor sharding layouts across device meshes, along with \CommSpec\ to execute collective operations like all-gather and all-to-all based on these specs.
colossalai/tensor · high confidence
Introduce Colossal Inference module with async engine and dynamic batching
Adds a new inference module under \colossalai/legacy/inference\ that provides a high-performance inference framework. This includes an \Async\_Engine\ for handling requests asynchronously, an \Async\_DynamicBatchManager\ for efficient dynamic batching of inference requests, and a \DynamicBatchManager\ base class. The module exposes \CaiInferEngine\ and \LlamaModelInferPolicy\ via its public API, integrating with tensor parallelism and supporting models like Llama and Bloom.
colossalai/legacy/inference · high confidence
Introduce Colossal-LLaMA application with LLaMA-3 support and training utilities
This change introduces the Colossal-LLaMA application, providing the infrastructure for continual pre-training and supervised fine-tuning of LLaMA-2 and LLaMA-3 models. It includes dataset loaders and tokenization logic for both pre-training and supervised fine-tuning (SFT), with specific conversation templates for LLaMA-2 and LLaMA-3. The application also provides utilities for expanding tokenizers and initializing model embeddings for new vocabularies, checkpoint I/O for saving/loading training states, and training enhancements such as NEFTune noise injection and streaming chat capabilities.
applications/Colossal-LLaMA · high confidence
Introduce ColossalAI-Inference module for accelerated LLM and Diffusion model serving
Adds a new \colossalai.inference\ package providing an \InferenceEngine\ and \InferenceConfig\ to accelerate Transformer and Diffusion model inference. The module introduces a unified API for offline generation and online serving (via \AsyncInferenceEngine\ and \RPCInferenceEngine\), supporting features such as continuous batching, paged attention, speculative decoding, CUDA graph execution, and tensor parallelism. It includes optimized modeling components, a KV cache manager, and configurable logit processors for sampling strategies like top-k, top-p, and repetition penalty.
colossalai/inference · high confidence
Introduce ColossalEval, an LLM evaluation framework
Adds the ColossalEval application, providing a uniform pipeline to evaluate language models on public datasets (such as MMLU, CMMLU, GSM8K, and AGIEval) using both classic metrics and GPT-based evaluation. The package includes installation instructions, configuration examples for inference and evaluation, and documentation detailing supported metrics and datasets.
applications/ColossalEval · high confidence
Introduce ColossalQA document retrieval QA application
Adds the ColossalQA application, a Langchain-based system for document retrieval and conversation. The package includes a web UI demo, Python scripts for Chinese, English, and bilingual retrieval QA, and support for various LLMs (ChatGLM2, LLaMa2, ChatGPT, Pangu) and vector stores (Chroma).
applications/ColossalQA · high confidence
Introduce DeviceMesh and AlphaBetaProfiler for logical device layout and communication profiling
The \colossalai/device\ module now provides a \DeviceMesh\ class that allows users to define a logical view of physical devices (e.g., shaping a 16-GPU cluster into a 2x2x4 mesh) and manage the underlying distributed process groups. It also includes an \AlphaBetaProfiler\ to automatically measure latency and bandwidth characteristics (alpha and beta values) for these device groups, which are used to estimate communication costs. Additionally, a new \calc\_pipeline\_strategy\ module implements the Alpa dynamic programming algorithm to automatically determine optimal pipeline parallel stage configurations based on these communication profiles.
colossalai/device · high confidence
Introduce Distributed Tensor (DTensor) and Layout Converter
This change introduces the Distributed Tensor (DTensor) subsystem, providing a new API to wrap PyTorch tensors with sharding specifications across a device mesh. Users can now create sharded tensors using \distribute\_tensor\ and \init\_as\_dtensor\, define sharding layouts via \ShardingSpec\, and convert between layouts using the new \LayoutConverter\. The module also includes \CommSpec\ to manage collective communication patterns (all-gather, all-to-all, all-reduce) required for layout transformations, and a \PaddedTensor\ utility to handle dynamic sequence lengths during sharding.
_colossalai/tensor/d\tensor · high confidence
Introduce FX graph analysis and partitioning passes for auto-parallelism
The \colossalai/fx/passes\ module now provides a suite of IR passes to enable automatic model partitioning and memory profiling. \MetaInfoProp\ and \ConcreteInfoProp\ analyze FX graphs to record tensor metadata, FLOPs, and memory usage per node, which powers the \metainfo\_trace\ API. These passes feed into partitioning strategies like \balanced\_split\_pass\ and \gpipe\_dp\_split\_pass\ (adapted from Alpa) to automatically split models across pipeline stages. Additionally, \shard\_1d\_pass\ annotates linear layers for 1D tensor parallelism, and \split\_module\ handles the structural decomposition of the graph into sub-modules.
colossalai/fx/passes · high confidence
Introduce FX-based memory and FLOPs profiler for PyTorch models
Adds a new \colossalai.fx.profiler\ module that estimates forward and backward memory usage, FLOPs, and execution time for PyTorch models by tracing their FX graphs. The profiler uses a custom \MetaTensor\ subclass to perform zero-allocation tracing on meta devices, calculating memory costs (including temporary activations and gradients) and operation counts for standard \torch.nn\ modules and functions. It provides \profile\_function\, \profile\_method\, and \profile\_module\ wrappers to instrument models, supporting both meta-device estimation and concrete-device measurement for validation, and includes utilities for handling sharding specs and autograd graph analysis.
colossalai/fx/profiler · high confidence
Introduce FX-based static graph analyzer and symbolic profiling
This change adds a new static graph analysis subsystem under \colossalai/\_analyzer/fx\, providing tools to inspect and optimize PyTorch models. It introduces \symbolic\_trace\ and \ColoGraphModule\ to extend Torch FX tracing with support for activation checkpointing (including nested regions) and custom code generation. Additionally, it provides \symbolic\_profile\ to compute per-node metrics such as FLOPs, memory footprint, and communication costs, which feed into the auto-parallel system's cost modeling.
_colossalai/\analyzer/fx · high confidence
Introduce FastAPI-based online inference server with chat and completion APIs
Users can now run a local HTTP API server for Colossal-Inference using FastAPI, accessible at http://127.0.0.1:8000/docs. This new server exposes endpoints for text completion (/completion), chat interactions (/chat), and raw generation (/generate), supporting both streaming and non-streaming responses. The service allows configuration of the Hugging Face model, data type (fp16, fp32, bf16), batch size, and input/output lengths via command-line arguments, and includes health check endpoints for monitoring server and engine status.
colossalai/inference/server · high confidence
Introduce GPT-based LLM evaluation pipeline
The \colossal\_eval.evaluate\ module now provides a new evaluation pipeline that uses GPT-3.5 and GPT-4 to assess Large Language Model outputs. This feature supports two modes: 'battle' evaluations, which compare the performance of two different models side-by-side, and standard evaluations, which rate model answers against pre-defined metrics (such as language organization, relevance, creativity, and correctness) using Chain-of-Thought prompting. The implementation includes configuration for Chinese and English capabilities, allowing users to define custom evaluation categories and metrics via JSON configs and prompt templates.
_applications/ColossalEval/colossal\eval/evaluate · high confidence
Introduce GRPO and RLVR support for PPO in experience maker
The experience maker now supports Group Relative Policy Optimization (GRPO) and Reinforcement Learning with Verifiable Rewards (RLVR) for PPO. Users can enable GRPO by setting \use\_grpo=True\, which disables the critic model requirement and generates multiple responses per prompt to calculate advantages. This change also includes infrastructure for handling inference rebatching via \inference\_batch\_size\ and \logits\_forward\_batch\_size\ parameters to manage memory usage during generation.
_applications/ColossalChat/coati/experience\maker · high confidence
Introduce Gemini ZeRO-3 memory management and optimizer components
This change introduces the core implementation of the Gemini ZeRO-3 engine within the \colossalai/zero/gemini\ package. It adds \GeminiDDP\ and \GeminiOptimizer\ to handle model wrapping and optimizer state sharding, alongside \GeminiManager\ and \ChunkManager\ for managing tensor chunks across CPU and GPU memory. The update includes \GeminiZeROHook\ for intercepting forward/backward operations to manage data movement, and implements \StaticPlacementPolicy\ and \AutoPlacementPolicy\ to control how tensors are evicted or prefetched based on memory constraints. Additionally, utility functions are provided to convert the sharded model into a standard PyTorch module for checkpointing or inference.
colossalai/zero/gemini · high confidence
Introduce Gemini memory tracer module
The \colossalai/zero/gemini/memory\_tracer\ package is now available, providing tools to trace and collect GPU and CPU memory statistics during model training. It includes \RuntimeMemTracer\ for dynamic tracing via forward/backward hooks, \StaticMemStatsCollector\ for static analysis using Torch FX, and various monitors (\AsyncMemoryMonitor\, \SyncCudaMemoryMonitor\) and collectors (\MemStatsCollector\, \ChunkMemStatsCollector\) to record memory usage patterns.
_colossalai/zero/gemini/memory\tracer · high confidence
Introduce JIT-optimized fused kernels and warmup utilities
Adds a new \colossalai.kernel.jit\ module providing PyTorch JIT-compiled implementations of common transformer operations, specifically fused bias+GELU and bias+dropout+add kernels, to improve inference and training performance. The module also includes a \set\_jit\_fusion\_options\ function to configure PyTorch's JIT fusion settings based on the installed version (using nvfuser for PyTorch 1.10+ and legacy fusers for earlier versions) and a \warmup\_jit\_fusion\ utility that pre-compiles these kernels with representative tensor shapes to reduce first-run latency during training.
colossalai/kernel/jit · high confidence
Introduce Low-Level ZeRO optimizer implementation
Adds a new \LowLevelZeroOptimizer\ class in the \colossalai/zero/low\_level\ module, providing a foundational implementation for ZeRO-1 and ZeRO-2 parallelism. This component introduces optimized gradient partitioning, communication bucketing, and mixed-precision handling (including FP8 support) to improve training efficiency and memory usage compared to previous implementations.
_colossalai/zero/low\level · high confidence
Introduce MTBench evaluation support in ColossalEval
The dataset evaluator now supports the MTBench benchmark via a new \mtbench\_single\_judge\ metric. This change adds a dedicated \gpt\_judge\ module that uses GPT-4 to generate single-turn and multi-turn judgements for model outputs, including specific handling for math, reasoning, and coding categories with reference answers. The evaluator's metric selection logic has been updated to recognize and route MTBench data to this new GPT-based evaluation path, while also maintaining existing support for label-based, loss-based, and other standard metrics.
_applications/ColossalEval/colossal\_eval/evaluate/dataset\evaluator · high confidence
Introduce MoeTensor API for expert parallelism metadata
This change introduces a new \moe\_tensor\ module that provides a programmatic interface for managing and querying expert parallelism (EP) and data parallelism (DP) metadata on tensors. Users can now check if a tensor is an MoE tensor, set and retrieve the expert parallel group, and access parallel sizes and ranks via helper functions like \get\_ep\_size\, \get\_dp\_rank\, and \get\_ep\_group\. The module also includes the \MoeParallelInfo\ class to encapsulate parallel topology information, supporting configurations where expert parallelism is nested inside or outside data parallelism.
_colossalai/tensor/moe\tensor · high confidence
Introduce Naive AMP mixed-precision optimizer with FP16/BF16 support
Users can now use the new \MixedPrecisionOptimizer\ in \colossalai.amp.naive\_amp\ to perform mixed-precision training. This optimizer wraps an existing PyTorch optimizer and supports both FP16 and BF16 precisions, automatically managing master weights and applying gradient scaling, clipping, and overflow checks via dedicated mixins (\NaiveFP16MixedPrecisionMixin\ and \BF16MixedPrecisionMixin\).
_colossalai/amp/naive\amp · high confidence
Introduce Paged Attention-based KV Cache Manager
The inference engine now uses a new block-based KV cache system to manage memory more efficiently. This change introduces a \KVCacheManager\ and \CacheBlock\ classes that implement Paged Attention, allowing the system to allocate physical KV cache tensors in fixed-size blocks rather than contiguous memory. This enables better memory utilization and supports variable-length sequences without pre-allocating excessive static memory, which is particularly beneficial for long-context inference scenarios.
_colossalai/inference/kv\cache · high confidence
Introduce Pipeline Inference module with micro-batch management
Adds a new Pipeline Inference module that enables model parallelism for large models that do not fit on a single GPU. This location introduces the \MicroBatchManager\ to track micro-batch states (prefill, generate, done) and KV-cache information, along with a benchmarking script and documentation to demonstrate usage and performance against Hugging Face pipelines.
colossalai/legacy/inference/pipeline · high confidence
Introduce ShardFormer module for automatic model parallelization
The new ShardFormer module automatically parallelizes mainstream models from HuggingFace and TIMM, allowing users to enable optimizations like tensor parallelism, pipeline parallelism, fused normalization, flash attention, and sequence parallelism via a simple \ShardConfig\. The module provides a \ShardFormer\ API to optimize models and supports custom sharding policies for user-defined architectures, with initial support documented for models such as BERT, T5, LLaMA, GPT-2, Bloom, ChatGLM, ViT, Whisper, SAM, BLIP2, Falcon, and Mistral.
colossalai/shardformer · high confidence
Introduce Zero Bubble distributed RL framework for GRPO and DAPO
Adds a new distributed reinforcement learning pipeline under \applications/ColossalChat/coati/distributed/zero\_bubble\ that implements a producer–consumer architecture to overlap rollout and training, thereby reducing GPU idle time. The implementation includes a \Distributor\ for broadcasting model weights, \Producer\ and \Consumer\ components for handling data and training steps, and specific consumer implementations for GRPO and DAPO algorithms. Users can enable this zero-bubble mode via the \--zero 2\ flag in the RL training example, with a configurable \data\_actor\_buffer\_size\_limit\ to manage the trade-off between throughput and data staleness.
_applications/ColossalChat/coati/distributed/zero\bubble · high confidence
Introduce auto-parallel runtime passes for shape consistency and metadata propagation
Adds a new set of FX graph passes in \colossalai/auto\_parallel/passes\ to handle distributed execution requirements. The \runtime\_preparation\_pass\ annotates nodes with sharding strategies and inserts placeholder nodes for sharding and communication dictionaries. The \runtime\_apply\_pass\ injects runtime calls to perform shape consistency conversions and apply communication specifications (like all-reduce) between nodes with mismatched sharding. The \meta\_info\_prop\ pass propagates memory and compute cost metadata through the graph, while \comm\_metainfo\_pass\ attaches specific metadata to communication nodes, enabling accurate profiling and optimization of the parallelized model.
_colossalai/auto\parallel/passes · high confidence
Introduce automatic parameter offloading to reduce GPU memory usage
Added a new auto-offload module in \colossalai/auto\_parallel/offload\ that automatically partitions a model's parameters into regions and offloads them between GPU and CPU to fit within a specified memory budget. The feature includes a \memory\_optimize\ entry point that analyzes the model graph, uses a solver to determine an efficient offloading strategy (supporting both synchronous and asynchronous modes), and wraps the model in a \BaseOffloadModule\ to handle runtime data movement. It also provides an \AMPOptimizer\ wrapper to manage mixed-precision training with dynamic gradient scaling and overflow detection in conjunction with the offloaded parameters.
_colossalai/auto\parallel/offload · high confidence
Introduce bookkeeping module for gradient and tensor bucket management
Added a new \bookkeeping\ package to the low-level ZeRO implementation, containing \BaseStore\, \BucketStore\, \GradientStore\, and \TensorBucket\. These classes provide the underlying infrastructure for managing gradient partitions, organizing parameters into communication buckets, and handling tensor flattening and all-gather operations, which are essential for the efficient execution of ZeRO optimization stages.
_colossalai/zero/low\level/bookkeeping · high confidence
Introduce centralized distributed training launch APIs
Colossal-AI now provides a unified entry point for initializing distributed environments through the new \colossalai.initialize\ module. Users can launch distributed training via \colossalai.launch\ for manual configuration, or use convenience wrappers \launch\_from\_slurm\, \launch\_from\_openmpi\, and \launch\_from\_torch\ that automatically detect rank and world size from their respective environment variables. The initialization process handles backend selection via the accelerator module, sets the CUDA device, configures random seeds, and ensures consistent kernel launch ordering by setting \CUDA\_DEVICE\_MAX\_CONNECTIONS=1\.
colossalai · high confidence
Introduce cluster module with distributed coordination and process group management
Adds a new \colossalai.cluster\ package that provides core infrastructure for distributed training. This includes a \DistCoordinator\ singleton for managing ranks, world size, and synchronized execution (e.g., printing only on master, blocking non-executor processes), a \ProcessGroupManager\ for creating and destroying named process groups, a \ProcessGroupMesh\ for organizing processes into multi-dimensional meshes with coordinate/rank conversion, and a \DeviceMeshManager\ for initializing and managing device meshes with automatic shape profiling.
colossalai/cluster · high confidence
Introduce core inference engine architecture with LLM and Diffusion support
This change introduces the foundational core components for the ColossalAI inference engine, establishing a unified interface for both Large Language Models (LLMs) and Diffusion models. The \InferenceEngine\ class acts as a dispatcher, routing requests to either \LLMEngine\ or \DiffusionEngine\ based on the model type. \LLMEngine\ provides features such as CUDA graph capture, speculative decoding support, and continuous batching via a \RequestHandler\. \DiffusionEngine\ enables inference for models like PixArt-Alpha and Stable Diffusion 3 using the \diffusers\ library. Additionally, an \RPCInferenceEngine\ is added to support multi-GPU online serving via RPyC, and an \AsyncInferenceEngine\ is provided for asynchronous generation. The module also includes \InferCheckpoint\_io\ for optimized model loading and \BaseEngine\ as the abstract foundation for these implementations.
colossalai/inference/core · high confidence
Introduce distributed RL training framework with GRPO and DAPO support
This location adds a new distributed Reinforcement Learning framework for fine-tuning large language models, supporting the GRPO and DAPO algorithms. The implementation uses a producer-consumer architecture orchestrated by Ray, where inference producers generate rollouts and training consumers optimize the policy. It includes specific consumer implementations for GRPO, communication utilities for Ray collective operations, and inference backends for vLLM, SGLang, and Hugging Face Transformers. The code also introduces a 'zero-bubble' training variant designed to reduce idle time between rollout generation and training steps.
applications/ColossalChat/coati/distributed · high confidence
Introduce distributed shardformer layer primitives
The \colossalai/shardformer/layer\ package now provides a comprehensive set of distributed-aware PyTorch modules to support model parallelism. This includes 1D parallel linear layers (\Linear1D\_Col\, \Linear1D\_Row\) with gradient accumulation and async communication, parallelized embeddings (\Embedding1D\, \VocabParallelEmbedding1D\), and distributed loss functions (\DistCrossEntropy\, \DistLogProb\) that compute cross-entropy and log-probability across tensor-parallel shards. The update also introduces parallel attention mechanisms (\ColoAttention\, \RingAttention\), fused normalization layers (\FusedLayerNorm\, \FusedRMSNorm\) with Apex/NPU support, and dropout layers (\DropoutForParallelInput\, \DropoutForReplicatedInput\) that ensure consistent randomization across ranks to prevent convergence issues.
colossalai/shardformer/layer · high confidence
Introduce distributed training launcher with multi-node support
The \colossalai run\ CLI command is now available to launch distributed training jobs on single or multiple nodes. It supports specifying hosts via command-line arguments or a hostfile, allows filtering devices with include/exclude options, and handles SSH-based execution across remote machines. The launcher automatically detects the PyTorch version to use the appropriate distributed launcher (\torch.distributed.launch\ for older versions, \torch.distributed.run\ for newer ones) and supports running user scripts as Python modules.
colossalai/cli/launcher · high confidence
Introduce document retrieval-based question answering
ColossalQA now supports retrieval-augmented generation (RAG) for question answering. Users can load documents (CSV, JSON, HTML, Markdown, PDF, TXT) and tables (CSV, Excel, Parquet, etc.) via new \DocumentLoader\ and \TableLoader\ components. The system builds vector indexes using a custom \CustomRetriever\ with incremental update support and answers queries using a \RetrievalQA\ chain that combines retrieved context with an LLM. A \ConversationBufferWithSummary\ memory class manages conversation history, summarizing older turns to stay within token limits while preserving recent context. Bilingual (English and Chinese) conversation wrappers are provided, including input disambiguation and rejection handling when answers cannot be derived from the provided information.
applications/ColossalQA/colossalqa · high confidence
Introduce lazy tensor initialization and Hugging Face model loading support
This change adds a new lazy initialization subsystem to ColossalAI, introducing \LazyTensor\ and \LazyInitContext\ to defer actual tensor allocation until materialization, which allows for memory-efficient model construction. It also integrates a \PretrainedManager\ that patches Hugging Face Transformers' \from\_pretrained\ method, enabling users to load pre-trained models directly into this lazy execution mode for optimized memory usage during initialization.
colossalai/lazy · high confidence
Introduce legacy ZeRO and Gemini memory management components
This change adds the foundational modules for the legacy ZeRO and Gemini memory management systems within \colossalai/legacy/zero\. It introduces the \convert\_to\_zero\_v2\ helper to integrate models and optimizers with ZeRO v2, along with \ShardedModelV2\ and \ShardedOptimizerV2\. It also implements the Gemini memory management infrastructure, including \StatefulTensor\ and \StatefulTensorMgr\ for tracking tensor states, \TensorPlacementPolicy\ classes (CPU, CUDA, Auto) for managing memory eviction and placement, and various operation hooks (\ophooks\) to handle parameter sharding and gradient memory tracing during forward and backward passes.
colossalai/legacy/zero · high confidence
Introduce legacy pipeline parallelism API
The \colossalai/legacy/pipeline\ module is now available, providing a complete API for pipeline-parallel training. This includes \PipelinableModel\ and \PipelinableContext\ for declaratively defining model layer specifications and execution sequences, a middleware system (\Topo\, \Partition\) to analyze graph topology and data dependencies, and RPC-based pipeline engines (\FillDrainPipelineEngine\, \OneFOneBPipelineEngine\, \ChimeraPipelineEngine\) that manage distributed stage execution, microbatch scheduling, and backward caching.
colossalai/legacy/pipeline · high confidence
Introduce legacy tensor distribution and specification modules
The \colossalai/legacy/tensor\ package now exposes a structured set of modules for managing tensor parallelism and distribution specifications. Users can now access \ColoTensorSpec\ to define tensor properties, \ComputeSpec\ and \ComputePattern\ to configure computation patterns (such as TP1D, TP2D, TP3D), and \DistSpecManager\ to handle the transformation of tensor distributions (e.g., sharding, gathering, all-to-all) across process groups. The package also provides \ProcessGroup\ for managing tensor and data parallel process group configurations, and \ShardSpec\/\ReplicaSpec\ for defining distributed placement patterns.
colossalai/legacy/tensor · high confidence
Introduce meta-profiler registry for auto-parallel cost estimation
The \colossalai/auto\_parallel/meta\_profiler\ module has been added to provide a structured registry for computing forward and backward compute and memory costs of PyTorch operations. This new component defines \ShardMetaInfo\ classes and a \meta\_register\ registry that map specific operations—including linear layers, convolutions, normalization, pooling, embeddings, and elementwise functions—to cost estimation functions. By integrating these meta-information handlers, the auto-parallel system can now accurately estimate resource requirements for different sharding strategies during the compilation phase.
_colossalai/auto\_parallel/meta\profiler · high confidence
Introduce mixed precision training plugin architecture
Users can now configure mixed precision training via a new plugin system in the booster. This change adds a \MixedPrecision\ base class and concrete implementations for FP16 (via PyTorch AMP, Apex, and naive strategies), BF16, and FP8. A factory function allows selecting the precision type by string key, enabling users to easily switch between different mixed precision strategies for their models and optimizers.
_colossalai/booster/mixed\precision · high confidence
Introduce modular CheckpointIO system with async and sharded support
The checkpointing subsystem has been refactored into a modular \checkpoint\_io\ package, introducing a base \CheckpointIO\ interface and specialized implementations for general, hybrid-parallel, and Mixture-of-Experts (MoE) workloads. Users can now save and load model and optimizer states in sharded formats compatible with Hugging Face conventions, including support for \safetensors\ and asynchronous I/O operations to improve performance. The new system handles distributed tensor gathering, pinned memory optimization, and index file management automatically, providing a unified API for checkpointing across different parallel strategies.
_colossalai/checkpoint\io · high confidence
Introduce new ShardFormer API for model sharding
The \colossalai/shardformer/shard\ module now provides a new user-facing API for sharding Hugging Face models. Users can configure sharding options (such as tensor parallelism, sequence parallelism, and gradient checkpointing) via the new \ShardConfig\ dataclass and apply them using the \ShardFormer.optimize()\ method, which returns the sharded model and any shared parameters. This replaces previous integration patterns with a dedicated, structured entry point for model parallelization.
colossalai/shardformer/shard · high confidence
Introduce new auto-parallel solver module with cost graph and ILP optimization
The \colossalai/auto\_parallel/tensor\_shard/solver\ package has been introduced, providing a new infrastructure for automatic parallelization strategy selection. This module includes a \CostGraph\ to linearize and simplify resharding costs between nodes, a \GraphAnalyser\ for variable liveness analysis, and a \StrategiesConstructor\ that generates candidate strategies for graph nodes (including support for distributed dataloaders and getattr handlers). The core \Solver\ class integrates these components to formulate and solve an Integer Linear Programming (ILP) problem using the \pulp\ library, aiming to find optimal strategy combinations that minimize compute and communication costs while respecting memory budgets.
_colossalai/auto\_parallel/tensor\shard/solver · high confidence
Introduce new pipeline parallel scheduling implementations
The \colossalai/pipeline/schedule\ module now provides a structured set of pipeline parallel execution strategies, including the standard 1F1B schedule (\OneForwardOneBackwardSchedule\), interleaved pipeline parallelism (\InterleavedSchedule\), and the Zero-Bubble V-pipe scheduler (\ZeroBubbleVPipeScheduler\). These schedules support advanced features such as FP8 communication casting, metadata caching for P2P efficiency, and arbitrary batch sizes in forward-only modes. Additionally, a new \GenerateSchedule\ class is introduced to handle pipeline parallel inference with micro-batch management and token generation loops.
colossalai/pipeline/schedule · high confidence
Introduce new pipeline parallelism infrastructure
The \colossalai/pipeline\ module has been restructured with new core components: \PipelineStageManager\ for managing stage topology and interleaved execution, \PipelineP2PCommunication\ for optimized peer-to-peer data transfer with device-aware serialization, and \WeightGradStore\ for buffering weight gradients. The module now exposes multiple scheduling strategies including \InterleavedSchedule\, \OneForwardOneBackwardSchedule\, and \ZeroBubbleVPipeScheduler\.
colossalai/pipeline · high confidence
Introduce new tensor-sharding autoparallel module with solver options and sharding strategies
A new \colossalai.auto\_parallel.tensor\_shard\ module has been added to provide an autoparallel system for tensor sharding. This includes a \SolverOptions\ dataclass that allows users to configure solver preferences (Standard, Data Parallel, Tensor Parallel), dataloader options (Replicated, Distributed), and shard options (Standard, Shard, Shard Last Axis, Full Shard). The module also introduces core data structures for defining sharding strategies, including \ShardingStrategy\, \CommAction\, and \OperationData\, along with constants mapping PyTorch operations to their types for parallelization analysis. Initialization logic is provided to set up device meshes and build strategy constructors for graph analysis.
_colossalai/auto\_parallel/tensor\shard · high confidence
Introduce new utility modules for timing, memory, and safetensors serialization
The \colossalai/utils\ package now exposes several new modules to support performance measurement, memory management, and checkpointing. Users can now use \Timer\ and \MultiTimer\ to measure execution times with accelerator synchronization, access \colo\_device\_memory\_capacity\ and \colo\_get\_cpu\_memory\_capacity\ for detailed CPU and GPU memory reporting (including container-aware CPU metrics), and leverage the new \safetensors\ utilities (\save\, \save\_nested\, \move\_and\_save\) to serialize model and optimizer state dictionaries using the safetensors format with optional async I/O support.
colossalai/utils · high confidence
Introduce speculative decoding drafter infrastructure
Added the \colossalai/inference/spec\ module containing the \Drafter\ class, which serves as a container for an assistant model used in speculative decoding. This component handles token speculation via greedy search, manages KV cache trimming, and ensures compatibility with Hugging Face \transformers\ versions prior to 4.38.0. The module also includes \DrafterOutput\ and \GlideInput\ dataclasses to structure drafter results and support GLIDE-style model interactions.
colossalai/inference/spec · high confidence
Introduce static graph analyzer with MetaTensor and symbolic profiling
The \colossalai/\_analyzer\ module is added, providing a static graph analysis toolkit built on a modified version of PyTorch FX. This new capability introduces \MetaTensor\ to enable ahead-of-time profiling, shape propagation, and ideal FLOP counting without executing the model on real hardware. It also adds \symbolic\_trace()\ for robust control-flow tracing and \symbolic\_profile()\ to simulate memory allocation and deallocation, allowing the auto-parallel system to gather performance metrics (such as FLOPs and memory usage) efficiently before deployment.
_colossalai/\analyzer · high confidence
Introduce unified plugin architecture for distributed training strategies
The \colossalai/booster/plugin\ module now provides a structured plugin system that abstracts distributed training configurations. Users can select from specific plugins—\TorchDDPPlugin\, \GeminiPlugin\, \LowLevelZeroPlugin\, \HybridParallelPlugin\, and \MoeHybridParallelPlugin\—to manage data, tensor, pipeline, and expert parallelism. The \HybridParallelPlugin\ and \MoeHybridParallelPlugin\ support advanced features like sequence parallelism, FP8 precision, and LoRA integration, while the \TorchFSDPPlugin\ is conditionally available for PyTorch 1.12+. Each plugin handles its own checkpointing logic and dataloader preparation, simplifying the setup for complex multi-dimensional parallel training.
colossalai/booster/plugin · high confidence
Introduces MetaTensor and FLOP counting infrastructure for static graph analysis
Adds a new \colossalai/\_analyzer/\_subclasses\ module that provides \MetaTensor\ and \MetaTensorMode\ to enable device-agnostic static graph analysis by executing operations on PyTorch's meta device while tracking logical device placement. This infrastructure includes monkey-patching for factory methods and aten operations, meta-registrations for convolutional layers, and a \FlopTensor\ implementation that allows users to count floating-point operations (FLOPs) for forward and backward passes of PyTorch models without performing actual hardware computation.
_colossalai/\_analyzer/\subclasses · high confidence
Introduces ZeRO DDP wrapper API with Gemini support
Users can now wrap models and optimizers for ZeRO Distributed Data Parallel training using the new \zero\_model\_wrapper\ and \zero\_optim\_wrapper\ functions. The model wrapper configures ZeRO stages 1 and 2 directly, while stage 3 automatically integrates the GeminiDDP engine for advanced memory management. The optimizer wrapper handles dynamic gradient scaling and selects the appropriate backend (\LowLevelZeroOptimizer\ for stages 1-2 or \GeminiOptimizer\ for stage 3) based on the model's configuration.
colossalai/zero · high confidence
Introduces extensible kernel loader for automatic hardware-aware kernel selection
The \colossalai/kernel\ module now provides a \KernelLoader\ system that automatically selects and loads the most appropriate CUDA or CPU kernel implementation based on the current machine's capabilities. This change introduces a registry-based architecture where specific loaders (such as \FlashAttentionLoader\, \MoeLoader\, and \CpuAdamLoader\) manage multiple backend extensions (e.g., NPU, CUDA, SDPA, X86, ARM). Users benefit from transparent hardware compatibility checks and priority-based fallbacks, ensuring the optimal kernel is loaded without manual configuration.
colossalai/kernel · high confidence
Introduces static graph analysis passes for shape and FLOP profiling
Adds new modules in \colossalai/\_analyzer/fx/passes\ that enable static analysis of Torch FX graphs. The \shape\_prop\ module provides a \ShapeProp\ interpreter to propagate tensor shapes and metadata through the graph without full execution, while \graph\_profile\ introduces \GraphProfiler\ and \FlopProfiler\ to estimate memory usage and floating-point operations (FLOPs) for each node. These tools allow users to analyze model complexity and resource requirements before deployment.
_colossalai/\analyzer/fx/passes · high confidence
Introduction of Auto Parallel documentation and module initialization
This change introduces the \colossalai/auto\parallel\ module by adding its \\\init\\_.py\ file and a new \README.md\. The documentation outlines the automatic parallel system's architecture, detailing the Analyzer, Solver, and Generator components, and explains how it optimizes distributed execution plans for PyTorch models based on research from systems like Alpa and Rotor.
_colossalai/auto\parallel · high confidence
Introduction of Config utility and thread-safe singleton pattern
The \colossalai.context\ package now provides a \Config\ class that allows configuration dictionaries to be accessed as attributes and loaded directly from Python files, alongside a \SingletonMeta\ metaclass that ensures thread-safe instantiation of singleton objects using double-checked locking. This change introduces new foundational utilities for managing configuration and global state within the library.
colossalai/context · high confidence
Introduction of MultiTensorApply utility for efficient tensor operations
A new MultiTensorApply utility has been added to the colossalai.utils.multi\_tensor\_apply module. This class allows users to apply operations to lists of tensors efficiently by chunking them, leveraging underlying C++/CUDA extensions if available (typically from NVIDIA Apex). If the required extensions are not installed, the utility gracefully degrades and raises a descriptive error when attempting to use the functionality, ensuring users are aware of the missing dependency.
_colossalai/utils/multi\_tensor\apply · high confidence
Introduction of colossalai.nn package with weight initialization utilities
The \colossalai/nn\ module has been introduced, providing a public API for neural network components. This change adds a new \init.py\ file that exposes standard weight initialization functions (such as \zeros\\, \ones\\, \uniform\\, \normal\\, \trunc\normal\\, \kaiming\uniform\\, \kaiming\normal\\, and \xavier\uniform\\) as factory functions returning initializers. These initializers wrap PyTorch's \torch.nn.init\ methods and include support for \fan\_in\ and \fan\_out\ parameters to assist with variance preservation during model construction.
colossalai/nn · high confidence
Introduction of core MoE communication operations
The \colossalai/moe\ module has been initialized with foundational communication primitives required for Mixture-of-Experts models. This includes the implementation of \AllGather\, \ReduceScatter\, and \AllToAll\ operations, which support both synchronous execution and asynchronous overlap modes to improve performance. Additionally, a new \HierarchicalAllToAll\ operation has been added to handle complex data dispatching across intra-node and inter-node process groups, enabling more efficient expert routing in distributed training environments.
colossalai/moe · high confidence
Introduction of generalized learning rate schedulers with warmup and delay capabilities
The \colossalai/nn/lr\_scheduler\ module now provides a suite of learning rate schedulers that extend PyTorch's standard implementations with support for linear warmup and flat delay phases. New classes such as \CosineAnnealingWarmupLR\, \MultiStepWarmupLR\, \PolynomialWarmupLR\, and \FlatAnnealingWarmupLR\ allow users to linearly warm up the learning rate before applying the base schedule, while \DelayerScheduler\ and \FlatAnnealingLR\ enable keeping the learning rate fixed for a specified number of steps before decay begins. These schedulers are designed to be compatible with checkpointing via proper \state\_dict\ and \load\_state\_dict\ implementations, ensuring that training can be resumed correctly after interruptions.
_colossalai/nn/lr\scheduler · high confidence
Introduction of legacy gradient handlers for distributed parallel strategies
The \colossalai/legacy/engine/gradient\_handler\ module has been added, providing specialized classes to manage gradient synchronization across different parallel execution modes. This includes handlers for Data Parallel, Pipeline Parallel (specifically for shared modules), Sequence Parallel, ZeRO optimization, and Mixture-of-Experts (MoE) models. These components implement bucketed all-reduce operations to optimize communication efficiency and are registered in the legacy registry, enabling users to select the appropriate gradient handling strategy based on their distributed training configuration.
_colossalai/legacy/engine/gradient\handler · high confidence
Introduction of meta-patch module for FX tracer customization
The \colossalai/fx/tracer/meta\_patch\ package has been introduced to support advanced tracing capabilities, including data-dependent control flow, pooling layer handling, and bias addition in modules. This new module provides the foundational structure (\patched\_function\ and \patched\_module\) that enables the FX tracer to correctly interpret and trace these specific PyTorch operations, ensuring accurate model representation during compilation.
_colossalai/fx/tracer/meta\patch · medium confidence
Legacy parallel context and process group initializers added
The \colossalai/legacy/context\ package now includes the \ParallelContext\ singleton, the \ParallelMode\ enumeration (covering data, tensor, pipeline, sequence, 1D, 2D, 2.5D, and 3D modes), and a suite of \ProcessGroupInitializer\ classes (for data, model, pipeline, tensor, sequence, 1D, 2D, 2.5D, and 3D parallelism) that register distributed process groups via the \DIST\_GROUP\_INITIALIZER\ registry. Additionally, a \SeedManager\ and helper functions are provided to manage random seeds and states across these parallel modes.
colossalai/legacy/context · high confidence
New BERT finetuning tutorial for GLUE tasks
Added a new tutorial example in \examples/tutorial/new\_api/glue\_bert\ that demonstrates how to finetune a BERT model on GLUE benchmark tasks using the ColossalAI Booster API. The example includes a training script (\finetune.py\) and data builder (\data.py\) that support multiple distributed plugins, including TorchDDP (with FP32 and FP16 options), Gemini, and LowLevelZero, allowing users to easily switch between different acceleration strategies.
_examples/tutorial/new\_api/glue\bert · high confidence
New BERT model implementation for sequence parallelism tutorial
The tutorial's sequence parallel model now uses a newly implemented BERT architecture located in \examples/tutorial/sequence\_parallel/model\. This change replaces the previous implementation with a custom model that includes dedicated layers for sequence-parallel attention (\TransformerSelfAttentionRing\), MLPs, embeddings, and heads. The new codebase supports both standard and pipeline-parallel training modes, allowing users to experiment with sequence parallelism strategies in the tutorial environment.
_examples/tutorial/sequence\parallel/model · high confidence
New CLI command to verify Colossal-AI installation and environment compatibility
Users can now run \colossalai check --installation\ to validate their setup. This command reports the versions of Colossal-AI, PyTorch, and CUDA, and checks compatibility between the system CUDA, PyTorch's required CUDA version, and the versions used for AOT-compiled CUDA extensions. It helps identify mismatches that could cause runtime errors, such as using a PyTorch build with a different CUDA version than the system or the library.
colossalai/cli/check · high confidence
New Coati trainer module with support for multiple RLHF algorithms
The \applications/ColossalChat/coati/trainer\ package has been introduced, providing a unified training interface for several reinforcement learning from human feedback (RLHF) and preference optimization algorithms. This includes supervised fine-tuning (SFT), Direct Preference Optimization (DPO), KTO, ORPO, PPO, and GRPO trainers, all inheriting from base classes for supervised (\SLTrainer\) and online learning (\OLTrainer\) workflows. The module also includes a callback system with a \PerformanceEvaluator\ to track throughput and latency metrics during training.
applications/ColossalChat/coati/trainer · high confidence
New ColossalQA examples for retrieval-based conversations and a WebUI demo
This update adds a suite of new example scripts in the ColossalQA examples directory, introducing retrieval-based conversation capabilities across multiple languages and models. Users can now run standalone conversation agents backed by ChatGPT (with SQL and vector retrieval tools), LLaMa2 (English), and ChatGLM (Chinese), featuring features like input disambiguation, conversation memory with summarization, and rejection handling for out-of-scope queries. Additionally, a new WebUI demo is provided, consisting of a FastAPI backend server and a Gradio-based chat interface, allowing users to upload documents, build a knowledge base, and interact with the RAG system via a browser.
applications/ColossalQA/examples · high confidence
New DreamBooth example with Colossal-AI integration and LoRA support
The \examples/images/dreambooth\ directory now provides a complete DreamBooth training example that integrates with Colossal-AI. This includes a new \train\_dreambooth\_colossalai.py\ script leveraging the Colossal-AI Booster API and plugins like Gemini for heterogeneous memory management, alongside a standard \train\_dreambooth.py\ for comparison. A new \train\_dreambooth\_colossalai\_lora.py\ script adds support for Low-Rank Adaptation (LoRA) to reduce memory usage. The example also includes scripts for inpainting (\train\_dreambooth\_inpaint.py\), inference, and CI testing, along with updated documentation and shell scripts for easy setup.
examples/images/dreambooth · high confidence
New GPT-2 Hybrid Parallelism examples for benchmarking and fine-tuning
Added a new set of example scripts in the \examples/language/gpt/hybridparallelism\ directory to demonstrate GPT-2 training with ColossalAI's hybrid parallelism plugins. The \benchmark.py\ script allows users to evaluate performance across various parallel strategies (Gemini, FSDP, and 3D hybrid parallel) and model sizes (118M to 6.21B). The \finetune.py\ script provides a complete workflow for fine-tuning GPT-2 on GLUE tasks, supporting plugins like \hybrid\_parallel\, \gemini\, and \low\_level\_zero\, along with utilities for data handling (\data.py\) and a launch script (\run.sh\).
examples/language/gpt/hybridparallelism · high confidence
New GPT-2 auto-parallelism example
Added a new demo in \examples/language/gpt/experiments/auto\_parallel\ that showcases how to run GPT-2 training with Colossal-AI's auto-parallelism engine. The entry includes a README with installation instructions for PyTorch, Colossal-AI, and required dependencies, along with a Python script (\auto\_parallel\_with\_gpt.py\) that initializes the model, applies the \autoparallelize\ function to distribute computation across devices, and runs a training loop with performance metrics.
_examples/language/gpt/experiments/auto\parallel · high confidence
New LLaMA pretraining benchmark example
Added a new example in examples/language/llama that provides benchmarking scripts and a smoke test for pretraining LLaMA-1, LLaMA-2, and LLaMA-3 models. The entry includes a README with usage instructions, a benchmark.py script supporting various parallelism plugins (Gemini, FSDP, Hybrid), and shell scripts for running benchmarks on 7B and 70B model configurations. A smoke test ensures basic functionality works with tensor parallelism.
examples/language/llama · high confidence
New Llama inference example and benchmarking scripts
The \examples/inference/llama\ directory now includes a new \llama\_generation.py\ script for running inference on Llama models (including Llama 3) with support for tensor parallelism, speculative decoding, and StreamingLLM. Additionally, \benchmark\_llama.py\ and \benchmark\_llama3.py\ scripts are provided to measure performance metrics like latency and throughput, along with a \run\_benchmark.sh\ helper script to automate benchmarking runs.
examples/inference/llama · high confidence
New LoRA fine-tuning and DPO/GRPO training examples for ColossalChat
This update adds new training scripts and configuration files to the ColossalChat examples. It introduces \lora\_finetune.py\ and \lora\_config.json\ for supervised fine-tuning with LoRA (using PiSSA initialization), along with sample data. It also adds \train\_dpo.py\ and \train\_dpo.sh\ for Direct Preference Optimization training, and \train\_grpo.py\ for Group Relative Policy Optimization with RLVR support. These scripts support various parallelism plugins (DDP, Gemini, Zero2, 3D, MoE) and allow users to fine-tune models using preference data or reinforcement learning objectives.
_applications/ColossalChat/examples/training\scripts · high confidence
New OPT inference tutorial with HTTP API and caching
Added a new tutorial example in \examples/tutorial/opt/inference\ that demonstrates running OPT model inference (125M to 175B parameters) via an HTTP server. The example provides two server implementations (\opt\_fastapi.py\ using FastAPI and \opt\_server.py\ using Sanic) that support tensor parallelism, request batching, and configurable response caching. It also includes scripts for processing and converting OPT-66B and OPT-175B checkpoints, along with a Locust-based benchmarking setup.
examples/tutorial/opt/inference · high confidence
New RPC-based inference executor for distributed model serving
The inference executor now supports remote procedure calls (RPC) via a new \rpc\_worker\ module, allowing inference tasks to be executed on remote workers. This change introduces an \rpcWorkerService\ that handles distributed environment initialization, model loading (with support for lazy initialization and specific policies for Llama and Baichuan models), KV-cache allocation, and forward execution with token sampling. This enables a client-server architecture for distributed inference workloads.
colossalai/inference/executor · high confidence
New Triton-accelerated inference kernels for attention, normalization, and rotary embeddings
This change introduces a new \colossalai/kernel/triton\ module that provides a suite of high-performance Triton GPU kernels to accelerate large language model inference. The module exports several key components: context and flash-decoding attention kernels (\context\_attn\_unpad\, \flash\_decoding\) that support paged/blocked KV caches for efficient memory management; fused rotary embedding kernels (\fused\_rotary\_embedding\, \no\_pad\_rotary\_embedding\) that apply position embeddings without padding; a fused RMS LayerNorm kernel (\rms\_layernorm\); and specialized kernels for KV cache copying (\kvcache\_copy\) and Llama-style activation combining (\llama\_act\_combine\_kernel\). These kernels are designed to reduce memory overhead and improve throughput for inference workloads by fusing operations and optimizing data movement on the GPU.
colossalai/kernel/triton · high confidence
New accelerator abstraction module for hardware portability
ColossalAI introduces a new \accelerator\ module that provides a unified abstraction layer for hardware backends, allowing users to easily switch between Nvidia GPUs, Huawei NPUs, and CPUs. This module exposes a simple \auto\_set\_accelerator()\ API that automatically detects and initializes the available hardware (checking CUDA, NPU, then CPU in priority order), as well as explicit \set\_accelerator()\ and \get\_accelerator()\ functions. The implementation includes specific accelerator classes (\CudaAccelerator\, \NpuAccelerator\, \CpuAccelerator\) that wrap PyTorch's native device APIs (including \torch.cuda\, \torch.npu\, and CPU operations) for device management, random number generation, and memory handling, aiming to make user code portable across different hardware platforms.
colossalai/accelerator · high confidence
New auto-parallel node handlers for PyTorch operations
The auto-parallel tensor shard module now includes a comprehensive set of node handlers to automatically generate sharding strategies for a wide range of PyTorch operations. This change introduces support for linear projections (torch.nn.Linear, torch.addmm), matrix multiplications (torch.matmul, torch.bmm, torch.addbmm), convolutions (torch.nn.Conv1d/2d/3d), normalization layers (torch.nn.LayerNorm, torch.nn.BatchNorm1d/2d/3d), embeddings (torch.nn.Embedding), and various elementwise and reshape operations (torch.flatten, torch.add, torch.pow, etc.). These handlers enable the framework to automatically parallelize these common neural network components across device meshes.
_colossalai/auto\_parallel/tensor\_shard/node\handler · high confidence
New benchmarking suite for RLHF alignment algorithms
The benchmarks directory now includes scripts and configuration to evaluate performance and memory usage for SFT, DPO, KTO, ORPO, and SimPO training strategies on OPT models. This addition introduces dedicated benchmark runners (e.g., benchmark\_dpo.sh, benchmark\_kto.sh) that utilize dummy dataset generation and support various ColossalAI parallelism plugins (zero2, zero2\_cpu, 3d) and LoRA with gradient checkpointing, enabling users to compare throughput and resource consumption across different alignment methods.
applications/ColossalChat/benchmarks · high confidence
New checkpoint solver API with Chen and Rotor strategies
The auto-parallel checkpoint module now exposes a user-friendly API for selecting optimal activation checkpointing strategies. Users can choose between the Chen greedy solver (focused on memory optimization) and the Rotor solver (a dynamic programming approach for heterogeneous chains) via the new \CheckpointSolverChen\ and \CheckpointSolverRotor\ classes. The Rotor solver includes a C extension (\rotorc\) for faster computation and supports linearizing graphs to handle common nodes and shape-consistency operations. This allows users to automatically generate optimized computation graphs that respect memory constraints while minimizing training time.
_colossalai/auto\parallel/checkpoint · high confidence
New convergence and performance benchmark examples for BERT and LLaMA
Added new example scripts in the shardformer examples directory to benchmark model optimization. The convergence benchmark (convergence\_benchmark.py) demonstrates fine-tuning a BERT model on GLUE tasks using tensor parallelism and all optimizations, including data loading, training loops, and evaluation metrics. The performance benchmark (performance\_benchmark.py) uses Triton to compare the inference speed of an original LLaMA model versus a sharded version with fused normalization and tensor parallelism. Supporting data utilities (data.py) and a shell script (convergence\_benchmark.sh) are also included to facilitate these benchmarks.
colossalai/shardformer/examples · high confidence
New dataset and GPT-based evaluation examples for ColossalEval
Added example scripts and configuration files for two new evaluation workflows in ColossalEval: dataset evaluation and GPT-based evaluation. The dataset evaluation examples provide inference and evaluation scripts for standard benchmarks (MMLU, CMMLU, AGIEval, GaoKaoBench, C-Eval) using local metrics, while the GPT evaluation examples enable model comparison and single-model assessment using OpenAI's GPT models (e.g., gpt-3.5-turbo, gpt-4) with configurable categories like chat, roleplay, and open QA.
applications/ColossalEval/examples · high confidence
New dataset infrastructure for KTO, GRPO, and prompt-based training
The dataset module now includes a new \Conversation\ class for managing chat templates and system messages, alongside dedicated data collators (\DataCollatorForKTODataset\, \DataCollatorForPreferenceDataset\, \DataCollatorForPromptDataset\, \DataCollatorForSupervisedDataset\) and tokenization utilities (\tokenize\_kto\, \tokenize\_prompt\, \tokenize\_sft\, \tokenize\_rlhf\). This enables support for new training paradigms like KTO and GRPO, as well as code generation tasks via prompt-based loading, replacing the previous custom dataloader setup with these built-in components.
applications/ColossalChat/coati/dataset · high confidence
New distributed logging module with Rich formatting and singleton management
The \colossalai/logging\ package now provides a \DistributedLogger\ class that manages singleton logger instances per name, automatically using the \rich\ library for formatted console output when available (falling back to standard stream handlers). Users can retrieve loggers via \get\_dist\_logger\, control verbosity with \set\_level\, and persist logs to files using \log\_to\_file\. A new \disable\_existing\_loggers\ utility allows selective silencing of other Python loggers to reduce noise, defaulting to keeping only the 'colossalai' logger active.
colossalai/logging · high confidence
New distributed optimizers and auto-cast to distributed versions
Colossal-AI now includes distributed implementations for the GaLore, CAME, AdaFactor, and LAMB optimizers, enabling memory-efficient training with Tensor Parallelism and ZeRO. To simplify usage, the optimizer module automatically converts standard local instances of these algorithms into their distributed counterparts when used in a distributed setting, removing the need for manual wrapping.
colossalai/nn/optimizer · high confidence
New dynamic batching inference engine and GPTQ/SmoothQuant quantization support
This change introduces a new dynamic batching inference architecture in \colossalai/legacy/inference/dynamic\_batching\, providing core components for request handling, batching, and Ray-based distributed execution. It also adds GPTQ and SmoothQuant quantization modules, enabling 4-bit and 8-bit inference optimizations for models like LLaMA.
_colossalai/legacy/inference/dynamic\batching · high confidence
New examples documentation and testing integration guide
Added a comprehensive README for the examples directory that outlines the folder structure, provides an invitation for community contributions, and details the requirements for integrating new examples with the automated CI testing workflow (including the creation of \test\_ci.sh\ scripts).
examples · high confidence
New extensions module for unified high-performance kernel management
The \extensions\ directory introduces a new module that provides a unified, hardware-aware loader for high-performance kernels (such as CPU Adam, Flash Attention, and LayerNorm). This module allows users to load optimized kernels that automatically adapt to the current hardware (CPU, GPU, NPU) and compiler backend (AOT or JIT) without manual configuration. It includes a base extension framework (\\_Extension\) for registering custom kernels and C++ source files with type traits and dispatch macros to support multi-backend compilation.
extensions · high confidence
New hybrid engine for pipeline-parallel LLaMA inference
The \colossalai/legacy/inference/hybridengine\ module introduces a new \CaiInferEngine\ class that enables pipeline-parallel inference for LLaMA models. This engine manages tensor and pipeline parallelism, handles KV-cache memory via \MemoryManager\, and executes generation using optimized Triton-based attention kernels (with optional \lightllm\ or \flash\_attn\ backends). It exposes a \LlamaModelInferPolicy\ to replace standard Hugging Face forward methods with custom inference forwards and supports data types fp16, bf16, and fp32.
colossalai/legacy/inference/hybridengine · high confidence
New inference layer components for attention, Baichuan TP, and diffusion models
This change introduces new files in the inference layers module: a pure PyTorch PagedAttention implementation with KV cache management and padding mask support, a tensor-parallel linear layer for Baichuan models that integrates lazy initialization, and wrappers for diffusion pipelines including a distributed 'Distrifusion' acceleration path for PixArt and SD3 transformers. Users gain access to these specific layer implementations for inference workloads, enabling paged attention handling, Baichuan tensor parallelism, and distributed diffusion inference within this module.
colossalai/inference/modeling/layers · high confidence
New inference model implementations for Llama, Baichuan, Glide, PixArt-Alpha, and Stable Diffusion 3
This change introduces new inference-specific model implementations in the \colossalai/inference/modeling/models\ directory. It adds \nopadding\_llama.py\ and \nopadding\_baichuan.py\, which provide optimized, no-padding forward passes for Llama and Baichuan models, integrating custom CUDA kernels for RMSNorm, rotary embeddings, and attention backends to improve inference performance. It also adds \glide\_llama.py\ to support speculative decoding with the Glide drafter model for Llama architectures. Furthermore, it introduces \pixart\alpha.py\ and \stablediffusion3.py\, enabling diffusion model inference for PixArt-Alpha and Stable Diffusion 3 by adapting their respective pipelines to the ColossalAI inference engine. The \\\init\\_.py\ file is created to expose these new model modules.
colossalai/inference/modeling/models · high confidence
New inference model policies for Llama, Baichuan, and diffusion models
This change introduces a new policy registry and specific inference policies for several model architectures. It adds support for NoPadding Llama and Baichuan models, enabling optimized tensor-parallel inference with fused linear operations and custom forward passes. Additionally, it introduces inference policies for PixArt-Alpha and Stable Diffusion 3, integrating Distrifusion acceleration layers (such as DistrifusionConv2D and DistrifusionFusedAttention) to support patched parallelism in diffusion models. A GlideLlama policy is also added to support GLIDE drafter models. These policies are registered in a central map for easy lookup by model name.
colossalai/inference/modeling/policy · high confidence
New language example utilities for data, models, and performance profiling
The \examples/language\ package now includes core utility modules to support language model examples. \data\_utils.py\ provides a \StatefulDistributedSampler\ for resuming data iteration and a \prepare\_dataloader\ helper for distributed training setups. \model\_utils.py\ adds context managers for low-precision initialization and helpers for formatting model parameter counts. \performance\_evaluator.py\ introduces a \PerformanceEvaluator\ class that tracks throughput and TFLOPS, along with configurable profilers (PyTorch, Nsight Systems, or dummy) to facilitate performance benchmarking.
examples/language · high confidence
New legacy MoE layer components and OpenMoE benchmarking tools
This change introduces the core implementation files for the legacy Mixture-of-Experts (MoE) system, including \colossalai/legacy/moe/layer/\_\init\\_.py\, \experts.py\ (defining \MLPExperts\), \layers.py\ (defining \SparseMLP\), \routers.py\, \load\_balance.py\, and \manager.py\. It also adds benchmarking scripts and documentation for the OpenMoE model under \colossalai/legacy/moe/openmoe/\, enabling users to train and evaluate OpenMoE with various parallel strategies (EP, EP-ZERO, Hybrid) and load balancing options.
colossalai/legacy/moe · high confidence
New model and optimizer wrapper interfaces for distributed training
The \colossalai/interface\ package now provides standardized wrapper classes (\ModelWrapper\, \OptimizerWrapper\) and mixins (\AMPModelMixin\, \PeftUnwrapMixin\) that define the common interface used by the Booster. \ModelWrapper\ handles model unwrapping and integrates with PEFT (LoRA) to manage state dictionaries for checkpoint saving and loading, while \OptimizerWrapper\ exposes standard optimization operations (step, zero\_grad, backward) and gradient clipping utilities. Additionally, \pretrained.py\ introduces helper functions to manage pretrained model paths.
colossalai/interface · high confidence
New pipeline and sequence-parallel modeling implementations for BERT, BLIP-2, Bloom, ChatGLM2, Command, DeepSeek, and DeepSeek-V3
The \colossalai/shardformer/modeling\ package now includes dedicated forward-pass implementations for several models, enabling pipeline parallelism and sequence parallelism optimizations. New files provide \BertPipelineForwards\, \BloomPipelineForwards\, \ChatGLMPipelineForwards\, \CommandPipelineForwards\, and \deepseek\_v3\_model\_forward\ to handle stage management, attention mask preparation, and sequence splitting/gathering. Additionally, \blip2.py\ introduces optimized attention and JIT-fused MLP/attention output functions, while \deepseek.py\ and \deepseek\_v3.py\ implement Expert Parallel (EP) modules (\EPDeepseekMoE\, \EpDeepseekV3MoE\) with support for FP8 communication and gradient scaling for Mixture-of-Experts models.
colossalai/shardformer/modeling · high confidence
New quantization module with BitsAndBytes and FP8 support
A new \colossalai.quantization\ package has been introduced, providing two distinct quantization capabilities. First, it integrates BitsAndBytes (bnb) to enable 4-bit and 8-bit model quantization via \BnbQuantizationConfig\ and \quantize\_model\, including version checks for \bitsandbytes\ compatibility. Second, it adds FP8 support, featuring FP8 linear layers, casting utilities, and optimized distributed communication hooks (all-reduce, all-gather) for FSDP and other parallel strategies, with optional \torch.compile\ acceleration for supported PyTorch versions.
colossalai/quantization · high confidence
New sharding strategy generators for core neural network operations
The auto-parallel system now includes dedicated strategy generators for a broad set of operations, enabling more flexible and accurate distributed training. This location introduces generators for BatchNorm, Conv, Embedding, LayerNorm, and various matrix multiplications (including Dot, Linear, and Batched MatMul), as well as elementwise operations (Binary, Unary, Where), reshaping (Split, Transpose, View, Permute), and data access (Getattr, GetItem). Each generator defines specific sharding strategies, computes forward and backward memory and compute costs, and handles necessary communication actions to maintain correctness across devices.
_colossalai/auto\_parallel/tensor\_shard/node\handler/strategy · high confidence
New symbolic and meta tracing APIs with bias-addition graph restructuring
The \colossalai.fx.tracer\ module now exposes \symbolic\_trace\ and \meta\_trace\ entry points, enabling users to trace models using meta tensors to support dynamic control flow and pre-tracing without concrete inputs. The underlying \ColoTracer\ has been extended to automatically restructure computation graphs for operations involving bias addition (such as \F.linear\, \torch.addmm\, and \torch.addbmm\), decomposing them into separate linear/batch matrix multiplication and bias addition nodes to ensure correct graph representation for auto-parallelism.
colossalai/fx/tracer · high confidence
New tensor sharding utility functions for broadcast, reshape, and cost estimation
The \colossalai/auto\_parallel/tensor\_shard/utils\ package now exposes a comprehensive set of utilities to support the auto-parallel sharding strategy solver. This includes \broadcast.py\ for handling tensor broadcasting logic (checking broadcastability, computing shapes, and generating communication actions), \reshape.py\ for detecting reshape mappings and inferring output sharding specs, \factory.py\ for generating sharding specs and calculating resharding costs, and \sharding.py\ for enumerating 1D/2D sharding configurations and manipulating partition dimensions. These functions provide the core logic for the parallelizer to analyze and optimize tensor distributions across devices.
_colossalai/auto\_parallel/tensor\shard/utils · high confidence
New tensor-parallel inference engine for Llama, Bloom, and ChatGLM2
This change introduces a new tensor-parallel inference architecture in \colossalai/legacy/inference/tensor\_parallel\. It adds a \TPInferEngine\ that shards models using \ShardFormer\ and manages KV-cache memory via a new \MemoryManager\. The engine includes optimized inference forward passes for Llama, Bloom, and ChatGLM2 models, replacing standard transformer methods with custom implementations that support batched inference, context/decode stage separation, and optional Triton/lightLLM kernels for attention and normalization.
_colossalai/legacy/inference/tensor\parallel · high confidence
New tutorial examples for training ResNet and ViT on CIFAR-10 using the Booster API
Added new tutorial examples in \examples/tutorial/new\_api/cifar\_resnet\ and \examples/tutorial/new\_api/cifar\_vit\ that demonstrate how to train ResNet-18 and Vision Transformer (ViT) models on the CIFAR-10 dataset using the ColossalAI Booster API. These examples showcase the new plugin-based architecture, supporting distributed training strategies such as \torch\_ddp\, \torch\_ddp\_fp16\ (mixed precision), and \low\_level\_zero\. The tutorials include complete training and evaluation scripts, configuration for checkpointing and resuming, and performance benchmarks comparing single-GPU baselines against the various Booster plugins.
_examples/tutorial/new\_api/cifar\_resnet, examples/tutorial/new\_api/cifar\vit · high confidence
New tutorial for large-batch training with Lamb and Lars optimizers
Added a new tutorial example that demonstrates large-batch training optimization using the Lamb and Lars optimizers from Colossal-AI. The example uses a synthetic dataset and ResNet-18 model, allowing users to quickly experiment with these optimizers via command-line arguments without needing to prepare external data.
_examples/tutorial/large\_batch\optimizer · high confidence
New verifiable reward functions for math and code tasks
This change introduces a new reward calculation module in the ColossalChat Coati distributed reward system. It adds specific reward functions (\math\_reward\_fn\, \boxed\_math\_reward\_fn\) that verify model answers using LaTeX parsing and exact matching, including support for boxed math solutions. It also includes utility functions for validating response structures (e.g., XML-style tags) and extracting solutions. A \VerifiableReward\ class is added to orchestrate these functions, allowing for batched evaluation of rewards based on ground truth answers or test cases, supporting both training and evaluation modes with length-based penalties.
applications/ColossalChat/coati/distributed/reward · high confidence
Shardformer policy system restructured with auto-discovery and new model support
The shardformer policy system has been restructured to support automatic policy discovery via a new \auto\_policy\ module that maps Hugging Face model classes to specific policy implementations, eliminating the need for manual policy selection. This change introduces a standardized \Policy\ base class in \base\_policy.py\ that defines the interface for model optimization, including preprocessing, module replacement, and postprocessing steps. New policies have been added for a wide range of models including BERT, LLaMA, T5, GPT-2, GPT-J, ViT, OPT, Bloom, Whisper, SAM, BLIP-2, ChatGLM2, DeepSeek, DeepSeekV3, Falcon, and Mistral, each implementing tensor parallelism, sequence parallelism, and pipeline parallelism optimizations. The policy system now supports advanced features such as FP8 communication, fused normalization, JIT-optimized operations, and various sequence parallelism modes (split\_gather, all\_to\_all, ring).
colossalai/shardformer/policies · high confidence
Support for tracing modules with bias addition
The FX tracer now correctly handles PyTorch modules that include bias terms (Linear, Conv1d, Conv2d, Conv3d) by restructuring their computation graphs. Instead of tracing the module as a single opaque call, the tracer decomposes the operation into separate nodes for weight retrieval, bias retrieval, the core computation (e.g., convolution or linear), optional bias reshaping for convolutions, and the final bias addition. This ensures that the traced graph explicitly represents the bias addition step, which is critical for accurate parallelization and optimization.
_colossalai/fx/tracer/bias\_addition\_patch/patched\_bias\_addition\module · high confidence
Removals
Empty initialization module for colossalai.amp
The colossalai.amp package initialization file has been replaced with an empty file, effectively removing any public API exports or initialization logic previously defined in this module.
colossalai/amp · high confidence
Loss module initialization cleared
The \_\init\\_.py file for the colossalai.nn.loss package has been emptied, removing all public exports from this module. Users relying on direct imports from colossalai.nn.loss will no longer have access to any loss functions or classes through this path.
colossalai/nn/loss · high confidence
Architecture
Refactored meta-patched modules into a modular structure
The meta-patched module implementations for the FX tracer have been reorganized from a single file into a structured package with separate modules for activation functions, convolutions, embeddings, linear layers, normalization, pooling, and RNNs. This change improves code maintainability and clarity without altering the functional behavior of the tracer's handling of these PyTorch operations.
_colossalai/fx/tracer/meta\_patch/patched\module · high confidence
Behavioural changes
Introduce modular chunk management for Gemini ZeRO
The Gemini ZeRO memory management logic is reorganized into a dedicated \colossalai/zero/gemini/chunk\ package, exposing \Chunk\ and \ChunkManager\ classes alongside \search\_utils\ for automatic configuration. This change introduces a state-machine-driven approach to tensor lifecycle management (tracking states like FREE, COMPUTE, HOLD, and READY\_FOR\_REDUCE) and provides utilities to automatically search for optimal chunk sizes based on model parameter sizes and data-parallel degree, aiming to reduce memory waste and improve communication efficiency.
colossalai/zero/gemini/chunk · high confidence
Introduces modular attention and pre-attention backends for inference
The inference modeling layer now uses a structured backend system in \colossalai/inference/modeling/backends\ to handle attention and pre-attention operations (such as RoPE and KV cache management). This change introduces \AttentionBackend\ and \PreAttentionBackend\ abstractions with concrete implementations for CUDA kernels (\CudaAttentionBackend\, \CudaPreAttentionBackend\) and pure Triton kernels (\TritonAttentionBackend\). The system automatically selects the appropriate backend based on configuration flags like \use\_cuda\_kernel\, \use\_flash\_attn\, and \use\_spec\_dec\, ensuring that speculative decoding uses Triton kernels while standard inference can leverage optimized CUDA/Flash-Attention paths when available.
colossalai/inference/modeling/backends · high confidence
Legacy communication module relocated to legacy namespace
The distributed communication primitives (collective operations like all\_reduce, p2p transfers, and ring communication) have been moved into the \colossalai.legacy.communication\ package. This change isolates the legacy implementation from the main codebase, ensuring that users relying on these specific communication patterns are directed to the maintained legacy path while the primary library evolves separately.
colossalai/legacy/communication · high confidence
Legacy distributed training entry points moved to colossalai.legacy
The distributed training initialization functions (launch, initialize, get\_default\_parser) and related constants have been relocated to the colossalai.legacy package. Users relying on these legacy APIs for 1D/2D/3D parallelism and older engine implementations must now import from colossalai.legacy instead of the top-level colossalai namespace.
colossalai/legacy · high confidence
Legacy engine and gradient accumulation components moved to legacy package
The training engine and gradient accumulation utilities have been relocated to the \colossalai.legacy\ namespace. This includes the \Engine\ class for managing training loops and schedules, as well as the \accumulate\_gradient\ helper and its associated wrappers (\GradAccumOptimizer\, \GradAccumDataloader\, \GradAccumLrSchedulerByStep\) for handling gradient accumulation. Users relying on these specific legacy implementations must now import them from the \colossalai.legacy.engine\ path.
_colossalai/legacy/engine, colossalai/legacy/engine/gradient\accumulation · high confidence
Legacy engine schedule module relocated and refactored
The training and inference scheduling logic has been moved into the \colossalai.legacy.engine.schedule\ package. This change introduces a modular structure with a \BaseSchedule\ abstract class, a \NonPipelineSchedule\ for standard parallelism, and \PipelineSchedule\ (including a \PipelineScheduleV2\ variant) for pipeline parallelism. Users relying on these scheduling components should now import them from the legacy engine path, and the pipeline schedules now utilize the updated \p2p\_v2\ communication primitives in their v2 implementation.
colossalai/legacy/engine/schedule · high confidence
Legacy neural network layers moved to legacy.nn
The neural network layer implementations (including linear, embedding, normalization, and dropout layers) and their associated 1D/2D/3D parallelism operations have been relocated to the \colossalai/legacy/nn\ package. This change organizes the codebase by isolating these existing parallel layer implementations into a legacy namespace, while exposing them via the new \colossalai/legacy/nn\ module structure.
colossalai/legacy/nn · high confidence
Move builder and registry modules to legacy
The builder and registry components have been relocated to the \colossalai.legacy\ namespace. This change moves the \Registry\ class, its associated module definitions (such as LAYERS, MODELS, OPTIMIZERS), and the builder utilities (\build\_from\_config\, \build\_from\_registry\, \build\_gradient\_handler\) to the legacy package, indicating these are no longer part of the active public API.
colossalai/legacy/builder, colossalai/legacy/registry · high confidence
New FX tracer with bias-splitting and custom proxy support
The FX tracer in colossalai/\_analyzer/fx/tracer has been replaced with a new implementation that introduces ColoProxy and ColoTracer classes to better handle metadata during graph construction. This change adds support for splitting bias addition operations (e.g., in Linear and Conv layers) into separate weight and bias nodes, which improves compatibility with auto-parallel all-reduce operations. It also registers custom leaf modules for Apex normalization layers and enables tracing of torch.utils.checkpoint usage, allowing the analyzer to correctly process models with activation checkpointing.
_colossalai/\analyzer/fx/tracer · high confidence
PyTorch meta-tensor compatibility and data-dependent control flow support
The \colossalai.fx\ module now supports tracing models with data-dependent control flow (e.g., if-statements) by introducing \ColoProxy\ and \ColoGraphModule\, which leverage PyTorch meta-tensors to infer shapes and conditions without executing the model. To ensure compatibility across PyTorch versions, the module includes version-specific meta-registration files (\\_meta\_regist\_12.py\ for PyTorch ≤1.12 and \\_meta\_regist\_13.py\ for PyTorch 1.13+) and a compatibility check (\\_compatibility.py\) that validates the runtime environment before enabling these features.
colossalai/fx · high confidence
Refactored gradient scaler into modular base and implementations
The gradient scaler logic in the naive AMP module has been restructured into a modular design. A new \BaseGradScaler\ abstract class now provides common functionality for state management and logging, while specific strategies are implemented in \ConstantGradScaler\ (fixed scale) and \DynamicGradScaler\ (adaptive scale with growth/backoff factors, hysteresis, and min/max bounds). This change allows users to choose between constant or dynamic loss scaling strategies and ensures the scaler state can be properly serialized and restored.
_colossalai/amp/naive\_amp/grad\scaler · high confidence
Structured meta-patching for FX tracer operations
The FX tracer's meta-patching logic has been reorganized into a modular package under \colossalai/fx/tracer/meta\_patch/patched\_function\, replacing the previous flat structure. This change introduces dedicated modules for activation functions (e.g., ReLU), arithmetic operations (including \matmul\, \addbmm\, \addmm\, and \linear\), convolution operations (1D/2D/3D and their transposes), normalization (LayerNorm, BatchNorm), embedding, and core torch operations (such as \arange\, \where\, \cat\, \index\_select\, \squeeze\, \unsqueeze\, \roll\, \full\, \max\, and device moves like \cpu\/\cuda\). By explicitly registering these handlers via the \meta\_patched\_function\ registry, the tracer can now correctly infer meta-tensor shapes and types for a broader range of PyTorch operations during graph tracing, improving compatibility with complex model architectures.
_colossalai/fx/tracer/meta\_patch/patched\function · high confidence
Trainer and hook system moved to legacy module
The \Trainer\ class and its associated hook system (including logging, checkpointing, metrics, and LR scheduling) have been relocated to the \colossalai.legacy.trainer\ package. This change isolates the existing training workflow in a legacy namespace, signaling that these components are part of the older engine-based API rather than the current primary interface.
colossalai/legacy/trainer · high confidence
Test coverage
Add ColossalChat test infrastructure and coverage; Added CI smoke test for RoBERTa example; Added CI test runner for the OPT tutorial example; Added CI test script for the new API tutorial example; Added tests for checkpoint compatibility and watermark decoding; Added tests for document loading, memory management, and retrieval QA; Introduce colossalai.testing module for test utilities; Introduce standardized test kit and model zoo.
Dependencies
Standardize and pin dependency versions across applications and examples
This change introduces explicit requirements.txt files for all application directories (Colossal-LLaMA, ColossalChat, ColossalEval, ColossalMoE, ColossalQA) and example tutorials, replacing implicit or missing dependency declarations. It pins specific versions for core libraries such as PyTorch (e.g., \>=2.1.0, \>=2.2.0), Transformers (e.g., \>=4.39.3, ==4.51.3), and ColossalAI (e.g., \>=0.4.0, \>=0.3.4) to ensure compatibility and reproducibility. It also adds or updates dependencies for specific features like Flash Attention, vLLM, and LangChain, and includes test requirements for CI validation.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 50.
Lenses
- Code Health 56
- Architecture 95
- Maturity 54
- Readiness 40
- Security 51
- Performance 100
Changes since last survey
- 300 commits — 180 feature/other, 120 fixes
By area
- applications/ColossalChat — 125 commits
- colossalai/shardformer — 39 commits
- .github/workflows — 28 commits
- (root) — 25 commits
- (repo) — 22 commits
- colossalai/booster — 11 commits
- colossalai/checkpoint_io — 9 commits
- colossalai/zero — 8 commits
- tests/kit — 5 commits
- .github/scripts — 4 commits
- colossalai/utils — 4 commits
- tests/test_shardformer — 4 commits
- .github/ISSUE_TEMPLATE — 1 commit
- applications/Colossal-LLaMA — 1 commit
- applications/ColossalEval — 1 commit
- colossalai/cli — 1 commit
- colossalai/inference — 1 commit
- colossalai/legacy — 1 commit
- colossalai/pipeline — 1 commit
- colossalai/quantization — 1 commit
Notable commits
- fix: Merge pull request #6096 from BurkeHulk/hotfix/lora_ckpt
- fix: Merge pull request #6149 from ver217/hotfix/ckpt
- fix: Merge pull request #6336 from BurkeHulk/fix/update-test-config
- fix: Merge pull request #6378 from hpcaitech/grpo-latest-rebase-fix-resume
- fix: Merge pull request #6460 from hpcaitech/fix/empty-example-test-ci
- fix: [Fix] Add L2 Regularization (#6372)
- fix: [HotFix] update load lora model Readme; (#6240)
- fix: [Hotfix] hotfix normalization (#6163)
- fix: [Inference]Fix example in readme (#6178)
- fix: [checkpointio] fix async io (#6155)
- fix: [checkpointio] fix checkpoint for 3d (#6187)
- fix: [checkpointio] fix for async io (#6189)
- fix: [checkpointio] fix hybrid plugin model save (#6106)
- fix: [checkpointio] fix performance issue (#6139)
- fix: [checkpointio] fix pinned state dict
- fix: [checkpointio] fix size compute
- fix: [checkpointio] fix zero optimizer async save memory (#6151)
- fix: [extension] hotfix compile check (#6099)
- fix: [fix] fix bug caused by perf version (#6156)
- fix: [fix] fix_lazy_init for deepseek model in transformers
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
hpcaitech/ColossalAI was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 11 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit df01070248aa7d0ff2621545f630f18b505f8b9a — the exact code this score is about.
- Scored under rubric-2026.10.5 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fe8540b5da9b.