deepspeedai/DeepSpeed
62.7
Adequate · 26 September 2026
109.3k
lines of production code
Python
with C, C++
3
measurements over time
What this system is
DeepSpeed is a deep learning optimization library that accelerates training and inference for large-scale models through advanced parallelism and memory management techniques. It provides a unified framework for ZeRO optimization, pipeline and tensor parallelism, and efficient offloading of parameters and optimizer states to CPU and NVMe storage. The system also features a high-performance inference engine with support for quantization, Mixture-of-Experts architectures, and ragged batching, alongside comprehensive tooling for benchmarking, profiling, and multi-accelerator hardware abstraction.
How it got here
2020–2021 — Infrastructure modernization and kernel optimization
47 changes.
This period focused on establishing a robust project foundation through standardized configuration, governance, and comprehensive test coverage. It introduced significant architectural improvements, including a unified operator builder system, Pydantic-based configuration validation, and a dedicated communication backend. The work also delivered major performance enhancements via fused CUDA/ROCm kernels, CPU-optimized optimizers, and the initial implementation of ZeRO-Infinity NVMe offloading.
2022–2023 — Inference V2 and hardware abstraction
76 changes.
This period focused on developing the DeepSpeed Inference V2 engine, introducing ragged batching, FP6 quantization, and optimized CUDA kernels for MoE and transformer models. It also established a unified accelerator abstraction layer to support diverse hardware backends, including AMD, Huawei Ascend, and CPU, while significantly expanding unit test coverage for these new components.
2024–2026 — DeepCompile and hardware expansion
57 changes.
This period focused on introducing DeepCompile, a JIT compilation backend for ZeRO stages, alongside significant expansions in hardware support for XPU, HPU, MLU, SDAA, and Apple Silicon. The work also delivered advanced memory optimization features such as SuperOffload, ZenFlow, and DeepNVMe, while establishing robust sequence parallelism via Arctic Long Sequence Training and Ulysses.
Features
Add Apple Silicon (MPS) support for DeepSpeed Adam optimizers
Users on Apple Silicon can now utilize DeepSpeed's Adam optimizer variants on the Metal Performance Shaders (MPS) backend. This change introduces builders for a Metal-based FusedAdam kernel (with a PyTorch foreach fallback) and a CPU-based Adam kernel for ZeRO-Offload scenarios, enabling optimized training workflows on macOS devices without requiring CUDA.
_op\builder/mps · high confidence
Add Biren SUPA accelerator support
Introduces a new \op\_builder/supa\ module that provides operator builders and Python wrappers for the Biren SUPA GPU accelerator. This includes builders for fused optimizers (Adam, LAMB, Lion), CPU optimizers (Adam, Lion, Adagrad), async I/O, and quantization, as well as a comprehensive inference module wrapping transformer kernels (workspace management, normalization, softmax, bias/GELU/ReLU ops, and residual additions). The implementation delegates to \torch.ops.deepspeed\ kernels provided by the \torch\_supa\_ext\ extension, with pure-PyTorch fallbacks for testing or environments where the compiled kernels are unavailable.
_op\builder/supa · high confidence
Add CPU and XPU-optimized optimizer kernels for Adagrad and Adam
This change introduces new C++ implementation files for the Adagrad and Adam optimizers within the XPU component. The Adagrad implementation (cpu\_adagrad.cpp) provides a CPU-based optimizer with AVX/AVX512 SIMD acceleration and half-precision support. The Adam implementation (multi\_tensor\_adam.dp.cpp, fused\_adam\_frontend.cpp) provides a SYCL-based multi-tensor kernel for XPU execution, supporting both L2 regularization and decoupled weight decay (AdamW) modes. These files also include necessary headers (simd.h, type\_shim.h, compat.h) for SIMD operations, type dispatching, and PyTorch version compatibility.
csrc/xpu · high confidence
Add CPU-based Adagrad optimizer
Introduces a new \DeepSpeedCPUAdagrad\ optimizer class located in \deepspeed/ops/adagrad\, enabling Adagrad optimization steps to be executed on the CPU. This implementation leverages a custom C++ extension (\CPUAdagradBuilder\) and is designed to work with ZeRO-Offload by asserting that model parameters reside on the CPU device. It supports standard Adagrad arguments such as learning rate, epsilon, weight decay, and AMSGrad, and includes explicit resource cleanup to prevent memory leaks when the optimizer is instantiated multiple times within the same process.
deepspeed/ops/adagrad · high confidence
Add CPU-optimized Adam and Adagrad optimizers with SIMD and ZenFlow support
DeepSpeed now provides CPU-based implementations for the Adam and Adagrad optimizers, enabling training on machines without GPUs or offloading optimizer steps to the CPU. The CPU Adam implementation includes SIMD-accelerated kernels for AVX2, AVX-512, ARM NEON, and ARM SVE architectures, supports fp16/bf16 precision, and offers a ZenFlow mode that runs the optimizer in a separate native process to overlap computation with the Python training thread. The CPU Adagrad implementation similarly provides AVX2/AVX-512 vectorized updates. These components are exposed via C++ extensions and PyTorch bindings for integration into the DeepSpeed engine.
csrc/adam · high confidence
Add CPU-optimized Lion optimizer implementation
Users can now utilize the Lion optimizer on CPU devices through a new fused C++ extension. This change introduces \fused\_lion.cpp\, which implements the \multi\_tensor\_lion\ function to compute and apply gradient updates for parameters, exposing the functionality to Python via PyTorch's extension mechanism.
csrc/cpu/lion · high confidence
Add CPU-optimized optimizer kernels and CUDA utility headers
This change introduces new C++ header files in csrc/includes that provide CPU-optimized implementations for the Adam, Adagrad, and Lion optimizers, featuring SIMD acceleration (AVX, NEON, and ARM SVE) for training on CPU. It also adds utility headers for CUDA context management, timing (StopWatch, GPUTimer), data type conversions, and custom CUDA layers, alongside dequantization utilities to support mixed-precision and quantized training workflows.
csrc/includes · high confidence
Add Cambricon MLU operator builders for CPU-based optimizers
New builder classes are introduced in the \op\_builder/mlu\ module to support Cambricon MLU hardware, specifically enabling the build and loading of CPU-based optimizer implementations (CPU Adam, CPU Adagrad, and Fused Adam) via the \MLUOpBuilder\ base class. This addition allows the framework to recognize and compile these specific operator variants for MLU environments, while providing a placeholder for unimplemented operations.
_op\builder/mlu · high confidence
Add DeepSpeed Inference v2 kernel utility headers
This change introduces a set of new header files in the \deepspeed/inference/v2/kernels/includes\ directory to support the Inference v2 kernel implementation. The added headers provide foundational utilities including \activation\_type.h\ for defining activation functions, \conversion\_utils.h\ for type conversion templates (including BF16 support), \ds\_kernel\_utils.h\ for hardware-specific constants and macros (with ROCm warp size fixes), \memory\_access\_utils.h\ for optimized memory load/store operations, and \reduction\_utils.h\ for parallel reduction primitives. These headers collectively establish the low-level building blocks required for the new inference kernel architecture.
deepspeed/inference/v2/kernels/includes · high confidence
Add DeepSpeed4Science Evoformer attention kernels
This change introduces the C++ and CUDA implementation for the Evoformer attention mechanism within the DeepSpeed4Science module. It adds the PyTorch extension bindings in \attention.cpp\ and the core forward (\attention\_cu.cu\) and backward (\attention\_back.cu\) CUDA kernels, which utilize custom CUTLASS-based epilogues (including pipelined and logsumexp variants) and multi-architecture dispatch logic to support efficient attention computation with optional bias inputs.
csrc/deepspeed4science · high confidence
Add DeepSpeed4Science Evoformer attention operator
The \deepspeed/ops/deepspeed4science\ module now exposes a fused Evoformer attention implementation (\DS4Sci\_EvoformerAttention\ and \EvoformerFusedAttention\). This new capability allows users to leverage a custom CUDA-accelerated attention kernel for scientific workloads, specifically supporting the OpenFold training pipeline referenced in the associated fix commit.
deepspeed/ops/deepspeed4science · high confidence
Add HPU op builder components
Added the op\_builder/hpu module, introducing builders for CPUAdam, FusedAdam, transformer inference, and FP quantization to support the HPU accelerator.
_op\builder/hpu · high confidence
Add Inference V2 support for EXAONE 4.0, EXAONE 4.5, and Falcon models
DeepSpeed Inference V2 now includes model implementations for EXAONE 4.0, EXAONE 4.5, and Falcon (7B, 40B, 180B). The EXAONE 4.0 implementation supports post-normalization and QK-Norm architectures. The EXAONE 4.5 implementation extends this to handle hybrid attention with sliding windows and Llama 3-style RoPE scaling, while intentionally excluding the vision tower and MTP head for text-only serving. The Falcon implementation supports both the original architecture (using fused QKV) and the new decoder architecture (using GQA Megatron-style QKV). These additions are exposed via the \model\implementations\ package and registered in the \\\init\\_.py\ module.
_deepspeed/inference/v2/model\implementations · high confidence
Add Lion optimizer with CPU and fused GPU implementations
DeepSpeed now includes the Lion optimizer, offering both a CPU-based implementation (DeepSpeedCPULion) and a fused GPU kernel (FusedLion). The CPU version supports fp32 states and warns about potential issues with FP16 on AMD CPUs, while the fused GPU version accelerates training by supporting fp16, bf16, and fp32 parameter types via multi-tensor application.
deepspeed/ops/lion · high confidence
Add Random LTD token sampling and gathering utilities
New Python modules have been added to the \deepspeed.ops.random\_ltd\ package to support Random Layer Token Dropping (LTD). The \dropping\_utils.py\ file introduces functions \gpt\_sample\_tokens\ and \bert\_sample\_tokens\ for sampling token indices based on sequence length and batch size, along with \GatherTokens\ and \ScatterTokens\ autograd functions that interface with the compiled \RandomLTDBuilder\ kernel for efficient token reordering and gradient scattering. These utilities enable selective token processing in GPT and BERT architectures.
_deepspeed/ops/random\ltd · high confidence
Add XPU accelerator support for DeepSpeed op builder
Users can now build and load optimized C++ extensions for Intel XPU devices. This change introduces a new \op\_builder/xpu\ module containing a SYCL-based builder (\SYCLOpBuilder\) that compiles operations like CPU Adam, CPU Adagrad, Fused Adam, Async I/O, and Flash Attention for XPU. The builder ensures ABI compatibility by linking against the SYCL headers and libraries present in the Python environment rather than the system-wide oneAPI installation, and it handles JIT loading via the Intel \icpx\ compiler.
_op\builder/xpu · high confidence
Add fused LAMB optimizer kernel with ROCm support
This change introduces the fused LAMB (Layer-wise Adaptive Moments optimizer for Batch normalization) optimizer kernel for PyTorch. The implementation includes the C++ interface and CUDA/HIP kernel code, enabling optimized training performance. Notably, the kernel code includes conditional compilation logic to support both NVIDIA CUDA and AMD ROCm platforms, allowing the optimizer to run on AMD hardware in addition to NVIDIA GPUs.
csrc/lamb · high confidence
Add ragged embedding implementation for inference
A new ragged embedding module (DSRaggedEmbedding) is introduced in the inference v2 embedding implementations, supporting fp16, bf16, and fp32 residual dtypes. It leverages a RaggedEmbeddingKernel to process ragged batch inputs with optional position embeddings, enabling efficient embedding lookups for variable-length token sequences during inference.
deepspeed/inference/v2/modules/implementations/embedding · high confidence
Added CUTLASS-based MoE implementation for DeepSpeed FastGen
DeepSpeed FastGen now includes a new Mixture-of-Experts (MoE) inference implementation based on CUTLASS multi-GEMM operations. This addition, located in the v2 inference modules, introduces the \DSMultiGemmMoE\ class which supports fp16 and bfloat16 data types and top-k gating values of 1, 2, 4, and 8. It utilizes optimized kernels for gating, scattering, gathering, and MLP execution to accelerate MoE model inference.
deepspeed/inference/v2/modules/implementations/moe · high confidence
Added DeepSpeed inference v2 CUTLASS kernels for MoE and mixed-precision GEMM
This change introduces new CUDA kernels and Python bindings in the DeepSpeed inference v2 module to accelerate Mixture-of-Experts (MoE) and mixed-precision matrix multiplications. The update adds \MoEGEMM\ and \MixedMoEGEMM\ classes that leverage CUTLASS runners to support FP16/BF16 activations with FP16/BF16/FP8/FP4 weight variants, including optional bias and activation functions (GELU, SILU, RELU, IDENTITY). It also includes a base \DSKernelBase\ interface for kernel management and a fast host buffer utility for CUDA memory allocation.
(repo-wide) · high confidence
Added Metal FusedAdam kernel for Apple Silicon support
A new Metal shader implementation for the FusedAdam optimizer step has been added to the MPS backend, enabling hardware-accelerated training on Apple Silicon devices. The kernel mirrors the numerical behavior of the existing CUDA implementation by performing all math in fp32, ensuring consistent results for fp16 and bf16 parameters. It supports standard Adam and AdamW modes and includes specific instantiations for float and half-precision types, with conditional support for bfloat16 on Metal 3.1+ (macOS 14).
csrc/mps · high confidence
Added NVMe mount utilities and tensor casting helpers
This update introduces new utility scripts for managing NVMe storage on DGX-2 systems, specifically adding \dgx2\_mount\_nvme.sh\ and \dgx2\_umount\_nvme.sh\ to handle mounting and unmounting NVMe devices. Additionally, it adds C++ and Python bindings for tensor casting utilities (\tensor\_cast.cpp\, \tensor\_cast.h\, \py\_ds\_utils.cpp\) that allow casting 1D and multi-dimensional PyTorch tensors to byte tensors without data movement.
csrc/aio/utils, csrc/utils · high confidence
Added SDAA accelerator op builder support
Added the \op\_builder/sdaa\ module to enable building and loading optimized operators for the SDAA (Smart Data Accelerator Architecture) backend. This includes builders for CPU Adam and Fused Adam operations, allowing DeepSpeed to compile and utilize these specific kernel implementations when running on SDAA hardware.
_op\builder/sdaa · high confidence
Added optimized fused bias-add CUDA kernels and PyTorch bindings
This change introduces a new set of optimized CUDA kernels for fused bias-add operations in the spatial module, supporting three variants: standard bias addition, addition with a secondary tensor, and addition with two pairs of tensors and biases. These kernels are implemented in \opt\_bias\_add.cu\ and exposed to Python via new PyTorch bindings (\nhwc\_bias\_add\, \nhwc\_bias\_add\_add\, \nhwc\_bias\_add\_bias\_add\) in \pt\_binding.cpp\, enabling more efficient computation for half-precision tensors in deep learning models.
csrc/spatial · high confidence
CPU support for FusedAdam optimizer
Adds a new C++ implementation for the FusedAdam optimizer on CPU, exposing the \multi\_tensor\_adam\ function via PyTorch's C++ extension interface. This enables users to run the fused Adam optimization step on CPU hardware, mirroring the behavior available in the CUDA implementation, with support for different optimization modes (Adam vs AdamW) controlled by the \mode\ parameter.
csrc/cpu/adam · high confidence
DeepCompile introduces structured optimization passes with validation contracts and new memory offloading capabilities
The DeepSpeed compilation pipeline now uses a formalized pass system where optimization steps are defined by lightweight contracts that declare required and provided capabilities, ensuring schedules are validated before execution to prevent invalid pass combinations. This update adds several new compilation passes to improve memory efficiency and performance: activation offloading moves forward-graph activations to host memory to reduce GPU usage; optimizer state offloading (Adam states) and parameter offloading move heavy optimizer tensors and ZeRO-3 gathered parameters to host memory; a prefetch pass reorders and fuses all-gather operations to overlap communication with computation; and a selective gather pass persists only high-value parameters on the device during the final backward pass. Additionally, the system now supports long-context checkpointing with sequence-aware recomputation banning and integrates AutoSP (Sequence Parallelism) and AutoTP (Tensor Parallelism) compilation passes to handle sharding and collective operations within the compiled graph.
deepspeed/compile/passes · high confidence
GPUDirect Storage (GDS) support for DeepNVMe tensor swapping
This change introduces a new C++ and Python binding layer for GPUDirect Storage (GDS) within the DeepSpeed GDS component, enabling direct, high-performance swapping of optimizer tensors to and from NVMe storage devices. The implementation registers CUDA buffers with the cuFile driver to bypass the CPU, supporting both synchronous and asynchronous parallel read/write operations (pread/pwrite) with configurable queue depths and intra-operation parallelism. It also provides utilities for managing pinned device and CPU-locked tensors required for zero-copy I/O, effectively adding the underlying GDS capability that allows DeepNVMe to offload optimizer states directly to persistent storage.
csrc/gds · high confidence
Initial project configuration and contributor guidelines
The repository now includes foundational configuration files to standardize development workflows and governance. This adds \.clang-format\ and \.style.yapf\ for C++ and Python code formatting, \.flake8\ and \.pylintrc\ for Python linting, and a \.pre-commit-config.yaml\ to enforce these checks automatically before commits. It also introduces \AGENTS.md\ and \CLAUDE.md\ with specific rules for AI coding agents, \CONTRIBUTING.md\ detailing the submission and testing process, \GOVERNANCE.md\ outlining the Technical Steering Committee structure, \CODEOWNERS\ for PR review assignments, \CODE\_OF\_CONDUCT.md\ adopting the PyTorch Foundation standards, and \COMMITTERS.md\ listing the initial TSC members.
(repo-wide) · high confidence
Initial support for Huawei Ascend NPU operations
This change introduces the \op\_builder/npu\ module, enabling DeepSpeed to build and load custom C++ extensions for Huawei Ascend NPUs. It adds builders for key components including FusedAdam, CPU-based optimizers (Adam, Adagrad, Lion), asynchronous I/O, and inference-specific kernels (such as attention and layer norm). The implementation relies on the CANN toolkit and \torch\_npu\, requiring the \ASCEND\_HOME\_PATH\ environment variable to locate the necessary libraries and headers for compilation.
_op\builder/npu · high confidence
Introduce AutoEP for automatic expert parallelism in MoE models
DeepSpeed now includes AutoEP, a new capability in the module injection layer that automatically detects Mixture-of-Experts (MoE) architectures and configures expert parallelism. This feature supports automatic layer replacement, topology folding with tensor parallelism, and configurable communication backends (defaulting to standard collectives with an opt-in DeepEP path). It ships with built-in presets for major model families including DeepSeek V2/V3, Mixtral, and Qwen3.5-MoE, allowing users to enable expert parallelism without manual policy definitions.
_deepspeed/module\inject · high confidence
Introduce AutoTP training module with configurable tensor parallelism
Adds a new \deepspeed.runtime.tensor\_parallel\ package that provides the \TpTrainingManager\ and configuration models for DeepSpeed's Automatic Tensor Parallelism (AutoTP) in training mode. This enables users to split models across GPUs using tensor parallelism with support for custom partitioning patterns, HuggingFace \tp\_plan\ integration, and optional vocabulary-parallel LM heads with cross-entropy acceleration via the Liger backend.
_deepspeed/runtime/tensor\parallel · high confidence
Introduce DataStates-LLM asynchronous checkpointing engine
DeepSpeed now supports an opt-in asynchronous checkpointing engine called DataStates-LLM. Users can enable this feature by adding a 'datastates\_ckpt' block to their configuration file and downloading the external DataStates-LLM library. This addition introduces new configuration classes and runtime modules to handle the integration, providing an alternative to the default synchronous checkpointing mechanism for potentially improved performance during checkpoint saves.
deepspeed/runtime · high confidence
Introduce DeepCompile C++ backend for ZeRO-1, ZeRO-2, and ZeRO-3
This change adds the C++ implementation for the DeepCompile compiler integration, providing the core execution logic for ZeRO-1, ZeRO-2, and ZeRO-3 optimization stages. The new files in \csrc/compile\ define custom PyTorch operators (via \TORCH\_LIBRARY\_IMPL\) for parameter all-gathering, gradient reduction, and activation offloading/reloading. The implementation includes specific executors for each ZeRO stage (\z1.cpp\, \z2.cpp\, \z3.cpp\) that handle NCCL communication, symmetric memory usage, and CUDA stream management, along with initialization and utility functions to support these distributed training optimizations.
csrc/compile · high confidence
Introduce DeepCompile profiling infrastructure
Added a new profiling subsystem for DeepCompile that captures forward and backward graph traces, memory usage, execution times, and tensor sizes. The module includes a communication profiler that benchmarks collective operations (such as all-gather) to build performance predictors, and a graph profiler that instruments FX graphs with detailed metadata while handling distributed profiling failures robustly across ranks.
deepspeed/compile/profilers · high confidence
Introduce DeepCompile, a new JIT compilation backend for DeepSpeed ZeRO
Adds the \deepspeed.compile\ package, providing a new compilation mode (\deepcompile\) that integrates with PyTorch's AOTAutograd and Inductor to optimize ZeRO-1, ZeRO-2, and ZeRO-3 training graphs. This feature introduces configurable optimization passes (including AutoSP and AutoTP), activation offloading to pinned host memory, and specific workarounds for PyTorch Inductor reduction heuristics to prevent kernel resource limit errors in ZeRO-3. Users can enable this via the \compile.offload\_activation\ and \compile.passes\ configuration options to potentially reduce memory usage and improve throughput for large-scale models.
deepspeed/compile · high confidence
Introduce DeepNVMe performance benchmarking and validation tools
Adds a new \deepspeed/nvme\ module containing scripts and utilities for benchmarking and validating asynchronous I/O performance on NVMe storage devices. This includes a performance sweep runner (\perf\_run\_sweep.py\) to test various I/O configurations, a parameter generator (\perf\_generate\_param.py\) to optimize settings based on logs, and validation scripts (\validate\_async\_io.py\, \test\_ds\_aio.py\) to ensure the AsyncIO and GDS builders are correctly configured. The module also provides specific I/O engine implementations (\ds\_aio\_handle.py\, \ds\_aio\_basic.py\, \torch\_io.py\, \torch\_fastio\_engine.py\) and argument parsing logic (\ds\_aio\_args.py\) to facilitate these tests.
deepspeed/nvme · high confidence
Introduce DeepSpeed Async I/O library with NVMe write warnings and CPU bounce buffers
This change adds the DeepSpeed Async I/O (AIO) C++/Python bindings in \csrc/aio/py\_lib\, exposing an \aio\_handle\ class and \aio\_read\/\aio\_write\ functions for swapping optimizer tensors to/from NVMe storage. The implementation includes a process-wide warning that fires once per process when writing to NVMe, alerting users that consumer SSDs may suffer reduced lifespan from heavy write traffic. It also introduces CPU bounce buffers for non-pinned tensors, ensuring data is copied to a pinned buffer before I/O, and supports file offsets for partial reads/writes. The library uses a multi-threaded worker model with configurable block size, queue depth, and intra-op parallelism, and integrates with DeepSpeed's pinned-tensor manager for efficient memory handling.
_csrc/aio/py\lib · high confidence
Introduce DeepSpeed Autotuning feature
DeepSpeed now includes an autotuning module that automatically discovers optimal training configurations, such as ZeRO optimization stages, micro-batch sizes, and other ZeRO settings, to maximize throughput. Users can enable this feature by setting \"autotuning": {"enabled": true}\ in their DeepSpeed configuration file and passing the \--autotuning=\[run\|tune\]\ flag to the launcher. The system profiles the model and hardware to explore a search space of configurations, allowing users to customize the tuning scope, metrics (throughput, latency, FLOPS), and early stopping criteria via the configuration JSON.
deepspeed/autotuning · high confidence
Introduce DeepSpeed Data Efficiency Library for curriculum learning and data analysis
Adds a new data efficiency library under \deepspeed/runtime/data\_pipeline/data\_sampling\ that enables curriculum learning and distributed data analysis. This includes a \DataAnalyzer\ for computing metrics across datasets, a \DeepSpeedDataSampler\ that supports curriculum learning schedules based on difficulty metrics, and utilities for variable batch sizing and learning rate scaling. The library also provides an \MMapIndexedDataset\ for efficient memory-mapped data access and distributed map-reduce operations for analyzing large-scale datasets.
_deepspeed/runtime/data\_pipeline/data\sampling · high confidence
Introduce DeepSpeed Flops Profiler for model performance analysis
Adds the DeepSpeed Flops Profiler, a new tool that measures the latency, floating-point operations (FLOPS), and parameter counts of PyTorch models. The profiler provides detailed breakdowns of performance metrics per module and at different model depths, helping users identify bottlenecks and optimize training or inference efficiency. It can be used as a standalone package or integrated into the DeepSpeed runtime configuration.
_deepspeed/profiling/flops\profiler · high confidence
Introduce DeepSpeed GDS operation builder
A new entry point for the DeepSpeed GDS (GPUDirect Storage) operations has been added. The file initializes the GDS builder, enabling the framework to construct and integrate the underlying GDS-specific operations.
deepspeed/ops/gds · high confidence
Introduce DeepSpeed Inference V2 engine with ragged batching and FP6 quantization
This change introduces the new \deepspeed.inference.v2\ module, providing a modern inference engine that supports ragged batching for more efficient processing of variable-length sequences. The engine includes a factory (\build\_hf\_engine\) that automatically detects and applies specific inference policies for a wide range of models, including LLaMA, Mistral, Mixtral, Falcon, Phi, Qwen, and EXAONE variants. It also introduces support for FP6 weight-only quantization (stored as FP16 with FP16 activations) and allows loading engines from either HuggingFace model paths or DeepSpeed checkpoints.
deepspeed/inference/v2 · high confidence
Introduce DeepSpeed autotuning tuner module
Added a new \deepspeed.autotuning.tuner\ package that provides a framework for automatically optimizing DeepSpeed configurations. This includes a \BaseTuner\ class for managing experiment trials and early stopping, alongside specific implementations: \RandomTuner\ and \GridSearchTuner\ for exhaustive or random search strategies, and \ModelBasedTuner\ which uses an XGBoost cost model to predict and select promising configurations. The module also includes utilities for converting configuration dictionaries into feature vectors and managing experiment resources.
deepspeed/autotuning/tuner · high confidence
Introduce DeepSpeed distributed training launcher
Adds the \deepspeed/launcher\ module, providing the command-line entry points (\deepspeed\ runner and \launch.py\) and backend implementations (PDSH, OpenMPI, MPICH, SLURM, IMPI, MVAPICH) required to start multi-node and multi-GPU distributed training jobs. This includes logic for parsing hostfiles, managing environment variables, handling process trees, and supporting features like core binding and per-rank logging.
deepspeed/launcher · high confidence
Introduce FP16 Optimizer with Dynamic Loss Scaling and Fused Kernels
DeepSpeed now provides a dedicated FP16 optimizer module that enables mixed-precision training with dynamic loss scaling and fused kernel support. The new FP16\_Optimizer and FP16\_UnfusedOptimizer classes manage the conversion of model parameters to float16 while maintaining master weights in float32, automatically handling gradient scaling and overflow detection to prevent numerical instability. This implementation includes optimized fused Adam and LAMB steps, configurable loss scale profiles (fused vs. unfused), and integration with DeepSpeed's existing infrastructure for gradient clipping and checkpointing, allowing users to train with reduced memory footprint and improved throughput without manual loss scale management.
deepspeed/runtime/fp16 · high confidence
Introduce Muon optimizer with Gram Newton-Schulz orthogonalization
DeepSpeed now includes the Muon optimizer, available as \MuonWithAuxAdam\ in \deepspeed.runtime.zero.muon\. This optimizer integrates Newton-Schulz orthogonalization to improve gradient updates, specifically introducing a Gram Newton-Schulz method that operates on the smaller Gram matrix for better performance on high-aspect-ratio matrices. The implementation supports an auxiliary Adam optimizer for parameters not using Muon, handles ZeRO partitioning logic, and allows configuration of orthogonalization methods and momentum buffers.
deepspeed/runtime/zero/muon · high confidence
Introduce SuperOffload optimizer for ZeRO-Stage 3
Adds the SuperOffload optimizer module for DeepSpeed ZeRO-Stage 3, which offloads optimizer state and computation to a separate CPU process to reduce GPU memory pressure. The new \superoffload\ package includes \SuperOffloadOptimizer\_Stage3\, which extends the standard ZeRO-3 optimizer to manage parameter partitioning and gradient accumulation, and \SuperOffloadCPUOptimizer\, which runs \DeepSpeedCPUAdam\ steps in a background worker process via multiprocessing queues. This allows users to enable CPU-based optimization for large models without increasing GPU VRAM usage during the optimizer step.
deepspeed/runtime/superoffload · high confidence
Introduce Triton-based inference kernels for attention, MLP, and primitives
Adds a new set of Triton-accelerated kernels for DeepSpeed inference, including self-attention, MLP, layer normalization, GELU, softmax, residual addition, and matrix multiplication. These kernels replace the previous non-Triton attention path and provide optimized, autotuned implementations for transformer inference operations, with support for caching autotune results and handling NFS paths for the cache directory.
deepspeed/ops/transformer/inference/triton · high confidence
Introduce ZeRO-Infinity NVMe offload with async swapping and GDS support
DeepSpeed now supports ZeRO-Infinity, allowing optimizer and parameter tensors to be offloaded to NVMe storage to extend memory capacity. This change introduces a new \deepspeed.runtime.swap\_tensor\ module containing asynchronous swappers (\AsyncTensorSwapper\, \PartitionedOptimizerSwapper\, \PipelinedOptimizerSwapper\) that manage data movement between GPU/CPU memory and NVMe. The implementation includes configurable AIO parameters (block size, queue depth, intra-op parallelism) and adds support for NVIDIA GPUDirect Storage (GDS) to accelerate I/O on CUDA devices. It also handles tensor alignment, buffer management, and gradient swapping to enable efficient large-scale model training.
_deepspeed/runtime/swap\tensor · high confidence
Introduce ZenFlow selective optimizer with native process offload
Adds the ZenFlow optimization module, enabling selective gradient updates to reduce optimizer memory and compute overhead. The feature introduces a \ZenFlowConfig\ for tuning parameters like \topk\_ratio\, \select\_strategy\, and \update\_interval\, and integrates with ZeRO stages 1, 2, and 3. A key behavioral change is the execution of the overlapped CPU optimizer in a dedicated native process (coordinated via shared-memory semaphores) rather than a Python thread, which keeps Adam state NUMA-local to optimizer cores and avoids Python IPC overhead. The implementation requires PyTorch 2.1+ and enforces CPU offload for ZeRO stages 1 and 2.
deepspeed/runtime/zenflow · high confidence
Introduce activation checkpointing module with CPU offload support
DeepSpeed now includes a dedicated activation checkpointing module that reduces GPU memory consumption by partitioning, offloading, or storing activations contiguously. The module provides configuration options for partitioned activations, CPU checkpointing, contiguous memory optimization, and profiling, and introduces an asynchronous CPU offload engine that manages a pinned buffer pool and keep-last logic to minimize data transfer overhead during backward propagation.
_deepspeed/runtime/activation\checkpointing · high confidence
Introduce checkpoint engine abstraction for Inference V2
The Inference V2 checkpoint module now provides a structured engine interface for loading model parameters. This includes a base abstract class defining the parameter iteration contract, a HuggingFace engine that handles downloading and loading checkpoints (prioritizing safetensors when available), and an in-memory engine for loading parameters from an already instantiated PyTorch model.
deepspeed/inference/v2/checkpoint · high confidence
Introduce common asynchronous I/O library for NVMe tensor swapping
Added a new shared asynchronous I/O (AIO) implementation in csrc/aio/common to handle swapping optimizer tensors to and from NVMe storage devices. This new codebase provides the core infrastructure for DeepSpeed's ZeRO-Offload and ZeRO-Infinity features, enabling efficient, non-blocking data movement between CPU memory and persistent storage to reduce GPU memory pressure during training.
csrc/aio/common · high confidence
Introduce deepspeed.comm as a unified communication backend
DeepSpeed now provides a new \deepspeed.comm\ package that acts as a unified wrapper around underlying communication backends (such as NCCL, MPI, Gloo, and CCL). This module exposes a public API designed to be fully compatible with \torch.distributed\, allowing users to replace \import torch.distributed as dist\ with \from deepspeed import comm as dist\ without breaking existing code. The implementation includes a backend abstraction layer (\backend.py\), specific backend handlers like \CCLBackend\ and \TorchBackend\, and optional performance features such as SDMA-based allgather via the \mori\ library (opt-in via \DS\_SDMA\_ALLGATHER\) and a configurable communication logger for profiling and debugging collective operations.
deepspeed/comm · high confidence
Introduce dense blocked attention implementation for DeepSpeed FastGen
A new \DSDenseBlockedAttention\ module has been added to the DeepSpeed inference v2 attention implementations. This component provides a dense, blocked self-attention mechanism optimized for fp16 and bf16 data types with causal masking. It supports both standard and rotary positional embeddings (including trained frequencies) and utilizes a blocked KV-cache structure to improve memory efficiency and performance during inference.
deepspeed/inference/v2/modules/implementations/attention · high confidence
Introduce diff-driven test selection and Modal-based CI execution
The CI system now uses a new test selector engine to run only the tests affected by code changes, significantly reducing CI runtime and resource usage. This is powered by a new Modal-based infrastructure that executes tests inside secure, isolated GPU sandboxes, ensuring that untrusted pull-request code is never run on the main CI runners. The change includes a new \ci/tests\_fetcher.py\ module for analyzing git diffs, a \ci/torch\_latest.py\ controller for orchestrating Modal sandbox execution, and a \ci/accelerate.py\ module for running external integration tests.
ci · high confidence
Introduce native DeepSpeed host-memory pinning backend
Added a new native backend for page-locked CPU memory management in the pin\_memory module, implemented via new C++ sources (deepspeed\_pin\_tensor, page\_alloc) and exposed through a Python extension (py\_ds\_pin\_memory). This provides a dedicated mechanism for allocating, freeing, and checking page-locked tensors independent of other I/O libraries, supporting accelerators that require host memory pinning.
_csrc/pin\memory · high confidence
Introduce pipeline parallelism support
DeepSpeed now supports pipeline parallelism, allowing models to be partitioned across multiple devices to reduce per-device memory usage and enable training of larger models. This change introduces the \PipelineModule\ for defining layer-wise partitions, the \PipelineEngine\ to manage the hybrid parallel training loop (combining pipeline, data, and model parallelism), and a scheduling system (\TrainSchedule\, \InferenceSchedule\) to orchestrate micro-batch execution. It also includes peer-to-peer communication utilities for passing activations and gradients between pipeline stages, and a \ProcessTopology\ class to manage the mapping of process ranks to parallel dimensions. Note that pipeline parallelism is currently incompatible with ZeRO-2 and ZeRO-3.
deepspeed/runtime/pipe · high confidence
Introduce ragged batching support for inference
Adds a new \deepspeed/inference/v2/ragged\ module that enables ragged batching for text generation. This includes a \DSStateManager\ for tracking sequences, a \BlockedKVCache\ with a \BlockedAllocator\ for efficient, block-based KV-cache memory management, and a \RaggedBatchWrapper\ to construct and manage batches of sequences with varying lengths. Configuration is handled via Pydantic-based models (\DSStateManagerConfig\, \KVCacheConfig\, \MemoryConfig\) allowing users to tune allocation modes, block sizes, and sequence limits.
deepspeed/inference/v2/ragged · high confidence
Introduces a unified accelerator abstraction layer
DeepSpeed now provides a standardized \DeepSpeedAccelerator\ interface that unifies support for diverse hardware backends. This change introduces dedicated implementations for CUDA, CPU, Apple Silicon (MPS), Habana HPU, Cambricon MLU, Ascend NPU, Intel XPU, Tecorigin SDAA, and Biren SUPA devices. The new \real\_accelerator\ module handles automatic device detection and allows users to explicitly override the target hardware via the \DS\_ACCELERATOR\ environment variable, ensuring consistent behavior across different accelerator types.
accelerator · high confidence
Introduces custom PyTorch ops for Sequence and Tensor Parallelism collectives
Adds a new \deepspeed.compile.custom\_ops\ module providing custom PyTorch operators (\autosp::all\_to\_all\, \autotp::copy\_to\_tp\_region\, \autotp::reduce\_from\_tp\_region\, \autotp::gather\_from\_tp\_region\) and their fake/autograd implementations to support sequence parallel (SP) and tensor parallel (TP) communication patterns. The module includes a compatibility check requiring PyTorch \>= 2.9 and a registry for managing SP/DP process groups, enabling efficient all-to-all shuffling for SP and all-reduce/all-gather operations for TP, including support for uneven TP shard widths.
_deepspeed/compile/custom\ops · high confidence
Introduction of Arctic Long Sequence Training (ALST) and Ulysses Sequence Parallelism
DeepSpeed introduces a new sequence parallelism module (\deepspeed/runtime/sequence\_parallel\) implementing Arctic Long Sequence Training (ALST) and Ulysses SP for Hugging Face Transformers. This addition provides \UlyssesSPAttentionHF\ for efficient long-sequence attention, \SequenceTiledCompute\ and \TiledMLP\ for memory-efficient MLP computation, and \TiledFusedLogitsLoss\ for loss calculation without materializing full logits. The module also includes \UlyssesSPDataLoaderAdapter\ for sharding data batches and manages dedicated sequence parallel process groups via \parallel\_state\_sp.py\, enabling scalable training for multi-million token sequences.
_deepspeed/runtime/sequence\parallel · high confidence
Introduction of AsyncIO operations module
A new module for asynchronous I/O operations has been added to the DeepSpeed library. This change introduces the \AsyncIOBuilder\ from the \op\_builder\ package, making asynchronous I/O capabilities available to users within the \deepspeed.ops.aio\ namespace.
deepspeed/ops/aio · high confidence
Introduction of DeepCompile module for compiler integration
A new \deepspeed.ops.compile\ module has been added to the codebase. This module serves as the entry point for DeepCompile functionality, initializing the \DeepCompileBuilder\ from the existing op builder infrastructure to facilitate enhanced compiler integration.
deepspeed/ops/compile · high confidence
Introduction of Domino cross-layer overlapping optimization
DeepSpeed introduces the Domino runtime module, which implements cross-layer overlapping to improve training performance. This change adds new PyTorch modules (\DominoAsyncColumnParallelLinear\, \RowParallelLinearNoComm\) and a transformer implementation (\DominoModule\, \ShardedAttention\) that utilize asynchronous communication operations. These components allow computation and communication to overlap across layers, specifically handling tensor-parallel linear layers and attention mechanisms with asynchronous all-reduce operations during the backward pass.
deepspeed/runtime/domino · high confidence
Introduction of FusedLamb optimizer and DeepSpeed Transformer ops
This change introduces the FusedLamb optimizer (deepspeed/ops/lamb) and the DeepSpeed Transformer operations (deepspeed/ops/transformer). The FusedLamb implementation provides a GPU-accelerated optimizer supporting layer-wise adaptive moments for large batch training, with configurable parameters like max/min coefficients and bias correction. The Transformer ops module exposes DeepSpeedTransformerLayer and DeepSpeedTransformerConfig, enabling users to utilize optimized transformer kernels with features such as FP16 support, pre/post-layer norm configurations, stochastic mode for performance, and configurable intermediate sizes. These components are new additions to the library, expanding available optimization and inference capabilities.
deepspeed/ops/transformer · high confidence
Introduction of new DeepSpeed CLI tools and launcher scripts
This change introduces a suite of new command-line interface scripts in the bin directory to enhance user interaction with DeepSpeed. The primary entry point is the \ds\ script (with \deepspeed\ and \deepspeed.pt\ symlinks), which launches the DeepSpeed runner using Python 3. New utility commands include \ds\_report\ for generating environment diagnostics, \ds\_bench\ for running communication benchmarks, \ds\_io\ for NVMe I/O operations, and \ds\_nvme\_tune\ for performance tuning of DeepNVMe storage. Additionally, \ds\_ssh\ is added to facilitate distributed training across hosts via SSH with configurable hostfiles, and Windows batch wrappers (\deepspeed.bat\, \ds\_report.bat\) are provided for cross-platform compatibility.
bin · high confidence
Introduction of the deepspeed.ops module
A new \deepspeed.ops\ package has been added to the library, providing a unified entry point for accessing DeepSpeed's optimized operations. This module exposes key components such as the Adam, Adagrad, Lamb, and Lion optimizers, as well as the DeepSpeed Transformer Layer and Configuration classes. It also includes a \fp\_quantizer\ and integrates with the \git\_version\_info\ to expose compatible operations, effectively centralizing access to these core computational kernels.
deepspeed/ops · high confidence
Modular Checkpoint Engine Architecture
DeepSpeed introduces a modular Checkpoint Engine system that abstracts checkpoint serialization, allowing users to switch between different backend implementations via configuration. The new architecture includes a base \CheckpointEngine\ interface and three concrete implementations: \TorchCheckpointEngine\ for standard PyTorch saving, \FastCheckpointEngine\ for optimized serialization with AIO support, and \DecoupledCheckpointEngine\ which offloads checkpointing to a separate process to improve reliability and prevent deadlocks. Additionally, support for the \DataStatesCheckpointEngine\ is added for asynchronous checkpointing, with automatic fallback to the standard engine if the DataStates library is not installed.
_deepspeed/runtime/checkpoint\engine · high confidence
New AutoEP expert parallelism implementation with optimized routing and computation
The \deepspeed/moe\ package now includes a new AutoEP (Auto Expert Parallelism) subsystem that introduces a token-choice top-K router (\ep\_router.py\) supporting node-limited routing to reduce cross-node communication, and a \GroupedExperts\ module (\ep\_experts.py\) that executes SwiGLU expert MLPs via Triton grouped GEMM kernels or \torch.\_grouped\_mm\ for improved performance on Ampere/Ada GPUs. This change also adds utilities for expert weight repacking (\ep\_repack.py\) to convert HuggingFace expert formats into grouped tensors, token permutation index generation (\ep\_kernels.py\) with Triton acceleration, and dispatch helpers (\ep\_tp\_dispatch.py\) for AutoEP + AutoTP parallel folding, alongside a new \Experts\ container (\experts.py\) and fused layout classification (\fused\_expert\_layout.py\).
deepspeed/moe · high confidence
New AutoSP module for multimodal sequence parallelism
Introduces the \deepspeed.sequence\ package to enable sequence parallelism for multimodal models. This includes an \auto\_wrap\_model\_for\_sp\ utility that automatically detects and wraps Vision Transformer (ViT) encoder attention layers with \UlyssesSPViTAttention\ to reduce memory usage, while providing adapters (e.g., \LlavaFusionAdapter\) to handle the sequence scatter/gather at the vision-language projection boundary. The module also adds vocabulary-parallel cross-entropy loss support and a new FPDT (Flash Parallel Distributed Training) attention layer.
deepspeed/sequence · high confidence
New CPU and ZenFlow Adam optimizer implementations
This change introduces new optimizer classes in the \deepspeed/ops/adam\ module: \DeepSpeedCPUAdam\ for fast CPU-based optimization (supporting FP16/BF16 parameters and configurable precision for optimizer states), \ZenFlowCPUAdam\ for overlapped CPU optimization steps, and \ZenFlowSelectiveAdamW\ for selective parameter updates with optional CPU offloading. It also includes \FusedAdam\ for GPU-accelerated training with multi-tensor apply support. These additions enable ZeRO-Offload and ZenFlow features by providing specialized optimizer backends that handle parameter offloading and selective gradient updates efficiently.
deepspeed/ops/adam · high confidence
New CPU op builder module with FusedAdam, CPUAdam, and AsyncIO support
A new \op\_builder/cpu\ package has been introduced to manage the compilation of CPU-specific DeepSpeed operations. This module provides builders for FusedAdam and CPUAdam (enabling CPU training with mixed precision support), AsyncIO (leveraging libaio for non-blocking I/O), and communication primitives like CCLComm and ShareMemComm. It also includes a PinMemory builder and a fallback NotImplementedBuilder, allowing users to build and utilize these CPU-optimized components directly.
_op\builder/cpu · high confidence
New CPU shared-memory communication backend with multi-architecture vectorization
A new shared-memory (SHM) based communication backend for CPU inference has been added, introducing low-latency all-reduce operations for single-node scenarios. The implementation includes architecture-specific optimized kernels for x86\_64 (AVX-512), ARM64 (NEON), and RISC-V (RVV), enabling efficient vectorized reduction of BF16, FP16, and FP32 data types. The backend registers \deepspeed::inference\_all\_reduce\ as a native CPU operator, allowing users to perform collective communication directly on CPU tensors without falling back to other backends.
csrc/cpu/comm · high confidence
New CUDA inference kernel library for core operations
This change introduces a new \core\_ops\ module within the DeepSpeed inference v2 kernels, providing optimized CUDA implementations for fundamental operations. It includes fused bias-activation kernels (supporting GELU, RELU, SILU, and IDENTITY with FP16/BF16), BLAS-based linear and 4D matrix multiplication layers, and fused LayerNorm variants (standard, pre-LayerNorm, and post-LayerNorm). Additionally, it adds support for Wf6Af16 (FP6 weight-only quantized) linear layers and RMSNorm, all exposed via a unified Python API and C++/CUDA bindings.
_deepspeed/inference/v2/kernels/core\ops · high confidence
New CUDA kernel implementations for inference operations
The inference engine now includes a new set of optimized CUDA kernels for core transformer operations, including rotary positional embeddings, dequantization, GELU, LayerNorm, ReLU, RMSNorm, Softmax, and pointwise vector addition. These kernels, exposed via a new PyTorch C++ binding (pt\_binding.cpp), provide the low-level compute primitives required for the Hybrid Engine's inference path, supporting float, half, and bfloat16 data types.
csrc/transformer/inference/csrc · high confidence
New CUDA-based quantization and dequantization kernels
This change introduces a new set of CUDA kernels for quantization and dequantization operations, supporting both 4-bit and 8-bit precision with symmetric and asymmetric quantization types. The implementation includes optimized kernels for standard quantization, dequantization, and swizzled quantization (which reorganizes quantized groups to facilitate better communication in distributed settings). It also provides PyTorch bindings (pt\_binding.cpp) to expose these operations to Python, allowing users to perform quantization and dequantization directly on CUDA tensors. Additionally, a new kernel for quantized reduction is included, enabling efficient reduction operations on quantized data.
csrc/quantization · high confidence
New CometML monitoring integration
DeepSpeed Monitor now supports logging to CometML. Users can enable this by setting the \comet\ configuration block in their DeepSpeed config, providing parameters such as \api\_key\, \project\, \workspace\, and \experiment\_name\. The integration handles experiment initialization, metric logging at specified sample intervals, and offline/online modes, complementing the existing TensorBoard, Weights & Biases, and CSV monitors.
deepspeed/monitor · high confidence
New DeepSpeed IO subsystem with FastFileWriter and buffer abstractions
Introduces a new IO module in deepspeed/io providing a suite of file writers (FastFileWriter, PyFileWriter, MockFileWriter) and IO buffer implementations (Single\_IO\_Buffer, Double\_IO\_Buffer) to handle high-performance tensor storage. FastFileWriter utilizes asynchronous I/O with pinned tensors and double-buffering for optimized write speeds, while PyFileWriter and MockFileWriter offer standard and testing alternatives. The module includes configuration classes, constants for statistics tracking, and utility functions for tensor serialization compatible with different PyTorch versions.
deepspeed/io · high confidence
New DeepSpeed Inference Transformer Ops package with MoE and Diffusers support
DeepSpeed introduces a new \deepspeed.ops.transformer.inference\ package that consolidates transformer inference kernels. This update adds native support for Mixture-of-Experts (MoE) models via \DeepSpeedMoEInference\ and \DeepSpeedMoEInferenceConfig\, and provides optimized inference layers for Stable Diffusion (Diffusers) through \DeepSpeedDiffusersAttention\ and \DeepSpeedDiffusersTransformerBlock\. The package also includes a Triton-based flash attention implementation (\triton\_flash\_attn\) for faster attention computation, alongside refactored standard attention (\DeepSpeedSelfAttention\) and MLP (\DeepSpeedMLP\) modules with updated configuration options like \rope\_theta\ and \scale\_attn\_by\_inverse\_layer\_idx\.
deepspeed/ops/transformer/inference · high confidence
New DeepSpeed Linear module with LoRA and quantization support
A new \deepspeed.linear\ package has been introduced, providing an \OptimizedLinear\ layer that supports Low-Rank Adaptation (LoRA) with base weight sharding and FP6/8/12 quantization. The module includes configuration dataclasses (\LoRAConfig\, \QuantizationConfig\) and a context manager (\Init\) that allows users to inject these optimized layers during model construction (e.g., via \transformers.AutoModelForCausalLM.from\_pretrained\), automatically handling LoRA initialization and weight sharding to reduce memory usage.
deepspeed/linear · high confidence
New DeepSpeed model implementations for Diffusers with CUDA graph support
This change introduces a new folder structure under \deepspeed/model\implementations\ to isolate model-specific code, specifically adding DeepSpeed wrappers for the Diffusers library's UNet and VAE components. These new implementations (\DSUNet\ and \DSVAE\) integrate with the \CUDAGraph\ feature, allowing users to enable CUDA graph capture and replay for these models to potentially improve inference performance. The entry also includes the base \CUDAGraph\ abstract class in the \features\ module and updates the \\\init\\_.py\ files to expose these new components.
_deepspeed/model\implementations/diffusers · high confidence
New DeepSpeed rollout engine for on-policy generation
A new \deepspeed/runtime/rollout\ module introduces a structured interface for on-policy generation during RL/distillation training. It provides a \RolloutEngine\ abstract base class and a concrete \HybridEngineRollout\ implementation that leverages DeepSpeed's hybrid engine. The rollout supports three generation paths: standard HuggingFace \generate\, CUDA graph capture with \DeepSpeedStaticCache\ for greedy decoding, and an experimental continuous batching mode. The module includes dataclasses for configuration (\RolloutConfig\, \SamplingConfig\) and I/O (\RolloutRequest\, \RolloutBatch\), along with a \ContinuousBatchScheduler\ for managing request admission and retirement in bounded, slot-based continuous batching scenarios.
deepspeed/runtime/rollout · high confidence
New FP quantization implementation with selective dequantization support
The \csrc/fp\_quantizer\ module now includes a new C++/CUDA implementation for floating-point quantization. This change introduces support for quantizing tensors to various bit-widths (FP4, FP6, FP8, and FP12) with configurable mantissa and exponent bits, including optional stochastic rounding for improved accuracy. A key behavioral addition is the new \selective\_dequantize\ function, which allows dequantizing only specific groups of data identified by an index array, rather than the entire tensor. The module also exposes \get\_scales\ to retrieve the scaling factors used during quantization, enabling users to inspect or reuse these values for downstream operations.
_csrc/fp\quantizer · high confidence
New FP8 quantization and fused GEMM operations
This change introduces a new \fp\_quantizer\ module providing FP8 quantization capabilities and fused matrix multiplication kernels. It adds \FP\_Quantize\ and \Quantizer\ classes for handling quantization/dequantization of tensors (supporting 4, 6, 8, and 12-bit precisions) and implements \matmul\_fp8\ which fuses GEMM with FP8 weight dequantization. The implementation includes a high-performance Triton-based kernel (\fp8\_gemm\_triton.py\) for bf16 and fp16 inputs, with a PyTorch fallback for systems where Triton is unavailable or unsupported.
_deepspeed/ops/fp\quantizer · high confidence
New Flops Profiler configuration schema
The deepspeed/profiling module now introduces a dedicated configuration structure for the Flops Profiler. Users can enable the profiler and tune its behavior via new JSON config keys: enabled, recompute\_fwd\_factor, profile\_step, module\_depth, top\_modules, detailed, and output\_file, each with defined defaults.
deepspeed/profiling · high confidence
New SDMA-accelerated ZeRO-3 example for AMD GPUs
The examples directory now includes a new \sdma\_allgather\ demo that showcases a hardware-accelerated AllGather path for ZeRO-3 on AMD MI300-series GPUs. By setting the \DS\_SDMA\_ALLGATHER=1\ environment variable, users can route collective traffic through the dedicated System DMA (SDMA) engines via the \mori\ library, bypassing the compute units to improve overlap with GEMM and attention workloads. The example provides scripts for both GPT-2-style and Qwen3-32B models, demonstrating approximately 10% step-time improvements while maintaining numerical accuracy compared to the standard RCCL/NCCL baseline.
examples · high confidence
New Triton-backed fused kernels for SwiGLU, grouped GEMM, and AutoEP token restore
DeepSpeed introduces a new \triton\_ops\ module providing fused Triton kernels to accelerate specific MoE and MLP operations. The \swiglu\ function now fuses the SiLU gate and up-projection into a single kernel, reducing intermediate memory traffic. A \group\_gemm\_triton\ implementation replaces the slow Python-side fallback for grouped GEMM on Ampere GPUs (sm80/sm86), launching a single fused kernel instead of per-group calls. Additionally, an opt-in \fused\_weighted\_sum\ restore for AutoEP eliminates eager scatter and FP32 intermediates during token restoration. All kernels include autograd support and fall back to eager PyTorch when Triton is unavailable.
_deepspeed/ops/triton\ops · high confidence
New benchmarking tools for pinned memory and multimodal sequence parallelism
Added new benchmark scripts to the benchmarks directory: a pinned-memory host-to-device/device-to-host bandwidth test (benchmarks/pin\_memory/h2d\_d2h\_bench.py) that compares torch and native pinned-memory backends, and an AutoSP multimodal sequence-parallelism benchmark (benchmarks/autosp/bench\_multimodal\_sp.py) for ViT+LLM architectures (InternVL, Qwen2VL) to measure latency, throughput, and memory usage. A README was also added to point users to external communication and inference benchmark suites.
benchmarks · high confidence
New fused CUDA transformer kernels for DeepSpeed
This change introduces a new set of highly optimized, fused CUDA kernels for the transformer module, including implementations for attention softmax, GELU activation, layer normalization, dropout, and matrix transposition. These kernels are designed to improve training and inference performance by reducing memory bandwidth usage and kernel launch overhead through fusions like bias-add-GELU and fused bias-residual-layer-norm. The implementation includes C++ wrappers and Python bindings to integrate these custom CUDA operations into the DeepSpeed training pipeline, supporting both FP32 and FP16 (half-precision) data types.
csrc/transformer · high confidence
New group-wise weight quantization for ZeRO-Inference
DeepSpeed ZeRO-Inference now supports group-wise weight quantization (INT4 and INT8) for Linear and Embedding layers. This change introduces a new quantization module that replaces standard PyTorch layers with quantized equivalents, allowing models to be compressed and run with reduced memory footprint during inference while maintaining accuracy through asymmetric fine-grained block quantization.
deepspeed/inference/quantization · high confidence
New inference context and CUDA layer abstractions for transformer models
This change introduces new header files in the inference includes directory that define the core runtime context and CUDA kernel launchers for transformer inference. The \InferenceContext\ class manages GPU workspace allocation, handling memory checks and cublas handle initialization, while the new \inference\_cublas\_wrappers.h\ provides a unified interface for GEMM operations that supports both NVIDIA CUDA and AMD ROCm backends (including backward compatibility for older Torch versions). Additionally, \inference\_cuda\_layers.h\ exposes template-based launchers for essential transformer operations such as attention softmax, bias-activation fusions (gelu, relu), layer normalization, residual connections, and rotary positional embeddings, enabling the underlying inference engine to execute these compute-intensive steps efficiently.
csrc/transformer/inference/includes · high confidence
New inference module interfaces and CUDA implementations for normalization layers
The DeepSpeed inference v2 module system now includes base interfaces and CUDA implementations for pre-normalization (LayerNorm and RMSNorm) and post-normalization layers. This adds \DSPreLNCUDAModule\ and \DSPreRMSCUDAModule\ for pre-norm operations, and \DSPostLNCUDAModule\ for post-norm, alongside their corresponding base classes (\DSPreNormBase\, \DSPostNormBase\) and registry structures. These components enable optimized, in-place normalization operations during inference, supporting both standard LayerNorm and RMSNorm variants with configurable epsilon and dtype constraints.
_deepspeed/inference/v2/modules/implementations/post\_norm, deepspeed/inference/v2/modules/implementations/pre\norm, deepspeed/inference/v2/modules/interfaces · high confidence
New modular container features for Hybrid Engine support
The \deepspeed/module\_inject/containers/features\ directory now introduces a set of new feature containers that extend the base model containers to support advanced Hybrid Engine capabilities. These include \HybridEngineContainer\ for managing LoRA parameter fusion/unfusion and training/inference transformations, \HybridGatedMLPContainer\ for handling models with separate gating and activation weights, \HybridSplitQKVContainer\ for supporting unfused QKV attention heads, \MetaTensorContainer\ for enabling checkpoint loading with PyTorch meta tensors, \MegatronContainer\ for handling Megatron-specific QKV layouts, and \HybridMegatronContainer\ combining both. These components provide the necessary infrastructure for models to leverage Hybrid Engine features like LoRA and meta-tensor optimizations within the DeepSpeed inference pipeline.
_deepspeed/module\inject/containers/features · high confidence
New op\_binding module for DeepSpeed inference operations
A new \op\_binding\ package has been introduced under \deepspeed/ops/transformer/inference\ to provide a unified, Python-based interface for DeepSpeed's inference kernels. This module defines a \BaseOp\ class that loads the underlying C++ inference module via \InferenceBuilder\ and exposes specific operations such as \LinearOp\, \QKVGemmOp\, \MLPGemmOp\, \SoftmaxOp\, and \RMSNormOp\. These bindings support multiple data types (float16, bfloat16, fp32, int8) and include fallback implementations using standard PyTorch functions when the optimized kernels are unavailable, ensuring compatibility across different hardware and configuration setups.
_deepspeed/ops/transformer/inference/op\binding · high confidence
New ragged-batch inference kernels for embeddings, attention, and KV-cache operations
DeepSpeed Inference V2 introduces a new \ragged\_ops\ module containing CUDA kernels that enable efficient inference on ragged (variable-length) batches. This location provides the low-level building blocks: a ragged-aware embedding lookup, a blocked FlashAttention-2 implementation that processes attention atoms over a blocked KV-cache, and kernels for applying rotary position embeddings and copying query/key/value tensors into the KV cache. These components allow the inference engine to handle variable sequence lengths and grouped-query attention patterns without padding overhead.
_deepspeed/inference/v2/kernels/ragged\ops · high confidence
Architecture
Introduction of the op\_builder package for accelerator-aware custom op compilation
The \op\builder\ directory now provides a unified, abstracted interface for building DeepSpeed's custom C++/CUDA/ROCm extensions. This change introduces a new \OpBuilder\ base class and a registry system (\\\init\\_.py\, \all\_ops.py\) that dynamically discovers and instantiates builders (such as \CPUAdamBuilder\, \FusedAdamBuilder\, \AsyncIOBuilder\, and \FPQuantizerBuilder\) based on the current accelerator. This architecture allows the build system to automatically handle platform-specific requirements (e.g., ROCm vs. CUDA, CPU vs. GPU) and ensures that compatibility checks and compilation flags are applied correctly for each operator without hardcoding platform logic in individual modules.
_op\builder · high confidence
ZeRO runtime restructured into a dedicated package with Pydantic configuration and new tiling/leaf-module APIs
The ZeRO optimization logic has been reorganized from flat files into a structured \\deepspeed/runtime/zero\\ package. Configuration is now validated via Pydantic models (\\DeepSpeedZeroConfig\\, \\DeepSpeedZeroOffloadParamConfig\\, etc.), replacing the previous dictionary-based setup. New public APIs are exposed through \\\_\init\\_.py\\, including \\TiledLinear\\ and \\TiledLinearReturnBias\\ for parameter memory release during forward passes, and \\DeepSpeedZeroLeafModuleConfig\\ to allow users to specify modules that bypass hook installation. The package also introduces \\DeepSpeedZeRoOffload\\ for managing parameter/optimizer offloading and \\ContiguousMemoryAllocator\\ for efficient memory management.
deepspeed/runtime/zero · high confidence
Behavioural changes
Automated release workflow with version validation and patch bumping
The release process now includes automated checks to ensure the version number provided matches the repository's current version before building and uploading the package. A new \release.sh\ script validates the input, verifies the version against \version.txt\, builds the source distribution using \python -m build --sdist\, uploads it via twine, and creates a git tag. Additionally, a \bump\_patch\_version.py\ script is introduced to automatically increment the patch version after a release, and a \check\_release\_version.py\ script ensures consistency between the release version and the repository state.
release · high confidence
Curriculum learning scheduler now prevents difficulty from dropping below minimum
The curriculum learning scheduler in the data pipeline has been updated to ensure that the difficulty level never falls below the configured minimum. Previously, when using fixed root or linear schedules, the difficulty calculation could floor to a value lower than the specified \min\_difficulty\ (for example, resulting in a zero-length sequence for sequence-length metrics). The new logic raises the calculated difficulty to the first valid multiple of the \difficulty\_step\ that is at or above the minimum, ensuring the schedule respects the lower bound while maintaining hardware alignment requirements.
_deepspeed/runtime/data\pipeline · high confidence
DeepSpeed initialization and environment reporting overhaul
DeepSpeed now prevents CUDA context initialization during import by defaulting the PYTORCH\_NVML\_BASED\_CUDA\_CHECK environment variable, ensuring fork-based multiprocessing remains stable. The distributed backend timeout has been reduced from 30 minutes to 10 minutes, and the \ds\_report\ utility now includes shared memory (/dev/shm) size diagnostics. Additionally, the library adds compatibility support for Python 3.14's new annotation handling and skips the Triton import on AMD hardware to avoid device API conflicts.
deepspeed · high confidence
DeepSpeed utils package reorganization and new capabilities
The deepspeed/utils package has been reorganized into a structured module, introducing a new CommsLogger for detailed communication profiling and logging, a backward-compatible API (bwc.py) for tensor and pipeline parallel group queries, and a new OnDevice context manager for initializing modules on specific devices and dtypes. The reorganization also includes new utilities for NUMA-aware process binding, NVTX instrumentation for profiling, and improved logging controls to reduce startup noise and prevent graph breaks during torch.compile.
deepspeed/utils · high confidence
Deprecation of deepspeed.compression module
The deepspeed.compression module has been removed and replaced with compatibility shims that emit deprecation warnings. Users importing recursive\_getattr or recursive\_setattr from this location will now see a FutureWarning directing them to import these utilities from deepspeed.utils.module\_utils instead.
deepspeed/compression · high confidence
Introduce DeepSpeed Communication Backend with Coalesced Collective Operations
The \deepspeed/runtime/comm\ module is introduced to provide a dedicated communication backend, replacing direct usage of \torch.distributed\ in favor of \deepspeed.comm\. This change includes the new \coalesced\_collectives.py\ module, which implements batched collective operations (such as \reduce\_scatter\_coalesced\ and \all\_to\_all\_quant\_reduce\) to amortize overhead and improve bandwidth utilization. The implementation supports quantized all-to-all operations with a fallback to standard reduce-scatter when tensor sizes are not divisible by the global world size, ensuring compatibility while optimizing performance for ZeRO++ and LoCo workflows.
deepspeed/runtime/comm · high confidence
Introduce FP quantizer header infrastructure with AMD BF16 build fix
This change adds the \fp\_context.h\ and \fp\_quantize.h\ headers to the \csrc/fp\_quantizer/includes\ directory, establishing the core context management and quantization/dequantization API for the FP quantizer. A key behavioral adjustment is the conditional inclusion of \cuda\_bf16.h\: it is now wrapped in a \BF16\_AVAILABLE\ check and explicitly excluded on AMD platforms (where it would otherwise resolve to \hip/hip\_bfloat16.h\), preventing host-only compiler build errors by using a forward declaration instead. The headers also define macros for quantization bit-switching and declare the launch functions for quantization, dequantization, and selective dequantization.
_csrc/fp\quantizer/includes · high confidence
Introduces geometric affine maps for universal checkpoint conversion
The checkpoint module now uses a geometric affine map representation to describe how parameters are sharded across ranks, replacing the previous regex-based name matching. This change enables universal checkpoints to correctly handle complex parallelism configurations, including AutoEP (Expert Parallelism) with ZeRO-3, AutoTP (Tensor Parallelism) folding, and uneven expert sharding, by explicitly tracking tensor offsets, strides, and scaling factors for invertible conversion between topologies.
deepspeed/checkpoint · high confidence
Introduction of DeepSpeed Inference Engine v2 with Pydantic-based configuration
The DeepSpeed Inference module has been replaced by a new v2 engine (InferenceEngineV2) that uses Pydantic for configuration validation. This introduces a structured \DeepSpeedInferenceConfig\ schema supporting explicit settings for tensor parallelism (AutoTP), Mixture-of-Experts (MoE), quantization, and CUDA graphs. Users must now provide configuration via this new typed schema rather than the previous dictionary-based or legacy config formats, ensuring stricter validation for parameters like data types, expert parallelism sizes, and kernel injection policies.
deepspeed/inference · high confidence
Introduction of standalone pin\_memory operation module
A new standalone module for the pin\_memory operation has been introduced in the deepspeed/ops directory. This change establishes a dedicated entry point for the pin\_memory functionality, separating it from other operations to provide a more modular structure for host memory pinning capabilities.
_deepspeed/ops/pin\memory · medium confidence
New modular container architecture for model inference
DeepSpeed introduces a new container-based architecture in the \module\_inject/containers\ module to manage model-specific tensors and policies for inference. This change replaces the previous monolithic approach with a structured hierarchy: a \BaseTransformerContainer\ handles common configuration and tensor management, while specific containers (e.g., \DS\_BERTContainer\, \DS\_LLAMAContainer\, \DS\_BloomContainer\) and their corresponding layer policies (e.g., \HFBertLayerPolicy\, \LLAMALayerPolicy\) provide model-specific implementations. This modular design supports advanced features like LoRA integration, hybrid engine compatibility, and meta-tensor loading across a wide range of models including BERT, GPT-2, Llama, BLOOM, and InternLM, ensuring that model-specific logic is cleanly encapsulated within these new container classes.
_deepspeed/module\inject/containers · high confidence
New modular transformer inference implementation structure
DeepSpeed introduces a new directory structure under \deepspeed/model\_implementations/transformers\ to isolate model-specific inference code. This change adds a base transformer class (\ds\_base.py\) and specific inference wrappers for BERT, Bloom, GPT, Llama 2, Megatron GPT, and OPT models. The core \ds\_transformer.py\ now handles the unified inference logic, including support for Triton kernels when available, and manages key-value caching for autoregressive generation. Additionally, a new \clip\_encoder.py\ provides CUDA graph support for CLIP model inference.
_deepspeed/model\implementations/transformers · high confidence
New module registry and heuristic-based instantiation for DeepSpeed Inference v2
DeepSpeed Inference v2 introduces a new modular architecture in the \modules\ package, featuring a registry system (\module\_registry.py\) and a base class (\ds\_module.py\) for defining inference components. A new \heuristics.py\ module centralizes the logic for selecting specific implementations (such as attention, embeddings, linear layers, MoE, and normalization) based on configuration and hardware constraints. Notably, linear layer instantiation now enforces hardware requirements for the \wf6af16\ quantization mode, restricting it to NVIDIA Ampere GPUs (compute capability 8.0+) and raising errors for unsupported architectures or non-CUDA environments.
deepspeed/inference/v2/modules · high confidence
New pre-commit scripts to enforce code standards and accelerator abstraction
Added four new pre-commit check scripts: \check-torchcuda.py\ enforces the use of \deepspeed.accelerator.get\_accelerator()\ instead of direct \torch.cuda\ calls to support hardware abstraction; \check-torchdist.py\ requires \deepspeed.comm\ instead of \torch.distributed\; \check-extraindexurl.py\ prevents the use of \--extra-index-url\ in favor of \--index-url\; and \check-license.py\ ensures all source files contain the correct Microsoft DeepSpeed copyright and Apache 2.0 license headers. A companion script, \replace\_copyright.py\, is also added to help automate the replacement of copyright headers across Python, C/C++, and Bash files.
scripts · high confidence
Optimized host buffer allocation for faster inference copies
The inference engine now uses a specialized host buffer allocation strategy to reduce latency during host-to-accelerator data transfers. New C++/CUDA source files implement a \get\_cuda\_fast\_buffer\ function that utilizes \cudaHostAlloc\ with specific flags (portable, mapped, write-combined) to optimize memory access, alongside PyTorch bindings (\allocate\_fast\_host\_buffer\, \allocate\_view\_on\, \allocate\_view\_like\) that create optimized host-side mirrors of accelerator tensors. This change targets the critical path of the forward pass to improve performance by minimizing copy overhead.
deepspeed/inference/v2/ragged/csrc · high confidence
Refactored checkpoint writer configuration and data-parallel resource assignment
The checkpointing subsystem has been restructured to support more granular configuration and complex parallelism topologies. A new \model\_checkpointing\ package introduces explicit configuration handling in \config.py\ and \constants.py\, allowing users to tune writer types (mock, python, fast), I/O buffer sizes, and data-parallel units (replica, socket, machine). The \DataParallelWriterFactory\ now dynamically assigns checkpoint-writing resources based on Universal Parallel Info, enabling support for 3D parallelism (pipeline, tensor, data) and expert parallelism (MoE) by partitioning ranks across sockets or machines. Additionally, the \CheckpointWriterFactory\ has been updated to route tensor pinning through the accelerator's \pin\_memory\ interface, ensuring DeepNVMe can efficiently skip bounce buffers on CUDA-capable hosts.
_deepspeed/runtime/model\checkpointing · high confidence
Updated base images and added Python 3.11/3.12 build environments
The Docker build environment has been updated to use newer base images, specifically switching the main CUDA image to nvidia/cuda:12.2.2-devel-ubuntu20.04 and the PyTorch version to 1.13.0. Additionally, new Dockerfiles have been introduced in the gh-builder directory to support building with Python 3.11 and Python 3.12, providing users with updated language runtime options for development and testing.
docker · high confidence
Test coverage
Added BingBertSquad model test suite; Added benchmark scripts for DeepSpeed4Science and Triton kernels; Added comprehensive test coverage for the Muon optimizer; Added debugging and validation tests for ZeRO-3 partial offload and memory behavior; Added initial PyTorch Lightning integration test; Added model sanity check test runner; Added performance benchmarks and accelerator initialization tests; Added tests for DeepCompile ZeRO-3 parameter management and graph break diagnostics; Added tests for DeepSpeed hybrid engine integration; Added tests for Gemma4 configuration validation; Added tests for Megatron GPT2 model integration; Added tests for PipelineModule checkpointing behavior; Added tests for configurable engine log levels and hybrid engine rollout profiling; Added unit tests for AutoEP + AutoTP folding; Added unit tests for AutoEP communication, AutoTP partitioning, and Llama RoPE configuration; Added unit tests for AutoTP custom patterns, multi-model isolation, training, and TP plans; Added unit tests for AutoTP universal checkpoint metadata collection; Added unit tests for BF16, mixed-precision, and gradient clipping behaviors; Added unit tests for CPU and ZenFlow optimizer operations; Added unit tests for DeepCompile compilation and memory management; Added unit tests for DeepSpeed Hybrid Engine capabilities; Added unit tests for DeepSpeed compile subsystem; Added unit tests for DeepSpeed launcher argument parsing and multinode runners; Added unit tests for DeepSpeed linear optimization components; Added unit tests for DeepSpeed runtime components; Added unit tests for DeepSpeed utility functions; Added unit tests for DeepSpeed-FastGen v2 ragged inference components; Added unit tests for DeepSpeed4Science EvoformerAttention; Added unit tests for FP quantizer operations; Added unit tests for FlopsProfiler accuracy and stability; Added unit tests for INTx quantization and ZeRO-3 inference integration; Added unit tests for Inference V2 core modules; Added unit tests for Inference V2 kernel operations; Added unit tests for Inference V2 model implementation internals; Added unit tests for NHWC bias add spatial inference operations; Added unit tests for NPU graph operations and MPS accelerator initialization; Added unit tests for NVMe async I/O, GDS, and pinned memory management; Added unit tests for Ulysses SP and ALST components; Added unit tests for ZeRO runtime behaviors; Added unit tests for ZeRO state offloading, gradient reduction, and checkpoint alignment; Added unit tests for ZenFlow configuration validation and runtime behavior; Added unit tests for accelerator forward and backward passes; Added unit tests for activation checkpointing and CPU offload; Added unit tests for autotuning components; Added unit tests for coalesced collective operations; Added unit tests for communication layer components; Added unit tests for distributed tensor partitioning utilities; Added unit tests for half-precision runtime behavior; Added unit tests for monitoring integrations; Added unit tests for native host memory pinning and accelerator integration; Added unit tests for new Triton-optimized operations; Added unit tests for op builder JIT and compatibility logic; Added unit tests for pipeline parallel runtime components; Added unit tests for sparse gradient handling and averaging; Added unit tests for the Lion optimizer; Added unit tests for the apply\_rotary\_pos\_emb sequence layer; Added unit tests for the quantizer inference operations; Added unit tests for transformer inference operators; Expanded unit test coverage for checkpointing, AutoTP, and AutoEP; New AIO benchmarking and validation test suite; New unit test infrastructure and fixtures; New unit tests for DeepSpeed inference capabilities; Standardize test infrastructure with pytest configuration and environment validation.
Dependencies
Standardize and pin Python dependencies across modular requirement files
The project now uses a set of modular requirement files (e.g., requirements.txt, requirements-dev.txt, requirements-readthedocs.txt) to manage dependencies more precisely. Key updates include pinning Pydantic to version 2.0.0 or higher across core and documentation builds, upgrading the development toolchain to use clang-format 18.1.3 and pre-commit 3.2.0, and constraining pytest to versions between 7.2.0 and 8.4.0. Additional changes include adding specific packages for autotuning (hjson, tabulate, xgboost), inference (lm-eval 0.3.0, safetensors), and Stable Diffusion support (diffusers \>=0.25.0, triton \>=2.1.0), while removing hard requirements for tensorboardX and megatron-lm.
(dependencies) · high confidence
Housekeeping
Added license header to quantizer module initialization
The \_\init\\_.py file for the deepspeed.ops.quantizer module now includes the standard Microsoft copyright notice and Apache 2.0 license identifier, aligning the module with the project's updated licensing standards.
deepspeed/ops/quantizer · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 50 → 63 (+13.1)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 82 → 80 (-2.4)
- Architecture 90 → 99 (+9.4)
- Maturity 81 → 71 (-9.5)
- Readiness 30 → 52 (+22.2)
- Security 52 → 66 (+14.2)
Resolved (175)
- Context/problem and consequences/trade-offs are absent; only the title and link point to a README describing DeepSpeed-FastGen (docs/_posts/2023-11-06-deepspeed-fastgen-chinese.md)
- Coverage not measured — test suite did not build
- Decision and consequences are present but the context/problem is only hinted by the title ('降低4倍网络通信') with no body explaining why this trade-off matters or what it replaces (docs/_posts/2023-06-22-zeropp-chinese.md)
- Dimension evaluation failed
- Duplicated block (10 lines × 2) (deepspeed/ops/sparse_attention/matmul.py)
- Duplicated block (10 lines × 2) (deepspeed/runtime/engine.py)
- Duplicated block (10 lines × 2) (deepspeed/runtime/zero/mics.py)
- Duplicated block (10 lines × 2) (tests/unit/sequence_parallelism/test_ulysses.py)
- Duplicated block (10 lines × 2) (tests/unit/ulysses_alst/test_ulysses_sp_hf.py)
- Duplicated block (10 lines × 2) (tests/unit/ulysses_alst/test_ulysses_sp_hf.py)
- Duplicated block (10 lines × 2) (tests/unit/v1/half_precision/test_with_autocast.py)
- Duplicated block (10 lines × 2) (tests/unit/v1/moe/test_moe.py)
- Duplicated block (10 lines × 3) (tests/unit/v1/half_precision/test_bf16.py)
- Duplicated block (11 lines × 2) (deepspeed/ops/transformer/inference/triton/matmul_ext.py)
- Duplicated block (11 lines × 2) (deepspeed/ops/transformer/inference/triton/matmul_ext.py)
- Duplicated block (11 lines × 2) (deepspeed/runtime/zero/stage3.py)
- Duplicated block (11 lines × 2) (tests/unit/ulysses_alst/test_ulysses_sp_hf.py)
- Duplicated block (11 lines × 3) (tests/unit/ops/muon/test_muon.py)
- Duplicated block (11 lines × 3) (tests/unit/runtime/half_precision/onebit/test_onebit.py)
- Duplicated block (12 lines × 2) (deepspeed/runtime/engine.py)
- …and 155 more
New (1146)
- AllGatherCoalescedHandle.wait (cognitive 20) (deepspeed/runtime/zero/partition_parameters.py)
- AutoEP.ep_parser (cognitive 107) (deepspeed/module_inject/auto_ep.py)
- AutoEP.ep_parser (cyclomatic 39) (deepspeed/module_inject/auto_ep.py)
- AutoEPMoELayer.init (cognitive 37) (deepspeed/module_inject/auto_ep_layer.py)
- AutoEPMoELayer.init (cyclomatic 29) (deepspeed/module_inject/auto_ep_layer.py)
- AutoEPMoELayer._finalize_output (cognitive 17) (deepspeed/module_inject/auto_ep_layer.py)
- AutoEPMoELayer._forward (cognitive 22) (deepspeed/module_inject/auto_ep_layer.py)
- AutoEPMoELayer._forward (cyclomatic 17) (deepspeed/module_inject/auto_ep_layer.py)
- AutoTP._configure_gathered_column_tie_fallbacks (cognitive 34) (deepspeed/module_inject/auto_tp.py)
- AutoTP._configure_gathered_column_tie_fallbacks (cyclomatic 23) (deepspeed/module_inject/auto_tp.py)
- AutoTP._replace (cognitive 26) (deepspeed/module_inject/auto_tp.py)
- AutoTP._replace (cyclomatic 27) (deepspeed/module_inject/auto_tp.py)
- AutoTP._replace_module (cognitive 55) (deepspeed/module_inject/auto_tp.py)
- AutoTP._replace_module (cyclomatic 28) (deepspeed/module_inject/auto_tp.py)
- AutoTP.tp_parser (cognitive 27) (deepspeed/module_inject/auto_tp.py)
- AutoTP.tp_parser (cyclomatic 27) (deepspeed/module_inject/auto_tp.py)
- Autotuner._generate_experiments (cognitive 16) (deepspeed/autotuning/autotuner.py)
- Autotuner.get_min_max_micro_batch_size (cognitive 47) (deepspeed/autotuning/autotuner.py)
- Autotuner.get_min_max_micro_batch_size (cyclomatic 21) (deepspeed/autotuning/autotuner.py)
- Autotuner.print_tuning_results (cognitive 23) (deepspeed/autotuning/autotuner.py)
- …and 1126 more
Changes since last survey
- 196 commits — 151 feature/other, 45 fixes
By area
- deepspeed/runtime — 65 commits
- tests/unit — 29 commits
- deepspeed/module_inject — 18 commits
- .github/workflows — 10 commits
- deepspeed/compile — 10 commits
- (root) — 7 commits
- deepspeed/comm — 6 commits
- deepspeed/checkpoint — 5 commits
- deepspeed/ops — 5 commits
- deepspeed/profiling — 4 commits
- deepspeed/utils — 4 commits
- deepspeed/elasticity — 3 commits
- op_builder/mps — 3 commits
- accelerator/abstract_accelerator.py — 2 commits
- csrc/cpu — 2 commits
- csrc/includes — 2 commits
- deepspeed/init.py — 2 commits
- deepspeed/inference — 2 commits
- deepspeed/launcher — 2 commits
- deepspeed/sequence — 2 commits
Notable commits
- fix: Fix 2 CPU failures in cpu-torch-latest (#8629)
- fix: Fix AutoEP ZeRO-1/2 universal conversion (#8198)
- fix: Fix AutoTP + deep compile collectives silently drop when AC is on (#8355)
- fix: Fix AutoTP metadata updates for unsharded modules (#8299)
- fix: Fix DeepCompile last use for unconsumed waits (#8254)
- fix: Fix HybridEngine OPT generation with modern cache interfaces (#8523)
- fix: Fix Muon aspect-ratio scale truncated when muon_update recompiles (#8641)
- fix: Fix Muon optimizer under ZeRO CPU offload and bound gather buffers (#8464)
- fix: Fix SequenceTiledCompute backward for empty trailing shards (#8434)
- fix: Fix Triton NFS detection crash when df wraps long device names (#8259)
- fix: Fix ZeRO parameter alignment for grouped_mm (#8277)
- fix: Fix ZeRO++ secondary shard copy for small params (#8210)
- fix: Fix ZeRO-1/2 with zero-sized parameters (#8280)
- fix: Fix ZeRO-3 crash in AutoTP universal-checkpoint metadata (#8270)
- fix: Fix ZeRO-3 synchronization during OPSD rollout (#8264)
- fix: Fix ZeroDivisionError in compute_elastic_config return_microbatch on non-0.2 elasticity (#8286)
- fix: Fix checkpoint rank selection for Ulysses sequence parallelism (#8226)
- fix: Fix comms logger KeyError when log_name is omitted (#8267)
- fix: Fix compilation error for torch 2.12+ (#8238)
- fix: Fix device mismatch in test_gate_up_partition_covers_the_whole_weight (#8308)
- …and 176 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
deepspeedai/DeepSpeed was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit bc778b8c3fd1fabc87a29202548d8b0ddfba2ace — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-9984f8053b7b.