Skip to content
CAI
Software that uses CAICheck a score

huggingface/diffusers

67.3

Adequate · 18 September 2026

758.3k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a comprehensive library for diffusion-based generative AI, providing a unified framework to load, fine-tune, and run inference on a wide variety of models for image, video, audio, and text generation. It supports extensive customization through modular pipeline architectures, parameter-efficient fine-tuning methods like LoRA, and advanced conditioning techniques such as ControlNet and IP-Adapter. The library also includes specialized tools for benchmarking, profiling, and deploying these models in production environments.

How it got here

2022–2023 — Model expansion and infrastructure modernization

68 changes.

This period focused on integrating a wide array of new model architectures, including Flux, Wan, Kandinsky, and DiT, while establishing modular pipeline structures and comprehensive training examples. The project simultaneously modernized its internal architecture by reorganizing utilities, implementing strict quality checks, and expanding test coverage to support these diverse new features.

2024 — Transformer expansion and quantization infrastructure

44 changes.

The project significantly expanded its model support by integrating numerous transformer-based architectures for image and video generation, including Stable Diffusion 3, Hunyuan, CogVideoX, and Mochi. Concurrently, a unified quantization infrastructure was introduced to enable efficient inference across multiple backends like bitsandbytes, GGUF, and TorchAO. The period also featured extensive research examples for fine-tuning and alignment techniques, alongside structural reorganizations of UNet and ControlNet modules.

2025–2026 — modular pipeline architecture and multimodal expansion

64 changes.

The project introduced a comprehensive modular pipeline framework, allowing users to compose diffusion workflows from individual model blocks rather than relying on monolithic classes. This architectural shift was accompanied by the integration of numerous new pipelines for image, video, audio, and text generation, significantly expanding the library's multimodal capabilities. Additionally, the period saw the addition of advanced inference optimizations, quantization backends, and research examples to support these new models.

Features

Add AnyFlow video diffusion pipelines

New \AnyFlowPipeline\ and \AnyFlowFARPipeline\ classes are now available for AnyFlow flow-map-distilled video generation. The bidirectional \AnyFlowPipeline\ supports any-step inference (1, 2, 4, 8... NFE) for text-to-video, while the causal \AnyFlowFARPipeline\ enables chunk-wise autoregressive generation with support for text-to-video, image-to-video, and video-to-video modes.

src/diffusers/pipelines/anyflow · high confidence

Add AnyText research example for multilingual visual text generation and editing

This change introduces the AnyText research project to the examples directory, providing a diffusion pipeline that generates or edits text within images. The implementation includes a custom \AnyTextControlNetModel\ to handle text glyph and position conditioning, alongside a bundled OCR recognition module (based on MobileNetV1 and SVTR) to encode stroke data. Users can now leverage this example to create images with specific text overlays or edit existing text in images, as demonstrated in the provided README and usage scripts.

_examples/research\projects/anytext · high confidence

Add BitsAndBytes quantization support for Diffusers models

Users can now load and run Diffusers models using 4-bit and 8-bit quantization via the \bitsandbytes\ library. This change introduces the \BnB4BitDiffusersQuantizer\ and \BnB8BitDiffusersQuantizer\ classes, which replace standard linear layers with \bnb.nn.Linear4bit\ or \bnb.nn.Linear8bitLt\ modules during model loading. The implementation includes validation for required dependencies (Accelerate \>= 0.26.0, bitsandbytes \>= 0.43.3) and support for GPU, XPU, and MPS devices, allowing for reduced memory usage and faster inference on compatible hardware.

src/diffusers/quantizers/bitsandbytes · high confidence

Add CogVideoX LoRA finetuning examples

New training scripts and documentation are provided for performing Low-Rank Adaptation (LoRA) on the CogVideoX model. The \train\_cogvideox\_lora.py\ script enables text-to-video finetuning, while \train\_cogvideox\_image\_to\_video\_lora.py\ supports image-to-video adaptation. The accompanying \README.md\ details data preparation formats, required dependencies (such as \peft\ and \diffusers\), and usage examples for distributed training via Accelerate.

examples/cogvideo · high confidence

Add CogVideoX video generation pipelines

Introduces the CogVideoX pipeline family for video generation, including \CogVideoXPipeline\ for text-to-video, \CogVideoXImageToVideoPipeline\ for image-to-video, \CogVideoXVideoToVideoPipeline\ for video-to-video, and \CogVideoXFunControlPipeline\ for controlled generation. These pipelines leverage the CogVideoX transformer model and T5 text encoder, supporting schedulers like CogVideoXDDIM and CogVideoXDPMS, and include lazy-loading infrastructure and output structures for seamless integration.

src/diffusers/pipelines/cogvideo · high confidence

Add CogView3Plus text-to-image pipeline

Users can now generate images using the CogView3Plus model via the new \CogView3PlusPipeline\. This addition includes the pipeline implementation, a dedicated output dataclass, and module exports, enabling integration with existing Diffusers workflows for this specific architecture.

src/diffusers/pipelines/cogview3 · high confidence

Add CogView4 Control training example

Added a new example in \examples/cogview4-control\ that demonstrates how to train Control LoRAs and perform full fine-tuning for the CogView4 model using structural controls like depth maps and poses. The entry includes the \train\_control\_lora\_cogview4.py\ and \train\_control\_cogview4.py\ scripts, along with documentation explaining the input feature expansion (64 to 128 channels) and providing inference examples using the \CogView4ControlPipeline\.

examples/cogview4-control · high confidence

Add CogView4 and CogView4 Control pipelines

Introduces the \CogView4Pipeline\ for text-to-image generation and the \CogView4ControlPipeline\ for control-conditioned generation. The implementation includes the core pipeline classes, a dedicated output dataclass, and lazy-loading infrastructure, enabling users to generate images using the CogView4 model architecture with support for both standard prompts and control images.

src/diffusers/pipelines/cogview4 · high confidence

Add ColossalAI DreamBooth training example

Introduces a new example script (\train\_dreambooth\_colossalai.py\) that enables training Stable Diffusion models using the ColossalAI framework. This addition allows users to leverage ColossalAI's Gemini memory management to reduce GPU memory usage and support larger model scales through heterogeneous CPU/GPU memory placement. The example includes a README with setup instructions, a requirements file, and an inference script, providing a complete workflow for personalizing text-to-image models with ColossalAI's distributed training capabilities.

_examples/research\projects/colossalai · high confidence

Add Consistency Training example script

Added a new community example in \examples/research\_projects/consistency\_training\ that provides a script (\train\_cm\_ct\_unconditional.py\) and documentation for training consistency models from scratch. The example implements the consistency training algorithm (including improvements from the iCT paper) and supports both unconditional and class-conditional training on datasets like CIFAR-10, utilizing the \ConsistencyModelPipeline\ and \CMStochasticIterativeScheduler\ from the diffusers library.

_examples/research\_projects/consistency\training · high confidence

Add Control-LoRA inference example for SDXL

Added a new example script and documentation in the research projects directory that demonstrates how to use Control-LoRA with Stable Diffusion XL. The example shows how to load Control-LoRA adapters into a ControlNet model to enable image conditioning (specifically using canny edge detection) for more efficient and compact model control on consumer GPUs.

_examples/research\_projects/control\lora · high confidence

Add Cosmos 3 example with multi-modal generation and multi-GPU parallelism

The \examples/cosmos3\ directory now provides a runnable smoke-test CLI (\inference\_cosmos3.py\) and multi-GPU helpers (\cosmos\_parallel.py\) for the Cosmos 3 pipeline. Users can generate text-to-image, text-to-video, image-to-video, and video-to-video content, as well as generate synchronized audio tracks. The example also supports action generation modes (forward/inverse dynamics and policy) for robotics and autonomous vehicle domains. For large models like Cosmos 3-Super, the example includes support for tensor parallelism to shard weights across GPUs, and context parallelism to shard sequences for lower latency on long videos.

examples/cosmos3 · high confidence

Add DiT image generation pipeline

Users can now generate images using the DiT (Diffusion Transformer) architecture via the new \DiTPipeline\. This pipeline replaces the traditional UNet backbone with a transformer-based model (\DiTTransformer2DModel\) and supports class-conditioned generation using ImageNet labels. It includes utilities for mapping label strings to class IDs and integrates with standard Diffusers components like \AutoencoderKL\ and \KarrasDiffusionSchedulers\.

src/diffusers/pipelines/bria, src/diffusers/pipelines/dit · high confidence

Add DiffusionGemma pipeline for block-diffusion text generation

Users can now generate text using the DiffusionGemma model via the new \DiffusionGemmaPipeline\. This pipeline supports block-diffusion generation with configurable schedulers (BlockRefinementScheduler, DiscreteDDIMScheduler, or EntropyBoundScheduler), handles both single and batched prompts (including multimodal inputs with images), and exposes parameters for controlling generation length, inference steps, and early stopping based on EOS tokens or stability thresholds.

_src/diffusers/pipelines/diffusion\gemma · high confidence

Add DreamBooth training scripts with Scheduled Pseudo-Huber Loss

New example scripts are available in the \examples/research\_projects/scheduled\_huber\_loss\_training/dreambooth\ directory for training DreamBooth models using a Scheduled Pseudo-Huber Loss objective. The release includes \train\_dreambooth.py\ for standard DreamBooth training, \train\_dreambooth\_lora.py\ for LoRA-based training on Stable Diffusion, and \train\_dreambooth\_lora\_sdxl.py\ for LoRA training on SDXL, providing users with research-oriented training configurations.

_examples/research\_projects/scheduled\_huber\_loss\training/dreambooth · high confidence

Add EasyAnimateV5.1 video generation pipelines

Introduces three new pipelines for EasyAnimateV5.1 video generation: \EasyAnimatePipeline\ for text-to-video, \EasyAnimateControlPipeline\ for control-to-video (using pose or other control signals), and \EasyAnimateInpaintPipeline\ for image-to-video inpainting. These pipelines support the EasyAnimateTransformer3DModel and Magvit autoencoder, utilize Qwen2-VL for text encoding, and include specific utilities for handling video latents and control inputs.

src/diffusers/pipelines/easyanimate · high confidence

Add Ernie-Image text-to-image pipeline

Introduces the ErnieImagePipeline, enabling text-to-image generation using the Ernie-Image architecture. This new pipeline integrates a custom DiT transformer, a Flux2-style VAE, and a Mistral3 text encoder, with optional support for a Prompt Enhancement (PE) model to rewrite prompts via a chat template. It also includes LoRA loading capabilities through ErnieImageLoraLoaderMixin and handles latent patchification/unpatchification for the specific model requirements.

_src/diffusers/pipelines/ernie\_image, src/diffusers/pipelines/nucleusmoe\_image, src/diffusers/pipelines/ovis\image · high confidence

Add Flux image generation examples for PyTorch/XLA on TPUs

New example scripts (\flux\_inference.py\ and \flux\_inference\_spmd.py\) and documentation have been added to demonstrate image generation using the FLUX model on TPU devices via PyTorch/XLA. The standard script runs inference on a single TPU chip using XLA multiprocessing and custom flash attention block sizes optimized for Trillium TPUs. An additional SPMD version is provided for configurations like v5e-8, where the transformer weights are sharded across multiple chips to handle memory constraints, while the VAE remains on the CPU and is moved to the TPU only for decoding.

_examples/research\_projects/pytorch\xla/inference · high confidence

Add GGUF quantization support for diffusion models

Users can now load and run diffusion models quantized in the GGUF format, which significantly reduces memory usage and improves inference speed on supported hardware. This change introduces the \GGUFQuantizer\ class and associated utilities to handle GGUF-specific weight loading, dequantization, and linear layer replacement. It includes support for various GGML quantization types (Q4\_0, Q8\_0, etc.) and optional CUDA kernel acceleration for faster dequantization when available. The implementation respects \modules\_to\_not\_convert\ and \keep\_in\_fp32\_modules\ configurations to ensure numerical stability for specific model components.

src/diffusers/quantizers/gguf · high confidence

Add GeoDiff molecule conformation research notebook

Added a new experimental research notebook in the \examples/research\_projects/geodiff\ directory that demonstrates how to generate stable 3D structures of molecules using the GeoDiff model and Diffusers. The notebook includes instructions for setting up the environment (including Conda and PyTorch) and running the pretrained models on GEOM datasets to produce physically accurate molecular conformations.

_examples/research\projects/geodiff · high confidence

Add HiDream Image generation pipeline

Introduces the HiDreamImagePipeline, enabling users to generate images using the HiDream architecture. This new pipeline integrates multiple text encoders (CLIP, T5, Llama) and supports LoRA fine-tuning via HiDreamImageLoraLoaderMixin, with output provided through the HiDreamImagePipelineOutput class.

_src/diffusers/pipelines/hidream\image · high confidence

Add Hunyuan-DiT ControlNet inference pipeline

Introduces the \HunyuanDiTControlNetPipeline\ and its module registration, enabling users to perform image generation using the Tencent Hunyuan-DiT model with ControlNet conditioning. The pipeline supports both English and Chinese prompts, utilizes \T5Tokenizer\ for text encoding, and handles standard image aspect ratios and shapes for 1024x1024 and other common resolutions.

_src/diffusers/pipelines/controlnet\hunyuandit · high confidence

Add HunyuanDiT pipeline for bilingual text-to-image generation

Introduces the HunyuanDiTPipeline, enabling image generation from English and Chinese prompts using the Tencent Hunyuan DiT model. The pipeline integrates a bilingual CLIP text encoder and an mT5 encoder (via T5Tokenizer) for robust text understanding, supports standard aspect ratios (1:1, 4:3, 3:4, 16:9, 9:16) with automatic shape mapping, and includes noise rescaling to improve image quality and reduce overexposure.

src/diffusers/pipelines/hunyuandit · high confidence

Add HunyuanImage and HunyuanImageRefiner pipelines

Users can now generate images using the Hunyuan-Image 2.1 model via two new pipelines: \HunyuanImagePipeline\ for text-to-image generation and \HunyuanImageRefinerPipeline\ for image refinement. These pipelines integrate the HunyuanImage transformer, a Hunyuan-specific VAE, and Qwen2.5-VL and T5 text encoders, with support for optional OCR guidance and XLA acceleration.

_src/diffusers/pipelines/hunyuan\image · high confidence

Add Ideogram4 modular pipeline with structured prompt upsampling and LoRA support

Introduces the \Ideogram4ModularPipeline\ in \src/diffusers/modular\_pipelines/ideogram4\, providing a modular text-to-image workflow for the Ideogram4 model. This pipeline includes an optional structured prompt upsampling step that rewrites prompts into Ideogram4's native JSON caption format using a Qwen3-VL text encoder and \Ideogram4PromptEnhancerHead\ (requiring the \outlines\ library for schema-constrained generation). The core denoising process utilizes an asymmetric CFG approach with conditional and unconditional transformers, a logit-normal flow-matching schedule, and a per-step guidance schedule that drops to 3.0 for the final polish steps. The pipeline also supports LoRA loading via \Ideogram4LoraLoaderMixin\ and decodes latents using \AutoencoderKLFlux2\.

_src/diffusers/modular\pipelines/ideogram4 · high confidence

Add Intel-optimized inference example for Stable Diffusion

Added a new example script and documentation in the Intel optimization research project that demonstrates accelerating Stable Diffusion inference on Intel platforms using Intel Extension for PyTorch. The example enables Bfloat16 support and memory format optimizations (channels\_last) to improve performance on Intel Xeon Scalable Processors, including support for the DPMSolverMultistepScheduler.

_examples/research\_projects/intel\opts · high confidence

Add JoyImage Edit and Edit Plus pipelines

Introduces two new image-editing pipelines: JoyImageEditPipeline for single-image editing and JoyImageEditPlusPipeline for multi-image instruction-guided editing. These pipelines leverage a Qwen3-VL text encoder, a 3-D transformer denoiser, and a WAN VAE, and include a custom image processor that handles bucket-based resolution selection and resize-center-crop preprocessing to optimize inference performance.

src/diffusers/pipelines/joyimage · high confidence

Add Kandinsky 2.1 text-to-image pipeline

Introduces the Kandinsky 2.1 pipeline suite, enabling text-to-image, image-to-image, and inpainting generation. This includes the core \KandinskyPipeline\, specialized variants (\KandinskyImg2ImgPipeline\, \KandinskyInpaintPipeline\), a \KandinskyPriorPipeline\ for generating image embeddings from text, and a \KandinskyCombinedPipeline\ that unifies these components. The implementation supports CPU offloading, lazy imports for optional dependencies (Torch, Transformers), and PyTorch/XLA acceleration.

src/diffusers/pipelines/kandinsky, src/diffusers/pipelines/lumina, src/diffusers/pipelines/lumina2 · high confidence

Add Kandinsky 2.2 text-to-image fine-tuning examples

New training scripts and documentation are provided for fine-tuning the Kandinsky 2.2 prior and decoder models. The \examples/kandinsky2\_2/text\_to\_image/\ directory now includes \train\_text\_to\_image\_decoder.py\ and \train\_text\_to\_image\_prior.py\ for full model fine-tuning, as well as \train\_text\_to\_image\_lora\_decoder.py\ and \train\_text\_to\_image\_lora\_prior.py\ for parameter-efficient LoRA fine-tuning. The accompanying README guides users through setting up the environment, configuring datasets (such as the Naruto example), and running distributed training with Accelerate, including instructions for saving and loading the resulting pipelines.

_examples/kandinsky2\2 · high confidence

Add Kolors text-to-image and image-to-image pipelines

Users can now generate images using the Kolors model via the new \KolorsPipeline\ (text-to-image) and \KolorsImg2ImgPipeline\ (image-to-image). These pipelines integrate the ChatGLM3-6B text encoder and tokenizer, support LoRA weight loading, and allow for IP-Adapter integration, providing a direct interface for Kwai-Kolors models within the library.

src/diffusers/pipelines/kolors · high confidence

Add Krea 2 Modular Pipeline support

Introduces a new modular pipeline implementation for the Krea 2 text-to-image model, providing both standard and distilled 'turbo' variants. This change adds the \Krea2ModularPipeline\ and \Krea2TurboModularPipeline\ classes along with their corresponding block definitions (\Krea2AutoBlocks\, \Krea2TurboAutoBlocks\) and step implementations for text encoding, denoising, and decoding. The standard pipeline utilizes a Qwen3-VL text encoder and symmetric classifier-free guidance, while the turbo variant runs guidance-free for faster inference.

_src/diffusers/modular\pipelines/krea2 · high confidence

Add Krea 2 text-to-image pipeline

Users can now generate images using the Krea 2 (K2) model via the new \Krea2Pipeline\. This pipeline integrates a Qwen3-VL text encoder, a Krea 2 transformer, and a Qwen-Image VAE, supporting both standard mid-training checkpoints and few-step distilled (TDM) modes. It includes built-in LoRA loading capabilities through \Krea2LoraLoaderMixin\ and exposes a dedicated output class for returned images.

src/diffusers/pipelines/krea2 · high confidence

Add LEdits++ pipelines for Stable Diffusion and Stable Diffusion XL

Introduces new \LEditsPPPipelineStableDiffusion\ and \LEditsPPPipelineStableDiffusionXL\ classes in the \ledits\pp\ module, enabling image editing via the LEdits++ algorithm. These pipelines provide an \invert\ method for image inversion and a \\\call\\_\ method for guided editing using parameters like \editing\_prompt\, \edit\_guidance\_scale\, and \edit\_threshold\. The implementation includes custom attention storage and processing components to facilitate the editing process, along with dedicated output dataclasses for both the final edited images and the intermediate inversion results.

_src/diffusers/pipelines/ledits\pp · high confidence

Add LLaDA2 discrete diffusion training and sampling examples

New example scripts and documentation have been added to the \examples/discrete\_diffusion\ directory to support the LLaDA2 pipeline. Users can now train causal language models using block-wise iterative refinement with a confidence-aware loss via \train\_llada2.py\, or generate text using the \LLaDA2Pipeline\ with block-level sampling and optional post-mask editing via \sample\_llada2.py\.

_examples/discrete\diffusion · high confidence

Add LLaDA2 pipeline for discrete diffusion text generation

Users can now generate text using the LLaDA2 pipeline, which implements discrete diffusion via block-wise iterative refinement. This new capability allows for text generation by maintaining a template sequence filled with mask tokens and refining it in blocks, sampling candidate tokens based on confidence. The pipeline integrates with the BlockRefinementScheduler and supports standard input methods including prompts, messages with chat templates, and direct input IDs.

src/diffusers/pipelines/llada2 · high confidence

Add LTX Video modular pipeline

Introduces a new modular pipeline for LTX Video, enabling text-to-video and image-to-video generation through a composable block-based architecture. This change adds the \LTXModularPipeline\ class along with specialized blocks for text encoding (\LTXTextEncoderStep\), latent preparation (\LTXPrepareLatentsStep\), denoising (\LTXDenoiseStep\), and video decoding (\LTXVaeDecoderStep\). It also includes a \LTXVideoPachifier\ for handling latent tensor reshaping and exposes the pipeline via the \diffusers.modular\_pipelines.ltx\ module.

_src/diffusers/modular\pipelines/ltx · high confidence

Add Latent Consistency Models pipelines

Introduces the \LatentConsistencyModelPipeline\ for text-to-image generation and the \LatentConsistencyModelImg2ImgPipeline\ for image-to-image generation. These new pipelines are designed to work with Latent Consistency Models (LCM), enabling significantly faster inference with fewer steps (typically 1–8 steps) compared to standard diffusion models. Both pipelines support standard features such as LoRA weights, IP-Adapters, and textual inversions, and are exported via the \diffusers.pipelines.latent\_consistency\_models\ module.

_src/diffusers/pipelines/latent\_consistency\models · high confidence

Add Latent Perceptual Loss (LPL) example for Stable Diffusion XL

This change introduces a new research example in \examples/research\_projects/lpl\ that implements Latent Perceptual Loss (LPL) for training Stable Diffusion XL models. The addition includes a training script (\train\_sdxl\_lpl.py\) and a loss module (\lpl\_loss.py\) that apply perceptual objectives in the VAE latent space to improve image quality, detail preservation, and training stability. Users can enable this feature via command-line arguments (e.g., \--use\_lpl\) to integrate these perceptual losses into their SDXL fine-tuning workflows.

_examples/research\projects/lpl · high confidence

Add LoRA fine-tuning example for Cosmos Predict 2.5

This change introduces a complete example for fine-tuning the Cosmos Predict 2.5 model using LoRA. It includes a training script (\train\_cosmos\_predict25\_lora.py\) that supports LoRA and DoRA adapters, an evaluation script (\eval\_cosmos\_predict25\_lora.py\) for inference with trained adapters, and helper scripts to download and preprocess the GR1 dataset. The example also provides shell scripts for launching training and evaluation, along with LLM-based evaluation prompts for assessing video instruction following and physical commonsense.

examples/cosmos · high confidence

Add LoRA fine-tuning example with text encoder support

A new training script and documentation are provided in the \examples/research\_projects/lora\ directory for fine-tuning Stable Diffusion using Low-Rank Adaptation (LoRA). This example extends standard LoRA training by adding support for applying LoRA layers to the text encoder, in addition to the UNet. The included README details how to run the training using the Naruto dataset and explains the inference process for loading the resulting LoRA weights.

_examples/research\projects/lora · high confidence

Add LongCat-Image and LongCat-Image Edit pipelines

Introduces two new pipelines, LongCatImagePipeline and LongCatImageEditPipeline, for generating and editing images using the LongCat-Image model. The generation pipeline supports text-to-image synthesis with automatic prompt optimization and language detection (English/Chinese), while the edit pipeline enables image modification based on text prompts. Both pipelines utilize the LongCatImageTransformer2DModel and Qwen2.5-VL text encoder, and include support for XLA acceleration.

_src/diffusers/pipelines/longcat\image · high confidence

Add LongCatAudioDiT pipeline for text-to-audio generation

Users can now generate audio from text prompts using the new LongCatAudioDiTPipeline. This pipeline integrates a UMT5 text encoder, a LongCatAudioDiT transformer, and a VAE, allowing for the synthesis of audio files (e.g., ambience, sound effects) based on natural language descriptions. The implementation includes automatic duration estimation from prompt text and supports standard diffusion sampling parameters like guidance scale and inference steps.

_src/diffusers/pipelines/longcat\_audio\dit · high confidence

Add LucyEditPipeline for video-to-video editing

Users can now use the new LucyEditPipeline to perform video-to-video editing by providing a condition video appended to the channel dimension. This pipeline, based on the Wan architecture, supports two-stage denoising via an optional second transformer and allows for LoRA loading, enabling specific visual edits guided by text prompts.

src/diffusers/pipelines/lucy · high confidence

Add MiniMax Music 3 modular pipeline

Users can now generate music using the MiniMax Music 3 model via a new modular pipeline located in \src/diffusers/modular\_pipelines/minimax\_music3\. This pipeline implements a three-stage workflow: an autoregressive Qwen3 language model generates per-frame semantic codes from a text prompt and optional lyrics; a flow-matching transformer denoises these codes into Flow-VAE latents in overlapping 200-frame windows; and a DAC-style vocoder decodes the latents into a stereo 44.1 kHz waveform. The implementation includes specific handling for prompt tokenization, chunk-based conditioning with overlap blending, and configurable inference steps.

_src/diffusers/modular\_pipelines/minimax\music3 · high confidence

Add MiniMax-H3 modular pipeline support

Introduces the \MiniMaxH3ModularPipeline\ and its constituent blocks in \src/diffusers/modular\_pipelines/minimax\_h3\, enabling text-to-video, image-to-video, and reference-to-video generation with the MiniMax-H3 model. This implementation includes specific handling for the model's unique requirements, such as a fixed 24 fps output, a configurable 768-pixel short-edge canvas, and a packed sequence layout that jointly processes video, audio, and text tokens. The pipeline integrates Qwen3-VL for text conditioning, supports keyframe anchoring for video generation, and includes dedicated decoders for both video and audio outputs.

_src/diffusers/modular\_pipelines/minimax\_h3, src/diffusers/pipelines/deepfloyd\if · high confidence

Add Mochi 1 text-to-video pipeline

Introduces the MochiPipeline for text-to-video generation, including the main pipeline implementation, a dedicated output class, and module initialization. This new capability allows users to generate videos using the Mochi model architecture, supporting features like LoRA loading and XLA acceleration.

src/diffusers/pipelines/allegro, src/diffusers/pipelines/mochi · high confidence

Add NVIDIA ModelOpt quantization backend

Users can now quantize models using the NVIDIA ModelOpt library. This change introduces a new \NVIDIAModelOptQuantizer\ in the \src/diffusers/quantizers/modelopt\ module, enabling support for pre-quantized ModelOpt models and runtime quantization with FP8 support. The implementation handles specific ModelOpt requirements such as calibration, compression, and managing quantization state, requiring the \nvidia\_modelopt\ package to be installed.

src/diffusers/quantizers/modelopt · high confidence

Add Nunchaku Lite single-file quantization support

Users can now quantize models using the Nunchaku Lite format, which supports SVDQ W4A4 and AWQ W4A16 precision modes. This change introduces the \NunchakuLiteQuantizer\ class and integrates with the Hugging Face \kernels\ package to download and apply optimized kernels for inference. The implementation enforces specific hardware requirements, requiring a CUDA-capable NVIDIA GPU (Turing or newer for INT4, Blackwell or newer for NVFP4) and validates the environment before loading checkpoints.

src/diffusers/quantizers/nunchaku · high confidence

Add ONNX Runtime unconditional image generation example

Introduces a new training script and documentation for unconditional image generation using ONNX Runtime acceleration. The example allows users to train a DDPM UNet model on datasets like Oxford Flowers, supporting features such as center cropping, random flipping, EMA, and mixed precision training via Accelerate.

_examples/research\_projects/onnxruntime/unconditional\_image\generation · high confidence

Add ONNX Runtime-accelerated Textual Inversion training example

Introduces a new example script (\textual\_inversion.py\) and documentation for fine-tuning Stable Diffusion models using Textual Inversion with ONNX Runtime acceleration. This allows users to leverage ONNX Runtime for faster training on custom datasets, including specific instructions for setting up the environment, authenticating with the Hugging Face Hub, and running the training pipeline via Accelerate.

_examples/research\_projects/onnxruntime/textual\inversion · high confidence

Add OmniGen multimodal-to-image generation pipeline

Introduces the OmniGen pipeline, enabling users to generate images from text instructions and input images. This new capability includes the \OmniGenPipeline\ class for orchestration, an \OmniGenMultiModalProcessor\ for handling multimodal inputs (text and images), and the necessary module exports in \\_\init\\_.py\ to make the pipeline accessible via the diffusers library.

src/diffusers/pipelines/omnigen · high confidence

Add PixArt-Alpha and PixArt-Sigma text-to-image pipelines

Introduces two new pipelines, \PixArtAlphaPipeline\ and \PixArtSigmaPipeline\, for text-to-image generation. The \PixArtAlphaPipeline\ supports resolution binning for 256px, 512px, and 1024px outputs, while \PixArtSigmaPipeline\ extends this with 2048px resolution support and allows custom timestep or sigma schedules. Both pipelines utilize the \PixArtTransformer2DModel\ and \T5\ text encoders, and include support for Torch/XLA inference.

_src/diffusers/pipelines/pixart\alpha · high confidence

Add PromptDiffusion research example with custom ControlNet and pipeline

Introduces a new example in \examples/research\_projects/promptdiffusion\ implementing the Prompt Diffusion model (based on the paper 'In-Context Learning Unlocked for Diffusion Models'). This includes a custom \PromptDiffusionControlNetModel\ that extends the standard ControlNet to support query-conditioned embeddings, a dedicated \PromptDiffusionPipeline\ for inference, and a conversion script to transform original checkpoints into the diffusers format.

_examples/research\projects/promptdiffusion · high confidence

Add PyTorch/XLA text-to-image training example for Stable Diffusion

A new example script and documentation have been added to demonstrate fine-tuning Stable Diffusion models on TPU devices using PyTorch/XLA. The \train\_text\_to\_image\_xla.py\ script implements Distributed Data Parallel via XLA's GSPMD feature to shard input batches across TPU cores, supporting multi-host training on v4 and v5p accelerators. The accompanying README provides setup instructions for Google Cloud TPUs, including environment configuration and dependency installation, along with performance benchmarks and usage examples for both training and inference.

_examples/research\_projects/pytorch\xla/training · high confidence

Add RealFill research example for text2image inpainting personalization

This change introduces the RealFill example, a training and inference script for personalizing Stable Diffusion inpainting models using just a few reference images. The entry includes \train\_realfill.py\ for training with support for LoRA, gradient checkpointing, 8-bit optimizers, and xformers to accommodate low-memory GPUs, as well as \infer.py\ for running inference on trained models. A \README.md\ provides setup instructions and usage examples.

_examples/research\projects/realfill · high confidence

Add Retrieval Augmented Diffusion (RDM) research example

A new research project example has been added to \examples/research\_projects/rdm\ that implements Retrieval Augmented Diffusion Models. This includes a \RDMPipeline\ for text-to-image generation that integrates a CLIP-based image retriever to inject retrieved visual features into the diffusion process, alongside a \Retriever\ module for managing FAISS-based image indexes and datasets.

_examples/research\projects/rdm · high confidence

Add SDXL example for JAX/Flax on TPU

The \examples/research\_projects/sdxl\_flax\ directory now includes a demonstration of running Stable Diffusion XL using JAX and Flax on Google Cloud TPUs. The example provides two Python scripts: \sdxl\_single.py\, which shows a standard JIT-compiled inference pipeline, and \sdxl\_single\_aot.py\, which demonstrates Ahead-of-Time (AOT) compilation for optimized parallel execution across multiple TPU devices. The accompanying README details setup requirements, including installing JAX with TPU support and using \diffusers\ versions prior to 0.40.0, as JAX/Flax support was removed in later releases.

_examples/research\_projects/sdxl\flax · high confidence

Add Sana Sprint training example for diffusers

Added a new example in \examples/research\_projects/sana\ that demonstrates how to train the Sana Sprint model using the diffusers library. The entry includes a Python training script (\train\_sana\_sprint\_diffusers.py\), a bash execution script, and a README with setup and usage instructions. The example supports training on the \brivangl/midjourney-v6-llava\ dataset using pre-trained teacher models (1.6B or 0.6B) and includes configuration for mixed precision, gradient checkpointing, and specific training strategies like \train\_largest\_timestep\ and \misaligned\_pairs\_D\.

_examples/research\projects/sana · high confidence

Add Shap-E 3D generation pipeline with text and image-to-3D support

Introduces the Shap-E pipeline, enabling users to generate 3D assets from text prompts or input images. This location provides the core implementation including \ShapEPipeline\ for text-to-3D, \ShapEImg2ImgPipeline\ for image-to-3D, a differentiable projective camera module for rendering, and a NeRF-based renderer that projects latents into 3D objects. The pipelines utilize a CLIP text or image encoder, a prior transformer, and a Heun discrete scheduler to generate latent representations which are then rendered into 3D visualizations.

_src/diffusers/pipelines/shap\e · high confidence

Add SkyReels V2 video generation pipelines

Introduces the SkyReels V2 pipeline module, providing five new video generation capabilities: standard text-to-video (\SkyReelsV2Pipeline\), image-to-video (\SkyReelsV2ImageToVideoPipeline\), and three variants using diffusion forcing for extended or asynchronous generation (\SkyReelsV2DiffusionForcingPipeline\, \SkyReelsV2DiffusionForcingImageToVideoPipeline\, and \SkyReelsV2DiffusionForcingVideoToVideoPipeline\). These pipelines leverage the \SkyReelsV2Transformer3DModel\, \AutoencoderKLWan\, and \UniPCMultistepScheduler\ to support high-resolution video synthesis with configurable parameters for frame count, resolution, and inference steps.

_src/diffusers/pipelines/skyreels\v2 · high confidence

Add Stable Diffusion 3 Modular Pipeline

Introduces a new modular pipeline implementation for Stable Diffusion 3, exposing \StableDiffusion3ModularPipeline\ and \StableDiffusion3AutoBlocks\ for programmatic control over the generation process. This change adds the core pipeline wiring and a suite of granular processing blocks—including text encoding, latent preparation, denoising loops, and image decoding—allowing users to compose and customize the inference steps rather than relying on a single monolithic call.

_src/diffusers/modular\_pipelines/anima, src/diffusers/modular\_pipelines/stable\_diffusion\_3, src/diffusers/pipelines/consistency\models · high confidence

Add T2I-Adapter training example for Stable Diffusion XL

The \examples/t2i\_adapter\ directory now includes a complete training example for T2I-Adapters on Stable Diffusion XL. This release adds the \train\_t2i\_adapter\_sdxl.py\ script, which allows users to train adapters using the fill50k dataset, along with \README\_sdxl.md\ documentation covering installation, training configuration, and inference usage. A corresponding test file (\test\_t2i\_adapter.py\) has been added to verify the training script's functionality, and the main \README.md\ has been updated to direct users to the SDXL-specific documentation.

_examples/t2i\adapter · high confidence

Add Wan-Animate-2 modular pipeline implementation

This change introduces the complete modular pipeline implementation for the Wan-Animate-2 model, enabling image-to-video character animation. The new \src/diffusers/modular\_pipelines/wan\_animate\_2\ module provides \WanAnimate2ModularPipeline\ and \WanAnimate2DistilledModularPipeline\ classes, along with the underlying modular blocks for encoding (text, image, video), segment-based denoising, and decoding. It includes a custom \WanAnimate2VideoProcessor\ to handle specific letterboxing and aspect-ratio preservation requirements, and supports both standard and distilled checkpoint variants.

_src/diffusers/modular\_pipelines/wan\_animate\2 · high confidence

Add experimental Diffusion DPO training examples for SD, SDXL, and SDXL Turbo

The \examples/research\_projects/diffusion\_dpo\ directory now includes training scripts (\train\_diffusion\_dpo.py\, \train\_diffusion\_dpo\_sdxl.py\) and documentation for performing Direct Preference Optimization (DPO) on Stable Diffusion models using LoRA. Users can align SD 1.5, SDXL, and SDXL Turbo models with preference data (e.g., PickaScore) via \accelerate launch\, with specific support for SDXL Turbo's low-step inference mode. These scripts are experimental and intended for low-data regime research.

_examples/research\_projects/diffusion\dpo · high confidence

Add modular pipeline for HunyuanVideo 1.5

Introduces a new modular pipeline implementation for the HunyuanVideo 1.5 model, enabling text-to-video and image-to-video generation. This change adds the \HunyuanVideo15ModularPipeline\ class and a suite of modular blocks (encoders, denoisers, decoders) that handle the specific architecture of HunyuanVideo 1.5, including dual text encoding via Qwen2.5-VL and ByT5, and integration with the \HunyuanVideo15Transformer3DModel\ and \AutoencoderKLHunyuanVideo15\.

_src/diffusers/modular\_pipelines/hunyuan\_video1\5 · high confidence

Add scheduled Pseudo-Huber loss training scripts for text-to-image models

New training scripts have been added to the \examples/research\_projects/scheduled\_huber\_loss\_training/text\_to\_image\ directory, providing examples for fine-tuning Stable Diffusion and Stable Diffusion XL models using the scheduled Pseudo-Huber loss. The release includes \train\_text\_to\_image.py\ and \train\_text\_to\_image\_sdxl.py\ for standard fine-tuning, as well as \train\_text\_to\_image\_lora.py\ and \train\_text\_to\_image\_lora\_sdxl.py\ for LoRA-based adaptation. These scripts support dataset loading from the Hugging Face Hub (defaulting to \lambdalabs/naruto-blip-captions\), validation logging via WandB or TensorBoard, and model card generation for uploaded results.

_examples/research\_projects/scheduled\_huber\_loss\_training/text\_to\image · high confidence

Add textual inversion example with Intel Extension for PyTorch optimizations

Introduces a new training script and documentation for performing textual inversion fine-tuning of Stable Diffusion models using Intel Extension for PyTorch. This addition enables users to leverage CPU-specific optimizations, including support for bfloat16 on compatible Intel Xeon processors, and provides instructions for both single-node and multi-node distributed training setups.

_examples/research\_projects/intel\_opts/textual\inversion · high confidence

Advanced DreamBooth LoRA training scripts now support DoRA, B-LoRA, and Flux.1

The advanced training scripts for Stable Diffusion 1.5, SDXL, and the new Flux.1 model have been updated with several new training capabilities. Users can now utilize DoRA (Direction-wise Low-Rank Adaptation) and B-LoRA as alternative LoRA optimization methods, in addition to standard LoRA. The SDXL script also adds support for EDM-style training schedulers. Furthermore, a new advanced training script for Flux.1 has been introduced, featuring support for Pivotal Tuning (textual inversion) on both CLIP and T5 text encoders, configurable target modules for the Transformer, and latent caching to reduce memory usage. These updates are accompanied by corresponding documentation and test coverage.

_examples/advanced\_diffusion\training · high confidence

Community Pipeline Examples documentation

The \examples/community/README.md\ file has been added to provide a comprehensive catalog of community-contributed pipeline examples. This documentation lists various pipelines such as CLIP Guided Stable Diffusion, Long Prompt Weighting, TensorRT acceleration, and AnimateDiff ControlNet, including descriptions, code links, and Colab notebooks to help users discover and utilize these community scripts.

examples/community · high confidence

ControlNet training examples now support FLUX, SD3/3.5, and SDXL

The ControlNet training examples have been expanded to include dedicated scripts and documentation for training on FLUX, Stable Diffusion 3/3.5, and Stable Diffusion XL, in addition to the existing Stable Diffusion 1.5 support. New training scripts (\train\_controlnet\_flux.py\, \train\_controlnet\_sd3.py\, \train\_controlnet\_sdxl.py\) and corresponding README guides are provided, along with updated tests to validate these new model families.

examples/controlnet · high confidence

Experimental value-guided RL pipeline for diffusion models

Adds a new \ValueGuidedRLPipeline\ in the experimental reinforcement learning module, enabling value-guided sampling from diffusion models trained on state sequences. Users can now integrate a value function (UNet1DModel) with a denoising UNet and an OpenAI Gym-compatible environment to generate and select actions based on trajectory value, supporting planning horizons and guide steps for improved decision-making in environments like Hopper.

src/diffusers/experimental/rl · high confidence

Introduce AudioLDM2 pipeline for text-to-audio and text-to-speech generation

Adds the AudioLDM2 pipeline, enabling users to generate audio waveforms from text prompts or synthesize speech from text and transcripts. The implementation includes the \AudioLDM2Pipeline\ class, a custom \AudioLDM2UNet2DConditionModel\ for denoising, and an \AudioLDM2ProjectionModel\ to map embeddings from multiple text encoders (CLAP, T5/VITS) into a shared latent space for the GPT-2 language model.

src/diffusers/pipelines/audioldm2 · high confidence

Introduce AuraFlow image generation pipeline

Adds the AuraFlow pipeline, enabling users to generate images using the AuraFlow model architecture. This new feature includes support for LoRA adapters via the AuraFlowLoraLoaderMixin, compatibility with Torch XLA for accelerated inference, and integration with the FlowMatchEulerDiscreteScheduler. The pipeline utilizes a T5-based text encoder and an AuraFlowTransformer2DModel to process prompts and denoise latents, providing a complete end-to-end solution for AuraFlow-based image synthesis.

_src/diffusers/pipelines/aura\flow · high confidence

Introduce Bria Fibo and Bria Fibo Edit pipelines

Added the \BriaFiboPipeline\ and \BriaFiboEditPipeline\ classes to the \diffusers\ library, enabling image generation and editing using the Bria Fibo model architecture. The edit pipeline supports multi-reference conditioning, allowing users to pass multiple reference images to guide the editing process, and includes utilities for handling JSON-based prompts and masks.

_src/diffusers/pipelines/bria\fibo · high confidence

Introduce Chroma image generation pipelines

Added the Chroma pipeline module (\src/diffusers/pipelines/chroma\) providing \ChromaPipeline\ for text-to-image generation, \ChromaImg2ImgPipeline\ for image-to-image tasks, and \ChromaInpaintPipeline\ for inpainting. These pipelines are designed for the \lodestones/Chroma1-HD\ model and support features such as LoRA loading, single-file checkpoint loading, and IP-Adapter integration.

src/diffusers/pipelines/chroma · high confidence

Introduce ConsisID pipeline for identity-consistent video generation

Added the ConsisID pipeline, which enables generating videos that preserve the identity of a person shown in a reference image. The implementation includes the main pipeline class, utility functions for face detection and embedding extraction (using insightface, facexlib, and consisid\_eva\_clip), and a dedicated output class. Users can now import ConsisIDPipeline to perform identity-consistent video synthesis, requiring optional dependencies like OpenCV, PyTorch, and Transformers.

src/diffusers/pipelines/consisid · high confidence

Introduce Custom Diffusion training example with real-image regularization

Adds a new \custom\_diffusion\ example that demonstrates how to customize text-to-image models (like Stable Diffusion) using only a few images of a subject. The training script (\train\_custom\_diffusion.py\) supports single and multi-concept training, integrates with Weights & Biases for experiment tracking, and allows pushing results to the Hugging Face Hub. To prevent overfitting, the example includes a \retrieve.py\ utility that downloads real regularization images from LAION using \clip-retrieval\, alongside a pytest-based test suite (\test\_custom\_diffusion.py\) to validate training and checkpointing behavior.

_examples/custom\diffusion · high confidence

Introduce DDIMPipeline for image generation

Adds a new DDIMPipeline class that wraps a UNet2DModel and DDIMScheduler to generate images. The pipeline supports configurable batch sizes, inference steps, and the eta parameter to interpolate between DDIM and DDPM sampling behaviors. It handles device execution (including PyTorch/XLA support), manages model offloading, and returns outputs in PIL, NumPy, or PyTorch tensor formats.

src/diffusers/pipelines/ddim · high confidence

Introduce DDPM pipeline with PyTorch/XLA support

The DDPM pipeline is now available as a dedicated module under \src/diffusers/pipelines/ddpm\, exposing the \DDPMPipeline\ class for image generation. This implementation supports PyTorch/XLA acceleration via \torch\_xla\ when available, ensuring compatibility with XLA devices during the denoising loop. Users can now import and use the pipeline directly from this location, leveraging the standard \DiffusionPipeline\ interface for loading models and generating images.

src/diffusers/pipelines/ddpm · high confidence

Introduce ErnieImage modular pipeline for text-to-image generation

Adds a new modular pipeline implementation for the ErnieImage model, enabling text-to-image generation with an optional prompt enhancer (Ministral3ForCausalLM) and support for LoRA loading. The pipeline is composed of distinct processing steps: prompt enhancement, text encoding (Mistral3Model), denoising (ErnieImageTransformer2DModel with classifier-free guidance), and decoding (AutoencoderKLFlux2). It requires transformers version 5.0.0 or higher for the prompt enhancer components.

_src/diffusers/modular\_pipelines/ernie\image · high confidence

Introduce FLUX.2 pipeline family with Klein variants and KV-caching

Adds the \src/diffusers/pipelines/flux2\ module, introducing the \Flux2Pipeline\ for text-to-image generation alongside specialized variants: \Flux2KleinPipeline\ for the 9B distilled model, \Flux2KleinInpaintPipeline\ for inpainting with optional reference conditioning, and \Flux2KleinKVPipeline\ which caches reference image attention keys and values to accelerate subsequent denoising steps. The module also includes \Flux2ImageProcessor\ for validating and resizing input images, \Flux2PipelineOutput\ for standardized results, and system message templates for prompt formatting.

src/diffusers/pipelines/flux2 · high confidence

Introduce Flux pipeline family

Adds the Flux pipeline module, providing a complete set of generation pipelines for the Flux model architecture. This includes the base text-to-image pipeline, along with specialized variants for image-to-image, inpainting, and controllable generation (ControlNet, Canny, Depth). The module also introduces the FluxFill pipeline for fill-in-the-middle tasks, the FluxKontext pipeline for context-aware generation, and the FluxPriorRedux pipeline for multi-image input scenarios, all exposed via a centralized \\_\init\\_.py\ for easy import.

_src/diffusers/pipelines/flux, src/diffusers/pipelines/stable\_diffusion\xl · high confidence

Introduce Flux2 modular pipeline implementation

Adds a new modular pipeline implementation for the Flux2 model family, including support for the standard Flux2-dev, the distilled Flux2-Klein, and the base Flux2-Klein variants. This change introduces a set of reusable pipeline blocks (encoders, denoisers, decoders, and input processors) and corresponding auto-block configurations that allow users to compose and customize the Flux2 inference workflow. The implementation includes specific handling for text embeddings via Mistral3/Qwen encoders, image conditioning via VAE encoding, and the core denoising loop with guidance and RoPE inputs, exposing these as importable components under \diffusers.modular\_pipelines.flux2\.

_src/diffusers/modular\_pipelines/flux2, src/diffusers/modular\_pipelines/stable\_diffusion\xl · high confidence

Introduce GLM-Image text-to-image pipeline

Users can now generate images using the GLM-Image model via the new \GlmImagePipeline\. This addition includes the pipeline implementation, output data structures, and module initialization, integrating the autoregressive vision-language encoder with a diffusion transformer for image decoding.

_src/diffusers/pipelines/glm\image · high confidence

Introduce Helios modular video generation pipeline

Adds the Helios modular pipeline to Diffusers, providing a new way to generate videos using the Helios model. This includes the core pipeline class (HeliosModularPipeline) and a set of modular blocks for text encoding, image/video encoding, chunk-based denoising (with support for standard, pyramid, and distilled pyramid modes), and decoding. The implementation supports text-to-video, image-to-video, and video-to-video workflows, leveraging the new HeliosTransformer3DModel and HeliosScheduler components.

_src/diffusers/modular\pipelines/helios · high confidence

Introduce HunyuanVideo 1.5 pipelines for text-to-video and image-to-video generation

Adds the \HunyuanVideo15Pipeline\ for text-to-video generation and the \HunyuanVideo15ImageToVideoPipeline\ for image-to-video generation, along with the \HunyuanVideo15ImageProcessor\ for handling input image preprocessing. These new components enable users to leverage the HunyuanVideo 1.5 model architecture within the diffusers library.

_src/diffusers/pipelines/hunyuan\_video1\5 · high confidence

Introduce HunyuanVideo pipeline family

Adds the HunyuanVideo pipeline module, providing four new pipelines for video generation: \HunyuanVideoPipeline\ for text-to-video, \HunyuanVideoImageToVideoPipeline\ for image-to-video, \HunyuanSkyreelsImageToVideoPipeline\ for SkyReels-based image-to-video, and \HunyuanVideoFramepackPipeline\ for framepack-based generation. These pipelines integrate HunyuanVideo-specific components, including the \HunyuanVideoTransformer3DModel\, \AutoencoderKLHunyuanVideo\, and \FlowMatchEulerDiscreteScheduler\, and support features like LoRA loading and PyTorch/XLA acceleration.

_src/diffusers/pipelines/hunyuan\video · high confidence

Introduce Ideogram4 text-to-image pipeline with structured prompt upsampling

Adds the Ideogram4Pipeline for generating images using the Ideogram4 flow-matching model, featuring asymmetric classifier-free guidance and a resolution-aware logit-normal sigma schedule. The pipeline includes an optional structured prompt upsampling capability that rewrites user prompts into a specific JSON schema via the Ideogram4PromptEnhancerHead, ensuring consistent and detailed image generation.

src/diffusers/pipelines/ideogram4 · high confidence

Introduce Kandinsky 3.0 text-to-image and image-to-image pipelines

Adds the Kandinsky3Pipeline for text-to-image generation and Kandinsky3Img2ImgPipeline for image-to-image generation, along with a conversion script to migrate U-Net weights to the Kandinsky3UNet format. These pipelines leverage T5 text encoders, Kandinsky3 U-Nets, and VQ models, support classifier-free guidance, and include utilities for CPU offloading and LoRA loading.

src/diffusers/pipelines/kandinsky3 · high confidence

Introduce Kandinsky 5.0 pipelines for text-to-image, image-to-image, text-to-video, and image-to-video generation

This change adds the Kandinsky 5.0 pipeline implementations to the library, providing four new generation modes: text-to-image (Kandinsky5T2IPipeline), image-to-image (Kandinsky5I2IPipeline), text-to-video (Kandinsky5T2VPipeline), and image-to-video (Kandinsky5I2VPipeline). These pipelines leverage the Kandinsky5Transformer3DModel, the Qwen2.5-VL text encoder, and the HunyuanVideo VAE (for video pipelines) or the FLUX.1-dev VAE (for image pipelines), and support LoRA loading via KandinskyLoraLoaderMixin. The module exposes lazy imports for all four pipelines and defines dedicated output classes (KandinskyPipelineOutput for video, KandinskyImagePipelineOutput for images) to standardize return values.

src/diffusers/pipelines/kandinsky5 · high confidence

Introduce Kandinsky v2.2 pipeline suite

Adds a new set of pipelines for the Kandinsky v2.2 model family, including \KandinskyV22Pipeline\ for text-to-image generation, \KandinskyV22Img2ImgPipeline\ for image-to-image, \KandinskyV22InpaintPipeline\ for inpainting, and \KandinskyV22PriorPipeline\ for generating image embeddings. The release also includes combined pipelines (\KandinskyV22CombinedPipeline\, \KandinskyV22Img2ImgCombinedPipeline\, \KandinskyV22InpaintCombinedPipeline\) that wrap the prior and decoder steps, as well as ControlNet variants (\KandinskyV22ControlnetPipeline\, \KandinskyV22ControlnetImg2ImgPipeline\) for depth-conditioned generation. These components are registered in the \diffusers\ package via lazy imports.

_src/diffusers/pipelines/kandinsky2\2 · high confidence

Introduce LTX Video pipeline suite and latent upsampler

Adds the LTX Video pipeline family to the library, including \LTXPipeline\ for text-to-video, \LTXConditionPipeline\ for text-and-frame conditioning, \LTXImageToVideoPipeline\ for image-to-video, \LTXI2VLongMultiPromptPipeline\ for long multi-prompt image-to-video with temporal tiling, and \LTXLatentUpsamplePipeline\ for latent-space upscaling. The suite introduces the \LTXLatentUpsamplerModel\ for spatial/temporal latent upsampling, exposes the \LTXVideoCondition\ dataclass for specifying frame-level conditioning, and registers all components in the \src/diffusers/pipelines/ltx\ module for lazy import.

src/diffusers/pipelines/ltx · high confidence

Introduce LTX-2 video generation pipelines and supporting components

Adds the LTX-2 pipeline module (\src/diffusers/pipelines/ltx2\) with a full suite of video generation capabilities, including the base \LTX2Pipeline\, image-to-video (\LTX2ImageToVideoPipeline\), condition-based (\LTX2ConditionPipeline\), latent upscaling (\LTX2LatentUpsamplePipeline\), and high-dynamic-range (HDR) and in-context LoRA variants (\LTX2HDRPipeline\, \LTX2InContextPipeline\). The module also introduces specialized components such as text connectors, a duration head for predicting shot length, a latent upsampler, and HDR-specific image processing and export utilities, enabling users to generate and process high-quality video content with LTX-2 models.

src/diffusers/pipelines/ltx2 · high confidence

Introduce Latent Diffusion pipelines for text-to-image and super-resolution

The \latent\_diffusion\ module now provides two new pipelines: \LDMTextToImagePipeline\ for generating images from text prompts and \LDMSuperResolutionPipeline\ for upscaling low-resolution images. These pipelines support optional Torch/XLA acceleration, allow users to pass pre-generated latents for reproducible tweaking, and default to returning PIL images while accepting various schedulers (DDIM, PNDM, LMS, Euler, DPMSolver).

_src/diffusers/pipelines/latent\diffusion · high confidence

Introduce Latte text-to-video pipeline

Adds a new \LattePipeline\ for text-to-video generation using the Latent Diffusion Transformer architecture. The pipeline integrates a T5 text encoder, an AutoencoderKL for video latents, and the \LatteTransformer3DModel\, supporting features like CPU offloading and PyTorch/XLA acceleration.

src/diffusers/pipelines/latte · high confidence

Introduce Modular Diffusers pipelines for composable model architectures

This release introduces the Modular Diffusers framework, allowing users to build and run pipelines by composing individual model blocks (such as UNets, VAEs, and text encoders) rather than loading monolithic pipeline classes. The \src/diffusers/modular\_pipelines\ module provides core infrastructure including \ModularPipeline\, \AutoPipelineBlocks\, and a \ComponentsManager\ that handles component loading, configuration, and custom offloading strategies. It ships with pre-configured modular pipelines for a wide range of models, including Stable Diffusion XL, Stable Diffusion 3, Flux, Flux 2 (with Klein support), Wan (image-to-video and animate), Helios, Ideogram 4, Krea 2, Qwen Image, LTX, LTX 2.5, Cosmos 3, Ernie Image, Hunyuan Video 1.5, MiniMax H3, MiniMax Music 3, Z-Image, and Anima. The system also supports integration with the Mellon node-based workflow editor and allows saving modular pipelines to the Hugging Face Hub.

_src/diffusers/modular\pipelines · high confidence

Introduce NVIDIA Cosmos video generation pipelines

This change adds a new \cosmos\ pipeline package to Diffusers, exposing inference support for NVIDIA Cosmos models. It includes \Cosmos2TextToImagePipeline\ and \Cosmos2VideoToWorldPipeline\ for the Cosmos Predict2 series, \Cosmos2\_5\_PredictBasePipeline\ and \Cosmos2\_5\_TransferPipeline\ for the Predict2.5 and Transfer2.5 series (with ControlNet support for edge, depth, segmentation, and blur), and \Cosmos3OmniPipeline\ for the Cosmos3 model. The Cosmos3 pipeline specifically introduces mixed W8A8/W8A16 precision handling for ModelOpt FP8 checkpoints, allowing high-precision execution for the first and last denoising steps to improve quality while maintaining efficiency.

src/diffusers/pipelines/cosmos · high confidence

Introduce PRX and PRXPixel text-to-image pipelines

Adds the \PRXPipeline\ and \PRXPixelPipeline\ classes to the \diffusers\ library, enabling text-to-image generation using Photoroom's PRX models. The standard \PRXPipeline\ supports aspect-ratio bins for 256px, 512px, and 1024px resolutions, while the new \PRXPixelPipeline\ operates directly in pixel space (without a VAE) at a default 1024px resolution. The module also includes a compatibility wrapper for \T5GemmaEncoder\ to handle composite config structures in transformers 5.x and defines the \PRXPipelineOutput\ dataclass for standardized results.

src/diffusers/pipelines/prx · high confidence

Introduce Qwen-Image pipeline family

Adds a new family of pipelines for the Qwen-Image model, including text-to-image, image editing, inpainting, img2img, and ControlNet variants. This release introduces \QwenImagePipeline\ for standard generation, \QwenImageEditPipeline\ and \QwenImageEditInpaintPipeline\ for editing tasks, \QwenImageImg2ImgPipeline\ for image-to-image translation, and \QwenImageControlNetPipeline\ and \QwenImageControlNetInpaintPipeline\ for ControlNet-guided generation. The module also includes \QwenImageEditPlusPipeline\ for future feature upgrades and \QwenImageLayeredPipeline\ for layered support, all utilizing the \Qwen2\_5\_VL\ text encoder and \QwenImageTransformer2DModel\.

src/diffusers/pipelines/qwenimage · high confidence

Introduce Stable Diffusion 3 ControlNet pipelines

Adds \StableDiffusion3ControlNetPipeline\ and \StableDiffusion3ControlNetInpaintingPipeline\ to the \diffusers\ library, enabling users to apply ControlNet conditioning and inpainting capabilities to Stable Diffusion 3 models. These new pipelines support single and multiple ControlNets, IP-Adapter integration, and LoRA loading, allowing for precise image generation and editing guided by structural inputs.

_src/diffusers/pipelines/controlnet\sd3 · high confidence

Introduce Stable Diffusion 3 pipeline suite

Adds the Stable Diffusion 3 pipeline family, including \StableDiffusion3Pipeline\ for text-to-image generation, \StableDiffusion3Img2ImgPipeline\ for image-to-image transformation, and \StableDiffusion3InpaintPipeline\ for inpainting. These pipelines leverage the SD3Transformer2DModel, FlowMatchEulerDiscreteScheduler, and a multi-encoder text setup (CLIP and T5), and support LoRA loading, IP-Adapter integration, and single-file model loading.

_src/diffusers/pipelines/stable\_diffusion\3 · high confidence

Introduce Stable Diffusion pipeline module with ONNX support

The \src/diffusers/pipelines/stable\_diffusion\ directory now contains the core Stable Diffusion pipeline implementations, including text-to-image, image-to-image, and inpainting variants. This update adds dedicated ONNX pipeline classes (\OnnxStableDiffusionPipeline\, \OnnxStableDiffusionImg2ImgPipeline\, \OnnxStableDiffusionInpaintPipeline\) that allow users to run Stable Diffusion models using the ONNX Runtime, alongside a conversion script (\convert\_from\_ckpt.py\) to migrate original Stable Diffusion checkpoints to the diffusers format. The module also includes a \CLIPImageProjection\ model for image embedding projection and a README with usage examples for local and Hub-based model loading.

_src/diffusers/pipelines/stable\diffusion · high confidence

Introduce Stable Video Diffusion pipeline for image-to-video generation

Adds the \StableVideoDiffusionPipeline\ and its associated output class, enabling users to generate short video clips from a single input image. The pipeline integrates a temporal VAE (\AutoencoderKLTemporalDecoder\), a spatio-temporal UNet, and a CLIP image encoder, utilizing a \VideoProcessor\ for resizing and frame handling. It supports standard diffusion scheduling via \EulerDiscreteScheduler\ and includes basic PyTorch/XLA compatibility for the timestep device.

_src/diffusers/pipelines/stable\_video\diffusion · high confidence

Introduce T2I-Adapter pipelines for Stable Diffusion and SDXL

Adds \StableDiffusionAdapterPipeline\ and \StableDiffusionXLAdapterPipeline\ to the library, enabling text-to-image generation conditioned on structural inputs (such as sketches, depth maps, or color palettes) via the T2I-Adapter architecture. These pipelines integrate with existing features like LoRA, IP-Adapter, and custom timesteps, allowing users to guide the diffusion process with additional image-based constraints alongside text prompts.

_src/diffusers/pipelines/t2i\adapter · high confidence

Introduce Textual Inversion training examples for Stable Diffusion and SDXL

The \examples/textual\_inversion\ directory now provides complete, runnable training scripts (\textual\_inversion.py\ and \textual\_inversion\_sdxl.py\) along with documentation and tests. Users can fine-tune Stable Diffusion v1.5 and SDXL models on custom image datasets to learn new concepts. The PyTorch scripts support multi-vector embeddings, safetensors serialization, and checkpointing, while the Flax/JAX script is included for TPU/GPU acceleration (noting that Flax support was removed in diffusers v0.40.0).

_examples/textual\inversion · high confidence

Introduce TorchAO quantizer with safetensors support

Adds a new TorchAO quantizer implementation that enables loading and saving models quantized with the TorchAO library, including support for safetensors serialization for TorchAO versions 0.16.0 and above. The minimum required version of the torchao dependency is set to 0.15.0, and the quantizer registers specific safe globals to ensure compatibility with PyTorch's serialization mechanism.

src/diffusers/quantizers/torchao · high confidence

Introduce Wan video generation pipelines

Adds a new \src/diffusers/pipelines/wan\ module exposing five pipelines for Wan models: \WanPipeline\ (text-to-video), \WanImageToVideoPipeline\ (image-to-video), \WanVideoToVideoPipeline\ (video-to-video), \WanVACEPipeline\ (controllable generation with first/last frames and masks), and \WanAnimatePipeline\ (character animation/replacement using pose and face videos). The module includes a lazy-loading \\_\init\\_.py\, a \WanPipelineOutput\ dataclass, and a \WanAnimateImageProcessor\ for reference image preprocessing, enabling users to generate and manipulate videos with Wan models.

src/diffusers/pipelines/animatediff, src/diffusers/pipelines/helios, src/diffusers/pipelines/wan · high confidence

Introduce \`diffusers-cli\` with environment, custom block, and agentic run commands

The \diffusers-cli\ tool is now available, providing a unified command-line interface for managing Diffusers workflows. The new \env\ command prints a comprehensive report of the Diffusers version, dependencies (including quantization backends like bitsandbytes and GGUF), and hardware details to assist with bug reports. A \custom\_blocks\ command allows users to package local \ModularPipelineBlocks\ subclasses for the Hub. The \run\ command serves as a single agentic entry point to execute pipelines, supporting remote execution via Hugging Face Sandboxes and automatic handling of image/video inputs. Additionally, the \schema\ command introspects pipeline inputs without downloading weights, and the \skills\ command installs AI agent skill bundles for tools like Claude Code and Cursor. The legacy \fp16\_safetensors\ command is marked as deprecated in favor of native pipeline serialization options.

src/diffusers/commands · high confidence

Introduce dedicated ControlNet pipeline module

The \diffusers.pipelines.controlnet\ package is now a first-class module, providing a centralized location for all ControlNet-related pipelines. This includes the standard Stable Diffusion pipelines (\StableDiffusionControlNetPipeline\, \Img2Img\, \Inpaint\), their Stable Diffusion XL counterparts (\StableDiffusionXLControlNetPipeline\, \Img2Img\, \Inpaint\), and the new ControlNet Union variants (\ControlNetUnionPipeline\, \Img2Img\, \Inpaint\). The module also exposes \MultiControlNetModel\ for applying multiple ControlNets simultaneously and includes the \BlipDiffusionControlNetPipeline\ for subject-driven generation. This reorganization consolidates the ControlNet ecosystem into a single, easily importable namespace.

src/diffusers/pipelines/controlnet · high confidence

Introduce modular pipeline architecture for Cosmos3

The Cosmos3 video generation model now supports a modular pipeline structure, breaking the generation process into distinct, composable blocks for encoding, denoising, and decoding. This change introduces new entry points \Cosmos3OmniModularPipeline\ and \Cosmos3DistilledModularPipeline\, along with specialized blocks for handling vision, sound, and action modalities. Users can now leverage this modular design to customize specific stages of the generation workflow, such as text encoding, latent preparation, or video decoding, while maintaining compatibility with existing Cosmos3 features like transfer learning and distilled inference.

_src/diffusers/modular\pipelines/cosmos · high confidence

Introduce modular pipeline architecture for Flux and Flux Kontext models

Users can now compose Flux and Flux Kontext image generation workflows using a modular block-based system. This change adds new pipeline classes (\FluxModularPipeline\, \FluxKontextModularPipeline\) and a suite of reusable components—including text encoding, image preprocessing, latent preparation, denoising loops, and decoding—that allow for flexible, step-by-step construction of text-to-image and image-to-image generation tasks.

_src/diffusers/modular\pipelines/flux · high confidence

Introduce modular pipeline support for LTX-2 and LTX-2.5

Adds a new modular pipeline implementation for the LTX-2 and LTX-2.5 models, enabling users to construct and customize generation workflows using composable blocks. This includes \LTX2ModularPipeline\ and \LTX25ModularPipeline\ classes, along with dedicated block modules for encoding, denoising, and decoding (including a new diffusion-based video decoder for LTX-2.5). The update also introduces a Gemma-4-based prompt enhancement step for LTX-2.5 and supports joint video and audio generation workflows.

_src/diffusers/modular\pipelines/ltx2 · high confidence

Introduce new schedulers for AMUSED, block refinement, and consistency models

The scheduler module now includes three new scheduling strategies: the AmusedScheduler for masked token generation using confidence-based unmasking, the BlockRefinementScheduler for iterative block-wise token commitment with optional editing, and the ConsistencyDecoderScheduler for two-step consistency model decoding. These schedulers are registered in the module's lazy import structure and are available for use in compatible pipelines.

src/diffusers/schedulers · high confidence

Introduce official pipeline callbacks for CFG cutoff and IP-Adapter scaling

The library now provides a \PipelineCallback\ base class and specific implementations in \src/diffusers/callbacks.py\ to manage step-wise pipeline behavior. Users can now apply \SDCFGCutoffCallback\, \SDXLCFGCutoffCallback\, and \SDXLControlnetCFGCutoffCallback\ to disable classifier-free guidance after a specified step ratio or index, and \IPAdapterScaleCutoffCallback\ to zero out IP-Adapter scales. These callbacks are registered in the main \\_\init\\_.py\ and support chaining via \MultiPipelineCallbacks\, allowing for more granular control over the generation process without modifying pipeline source code.

src/diffusers · high confidence

Introduction of FreeInit and FreeNoise utilities for improved generation quality and memory efficiency

The \src/diffusers/pipelines\ area now includes \FreeInitMixin\ and \FreeNoiseTransformerBlock\ utilities. FreeInit allows users to enable a noise re-initialization mechanism (via \enable\_free\_init\) to improve generation quality, while FreeNoise provides a transformer block variant designed to reduce memory usage during inference by splitting inputs for more efficient processing.

src/diffusers/pipelines · high confidence

Introduction of the controlnet module with new model implementations

The diffusers library now includes a dedicated \controlnet\ module that centralizes ControlNet model implementations. This change introduces new model classes for various architectures, including \ControlNetModel\ for standard Stable Diffusion, \FluxControlNetModel\ for Flux, \SD3ControlNetModel\ for Stable Diffusion 3, \HunyuanDiT2DControlNetModel\ for Hunyuan, \QwenImageControlNetModel\ for Qwen-Image, \SanaControlNetModel\ for Sana, \CosmosControlNetModel\ for Cosmos Transfer2.5, and \ZImageControlNetModel\ for Z-Image. It also adds support for multi-controlnet scenarios with \MultiControlNetModel\ and \MultiControlNetUnionModel\, as well as specialized adapters like \ControlNetUnionModel\, \ControlNetXSAdapter\, and \SparseControlNetModel\. These models are now accessible via the \diffusers.models.controlnets\ package, enabling users to leverage ControlNet conditioning across a wider range of base models and architectures.

src/diffusers/models/controlnets · high confidence

Marigold pipeline now supports intrinsic image decomposition alongside depth and normals estimation

The Marigold pipeline module has been expanded to include a new \MarigoldIntrinsicsPipeline\ for intrinsic image decomposition, in addition to the existing \MarigoldDepthPipeline\ and \MarigoldNormalsPipeline\. This update introduces the \MarigoldIntrinsicsOutput\ class and the \MarigoldImageProcessor\ to handle the new modality, allowing users to extract properties like albedo, roughness, and metallicity from single images using the v1-1 model variants.

src/diffusers/pipelines/marigold · high confidence

New ACE-Step pipeline for text-to-music generation

Added the AceStepPipeline and its supporting model components (AceStepConditionEncoder, AceStepAudioTokenizer, etc.) to enable text-to-music generation. The pipeline supports multiple task types including text2music, cover, repaint, extract, lego, and complete, and integrates with the existing LoRA loading infrastructure via AceStepLoraLoaderMixin.

_src/diffusers/pipelines/ace\step · high confidence

New Amused training example with LoRA and 8-bit optimizer support

Added a new training example for the Amused model in the \examples/amused\ directory, including \train\_amused.py\ and \README.md\. This example provides recipes for finetuning Amused on datasets like nouns and Minecraft, supporting full finetuning, LoRA (Low-Rank Adaptation), and 8-bit Adam optimizers to enable training on lower-memory configurations (as low as 5.5 GB). The documentation includes specific hyperparameter recommendations and memory usage estimates for various batch sizes and gradient accumulation steps.

examples/amused · high confidence

New AutoencoderKL training example for CIFAR-10 and ImageNet

Added a new training script and documentation in the \examples/research\_projects/autoencoderkl\ directory, enabling users to fine-tune the AutoencoderKL model on datasets like CIFAR-10 and ImageNet. The example includes a \train\_autoencoderkl.py\ script with support for validation logging via Weights & Biases or TensorBoard, mixed-precision training, and Hugging Face Hub integration for model saving.

_examples/research\projects/autoencoderkl · high confidence

New DreamBooth training documentation for FLUX.2, Ideogram 4, Krea 2, and HiDream

The examples/dreambooth directory now includes dedicated README guides for training LoRA adapters on several new model families. Users can follow the new README\_flux2.md for FLUX.2 \[dev\] and FLUX 2 \[klein\], which details memory optimizations like remote text encoding, FSDP, and FP8/QLoRA quantization. The README\_ideogram4.md covers Ideogram 4, highlighting its JSON caption format, SDNQ FP8 base support, and the need to disable training autocast. The README\_krea2.md explains the workflow for Krea 2, specifically training on the RAW checkpoint and validating on the distilled Turbo checkpoint. Additionally, README\_hidream.md provides instructions for HiDream Image, including NF4 quantization support.

examples/dreambooth · high confidence

New Dreambooth inpainting training examples with LoRA support

Added \train\_dreambooth\_inpaint.py\ and \train\_dreambooth\_inpaint\_lora.py\ scripts to the research examples, enabling users to fine-tune Stable Diffusion inpainting models using Dreambooth. The standard script supports prior-preservation loss, gradient checkpointing, 8-bit optimizers, and optional text encoder training, while the LoRA variant introduces Low-Rank Adaptation for more memory-efficient fine-tuning. Both scripts include documentation on usage, required arguments, and hardware considerations.

_examples/research\_projects/dreambooth\inpaint · high confidence

New Flux Control training examples for LoRA and full fine-tuning

Added \train\_control\_lora\_flux.py\ and \train\_control\_flux.py\ scripts to the \examples/flux-control\ directory, enabling users to train Control LoRAs or perform full fine-tuning on the FLUX.1-dev model using structural conditions like depth maps or poses. The examples demonstrate expanding the model's input channels to incorporate control latents, with support for features such as gradient checkpointing, mixed precision, DeepSpeed optimization, and validation logging via WandB or TensorBoard.

examples/flux-control · high confidence

New Flux.1 LoRA quantization training example

Added a new research project example in \examples/research\_projects/flux\_lora\_quantization\ that demonstrates how to fine-tune the Flux.1 Dev model using LoRA and NF4 quantization. The example includes a training script (\train\_dreambooth\_lora\_flux\_miniature.py\) that leverages \bitsandbytes\ for 4-bit quantization, 8-bit Adam optimization, and gradient checkpointing to reduce memory usage, along with a helper script (\compute\_embeddings.py\) for precomputing text embeddings. It also provides configuration files for Accelerate and DeepSpeed Zero2 to support memory-optimized training workflows.

_examples/research\_projects/autoencoder\_rae, examples/research\_projects/flux\_lora\_quantization, examples/research\_projects/multi\_subject\dreambooth · high confidence

New GLIGEN training example for grounded text-to-image generation

Added a new research example in \examples/research\_projects/gligen\ that provides scripts to train the GLIGEN model on the COCO dataset. This includes \train\_gligen\_text.py\ for the training loop, \dataset.py\ for handling grounding data, \make\_datasets.py\ for preparing datasets using RAM, Grounding DINO, and BLIP2, along with a \demo.ipynb\ notebook and \README.md\ documentation to guide users through installation, data preparation, and inference.

_examples/research\projects/gligen · high confidence

New InstructPix2Pix LoRA fine-tuning example

Added a new training script and documentation for fine-tuning Stable Diffusion InstructPix2Pix using LoRA adapters. This example allows users to edit images based on text instructions by training Low-Rank Adaptation layers on the UNet, leveraging the PEFT library for efficient parameter-efficient fine-tuning.

_examples/research\_projects/instructpix2pix\lora · high confidence

New InstructPix2Pix training examples for Stable Diffusion and SDXL

Added \train\_instruct\_pix2pix.py\ and \train\_instruct\_pix2pix\_sdxl.py\ scripts, along with corresponding \README.md\ and \README\_sdxl.md\ documentation, enabling users to fine-tune Stable Diffusion and Stable Diffusion XL models for instruction-based image editing. The examples include support for multi-GPU distributed training via Accelerate, validation logging with Weights & Biases, and specific configuration options such as \--vae\_precision\ for SDXL to handle VAE stability. A test suite (\test\_instruct\_pix2pix.py\) is also included to verify checkpoint management behavior.

_examples/instruct\pix2pix · high confidence

New Latent Consistency Distillation examples for Stable Diffusion and SDXL

The examples/consistency\_distillation directory now includes training scripts and documentation for distilling Stable Diffusion and SDXL models using Latent Consistency Models (LCM). Users can now distill full models or train LCM-LoRAs using WebDataset-based pipelines (train\_lcm\_distill\_sd\_wds.py, train\_lcm\_distill\_sdxl\_wds.py) or the Hugging Face datasets library (train\_lcm\_distill\_lora\_sdxl.py). The location also provides a test suite (test\_lcm\_lora.py) to verify LCM-LoRA training and checkpointing behavior.

_examples/consistency\distillation · high confidence

New ONNX Runtime text-to-image fine-tuning example

Added a new experimental training script (\train\_text\_to\_image.py\) and documentation for fine-tuning Stable Diffusion models using ONNX Runtime to accelerate training. The example demonstrates how to configure the environment, authenticate with the Hugging Face Hub, and run training on the Naruto dataset with specific hyperparameters, while noting that fine-tuning the whole model may lead to overfitting or catastrophic forgetting.

_examples/research\_projects/onnxruntime/text\_to\image · high confidence

New PAG pipelines for Stable Diffusion, SDXL, HunyuanDiT, Kolors, and others

The \diffusers\ library now includes a dedicated \pag\ module providing Perturbed Attention Guidance (PAG) variants for multiple model architectures. This location introduces the \PAGMixin\ utility and specific pipeline classes—including \StableDiffusionPAGPipeline\, \StableDiffusionXLControlNetPAGPipeline\, \HunyuanDiTPAGPipeline\, and \KolorsPAGPipeline\—which allow users to enable PAG via \enable\_pag=True\ and configure guidance layers using \pag\_applied\_layers\ and \pag\_scale\ to improve image quality and reduce artifacts.

src/diffusers/pipelines/pag · high confidence

New SANA-Video pipelines for text-to-video and image-to-video generation

Added the \SanaVideoPipeline\ for text-to-video generation and the \SanaImageToVideoPipeline\ for image-to-video generation within the \diffusers\ library. These pipelines support the SANA model architecture, allowing users to generate video content from text prompts or animate static images, with support for aspect ratio binning (480p and 720p), LoRA loading, and configurable motion scores.

_src/diffusers/pipelines/sana\video · high confidence

New SD3 DreamBooth LoRA training example for low-VRAM environments

Added a new research project in \examples/research\_projects/sd3\_lora\_colab\ that enables Stable Diffusion 3 DreamBooth LoRA fine-tuning on hardware with under 16GB of GPU VRAM, such as free-tier Google Colab instances. The example includes a Colab notebook (\sd3\_dreambooth\_lora\_16gb.ipynb\), a Python script for pre-computing and serializing text embeddings (\compute\_embeddings.py\) to reduce memory usage, and a miniature training script (\train\_dreambooth\_lora\_sd3\_miniature.py\) that utilizes 8-bit T5 encoding, 8-bit Adam optimization, gradient checkpointing, and Flash Attention to fit the workload within consumer GPU constraints.

_examples/research\_projects/sd3\_lora\colab · high confidence

New SDXL ControlNet training script with WebDataset support

A new training script (\train\_controlnet\_webdataset.py\) has been added to the research projects examples, enabling users to train Stable Diffusion XL ControlNet models using WebDataset for efficient data loading. The script supports both Canny edge and depth map conditioning, utilizing \DPTImageProcessor\ for depth extraction and \CLIPImageProcessor\ for image processing, and includes built-in filtering for dataset quality based on image size and watermark probability.

_examples/research\projects/controlnet · high confidence

New Sana pipeline family for high-resolution text-to-image generation

This change introduces the Sana pipeline module, adding \SanaPipeline\ for text-to-image generation, \SanaControlNetPipeline\ for controlled generation, \SanaSprintPipeline\ for faster inference, and \SanaSprintImg2ImgPipeline\ for image-to-image tasks. These pipelines support resolutions up to 4096x4096 pixels, leverage the Gemma-2 text encoder, and include LoRA loading capabilities. The implementation also adds Torch XLA support for these pipelines and defines a dedicated \SanaPipelineOutput\ class for results.

src/diffusers/pipelines/sana · high confidence

New Stable Audio pipeline for text-to-audio generation

Users can now generate audio from text prompts using the StableAudio pipeline. This new feature introduces the \StableAudioPipeline\ class and the \StableAudioProjectionModel\ component, which handle conditioning on text and time boundaries (start/end seconds). The pipeline integrates with the \AutoencoderOobleck\ VAE and \StableAudioDiTModel\ transformer, supporting optional Torch XLA acceleration and requiring Transformers version 4.27.0 or later.

_src/diffusers/pipelines/stable\audio · high confidence

New VAE roundtrip example script

Added \vae\_roundtrip.py\ to the VAE research example directory, which demonstrates encoding an input image into latent space and decoding it back to a reconstructed image using either \AutoencoderKL\ or \AutoencoderTiny\. The script displays the original and reconstructed images side-by-side for visual comparison.

_examples/research\projects/vae · high confidence

New VQGAN training example with discriminator support

The examples/vqgan directory now includes a complete VQGAN training script (train\_vqgan.py) that supports training with a GAN discriminator (discriminator.py), allowing users to leverage perceptual and adversarial losses for higher-quality image synthesis. The example provides documentation on installing dependencies, configuring the VQModel architecture (including modifying block types and vocabulary size), and running training on datasets like CIFAR10. It also includes comprehensive tests (test\_vqgan.py) verifying model saving, checkpointing, and resumption functionality.

examples/vqgan · high confidence

New VisualCloze pipeline for image generation with visual in-context examples

Added the VisualCloze pipeline, which generates images based on visual in-context examples. This new capability includes the \VisualClozePipeline\ for combined generation and upsampling, the \VisualClozeGenerationPipeline\ for core generation, and the \VisualClozeProcessor\ for handling image preprocessing, resizing, and mask generation specific to the visual cloze task.

src/diffusers/pipelines/visualcloze · high confidence

New Wuerstchen text-to-image fine-tuning examples with LoRA support

Added new example scripts and documentation for fine-tuning the Würstchen text-to-image model, including support for Low-Rank Adaptation (LoRA) to efficiently adapt the prior model. The update introduces \train\_text\_to\_image\_lora\_prior.py\ for LoRA-based training, \train\_text\_to\_image\_prior.py\ for standard fine-tuning, and a custom \EfficientNetEncoder\ component, allowing users to customize the model on their own datasets with detailed setup instructions.

_examples/research\projects/wuerstchen · high confidence

New Z-Image pipeline family for image generation, inpainting, and ControlNet conditioning

Adds a new \z\_image\ pipeline module providing \ZImagePipeline\ for text-to-image generation, \ZImageImg2ImgPipeline\ for image-to-image transformation, \ZImageInpaintPipeline\ for inpainting, \ZImageControlNetPipeline\ and \ZImageControlNetInpaintPipeline\ for ControlNet-guided generation and inpainting, and \ZImageOmniPipeline\ which integrates a SigLIP-2 vision model for enhanced conditioning. All pipelines support LoRA loading via \ZImageLoraLoaderMixin\, use the \FlowMatchEulerDiscreteScheduler\, and expose a \ZImagePipelineOutput\ dataclass for results.

_src/diffusers/pipelines/z\image · high confidence

New async server example for concurrent Stable Diffusion 3 inference

The \examples/server-async\ directory now provides a complete FastAPI-based server example that enables safe, concurrent text-to-image inference using Stable Diffusion 3 and 3.5 models. This example introduces a \RequestScopedPipeline\ mechanism that keeps a single copy of the heavy model weights in memory while creating lightweight, per-request views. It ensures thread safety by cloning mutable state (such as the scheduler and RNG) and wrapping tokenizers, VAEs, and image processors with locks to prevent race conditions like 'Already borrowed' errors. The server exposes an inference endpoint at \/api/diffusers/inference\ and includes utilities for automatic model loading, memory management, and image saving.

examples/server-async · high confidence

New autoencoder module with expanded model support

The \src/diffusers/models/autoencoders\ directory has been restructured into a dedicated module, introducing a comprehensive \\_\init\\_.py\ that exposes a wide range of new autoencoder implementations. This update adds support for numerous models including AsymmetricAutoencoderKL, Cosmos3 AVAE Audio Tokenizer, Deep Compression Autoencoder (DC), Allegro, CogVideoX, Cosmos, Flux2, Hunyuan Video, KVAE, LTX Video, MiniMax H3, Mochi, Qwen-Image, LongCat Audio DiT, Oobleck, RAE, SAME, VidTok, ConsistencyDecoderVAE, LTX2 Video Diffusion Decoder, MiniMax Music 3 Vocoder, and VQModel. Users can now utilize these specific autoencoder classes for encoding and decoding latents in their respective pipelines.

src/diffusers/models/autoencoders · high confidence

New benchmarking suite for Flux, LTX, SDXL, and Wan models

A new benchmarking infrastructure has been added to the \benchmarks\ directory, enabling users to measure latency and memory usage for popular models including Flux, LTX-Video, SDXL, and Wan. The suite provides scripts to benchmark specific scenarios such as base performance with \torch.compile()\, NF4 quantization, layerwise upcasting, and group offloading. Results are collected into CSV files and can be automatically pushed to the \diffusers/benchmarks\ Hugging Face dataset via the \push\_results.py\ utility, with a CI workflow defined to run these benchmarks weekly.

benchmarks · high confidence

New documentation and repository quality checks

Added a suite of new utility scripts in the \utils/\ directory to enforce consistency and quality across the repository. These include \check\_ai.py\ to validate agent skill links and metadata, \check\_config\_docstrings.py\ to ensure config classes reference valid checkpoints, \check\_copies.py\ to verify copied code remains consistent with its source, \check\_doc\_toc.py\ to manage and sort the documentation table of contents, \check\_dummies.py\ to keep dummy backend objects in sync with imports, \check\_forward\_call\_docstrings.py\ to validate method signatures against docstrings, \check\_inits.py\ to analyze import structures, \check\_repo.py\ to enforce repository-wide model and test conventions, \check\_support\_list.py\ to ensure API classes are documented, and \check\_table.py\ to validate documentation tables. These tools help maintainers catch documentation drift, broken links, and structural inconsistencies automatically.

utils · high confidence

New documentation and test infrastructure for examples

The \examples\ directory now includes a comprehensive README that outlines the philosophy for official, community, and research examples, along with a table of supported training tasks (such as Unconditional Image Generation, Text-to-Image fine-tuning, and ControlNet) and links to Colab notebooks. Additionally, a new pytest configuration (\conftest.py\) and utility test module (\test\_examples\_utils.py\) have been added to support running and validating these example scripts.

examples · high confidence

New example for distillation-based quantization of Textual Inversion models

Added a new research example in \examples/research\_projects/intel\_opts/textual\_inversion\_dfq\ that demonstrates how to apply distillation for quantization to Textual Inversion models. The entry includes \textual\_inversion.py\ for training, \text2images.py\ for inference, and a \README.md\ with instructions. This allows users to generate smaller, INT8-quantized Stable Diffusion models personalized via Textual Inversion while maintaining quality through knowledge distillation.

_examples/research\_projects/intel\_opts/textual\_inversion\dfq · high confidence

New experimental parallelism and model architecture modules

The \src/diffusers/models\ package now includes new modules for experimental context and tensor parallelism (\\_modeling\_parallel.py\), a centralized activation function registry (\activations.py\), and a refactored attention system (\attention.py\, \attention\_dispatch.py\) that supports multiple backends like FlashAttention and SageAttention. Additionally, a new \MultiAdapter\ model is introduced to merge outputs from multiple T2I adapters with configurable weights, and an \experimental\ subpackage is added to host novel applications such as Reinforcement Learning pipelines.

src/diffusers/models · high confidence

New model conversion scripts for LDM, ACE-Step, Anima, AnyFlow, Asymmetric VQGAN, AuraFlow, and BLIP Diffusion

Added new conversion scripts in the \scripts\ directory to migrate external model checkpoints into the Diffusers format. These include \convert\_ace\_step\_to\_diffusers.py\ for ACE-Step audio models, \convert\_amused.py\ for MUSE, \convert\_anima\_to\_diffusers.py\ for Anima, \convert\animatediff\\*.py\ for AnimateDiff motion modules and LoRAs, \convert\_anyflow\_to\_diffusers.py\ for AnyFlow video models, \convert\_asymmetric\_vqgan\_to\_diffusers.py\ for Asymmetric VQGAN, \convert\_aura\_flow\_to\_diffusers.py\ for AuraFlow, and \convert\_blipdiffusion\_to\_diffusers.py\ for BLIP Diffusion. Additional scripts handle LDM uncond conversion (\conversion\_ldm\_uncond.py\), naming config updates (\change\_naming\_configs\_and\_checkpoints.py\), and SparseCtrl conversion (\convert\_animatediff\_sparsectrl\_to\_diffusers.py\).

scripts · high confidence

New model search example with Civitai and Hugging Face integration

The \examples/model\_search\ directory now includes a new example demonstrating how to use the \auto\_diffusers\ library to search for and load models from Civitai and the Hugging Face Hub. This addition provides \EasyPipeline\ classes that allow users to easily load models for text-to-image, image-to-image, and inpainting tasks directly from these sources, including support for loading LoRA and TextualInversion weights automatically.

_examples/model\search · high confidence

New modular Z-Image pipeline implementation

Adds a new modular pipeline implementation for the Z-Image model, introducing \ZImageModularPipeline\ and \ZImageAutoBlocks\ to support text-to-image and image-to-image workflows. This change includes dedicated pipeline blocks for text encoding (using Qwen2/Qwen3), latent preparation, denoising with classifier-free guidance, and VAE decoding, allowing users to compose and customize the generation process via the modular pipeline API.

_src/diffusers/modular\_pipelines/z\image · high confidence

New modular guider system with multiple guidance implementations

The \src/diffusers/guiders\ package introduces a new modular architecture for diffusion model guidance, replacing the previous flat structure with a dedicated folder and \\_\init\\_.py\ that exports a comprehensive set of guidance classes. This change adds several new capabilities for users: Adaptive Projected Guidance (APG) and a combined APG+CFG variant (AdaptiveProjectedMixGuidance) for improved image quality; AutoGuidance which applies skip-layer guidance via hooks; Frequency-Decoupled Guidance (FDG) that separates low- and high-frequency components for better detail retention; and ClassifierFreeZeroStarGuidance for optimized early-step noise prediction. Existing guidance methods like ClassifierFreeGuidance, PerturbedAttentionGuidance, SkipLayerGuidance, SmoothedEnergyGuidance, MagnitudeAwareGuidance, and TangentialClassifierFreeGuidance are also included in this new module, along with a \BaseGuidance\ utility class and \LTX2Guidance\ for specific model support. Users can now import and configure these diverse guidance techniques from a single, organized location.

src/diffusers/guiders · high confidence

New modular hook system for inference optimization and parallelism

The \src/diffusers/hooks\ module has been introduced to provide a unified, modular system for applying advanced inference optimizations and parallelism strategies to diffusion models. This new location exposes a registry-based hook architecture that enables users to apply features such as FasterCache, First Block Cache, SeaCache, MagCache, and TaylorSeer Cache to skip redundant computations, as well as Context Parallelism, Tensor Parallelism, and Group Offloading (including disk offloading) to manage memory and scale across devices. The module also supports Layerwise Casting, Layer Skipping, and Text KV Caching, allowing for fine-grained control over model execution behavior and resource usage.

src/diffusers/hooks · high confidence

New modular pipeline blocks for Wan 2.1 and Wan 2.2 video models

This change introduces a new modular pipeline implementation for the Wan family of video generation models, located in \src/diffusers/modular\_pipelines/wan\. It adds a set of reusable pipeline blocks (e.g., \WanBlocks\, \Wan22Blocks\, \Wan22Image2VideoBlocks\) and corresponding modular pipelines (e.g., \WanModularPipeline\, \Wan22Image2VideoModularPipeline\) that decompose the generation process into distinct steps such as text encoding, denoising, and decoding. The implementation supports both Wan 2.1 and Wan 2.2 architectures, covering text-to-video, image-to-video, and video-to-video (VACE) workflows, and is exposed via the \diffusers.modular\_pipelines.wan\ module for lazy loading.

_src/diffusers/modular\pipelines/wan · high confidence

New modular pipeline implementation for QwenImage models

Users can now access a modular pipeline architecture for QwenImage, QwenImage-Edit, QwenImage-Edit Plus, and QwenImage-Layered models. This change introduces a new directory structure under \src/diffusers/modular\_pipelines/qwenimage\ containing reusable pipeline blocks (encoders, decoders, denoisers, and input processors) and auto-pipeline classes that allow for more flexible composition and customization of the generation workflow compared to the previous monolithic pipeline implementation.

_src/diffusers/modular\pipelines/qwenimage · high confidence

New multi-subject Dreambooth inpainting example

Added a new research example that demonstrates how to fine-tune Stable Diffusion inpainting models for multiple subjects simultaneously. The entry includes a Python training script (\train\_multi\_subject\_dreambooth\_inpainting.py\) that supports loading multiple Hugging Face datasets, optional text encoder training, and Weights & Biases logging, along with a companion notebook for creating prompt-image-mask pairs.

_examples/research\_projects/multi\_subject\_dreambooth\inpainting · high confidence

New profiling examples for DiffusionPipeline optimization

Added a new \examples/profiling\ directory containing a CLI tool (\profiling\_pipelines.py\) and supporting utilities to profile popular Diffusion pipelines (Flux, Wan, LTX2, QwenImage) using \torch.profiler\. This tool helps users identify CPU-GPU synchronization bottlenecks and Python overhead to optimize pipelines for \torch.compile\, providing Chrome trace exports and summary tables for analysis.

examples/profiling · high confidence

New reinforcement learning examples for diffusion-based policy and locomotion

Added two new example scripts in the reinforcement learning directory: \diffusion\_policy.py\ and \run\_diffuser\_locomotion.py\, along with a \README.md\ documenting them. The diffusion policy example demonstrates a robot control model that uses a diffusion model to predict smooth movement trajectories for pushing a T-shaped block, requiring specific dependencies like \free-mujoco-py\ and \d4rl\. The locomotion example showcases the use of \ValueGuidedRLPipeline\ from the \diffusers.experimental\ module to run a hopper environment, allowing users to adjust guide steps for reward maximization versus pure sampling.

_examples/reinforcement\learning · high confidence

New server example for concurrent image generation

Added a new example server in \examples/server\ that demonstrates how to use Diffusers pipelines as an inference engine for concurrent, multithreaded image generation requests. The server, built with FastAPI and Uvicorn, exposes a \/v1/images/generations\ endpoint that accepts text prompts and returns generated image URLs. It uses \StableDiffusion3Pipeline\ and implements thread-safe scheduling by creating a new scheduler instance per request to avoid errors from sharing a single scheduler across threads, while reusing the shared model pipeline to minimize GPU memory usage.

examples/server · high confidence

New transformer model implementations added to the library

The \src/diffusers/models/transformers\ module has been populated with new transformer model classes, making them available for import and use. This includes the \\_\init\\_.py\ file which now exports a wide range of models such as \AceStepTransformer1DModel\, \AuraFlowTransformer2DModel\, \CogVideoXTransformer3DModel\, \DiTTransformer2DModel\, \HunyuanDiT2DModel\, \SD3Transformer2DModel\, \FluxTransformer2DModel\, and many others. Corresponding implementation files (e.g., \ace\_step\_transformer.py\, \auraflow\_transformer\_2d.py\, \cogvideox\_transformer\_3d.py\) have been added to provide the underlying architecture for these models.

src/diffusers/models/transformers · high confidence

New unconditional image generation training example and tests

This change introduces a new \examples/unconditional\_image\_generation\ directory containing \train\_unconditional.py\, a script for training DDPM UNet models on datasets like Oxford Flowers or Pokemon, along with comprehensive pytest tests in \test\_unconditional.py\ to verify training, checkpointing limits, and model saving. The example supports multi-GPU training via Accelerate, logging to Weights & Biases, and optional precision-preserving preprocessing for high-bit-depth images.

_examples/unconditional\_image\generation · high confidence

New unified quantization infrastructure with auto-discovery and pipeline-level configuration

The \src/diffusers/quantizers\ module has been restructured to provide a centralized, auto-discovering quantization system. A new \DiffusersAutoQuantizer\ class automatically resolves and instantiates the correct quantizer backend (such as bitsandbytes, GGUF, TorchAO, Quanto, or Nunchaku Lite) based on the model's configuration, simplifying the loading process for users. Additionally, a new \PipelineQuantizationConfig\ class enables on-the-fly, pipeline-level quantization, allowing users to specify which components of a diffusion pipeline to quantize and which backend to use directly in \from\_pretrained\. This change introduces a standardized base \DiffusersQuantizer\ interface and a unified \QuantizationConfigMixin\, replacing scattered quantization logic with a consistent, extensible architecture.

src/diffusers/quantizers · high confidence

PixArt-Alpha ControlNet example added

The \examples/research\_projects/pixart\ directory now includes a complete example for running PixArt-Alpha with ControlNet conditioning. This addition introduces the \PixArtControlNetAdapterModel\ and \PixArtControlNetTransformerModel\ classes to integrate ControlNet blocks into the PixArt architecture, a dedicated \PixArtAlphaControlnetPipeline\ for inference, and scripts (\train\_pixart\_controlnet\_hf.py\ and \run\_pixart\_alpha\_controlnet\_pipeline.py\) to fine-tune and run the model with edge-conditioned inputs.

_examples/research\projects/pixart · high confidence

Repository initialization with project scaffolding and documentation

The repository is initialized with essential project files including a comprehensive .gitignore, Apache 2.0 LICENSE, CITATION.cff, CODE\_OF\_CONDUCT.md, CONTRIBUTING.md, SECURITY.md, and a Makefile for development workflows. The README.md is updated with installation instructions, quickstart examples, and documentation links, while symlinks (AGENTS.md, CLAUDE.md, PHILOSOPHY.md) point to the .ai/ directory for AI agent conventions.

(repo-wide) · high confidence

Removals

Stable Cascade pipelines are deprecated

The Stable Cascade pipeline implementations (\StableCascadePriorPipeline\, \StableCascadeDecoderPipeline\, and \StableCascadeCombinedPipeline\) in \src/diffusers/pipelines/stable\_cascade\ are now marked as deprecated. These classes inherit from \DeprecatedPipelineMixin\ and include a \\_last\_supported\_version\ of \0.35.2\, indicating that they will be removed in a future release. Users relying on these pipelines for Stable Cascade image generation should prepare for migration to alternative solutions.

_src/diffusers/pipelines/stable\cascade · high confidence

Architecture

Diffusers utilities are reorganized into a structured package

The \src/diffusers/utils\ directory has been reorganized into a formal Python package with a dedicated \\_\init\\.py\ that explicitly exports all utility modules. This change introduces lazy-loading support via new \dummy\\*\ modules for optional backends (such as \auto\_round\, \bitsandbytes\, \gguf\, \nvidia\_modelopt\, and \optimum\_quanto\), ensuring that imports fail gracefully with clear error messages when dependencies are missing. Additionally, the package now includes new utility modules for distributed environment checks (\distributed\_utils\), documentation string replacement (\doc\_utils\), and accelerated offloading hooks (\accelerate\_utils\), while consolidating constants, deprecation logic, and state-dict conversion helpers into clearly defined submodules.

src/diffusers/utils · high confidence

Refactored loader module into a modular package structure

The \src/diffusers/loaders\ directory has been reorganized from a single monolithic file into a structured package with dedicated modules for specific capabilities. The public API is now exposed through \\_\init\\_.py\, which imports mixins from specialized files: \lora\_pipeline.py\ (containing model-specific LoRA loaders like \FluxLoraLoaderMixin\ and \WanLoraLoaderMixin\), \lora\_base.py\ (core LoRA logic and utilities), \lora\_conversion\_utils.py\ (state-dict conversion helpers), \ip\_adapter.py\ (IP-Adapter handling), \peft.py\ (PEFT adapter management via \PeftAdapterMixin\), and \single\_file.py\ (single-file checkpoint loading via \FromSingleFileMixin\). This change improves code maintainability and separation of concerns without altering the external loader API.

src/diffusers/loaders · high confidence

UNet models reorganized into a dedicated \`unets\` module

The UNet model implementations (including UNet1D, UNet2D, UNet2DCondition, UNet3DCondition, and others) have been moved from the top-level \models\ directory into a new \src/diffusers/models/unets\ subpackage. This structural change groups all UNet variants and their associated blocks together, and the \\_\init\\_.py\ file now exposes these models from the new location, which may require updating import paths if you reference these classes directly.

src/diffusers/models/unets · high confidence

Behavioural changes

Deprecate KarrasVeScheduler and ScoreSdeVpScheduler

The \KarrasVeScheduler\ and \ScoreSdeVpScheduler\ classes have been moved to the \src/diffusers/schedulers/deprecated\ module. Users relying on these schedulers will now encounter deprecation warnings, indicating that these specific variance-expanding and variance-preserving SDE schedulers are scheduled for removal in a future release and should be migrated to supported alternatives.

src/diffusers/schedulers/deprecated · high confidence

Deprecated multi-token textual inversion research example

The multi-token textual inversion example in this directory is now deprecated and users are directed to the official textual inversion example in the main examples folder, which natively supports multi-token features. The README explicitly states this deprecation and provides migration guidance. The code files (multi\_token\_clip.py, textual\_inversion.py, textual\_inversion\_flax.py) remain present but are marked as research projects that are no longer actively maintained.

_examples/research\_projects/multi\_token\_textual\inversion · high confidence

Deprecated pipelines moved to dedicated module with lazy loading and deprecation warnings

Pipelines with low usage (such as AltDiffusion, Amused, and others) have been moved into a new \src/diffusers/pipelines/deprecated\ directory. These pipelines are now exposed via a lazy-loading module structure that respects optional dependencies (Torch, Transformers, Librosa, Note-seq) and includes a \DeprecatedPipelineMixin\ to emit warnings when used. The \README\ in this folder clarifies that while these pipelines remain functional, they will no longer be tested or accepted for changes, signaling a shift in maintenance priority for these specific models.

src/diffusers/pipelines/deprecated · high confidence

Deprecation of legacy inference example scripts

The \image\_to\_image.py\ and \inpainting.py\ scripts in the \examples/inference\ directory are now deprecated and will be removed in a future version. Running these scripts now triggers a warning directing users to use the \StableDiffusionImg2ImgPipeline\ and \StableDiffusionInpaintPipeline\ classes directly from the \diffusers\ library instead. Users are advised to migrate to the official pipeline examples located in the \src/diffusers/pipelines\ folder.

examples/inference · high confidence

Diffusion ORPO example redirects to MaPO project

The training scripts for diffusion ORPO alignment in this directory are now deprecated; the README directs users to the MaPO project (mapo-t2i.github.io) for the official codebase, models, datasets, and paper.

_examples/research\_projects/diffusion\orpo · high confidence

IP-Adapter training examples moved to research projects

The IP-Adapter training scripts and documentation have been relocated to the \examples/research\_projects/ip\_adapter\ directory. This area now contains the \README.md\ and specific training tutorials for FaceID (\tutorial\_train\_faceid.py\), standard IP-Adapter (\tutorial\_train\_ip-adapter.py\), IP-Adapter Plus (\tutorial\_train\_plus.py\), and SDXL (\tutorial\_train\_sdxl.py\), providing users with the necessary code to fine-tune Stable Diffusion models using image prompts.

_examples/research\_projects/ip\adapter · high confidence

Stable Diffusion text-to-image training examples updated to v0.41.0.dev0

The training scripts and documentation in the \examples/text\_to\_image\ directory have been updated to require diffusers version 0.41.0.dev0. This update includes the \train\_text\_to\_image.py\ and \train\_text\_to\_image\_lora.py\ scripts for Stable Diffusion, as well as their SDXL counterparts (\train\_text\_to\_image\_sdxl.py\ and \train\_text\_to\_image\_lora\_sdxl.py\). The examples now support advanced training features such as Min-SNR weighting, EMA (Exponential Moving Average) tracking with CPU offloading, and DREAM training. The documentation has been expanded to cover these new options, along with improved guidance on using DeepSpeed for memory efficiency and PyTorch XLA for inference. Tests have been added to verify checkpointing behavior and LoRA weight loading for both standard and SDXL models.

_examples/text\_to\image · high confidence

Test coverage

Extensive test suite expansion and refactoring across 1348 commits

This massive update adds and updates tests for a wide variety of new and existing components, including pipelines (Wan, Ideogram4, LTX-2.X, Cosmos3, Qwen-Image, SANA, Flux, SD3, HunyuanVideo, etc.), models (transformers, autoencoders, schedulers), and features (LoRA, IP-Adapter, ControlNet, single-file loading, quantization with TorchAO/bitsandbytes, group offloading, torch.compile compatibility). The tests have been significantly refactored to use new mixin structures (PipelineTestMixin, LoraBaseMixin, etc.), made device-agnostic (supporting XPU, MPS, CPU, CUDA), and optimized for speed and reliability (reducing model sizes, fixing flaky tests, standardizing assertions).

tests · high confidence

Dependencies

Standardized dependency requirements for all example scripts

This change introduces explicit \requirements.txt\ files for every example and research project directory (such as Dreambooth, ControlNet, Flux, and CogVideo), replacing previously implicit or missing dependency declarations. By pinning or specifying minimum versions for core libraries like \transformers\, \accelerate\, \peft\, and \torchvision\, users can now reliably install the correct environment for each specific training or inference script without encountering version conflicts or missing package errors.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 67.

Lenses

  • Code Health 71
  • Architecture 95
  • Maturity 64
  • Readiness 66
  • Security 70

Changes since last survey

  • 300 commits — 233 feature/other, 67 fixes

By area

  • src/diffusers — 117 commits
  • tests/pipelines — 73 commits
  • docs/source — 20 commits
  • tests/models — 20 commits
  • .github/workflows — 12 commits
  • tests/lora — 11 commits
  • examples/dreambooth — 8 commits
  • tests/modular_pipelines — 8 commits
  • .ai/AGENTS.md — 4 commits
  • .ai/models.md — 4 commits
  • (root) — 3 commits
  • .ai/references — 3 commits
  • tests/quantization — 3 commits
  • .ai/modular.md — 2 commits
  • tests/others — 2 commits
  • .ai/skills — 1 commit
  • .ai/testing.md — 1 commit
  • docker/diffusers-pytorch-cuda — 1 commit
  • examples/cogview4-control — 1 commit
  • examples/community — 1 commit

Notable commits

  • fix: FIX LoRA tests warning about unexpected keys (#14476)
  • fix: Fix AuraFlow VAE dtype mismatch on pipeline reuse (#14184)
  • fix: Fix AuraFlow model parallelism device mismatch and update XPU IP-Adap… (#14273)
  • fix: Fix Cosmos3 Edge generator K normalization (#14246)
  • fix: Fix DreamLite legacy block type aliases (#14066)
  • fix: Fix FA3 varlen wrapper when hub kernel returns single tensor (#14102)
  • fix: Fix Helios auto offload decode (#14140)
  • fix: Fix Ideogram 4 Callback Handling and Tests (#14621)
  • fix: Fix Kohya UNet LoRA key conversion for conv_in/conv_out/time_embedding (#14006)
  • fix: Fix LoRA hot-swapping recompilation with different_shapes_for_compilation (#14297)
  • fix: Fix PriorTransformer Group Offloading Bug (#14695)
  • fix: Fix Wan and Motif video pipeline fast test failures (#14269)
  • fix: Fix batched DiffusionGemma adaptive stopping (#14386)
  • fix: Fix callback tensor inputs that are never bound in the denoising loop (#14416)
  • fix: Fix duplicated words and a misspelling in messages and docstrings (#14287)
  • fix: Fix local LoRA weight auto-discovery in offline mode (#14204)
  • fix: Fix model cuda tests (#13975)
  • fix: Fix model offloading and training tests + prevent examples timeout (#14091)
  • fix: Fix mutable default args in lora_base.py (#14064)
  • fix: Revert "deprecate dduf."
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/diffusers was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit a3e0b8ec235c27a6c17a21976daf7fd32d819d05 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.