Skip to content
CAI
Software that uses CAICheck a score

huggingface/peft

65.4

Weak · 19 September 2026

86.6k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is the PEFT library, a toolkit for Parameter-Efficient Fine-Tuning that enables users to adapt large pretrained models by training only a small subset of parameters. It provides a comprehensive registry of diverse adapter methods—including LoRA, IA3, BOFT, and numerous specialized variants—along with utilities for quantization, distributed training, and multi-adapter management. The library supports a wide range of model types, from language models to diffusion networks, allowing for efficient fine-tuning across various hardware accelerators.

How it got here

2022–2023 — Library initialization and tuner expansion

34 changes.

This period established the PEFT library's core architecture, including its modular tuner infrastructure, public API, and comprehensive test suite. It simultaneously expanded the library's capabilities by introducing numerous new adapter types such as AdaLoRA, LoHa, LoKr, and OFT, alongside extensive documentation and examples for diverse fine-tuning workflows.

2024–2025 — expansion of PEFT adapter methods

49 changes.

This period focused on integrating a wide variety of new parameter-efficient fine-tuning adapters, including BOFT, VeRA, HRA, and TinyLoRA, into the core library. It also involved adding specialized optimizers and comprehensive examples for these methods, alongside establishing robust benchmarking suites to compare their performance.

2026 — expansion of PEFT method catalog

33 changes.

This period focused on integrating a wide variety of new parameter-efficient fine-tuning adapters into the PEFT library, including UniLoRA, PSOFT, Lily, PEANuT, AdaMSS, HiRA, FRoD, GLoRA, DEFT, KaSA, Super-Tuning, and ShadowPEFT. Each new tuner was accompanied by dedicated example scripts and documentation to demonstrate usage across text, vision, and diffusion models. The work also included comprehensive distributed training tests and a benchmarking suite to evaluate and compare these diverse methods.

Features

Add BEFT (Bias-Efficient Fine-Tuning) example for low-data regimes

This change introduces a new example demonstrating Bias-Efficient Fine-Tuning (BEFT), a technique for efficiently fine-tuning large language models in low-data regimes by targeting specific bias terms. The example script (\beft\_finetuning.py\) shows how to apply the \BeftConfig\ to a sequence-to-sequence model (using \bigscience/mt0-small\) and fine-tune only the bias terms of the query, key, or value projections (defaulting to the value projection \v\). It uses the \gtfintechlab/financial\_phrasebank\_sentences\_allagree\ dataset, subsampling 500 training examples to illustrate performance in data-scarce scenarios, and includes a README with usage instructions and citation details.

_examples/beft\finetuning · high confidence

Add BOFT ControlNet example utility modules

The examples/boft\_controlnet/utils directory now includes a suite of helper modules to support the BOFT ControlNet training example. This includes argument parsing (args\_loader.py), dataset loading and preprocessing (dataset.py), a lightweight ControlNet model implementation (light\_controlnet.py), a custom inference pipeline (pipeline\_controlnet.py), memory tracking utilities (tracemalloc.py), and a modified UNet2DConditionModel (unet\_2d\_condition.py) that supports guided hints for ControlNet conditioning.

_examples/boft\controlnet/utils · high confidence

Add BOFT ControlNet fine-tuning example with XPU support

This location introduces a new example for fine-tuning Stable Diffusion using BOFT (Orthogonal Finetuning via Butterfly Factorization) for controllable generation with ControlNet. The package includes the training script (train\_controlnet.py), inference script (test\_controlnet.py), evaluation script (eval.py), and corresponding shell scripts, along with documentation. The implementation adds support for Intel XPU devices in the inference and evaluation scripts, allowing users to run the example on XPU hardware in addition to CUDA and CPU.

_examples/boft\controlnet · high confidence

Add C3A (Circular Convolution Adaptation) adapter

Introduces a new C3A adapter type that applies circular convolution adaptation to linear layers. Users can now configure and apply C3A via \C3AConfig\, specifying parameters such as \block\_size\ (which must divide input/output dimensions), \target\_modules\, and bias handling. The implementation registers \c3a\ as a valid PEFT method, allowing it to be used alongside existing adapters for efficient parameter-efficient fine-tuning.

src/peft/tuners/c3a · high confidence

Add CUDA kernel for fast block diagonal operations in BOFT

This change introduces a new CUDA extension (\fbd\) within the BOFT tuner to accelerate fast block diagonal matrix operations. The addition includes a C++ wrapper and a CUDA kernel implementation that handle both forward and backward passes for these specific tensor transformations, enabling more efficient computation for users employing the Butterfly Factorization tuning method.

src/peft/tuners/boft/fbd · high confidence

Add CorDA example for context-aware LoRA initialization

The \examples/corda\_finetuning\ directory now includes a complete example for CorDA (Context-Oriented Decomposition Adaptation), a method that initializes LoRA adapters using task-aware covariance matrices to improve fine-tuning performance. The example provides scripts for preprocessing (to collect covariance and perform SVD) and fine-tuning, supporting two modes: Knowledge-Preserved Mode (KPM) to maintain existing world knowledge, and Instruction-Previewed Mode (IPM) to accelerate convergence on new tasks. It also includes utilities for loading calibration datasets and documentation on reducing memory consumption during the preprocessing step.

_examples/corda\finetuning · high confidence

Add DEFT (Decompositional Efficient Fine-Tuning) adapter

Users can now apply the DEFT tuning method, which injects knowledge into frozen base weights via a low-rank projection and injection matrix. This new adapter supports configurable decomposition methods (ReLU or QR), an optional pure subspace removal mode (PaRa), and specific handling for Conv1D layers. The implementation includes a dedicated configuration class, layer logic for managing projection/injection parameters, and model integration to replace target modules with DEFT-aware layers.

src/peft/tuners/deft · high confidence

Add DEFT (Decompositional Efficient Fine-Tuning) example

Added a new example script and documentation for DEFT, a parameter-efficient fine-tuning method that adapts frozen weights by removing a learned sub-space and injecting a low-rank update. The example demonstrates how to configure and train a model using \DeftConfig\ (with options for \relu\ or \qr\ decomposition methods) and shows how to load the resulting adapter with \PeftModel\.

_examples/deft\finetuning · high confidence

Add FRoD fine-tuning examples for text and image classification

Added new example scripts in \examples/frod\_finetuning\ that demonstrate how to fine-tune models using the FRoD (Fast Rotational Diagonal) parameterization via the PEFT library. The \frod\_text\_classification.py\ script shows fine-tuning \bert-base-uncased\ on the GLUE SST-2 dataset, while \frod\_image\_classification.py\ demonstrates fine-tuning \clip-vit-base-patch32\ on the Stanford Cars dataset. Both examples configure separate learning rates for FRoD diagonal coefficients, sparse coefficients, and the classification head, and include instructions for running the examples and using local model/dataset mirrors.

_examples/frod\finetuning · high confidence

Add FourierFT adapter support

Introduces FourierFT, a new parameter-efficient fine-tuning method that updates model weights via learnable frequency spectra in the Discrete Fourier Transform domain. Users can now apply this technique by configuring \FourierFTConfig\ (with parameters like \n\_frequency\ and \scaling\) and registering the method via \register\_peft\_method\. The implementation includes \FourierFTModel\ for injection, \FourierFTLayer\ and \FourierFTLinear\ for handling the spectral weight updates and merging, and supports standard PEFT features such as target module selection, layer patterns, and adapter management.

src/peft/tuners/delora, src/peft/tuners/fourierft · high confidence

Add GraLoRA finetuning example

Added a new example script and documentation for GraLoRA (Granular Low-Rank Adaptation), a PEFT method that enhances low-rank adaptation expressivity and robustness to outlier activations. The \gralora\_finetuning.py\ script demonstrates how to fine-tune models using the \GraloraConfig\ from the PEFT library, supporting configuration of parameters such as rank, alpha, dropout, target modules, and the granularity factor \k\. Users can run this example to train models on datasets like \openassistant-guanaco\ and optionally push the results to the Hugging Face Hub.

(repo-wide) · high confidence

Add Householder Reflection Adaptation (HRA) method

Introduces HRA, a new parameter-efficient fine-tuning adapter that uses Householder reflections to modify model weights. This change adds the HRAConfig, HRAModel, and specific layer implementations (HRALinear, HRAConv2d) to the PEFT library, registering HRA as a selectable method alongside existing adapters like LoRA. Users can now apply HRA to linear and convolutional layers with configurable rank, Gram-Schmidt orthogonalization, and bias handling.

src/peft/tuners/hra · high confidence

Add KappaTune example demonstrating selective LoRA fine-tuning

Added an example script and documentation in \examples/KappaTune\ that compares standard LoRA, a no-adaptation baseline, and the new KappaTune strategy. The example uses the \find\_kappa\_target\_modules\ helper to select a subset of model parameters (via \top\_p\) for fine-tuning, aiming to reduce catastrophic forgetting on general knowledge (measured via WikiText perplexity) while adapting to a downstream task (GSM8K).

examples/KappaTune · high confidence

Add Lily PEFT finetuning example

Added a new example in \examples/lily\_finetuning\ that demonstrates how to fine-tune models using the Lily (Low-Rank Interconnected Adaptation Across Layers) PEFT method. The entry includes a \lily\_finetuning.py\ script and a \README.md\ explaining Lily's cross-layer parameter sharing mechanism, providing users with a ready-to-run workflow for swapping \LoraConfig\ for \LilyConfig\ to improve parameter efficiency.

_examples/lily\finetuning · high confidence

Add Lily low-rank interconnected adapter support

Introduces the Lily adapter, a new parameter-efficient fine-tuning method that shares A adapters across consecutive layers (controlled by \stride\_A\) and combines multiple B adapters via a learned router (controlled by \num\_B\). This feature adds the \LilyConfig\, \LilyLayer\, and \LilyModel\ classes to the PEFT library, registering 'lily' as a new tunable method. Users can now apply this architecture to supported linear layers to reduce trainable parameters while maintaining capacity through shared representations and expert combination.

src/peft/tuners/lily · high confidence

Add LoRA Dreambooth conversion scripts and Colab notebook

The examples/lora\_dreambooth directory now includes utilities to convert LoRA models between the PEFT format and the kohya\_ss format, along with a Colab notebook for running the training. Users can convert kohya\_ss-trained LoRAs to PEFT using convert\_kohya\_ss\_sd\_lora\_to\_peft.py and convert PEFT-trained LoRAs back to kohya\_ss format using convert\_peft\_sd\_lora\_to\_kohya\_ss.py, enabling interoperability between different training ecosystems. The new colab\_notebook.ipynb provides a ready-to-run environment for training LoRA Dreambooth models.

_examples/lora\dreambooth · high confidence

Add LyCORIS LoHa adapter support for Stable Diffusion models

This change introduces the LoHa (Low-Rank Hadamard Product) adapter, a new fine-tuning method based on the LyCORIS library, specifically targeting Stable Diffusion and SDXL models. The implementation adds \LoHaConfig\, \LoHaModel\, and \LoHaLayer\ classes to the \src/peft/tuners/loha\ package, enabling users to apply LoHa adapters to Linear, Conv1d, and Conv2d layers. Key features include support for effective convolution decomposition (FedPara), configurable rank and alpha patterns per layer, and dropout options for rank and modules. The adapter is registered as a new PEFT method named 'loha', allowing it to be used alongside existing methods like LoRA for model adaptation.

src/peft/tuners/loha · high confidence

Add MetaMathQA method comparison benchmark suite

Introduces a new benchmarking framework in method\_comparison/MetaMathQA for evaluating and comparing different PEFT methods on the MetaMathQA and GSM8K datasets. The suite includes a Makefile for automated experiment discovery and execution, a training script (run.py) with support for torch.compile, gradient checkpointing, and mixed precision, and utilities for data handling, tokenization, and result logging. It also provides a helper script (run\_with\_jobs.py) to run experiments on Hugging Face Jobs runners, enabling reproducible and version-controlled benchmarking of methods like LoRA, IA3, and others.

_method\comparison/MetaMathQA · high confidence

Add MonteCLoRA fine-tuning example for sequence classification

Added a new example script in \examples/monteclora\_finetuning\ that demonstrates how to fine-tune a sequence classification model (specifically RoBERTa on the GLUE MRPC dataset) using the MonteCLoRA adapter technique. The script integrates the \MontecloraConfig\ and \MonteCLoRATrainerMixin\ from the PEFT library to apply variational regularization during training, allowing users to configure parameters such as LoRA rank, alpha, target modules, and the number of Monte Carlo samples.

_examples/monteclora\finetuning · high confidence

Add Orthogonal Subspace Fine-tuning (OSF) continual learning example

A new example demonstrating Orthogonal Subspace Fine-tuning (OSF) for parameter-efficient continual learning has been added to the \examples/orthogonal\_subspace\_learning\ directory. This example shows how to train a model sequentially on multiple tasks (ScienceQA, NumGLUE, and FOMC) while preventing catastrophic forgetting by constraining parameter updates to be orthogonal to previously important directions. The provided \osf\_continual\_learning.py\ script and \utils.py\ module illustrate OSF's progressive budget allocation strategy, allowing users to compare OSF's retention capabilities against standard full fine-tuning baselines.

_examples/orthogonal\_subspace\learning · high confidence

Add PSOFT tuner for efficient orthogonal fine-tuning

Introduces the PSOFT (Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation) adapter, allowing users to insert an r\*r orthogonal transformation R between low-rank matrices A and B. This new capability lets you fine-tune models by training only the orthogonal transformation and optional scaling vectors, while keeping the principal subspace frozen. The implementation includes the configuration class (PsoftConfig), the layer logic (OrthLayer and PsoftLayer), and the model integration (PsoftModel), registered under the 'psoft' method name.

src/peft/tuners/psoft · high confidence

Add Poly Adapter for multi-task fine-tuning

Introduces the Poly adapter, a new parameter-efficient fine-tuning method that enables multi-task learning by routing inputs through a mixture of LoRA experts. Users can now configure the number of tasks, skills (LoRAs), and splits via \PolyConfig\, allowing the model to dynamically mix adapter outputs based on provided \task\_ids\ during forward passes.

src/peft/tuners/poly · high confidence

Add SHiRA adapter finetuning example

Users can now finetune models using Sparse High Rank Adapters (SHiRA) via the new \examples/shira\_finetuning\ directory. This includes a Python script (\shira\_finetuning.py\) and documentation demonstrating how to configure \ShiraConfig\, apply sparse masking (including custom mask functions), and train with the Hugging Face Trainer, offering an alternative to low-rank adapters for vision and language tasks.

_examples/shira\finetuning · high confidence

Add VB-LoRA adapter support

Introduces a new VB-LoRA (Variational Bayes Low-Rank Adaptation) tuning method, allowing users to apply this specific adapter type to Hugging Face models. The implementation includes the \VBLoRAConfig\ for configuring parameters such as rank, vector bank size, and top-K selection, the \VBLoRAModel\ tuner for injecting the adapters into the model architecture, and the \VBLoRALayer\ and \Linear\ classes that handle the actual weight transformations and merging logic.

src/peft/tuners/vblora · high confidence

Add WaveFT parameter-efficient fine-tuning method

Introduces WaveFT (Wavelet-based Fine-Tuning), a new adapter method that leverages the sparsity of wavelet transforms to efficiently fine-tune pretrained models. This update adds the \WaveFTConfig\, \WaveFTLayer\, and \WaveFTModel\ classes to the \peft.tuners.waveft\ module, registering the method for use. Users can now apply wavelet-based adaptation by specifying parameters such as \n\_frequency\ (number of learnable coefficients), \wavelet\_family\ (e.g., 'db1', 'sym2'), and \scaling\, allowing for controlled parameter updates via Discrete Wavelet Transforms (DWT) and Inverse DWT (IDWT).

src/peft/tuners/waveft · high confidence

Add X-LoRA inference example using mistral.rs

Added a new example script \xlora\_inference\_mistralrs.py\ and accompanying documentation in \examples/xlora/\ that demonstrates how to perform inference with X-LoRA models using the mistral.rs engine. The example shows how to configure the \Runner\ with X-LoRA specific parameters, such as the ordering file and non-granular scaling index, to leverage mistral.rs features like dual-KV cache and continuous batching for improved throughput.

examples/xlora · high confidence

Add aLoRA finetuning example with invocation token support

Added a new example script and documentation for finetuning models using Activated LoRA (aLoRA). This example allows users to configure an invocation string (e.g., "\[/INST\]") so that the adapter weights remain inactive until that specific token sequence is encountered, enabling KV cache reuse with the base model. The script supports training on Hugging Face datasets, 4-bit quantization, and pushing results to the Hub, with device detection for both CUDA and XPU accelerators.

_examples/alora\finetuning · high confidence

Added example demonstrating DoRA loading with ephemeral GPU offloading

A new example script (\load\_with\_dora.py\) has been added to the \examples/ephemeral\_gpu\_offloading\ directory. This script demonstrates how to load a model with a DoRA adapter using the ephemeral GPU offloading feature, allowing users to compare loading times between pure CPU and offloaded scenarios.

_examples/ephemeral\_gpu\offloading · high confidence

Added placeholder directories for image generation results

Empty placeholder directories (\.gitkeep\) have been added to \method\_comparison/image-gen/cancelled\_results\, \method\_comparison/image-gen/sample-images/cancelled\_results\, \method\_comparison/image-gen/sample-images/results\, \method\_comparison/image-gen/sample-images/temporary\_results\, and \method\_comparison/image-gen/temporary\_results\. These ensure the directory structure is preserved in version control, likely to support the storage of benchmark output files for cancelled, sample, and temporary image generation tasks.

_method\_comparison/image-gen/cancelled\_results, method\_comparison/image-gen/sample-images, method\_comparison/image-gen/temporary\results · high confidence

Initial repository structure and project configuration

The repository is initialized with the core project configuration files, including \setup.py\ (defining the \peft\ package, version 0.21.1.dev0, and dependencies like \torch\, \transformers\, and \accelerate\), \README.md\ (providing documentation and usage examples for Parameter-Efficient Fine-Tuning), and \Makefile\ (defining targets for code quality, styling, and testing). Additionally, development tooling is configured via \.gitignore\, \.pre-commit-config.yaml\ (using ruff), and \SECURITY.md\ (outlining vulnerability reporting procedures).

(repo-wide) · high confidence

Introduce AdaLoRA adapter support with adaptive rank allocation

Users can now apply the AdaLoRA (Adaptive LoRA) parameter-efficient fine-tuning method. This adds the \AdaLoraConfig\ and \AdaLoraModel\ classes, which implement adaptive rank allocation by decomposing weight matrices into singular values and vectors (lora\_A, lora\_B, lora\_E) and dynamically adjusting rank budgets during training via a \RankAllocator\. The implementation includes specific layer classes for standard linear layers (\SVDLinear\), 8-bit quantized layers (\SVDLinear8bitLt\), 4-bit quantized layers (\SVDLinear4bit\), and GPTQ-quantized layers (\SVDQuantLinear\), and registers \adalora\ as a selectable PEFT method.

src/peft/tuners/adalora · high confidence

Introduce AdaMSS parameter-efficient fine-tuning method

Adds the AdaMSS (Adaptive Multi-Subspaces) tuner, a new parameter-efficient fine-tuning method that decomposes weight matrices using SVD and clusters the decomposed space into multiple trainable subspaces. This location provides the core implementation including the \AdamssConfig\ for configuration (e.g., rank, number of subspaces, Adaptive Subspace Allocation settings), the \AdamssLayer\ and \AdamssModel\ for applying the adaptation to \torch.nn.Linear\ modules, and the \AdamssAsaCallback\ for integrating with the HuggingFace \Trainer\. The method is registered as a new PEFT method named 'adamss' and requires \scikit-learn\ for the clustering step.

src/peft/tuners/adamss · high confidence

Introduce Adaption Prompt tuner for Llama, Mistral, and GPT-2 models

This change adds the Adaption Prompt adapter implementation to PEFT, enabling users to fine-tune Llama, Mistral, and GPT-2 models by injecting trainable prompts into the top attention layers. The new \AdaptionPromptModel\ and \AdaptedAttention\ modules handle the insertion of these prompts, while the \AdaptionPromptConfig\ allows specifying the number of adapter tokens and layers. The implementation includes model-specific logic for computing query states and handling rotary embeddings across various \transformers\ library versions, ensuring compatibility with both older and newer releases.

_src/peft/tuners/adaption\prompt · high confidence

Introduce BOFT (Butterfly Orthogonal Finetuning) adapter

Adds support for the BOFT parameter-efficient fine-tuning method, based on the ICLR 2024 paper 'Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization'. This new adapter allows users to apply orthogonal transformations using butterfly factorization to \torch.nn.Linear\ and \torch.nn.Conv2d\ layers. The implementation includes a configuration class (\BOFTConfig\) with parameters for block size, butterfly factors, and dropout, as well as a CUDA-accelerated extension for efficient block-diagonal operations. The adapter is registered in the PEFT framework and can be used via \get\_peft\_model\.

src/peft/tuners/boft · high confidence

Introduce Context-aware Prompt Tuning (CPT) adapter

Users can now apply Context-aware Prompt Tuning (CPT) to causal language models via the new \cpt\ PEFT method. This feature adds \CPTConfig\ and \CPTEmbedding\ classes that support token type masks, loss weighting with exponential decay, and epsilon-based projection for delta embeddings, enabling more controlled prompt tuning for sequence generation tasks.

src/peft/tuners/cpt · high confidence

Introduce FRoD adapter for sparse rotation-based parameter-efficient fine-tuning

Adds the FRoD (Factorized Rotation Decomposition) adapter, a new PEFT method that applies sparse trainable rotation matrices to linear layers. Users can now configure FRoD via \FrodConfig\ (supporting options like \sparse\_rate\, \runtime\_offload\_base\_weight\, and \layers\_pattern\) and apply it to models using the \frod\ method. The implementation includes \FrodLayer\ for per-layer adaptation and \FrodModel\ for managing shared projection buffers across module categories, enabling efficient fine-tuning with reduced trainable parameters compared to dense adapters.

src/peft/tuners/frod · high confidence

Introduce GLoRA adapter support with safe weight merging

Users can now apply the GLoRA (Generalized Low Rank Adapter) technique to linear layers in their models. This new adapter type allows for flexible parameterization of weight and bias corrections via \config\_A\_B\, \config\_C\, and \config\_D\_E\ settings (supporting low-rank, vector, constant, or none modes). The implementation includes a fix to the \safe\_merge\ process to ensure that merging multiple adapters does not mutate the base weights incorrectly, preventing sequential composition errors where subsequent adapters are computed against already-modified weights rather than the original base.

src/peft/tuners/glora · high confidence

Introduce HiRA adapter support for linear, convolutional, and embedding layers

Added the HiRA (Hierarchical Rank Adaptation) tuner to PEFT, enabling users to apply HiRA adapters to Linear, Conv1d/2d/3d, and Embedding layers. The implementation includes a dedicated HiraConfig for tuning parameters like rank and dropout, layer-specific modules (HiraLayer, Linear, Conv2d, etc.) that handle weight merging and unmerging, and a HiraModel wrapper for integrating with pretrained models. Additionally, 8-bit and 4-bit quantized linear layers (Linear8bitLt, Linear4bit) are supported via bitsandbytes integration, allowing HiRA adapters to be used with quantized models.

src/peft/tuners/hira · high confidence

Introduce LayerNorm (LN) Tuning adapter

Users can now apply LayerNorm tuning to transformer models by using the new \LNTuningConfig\ and \LNTuningModel\ classes. This feature, registered as the \ln\_tuning\ PEFT method, replaces specified target modules (or all modules if not specified) with tunable LayerNorm copies. The implementation supports single-adapter inference, allows excluding specific modules, and handles adapter merging by swapping the base layer with the tuned LayerNorm layer.

_src/peft/tuners/ln\tuning · high confidence

Introduce LoKr adapter support for Stable Diffusion models

Adds the LoKr (Low-Rank Kronecker Product) adapter implementation to PEFT, enabling efficient fine-tuning of Stable Diffusion and SDXL models. This new feature includes the \LoKrConfig\, \LoKrModel\, and \LoKrLayer\ classes, which support linear and convolutional layers (Conv1d/Conv2d) with options for effective decomposition and LyCORIS-style initialization, and registers the method for use via the standard PEFT interface.

src/peft/tuners/lokr · high confidence

Introduce MiSS adapter as a replacement for Bone

Adds the MiSS (Householder reflection adaptation) tuner to PEFT, providing a new method to adapt models alongside existing options like LoRA. This feature includes the \MissConfig\ for configuration (e.g., rank, dropout, initialization modes like 'bat' and 'mini'), the \MissLayer\ and \MissLinear\ implementations for applying the adapter to linear layers, and the \MissModel\ tuner class for integrating the adapter into pretrained models. The adapter is registered under the name 'miss' and supports generic quantization backends.

src/peft/tuners/miss · high confidence

Introduce Multitask Prompt Tuning adapter

Adds a new Multitask Prompt Tuning adapter that enables initializing prompt embeddings from source task weights. Users can now configure this method via \MultitaskPromptTuningConfig\, selecting initialization strategies such as \RANDOM\, \TEXT\, \AVERAGE\_SOURCE\_TASKS\, \EXACT\_SOURCE\_TASK\, or \ONLY\_SOURCE\_SHARED\. The implementation supports loading source state dictionaries from both standard PyTorch and safetensor formats, allowing for transfer learning from pretrained source prompts to downstream tasks.

_src/peft/tuners/multitask\_prompt\tuning · high confidence

Introduce Orthogonal Finetuning (OFT) adapter

Adds a new Orthogonal Finetuning (OFT) adapter to PEFT, allowing users to apply orthogonal transformations to linear, convolutional, and embedding layers. The implementation includes OFTConfig for tuning parameters like rank, block size, and constrained variants (COFT), OFTLayer for the core rotation logic, and OFTModel to integrate the adapter into pretrained models. This feature is registered as a standard PEFT method named 'oft'.

src/peft/tuners/oft · high confidence

Introduce Orthogonal Subspace Fine-Tuning (OSF) tuner

Users can now apply the Orthogonal Subspace Fine-Tuning (OSF) method for parameter-efficient continual learning. This new tuner preserves the top singular directions of weight matrices (the 'high' subspace) while training only the remaining low-rank components, helping to mitigate catastrophic forgetting. The implementation includes the OSFConfig for specifying the effective rank and target modules, the OSFLayer for handling SVD decomposition and gradient projection, and the OSFModel to integrate the tuner into the PEFT workflow.

src/peft/tuners/osf · high confidence

Introduce PEANuT adapter for weight-aware parameter-efficient fine-tuning

Adds the PEANuT (PEANut) adapter, a new parameter-efficient fine-tuning method that conditions the delta weight on the base layer's weights. Users can now apply PEANuT to supported transformer models via \PeanutConfig\, configuring parameters such as rank (\r\), depth of hidden adapter layers, activation functions, and scaling. The implementation registers \peanut\ as a new \PeftType\, allowing it to be used alongside existing adapters like LoRA, with support for merging weights and handling residual encoder/decoder blocks within the target modules.

src/peft/tuners/peanut · high confidence

Introduce PEFT Method Comparison Gradio app with secure filtering and Pareto analysis

A new interactive Gradio application is added to compare the performance of different PEFT methods across tasks like MetaMathQA and image generation. The app visualizes results using Pareto frontier plots, allowing users to filter experiments by task, model, and metrics. To ensure security, the app uses a custom AST-based sanitizer for filtering instead of unsafe query methods, and includes tests to verify that arbitrary code execution attempts are blocked.

_method\comparison · high confidence

Introduce PEFT Shop Gradio app for browsing and comparing PEFT methods

Adds a new Gradio-based 'PEFT Shop' application that allows users to browse Parameter-Efficient Fine-Tuning (PEFT) methods as if in an online store. The app enables filtering methods by capabilities (such as merging, multi-adapter support, and quantization backends) and by minimum star ratings derived from benchmark results. Users can view benchmark metrics (e.g., test accuracy, DINO similarity) for benchmarks like MetaMathQA and image generation, add methods to a cart to view usage code snippets and feature comparison tables, and access official documentation. The app automatically extracts method descriptions from PEFT documentation and aggregates benchmark data from the method comparison suite.

_method\comparison/peft-shop · high confidence

Introduce PEFT text generation benchmarking suite

A new benchmarking framework has been added to the \method\_comparison/text\_generation\_benchmark\ directory to measure inference performance, memory usage, and parameter efficiency for Parameter-Efficient Fine-Tuning (PEFT) methods. The suite includes a main runner (\run.py\) and a base model caching script (\run\_base.py\) to establish consistent baselines. It supports configurable model precision (float16, float32, bfloat16), 4-bit and 8-bit quantization via BitsAndBytes, and evaluates performance across short, medium, and long prompt categories. The implementation explicitly supports Intel XPU devices alongside CUDA, and uses a structured JSON output format to log detailed metadata, hardware information, and per-category metrics.

_method\_comparison/text\_generation\benchmark · high confidence

Introduce RoAd (2D Rotary Adaptation) as a new PEFT tuning method

Adds the RoAd adapter, a new parameter-efficient fine-tuning method that applies 2D rotations and learnable scales to linear layer weights. This change introduces the \RoadConfig\ (with variants \road\_1\, \road\_2\, \road\_4\ and configurable \group\_size\), the \RoadLayer\ implementation for standard and bitsandbytes quantized (8-bit/4-bit) linear modules, and the \RoadModel\ tuner class that registers \road\ as a new PEFT method. Users can now apply RoAd adapters to models to improve training efficiency or performance, with support for mixed-adapter inference and safe merging of adapter weights into base layers.

src/peft/tuners/road · high confidence

Introduce SHiRA sparse high-rank adapter support

Users can now apply the SHiRA (Sparse High-Rank Adapter) tuning method to linear layers in their models. This new adapter type uses a sparse mask to select a subset of weights for adaptation, with the number of parameters scaled by r(m+n) to match LoRA sizing. The implementation includes a configuration class (ShiraConfig) to control parameters like rank, mask type (defaulting to random sparse), and target modules, as well as model integration that registers SHiRA as a valid PEFT method for injection and state-dict handling.

src/peft/tuners/shira · high confidence

Introduce ShadowPEFT adapter for parallel shadow-network adaptation

Adds the ShadowPEFT tuner to the PEFT library, enabling users to augment a frozen base model with a small, trainable parallel 'shadow' network. This mechanism injects discrepancy signals into targeted transformer blocks and advances a shadow state via gated residual updates, allowing adaptation without merging weights into the base model. The implementation includes a \ShadowConfig\ for tuning parameters (such as rank, alpha, and dropout), a \ShadowModel\ wrapper, and specialized components like \ShadowCache\ for paired KV-cache handling during autoregressive generation. It also provides architecture-aware backends for Diffusers models (e.g., Flux2) and a generic token-wise MLP fallback for other diffusion transformers.

src/peft/tuners/shadow · high confidence

Introduce Super-Tuning PEFT method with sparse weight support and optional LoRA hybrid

A new Super-Tuning adapter is now available in PEFT, implementing the method from arXiv:2607.09287. This approach freezes base model weights and trains only a sparse subset of scalar entries selected by weight magnitude, configurable via the \sparsity\ parameter (default 0.99). Users can also enable the 'Supra' hybrid mode by setting a LoRA rank (\r\), which adds low-rank A/B parameters that are composed additively with the sparse support. The implementation includes \SupertuningConfig\ for configuration, \SupertuningLayer\ and \Linear\ for layer adaptation, and \SupertuningModel\ for model integration, with support for saving precomputed sparse indices to reduce checkpoint size or reconstructing them deterministically at load time.

src/peft/tuners/supertuning · high confidence

Introduce TinyLoRA parameter-efficient fine-tuning adapter

Adds the TinyLoRA adapter, a new parameter-efficient fine-tuning method based on the paper "Learning to Reason in 13 Parameters" (arXiv:2602.04118). This feature allows users to adapt models using SVD decomposition of frozen weights and a tiny trainable vector projected through fixed random tensors. The implementation includes the \TinyLoraConfig\ for specifying parameters like SVD rank (\r\), trainable vector dimension (\u\), and weight tying (\weight\_tying\), the \TinyLoraModel\ for injecting these adapters into pretrained models, and the \TinyLoraLayer\ (with \Embedding\ and \Linear\ variants) for the actual layer replacement. The adapter supports features such as deterministic projection seeding, optional saving of projection tensors for reproducibility, and configurable initialization of the trainable vectors.

src/peft/tuners/gralora, src/peft/tuners/randlora, src/peft/tuners/tinylora · high confidence

Introduce Trainable Tokens tuner for efficient embedding adaptation

A new \TrainableTokens\ PEFT method is introduced, allowing users to train only specific token embeddings (identified by indices) rather than the entire embedding matrix. This reduces memory and storage overhead when adding new tokens or fine-tuning existing ones. The implementation includes a dedicated \TrainableTokensConfig\ and \TrainableTokensLayer\, supports weight tying for models with tied embeddings, and handles initialization strategies (copying original weights vs. random initialization). It is registered as a standard PEFT adapter method.

_src/peft/tuners/trainable\tokens · high confidence

Introduce UniLoRA adapter support in PEFT

Adds the UniLoRA tuner to the PEFT library, enabling users to apply the Uni-LoRA adaptation method (based on a single shared vector) to compatible models. This change introduces the \UniLoraConfig\ for specifying parameters such as rank, shared vector length, and dropout, registers the \UniLoraModel\ and \UniLoraLayer\ classes, and implements the logic for injecting these adapters into transformer models, including handling of shared state and index assignment.

src/peft/tuners/pvera, src/peft/tuners/unilora · high confidence

Introduce VeRA (Vector-based Random Matrix Adaptation) tuning method

Adds support for the VeRA parameter-efficient fine-tuning method, which uses shared random projection matrices (vera\_A and vera\_B) and per-layer trainable scaling vectors (vera\_lambda\_b and vera\_lambda\_d) to reduce memory overhead compared to LoRA. This change introduces the VeraConfig, VeraLayer, and VeraModel components, registers the 'vera' method within the PEFT framework, and allows users to apply VeRA adapters to linear layers with configurable rank, dropout, and PRNG initialization keys.

src/peft/tuners/vera · high confidence

Introduce X-LoRA, a Mixture-of-Experts adapter for dynamic LoRA scaling

Adds the X-LoRA tuner, which enables a Mixture-of-Experts approach by dynamically selecting and scaling LoRA adapters based on input context. This feature introduces a new \XLoraConfig\ for configuring expert selection (including top-k and softmax options), an \XLoraClassifier\ to predict scaling factors from hidden states, and wrapper layers (\XLoraLinearLayer\, \XLoraEmbeddingLayer\) that apply these scalings to standard LoRA operations. The implementation registers \xlora\ as a new PEFT method, allowing users to load multiple LoRA adapters and have the model automatically route inputs to the most relevant experts.

src/peft/tuners/xlora · high confidence

Introduction of (IA)^3 adapter support for linear and convolutional layers

The PEFT library now includes the (IA)^3 (Infused Adapter by Inhibiting and Amplifying Inner Activations) tuning method. This change adds the \IA3Model\, \IA3Config\, and corresponding layer implementations (\Linear\, \Conv2d\, \Conv3d\) to the \src/peft/tuners/ia3\ package. Users can now apply (IA)^3 adapters to target modules, with specific support for feedforward modules where scaling is applied to inputs rather than outputs. The implementation also includes lazy imports for Bitsandbytes to support 8-bit and 4-bit quantized layers (\Linear8bitLt\, \Linear4bit\) when available, allowing (IA)^3 to be used with quantized models.

src/peft/tuners/ia3 · high confidence

Introduction of Prefix Tuning adapter with zero-initialization support

The Prefix Tuning method is now available as a distinct PEFT adapter. Users can configure it via \PrefixTuningConfig\, which includes a new \init\_weights\ option allowing weights to be initialized to zero (ensuring activations start as a no-op) in addition to the default random initialization. The implementation registers the \PrefixEncoder\ model class and exposes the configuration and encoder through the \src/peft/tuners/prefix\_tuning\ module.

_src/peft/tuners/prefix\tuning · high confidence

LoRA tuner restructured with new initialization methods and quantization support

The LoRA tuner module has been reorganized into a dedicated subpackage, introducing support for several new initialization strategies including Arrow, VeLoRA, CorDA, BD-LoRA, and MonteCLoRA. The update adds native integration with quantization frameworks such as BitsAndBytes (8-bit/4-bit), AWQ, AQLM, and GPTQModel, enabling LoRA adapters on quantized linear layers. A new conversion utility allows non-LoRA PEFT adapters to be converted into LoRA format for broader compatibility, while the module now exposes specific layer types like Conv2d, Conv3d, and Embedding for targeted fine-tuning.

src/peft/tuners/lora · high confidence

Mixed adapter models support mixing different adapter types

Users can now apply different adapter types (LoRA, LoHa, LoKr, AdaLoRA, OFT, and SHiRA) to different layers within the same model by using the \MixedModel\ tuner. This capability allows for more flexible model tuning strategies where specific layers can be optimized with the most suitable adapter technique, rather than being restricted to a single adapter type across the entire model.

src/peft/tuners/mixed · high confidence

New AdaMSS fine-tuning examples with Adaptive Subspace Allocation

Added example scripts and documentation for AdaMSS (Adaptive Matrix Decomposition with Subspace Selection), a parameter-efficient fine-tuning method that uses SVD to decompose weight matrices into low-rank subspaces. The new files include \glue\_adamss\_asa\_example.py\ and \glue\_adamss\_asa\_manual\_example.py\ for NLU tasks (GLUE benchmark) and \image\_classification\_adamss\_asa.py\ for vision tasks (Image Classification), demonstrating both callback-based and manual ASA (Adaptive Subspace Allocation) integration. The \README.md\ provides configuration details, usage instructions, and experimental results showing significant parameter reduction (\~0.07% of original trainable parameters) while maintaining competitive performance.

_examples/adamss\finetuning · high confidence

New Arrow multitask evaluation example for Phi-3 Mini

Added a new example script (\arrow\_phi3\_mini.py\) that demonstrates evaluating the Phi-3 Mini model on multiple-choice reasoning datasets (such as BoolQ, HellaSwag, and ARC) using three strategies: base, Arrow modular routing, and Arrow with GenKnowSub (general knowledge subtraction). The script supports 4-bit quantization and allows users to compare accuracy improvements from task-specific adapters and knowledge subtraction techniques.

_examples/arrow\multitask · high confidence

New BD-LoRA finetuning example with multi-GPU serving scripts

Added an example in \examples/bdlora\_finetuning\ that demonstrates Block-Diagonal LoRA (BD-LoRA) finetuning and inference. The \bdlora\_peft\_demo.ipynb\ notebook shows how to configure and train a model using \BdLoraConfig\ to optimize for tensor-parallel serving. The folder also includes \vllm\_server.bash\ to launch a vLLM server with block-diagonal sharded LoRA support and \chat.py\ to query the server, enabling users to experience the inference speed-ups from reduced communication overhead across multiple GPUs.

_examples/bdlora\finetuning · high confidence

New BOFT DreamBooth example with Intel XPU support

This location introduces a new example for fine-tuning Stable Diffusion models using BOFT (Butterfly Orthogonal Finetuning), providing a training script, inference notebook, and documentation. The implementation allows users to significantly reduce trainable parameters by integrating full-rank orthogonal matrices with a butterfly structure into attention blocks. A key behavioral addition is explicit support for Intel XPU hardware, enabling the training and inference workflows to run on XPU devices alongside standard CUDA setups.

_examples/boft\dreambooth · high confidence

New Cartridge self-study distillation example with XPU support

The \examples/cartridge\_self\_study\ directory now includes a complete workflow for training CARTRIDGE adapters via self-study context distillation. This example provides scripts to synthesize training data (using vLLM with prefix caching or standard Hugging Face generation) and train the adapter, along with convenience wrappers for the arXiv paper's LaTeX source. The training scripts have been updated to support Intel XPU devices alongside CUDA, MPS, and CPU, allowing users to run the distillation process on XPU hardware.

_examples/cartridge\_self\study · high confidence

New Context-aware Prompt Tuning (CPT) example

Added a new example demonstrating Context-aware Prompt Tuning (CPT), a method that combines In-Context Learning with adversarial-style optimization to refine context embeddings for few-shot learning. The entry includes a detailed README explaining the dataset preparation, template-based tokenization, and the use of projected gradient descent to update specific context tokens while preserving label integrity, along with a Jupyter notebook (\cpt\_train\_and\_inference.ipynb\) that provides a complete workflow for setup, data preparation, model training, and evaluation using the Hugging Face Trainer.

_examples/cpt\finetuning · high confidence

New DEFT DreamBooth fine-tuning example

Added a new example in \examples/deft\_dreambooth\ that demonstrates how to fine-tune a diffusion model using Decompositional Efficient Fine-Tuning (DEFT). The directory includes a training script (\train\_dreambooth.py\) with arguments to enable DEFT (e.g., \--use\_deft\, \--deft\_r\), a Jupyter notebook (\deft\_dreambooth\_inference.ipynb\) for loading the trained adapters, and documentation (\README.md\) explaining the method and setup. This allows users to personalize models while preserving base model editability by injecting a new low-rank sub-space.

(repo-wide) · high confidence

New DNA Language Model fine-tuning example

Added a new notebook example in the examples/dna\_language\_models directory that demonstrates how to use parameter-efficient fine-tuning (PEFT) techniques to adapt the SpeciesLM DNA language model for nucleotide benchmark tasks.

_examples/dna\_language\models · high confidence

New DoRA finetuning example with quantization and caching benchmarks

Added a new example in \examples/dora\_finetuning\ that demonstrates fine-tuning models using DoRA (Weight-Decomposed Low-Rank Adaptation). The entry includes a Python script (\dora\_finetuning.py\) and a Jupyter notebook (\QDoRA\_finetuning.ipynb\) that support 4-bit quantization (QDoRA) and allow users to toggle DoRA via the \--use\_dora\ flag. The example also introduces a benchmarking script (\dora-caching.py\) to measure the inference performance and memory overhead of DoRA with and without the new caching optimization.

_examples/dora\finetuning · high confidence

New DreamBooth example using Householder Reflection Adaptation (HRA)

A new example has been added to demonstrate fine-tuning the Stable Diffusion 2.1 model using the Householder Reflection Adaptation (HRA) method. This includes a training script and a Jupyter notebook for inference, allowing users to apply HRA adapters to the UNet and text encoder components to reduce parameter counts and computation costs while preserving pre-training knowledge.

_examples/hra\dreambooth · high confidence

New EVA finetuning example demonstrating Explained Variance Adaptation

The \examples/eva\_finetuning\ directory now contains a complete example for using Explained Variance Adaptation (EVA) with LoRA. This includes a README explaining the method, a standard finetuning script (\eva\_finetuning.py\), a multi-accelerator distributed training script (\eva\_finetuning\_multi\_accelerator.py\), and utility functions for tokenization and data collation. Users can now follow this example to initialize LoRA adapter weights using EVA's data-driven SVD approach, supporting both single-GPU and distributed setups.

_examples/eva\finetuning · high confidence

New FP4 QLoRA finetuning example for OPT models

Added a new example script (\finetune\_fp4\_opt\_bnb\_peft.py\) demonstrating how to fine-tune the \facebook/opt-350m\ model using 4-bit quantization via \bitsandbytes\ and Low-Rank Adapters (LoRA) from the \peft\ library. The script shows how to load the model with \BitsAndBytesConfig\ for NF4 quantization, apply LoRA adapters, and train on the \Abirate/english\_quotes\ dataset, supporting both CUDA and Intel XPU devices.

_examples/fp4\finetuning · high confidence

New GLoRA fine-tuning example added

Added a new example script and documentation in \examples/glora\_finetuning\ that demonstrates how to fine-tune causal language models using GLoRA adapters. The example uses the \yahma/alpaca-cleaned\ dataset and allows users to configure specific GLoRA parameterization modes (LoRA, vector, constant, or none) for different weight correction paths to balance expressiveness and parameter count.

_examples/glora\finetuning · high confidence

New KaSA fine-tuning example demonstrating spectral-domain adapter training

An example script and documentation have been added for KaSA (Knowledge-aware Singular-value Adaptation), a parameter-efficient fine-tuning method that operates in the spectral domain of base weights. The example shows how to configure \KasaConfig\ within \LoraConfig\ to apply SVD-based truncation and learnable singular-value updates, and demonstrates subclassing \SFTTrainer\ to inject KaSA-specific auxiliary regularizers (L2 penalty on singular values and orthogonal regularization on adapter factors) into the training loss.

_examples/kasa\finetuning · high confidence

New LoRA fine-tuning examples for image classification

Added two new Jupyter notebooks in the \examples/image\_classification\ directory demonstrating how to fine-tune image classification models using Low-Rank Adaptation (LoRA) from PEFT. The first notebook covers fine-tuning a Vision Transformer model, while the second showcases fine-tuning a PoolFormer model from the \timm\ library, both highlighting significant parameter reduction during training.

_examples/image\classification · high confidence

New LoRA model evaluation notebook using lm-eval-harness

Added a new Jupyter notebook in the examples/evaluation directory that demonstrates how to evaluate a fine-tuned LoRA model on the HellaSwag task using the lm-eval-harness toolkit. The notebook provides a step-by-step guide, including installing necessary dependencies and running the evaluation pipeline.

examples/evaluation · high confidence

New LoRA-FA fine-tuning example with XPU support

Added a new example script and documentation for fine-tuning large language models using the LoRA-FA (Low-rank Adaptation with Frozen A) method. This example demonstrates how to use the \create\_lorafa\_optimizer\ from PEFT to freeze the adapter A matrix, reducing memory consumption during training. The script supports CPU, single-accelerator, and multi-accelerator setups via Accelerate, and includes specific support for Intel XPU devices alongside CUDA.

_examples/lorafa\finetune · high confidence

New LoRA-GA finetuning example with gradient-aligned initialization

Added a new example script and documentation for LoRA-GA (Low-Rank Adaptation with Gradient Approximation) in the \examples/lora\_ga\_finetuning\ directory. This example demonstrates how to use gradient information during the initialization phase to align LoRA adapters with full fine-tuning directions, achieving faster convergence. The script includes a gradient estimation phase, configurable direction and scaling strategies, and specific instructions for saving adapters while handling base weight modifications.

_examples/lora\_ga\finetuning · high confidence

New LoRA-specific optimizers: LoRA-FA, LoRA+, and Riemannian-preconditioned

The \src/peft/optimizers\ module now exposes three new factory functions—\create\_lorafa\_optimizer\, \create\_loraplus\_optimizer\, and \create\_riemannian\_optimizer\—allowing users to fine-tune PEFT models with specialized optimization strategies. LoRA-FA applies a gradient projection based on the A-matrix inverse to improve B-matrix updates. LoRA+ configures separate learning rates for LoRA A and B adapters (and optionally embeddings) via a configurable ratio. The Riemannian-preconditioned optimizer wraps a base optimizer to apply an r×r preconditioner to LoRA A and B gradients on every step, with support for damping to stabilize near rank-deficient cases.

src/peft/optimizers · high confidence

New LoftQ finetuning examples and documentation

Added a new \examples/loftq\_finetuning\ directory containing a README, a Jupyter notebook (\LoftQ\_weight\_replacement.ipynb\), and Python scripts (\quantize\_save\_load.py\, \int8\_correction.py\, \train\_gsm8k\_llama.py\). These resources demonstrate how to apply LoftQ initialization to reduce quantization error in QLoRA models, including instructions for generating custom LoftQ initializations, loading pre-initialized models from the Hugging Face Hub, and fine-tuning on the GSM8K dataset.

_examples/loftq\finetuning · high confidence

New MetaMathQA benchmark results for Llama-3.2-3B

Added a comprehensive set of evaluation results for the MetaMathQA task using the Llama-3.2-3B model. The new files in \method\_comparison/MetaMathQA/results/\ document the performance of various parameter-efficient fine-tuning (PEFT) methods, including AdaLoRA, AdamSS, Adaption Prompt, BEFT, BOFT, C3A, and DEFT. Each result file contains detailed training metrics, configuration parameters, and accuracy logs, providing a baseline for comparing different adaptation strategies on this specific model and dataset.

_method\comparison/MetaMathQA/results · high confidence

New OFT Dreambooth inference notebook

Added a new Jupyter notebook example demonstrating how to perform inference using a Dreambooth adapter trained with the OFT (Orthogonal Fine-Tuning) method. The notebook provides a complete workflow for loading a base Stable Diffusion model, attaching the trained PEFT adapter to the UNet and text encoder, and generating images, including support for Apple Silicon GPU acceleration.

_examples/oft\dreambooth · high confidence

New OLoRA finetuning example with orthonormal initialization

Added a new example script and documentation for OLoRA (Orthonormal Low Rank Adaptation), a finetuning approach that uses QR decomposition to initialize LoRA weights for faster convergence and more stable training. The example demonstrates how to use the \init\_lora\_weights="olora"\ configuration option, supports 4-bit quantization via BitsAndBytes, and includes instructions for running with DDP using Accelerate.

_examples/olora\finetuning · high confidence

New PEFT LoRA example for semantic search and similarity

Added a new Python script and Jupyter notebook example in the feature extraction directory that demonstrates how to fine-tune embedding models using Parameter-Efficient Fine-Tuning (PEFT) with LoRA for semantic search and similarity tasks. The example provides a complete workflow for training, evaluation, and inference, allowing users to leverage PEFT techniques to adapt pre-trained models for specific retrieval use cases with reduced computational cost.

_examples/feature\extraction · high confidence

New PEFT LoRA examples for NER and LayoutLM token classification

Added two new Jupyter notebooks in the token classification examples directory: one demonstrating Named Entity Recognition (NER) on the CoNLL-2003 dataset using PEFT (formerly PEFT) LoRA, and another showing fine-tuning of the LayoutLM model on the FUNSD dataset for form understanding. These notebooks provide updated, practical guides for users looking to apply parameter-efficient fine-tuning techniques to token classification tasks.

_examples/token\classification · high confidence

New PEFT examples for LayerNorm tuning, additional tokens, and DeepSpeed offload

The causal language modeling examples directory now includes new notebooks and scripts demonstrating advanced PEFT techniques: LayerNorm tuning (peft\_ln\_tuning\_clm.ipynb), training with new tokens added to the embedding layers and tokenizer (peft\_lora\_clm\_with\_additional\_tokens.ipynb), and a LoRA example configured for DeepSpeed ZeRO-3 with CPU offload (peft\_lora\_clm\_accelerate\_ds\_zero3\_offload.py), accompanied by its specific Accelerate configuration file (accelerate\_ds\_zero3\_cpu\_offload\_config.yaml).

_examples/causal\_language\modeling · high confidence

New PEFT examples for conditional generation tasks

Added new example scripts and notebooks demonstrating how to use Parameter-Efficient Fine-Tuning (PEFT) methods on sequence-to-sequence models. The new files include \peft\_adalora\_seq2seq.py\ for AdaLoRA, \peft\_ia3\_seq2seq.ipynb\ for IA3, \multitask\_prompt\_tuning.ipynb\ for multitask prompt tuning, and updated or new examples for LoRA, prefix tuning, and prompt tuning. Additionally, a DeepSpeed ZeRO-3 CPU offload configuration file (\accelerate\_ds\_zero3\_cpu\_offload\_config.yaml\) and an Accelerate-based LoRA example with DeepSpeed offload (\peft\_lora\_seq2seq\_accelerate\_ds\_zero3\_offload.py\) and FSDP support (\peft\_lora\_seq2seq\_accelerate\_fsdp.py\) were added to facilitate distributed training setups.

_examples/conditional\generation · high confidence

New PEFT utility module for adapter management and configuration

The \src/peft/utils\ package has been introduced to centralize core utilities for the PEFT library. This module provides essential configuration mappings for target modules across various transformer architectures (e.g., Llama, Mistral, Gemma) for methods like LoRA, IA3, and BOFT. It also includes critical infrastructure for adapter hot-swapping, allowing users to switch between adapters without reloading the model, as well as utilities for LoftQ initialization, incremental PCA for EVA, and state-dict management for saving and loading adapters.

src/peft/utils · high confidence

New Poly Adapter example for sequence-to-sequence models

Added a new Jupyter notebook example (\peft\_poly\_seq2seq\_with\_generate.ipynb\) demonstrating how to use the Poly Adapter (PEFT) with sequence-to-sequence models like \google/flan-t5-xl\. The example shows how to configure PolyConfig parameters such as rank, number of tasks, skills, and splits, and how to integrate this with the Hugging Face \Seq2SeqTrainer\ for multi-task learning.

examples/poly · high confidence

New QALoRA GPTQ finetuning example

Added a new example script and documentation for fine-tuning large language models using Quantization-Aware LoRA (QALoRA) with GPTQ-quantized models. This feature allows users to efficiently fine-tune quantized models (such as Llama-2-7b-GPTQ) by leveraging input feature pooling and specialized grouping techniques, significantly reducing memory requirements compared to standard LoRA. The provided script supports automatic detection of existing GPTQ quantization, caching of quantized models to avoid redundant processing, and configuration of QALoRA-specific parameters like group size to balance memory usage and performance.

_examples/qalora\finetuning · high confidence

New RandLora finetuning example with quantization and sparsity support

Added a new example demonstrating how to fine-tune large language models using the RandLora parameter-efficient technique, which performs full-rank updates via random bases. The provided script and notebook allow users to swap standard LoRA for RandLora, with support for 4-bit quantization (via BitsAndBytes), sparse or very sparse random bases to reduce overfitting, and configurable target modules. The example includes a Colab notebook for QRandLora (quantized RandLora) on T4 GPUs and a standalone Python script that handles dataset loading, tokenization, and training with the PEFT library.

_examples/randlora\finetuning · high confidence

New RoAd 2D Rotary Adaptation finetuning example

Added a new example script and documentation for fine-tuning models using RoAd (2D Rotary Adaptation). The \road\_finetuning.py\ script demonstrates how to apply the \RoadConfig\ via PEFT to adapt LLMs with less than 0.1% trainable parameters, supporting features like 4-bit quantization, efficient batching, and configurable variants (road\_1, road\_2, road\_4).

_examples/road\finetuning · high confidence

New ShadowPEFT example for parameter-efficient fine-tuning

Added a new example in \examples/shadow\_finetuning\ demonstrating ShadowPEFT, a technique that augments a frozen base model with a trainable 'shadow' network. The provided script (\shadow\_finetuning.py\) and documentation show how to configure a shadow backbone (either mirrored from the base model or initialized from a separate pretrained model) and train only the shadow components. It includes logic to save the adapter, reload it, and extract a standalone shadow model for independent evaluation.

_examples/shadow\finetuning · high confidence

New Super-Tuning example for sparse fine-tuning

Added a new example script and documentation for Super-Tuning (also referred to as Supra), a sparse fine-tuning method that freezes base weights and trains only a small subset of individual scalar weight entries selected by magnitude. The example demonstrates how to use the \SupertuningConfig\ to apply this method, including support for a hybrid 'Supra' mode that combines sparse support with a LoRA-style low-rank adapter, and provides usage instructions for both Python API and command-line execution.

_examples/supertuning\finetuning · high confidence

New Supervised Fine-Tuning (SFT) example with PEFT and distributed training support

The \examples/sft\ directory now provides a complete, standalone example for performing Supervised Fine-Tuning using PEFT (LoRA/QLoRA). This includes a refactored \train.py\ script that leverages \SFTConfig\ and \SFTTrainer\ from TRL, supporting 4-bit and 8-bit quantization via BitsAndBytes, Flash Attention 2, and Unsloth optimizations. The example is designed for various distributed setups, with dedicated launch scripts and configuration files for single-GPU, multi-GPU DDP, DeepSpeed ZeRO-3, and FSDP, including specific configurations for QLoRA with FSDP and DeepSpeed. It also supports GPT-Q quantized models with FSDP and handles chat template formatting for ChatML and Zephyr styles.

examples/sft · high confidence

New UniLoRA finetuning example

Added a new example script and documentation for finetuning models using the UniLoRA adapter technique. The example demonstrates how to configure and train a model with shared trainable vector banks to reduce adapter parameters, including usage with the Hugging Face Transformers and TRL libraries.

_examples/unilora\finetuning · high confidence

New example for LoRA fine-tuning ESM2 with Transformer Engine

Adds a new example demonstrating how to fine-tune the Transformer Engine-accelerated ESM2 model for protein token classification using Low-Rank Adaptation (LoRA). The entry includes the training script, a Dockerfile based on the NVIDIA PyTorch container, and documentation covering setup via Docker or virtual environment, synthetic dataset generation, and usage of the Hugging Face Trainer.

_examples/lora\_finetuning\_transformer\engine · high confidence

New example for fine-tuning custom models with LoRA

Added a new notebook example in the \examples/multilayer\_perceptron\ directory that demonstrates how to apply Low-Rank Adaptation (LoRA) via PEFT to a custom multilayer perceptron model rather than a standard transformers model. The example includes a complete walkthrough for training this custom architecture on a classification task, highlighting that PEFT supports fine-tuning any model type with supported layers.

_examples/multilayer\perceptron · high confidence

New image generation benchmark for comparing PEFT methods

Added a new benchmark in \method\_comparison/image-gen\ to evaluate and compare Parameter-Efficient Fine-Tuning (PEFT) methods on a DreamBooth-style image generation task using the FLUX.2-klein-base-4B model. The suite includes scripts (\run.py\, \evaluate.py\) to train adapters, measure metrics like DINOv2 similarity and drift, and generate sample images, along with a Makefile to automate experiment sweeps and a configuration file (\default\_training\_params.json\) defining the dataset, hyperparameters, and evaluation settings.

_method\comparison/image-gen · high confidence

New int8 training examples for BLIP-2, Whisper, and OPT models

Added new examples in the \examples/int8\_training\ directory demonstrating how to fine-tune large models using 8-bit quantization with \bitsandbytes\ and parameter-efficient fine-tuning (PEFT). The new files include a Python script and configuration for fine-tuning BLIP-2 with LoRA, a Python script and shell script for fine-tuning Whisper-large-v2 with AdaLora, and Jupyter notebooks for fine-tuning FLAN-T5 and OPT models using LoRA.

_examples/int8\training · high confidence

New multi-adapter examples for LoRA merging and Diffusers integration

Added three new Jupyter notebooks in the multi-adapter examples directory to demonstrate advanced adapter usage. The Lora\_Merging notebook shows how to load multiple LoRA adapters, merge them using weighted combinations (including ties and density strategies), and perform inference. The PEFT\_Multi\_LoRA\_Inference notebook provides a guide for loading and switching between multiple adapters on a single model. The multi\_adapter\_weighted\_inference\_diffusers notebook demonstrates how to apply PEFT adapter merging methods to image generation models using Diffusers.

_examples/multi\_adapter\examples · high confidence

New semantic segmentation example using LoRA and PEFT

Added a new Jupyter notebook example in the \examples/semantic\_segmentation\ directory that demonstrates how to fine-tune a SegFormer model for semantic segmentation using Low-Rank Adaptation (LoRA) via the PEFT library. The example shows how to achieve this with only 14% of the original trainable parameters by adding low-rank update matrices to attention blocks, and includes steps for installing dependencies, authenticating with Hugging Face, loading the SceneParse150 dataset, and preparing data for training.

_examples/semantic\segmentation · high confidence

New sequence classification examples for C3A, FourierFT, IA3, VB-LoRA, VeRA, and torchao integration

The examples/sequence\_classification directory now includes new Jupyter notebooks demonstrating fine-tuning with C3A, FourierFT, IA3, VB-LoRA, and VeRA adapters, alongside updated examples for LoRA with torchao 8-bit quantization (both int8\_dynamic\_activation\_int8\_weight and int8\_weight\_only modes). These notebooks provide ready-to-run workflows for applying these specific parameter-efficient fine-tuning methods to sequence classification tasks.

_examples/sequence\classification · high confidence

New utility scripts for documentation, CI, and PEFT method management

This update introduces several new automation and utility scripts to the \scripts/\ directory. \check\_doc\_coverage.py\ provides a tool to verify that public API objects are documented by scanning markdown files for references. \ci\_clean\_cache.py\ allows for the automated cleanup of Hugging Face Hub cache files based on age. \convert-bone-to-miss.py\ and \evaluate-lora-conversion.py\ support the new MiSS adapter format and the evaluation of PEFT-to-LoRA conversions, respectively. Additionally, \generate\_method\_capabilities.py\ creates a machine-readable matrix of PEFT method features, \train\_memory.py\ estimates training memory usage, and \triage\_prs.py\ enforces approval workflows for new pull requests. The \stale.py\ script was also updated to use timezone-aware datetime handling.

scripts · high confidence

PEFT library initialization and core API structure

The \src/peft\ package is initialized with version 0.21.1.dev0, exposing the full public API including \AutoPeftModel\ classes for automatic task-based model loading, the \get\_peft\_model\ factory function, and a comprehensive registry of PEFT tuners (such as LoRA, IA3, and Prompt Tuning) and their configurations. The release introduces a new \functional.py\ module that provides a stable, public API for low-level adapter operations like \inject\_adapter\_in\_model\, \set\_adapter\, and \delete\_adapter\, intended for integration with non-PeftModel models. Additionally, \auto.py\ implements a secure import allowlist mechanism to prevent arbitrary code execution when loading adapter configurations from the Hub, and \config.py\ now stores the PEFT version (including commit hash for dev builds) in saved adapter configs to ensure forward compatibility.

src/peft · high confidence

PVeRA example now supports Monte Carlo confidence interval estimation

The PVeRA example in \examples/pvera\ has been updated to demonstrate how to generate confidence intervals for model predictions. By setting \sample\_at\_inference=True\ in the \PveraConfig\, users can enable Monte Carlo sampling during inference, allowing for the estimation of prediction uncertainty through multiple passes. The provided \confidence\_interval\_generation.py\ script illustrates this workflow, including training a model with adapters, saving the configuration, reloading it with sampling enabled, and computing confidence intervals on the output predictions.

examples/pvera · high confidence

Prompt tuning now supports initialization from sampled vocabulary tokens

The prompt tuning implementation has been updated to support initializing virtual token embeddings from randomly sampled tokens in the model's vocabulary. Users can now set \prompt\_tuning\_init\ to \SAMPLE\_VOCAB\ in their configuration, which replaces the previous limitation of only supporting \RANDOM\ (continuous soft tokens) or \TEXT\ (text-based initialization). This change is implemented in the \PromptEmbedding\ model class within the \prompt\_tuning\ tuner, alongside the addition of the \PromptTuningInit\ enum and corresponding configuration fields.

_src/peft/tuners/prompt\tuning · high confidence

Support for LoHa, LoKr, and Intel INC FP8 quantization in Stable Diffusion examples

The Stable Diffusion examples now support LyCORIS LoHa and LoKr adapter formats in addition to standard LoRA, allowing users to train and convert these specific adapter types via the dreambooth training script and the new adapter conversion utility. Additionally, a new example demonstrates loading LoRA adapters into FP8-quantized FLUX models using Intel Neural Compressor on HPU devices.

_examples/stable\diffusion · high confidence

Architecture

Introduction of PEFT Tuner Subpackage and Core Infrastructure

The \src/peft/tuners\ directory has been restructured into a dedicated subpackage, introducing a centralized \\_\init\\_.py\ that exposes a comprehensive registry of parameter-efficient fine-tuning methods (including LoRA, LyCORIS variants like LoHa/LoKr, BOFT, IA3, and many others) along with their configuration classes. This change establishes the foundational infrastructure for the library by adding core utility modules such as \tuners\_utils.py\ (providing base tuner logic, adapter injection, and offload handling) and \\_buffer\_dict.py\ (a custom ordered dictionary for managing model buffers), effectively organizing the previously flat tuner implementations into a modular, extensible architecture.

src/peft/tuners · high confidence

Behavioural changes

Benchmark results for new PEFT methods on FLUX.2-klein

Added image generation benchmark results for multiple PEFT methods (including AdaLoRA, AdaMSS, BEFT, BOFT, C3A, DEFT, DeLoRA, and DoRA) trained on the FLUX.2-klein-base-4B model using the internal cat-image-dataset. These JSON files document training configurations, hyperparameters, and performance metrics (such as DINO similarity and training time) to support method comparison.

_method\comparison/image-gen/results · high confidence

Test coverage

Added distributed training tests for LoRA and FSDP; Added regression test suite for PEFT model outputs and state dict serialization; Introduce official PEFT test suite documentation and infrastructure.

Dependencies

Standardize example dependencies and configure project tooling

This change introduces explicit \requirements.txt\ files for all example directories (such as \sft\, \lora\_dreambooth\, \boft\_controlnet\, and \method\_comparison\), ensuring each example installs the precise library versions needed to run correctly. It also adds a \pyproject.toml\ to the project root to centralize configuration for \ruff\ (set to version 0.16.4) and \pytest\, while establishing a root \requirements.txt\ for the core library's development dependencies.

(dependencies) · high confidence

Housekeeping

Added placeholder directories for MetaMathQA evaluation results

Empty .gitkeep files were added to the method\_comparison/MetaMathQA/cancelled\_results and method\_comparison/MetaMathQA/temporary\_results directories. This ensures these folders are tracked by version control, likely to reserve space for storing evaluation outputs or logs in future iterations of the method comparison suite.

_method\_comparison/MetaMathQA/cancelled\_results, method\_comparison/MetaMathQA/temporary\results · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 65.

Lenses

  • Code Health 78
  • Architecture 90
  • Maturity 63
  • Readiness 60
  • Security 74

Changes since last survey

  • 300 commits — 173 feature/other, 127 fixes

By area

  • src/peft — 164 commits
  • .github/workflows — 22 commits
  • docs/source — 22 commits
  • method_comparison/MetaMathQA — 21 commits
  • (root) — 15 commits
  • method_comparison/image-gen — 14 commits
  • docker/peft-gpu — 8 commits
  • tests/test_common_gpu.py — 4 commits
  • .ai/AGENTS.md — 3 commits
  • tests/test_gpu_examples.py — 3 commits
  • method_comparison/peft-shop — 2 commits
  • tests/test_decoder_models.py — 2 commits
  • tests/test_stablediffusion.py — 2 commits
  • tests/testing_common.py — 2 commits
  • (repo) — 1 commit
  • .ai/skills — 1 commit
  • .github/dependabot.yml — 1 commit
  • examples/beft_finetuning — 1 commit
  • examples/hra_dreambooth — 1 commit
  • examples/int8_training — 1 commit

Notable commits

  • fix: CI FIX Some tests require torchvision (#3135)
  • fix: CI Fix GPTQmodel install in Docker build (#3188)
  • fix: CI Fix error with Windows loading stable diffusion models (#3399)
  • fix: CI Fix ruff version to 0.16.4 (#3613)
  • fix: DOC FIX Incorrect class references in docstrings (#3286)
  • fix: DOC Fix DeepSpeed guide's example script path (#3494)
  • fix: DOC Fix GraloraConfig docstring (#3568)
  • fix: DOC Fix LoRA-GA 'Usage Tips' subsection (#3331)
  • fix: DOC Fix broken docstring examples in peft_model and tuner models (#3657)
  • fix: DOC Fix code blocks in README.md examples (#3392)
  • fix: DOC Fix incorrect LoRA references in hira and ln_tuning model docstrings (#3291)
  • fix: DOC Fix incorrect imports in helpers.py docstring examples (#3249)
  • fix: DOC Fix wrong function name in docstring (#3281)
  • fix: DOC Fix wrong paper links for LoHa and LoKr (#3547)
  • fix: FIX Accept layers_to_transform=0 together with layers_pattern (#3425)
  • fix: FIX AdaMSS, PSOFT: Allow seeding the RNG (#3310)
  • fix: FIX AdaMSS, PSOFT: fork the RNG of the actual device (#3598)
  • fix: FIX AutoPeftModel forwards revision and token (#3442)
  • fix: FIX Auxiliary modules requiring x argument (#3199)
  • fix: FIX Avoid bare Exception in peft_model.py (#3462)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/peft was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 50a277e7c87db460ef7444055788f9da29f2da71 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.