hiyouga/LlamaFactory
53.0
Adequate · 2 August 2026
44.6k
lines of production code
Python
primary language
3
measurements over time
What this system is
LLaMA-Factory is a comprehensive framework for training and serving large language models, supporting pretraining, supervised fine-tuning, and reinforcement learning methods like DPO, KTO, and PPO. It provides a modular architecture for distributed training (FSDP2, DeepSpeed, Megatron Bridge) and flexible inference via Hugging Face, vLLM, and SGLang backends. The system includes a unified CLI, an OpenAI-compatible API server, and a Gradio-based web UI for model management and evaluation.
How it got here
2023–2024 — Modular training architecture and API server
40 changes.
The project underwent a comprehensive architectural overhaul, replacing legacy scripts with a unified CLI and modular training backends for SFT, DPO, KTO, and pretraining. This period also introduced an OpenAI-compatible API server, a Gradio-based Web UI, and extensive unit and end-to-end tests to support the new modular design.
2025 — LlamaFactory v1 architecture and plugin system
24 changes.
This period focused on building the new v1 training and inference engine, introducing a modular plugin system for models, kernels, and data processing. The work established the core infrastructure for distributed training, optimized hardware-specific kernels, and added comprehensive test coverage for the new architecture.
2026 — v1 training infrastructure and parallelism
11 changes.
This period focused on building the core v1 training pipeline, introducing new utilities for batching, checkpointing, and inference. Significant work was done to support advanced distributed training via FSDP2, Ulysses sequence parallelism, and the HyperParallel backend. The changes also included adding the Muon optimizer and structured callback systems to enhance training extensibility and logging.
Features
Add HyperParallel distributed training support
Users can now train models using the HyperParallel backend, which provides HSDP, Context Parallel, and activation optimization features. The new \HyperParallelTrainer\ replaces the standard Accelerate FSDP2 integration, handling model preparation, gradient clipping, and data sampling specifically for HyperParallel's distributed execution.
_src/llamafactory/train/hyper\parallel · high confidence
Add LLaMA-PRO expansion example and configuration
A new example for LLaMA-PRO is added to the extras directory, including an expand.sh script for model expansion and a YAML configuration file for freeze-based supervised fine-tuning (SFT) using the LLaMA-PRO method. This provides users with a ready-to-use template for expanding LLaMA-3 models with frozen layers.
_examples/extras/llama\pro · high confidence
Add MCA training workflow and trainer
The \src/llamafactory/train/mca\ directory has been added, introducing a new training workflow for MCA (Megatron Core Adapter) integration. This includes a custom trainer (\trainer.py\) that overrides padding logic to handle 3D position IDs, and a workflow module (\workflow.py\) that orchestrates pre-training, supervised fine-tuning, and DPO training using the MCA adapter. The addition enables users to train models using the MCA framework, with specific handling for Qwen-VL series models and other supported architectures.
src/llamafactory/train/mca · high confidence
Add Muon optimizer and Triton-based chunked attention kernels
Users can now train models using the Muon optimizer, which applies orthogonalization to parameter updates, and benefit from new Triton-optimized chunked attention mechanisms that support variable-length sequences and gated delta rules for improved performance.
_src/llamafactory/third\party · high confidence
Add Muon optimizer plugin for training
Users can now select the Muon optimizer for training. This optimizer applies a Newton-Schulz orthogonalization step to 2D hidden weight matrices while routing 1D biases, embeddings, head layers, and LoRA factors to AdamW. The implementation is DTensor-aware, ensuring correctness under FSDP2 and sequence parallelism.
_src/llamafactory/v1/plugins/trainer\plugins/optimizers · high confidence
Add Pissa LoRA fine-tuning example for LLaMA-3
A new example for fine-tuning Meta's LLaMA-3-8B-Instruct model using the Pissa (Principal Singular Value Iterative SVD Adaptation) method has been added to the extras directory. The entry includes a shell script (init.sh) to initialize the model and a YAML configuration file (llama3\_lora\_sft.yaml) that sets up a supervised fine-tuning (SFT) workflow with LoRA, enabling users to train the model with Pissa-specific parameters such as pissa\_init, pissa\_iter, and pissa\_convert.
examples/extras/pissa · high confidence
Add Qwen3 and Qwen3-VL LoRA training examples
Added configuration files and a shell script for fine-tuning Qwen3 and Qwen3-VL models using LoRA. The new examples cover various training stages including supervised fine-tuning (SFT), direct preference optimization (DPO), knowledge-based training (KTO), pretraining, and reward modeling. The set includes standard single-GPU configurations, DeepSpeed ZeRO-3 setup, and a Ray-based multi-GPU example.
_examples/train\lora · high confidence
Add Ulysses sequence parallel support for FSDP2
Introduced new parallelization plugins for sequence parallelism using the Ulysses algorithm, enabling support for FSDP2. The change adds \seq\_comm.py\ (containing \SeqAllToAll4D\ for all-to-all tensor communication), \ulysses.py\ (implementing \UlyssesAttention\ and group management), and \sequence\_parallel.py\ (registering the \SequenceParallelModelPlugin\ and \SequenceParallelLossPlugin\ to apply sequence parallelism to the model and loss calculations).
_src/llamafactory/v1/plugins/model\plugins/parallelization · high confidence
Add automated evaluation framework for model benchmarking
The LlamaFactory project now includes a new evaluation module under src/llamafactory/eval, introducing an automated benchmarking capability. This addition provides an Evaluator class that loads models and tokenizers, processes evaluation datasets (supporting English and Chinese templates), and computes accuracy scores across subjects and categories. Users can now run standardized tests to assess model performance on multiple-choice tasks, with results saved as JSON and log files.
src/llamafactory/eval · high confidence
Add multilingual and multimodal dataset examples and documentation
Added Chinese README documentation and demo datasets (Alpaca and C4 formats) to the data directory, providing users with concrete examples of how to structure instruction-tuning and pre-training data in both English and Chinese. The update also introduces a JSON-based configuration schema that supports multimodal columns (images, videos, audios) and custom header mappings, enabling users to easily integrate custom datasets by defining their column mappings in \dataset\_info.json\.
data · high confidence
Added FSDP+QLoRA example for LLaMA-3
Users can now train a LLaMA-3 model using 4-bit quantization with FSDP and LoRA, as demonstrated by the new configuration and launch script in the examples directory.
_examples/extras/fsdp\qlora · high confidence
Added KTO (Kahneman-Tversky Optimization) training support
The KTO training workflow and trainer have been added to LlamaFactory, enabling users to perform KTO-based fine-tuning. The new \src/llamafactory/train/kto\ module provides a \run\_kto\ entry point and a \CustomKTOTrainer\ that integrates with the existing training pipeline, supporting features like reference model handling, custom optimizers/schedulers, and evaluation metrics specific to KTO.
src/llamafactory/train/kto · high confidence
Added model conversion and test utilities for LLaMA 4 and Qwen3 architectures
New scripts were added to the \scripts/convert\_ckpt\ directory to support converting model checkpoints and generating tiny test models for specific architectures. \llamafy\_baichuan2.py\ and \llamafy\_qwen.py\ provide conversion logic to transform Baichuan2 and Qwen model weights into the LLaMA2-compatible format. Additionally, \tiny\_llama4.py\ and \tiny\_qwen3.py\ were introduced to generate minimal configuration and model instances for LLaMA 4 and Qwen3, facilitating testing and development workflows for these newer model types.
_scripts/convert\ckpt · high confidence
Added utility scripts for model statistics and profiling
New scripts have been added to the \scripts/stat\_utils\ directory to help users analyze and profile their models. These include \cal\_flops.py\ for calculating FLOPs and MACs, \cal\_lr.py\ for estimating optimal learning rates, \cal\_mfu.py\ for measuring Model FLOPs Utilization (MFU), \cal\_ppl.py\ for computing perplexity on datasets, and \length\_cdf.py\ for analyzing input length distributions. These tools provide command-line interfaces to gather key metrics about model performance and dataset characteristics.
_scripts/stat\utils · high confidence
Added v1 plugin and sampler infrastructure
The v1 plugins directory now includes empty \_\init\\_.py files for the main plugins module and the sampler\_plugins subpackage, alongside a new vllm.py file within sampler\_plugins. This establishes the structural foundation for v1-specific plugin and sampler implementations, preparing the codebase for future v1 feature development.
_src/llamafactory/v1/plugins, src/llamafactory/v1/plugins/sampler\plugins · low confidence
Centralized configuration system for training, data, and model parameters
The \src/llamafactory/hparams\ directory has been refactored to use a new modular argument structure. Users can now configure data loading, evaluation, fine-tuning methods (including LoRA, PiSSA, and OFT), model settings, and generation parameters through dedicated dataclasses. The parser now supports loading configurations from YAML and JSON files, and integrates with OmegaConf for config merging. This change also introduces support for distributed training via Ray, FP8 mixed precision training, and profiling capabilities.
src/llamafactory/hparams · high confidence
Introduce OpenAI-compatible API server
Adds a new FastAPI-based HTTP server at src/llamafactory/api that exposes an OpenAI-style interface for chat completions, streaming, and score evaluations. The server handles multimodal inputs (images, video, audio) and tool calls, enforces API key authentication, and validates requests against local file and SSRF risks.
src/llamafactory/api · high confidence
Introduce PPO training workflow and trainer
Adds a new PPO (Proximal Policy Optimization) training workflow for reinforcement learning from human feedback (RLHF). The change introduces a custom PPO trainer and supporting utilities in src/llamafactory/train/ppo, including workflow orchestration, model/reward model setup, and DeepSpeed Zero-3 compatibility. This enables users to perform PPO-based fine-tuning using the TRL library (versions 0.8.6 to 0.9.6).
src/llamafactory/train/dpo, src/llamafactory/train/ppo · high confidence
Introduce dedicated Reward Model (RM) training workflow
A new \src/llamafactory/train/rm\ module has been added, providing a complete workflow for training reward models. This includes a \PairwiseTrainer\ that computes pairwise loss, a \ComputeAccuracy\ metric for evaluation, and a \run\_rm\ entry point that handles data loading, model initialization, and training/evaluation/prediction steps. This change introduces a distinct training path for reward models, separate from the existing supervised or RLHF workflows.
src/llamafactory/train/rm · high confidence
Introduce modular training backends: Megatron Bridge, HyperParallel, and FP8 support
The training subsystem in src/llamafactory/train has been restructured to support multiple training backends. A new Megatron Bridge integration is added, providing a bridge to run pretraining and SFT using Megatron-LM's training loop, including configuration building, dataset export, and workflow management. Additionally, support for HyperParallel FSDP2 and MCA (mcore\_adapter) backends is introduced, routing training through specialized run functions. The codebase also adds FP8 training support, including utilities for configuring FP8 with both TorchAO and Transformer Engine backends, and a callback to fix value-head checkpoints. The training entry point (tuner.py) and utility modules (trainer\_utils.py, callbacks.py, test\_utils.py) are updated to orchestrate these new backends and callbacks.
src/llamafactory/train · high confidence
Introduce new accelerator and distributed computing infrastructure
A new accelerator module has been added to the v1 codebase, providing a unified interface for model parallelism (mp\_replicate, mp\_shard) and data parallelism (dp, cp). This includes helper functions for device detection, process group management, and collective communication, alongside a DistributedInterface class that initializes device meshes for training.
src/llamafactory/v1/accelerator · high confidence
Introduce new distributed training backends: FSDP2 and DeepSpeed
The training system now supports two new distributed training backends: FSDP2 and DeepSpeed. This change introduces a new plugin architecture for distributed training, allowing users to select between FSDP2 and DeepSpeed for their training jobs. The FSDP2 backend uses PyTorch's fully sharded data parallelism with checkpointing and model saving capabilities, while the DeepSpeed backend leverages Accelerate's built-in DeepSpeed integration for memory-efficient training. Both backends support model sharding, checkpointing, and saving, enabling users to leverage advanced distributed training techniques for larger models or faster training.
_src/llamafactory/v1/plugins/trainer\plugins/distributed · high confidence
Introduce pluggable kernel optimization framework with Liger and device-specific kernels
Users can now apply model kernel optimizations via a new plugin-based system. The change adds a base plugin structure and an interface that supports automatic kernel selection for NPU devices (including fused MoE, RMS norm, RoPE, and SwiGLU operations) as well as manual selection. It also introduces the Liger kernel plugin, which applies optimized operations to supported Qwen3 and Qwen3.5 model types, with configurable toggles for specific fused operations like cross-entropy and RMS norm.
_src/llamafactory/v1/plugins/model\plugins/kernels · high confidence
Introduce pretraining workflow and custom trainer for PyTorch
The \src/llamafactory/train/pt\ directory now contains the implementation for the pretraining (pt) stage. This includes a new \CustomTrainer\ that supports FP8 training, BAdam optimizer, and disabled shuffling, alongside a \workflow.py\ that orchestrates the pretraining loop, evaluation, and loss plotting. This change adds the foundational training logic for the pretraining stage, distinct from other fine-tuning stages.
src/llamafactory/train/pt · high confidence
Introduce structured training callbacks for logging and extensibility
A new callback system has been added to the v1 training pipeline, introducing a \TrainerState\ dataclass and a \TrainerCallback\ base class that standardizes training hooks (e.g., \on\_log\, \on\_step\_begin\). The \LoggingCallback\ implementation now persists training metrics to a \trainer\_log.jsonl\ file and formats learning rate logs with improved precision, ensuring that training history survives crashes and providing more readable console output.
src/llamafactory/v1/utils/callbacks · medium confidence
Introduce unified CLI entrypoint and v1 launcher routing
The LLaMA-Factory package now exposes a structured command-line interface via \llamafactory-cli\ (and the \lmf\ alias), providing commands for API, chat, training, and web UI. The \\_\init\\_.py\ file was added to expose the package version, while \cli.py\ and \launcher.py\ were introduced to handle command dispatching. Notably, the launcher now supports a \USE\_V1\ environment variable to route execution to the new v1 backend, enabling elastic and fault-tolerant distributed training via \torchrun\ with configurable node counts and restart policies.
src/llamafactory · high confidence
Introduce v1 CLI launcher and backend routing
The LlamaFactory v1 module now includes a new CLI launcher that routes commands (such as sft, dpo, rm, chat, and merge) to their respective handlers. This change introduces a new entry point for the v1 architecture, enabling distributed training via torchrun and supporting specific training stages like DPO and reward modeling within the v1 framework.
src/llamafactory/v1 · high confidence
Introduce v1 configuration and CLI sampler for training and inference
A new v1 configuration system has been added, introducing dedicated data, model, training, and sampling argument classes that are parsed from command-line or YAML/JSON config files. This includes support for distributed training parameters (e.g., DeepSpeed, FSDP2), batching strategies, and seed management. Additionally, a CLI-based synchronous sampler is added to enable interactive chat and batch inference using the new configuration structure.
src/llamafactory/v1/config · high confidence
Introduce v1 core training and inference engine
The v1 training and inference pipeline is now available. This adds a new core module that provides a \BaseTrainer\ for training, a \DataEngine\ for dataset loading and indexing, and a \ModelEngine\ for model and processor initialization. The update also introduces a \Renderer\ for converting messages into model inputs, along with a \BaseSampler\ for generation. These components form the foundation for the new v1 training workflow, supporting features like DeepSpeed, FSDP2, and multimodal data.
src/llamafactory/v1/core · high confidence
Introduce v1 trainer plugins for batching and learning rate scheduling
The v1 trainer plugin system is introduced, adding new modules for batching strategies and learning rate scheduling. The \batching.py\ file implements a \PaddingFreeBatcher\ that packs samples into contiguous sequences without padding, alongside helper functions for dynamic micro-batch sizing. The \lr\_scheduler.py\ file adds a placeholder \LRSchedulerPlugin\ to manage learning rate adjustments. These changes provide a new, modular way to configure training behavior in the v1 API.
_src/llamafactory/v1/plugins/trainer\plugins · high confidence
NPU-accelerated kernels for MoE, SwiGLU, RMSNorm, and RoPE
Added NPU-optimized implementations for several core operations: fused MoE (MLP) kernels for both CUDA (Triton-based grouped GEMM) and NPU (using torch\_npu.npu\_grouped\_matmul); NPU fused SwiGLU; NPU fused RMSNorm (including gated and residual variants); and NPU fused RoPE (including partial and multimodal sections). These changes replace default PyTorch/transformers implementations with hardware-specific kernels to improve performance on NPU devices.
_src/llamafactory/v1/plugins/model\plugins/kernels/ops · high confidence
New Dockerfiles for CUDA-based LLaMA Factory and Megatron Bridge
Added new Docker configuration files (Dockerfile, Dockerfile.base, Dockerfile.mbridge, Dockerfile.megatron, docker-compose.yml, and README.md) to support running LLaMA Factory on NVIDIA GPUs. The standard Dockerfile provides a base environment with PyTorch 2.6.0, CUDA 12.4, and optional Flash-Attention support. A separate Dockerfile.megatron provides a runtime for LLaMA-Factory combined with the NVIDIA Megatron Bridge, including TransformerEngine and Megatron Core. The setup includes pre-installed dependencies like flash-attn and flashinfer, and exposes ports 7860 and 8000.
docker/docker-cuda · high confidence
New Web UI and Chat interface for LLaMA Factory
The LLaMA Factory project now includes a complete, standalone Web UI and a lightweight chat interface, both built on Gradio. Users can now launch the full-featured Web UI via \llamafactory-cli webui\ or the minimal chat demo via \llamafactory-cli web-demo\. The Web UI provides a graphical interface for training, evaluation, inference, and export, while the chat demo offers a quick way to test models. This addition makes the framework significantly more accessible for users who prefer a graphical interface over command-line tools.
src/llamafactory/webui · high confidence
New \`llamafactory/extras\` module for shared utilities and configuration
The \llamafactory/extras\ package has been introduced to centralize shared utilities, constants, and configuration logic. This includes a new \constants.py\ defining model templates, supported architectures, and training stages; an \env.py\ for environment and dependency version checking; a \logging.py\ module providing thread-safe, rank-0-aware logging; a \misc.py\ with helper functions for device detection and parameter counting; and a \packages.py\ for optional dependency availability checks. These changes restructure the codebase to improve maintainability and provide a single source of truth for model and training configurations.
src/llamafactory/extras · high confidence
New data loading and conversion plugins for v1
The \src/llamafactory/v1/plugins/data\_plugins\ directory now includes new \converter.py\ and \loader.py\ modules. The converter introduces support for parsing Alpaca and ShareGPT formats, handling multimodal content (images, videos, audio) via inline placeholders. The loader adds a plugin-based system for loading local datasets (arrow, csv, json, parquet, text) and provides utilities to adjust dataset indices by size and weight.
_src/llamafactory/v1/plugins/data\plugins · high confidence
New modular SFT training pipeline with advanced loss functions and metrics
The SFT training workflow has been refactored into a new modular structure under \src/llamafactory/train/sft\, introducing a dedicated \CustomSeq2SeqTrainer\ and workflow entry point. This update enables support for advanced loss functions including DFT, EAFT, and ASFT, as well as FP8 training and BAdam optimizer integration. Users can now configure custom loss functions via \use\_dft\_loss\, \use\_eaft\_loss\, and \use\_asft\_loss\ arguments, and benefit from improved evaluation metrics such as accuracy and text similarity scores (BLEU, ROUGE) during training.
src/llamafactory/train/sft · high confidence
New utility scripts for model conversion, fine-tuning, and evaluation
Added several new scripts in the \scripts/\ directory to support advanced model workflows. \bench\_qwen.py\ provides a benchmarking script for Qwen2-VL models. \dcp2hf.py\ and \hf2dcp.py\ enable conversion between DCP and HuggingFace model formats. \llama\_pro.py\ implements block expansion for LLaMA, Mistral, Qwen2, and Yi models. \loftq\_init.py\ and \pissa\_init.py\ initialize LoRA weights using LoftQ and PiSSA techniques, respectively. \megatron\_merge.py\ handles merging Megatron checkpoints to HuggingFace format. \qwen\_omni\_merge.py\ merges the 'Talker' and 'Thinker' parts of Qwen-omni models. \eval\_bleu\_rouge.py\ calculates BLEU and ROUGE metrics for generated predictions. \vllm\_infer.py\ provides a batched inference script using vLLM, supporting image, video, and audio inputs.
scripts · high confidence
New v1 training core utilities for batching, checkpointing, and inference
Added new utility modules in src/llamafactory/v1/core/utils to support the v1 training pipeline. batching.py introduces a stateful batch generator that handles micro-batching, dynamic batching, and multimodal feature alignment. checkpoint.py provides helpers for saving, loading, and rotating training checkpoints, including metadata and RNG state. collation.py implements padding, truncation, and multimodal feature alignment for batch collation. inference\_engine.py defines a base class and HuggingFace implementation for asynchronous text generation. These changes enable improved training performance, checkpoint management, and inference capabilities in the v1 architecture.
src/llamafactory/v1/core/utils · high confidence
New v1 utility modules for precision, reproducibility, and plugin routing
The v1 training pipeline now includes a suite of new utility modules in src/llamafactory/v1/utils. A new dtype module provides a DtypeInterface that checks device availability for fp16, fp32, and bfloat16, and offers a context manager to set the default torch dtype. A helper module adds enable\_full\_determinism to fix gradient checkpointing and ensure reproducible training by seeding Python, NumPy, and PyTorch. A plugin module introduces a lightweight routing and parameter-parsing system for extensible components. Additional utilities include a StatefulBuffer for managing model inputs, a logging module with rank-0 aware methods, and type definitions for samples and batches. These changes support more robust, deterministic, and extensible training workflows.
src/llamafactory/v1/utils · high confidence
Repository initialization and documentation updates
The repository was initialized with essential configuration files including \.dockerignore\, \.env.local\, \.gitignore\, \.pre-commit-config.yaml\, \CITATION.cff\, \CLAUDE.md\, \MANIFEST.in\, and \Makefile\. Additionally, the \README.md\ and \README\_zh.md\ files were updated to reflect the project's features, supported models, and usage instructions.
(repo-wide) · high confidence
Unified chat interface with pluggable inference backends
The chat module has been refactored to support multiple inference backends (Hugging Face, vLLM, and SGLang) through a common \BaseEngine\ interface. Users can now select their preferred backend via the \--infer\_backend\ argument, enabling flexible model serving and generation across different libraries. The \ChatModel\ class now dynamically instantiates the appropriate engine (e.g., \HuggingfaceEngine\, \VllmEngine\, or \SGLangEngine\) based on the configuration, providing a consistent API for chat, streaming, and scoring regardless of the underlying engine.
src/llamafactory/chat · high confidence
Removals
Removal of HH-RLHF dataset loader
The dataset loader for the Human preference data about helpfulness and harmlessness (hh\_rlhf\_en) has been removed from the codebase. This change eliminates the previous implementation that parsed JSONL files from the Anthropic hh-rlhf dataset, specifically handling multi-turn chat formats for helpfulness and harmlessness evaluation.
_data/hh\_rlhf\en · high confidence
Removed UltraChat dataset loader
The UltraChat dataset loader (ultra\_chat.py) has been removed from the data directory. This means the UltraChat multi-turn dialogue dataset is no longer available for use in the system.
_data/ultra\chat · high confidence
Removed example dataset builder
The ExampleDataset builder, which previously loaded training data from an examples.json file to provide instruction, input, output, and history fields, has been removed from the codebase.
_data/example\dataset · high confidence
Architecture
Refactor model utilities into a dedicated module
The \src/llamafactory/model/model\_utils\ directory has been reorganized into a dedicated \model\_utils\ package, splitting the previous monolithic \modeling.py\ into specialized modules for attention, checkpointing, embedding, KV cache, Liger kernel, long-LoRA, misc utilities, MoD, MoE, packing, and quantization. This refactoring improves code modularity and maintainability without changing user-facing behavior.
_src/llamafactory/model/model\utils · high confidence
Refactored model loading and adapter setup into a modular structure
The model loading and adapter configuration logic has been reorganized into dedicated modules: \loader.py\ now handles the loading of configs, models, and tokenizers; \adapter.py\ manages full, freeze, and LoRA/OFT tuning setups; and \patcher.py\ centralizes model patching for attention, KV cache, and quantization. This refactoring improves code maintainability and clarity for users configuring fine-tuning methods.
src/llamafactory/model · high confidence
Behavioural changes
4 commits (1 fix) modifying data/mllm\_demo\_data
A change to existing behaviour in data/mllm\_demo\_data — 4 commits (1 fix), 11 files.
_data/mllm\_demo\data · medium confidence · unverified
New v1 model loading and PEFT plugins
The v1 model loading and parameter handling has been restructured into a new plugin system under \src/llamafactory/v1/plugins/model\_plugins\. This introduces dedicated plugins for DeepSpeed ZeRO-3 low-memory weight loading, LoRA/Freeze PEFT support with merge workflows, and quantization via BitsAndBytes (4/8-bit). Users can now configure model initialization strategies (meta, rank0, default), apply LoRA or freeze adapters, and load quantized models with specific compute dtypes and double-quantization settings.
_src/llamafactory/v1/plugins/model\plugins · medium confidence
Refactored data pipeline with modular components
The data loading and preprocessing logic has been refactored into a modular architecture. The previous monolithic structure has been split into distinct files for collation, conversion, formatting, and template handling. This change improves code maintainability and allows for more flexible dataset processing, such as supporting cloud storage paths and various data formats (Alpaca, ShareGPT, OpenAI) through dedicated converter classes.
src/llamafactory/data · high confidence
Refactored data pipeline with new processor classes
The data processing logic has been refactored into a modular structure under the \src/llamafactory/data/processor\ directory. This change introduces a new \DatasetProcessor\ base class and specific implementations for different training modes: \SupervisedDatasetProcessor\ and \PackedSupervisedDatasetProcessor\ for supervised fine-tuning, \PairwiseDatasetProcessor\ for preference optimization, \FeedbackDatasetProcessor\ for KTO/feedback tasks, \PretrainDatasetProcessor\ for pretraining, and \UnsupervisedDatasetProcessor\ for generation. The refactoring also adds utility functions like \greedy\_knapsack\ and \infer\_seqlen\ to handle sequence length calculations and packing logic.
src/llamafactory/data/processor · high confidence
Refactored entry points and removed legacy scripts
The repository has been refactored to use new entry points: \src/api.py\ for the API server, \src/train.py\ for training, and \src/webui.py\ for the web interface. The previous standalone scripts \cli\_demo.py\, \export\_model.py\, \train\_ppo.py\, \train\_rm.py\, \train\_sft.py\, and \web\_demo.py\ have been removed.
src · high confidence
Removal of legacy LLaMA-specific utility modules
The \src/utils\ directory has been restructured by deleting several files that contained LLaMA-specific implementations, including \common.py\ (model loading and adapter initialization), \config.py\ (data and finetuning arguments), \data\_collator.py\ (batching logic), \other.py\ (logging and helper functions), \pairwise.py\ (pairwise training), \peft\_trainer.py\ (PEFT trainer), \ppo.py\ (PPO training), and \seq2seq.py\ (generation metrics). This removes the previous LLaMA-centric codebase in favor of a new architecture.
src/utils · high confidence
Web UI components reorganized into a modular structure
The LlamaFactory web UI has been refactored by splitting the interface into distinct, reusable components. The \\_\init\\_.py\ file now imports and exports functions from separate modules for each UI section: \chatbot\, \data\, \eval\, \export\, \footer\, \infer\, \top\, and \train\. This modularization improves code organization and maintainability for the web UI.
src/llamafactory/webui/components · high confidence
Test coverage
Added comprehensive tests for data processing and formatting; Added end-to-end tests for chat, training, and SGLang backend; Added example scripts for image and tool-calling API tests; Added test for SFT trainer shuffling behavior; Added tests for DPO and FSDP2 training workflows; Added tests for FSDP2 meta-device loading and weight conversion; Added tests for core data and model loading components; Added tests for model plugins; Added tests for new batching strategies; Added tests for the v1 accelerator interface; Added tests for the v1 configuration argument parser; Added tests for the v1 rendering and message rendering logic; Added unit tests for Megatron Bridge training support; Added unit tests for data conversion plugins; Added unit tests for data processor modules; Added unit tests for evaluation template formatting; Added unit tests for model finetuning strategies; Added unit tests for model utilities; Added unit tests for the CLI sampler; Introduce automated license header checks and test configuration; New test configuration for LlamaFactory v1.
Dependencies
Project build system and dependencies modernized
The project has migrated from a legacy setuptools-based build to a modern pyproject.toml configuration using the hatchling build backend and uv for dependency management. This change enforces a minimum Python 3.11 runtime, updates core AI dependencies (including transformers, trl, and accelerate) to their latest supported versions, and introduces a new docs/requirements.txt for documentation builds. Users will benefit from improved build reliability, stricter dependency resolution, and support for newer Python versions.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 52 → 53 (+1.3)
Lenses
- Code Health 63 → 63 (-0.2)
- Architecture 89 → 88 (-1.4)
- Maturity 65 → 65 (+0.1)
- Readiness 39 → 39 (+0.0)
- Security 59 → 66 (+6.8)
Resolved (5)
- Build action pinned to a mutable branch
- Further orphaned files (smaller)
- Off-boarding risk: anonymized user #1
- Off-boarding risk: anonymized user #2
- Scanner failed to run — not a clean result
New (341)
- AlpacaDatasetConverter.call (cognitive 21) (src/llamafactory/data/converter.py)
- AlpacaDatasetConverter.call (cyclomatic 20) (src/llamafactory/data/converter.py)
- BaseTrainer.init (cognitive 19) (src/llamafactory/v1/core/base_trainer.py)
- BaseTrainer.fit (cognitive 50) (src/llamafactory/v1/core/base_trainer.py)
- BaseTrainer.fit (cyclomatic 16) (src/llamafactory/v1/core/base_trainer.py)
- Change coupling: cal_lr.py ↔ length_cdf.py (scripts/stat_utils/cal_lr.py)
- Change coupling: chatter.py ↔ chatbot.py (src/llamafactory/webui/chatter.py)
- Change coupling: chatter.py ↔ infer.py (src/llamafactory/webui/chatter.py)
- Change coupling: eval.py ↔ train.py (src/llamafactory/webui/components/eval.py)
- Change coupling: top.py ↔ runner.py (src/llamafactory/webui/components/top.py)
- Change coupling: train.py ↔ runner.py (src/llamafactory/webui/components/train.py)
- Change coupling: workflow.py ↔ workflow.py (src/llamafactory/train/dpo/workflow.py)
- Change coupling: workflow.py ↔ workflow.py (src/llamafactory/train/pt/workflow.py)
- Change coupling: workflow.py ↔ workflow.py (src/llamafactory/train/pt/workflow.py)
- Change coupling: workflow.py ↔ workflow.py (src/llamafactory/train/rm/workflow.py)
- CustomDPOTrainer.init (cognitive 20) (src/llamafactory/train/dpo/trainer.py)
- CustomKTOTrainer.init (cognitive 22) (src/llamafactory/train/kto/trainer.py)
- CustomKTOTrainer.log (cognitive 17) (src/llamafactory/train/kto/trainer.py)
- CustomPPOTrainer.init (cognitive 20) (src/llamafactory/train/ppo/trainer.py)
- CustomPPOTrainer.batched_forward_pass (cognitive 18) (src/llamafactory/train/ppo/trainer.py)
- …and 321 more
Changes since last survey
- 3 commits — 1 feature/other, 2 fixes
By area
- src/llamafactory — 2 commits
- .github/workflows — 1 commit
Notable commits
- fix: [train] Fix hyper parallel tail accumulation loss scaling (#10705)
- fix: fix(ci): align workflow Python version with requires-python (#10707)
- change: [v1] Support multimodal data training (#10656)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
hiyouga/LlamaFactory was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 2 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 62ae362455801d4900a5132c7a30b23dc5fc3802 — the exact code this score is about.
- Scored under rubric-2026.08.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer latest.