Skip to content
CAI
Software that uses CAICheck a score

Lightning-AI/pytorch-lightning

70.7

Strong · 18 September 2026

99.8k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a unified deep learning framework that provides two distinct APIs: PyTorch Lightning for high-level, declarative model training and Lightning Fabric for low-level, customizable training loops. It supports distributed training across various hardware accelerators, including GPUs, TPUs, and Apple Silicon, utilizing strategies like DDP, FSDP, and tensor parallelism. The library includes extensive tooling for experiment logging, checkpoint management, and hyperparameter tuning, alongside a comprehensive suite of examples and domain-specific templates for tasks ranging from computer vision to reinforcement learning.

How it got here

2019–2022 — Unified packaging and test expansion

34 changes.

This period focused on restructuring the project into a unified 'lightning' meta-package and modernizing its build infrastructure with centralized dependency management. A massive expansion of the test suite was conducted to cover core components, strategies, and plugins, while also establishing robust backward compatibility checks for legacy checkpoints.

2023–2024 — Fabric expansion and Lightning restructuring

65 changes.

This period focused on significantly expanding Lightning Fabric with comprehensive test coverage, new cluster environment plugins, and diverse training examples, while simultaneously restructuring the PyTorch Lightning codebase into a modular package architecture. The work introduced advanced distributed training capabilities, including tensor parallelism and FP8 support, alongside a unified plugin system for precision and checkpointing to improve maintainability and user experience.

Features

Add K-Fold Cross Validation example for Lightning Fabric

A new example demonstrating K-Fold cross validation using Lightning Fabric has been added to the examples directory. The \train\_fabric.py\ script shows how to use \scikit-learn\'s \KFold\ to split the MNIST dataset, train multiple CNN models in parallel across folds, and manage distributed training via Fabric's setup utilities. The accompanying \README.md\ provides instructions for running the example on CPU, GPU, or multiple GPUs using the \fabric run\ command.

_examples/fabric/kfold\cv · high confidence

Add MAML meta-learning example with Lightning Fabric and raw PyTorch implementations

A new example demonstrating Model Agnostic Meta Learning (MAML) on the Omniglot dataset has been added to the fabric examples. It provides two distributed training implementations: a raw PyTorch version using \torchrun\ and a Lightning Fabric version using \fabric run\, allowing users to compare boilerplate-heavy distributed code against the accelerated Fabric API for few-shot learning tasks.

_examples/fabric/meta\learning · high confidence

Add minimal Transformer language model example

A new example has been added to the Fabric examples directory that demonstrates a simple training loop for next-word prediction using a Transformer model on the WikiText2 dataset. The example includes a \train.py\ script and a \README.md\ with instructions for running the training on CPU, GPU, or multiple GPUs using the \fabric run\ command.

_examples/fabric/language\model · high confidence

Decoupled PPO Reinforcement Learning example for Fabric

A new decoupled Reinforcement Learning example has been added to the Fabric examples, implementing a Proximal Policy Optimization (PPO) agent using PyTorch and Gymnasium. This change replaces previous NumPy-based implementations with a pure PyTorch backend, introducing modular components for the agent (\PPOAgent\ and \PPOLightningAgent\), loss functions, and utilities. The example supports configurable activation functions (ReLU or Tanh), orthogonal initialization, and distributed training via a \share\_data\ flag, providing a modern, framework-native reference for training RL models with Lightning Fabric.

_examples/fabric/reinforcement\learning/rl · high confidence

Experimental API to validate model serving with FastAPI

A new experimental feature allows users to verify that their PyTorch Lightning models can be correctly served as REST APIs. By subclassing the new \ServableModule\ and implementing specific hooks (\configure\_payload\, \configure\_serialization\, \serve\_step\, and optionally \configure\_response\), users can attach the \ServableModuleValidator\ callback to their Trainer. This validator automatically spins up a local FastAPI server, sends a test request using the configured payload, and checks if the model's output matches the expected response, ensuring the model is ready for deployment before full integration.

src/lightning/pytorch/serve · high confidence

Fabric strategies module restructured and expanded

The \src/lightning/fabric/strategies\ package has been reorganized into a dedicated module with a public \\_\init\\_.py\ that explicitly exports all strategy classes (DDP, DeepSpeed, DataParallel, FSDP, ModelParallel, Parallel, SingleDevice, SingleXLA, XLA, XLAFSDP) and registers them in a global \STRATEGY\_REGISTRY\. This change introduces the \ModelParallelStrategy\ for 2D parallelism (FSDP2 + Tensor Parallelism) and the \XLAFSDPStrategy\ for TPU support, while also adding a dedicated \SingleDeviceXLAStrategy\ for single-XLA-device training. The \DeepSpeedStrategy\ now supports \exclude\_frozen\_parameters\ and learning rate scheduling, and the \FSDPStrategy\ exposes \sharding\_strategy\ as a string and supports \device\_mesh\ for hybrid sharding.

src/lightning/fabric/strategies · high confidence

Introduce Fabric loggers (CSV and TensorBoard)

The \src/lightning/fabric/loggers\ module now provides \CSVLogger\ and \TensorBoardLogger\ for Fabric experiments. \CSVLogger\ saves metrics to local CSV files with automatic versioning and optional experiment names, while \TensorBoardLogger\ writes to TensorBoard format, supporting both \tensorboard\ and \tensorboardX\ backends, subdirectories, and hyperparameter logging.

src/lightning/fabric/loggers · high confidence

Introduce \`lightning.pytorch.core\` package with core module definitions

The \lightning.pytorch.core\ package has been introduced, consolidating the foundational components of the PyTorch integration. This location now explicitly exports \LightningModule\ and \LightningDataModule\ via its \\_\init\\_.py\, while housing the implementation of these classes, their associated hooks (\hooks.py\), optimizer wrapping logic (\optimizer.py\), and checkpoint saving/loading utilities (\saving.py\). This structure centralizes the core abstractions that users interact with when defining models and data pipelines.

src/lightning/pytorch/core · high confidence

Introduce experimental collective communication plugin interface

A new experimental API for collective operations has been added to Lightning Fabric under \lightning.fabric.plugins.collectives\. This introduces a \Collective\ abstract base class defining standard distributed operations (broadcast, all\_reduce, gather, etc.), along with a \TorchCollective\ implementation that wraps \torch.distributed\ and a \SingleDeviceCollective\ stub for single-device scenarios. Users can now access these classes directly from the \lightning.fabric.plugins.collectives\ module to perform collective communications in a framework-agnostic way.

src/lightning/fabric/plugins/collectives · high confidence

Introduces dedicated launcher components for Fabric distributed training

The \src/lightning/fabric/strategies/launchers\ package now provides a structured set of classes to manage process creation for distributed training. This includes \\_SubprocessScriptLauncher\ for spawning child processes on a single node (supporting Hydra configuration for working directories), \\_MultiProcessingLauncher\ for multi-process scenarios, and \\_XLALauncher\ for launching parallel workers on XLA hardware. These components centralize the logic for environment variable setup, process monitoring, and inter-process communication within the Fabric strategy layer.

src/lightning/fabric/strategies/launchers · high confidence

Introduction of the unified 'lightning' meta-package

The \lightning\ package is now available as a single entry point that aggregates core components from PyTorch Lightning and Lightning Fabric. Users can install this unified package to access top-level imports such as \Trainer\, \LightningModule\, \LightningDataModule\, \Callback\, \Fabric\, and \seed\_everything\ directly from the \lightning\ namespace. The package enforces a minimum Python version of 3.10 and includes a CLI entry point for the Fabric tool.

src/lightning · high confidence

Lightning Fabric package structure and build configuration

The \src/lightning\fabric\ directory now contains the core package metadata and build configuration files, including \MANIFEST.in\ to define included assets (such as version info, changelog, and requirements), \\\about\\.py\ for package metadata, \\\setup\\.py\ for the setuptools configuration (specifying Python 3.10+ support and entry points), and \\\version\\_.py\ for version management. Additionally, a \py.typed\ marker file is added to enable PEP 561 type checking support for consumers of the library.

_src/lightning\fabric · high confidence

New 'Build Your Own Trainer' example for Lightning Fabric

A new example located at examples/fabric/build\_your\_own\_trainer demonstrates how to construct a fully customizable training loop using Lightning Fabric. The example provides a \MyCustomTrainer\ implementation in \trainer.py\ and a usage script \run.py\ that trains an MNIST model, illustrating key Fabric features such as hardware acceleration selection, distributed sampler handling, gradient accumulation, and checkpointing without relying on the full PyTorch Lightning Trainer.

_examples/fabric/build\_your\_own\trainer · high confidence

New DCGAN example demonstrating Lightning Fabric acceleration

Added a new example in the Fabric section that implements a Deep Convolutional Generative Adversarial Network (DCGAN) for generating realistic face images using the CelebA dataset. The example provides two implementations: a standard raw PyTorch version and a Lightning Fabric version, allowing users to compare the boilerplate reduction and hardware acceleration capabilities (including Apple Silicon support) offered by Fabric.

examples/fabric/dcgan · high confidence

New FP8 distributed transformer examples for Fabric and PyTorch Lightning

Added new example scripts and documentation for both Lightning Fabric and PyTorch Lightning that demonstrate training a Transformer model using FP8 low-precision quantization, FSDP2 (DTensor) for distributed sharding, and torch.compile for optimization. The Fabric example uses the ModelParallelStrategy with a configure\_model hook to apply quantization and compilation, while the PyTorch Lightning example implements the same logic within a LightningModule's configure\_model method, providing users with reference implementations for memory-efficient, high-throughput distributed training on supported hardware.

_examples/fabric/fp8\_distributed\_transformer, examples/pytorch/fp8\_distributed\transformer · high confidence

New Fabric MNIST example demonstrating scalable training

Added a new example in examples/fabric/image\_classifier that contrasts a vanilla PyTorch MNIST classifier with a Lightning Fabric version. The Fabric script (train\_fabric.py) shows how to use Fabric to enable GPU and multi-GPU training via the \fabric run\ CLI command, handle device placement automatically, and use \rank\_zero\_first\ for data downloads, providing a concrete reference for scaling PyTorch code with Fabric.

_examples/fabric/image\classifier · high confidence

New Fabric Reinforcement Learning example with coupled and decoupled PPO architectures

The \examples/fabric/reinforcement\_learning\ directory now contains a complete Proximal Policy Optimization (PPO) implementation accelerated by Lightning Fabric. This example demonstrates two distributed training patterns: a coupled architecture where the agent and environment run on the same processes, and a decoupled architecture that separates the 'Player' (environment interaction) from the 'Trainers' (model optimization) using Fabric collectives. The entry includes the \train\_fabric.py\ and \train\_fabric\_decoupled.py\ scripts, a raw PyTorch baseline (\train\_torch.py\) for comparison, and a \README.md\ detailing usage, hyperparameters, and visualization steps.

_examples/fabric/reinforcement\learning · high confidence

New Fabric example for 2D tensor and data parallelism

Added a new example in \examples/fabric/tensor\_parallel\ demonstrating how to combine tensor parallelism (TP) and data parallelism (DP) using the \ModelParallelStrategy\ with PyTorch Fabric. The example shows how to parallelize a Llama 3 7B model across multiple GPUs by applying tensor-parallel strategies to model layers and using Fully Sharded Data Parallel (FSDP) for data parallelism, including configuration of mixed precision, activation checkpointing, and distributed checkpointing.

_examples/fabric/tensor\parallel · high confidence

New PyTorch Lightning domain templates for computer vision, GANs, and reinforcement learning

Added new runnable example scripts in \examples/pytorch/domain\_templates\ that demonstrate specific deep learning patterns using PyTorch Lightning. These include a computer vision fine-tuning template using \BaseFinetuning\ and \LightningDataModule\, a Generative Adversarial Network (GAN) template with manual optimization, an ImageNet training template using \LightningCLI\, and reinforcement learning templates for Deep Q-Networks (DQN) and Proximal Policy Optimization (PPO) using OpenAI Gym.

_examples/pytorch/domain\templates · high confidence

New PyTorch basic examples and bug-report template

The \examples/pytorch/basics\ directory now includes runnable starter scripts for common patterns: an MNIST autoencoder (\autoencoder.py\), a backbone image classifier (\backbone\_image\_classifier.py\), a Transformer-based next-word predictor on WikiText2 (\transformer.py\), and a PyTorch Profiler demo (\profiler\_example.py\). Each example is documented in \README.md\ with run instructions for CPU, GPU, and DDP. Additionally, a bug-report template is provided in \examples/pytorch/bug\_report\ (notebook and Python script) to help users reproduce issues with a minimal Lightning model.

examples/pytorch/basics · high confidence

New PyTorch tensor-parallel example with 2D parallelism support

Added a new example in \examples/pytorch/tensor\_parallel\ demonstrating how to apply tensor parallelism to a Llama 3 7B model using PyTorch Lightning's \ModelParallelStrategy\. The example shows how to combine tensor parallelism with FSDP for 2D parallelism, requiring PyTorch 2.3+ and at least 4 GPUs with 24 GB memory each. It includes a complete training script (\train.py\) that configures the \ModelParallelStrategy\ with \data\_parallel\_size=2\ and \tensor\_parallel\_size=2\, along with helper modules for data handling (\data.py\) and parallelization logic (\parallelism.py\).

_examples/pytorch/tensor\parallel · high confidence

New build assistant and legacy checkpoint retrieval tooling

The repository now includes a new Python build assistant (\.actions/assistant.py\) that manages requirements for PyTorch, Fabric, and Data packages, featuring logic to adjust version constraints (freezing, unfreezing, or strict enforcement) and to generate package descriptions from READMEs. Additionally, a new shell script (\.actions/pull\_legacy\_checkpoints.sh\) has been added to download and extract legacy model checkpoints from S3 into the test directory, supporting tests that depend on historical model artifacts.

.actions · high confidence

New callbacks module exposes EMA weight averaging and device stats filtering

The \src/lightning/pytorch/callbacks\ package now includes the \EMAWeightAveraging\ and \WeightAveraging\ callbacks for generic weight averaging, and the \DeviceStatsMonitor\ callback now supports a \filter\keys\ argument to log only specified device statistics. These additions are exported via the \\\init\\_.py\ module, making them available for import from \lightning.pytorch.callbacks\.

src/lightning/pytorch/callbacks · high confidence

New cluster environment plugins for Fabric

The \lightning.fabric.plugins.environments\ module now provides a set of cluster environment plugins that allow Fabric to automatically detect and configure distributed training across various schedulers and runtimes. This includes support for SLURM, LSF, MPI, TorchElastic, Kubeflow, and Google Cloud TPUs (via XLA). These environments handle the discovery of main addresses, ports, ranks, and world sizes, enabling seamless distributed execution on HPC clusters and cloud platforms without manual configuration.

src/lightning/fabric/plugins/environments · high confidence

New demo modules for LSTM and Transformer models

The \lightning.pytorch.demos\ package now includes \LightningLSTM\ and \LightningTransformer\ classes, providing ready-to-use examples for sequence modeling tasks. The LSTM demo implements a language model using PyTorch's LSTM module with a custom \SequenceSampler\, while the Transformer demo provides a \LightningTransformer\ wrapper around a standard Transformer architecture with positional encoding, along with a \WikiText2\ dataset class for downloading and tokenizing text data. These additions are exported via the package's \\_\init\\_.py\ alongside existing boring classes and MNIST utilities.

src/lightning/pytorch/demos · high confidence

New distributed training overrides and utilities in PyTorch Lightning

The \src/lightning/pytorch/overrides\ module has been introduced to centralize PyTorch-specific overrides for distributed training. This includes a \DistributedDataParallel\ wrapper that ensures \prepare\_for\_backward\ is called correctly during manual optimization and handles unused parameter detection, a helper to register custom DDP communication hooks (such as FP16 compression or PowerSGD), and a utility to sync module states across processes. Additionally, an \UnrepeatedDistributedSampler\ is provided to allow deterministic, non-repeating data sampling during prediction phases, addressing issues where processes might run different numbers of batches.

src/lightning/pytorch/overrides · high confidence

New examples and core loop infrastructure for tensor parallelism and production serving

This release introduces new example scripts demonstrating tensor-parallel model implementation for both Fabric and PyTorch Lightning, alongside a production-serving example using \ServableModule\. Internally, the core training and evaluation loops have been refactored to support these patterns, including changes to how data is fetched and processed, and updates to the multiprocessing launcher to better handle global state snapshots and CUDA context checks. The Rich progress bar has also been enhanced with custom columns and infinite task support to improve visibility during training.

python · high confidence

New utilities module for PyTorch Lightning

The \src/lightning/pytorch/utilities\ package has been introduced to centralize general-purpose helper functions and classes. This new module exposes utilities such as \CombinedLoader\ for managing multiple data loaders, \AttributeDict\ for structured argument handling, and \is\_picklable\ for serialization checks. It also includes tools for gradient inspection (\grad\_norm\), shared parameter management, and environment-variable-based configuration defaults. By consolidating these components, the library provides a single, organized location for common utilities used across the Trainer, Fabric, and callbacks.

src/lightning/pytorch/utilities · high confidence

Project initialization with Apache 2.0 license and development tooling

The repository has been initialized with a new project structure, including a Makefile for managing setup, tests, and documentation builds, and a .pre-commit-config.yaml enforcing code quality via Ruff, mdformat, and other linters. The project license has been changed from MIT to Apache 2.0, and a CITATION.cff file has been added to facilitate academic citation. Additionally, configuration files for Codecov, Read the Docs, and Git (including .gitignore and .gitmodules for the tutorials submodule) are now present to support continuous integration and documentation generation.

(repo-wide) · high confidence

Architecture

Trainer module restructured into a dedicated package

The \Trainer\ core logic has been reorganized from a single monolithic file into a structured package under \src/lightning/pytorch/trainer\. This change introduces dedicated modules for specific responsibilities: \call.py\ handles hook execution and interrupt management, \setup.py\ manages initialization flags and profiler setup, and \configuration\validator.py\ enforces loop configuration rules. The main \Trainer\ class is now imported from this new package structure, and the \\\init\\_.py\ exposes \Trainer\ and \seed\_everything\ for public use.

src/lightning/pytorch/trainer · high confidence

Behavioural changes

Accelerator implementations moved to lightning.pytorch.accelerators

The PyTorch Lightning accelerator implementations (CPU, CUDA, MPS, and XLA) have been relocated from the legacy \\pytorch\_lightning.accelerators\\ namespace to \\lightning.pytorch.accelerators\\. This change consolidates the accelerator logic within the new \\lightning\\ package structure, ensuring that users importing from the new namespace will access the correct, updated implementations for device management and statistics.

src/lightning/pytorch/accelerators · high confidence

Automatic legacy checkpoint migration and unpickling support

Lightning now automatically upgrades legacy checkpoints when loading them, applying version-specific migrations (e.g., for loop structures, progress tracking, and callback keys) to match the current format. The system also handles unpickling of older checkpoints that contain removed classes or modules (such as \\_FaultTolerantMode\ or \pytorch\_lightning\ imports) by temporarily patching the environment, ensuring seamless resumption of training from older saved states.

src/lightning/pytorch/utilities/migration · high confidence

Backward compatibility shims for renamed and removed components

The library now includes a compatibility layer in the \\_graveyard\ module to support legacy imports and class names. Deprecated precision plugins (e.g., \PrecisionPlugin\, \HalfPrecisionPlugin\) are aliased to their new counterparts (e.g., \Precision\, \HalfPrecision\) with deprecation warnings. TPU-specific classes (\TPUAccelerator\, \SingleTPUStrategy\) are mapped to their XLA equivalents (\XLAAccelerator\, \SingleDeviceXLAStrategy\). Removed Habana (HPU) integrations are replaced with stub classes that raise \NotImplementedError\ to clearly indicate removal. Additionally, a patch ensures older versions of \torchmetrics\ (\<0.8.0) correctly resolve internal imports to the unified \lightning.pytorch\ package.

_src/lightning/pytorch/\graveyard · high confidence

Enhanced ModelSummary with training mode, FLOPs, and Deepspeed support

The model summary output now includes columns for the module's training mode (train/eval) and estimated FLOPs, providing deeper insight into model behavior and computational cost. It also correctly handles non-layer parameters by listing them separately and supports PyTorch DTensors. For users of the DeepSpeed strategy, a specialized summary is available that displays parameters per device and fixes column alignment issues.

_src/lightning/pytorch/utilities/model\summary · high confidence

Fabric 2.6.5 release with remote checkpoint support and bug fixes

This release introduces support for remote storage (fsspec URLs) when saving and loading distributed checkpoints with FSDPStrategy, ModelParallelStrategy, and the checkpoint consolidation CLI. It also adds \using\_sparse\_model\ and \sparse\_cuda\_acceleration\_factor\ parameters to the \Throughput\ utility to correctly default to dense peak FLOPs and allow explicit sparse peak opt-in. Several bugs are fixed: FSDP mixed precision now correctly initializes parameters in fp32 instead of half precision, the \BitsandbytesPrecision\ plugin type checking is corrected for bitsandbytes 0.48/0.49, DeepSpeed checkpoint path validation now accepts remote filesystem URIs (S3, GCS, HDFS), and a RuntimeError under PyTorch 2.5+ when FSDP's \root\_device\ is CPU is resolved by passing an explicit \torch.device("cpu")\. Additional fixes address an AccumulateGrad stream mismatch warning with DDP, incorrect CUDA device initialization in \CUDAAccelerator.setup\_device\, and \PermissionError\ swallowing during atomic checkpoint saves. Dead code paths for PyTorch below 2.6 have been removed.

src/lightning/fabric · high confidence

Fabric accelerator implementations and registry are restructured

The \src/lightning/fabric/accelerators\ module now provides the concrete accelerator classes (CPU, CUDA, MPS, XLA) and the \\_AcceleratorRegistry\ that Fabric uses to select hardware. This change introduces a registry-based lookup for accelerators, ensuring that the correct hardware implementation is instantiated based on the user's configuration. It also includes specific fixes for CUDA device initialization (ensuring \set\_device\ is called before precision checks) and XLA runtime detection (supporting \torch-xla\>=2.5\ PJRT checks).

src/lightning/fabric/accelerators · high confidence

Fabric precision plugin API restructured with new context managers and plugin classes

The precision plugin system in Fabric has been reorganized to support a more granular control over model initialization and data conversion. The base \Precision\ class now exposes \tensor\_init\_context\ and \module\_init\context\ hooks, allowing plugins to control tensor creation and module instantiation separately. New concrete plugins have been introduced: \BitsandbytesPrecision\ for quantization (supporting modes like \nf4\, \fp4\, and \int8\), \TransformerEnginePrecision\ for FP8 training via NVIDIA's Transformer Engine, and \XLAPrecision\ for XLA devices. Existing plugins like \MixedPrecision\, \FSDPPrecision\, \DeepSpeedPrecision\, \HalfPrecision\, and \DoublePrecision\ have been updated to align with the new API, including changes to how they handle input/output conversion and optimizer steps (e.g., \FSDPPrecision\ now uses \ShardedGradScaler\). The \\\init\\_.py\ module now explicitly exports all these precision plugins.

src/lightning/fabric/plugins/precision · high confidence

Fabric utilities module reorganization and new checkpointing features

The \src/lightning/fabric/utilities\ package has been reorganized into a dedicated module, exposing a consolidated set of utilities including \AttributeDict\, \Throughput\ monitoring, and data-handling helpers like \move\_data\_to\_device\. This change introduces atomic checkpoint saving via \fsspec\ to prevent corruption during interruptions, supports remote storage URLs for FSDP checkpoints, and adds a CLI tool to consolidate sharded FSDP checkpoints into a single file. Additionally, the \\_DeviceDtypeModuleMixin\ now respects PyTorch 2.8's default device context, and the \load\ module implements lazy checkpoint loading to reduce memory overhead when loading large models.

src/lightning/fabric/utilities · high confidence

Hyperparameter handling logic moved to a dedicated mixin

The \save\_hyperparameters\ functionality and related hyperparameter management (including \hparams\ and \hparams\_initial\ properties) have been extracted from the core model classes into a new \HyperparametersMixin\ in \src/lightning/pytorch/core/mixins\. This change restructures how hyperparameters are stored and accessed, making the logic available for reuse and potentially altering inheritance patterns for classes that previously relied on the previous implementation location.

src/lightning/pytorch/core/mixins · medium confidence

Introduce Lightning Data as a wrapper for the litdata package

The \src/lightning/data\ module now acts as a compatibility layer that re-exports the core functionality of the \litdata\ package. Users can continue to import \StreamingDataset\, \optimize\, \map\, and other data processing tools from \lightning.data\, but the implementation is now sourced from the external \litdata\ library. This change requires the \litdata\ package to be installed, as the module explicitly checks for its presence and raises an error if it is missing, effectively decoupling the data loading logic from the main Lightning repository.

src/lightning/data · high confidence

Introduction of structured logging validation and integration suggestions in Trainer logs

The logger connector now includes an \fx\_validator\ that strictly defines which Lightning hooks allow \self.log()\ calls and enforces the \on\_step\/\on\_epoch\ logging levels, raising clear errors for invalid usage. Additionally, when \Trainer(suggest\_integrations=True)\ (the default), the connector checks if a \LitLogger\ is present and, if not, suggests installing the \litlogger\ package for seamless cloud logging.

_src/lightning/pytorch/trainer/connectors/logger\connector · high confidence

New CheckpointIO plugin interface with weights\_only support

The \src/lightning/fabric/plugins/io\ module now exposes a new \CheckpointIO\ interface along with \TorchCheckpointIO\ and \XLACheckpointIO\ implementations. This change introduces a \weights\_only\ parameter to checkpoint loading, allowing users to restrict loading to safe state dicts for improved security when handling untrusted sources. The \XLACheckpointIO\ implementation specifically handles TPU-specific requirements, including OmegaConf compatibility workarounds.

src/lightning/fabric/plugins/io · high confidence

New plugin architecture for layer synchronization and precision management

The \src/lightning/pytorch/plugins\ module has been restructured to introduce a dedicated plugin system for managing model layer synchronization and precision. This change adds a new \LayerSync\ abstract base class and a concrete \TorchSyncBatchNorm\ implementation, allowing users to wrap batch normalization layers with synchronization logic for multi-GPU and multi-node training scenarios. Additionally, the module now explicitly exports a comprehensive set of precision plugins—including \BitsandbytesPrecision\, \TransformerEnginePrecision\, \FSDPPrecision\, and \XLAPrecision\—alongside standard options like \MixedPrecision\ and \HalfPrecision\, centralizing how these configurations are accessed and applied within the PyTorch Lightning Trainer.

src/lightning/pytorch/plugins · high confidence

New precision plugin architecture in PyTorch Lightning

The precision handling in PyTorch Lightning has been reorganized into a new plugin layout under \src/lightning/pytorch/plugins/precision\. This change introduces a unified \Precision\ base class and dedicated plugins for various precision strategies, including \MixedPrecision\ (AMP), \HalfPrecision\, \DoublePrecision\, \FSDPPrecision\, \DeepSpeedPrecision\, \XLAPrecision\, \BitsandbytesPrecision\, and \TransformerEnginePrecision\. Users can now configure these specific precision modes via the \Trainer\'s \precision\ argument, which routes to the appropriate plugin to handle dtype conversion, gradient scaling, and optimizer stepping logic for each strategy.

src/lightning/pytorch/plugins/precision · high confidence

Profiler module relocated to lightning.pytorch.profilers

The profiler implementations (AdvancedProfiler, PyTorchProfiler, SimpleProfiler, XLAProfiler, and PassThroughProfiler) have been moved from the legacy pytorch\_lightning namespace to the new lightning.pytorch.profilers package. Users should update their imports to use the new location, for example: from lightning.pytorch.profilers import PyTorchProfiler.

src/lightning/pytorch/profilers · high confidence

Progress bar logic reorganized into a dedicated module

The progress bar implementation has been restructured into a new \src/lightning/pytorch/callbacks/progress\ package, separating the base \ProgressBar\ class, the default \TQDMProgressBar\, and the \RichProgressBar\ into distinct files. This change consolidates the progress bar code, making it easier to maintain and extend, while preserving the existing functionality for tracking training, validation, and testing progress.

src/lightning/pytorch/callbacks/progress · high confidence

PyTorch Lightning loggers module restructured and LitLogger added

The \src/lightning/pytorch/loggers\ package has been reorganized to expose a unified \\_\init\\_.py\ that imports and exports all available loggers, including the newly added \LitLogger\ for remote experiment tracking on Lightning AI. The \NeptuneLogger\ has been removed and replaced with a deprecated stub that raises a \RuntimeError\ directing users to migrate to \LitLogger\. The module now includes updated implementations for \CometLogger\, \CSVLogger\, \MLFlowLogger\, \TensorBoardLogger\, and \WandbLogger\, along with shared utilities for checkpoint scanning and hyperparameter logging.

src/lightning/pytorch/loggers · high confidence

PyTorch Lightning package structure and metadata reorganization

The \src/pytorch\lightning\ directory has been restructured to support the new \src/\ layout and unified packaging strategy. A new \MANIFEST.in\ ensures that version info, changelogs, READMEs, and requirement files for both PyTorch Lightning and Lightning Fabric are correctly included in the distribution. The package now includes a \py.typed\ marker file to enable PEP 561 type checking support. Metadata files (\\\about\\.py\, \\\version\\.py\, \\\setup\\_.py\) have been updated to reflect the new source root, load version information from \version.info\, and configure the package to include both \pytorch\_lightning\ and \lightning\_fabric\ source code. The setup script now explicitly requires Python 3.10+ and defines classifiers for Python 3.10 through 3.13.

_src/pytorch\lightning · high confidence

Reorganized cluster environment plugins under the new lightning package structure

The cluster environment plugins have been reorganized to align with the new package structure. The \src/lightning/pytorch/plugins/environments/\_\init\\_.py\ file now serves as a central re-export point, importing \ClusterEnvironment\ from \lightning.fabric.plugins\ and exposing specific environment classes (\KubeflowEnvironment\, \LightningEnvironment\, \LSFEnvironment\, \MPIEnvironment\, \SLURMEnvironment\, \TorchElasticEnvironment\, \XLAEnvironment\) from \lightning.fabric.plugins.environments\. This change consolidates the availability of these environment configurations for users migrating to the new \lightning\ namespace.

src/lightning/pytorch/plugins/environments · high confidence

Reorganized distributed training strategies into a new package structure

The distributed training strategies (DDP, DeepSpeed, FSDP, ModelParallel, Parallel, SingleDevice, and XLA) have been reorganized into the new \src/lightning/pytorch/strategies\ package. This change introduces a centralized \\_\init\\_.py\ that registers these strategies with the global strategy registry, making them available for import from \lightning.pytorch.strategies\. The individual strategy files have been updated to import from the new location and register themselves, ensuring that users can continue to use these strategies via the standard Lightning entry points without code changes.

src/lightning/pytorch/strategies · high confidence

Restructure PyTorch Lightning launchers into a dedicated module

The launcher implementations for multiprocessing, subprocess scripts, and XLA have been moved into a new \src/lightning/pytorch/strategies/launchers\ package. This change introduces a base \\_Launcher\ class that extends the Fabric launcher with a \kill\ method for process termination, and exposes the specific launchers (\\_MultiProcessingLauncher\, \\_SubprocessScriptLauncher\, \\XLALauncher\) via a new \\\init\\_.py\. The \\_SubprocessScriptLauncher\ now validates cluster environment settings before spawning child processes, and the \\_XLALauncher\ enforces a restriction against calling \trainer.fit()\ twice on the same instance when using spawn-based strategies.

src/lightning/pytorch/strategies/launchers · high confidence

Restructured and refactored internal training loops and data fetching

The internal loop architecture in \src/lightning/pytorch/loops\ has been reorganized: the \\_\init\\_.py\ now explicitly exports core loop classes (\\_FitLoop\, \\_EvaluationLoop\, \\_PredictionLoop\, \\_TrainingEpochLoop\, and optimization loops), and a new \fetchers.py\ module centralizes data fetching logic with dedicated classes like \\_PrefetchDataFetcher\ and \\_DataLoaderIterDataFetcher\. This refactoring moves data-fetching responsibilities out of the loops themselves, introduces a unified \CombinedLoader\ assumption for fetcher inputs, and adds utilities for selecting the appropriate fetcher based on hook signatures (e.g., \dataloader\_iter\ support). Users benefit from a more modular internal structure that supports flexible data iteration patterns while maintaining backward compatibility with existing training, validation, and prediction workflows.

src/lightning/pytorch/loops · high confidence

Restructured examples with Fabric and Trainer scripts

The examples directory has been reorganized to clearly separate Lightning Fabric and Lightning Trainer use cases. A new README provides navigation to Fabric examples (MNIST, DCGAN) and Trainer examples (Image Classifier, Autoencoder). Additionally, shell scripts (run\_fabric\_examples.sh, run\_pl\_examples.sh) have been added to simplify running these examples with default parameters.

examples · high confidence

Trainer connectors restructured with new default callbacks and signal handling

The \src/lightning/pytorch/trainer/connectors\ module has been reorganized into distinct connector classes (\\_AcceleratorConnector\, \\_CallbackConnector\, \\_CheckpointConnector\, \\_SignalConnector\) to manage specific Trainer configurations. By default, the Trainer now uses \RichProgressBar\ and \RichModelSummary\ when the \rich\ library is available, falling back to \TQDMProgressBar\ and \ModelSummary\ otherwise. A new \LitLogger\ integration is supported, with a tip displayed to users if \litmodels\ is installed but not actively used for model logging. Additionally, signal handling for SLURM auto-requeueing and SIGTERM has been centralized in \\_SignalConnector\ to improve graceful shutdown and job requeuing reliability.

src/lightning/pytorch/trainer/connectors · high confidence

Tuner module relocated to lightning.pytorch.tuner with security and stability fixes

The tuner functionality (LearningRateFinder and BatchSizeFinder) has been moved into the \lightning.pytorch.tuner\ package, exposing the \Tuner\ class for model tuning. This change includes a security fix for the LearningRateFinder to support the \weights\_only\ parameter when restoring checkpoints, and addresses stability issues such as resetting the epoch loop restarting flag and fixing gradient calculation for exponential mode.

src/lightning/pytorch/tuner · high confidence

Fixes

AsyncCheckpointIO snapshots tensors to prevent race conditions

The AsyncCheckpointIO plugin now clones checkpoint tensors before submitting them to the background thread. This snapshotting prevents race conditions where model parameters might be mutated by the training loop while the checkpoint is being saved asynchronously, ensuring data integrity during concurrent save operations.

src/lightning/pytorch/plugins/io · high confidence

Test coverage

Added comprehensive test coverage for Rich and TQDM progress bars; Added comprehensive test suite for PyTorch precision plugins; Added comprehensive tests for PyTorch trainer logging behavior; Added internal testing utility for conditional test execution; Added parity and benchmarking tests for PyTorch integration; Added parity tests for Lightning Fabric against raw PyTorch; Added test coverage for Fabric accelerator implementations; Added test coverage for PyTorch Tuner components; Added test coverage for PyTorch plugins; Added test helper utilities for Fabric; Added test infrastructure for legacy checkpoint compatibility; Added test script for LightningCLI integration; Added test suite for PyTorch Trainer configuration and validation; Added test suite for PyTorch core components; Added test suite for Trainer connectors; Added tests for Fabric CSV and TensorBoard loggers; Added tests for Fabric cluster environment plugins; Added tests for Fabric collectives plugins; Added tests for Fabric multiprocessing and subprocess launchers; Added tests for LSTM and Transformer demo utilities; Added tests for PyTorch Lightning utilities; Added tests for PyTorch accelerator implementations; Added tests for PyTorch checkpoint migration utilities; Added tests for PyTorch distributed overrides; Added tests for PyTorch multiprocessing and subprocess script launchers; Added tests for RunIf utility; Added tests for ServableModuleValidator callback; Added tests for Trainer flag configurations; Added tests for deprecated and removed PyTorch Lightning components; Added tests for deprecated precision plugins and API behaviors; Added tests for trainer optimization and dynamic arguments; Fabric test suite reorganization and isolation improvements; New PyTorch test suite structure and fixtures; New test helper utilities and fixtures for PyTorch Lightning; New test suite for Fabric distributed strategies; New test suite for PyTorch Lightning strategies; New test suite for PyTorch loggers; New testing utilities for conditional test skipping.

Dependencies

Centralized and structured requirements management

The project has reorganized its dependency management by introducing a structured \requirements/\ directory. This change consolidates dependencies into specific files for different environments: \ci.txt\ for continuous integration, \docs.txt\ for documentation building (including Sphinx and related tools), \doctests.txt\ for testing, and \typing.txt\ for type checking (pinning MyPy and Torch). A new \collect\_env\_details.py\ script has been added to help diagnose system issues by reporting OS, CUDA, and installed package versions. Additionally, a \README.md\ explains the new structure and the use of upper bounds for CI stability.

requirements · high confidence

Introduce pyproject.toml and structured requirements files

The project now uses a centralized pyproject.toml for build configuration and tooling settings (including Ruff, mypy, and pytest), replacing ad-hoc configurations. Dependency management is restructured into a modular system: the root requirements.txt now aggregates base requirements from ./requirements/fabric/base.txt and ./requirements/pytorch/base.txt, while specific examples (such as FP8 distributed transformers and reinforcement learning) have their own dedicated requirements.txt files to isolate their dependencies.

(dependencies) · high confidence

PyTorch requirements updated to support PyTorch 2.6 and newer dependency versions

The PyTorch package requirements have been updated to support PyTorch 2.6 (minimum) and up to 2.14.0, along with corresponding updates to torchvision (\>=0.21.0), torchmetrics (\>0.7.0), and other core dependencies like fsspec, tqdm, and typing-extensions. The DeepSpeed strategy requirement has been bumped to \>=0.16.0, and test dependencies have been updated to use pytest 9.0.2 and coverage 7.13.5. These changes ensure compatibility with the latest PyTorch releases and their associated ecosystem libraries.

requirements/pytorch · high confidence

Housekeeping

Lightning PyTorch 2.0.4 and 2.0.5 changelog update; Update development version to 2.7.0.dev.

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 71.

Lenses

  • Code Health 86
  • Architecture 98
  • Maturity 64
  • Readiness 72
  • Security 82
  • Domain Modelling 77

Changes since last survey

  • 300 commits — 229 feature/other, 71 fixes

By area

  • src/lightning — 101 commits
  • .github/workflows — 40 commits
  • requirements/pytorch — 31 commits
  • docs/source-pytorch — 22 commits
  • (root) — 21 commits
  • requirements/fabric — 20 commits
  • tests/tests_pytorch — 16 commits
  • requirements/docs.txt — 12 commits
  • requirements/ci.txt — 6 commits
  • .lightning/workflows — 5 commits
  • requirements/typing.txt — 5 commits
  • requirements/doctests.txt — 4 commits
  • tests/legacy — 4 commits
  • .github/checkgroup.yml — 3 commits
  • .github/markdown-links-config.json — 2 commits
  • docs/source-fabric — 2 commits
  • .actions/assistant.py — 1 commit
  • .azure/README.md — 1 commit
  • .azure/gpu-benchmarks.yml — 1 commit
  • .github/CODEOWNERS — 1 commit

Notable commits

  • fix: Bugfix for BackboneFinetuning + LearningRateFinder (#21224)
  • fix: CI: fix doctest failure from PyTorch LeafSpec FutureWarning (#21502)
  • fix: CUDAAccelerator.setup_device: fix unrelated device init by matmul precision check (#21726)
  • fix: Fix DeepSpeed checkpoint path validation for remote filesystem URIs (#21636)
  • fix: Fix EADDRINUSE errors in distributed tests with port manager and retry logic (#21309)
  • fix: Fix FLOPs inconsistency and add using_sparse_model flag (#21743)
  • fix: Fix FSDP CPU offloading when shards are already on CPU (#21805)
  • fix: Fix FSDP mixed precision semantics and add user warning (#21361)
  • fix: Fix FSDPPrecision rejecting a scaler for 16-mixed precision (#21831)
  • fix: Fix LearningRateMonitor TypeError on MPS with float64 lr values (#21545)
  • fix: Fix LightningCLI loading of hyperparameters from ckpt_path failing for subclass model mode (#21246)
  • fix: Fix LightningDataModule.load_from_checkpoint to restore subclass and hyperparameters (#21478)
  • fix: Fix ModelCheckpoint file_exists OOM in DDP (#21380)
  • fix: Fix ModelParallelStrategy fails with non-distributed checkpoint. (#21384)
  • fix: Fix PT011 (#21197)
  • fix: Fix PermissionError swallowing in _atomic_save (#21799)
  • fix: Fix RichModelSummary model size display (#21467)
  • fix: Fix ModelCheckpoint with manual optimization and every_n_train_steps (#21239)
  • fix: Fix StochasticWeightAveraging with infinite epochs (#21396)
  • fix: Fix _generate_seed_sequence sampling (#21399)
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

Lightning-AI/pytorch-lightning was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit fab6fb2dcf25c0920b8c2a6174258ffa66ea3337 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.