Lightning-AI/pytorch-lightning
70.7
Strong · 18 September 2026
99.8k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a unified deep learning framework that provides two distinct APIs: PyTorch Lightning for high-level, declarative model training and Lightning Fabric for low-level, customizable training loops. It supports distributed training across various hardware accelerators, including GPUs, TPUs, and Apple Silicon, utilizing strategies like DDP, FSDP, and tensor parallelism. The library includes extensive tooling for experiment logging, checkpoint management, and hyperparameter tuning, alongside a comprehensive suite of examples and domain-specific templates for tasks ranging from computer vision to reinforcement learning.
How it got here
2019–2022 — Unified packaging and test expansion
34 changes.
This period focused on restructuring the project into a unified 'lightning' meta-package and modernizing its build infrastructure with centralized dependency management. A massive expansion of the test suite was conducted to cover core components, strategies, and plugins, while also establishing robust backward compatibility checks for legacy checkpoints.
2023–2024 — Fabric expansion and Lightning restructuring
65 changes.
This period focused on significantly expanding Lightning Fabric with comprehensive test coverage, new cluster environment plugins, and diverse training examples, while simultaneously restructuring the PyTorch Lightning codebase into a modular package architecture. The work introduced advanced distributed training capabilities, including tensor parallelism and FP8 support, alongside a unified plugin system for precision and checkpointing to improve maintainability and user experience.
Features
Add K-Fold Cross Validation example for Lightning Fabric
A new example demonstrating K-Fold cross validation using Lightning Fabric has been added to the examples directory. The \train\_fabric.py\ script shows how to use \scikit-learn\'s \KFold\ to split the MNIST dataset, train multiple CNN models in parallel across folds, and manage distributed training via Fabric's setup utilities. The accompanying \README.md\ provides instructions for running the example on CPU, GPU, or multiple GPUs using the \fabric run\ command.
_examples/fabric/kfold\cv · high confidence
Add MAML meta-learning example with Lightning Fabric and raw PyTorch implementations
A new example demonstrating Model Agnostic Meta Learning (MAML) on the Omniglot dataset has been added to the fabric examples. It provides two distributed training implementations: a raw PyTorch version using \torchrun\ and a Lightning Fabric version using \fabric run\, allowing users to compare boilerplate-heavy distributed code against the accelerated Fabric API for few-shot learning tasks.
_examples/fabric/meta\learning · high confidence
Add minimal Transformer language model example
A new example has been added to the Fabric examples directory that demonstrates a simple training loop for next-word prediction using a Transformer model on the WikiText2 dataset. The example includes a \train.py\ script and a \README.md\ with instructions for running the training on CPU, GPU, or multiple GPUs using the \fabric run\ command.
_examples/fabric/language\model · high confidence
Decoupled PPO Reinforcement Learning example for Fabric
A new decoupled Reinforcement Learning example has been added to the Fabric examples, implementing a Proximal Policy Optimization (PPO) agent using PyTorch and Gymnasium. This change replaces previous NumPy-based implementations with a pure PyTorch backend, introducing modular components for the agent (\PPOAgent\ and \PPOLightningAgent\), loss functions, and utilities. The example supports configurable activation functions (ReLU or Tanh), orthogonal initialization, and distributed training via a \share\_data\ flag, providing a modern, framework-native reference for training RL models with Lightning Fabric.
_examples/fabric/reinforcement\learning/rl · high confidence
Experimental API to validate model serving with FastAPI
A new experimental feature allows users to verify that their PyTorch Lightning models can be correctly served as REST APIs. By subclassing the new \ServableModule\ and implementing specific hooks (\configure\_payload\, \configure\_serialization\, \serve\_step\, and optionally \configure\_response\), users can attach the \ServableModuleValidator\ callback to their Trainer. This validator automatically spins up a local FastAPI server, sends a test request using the configured payload, and checks if the model's output matches the expected response, ensuring the model is ready for deployment before full integration.
src/lightning/pytorch/serve · high confidence
Fabric strategies module restructured and expanded
The \src/lightning/fabric/strategies\ package has been reorganized into a dedicated module with a public \\_\init\\_.py\ that explicitly exports all strategy classes (DDP, DeepSpeed, DataParallel, FSDP, ModelParallel, Parallel, SingleDevice, SingleXLA, XLA, XLAFSDP) and registers them in a global \STRATEGY\_REGISTRY\. This change introduces the \ModelParallelStrategy\ for 2D parallelism (FSDP2 + Tensor Parallelism) and the \XLAFSDPStrategy\ for TPU support, while also adding a dedicated \SingleDeviceXLAStrategy\ for single-XLA-device training. The \DeepSpeedStrategy\ now supports \exclude\_frozen\_parameters\ and learning rate scheduling, and the \FSDPStrategy\ exposes \sharding\_strategy\ as a string and supports \device\_mesh\ for hybrid sharding.
src/lightning/fabric/strategies · high confidence
Introduce Fabric loggers (CSV and TensorBoard)
The \src/lightning/fabric/loggers\ module now provides \CSVLogger\ and \TensorBoardLogger\ for Fabric experiments. \CSVLogger\ saves metrics to local CSV files with automatic versioning and optional experiment names, while \TensorBoardLogger\ writes to TensorBoard format, supporting both \tensorboard\ and \tensorboardX\ backends, subdirectories, and hyperparameter logging.
src/lightning/fabric/loggers · high confidence
Introduce \`lightning.pytorch.core\` package with core module definitions
The \lightning.pytorch.core\ package has been introduced, consolidating the foundational components of the PyTorch integration. This location now explicitly exports \LightningModule\ and \LightningDataModule\ via its \\_\init\\_.py\, while housing the implementation of these classes, their associated hooks (\hooks.py\), optimizer wrapping logic (\optimizer.py\), and checkpoint saving/loading utilities (\saving.py\). This structure centralizes the core abstractions that users interact with when defining models and data pipelines.
src/lightning/pytorch/core · high confidence
Introduce experimental collective communication plugin interface
A new experimental API for collective operations has been added to Lightning Fabric under \lightning.fabric.plugins.collectives\. This introduces a \Collective\ abstract base class defining standard distributed operations (broadcast, all\_reduce, gather, etc.), along with a \TorchCollective\ implementation that wraps \torch.distributed\ and a \SingleDeviceCollective\ stub for single-device scenarios. Users can now access these classes directly from the \lightning.fabric.plugins.collectives\ module to perform collective communications in a framework-agnostic way.
src/lightning/fabric/plugins/collectives · high confidence
Introduces dedicated launcher components for Fabric distributed training
The \src/lightning/fabric/strategies/launchers\ package now provides a structured set of classes to manage process creation for distributed training. This includes \\_SubprocessScriptLauncher\ for spawning child processes on a single node (supporting Hydra configuration for working directories), \\_MultiProcessingLauncher\ for multi-process scenarios, and \\_XLALauncher\ for launching parallel workers on XLA hardware. These components centralize the logic for environment variable setup, process monitoring, and inter-process communication within the Fabric strategy layer.
src/lightning/fabric/strategies/launchers · high confidence
Introduction of the unified 'lightning' meta-package
The \lightning\ package is now available as a single entry point that aggregates core components from PyTorch Lightning and Lightning Fabric. Users can install this unified package to access top-level imports such as \Trainer\, \LightningModule\, \LightningDataModule\, \Callback\, \Fabric\, and \seed\_everything\ directly from the \lightning\ namespace. The package enforces a minimum Python version of 3.10 and includes a CLI entry point for the Fabric tool.
src/lightning · high confidence
Lightning Fabric package structure and build configuration
The \src/lightning\fabric\ directory now contains the core package metadata and build configuration files, including \MANIFEST.in\ to define included assets (such as version info, changelog, and requirements), \\\about\\.py\ for package metadata, \\\setup\\.py\ for the setuptools configuration (specifying Python 3.10+ support and entry points), and \\\version\\_.py\ for version management. Additionally, a \py.typed\ marker file is added to enable PEP 561 type checking support for consumers of the library.
_src/lightning\fabric · high confidence
New 'Build Your Own Trainer' example for Lightning Fabric
A new example located at examples/fabric/build\_your\_own\_trainer demonstrates how to construct a fully customizable training loop using Lightning Fabric. The example provides a \MyCustomTrainer\ implementation in \trainer.py\ and a usage script \run.py\ that trains an MNIST model, illustrating key Fabric features such as hardware acceleration selection, distributed sampler handling, gradient accumulation, and checkpointing without relying on the full PyTorch Lightning Trainer.
_examples/fabric/build\_your\_own\trainer · high confidence
New DCGAN example demonstrating Lightning Fabric acceleration
Added a new example in the Fabric section that implements a Deep Convolutional Generative Adversarial Network (DCGAN) for generating realistic face images using the CelebA dataset. The example provides two implementations: a standard raw PyTorch version and a Lightning Fabric version, allowing users to compare the boilerplate reduction and hardware acceleration capabilities (including Apple Silicon support) offered by Fabric.
examples/fabric/dcgan · high confidence
New FP8 distributed transformer examples for Fabric and PyTorch Lightning
Added new example scripts and documentation for both Lightning Fabric and PyTorch Lightning that demonstrate training a Transformer model using FP8 low-precision quantization, FSDP2 (DTensor) for distributed sharding, and torch.compile for optimization. The Fabric example uses the ModelParallelStrategy with a configure\_model hook to apply quantization and compilation, while the PyTorch Lightning example implements the same logic within a LightningModule's configure\_model method, providing users with reference implementations for memory-efficient, high-throughput distributed training on supported hardware.
_examples/fabric/fp8\_distributed\_transformer, examples/pytorch/fp8\_distributed\transformer · high confidence
New Fabric MNIST example demonstrating scalable training
Added a new example in examples/fabric/image\_classifier that contrasts a vanilla PyTorch MNIST classifier with a Lightning Fabric version. The Fabric script (train\_fabric.py) shows how to use Fabric to enable GPU and multi-GPU training via the \fabric run\ CLI command, handle device placement automatically, and use \rank\_zero\_first\ for data downloads, providing a concrete reference for scaling PyTorch code with Fabric.
_examples/fabric/image\classifier · high confidence
New Fabric Reinforcement Learning example with coupled and decoupled PPO architectures
The \examples/fabric/reinforcement\_learning\ directory now contains a complete Proximal Policy Optimization (PPO) implementation accelerated by Lightning Fabric. This example demonstrates two distributed training patterns: a coupled architecture where the agent and environment run on the same processes, and a decoupled architecture that separates the 'Player' (environment interaction) from the 'Trainers' (model optimization) using Fabric collectives. The entry includes the \train\_fabric.py\ and \train\_fabric\_decoupled.py\ scripts, a raw PyTorch baseline (\train\_torch.py\) for comparison, and a \README.md\ detailing usage, hyperparameters, and visualization steps.
_examples/fabric/reinforcement\learning · high confidence
New Fabric example for 2D tensor and data parallelism
Added a new example in \examples/fabric/tensor\_parallel\ demonstrating how to combine tensor parallelism (TP) and data parallelism (DP) using the \ModelParallelStrategy\ with PyTorch Fabric. The example shows how to parallelize a Llama 3 7B model across multiple GPUs by applying tensor-parallel strategies to model layers and using Fully Sharded Data Parallel (FSDP) for data parallelism, including configuration of mixed precision, activation checkpointing, and distributed checkpointing.
_examples/fabric/tensor\parallel · high confidence
New PyTorch Lightning domain templates for computer vision, GANs, and reinforcement learning
Added new runnable example scripts in \examples/pytorch/domain\_templates\ that demonstrate specific deep learning patterns using PyTorch Lightning. These include a computer vision fine-tuning template using \BaseFinetuning\ and \LightningDataModule\, a Generative Adversarial Network (GAN) template with manual optimization, an ImageNet training template using \LightningCLI\, and reinforcement learning templates for Deep Q-Networks (DQN) and Proximal Policy Optimization (PPO) using OpenAI Gym.
_examples/pytorch/domain\templates · high confidence
New PyTorch basic examples and bug-report template
The \examples/pytorch/basics\ directory now includes runnable starter scripts for common patterns: an MNIST autoencoder (\autoencoder.py\), a backbone image classifier (\backbone\_image\_classifier.py\), a Transformer-based next-word predictor on WikiText2 (\transformer.py\), and a PyTorch Profiler demo (\profiler\_example.py\). Each example is documented in \README.md\ with run instructions for CPU, GPU, and DDP. Additionally, a bug-report template is provided in \examples/pytorch/bug\_report\ (notebook and Python script) to help users reproduce issues with a minimal Lightning model.
examples/pytorch/basics · high confidence
New PyTorch tensor-parallel example with 2D parallelism support
Added a new example in \examples/pytorch/tensor\_parallel\ demonstrating how to apply tensor parallelism to a Llama 3 7B model using PyTorch Lightning's \ModelParallelStrategy\. The example shows how to combine tensor parallelism with FSDP for 2D parallelism, requiring PyTorch 2.3+ and at least 4 GPUs with 24 GB memory each. It includes a complete training script (\train.py\) that configures the \ModelParallelStrategy\ with \data\_parallel\_size=2\ and \tensor\_parallel\_size=2\, along with helper modules for data handling (\data.py\) and parallelization logic (\parallelism.py\).
_examples/pytorch/tensor\parallel · high confidence
New build assistant and legacy checkpoint retrieval tooling
The repository now includes a new Python build assistant (\.actions/assistant.py\) that manages requirements for PyTorch, Fabric, and Data packages, featuring logic to adjust version constraints (freezing, unfreezing, or strict enforcement) and to generate package descriptions from READMEs. Additionally, a new shell script (\.actions/pull\_legacy\_checkpoints.sh\) has been added to download and extract legacy model checkpoints from S3 into the test directory, supporting tests that depend on historical model artifacts.
.actions · high confidence
New callbacks module exposes EMA weight averaging and device stats filtering
The \src/lightning/pytorch/callbacks\ package now includes the \EMAWeightAveraging\ and \WeightAveraging\ callbacks for generic weight averaging, and the \DeviceStatsMonitor\ callback now supports a \filter\keys\ argument to log only specified device statistics. These additions are exported via the \\\init\\_.py\ module, making them available for import from \lightning.pytorch.callbacks\.
src/lightning/pytorch/callbacks · high confidence
New cluster environment plugins for Fabric
The \lightning.fabric.plugins.environments\ module now provides a set of cluster environment plugins that allow Fabric to automatically detect and configure distributed training across various schedulers and runtimes. This includes support for SLURM, LSF, MPI, TorchElastic, Kubeflow, and Google Cloud TPUs (via XLA). These environments handle the discovery of main addresses, ports, ranks, and world sizes, enabling seamless distributed execution on HPC clusters and cloud platforms without manual configuration.
src/lightning/fabric/plugins/environments · high confidence
New demo modules for LSTM and Transformer models
The \lightning.pytorch.demos\ package now includes \LightningLSTM\ and \LightningTransformer\ classes, providing ready-to-use examples for sequence modeling tasks. The LSTM demo implements a language model using PyTorch's LSTM module with a custom \SequenceSampler\, while the Transformer demo provides a \LightningTransformer\ wrapper around a standard Transformer architecture with positional encoding, along with a \WikiText2\ dataset class for downloading and tokenizing text data. These additions are exported via the package's \\_\init\\_.py\ alongside existing boring classes and MNIST utilities.
src/lightning/pytorch/demos · high confidence
New distributed training overrides and utilities in PyTorch Lightning
The \src/lightning/pytorch/overrides\ module has been introduced to centralize PyTorch-specific overrides for distributed training. This includes a \DistributedDataParallel\ wrapper that ensures \prepare\_for\_backward\ is called correctly during manual optimization and handles unused parameter detection, a helper to register custom DDP communication hooks (such as FP16 compression or PowerSGD), and a utility to sync module states across processes. Additionally, an \UnrepeatedDistributedSampler\ is provided to allow deterministic, non-repeating data sampling during prediction phases, addressing issues where processes might run different numbers of batches.
src/lightning/pytorch/overrides · high confidence
New examples and core loop infrastructure for tensor parallelism and production serving
This release introduces new example scripts demonstrating tensor-parallel model implementation for both Fabric and PyTorch Lightning, alongside a production-serving example using \ServableModule\. Internally, the core training and evaluation loops have been refactored to support these patterns, including changes to how data is fetched and processed, and updates to the multiprocessing launcher to better handle global state snapshots and CUDA context checks. The Rich progress bar has also been enhanced with custom columns and infinite task support to improve visibility during training.
python · high confidence
New utilities module for PyTorch Lightning
The \src/lightning/pytorch/utilities\ package has been introduced to centralize general-purpose helper functions and classes. This new module exposes utilities such as \CombinedLoader\ for managing multiple data loaders, \AttributeDict\ for structured argument handling, and \is\_picklable\ for serialization checks. It also includes tools for gradient inspection (\grad\_norm\), shared parameter management, and environment-variable-based configuration defaults. By consolidating these components, the library provides a single, organized location for common utilities used across the Trainer, Fabric, and callbacks.
src/lightning/pytorch/utilities · high confidence
Project initialization with Apache 2.0 license and development tooling
The repository has been initialized with a new project structure, including a Makefile for managing setup, tests, and documentation builds, and a .pre-commit-config.yaml enforcing code quality via Ruff, mdformat, and other linters. The project license has been changed from MIT to Apache 2.0, and a CITATION.cff file has been added to facilitate academic citation. Additionally, configuration files for Codecov, Read the Docs, and Git (including .gitignore and .gitmodules for the tutorials submodule) are now present to support continuous integration and documentation generation.
(repo-wide) · high confidence
Architecture
Trainer module restructured into a dedicated package
The \Trainer\ core logic has been reorganized from a single monolithic file into a structured package under \src/lightning/pytorch/trainer\. This change introduces dedicated modules for specific responsibilities: \call.py\ handles hook execution and interrupt management, \setup.py\ manages initialization flags and profiler setup, and \configuration\validator.py\ enforces loop configuration rules. The main \Trainer\ class is now imported from this new package structure, and the \\\init\\_.py\ exposes \Trainer\ and \seed\_everything\ for public use.
src/lightning/pytorch/trainer · high confidence
Behavioural changes
Accelerator implementations moved to lightning.pytorch.accelerators
The PyTorch Lightning accelerator implementations (CPU, CUDA, MPS, and XLA) have been relocated from the legacy \\pytorch\_lightning.accelerators\\ namespace to \\lightning.pytorch.accelerators\\. This change consolidates the accelerator logic within the new \\lightning\\ package structure, ensuring that users importing from the new namespace will access the correct, updated implementations for device management and statistics.
src/lightning/pytorch/accelerators · high confidence
Automatic legacy checkpoint migration and unpickling support
Lightning now automatically upgrades legacy checkpoints when loading them, applying version-specific migrations (e.g., for loop structures, progress tracking, and callback keys) to match the current format. The system also handles unpickling of older checkpoints that contain removed classes or modules (such as \\_FaultTolerantMode\ or \pytorch\_lightning\ imports) by temporarily patching the environment, ensuring seamless resumption of training from older saved states.
src/lightning/pytorch/utilities/migration · high confidence
Backward compatibility shims for renamed and removed components
The library now includes a compatibility layer in the \\_graveyard\ module to support legacy imports and class names. Deprecated precision plugins (e.g., \PrecisionPlugin\, \HalfPrecisionPlugin\) are aliased to their new counterparts (e.g., \Precision\, \HalfPrecision\) with deprecation warnings. TPU-specific classes (\TPUAccelerator\, \SingleTPUStrategy\) are mapped to their XLA equivalents (\XLAAccelerator\, \SingleDeviceXLAStrategy\). Removed Habana (HPU) integrations are replaced with stub classes that raise \NotImplementedError\ to clearly indicate removal. Additionally, a patch ensures older versions of \torchmetrics\ (\<0.8.0) correctly resolve internal imports to the unified \lightning.pytorch\ package.
_src/lightning/pytorch/\graveyard · high confidence
Enhanced ModelSummary with training mode, FLOPs, and Deepspeed support
The model summary output now includes columns for the module's training mode (train/eval) and estimated FLOPs, providing deeper insight into model behavior and computational cost. It also correctly handles non-layer parameters by listing them separately and supports PyTorch DTensors. For users of the DeepSpeed strategy, a specialized summary is available that displays parameters per device and fixes column alignment issues.
_src/lightning/pytorch/utilities/model\summary · high confidence
Fabric 2.6.5 release with remote checkpoint support and bug fixes
This release introduces support for remote storage (fsspec URLs) when saving and loading distributed checkpoints with FSDPStrategy, ModelParallelStrategy, and the checkpoint consolidation CLI. It also adds \using\_sparse\_model\ and \sparse\_cuda\_acceleration\_factor\ parameters to the \Throughput\ utility to correctly default to dense peak FLOPs and allow explicit sparse peak opt-in. Several bugs are fixed: FSDP mixed precision now correctly initializes parameters in fp32 instead of half precision, the \BitsandbytesPrecision\ plugin type checking is corrected for bitsandbytes 0.48/0.49, DeepSpeed checkpoint path validation now accepts remote filesystem URIs (S3, GCS, HDFS), and a RuntimeError under PyTorch 2.5+ when FSDP's \root\_device\ is CPU is resolved by passing an explicit \torch.device("cpu")\. Additional fixes address an AccumulateGrad stream mismatch warning with DDP, incorrect CUDA device initialization in \CUDAAccelerator.setup\_device\, and \PermissionError\ swallowing during atomic checkpoint saves. Dead code paths for PyTorch below 2.6 have been removed.
src/lightning/fabric · high confidence
Fabric accelerator implementations and registry are restructured
The \src/lightning/fabric/accelerators\ module now provides the concrete accelerator classes (CPU, CUDA, MPS, XLA) and the \\_AcceleratorRegistry\ that Fabric uses to select hardware. This change introduces a registry-based lookup for accelerators, ensuring that the correct hardware implementation is instantiated based on the user's configuration. It also includes specific fixes for CUDA device initialization (ensuring \set\_device\ is called before precision checks) and XLA runtime detection (supporting \torch-xla\>=2.5\ PJRT checks).
src/lightning/fabric/accelerators · high confidence
Fabric precision plugin API restructured with new context managers and plugin classes
The precision plugin system in Fabric has been reorganized to support a more granular control over model initialization and data conversion. The base \Precision\ class now exposes \tensor\_init\_context\ and \module\_init\context\ hooks, allowing plugins to control tensor creation and module instantiation separately. New concrete plugins have been introduced: \BitsandbytesPrecision\ for quantization (supporting modes like \nf4\, \fp4\, and \int8\), \TransformerEnginePrecision\ for FP8 training via NVIDIA's Transformer Engine, and \XLAPrecision\ for XLA devices. Existing plugins like \MixedPrecision\, \FSDPPrecision\, \DeepSpeedPrecision\, \HalfPrecision\, and \DoublePrecision\ have been updated to align with the new API, including changes to how they handle input/output conversion and optimizer steps (e.g., \FSDPPrecision\ now uses \ShardedGradScaler\). The \\\init\\_.py\ module now explicitly exports all these precision plugins.
src/lightning/fabric/plugins/precision · high confidence
Fabric utilities module reorganization and new checkpointing features
The \src/lightning/fabric/utilities\ package has been reorganized into a dedicated module, exposing a consolidated set of utilities including \AttributeDict\, \Throughput\ monitoring, and data-handling helpers like \move\_data\_to\_device\. This change introduces atomic checkpoint saving via \fsspec\ to prevent corruption during interruptions, supports remote storage URLs for FSDP checkpoints, and adds a CLI tool to consolidate sharded FSDP checkpoints into a single file. Additionally, the \\_DeviceDtypeModuleMixin\ now respects PyTorch 2.8's default device context, and the \load\ module implements lazy checkpoint loading to reduce memory overhead when loading large models.
src/lightning/fabric/utilities · high confidence
Hyperparameter handling logic moved to a dedicated mixin
The \save\_hyperparameters\ functionality and related hyperparameter management (including \hparams\ and \hparams\_initial\ properties) have been extracted from the core model classes into a new \HyperparametersMixin\ in \src/lightning/pytorch/core/mixins\. This change restructures how hyperparameters are stored and accessed, making the logic available for reuse and potentially altering inheritance patterns for classes that previously relied on the previous implementation location.
src/lightning/pytorch/core/mixins · medium confidence
Introduce Lightning Data as a wrapper for the litdata package
The \src/lightning/data\ module now acts as a compatibility layer that re-exports the core functionality of the \litdata\ package. Users can continue to import \StreamingDataset\, \optimize\, \map\, and other data processing tools from \lightning.data\, but the implementation is now sourced from the external \litdata\ library. This change requires the \litdata\ package to be installed, as the module explicitly checks for its presence and raises an error if it is missing, effectively decoupling the data loading logic from the main Lightning repository.
src/lightning/data · high confidence
Introduction of structured logging validation and integration suggestions in Trainer logs
The logger connector now includes an \fx\_validator\ that strictly defines which Lightning hooks allow \self.log()\ calls and enforces the \on\_step\/\on\_epoch\ logging levels, raising clear errors for invalid usage. Additionally, when \Trainer(suggest\_integrations=True)\ (the default), the connector checks if a \LitLogger\ is present and, if not, suggests installing the \litlogger\ package for seamless cloud logging.
_src/lightning/pytorch/trainer/connectors/logger\connector · high confidence
New CheckpointIO plugin interface with weights\_only support
The \src/lightning/fabric/plugins/io\ module now exposes a new \CheckpointIO\ interface along with \TorchCheckpointIO\ and \XLACheckpointIO\ implementations. This change introduces a \weights\_only\ parameter to checkpoint loading, allowing users to restrict loading to safe state dicts for improved security when handling untrusted sources. The \XLACheckpointIO\ implementation specifically handles TPU-specific requirements, including OmegaConf compatibility workarounds.
src/lightning/fabric/plugins/io · high confidence
New plugin architecture for layer synchronization and precision management
The \src/lightning/pytorch/plugins\ module has been restructured to introduce a dedicated plugin system for managing model layer synchronization and precision. This change adds a new \LayerSync\ abstract base class and a concrete \TorchSyncBatchNorm\ implementation, allowing users to wrap batch normalization layers with synchronization logic for multi-GPU and multi-node training scenarios. Additionally, the module now explicitly exports a comprehensive set of precision plugins—including \BitsandbytesPrecision\, \TransformerEnginePrecision\, \FSDPPrecision\, and \XLAPrecision\—alongside standard options like \MixedPrecision\ and \HalfPrecision\, centralizing how these configurations are accessed and applied within the PyTorch Lightning Trainer.
src/lightning/pytorch/plugins · high confidence
New precision plugin architecture in PyTorch Lightning
The precision handling in PyTorch Lightning has been reorganized into a new plugin layout under \src/lightning/pytorch/plugins/precision\. This change introduces a unified \Precision\ base class and dedicated plugins for various precision strategies, including \MixedPrecision\ (AMP), \HalfPrecision\, \DoublePrecision\, \FSDPPrecision\, \DeepSpeedPrecision\, \XLAPrecision\, \BitsandbytesPrecision\, and \TransformerEnginePrecision\. Users can now configure these specific precision modes via the \Trainer\'s \precision\ argument, which routes to the appropriate plugin to handle dtype conversion, gradient scaling, and optimizer stepping logic for each strategy.
src/lightning/pytorch/plugins/precision · high confidence
Profiler module relocated to lightning.pytorch.profilers
The profiler implementations (AdvancedProfiler, PyTorchProfiler, SimpleProfiler, XLAProfiler, and PassThroughProfiler) have been moved from the legacy pytorch\_lightning namespace to the new lightning.pytorch.profilers package. Users should update their imports to use the new location, for example: from lightning.pytorch.profilers import PyTorchProfiler.
src/lightning/pytorch/profilers · high confidence
Progress bar logic reorganized into a dedicated module
The progress bar implementation has been restructured into a new \src/lightning/pytorch/callbacks/progress\ package, separating the base \ProgressBar\ class, the default \TQDMProgressBar\, and the \RichProgressBar\ into distinct files. This change consolidates the progress bar code, making it easier to maintain and extend, while preserving the existing functionality for tracking training, validation, and testing progress.
src/lightning/pytorch/callbacks/progress · high confidence
PyTorch Lightning loggers module restructured and LitLogger added
The \src/lightning/pytorch/loggers\ package has been reorganized to expose a unified \\_\init\\_.py\ that imports and exports all available loggers, including the newly added \LitLogger\ for remote experiment tracking on Lightning AI. The \NeptuneLogger\ has been removed and replaced with a deprecated stub that raises a \RuntimeError\ directing users to migrate to \LitLogger\. The module now includes updated implementations for \CometLogger\, \CSVLogger\, \MLFlowLogger\, \TensorBoardLogger\, and \WandbLogger\, along with shared utilities for checkpoint scanning and hyperparameter logging.
src/lightning/pytorch/loggers · high confidence
PyTorch Lightning package structure and metadata reorganization
The \src/pytorch\lightning\ directory has been restructured to support the new \src/\ layout and unified packaging strategy. A new \MANIFEST.in\ ensures that version info, changelogs, READMEs, and requirement files for both PyTorch Lightning and Lightning Fabric are correctly included in the distribution. The package now includes a \py.typed\ marker file to enable PEP 561 type checking support. Metadata files (\\\about\\.py\, \\\version\\.py\, \\\setup\\_.py\) have been updated to reflect the new source root, load version information from \version.info\, and configure the package to include both \pytorch\_lightning\ and \lightning\_fabric\ source code. The setup script now explicitly requires Python 3.10+ and defines classifiers for Python 3.10 through 3.13.
_src/pytorch\lightning · high confidence
Reorganized cluster environment plugins under the new lightning package structure
The cluster environment plugins have been reorganized to align with the new package structure. The \src/lightning/pytorch/plugins/environments/\_\init\\_.py\ file now serves as a central re-export point, importing \ClusterEnvironment\ from \lightning.fabric.plugins\ and exposing specific environment classes (\KubeflowEnvironment\, \LightningEnvironment\, \LSFEnvironment\, \MPIEnvironment\, \SLURMEnvironment\, \TorchElasticEnvironment\, \XLAEnvironment\) from \lightning.fabric.plugins.environments\. This change consolidates the availability of these environment configurations for users migrating to the new \lightning\ namespace.
src/lightning/pytorch/plugins/environments · high confidence
Reorganized distributed training strategies into a new package structure
The distributed training strategies (DDP, DeepSpeed, FSDP, ModelParallel, Parallel, SingleDevice, and XLA) have been reorganized into the new \src/lightning/pytorch/strategies\ package. This change introduces a centralized \\_\init\\_.py\ that registers these strategies with the global strategy registry, making them available for import from \lightning.pytorch.strategies\. The individual strategy files have been updated to import from the new location and register themselves, ensuring that users can continue to use these strategies via the standard Lightning entry points without code changes.
src/lightning/pytorch/strategies · high confidence
Restructure PyTorch Lightning launchers into a dedicated module
The launcher implementations for multiprocessing, subprocess scripts, and XLA have been moved into a new \src/lightning/pytorch/strategies/launchers\ package. This change introduces a base \\_Launcher\ class that extends the Fabric launcher with a \kill\ method for process termination, and exposes the specific launchers (\\_MultiProcessingLauncher\, \\_SubprocessScriptLauncher\, \\XLALauncher\) via a new \\\init\\_.py\. The \\_SubprocessScriptLauncher\ now validates cluster environment settings before spawning child processes, and the \\_XLALauncher\ enforces a restriction against calling \trainer.fit()\ twice on the same instance when using spawn-based strategies.
src/lightning/pytorch/strategies/launchers · high confidence
Restructured and refactored internal training loops and data fetching
The internal loop architecture in \src/lightning/pytorch/loops\ has been reorganized: the \\_\init\\_.py\ now explicitly exports core loop classes (\\_FitLoop\, \\_EvaluationLoop\, \\_PredictionLoop\, \\_TrainingEpochLoop\, and optimization loops), and a new \fetchers.py\ module centralizes data fetching logic with dedicated classes like \\_PrefetchDataFetcher\ and \\_DataLoaderIterDataFetcher\. This refactoring moves data-fetching responsibilities out of the loops themselves, introduces a unified \CombinedLoader\ assumption for fetcher inputs, and adds utilities for selecting the appropriate fetcher based on hook signatures (e.g., \dataloader\_iter\ support). Users benefit from a more modular internal structure that supports flexible data iteration patterns while maintaining backward compatibility with existing training, validation, and prediction workflows.
src/lightning/pytorch/loops · high confidence
Restructured examples with Fabric and Trainer scripts
The examples directory has been reorganized to clearly separate Lightning Fabric and Lightning Trainer use cases. A new README provides navigation to Fabric examples (MNIST, DCGAN) and Trainer examples (Image Classifier, Autoencoder). Additionally, shell scripts (run\_fabric\_examples.sh, run\_pl\_examples.sh) have been added to simplify running these examples with default parameters.
examples · high confidence
Trainer connectors restructured with new default callbacks and signal handling
The \src/lightning/pytorch/trainer/connectors\ module has been reorganized into distinct connector classes (\\_AcceleratorConnector\, \\_CallbackConnector\, \\_CheckpointConnector\, \\_SignalConnector\) to manage specific Trainer configurations. By default, the Trainer now uses \RichProgressBar\ and \RichModelSummary\ when the \rich\ library is available, falling back to \TQDMProgressBar\ and \ModelSummary\ otherwise. A new \LitLogger\ integration is supported, with a tip displayed to users if \litmodels\ is installed but not actively used for model logging. Additionally, signal handling for SLURM auto-requeueing and SIGTERM has been centralized in \\_SignalConnector\ to improve graceful shutdown and job requeuing reliability.
src/lightning/pytorch/trainer/connectors · high confidence
Tuner module relocated to lightning.pytorch.tuner with security and stability fixes
The tuner functionality (LearningRateFinder and BatchSizeFinder) has been moved into the \lightning.pytorch.tuner\ package, exposing the \Tuner\ class for model tuning. This change includes a security fix for the LearningRateFinder to support the \weights\_only\ parameter when restoring checkpoints, and addresses stability issues such as resetting the epoch loop restarting flag and fixing gradient calculation for exponential mode.
src/lightning/pytorch/tuner · high confidence
Fixes
AsyncCheckpointIO snapshots tensors to prevent race conditions
The AsyncCheckpointIO plugin now clones checkpoint tensors before submitting them to the background thread. This snapshotting prevents race conditions where model parameters might be mutated by the training loop while the checkpoint is being saved asynchronously, ensuring data integrity during concurrent save operations.
src/lightning/pytorch/plugins/io · high confidence
Test coverage
Added comprehensive test coverage for Rich and TQDM progress bars; Added comprehensive test suite for PyTorch precision plugins; Added comprehensive tests for PyTorch trainer logging behavior; Added internal testing utility for conditional test execution; Added parity and benchmarking tests for PyTorch integration; Added parity tests for Lightning Fabric against raw PyTorch; Added test coverage for Fabric accelerator implementations; Added test coverage for PyTorch Tuner components; Added test coverage for PyTorch plugins; Added test helper utilities for Fabric; Added test infrastructure for legacy checkpoint compatibility; Added test script for LightningCLI integration; Added test suite for PyTorch Trainer configuration and validation; Added test suite for PyTorch core components; Added test suite for Trainer connectors; Added tests for Fabric CSV and TensorBoard loggers; Added tests for Fabric cluster environment plugins; Added tests for Fabric collectives plugins; Added tests for Fabric multiprocessing and subprocess launchers; Added tests for LSTM and Transformer demo utilities; Added tests for PyTorch Lightning utilities; Added tests for PyTorch accelerator implementations; Added tests for PyTorch checkpoint migration utilities; Added tests for PyTorch distributed overrides; Added tests for PyTorch multiprocessing and subprocess script launchers; Added tests for RunIf utility; Added tests for ServableModuleValidator callback; Added tests for Trainer flag configurations; Added tests for deprecated and removed PyTorch Lightning components; Added tests for deprecated precision plugins and API behaviors; Added tests for trainer optimization and dynamic arguments; Fabric test suite reorganization and isolation improvements; New PyTorch test suite structure and fixtures; New test helper utilities and fixtures for PyTorch Lightning; New test suite for Fabric distributed strategies; New test suite for PyTorch Lightning strategies; New test suite for PyTorch loggers; New testing utilities for conditional test skipping.
Dependencies
Centralized and structured requirements management
The project has reorganized its dependency management by introducing a structured \requirements/\ directory. This change consolidates dependencies into specific files for different environments: \ci.txt\ for continuous integration, \docs.txt\ for documentation building (including Sphinx and related tools), \doctests.txt\ for testing, and \typing.txt\ for type checking (pinning MyPy and Torch). A new \collect\_env\_details.py\ script has been added to help diagnose system issues by reporting OS, CUDA, and installed package versions. Additionally, a \README.md\ explains the new structure and the use of upper bounds for CI stability.
requirements · high confidence
Introduce pyproject.toml and structured requirements files
The project now uses a centralized pyproject.toml for build configuration and tooling settings (including Ruff, mypy, and pytest), replacing ad-hoc configurations. Dependency management is restructured into a modular system: the root requirements.txt now aggregates base requirements from ./requirements/fabric/base.txt and ./requirements/pytorch/base.txt, while specific examples (such as FP8 distributed transformers and reinforcement learning) have their own dedicated requirements.txt files to isolate their dependencies.
(dependencies) · high confidence
PyTorch requirements updated to support PyTorch 2.6 and newer dependency versions
The PyTorch package requirements have been updated to support PyTorch 2.6 (minimum) and up to 2.14.0, along with corresponding updates to torchvision (\>=0.21.0), torchmetrics (\>0.7.0), and other core dependencies like fsspec, tqdm, and typing-extensions. The DeepSpeed strategy requirement has been bumped to \>=0.16.0, and test dependencies have been updated to use pytest 9.0.2 and coverage 7.13.5. These changes ensure compatibility with the latest PyTorch releases and their associated ecosystem libraries.
requirements/pytorch · high confidence
Housekeeping
Lightning PyTorch 2.0.4 and 2.0.5 changelog update; Update development version to 2.7.0.dev.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 71.
Lenses
- Code Health 86
- Architecture 98
- Maturity 64
- Readiness 72
- Security 82
- Domain Modelling 77
Changes since last survey
- 300 commits — 229 feature/other, 71 fixes
By area
- src/lightning — 101 commits
- .github/workflows — 40 commits
- requirements/pytorch — 31 commits
- docs/source-pytorch — 22 commits
- (root) — 21 commits
- requirements/fabric — 20 commits
- tests/tests_pytorch — 16 commits
- requirements/docs.txt — 12 commits
- requirements/ci.txt — 6 commits
- .lightning/workflows — 5 commits
- requirements/typing.txt — 5 commits
- requirements/doctests.txt — 4 commits
- tests/legacy — 4 commits
- .github/checkgroup.yml — 3 commits
- .github/markdown-links-config.json — 2 commits
- docs/source-fabric — 2 commits
- .actions/assistant.py — 1 commit
- .azure/README.md — 1 commit
- .azure/gpu-benchmarks.yml — 1 commit
- .github/CODEOWNERS — 1 commit
Notable commits
- fix: Bugfix for BackboneFinetuning + LearningRateFinder (#21224)
- fix: CI: fix doctest failure from PyTorch LeafSpec FutureWarning (#21502)
- fix: CUDAAccelerator.setup_device: fix unrelated device init by matmul precision check (#21726)
- fix: Fix DeepSpeed checkpoint path validation for remote filesystem URIs (#21636)
- fix: Fix EADDRINUSE errors in distributed tests with port manager and retry logic (#21309)
- fix: Fix FLOPs inconsistency and add using_sparse_model flag (#21743)
- fix: Fix FSDP CPU offloading when shards are already on CPU (#21805)
- fix: Fix FSDP mixed precision semantics and add user warning (#21361)
- fix: Fix FSDPPrecision rejecting a scaler for 16-mixed precision (#21831)
- fix: Fix LearningRateMonitor TypeError on MPS with float64 lr values (#21545)
- fix: Fix LightningCLI loading of hyperparameters from ckpt_path failing for subclass model mode (#21246)
- fix: Fix LightningDataModule.load_from_checkpoint to restore subclass and hyperparameters (#21478)
- fix: Fix ModelCheckpoint file_exists OOM in DDP (#21380)
- fix: Fix ModelParallelStrategy fails with non-distributed checkpoint. (#21384)
- fix: Fix PT011 (#21197)
- fix: Fix PermissionError swallowing in _atomic_save (#21799)
- fix: Fix RichModelSummary model size display (#21467)
- fix: Fix ModelCheckpoint with manual optimization and every_n_train_steps (#21239)
- fix: Fix StochasticWeightAveraging with infinite epochs (#21396)
- fix: Fix _generate_seed_sequence sampling (#21399)
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Lightning-AI/pytorch-lightning was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit fab6fb2dcf25c0920b8c2a6174258ffa66ea3337 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.