Stability-AI/generative-models
61.5
Adequate · 18 September 2026
15.3k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a Python-based library for training and inferring generative AI models, specifically focusing on Stable Diffusion XL for images and Stable Video Diffusion for video. It provides core infrastructure for diffusion architectures, including novel-view video generation (SV4D), autoencoding, and perceptual loss modules. The project offers high-level inference APIs and local demo interfaces to facilitate image and video synthesis, alongside utilities for memory-efficient processing and model licensing management.
Features
Add SV4D and SV4D++ novel-view video generation scripts
New sampling scripts (\simple\_video\_sample\_4d.py\ and \simple\_video\_sample\_4d2.py\) are added to the \scripts/sampling\ directory, enabling the generation of novel-view videos from input video or image sequences. The SV4D script supports generating 8 novel views from a single input view using SV3D for initial view synthesis, while the SV4D++ script supports two configurations (sv4d2 and sv4d2\_8views) for generating 4 or 8 novel views respectively, with configurable camera elevations and azimuths. These scripts handle background removal, video preprocessing, and VRAM-optimized encoding/decoding via \encoding\_t\ and \decoding\_t\ parameters.
scripts/sampling · high confidence
Introduces Gradio demos for Stable Video Diffusion and SV4D, and updates Streamlit sampling scripts
Users can now run local Gradio interfaces for Stable Video Diffusion (SVD) and the new SV4D model, which support image-to-video and video-to-video generation with automatic checkpoint downloading. The existing Streamlit-based sampling scripts have been updated to support SDXL Turbo, SDXL 1.0, and SV3D, while removing deprecated SD 2.1 models. Additionally, the watermark detection logic has been fixed to correctly handle BGR channel order, and the NSFW/watermark filtering now supports configurable device placement to improve performance.
scripts/demo · high confidence
New autoencoding loss and regularization modules for video generation
The \sgm/modules/autoencoding\ package now includes new components to support Stable Video Diffusion training and inference. This adds a \GeneralLPIPSWithDiscriminator\ loss module that integrates a PatchGAN discriminator (NLayerDiscriminator) and LPIPS perceptual loss, including visualization capabilities for discriminator logits. It also introduces a \LatentLPIPS\ loss for computing perceptual errors in latent space by decoding latents to images. Additionally, the package now provides a suite of vector quantization regularizers, including \GumbelQuantizer\ and \VectorQuantizer\, alongside base regularizer utilities. These modules are wired into the autoencoding pipeline to improve visual fidelity and latent structure during video model training.
sgm/modules/autoencoding · high confidence
New structured inference API for Stable Diffusion XL
Users can now generate images using a high-level \SamplingPipeline\ class that simplifies configuration for Stable Diffusion XL (SDXL) v0.9 and v1.0 base and refiner models. This new API in \sgm/inference/api.py\ replaces manual script usage by providing explicit enums for model architectures, samplers (such as DPMPP2M and EulerEDM), and discretization methods, along with a \SamplingParams\ dataclass for standardizing generation settings. The accompanying \helpers.py\ module supports this by providing core sampling logic, image-to-image strength handling, and automatic watermark embedding for generated outputs.
sgm/inference · high confidence
Support for SV4D video generation with new attention modules and conditioning
This update introduces core components for SV4D (Stable Video 4D) generation. New modules \spacetime\_attention.py\ and \video\_attention.py\ add temporal mixing capabilities to transformer blocks, enabling the model to process video sequences. The \GeneralConditioner\ in \encoders/modules.py\ is extended to handle new embedding keys (\cond\_view\, \cond\_motion\) and includes a new \VideoPredictionEmbedderWithEncoder\ to process video inputs with sigma conditioning. Additionally, \attention.py\ is updated to fix a memory issue in xformers by batching large attention computations and replaces print statements with proper logging.
sgm/modules · high confidence
Architecture
Repository structure and configuration updates
The repository has been restructured to improve maintainability and clarity. A CODEOWNERS file is introduced to direct pull request reviews to the infrastructure team. The project license is explicitly defined in a new LICENSE-CODE file. The legacy setup.py installation script has been removed, and the .gitignore file has been updated to exclude additional build artifacts and environment directories. Additionally, pytest configuration has been added to support marking inference tests, and minor code style adjustments were made to main.py.
(repo-wide) · high confidence
Behavioural changes
Added license agreements for SDXL-Turbo, SDXL 1.0, SV3D, and SVD models
The repository now includes specific legal terms for four Stability AI models in the model\_licenses directory. SDXL 1.0 is governed by the CreativeML Open RAIL++-M License, while SDXL-Turbo, SV3D, and SVD are covered by Stability AI Non-Commercial Community License Agreements. These licenses explicitly restrict usage to non-commercial research and prohibit production use or hosting via APIs without separate commercial arrangements. The previous default LICENSE file has been renamed to LICENSE-SDXL0.9 to reflect the older model version.
_model\licenses · high confidence
Improved memory efficiency and configurable checkpointing in autoencoder and diffusion models
The autoencoder models now use Python's logging module instead of print statements for status updates, and replace the direct \init\_from\_ckpt\ method with a pluggable \apply\_ckpt\ engine, allowing for more flexible checkpoint loading strategies. Additionally, the diffusion engine introduces an \en\_and\_decode\_n\_samples\_a\_time\ parameter that enables chunked encoding and decoding of latent tensors, significantly reducing peak GPU memory usage during inference and training by processing samples in batches rather than all at once.
sgm/models · high confidence
SGM package initialization and utility updates
The sgm package now exposes a version string (0.1.0) and a new get\_configs\_path utility to locate configuration directories in both development and installed environments. The load\_model\_from\_config function has been simplified by removing an intermediate variable assignment during state dict loading. Additionally, a new get\_nested\_attribute helper has been added to support recursive attribute access on objects.
sgm · high confidence
Support for video diffusion architectures and advanced sampling controls
This update introduces core modules for video generation models, including \VideoUNet\ and \VideoResBlock\ in \video\_model.py\ which add temporal processing via time stacks and mixers, and \BasicTransformerTimeMixBlock\ integration in \openaimodel.py\. It refactors the denoiser pipeline by decoupling loss weighting and scaling logic into new \loss\_weighting.py\ and \denoiser\_scaling.py\ modules, allowing for more flexible noise scheduling. The change also adds new guidance strategies in \guiders.py\ (such as \LinearPredictionGuider\ and \TrianglePredictionGuider\) for dynamic classifier-free guidance scaling over time, and fixes the \EDMDiscretization\ sigma minimum to correct sampling noise schedules.
sgm/modules/diffusionmodules · high confidence
Test coverage
Added attention benchmarking tests; Added tests for inference helpers and pipeline execution.
Dependencies
Migrate to Hatch-based Python packaging and consolidate requirements
The project now uses a \pyproject.toml\ file with the Hatch build system to manage the package metadata, versioning, and distribution, replacing the previous manual requirements files. The legacy \requirements\_pt13.txt\ (PyTorch 1.13) and \requirements\_pt2.txt\ (PyTorch 2.0) files have been removed. The new configuration explicitly sets the minimum Python version to 3.8, defines the build dependencies (hatchling), and configures Hatch to include the \configs\ directory in both source distributions and wheels. CI testing is now driven by the Hatch environment, which installs specific PyTorch 2.0.1 CUDA 11.8 packages and runs inference tests.
(dependencies) · high confidence
Housekeeping
Standardized import ordering in data and distribution modules
The import statements in the CIFAR-10 and MNIST data wrappers, as well as the distributions module, have been reordered to follow a consistent style. Specifically, third-party library imports (torchvision, torch) are now grouped together and placed after the local project imports (pytorch\_lightning, numpy). This change improves code readability and maintainability by ensuring a uniform import structure across these components.
sgm/data, sgm/modules/distributions · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 61.
Lenses
- Code Health 87
- Architecture 98
- Maturity 51
- Readiness 55
- Security 82
Changes since last survey
- 82 commits — 59 feature/other, 23 fixes
By area
- (root) — 25 commits
- sgm/modules — 16 commits
- (repo) — 13 commits
- scripts/demo — 11 commits
- scripts/sampling — 4 commits
- assets/sv4d_videos — 2 commits
- sgm/models — 2 commits
- .github/workflows — 1 commit
- assets/sv4d_example_video — 1 commit
- configs/inference — 1 commit
- model_licenses/LICENSE-SDV — 1 commit
- requirements/pt13.txt — 1 commit
- scripts/tests — 1 commit
- scripts/util — 1 commit
- sgm/inference — 1 commit
- sgm/util.py — 1 commit
Notable commits
- fix: Fix HEAD (#306)
- fix: Fix crashing line in logging in sgm/models/diffusion.py (#64)
- fix: Fix instruction
- fix: Fix license-files setting for project (#71)
- fix: Fix link (#24)
- fix: Fix loading safetensors with load_model_from_config
- fix: Fix to SV3D link
- fix: Fixes azimuth, adds simple instruction (#307)
- fix: Fixes links
- fix: Merge pull request #25 from pharmapsychotic/bugfix/watermark
- fix: Merge pull request #28 from jenuk/fix-samples_z
- fix: Merge pull request #43 from akx/fix-safetensors-load
- fix: Revert "Dead code removal (#48)" (#62)
- fix: Revert "Minimize re-exports from init files (#44)" (#63)
- fix: Revert "Replace most print()s with logging calls (#42)" (#65)
- fix: Revert "fall back to vanilla if xformers is not available (#51)" (#61)
- fix: SV4D 2.0 bug fix
- fix: Watermark encoder expects images in BGR channel order (matching cv2 imread). This fix reduces the watermark artifacts.
- fix: fix EDMDiscretization sigma_min for correct sampling noise scheduling (#114)
- fix: sv4d: fix readme
- …and 62 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Stability-AI/generative-models was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit e8cd657656fa5d61688191730d0e03242bf4ed44 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.