hpcaitech/Open-Sora
49.0
Weak · 18 September 2026
15.8k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is an open-source video generation framework, currently at version 2.0, that utilizes Flux-based and MMDiT architectures to perform text-to-video, image-to-video, and video-to-video synthesis. It provides a complete infrastructure for training and inference, featuring distributed training optimizations, memory-efficient data loading, and support for various autoencoders like DC-AE and Hunyuan VAE. The project includes a web-based Gradio interface for interactive generation and is packaged for standard installation with comprehensive contributor guidelines.
Features
Initial release of Open-Sora v1.3 Gradio demo application
This change introduces the \gradio/app.py\ script, providing a web-based interface for the Open-Sora v1.3 model. The application supports text-to-video (t2v) and image-to-video (i2v) generation modes, loading specific STDiT3 model weights from Hugging Face based on the selected resolution (360p or 720p) and mode. It includes logic to automatically install optional performance optimizations like Flash Attention and Apex, handles configuration parsing for the v1.3 inference pipeline, and implements the core inference loop with watermarking and output saving capabilities.
gradio · high confidence
Initial release of core inference and training utilities
This change introduces the foundational utility modules for the Open-Sora project, establishing the infrastructure for model execution and management. It adds distributed inference support via ColossalAI (including tensor and sequence parallelism handling in \cai.py\), comprehensive checkpoint loading and sharding logic (\ckpt.py\), and configuration parsing with command-line argument merging (\config.py\). The update also provides the main inference pipeline (\inference.py\) with support for text-to-image, text-to-video, and image-to-video workflows, alongside prompt refinement capabilities using OpenAI's GPT-4o (\prompt\_refine.py\). Furthermore, it implements the core sampling algorithms (\sampling.py\), training loop helpers including EMA updates and optimizer creation (\train.py\, \optimizer.py\), and essential logging and memory monitoring tools (\logger.py\, \misc.py\).
opensora/utils · high confidence
Introduce MMDiT model architecture with distributed training support
The \opensora/models/mmdit\ package now provides a new MMDiT (Multi-Modal Diffusion Transformer) model implementation, adapted from the Flux architecture. This includes the core model definition (\model.py\), specialized layers such as DoubleStreamBlock and SingleStreamBlock (\layers.py\), and mathematical utilities for attention and RoPE (\math.py\). To support large-scale generation, the package also introduces distributed training capabilities via \distributed.py\, which implements tensor and context parallelism using ColossalAI and Flash Attention, along with a corresponding sharding policy in \policy.py\.
opensora/models/mmdit · high confidence
Introduce unified model package exports
The opensora.models package now exposes a consolidated set of model components, including DC-AE, Hunyuan VAE, MM-DiT, text models, and VAE, making them directly accessible via the main package namespace.
opensora/models · high confidence
Introduction of module and dataset registry system
The OpenSora codebase now includes a registry system for managing models and datasets. This change introduces a \build\_module\ utility that allows components to be instantiated either from configuration dictionaries or directly from module objects, facilitating a more flexible and configurable architecture for model and dataset registration within the \opensora.models\ and \opensora.datasets\ locations.
opensora · high confidence
New VAE module with 2D autoencoder, 3D discriminator, and tensor-parallel support
The \opensora/models/vae\ package has been introduced, providing the core components for video generation training and inference. This includes a 2D autoencoder (\AutoEncoderFlux\) adapted from Flux, a 3D PatchGAN discriminator (\N\_LAYER\_DISCRIMINATOR\_3D\) for adversarial training, and a comprehensive loss module (\VAELoss\, \GeneratorLoss\, \DiscriminatorLoss\) that combines reconstruction, perceptual (LPIPS), and KL losses. To support large-scale training, the module also includes tensor-parallel implementations for 3D convolutions (\Conv3dTPCol\, \Conv3dTPRow\) and utility functions for channel-chunked 3D convolutions and diagonal Gaussian distributions, enabling memory-efficient processing of high-dimensional video data.
opensora/models/vae · high confidence
New acceleration module for distributed training optimizations
A new \opensora/acceleration\ package has been introduced to support distributed training performance. This includes an activation checkpointing system with CPU offloading capabilities in \checkpoint.py\ to reduce GPU memory usage, distributed communication primitives (All-to-All, Gather-Split) in \communications.py\, and parallel state management in \parallel\_states.py\. Additionally, it provides Shardformer integration for the T5 model, featuring a custom T5LayerNorm implementation and JIT-optimized policy replacements for T5 encoder layers.
opensora/acceleration · high confidence
Open-Sora 2.0 release with official packaging and contributor guidelines
This update introduces Open-Sora 2.0, featuring an 11B model that achieves performance on par with larger proprietary models. The release includes a formal \setup.py\ package definition (version 2.0.0) for standard installation, a \CONTRIBUTING.md\ guide detailing the development environment and pre-commit hooks, and an updated \LICENSE\ that explicitly incorporates the Tencent Hunyuan Community License Agreement alongside existing open-source licenses.
(repo-wide) · high confidence
Behavioural changes
Introduces new dataset infrastructure with aspect-ratio bucketing and memory-efficient video loading
The \opensora/datasets\ module has been refactored to support a new data loading pipeline. Users can now train on datasets with varied aspect ratios, as the new \aspect.py\ and \bucket.py\ components automatically map resolutions to predefined aspect ratios (e.g., 16:9, 9:16) and bucket videos by frame count and resolution for optimized batching. The \dataloader.py\ and \pin\_memory\_cache.py\ files introduce a custom multiprocessing iterator with a pin-memory cache to reduce memory overhead during training. Additionally, \read\_video.py\ implements a PyAV-based video reader that explicitly manages container lifecycle and garbage collection to prevent thread and memory leaks, while \datasets.py\ adds support for efficient Parquet file reading via sharding.
opensora/datasets · high confidence
Open-Sora 2.0 release with Flux-based video generation and DC-AE support
This update introduces Open-Sora 2.0, shifting the core video generation model to a Flux-based architecture (configurable via \type="flux"\) and adding support for the DC-AE (Diffusion Causal Autoencoder) as an alternative to the Hunyuan VAE. The release includes new inference configurations for 256px and 768px resolutions, enabling text-to-video (t2v), image-to-video (i2v), and video-to-video (v2v) generation with tensor and sequence parallelism plugins. Training configurations are provided for multi-stage training (stage1, stage2) and high-compression modes, alongside new example prompt datasets for text-to-video and image-to-video tasks.
(repo-wide) · high confidence
Dependencies
Initial dependency specification for Open-Sora 2.0
Added a new requirements.txt file defining the project's third-party dependencies, including PyTorch 2.4.0, TorchVision 0.19.0, ColossalAI, and various libraries for model execution, data processing, and monitoring.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 49.
Lenses
- Code Health 86
- Architecture 99
- Maturity 55
- Readiness 27
- Security 74
Changes since last survey
- 300 commits — 238 feature/other, 62 fixes
By area
- (repo) — 78 commits
- (root) — 58 commits
- docs/zh_CN — 22 commits
- tools/caption — 20 commits
- docs/report_03.md — 18 commits
- opensora/models — 16 commits
- gradio/app.py — 14 commits
- configs/opensora-v1-2 — 12 commits
- opensora/datasets — 8 commits
- opensora/utils — 8 commits
- docs/commands.md — 4 commits
- eval/sample.sh — 4 commits
- docs/data_processing.md — 3 commits
- docs/datasets.md — 3 commits
- scripts/inference.py — 3 commits
- configs/vae — 2 commits
- docs/report_02.md — 2 commits
- docs/structure.md — 2 commits
- eval/README.md — 2 commits
- eval/vbench — 2 commits
Notable commits
- fix: Docs/fix zangwei (#471)
- fix: Docs/fix zangwei (#474)
- fix: Docs/fix zw (#476)
- fix: Hotfix/forcehf (#483)
- fix: Hotfix/readme (#789)
- fix: Hotfix/t5 load (#486)
- fix: Hotfix/t5 load (#487)
- fix: Hotfix/vae (#502)
- fix: Merge branch 'main' of https://github.com/hpcaitech/Open-Sora-Dev into docs-fix
- fix: Merge pull request #154 from hpcaitech/hotfix/t5_assert
- fix: Merge pull request #161 from hpcaitech/hotfix/dataset
- fix: Merge pull request #162 from hpcaitech/hotfix/fix-sp
- fix: Merge pull request #164 from hpcaitech/hotfix/readme
- fix: Merge pull request #165 from hpcaitech/hotfix/sample_task_type
- fix: Merge pull request #169 from hpcaitech/hotfix/cut
- fix: Merge pull request #458 from hpcaitech/hotfix/license
- fix: Merge pull request #506 from hpcaitech/docs-fix
- fix: Merge pull request #515 from hpcaitech/hotfix/gradio
- fix: Merge pull request #523 from BurkeHulk/hotfix/fp16_nan_output
- fix: Merge pull request #526 from hpcaitech/fix/memory_leak
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
hpcaitech/Open-Sora was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 7ad6a96a135feb81f755c84fb391818718f6beb2 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.