Dao-AILab/flash-attention
61.4
Adequate · 18 September 2026
75.1k
lines of production code
Python
with C, C++
1
measurement over time
What this system is
This system is a high-performance library for implementing FlashAttention and related neural network kernels, optimized for modern NVIDIA and AMD GPUs. It provides CUDA and Triton-based implementations for attention mechanisms, fused transformer operations, and training utilities, supporting architectures from Hopper to Blackwell. The codebase includes comprehensive benchmarking, debugging tools, and distributed training infrastructure to facilitate efficient model development and evaluation.
How it got here
2022 — FlashAttention-4 migration and training infrastructure
22 changes.
The project restructured its codebase to prioritize FlashAttention-4 for Hopper and Blackwell GPUs, removing legacy CUDA implementations and updating build tools. It expanded its scope by introducing comprehensive training infrastructure, including PyTorch Lightning callbacks and data modules, alongside a wide array of fused operations and Hugging Face model converters.
2023–2026 — Flash Attention 4 and multi-architecture expansion
15 changes.
This period focused on releasing Flash Attention 4 with standalone packaging and comprehensive testing, while expanding hardware support to include AMD ROCm via ComposableKernel and Aiter. The work also introduced Flash Attention 3 for NVIDIA Hopper and Blackwell architectures, alongside significant CI infrastructure improvements and debugging tooling.
Features
Add Hopper-specific Flash Attention 3 interface and configuration
This change introduces the \hopper/flash\_attn\_3\ package, providing a dedicated interface for Flash Attention 3 on Hopper architecture. The package exposes version metadata, re-exports configuration from the base \flash\_attn\_config\, and wraps the low-level \\_flash\_attn\_forward\ and \\_flash\_attn\_backward\ operations from the underlying C++/CUDA interface for direct use.
_hopper/flash\_attn\3 · high confidence
Add Hugging Face model converters for Baichuan, BERT, BigCode, BTLM, Falcon, GPT-NeoX, GPT-J, LLaMA, and OPT
The \flash\_attn/models\ package now includes dedicated converter modules for a wide range of Hugging Face models, enabling users to load and run these architectures using the flash attention implementation. New files provide state-dict remapping and configuration translation for Baichuan, BERT, BigCode (StarCoder), BTLM, Falcon, GPT-NeoX, GPT-J, LLaMA (including Meta and Hugging Face formats), and OPT. These converters handle specific architectural differences such as weight layout transformations (e.g., QKV concatenation), layer norm naming, and embedding padding, allowing seamless interoperability between Hugging Face checkpoints and the flash attention engine.
_flash\attn/models · high confidence
Add sass\_diff tool for comparing SASS binaries
A new Python utility, sass\_diff.py, has been added to the tools directory to compare two SASS (SASS Assembly) files. The tool normalizes register assignments so that instructions using different registers but performing the same operation are treated as identical, allowing users to focus on logical differences rather than register allocation changes. It supports unified-diff-style output with configurable context, color coding, and a summary mode to quickly see instruction counts and differences.
tools · high confidence
Flash Attention 3 (Hopper) v3.0.0 release
This update introduces Flash Attention 3 (FA3) as the new default attention implementation for NVIDIA Hopper (SM90) architectures, marking the release of version 3.0.0. The new kernel brings significant performance improvements and expanded capabilities, including native support for FP8 data types, Multi-Head Latent Attention (MLA), and Paged Attention with variable-length sequences. It also introduces a new tiled scheduler for better L2 cache utilization and supports diverse attention patterns such as local, sliding window, and grouped-query attention (GQA). The release includes comprehensive benchmarking scripts to compare FA3 against FA2 and other backends, alongside robust build configurations for CUDA 12.x and Python 3.10+.
hopper · high confidence
FlashAttention-2 CUDA kernel implementation with optimized backward pass and modular structure
The \csrc/flash\_attn\ directory now contains the core CUDA kernel implementation for FlashAttention-2, featuring a refactored architecture that splits the backward pass into separate files per head dimension (32, 64, 96, 128, 192, 256) and data type (FP16, BF16) to significantly reduce compilation times. This update introduces a modular design with dedicated headers for Alibi bias, dropout, and masking logic, while the Python API (\flash\_api.cpp\) has been updated to support new features such as local attention windows, softcapping, and paged KV caches. The implementation also includes fixes for various edge cases in variable-length sequences and multi-query attention, ensuring correct behavior for causal and non-causal masking scenarios.
_csrc/flash\attn · high confidence
FlashAttention-2 v2.8.4 release with Triton support and block-sparse attention
This release updates the flash\_attn package to version 2.8.4, introducing a new \flash\_attn\ namespace that allows co-installation with FlashAttention-4. It adds an experimental Triton-based implementation (\flash\_attn\_triton.py\) supporting head dimensions up to 128, arbitrary sequence lengths, and attention bias, alongside a fallback mechanism for ROCm/AI-ter backends. The package also includes a new block-sparse attention module (\FlashBlocksparseAttention\) and refactors the sparse attention interface to use the \flash\_attn\_cuda\ kernel instead of the legacy \stream\_attn\_cuda\.
_flash\attn · high confidence
Initial release of optimized training infrastructure and documentation
This change introduces the training environment for the first time, providing a Dockerfile that sets up a PyTorch-based container with specific dependencies including FlashAttention 2.6.3, DeepSpeed 0.7.7, and various ML libraries. It also adds the main entry point script (run.py) which uses Hydra for configuration management to launch training or evaluation modes, and a README.md documenting the optimized Transformer implementation, model components, and usage instructions for training GPT models on datasets like Openwebtext and The Pile.
training · high confidence
Introduce AMD ROCm FlashAttention backend via ComposableKernel
Adds a new C++ implementation layer in csrc/flash\_attn\_ck that enables FlashAttention operations on AMD ROCm hardware using the ComposableKernel (CK) library. This change introduces the core forward and backward kernels (including variable-length and KV-cache variants) and exposes them as a PyTorch extension module, providing the necessary backend support for ROCm users.
_csrc/flash\_attn\ck · high confidence
Introduce Triton-based CrossEntropyLoss with advanced scaling and parallel support
The \flash\_attn/losses\ module now provides a new \CrossEntropyLoss\ implementation powered by Triton kernels. This component supports label smoothing, logit scaling, and a configurable LSE square scale (z-loss) for improved training stability. It also enables tensor parallelism via \process\_group\ for distributed vocabularies, allows passing precomputed LSE values, and can optionally return the z-loss component for logging purposes.
_flash\attn/losses · high confidence
Introduce fused activation, dense, and normalization operations
This change adds new high-performance PyTorch operations to the \flash\_attn/ops\ module, including fused GELU and SwiGLU activations, a \FusedDense\ linear layer with Tensor Parallel support, and fused LayerNorm/RMSNorm variants with dropout and residual addition. These components provide optimized building blocks for transformer architectures, enabling features like sequence parallelism, parallel residuals, and subset processing while maintaining compatibility with mixed-precision training.
_flash\attn/ops · high confidence
Introduce fused dense CUDA extension with bfloat16 support
This location adds a new CUDA extension (fused\_dense\_lib) that provides fused matrix multiplication with bias and activation (GELU or ReLU) operations for both fp16 and bfloat16 data types. The implementation leverages cuBLASLt for optimized performance, automatically allocating a larger workspace (32MB) on Hopper architecture GPUs (compute capability \>= 9) and a smaller one (4MB) for others. It includes Python bindings for forward and backward passes, allowing users to install and use these fused operations via pip within this directory.
_csrc/fused\_dense\lib · high confidence
Introduces core Transformer building blocks for FlashAttention models
The \flash\_attn/modules\ package now provides foundational PyTorch modules for constructing Transformer-based architectures, including \Block\ (with pre/post-norm and fused dropout/add/layer-norm support), \MHA\ (Multi-Head Attention with FlashAttention kernels, ALiBi, and windowing), \Mlp\ (standard and Gated variants like SwiGLU), and \Embedding\ layers (standard and tensor-parallel variants like \VocabParallelEmbedding\). These components enable users to build efficient, scalable models (e.g., LLaMA, GPT, BERT) that leverage FlashAttention's performance benefits and support distributed training via tensor parallelism.
_flash\attn/modules · high confidence
New AI debugging documentation and trace parsing tools
Added a suite of new documentation files in the AI directory to support kernel development and debugging: CLC\_TRACE\_DEBUG.md explains how to capture and interpret CLC scheduler traces; DEBUG\_2CTA.md details strategies for debugging GPU kernel hangs and deadlocks in 2CTA configurations; DEBUG\_METHODOLOGY.md establishes a root-cause investigation protocol; RACECHECK\_TMA\_HAZARD.md documents a false-positive racecheck issue with cp.async.bulk and its fix; SASS\_MMA\_ANALYSIS.md provides a guide for analyzing HGMMA instructions in SASS output; SM90\_BLOCK\_SIZE\_TUNING.md offers a tuning guide for SM90 block sizes; SM90\_R2P\_MASKING\_SASS.md analyzes the performance impact of R2P masking; SPARSE\_MLA\_RECOMPUTE\_P.md outlines the design for in-kernel recompute-P and token-chunked backward for sparse MLA; and VARLEN\_PREPROCESS\_TILE\_BUG.md records a fix for a varlen preprocess tile mismatch. Additionally, parse\_clc\_log.py was added to parse and visualize CLC trace logs.
AI · high confidence
New CUDA LayerNorm extension with fused operations and large dimension support
This location introduces a new CUDA extension for fused dropout, residual addition, and LayerNorm (including RMSNorm). The implementation supports hidden dimensions up to 8192 (divisible by 8) and works with both pre-norm and post-norm architectures. It provides forward and backward kernels for various data types (fp32, fp16, bf16) and registers launchers for specific hidden sizes (256 to 8192). The API allows for optional row/column scaling and subset input/output handling. Note that the README indicates this specific CUDA extension is no longer used in the FlashAttention repo as of 2024-01-05, having been replaced by a Triton-based implementation.
_csrc/layer\norm · high confidence
New FA4 CI container tooling for registry-free GPU jobs
Added a new Dockerfile and build scripts in tools/ci/docker to support a registry-free CI workflow. The Dockerfile provisions a CUDA 13.0 environment with PyTorch nightly (cu130) and Flash Attention 4 dependencies, while the accompanying shell scripts enable building the container locally into an Apptainer SIF image, bypassing the need for a container registry. This allows CI runners to provision the overlay from a fresh session using a locally built SIF file.
tools/ci/docker · high confidence
New FA4 CI infrastructure with Apptainer and runtime dependency enforcement
Added a new continuous integration workflow for Flash Attention 4 (FA4) that runs on self-hosted GPU runners using Apptainer (SIF) containers. The system implements a two-pass test strategy (compilation via FakeTensorMode followed by GPU execution) and supports a registry-free mode using runner-local SIF images. To prevent version mismatches between the baked container image and the repository code, the CI now installs cutlass-dsl, quack, and FA4 at runtime and includes a pre-flight check to enforce minimum dependency versions defined in pyproject.toml, failing early if the image is stale.
tools/ci · high confidence
New Triton-based implementations for core attention operations
The \flash\_attn/ops/triton\ module now provides native Triton implementations for LayerNorm, CrossEntropy, Rotary embeddings, and MLP activations. Users benefit from optimized performance on modern GPUs, including support for zero-centered LayerNorm weights, logit scaling and label smoothing in CrossEntropy, variable-length Rotary embeddings, and squared ReLU activations in FusedMLP layers.
_flash\attn/ops/triton · high confidence
New suite of attention and GEMM benchmarking scripts
Added a comprehensive set of new benchmarking scripts to the benchmarks directory to evaluate FlashAttention and related kernels across various configurations and backends. bench\_sm90.py provides a unified benchmark for SM90 forward and backward passes with support for headdim sweeps and tile size tuning. benchmark\_attn.py introduces a unified framework to compare multiple backends (Standard, FA2, cuDNN, FA3, FA4) with detailed MFU% reporting. Additional scripts include benchmark\_alibi.py for ALiBi attention, benchmark\_causal.py and benchmark\_flash\_attention.py for causal and general attention performance, benchmark\_gemm.py for GPU matrix multiplication throughput, benchmark\_mla\_paged\_kv.py for MLA paged key-value cache performance, benchmark\_varlen\_sched.py for variable-length tile scheduler comparisons, and clc\_bench.py for CLC (Chunked Local Context) sweeps.
benchmarks · high confidence
New training infrastructure: callbacks, data modules, and distributed utilities
This change introduces a comprehensive set of new components for the training pipeline. It adds a suite of PyTorch Lightning callbacks for monitoring and optimization, including Exponential Moving Average (EMA), FLOP counting, GPU affinity management, loss scale monitoring, norm tracking, speed monitoring, and Weights & Biases integration (model watching, code/checkpoint uploading, confusion matrix and F1/precision/recall heatmaps). It also introduces fault-tolerant random and distributed samplers to ensure reproducible resumption, along with data modules for ImageNet and Hugging Face language modeling datasets (featuring shared memory optimization and detokenization). Finally, it adds distributed communication hooks for FP16 gradient compression and a custom Mixup wrapper for timm.
training/src · high confidence
New utility modules for benchmarking, distributed training, and model generation
The \flash\_attn/utils\ package now includes several new modules to support advanced usage patterns. \benchmark.py\ provides helper functions to measure forward, backward, and combined pass performance using PyTorch's benchmarking tools, with support for automatic mixed precision. \distributed.py\ introduces utilities for tensor-parallel operations, including \all\_gather\, \reduce\_scatter\, and \all\_reduce\ functions that handle both legacy and modern PyTorch distributed APIs, along with helpers for synchronizing shared parameters and reducing sequence-parallel gradients. \generation.py\ adds a comprehensive text-generation loop supporting greedy, top-k, and top-p sampling, with optional CUDA graph caching for accelerated decoding. \pretrained.py\ enables loading model weights from Hugging Face Hub or local directories, supporting both standard and safetensors formats, and handling sharded checkpoints. \library.py\ provides a \triton\_op\ decorator to register Triton kernels as PyTorch custom operations, improving compatibility with \torch.compile\. \torch.py\ ensures compatibility with newer PyTorch versions by adapting AMP decorators. \testing.py\ offers utilities for generating random padding masks and preparing QKV tensors for attention tests.
_flash\attn/utils · high confidence
Standalone FlashAttention-4 (CuTeDSL) package with SM90/SM100 support
The \flash\_attn/cute\ directory is now a standalone installable package (\flash-attn-4\) providing a CuTeDSL-based FlashAttention implementation for Hopper (SM90) and Blackwell (SM100) GPUs. This release introduces new kernel implementations for forward and backward passes, including support for variable sequence lengths, paged KV attention, block-sparse tensors, and FP8 precision. It also includes comprehensive benchmarking utilities for performance analysis and reference implementations, along with the necessary build manifests, license, and author documentation to function as an independent library.
_flash\attn/cute · high confidence
Removals
Removal of legacy FMHA CUDA implementation from csrc/stream\_attn
The \csrc/stream\_attn\ directory has been cleaned up by removing the original NVIDIA Apex-based Flash Attention (FMHA) source files, including the PyTorch C++ API (\fmha\_api.cpp\), core CUDA headers (\fmha.h\, \gemm.h\, \gmem\_tile.h\, \smem\_tile.h\, \softmax.h\, \utils.h\, \mask.h\, \kernel\_traits.h\), and the associated README. This change eliminates the legacy implementation code from this location, likely in preparation for or following a replacement with a different attention mechanism or optimized kernel.
_csrc/stream\attn · high confidence
Behavioural changes
FlashAttention-3 kernel instantiations split into separate compilation units
The auto-generated kernel instantiation files in the \hopper/instantiations\ directory have been reorganized into individual \.cu\ files (e.g., \flash\_bwd\_hdim128\_bf16\_sm80.cu\, \flash\_fwd\_hdim128\_bf16\_sm100.cu\) to speed up compilation. This change introduces support for the new SM100 architecture alongside existing SM80, SM86, and SM90 targets, and expands backward pass support to include head dimensions of 64, 96, 128, 192, and 256. The refactoring also ensures that features like PagedKV and PackGQA are consistently enabled for relevant architectures to reduce binary size and compilation overhead.
hopper/instantiations · high confidence
Migrate AMD backend to Aiter
The AMD backend has been migrated to use the Aiter library, implemented by adding the Aiter submodule to the third\_party directory. This change replaces the previous implementation with the new Aiter-based backend for AMD hardware support.
_third\party · medium confidence
New PatchEmbed layer and refactored Rotary Embedding implementation
The \flash\attn/layers\ module now includes a new \PatchEmbed\ class for converting 2D images to patch embeddings, offering an API compatible with timm but using \nn.Linear\ for approximately 8x faster performance. Additionally, the rotary embedding logic has been significantly refactored: the \ApplyRotaryEmbQKV\\ operation is now wrapped as a custom Triton op, support for interleaved (GPT-J style) rotary embeddings has been added, and the implementation now handles variable sequence lengths and sequence length offsets for inference with KV caches. The \pos\_idx\_in\_fp32\ option has been removed, and frequency calculations are ensured to be in fp32.
_flash\attn/layers · high confidence
Repository restructured for FlashAttention-4 with legacy code removal
The repository has been reorganized to prioritize FlashAttention-4 (FA4), which is now the active development branch written in CuTeDSL for Hopper and Blackwell GPUs. Legacy components from earlier generations, including the streaming attention modules (\streaming\_attention.py\, \stream\_attn\_interface.py\), BERT padding utilities (\bert\_padding.py\), and rotary embeddings (\rotary.py\), have been removed. The project now uses \flash-attn-4\ as the PyPI package name and relies on submodules for NVIDIA CUTLASS and AMD Composable Kernel. Build infrastructure has been updated to use \pyproject.toml\ and \uv\ for dependency management, and the license has been changed to BSD 3-Clause.
(repo-wide) · high confidence
Test coverage
Add comprehensive test suites for FlashAttention kernels and rotary embeddings; Added comprehensive tests for Triton LayerNorm operations; Added model verification tests for Baichuan, BERT, BigCode, BTLM, Falcon, GPT, GPT-NeoX, and GPT-J; Added parallel module tests; Added tests for CrossEntropyLoss features and parallel execution; Added tests for NeoX and GPT-J style rotary embeddings; Added tests for fused linear, MLP, and layer normalization operations; Added tests for the Hugging Face language modeling data module; New CuTe DSL test suite and benchmarking infrastructure.
Dependencies
Flash Attention 4 CUTE implementation made installable as a standalone package
The \flash\_attn/cute\ module is now a standalone installable package (\flash-attn-4\) with its own \pyproject.toml\, allowing users to install the CUDA Template Engine implementation independently. This package requires Python 3.10+ and introduces specific dependencies including \nvidia-cutlass-dsl\>=4.6.2\ and \quack-kernels\>=0.5.3\, while also configuring build tools like setuptools and linting via ruff.
(dependencies) · high confidence
Updated Cutlass and Composable Kernel submodules
The csrc directory now points to new commits for the Cutlass and Composable Kernel submodules, incorporating upstream changes to these underlying libraries.
csrc · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 61.
Lenses
- Code Health 69
- Architecture 100
- Maturity 60
- Readiness 56
- Security 68
Changes since last survey
- 300 commits — 227 feature/other, 73 fixes
By area
- flash_attn/cute — 213 commits
- (root) — 16 commits
- .github/workflows — 10 commits
- tests/cute — 10 commits
- benchmarks/benchmark_attn.py — 6 commits
- csrc/flash_attn — 6 commits
- csrc/flash_attn_ck — 4 commits
- tools/ci — 4 commits
- .github/actions — 2 commits
- AI/RACECHECK_TMA_HAZARD.md — 2 commits
- benchmarks/tune_ex2_emu.py — 2 commits
- flash_attn/flash_attn_interface.py — 2 commits
- hopper/setup.py — 2 commits
- (repo) — 1 commit
- .github/scripts — 1 commit
- AI/DEBUG_2CTA.md — 1 commit
- AI/SM90_BLOCK_SIZE_TUNING.md — 1 commit
- AI/SM90_R2P_MASKING_SASS.md — 1 commit
- assets/fa4_paper.pdf — 1 commit
- benchmarks/bench_sm90.py — 1 commit
Notable commits
- fix: Build Fix: Update abi3 tag to cp310 and minimum python version to 3.10 (#2532)
- fix: Disable 2CTA fwd non-causal on CUDA 12 to work around codegen regression (#2461)
- fix: Fix (#2338)
- fix: Fix (#2505)
- fix: Fix CLC fuzz scheduler expectations (#2766)
- fix: Fix CuTe SM120 compile-time argument handling (#2671)
- fix: Fix FA2 + FA4 co-existence (#2331)
- fix: Fix GQA crash in cute FLASH backend: init load_Q before conditional (#2301)
- fix: Fix SM100 FP8 fwd with cutlass-dsl >=4.5.2 (MmaF8F6F4Op) (#2640)
- fix: Fix SM90 bwd crash for head_dim in (128, 192] (#2482)
- fix: Fix ZeroDivisionError in num_splits_heuristic for empty Q workloads (#2515)
- fix: Fix seqlen_k_loaded treating a window bound of 0 as unbounded (sliding window) (#2624)
- fix: Fix clang parser error of missing 'typename' prior to dependent type name occurs because LLVM/Clang is strictly adhering to C++ standards (#2295)
- fix: Fix clc scheduling request bug (#2508)
- fix: Fix compatibility issues with CuTe DSL 4.6.0+ (#2648)
- fix: Fix duplicated word in layer norm comment (#2744)
- fix: Fix edge case when tag has no delta from previous (#2394)
- fix: Fix long MSVC linker commands on Windows (#2517)
- fix: Fix removed Quack packed subtraction API (#2787)
- fix: Fix sm100 fwd missing tSrQs init regression (#2293)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Dao-AILab/flash-attention was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 1bda8f9290cd48d030f1516f0e680cd464ef3554 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.