tracel-ai/burn
67.3
Adequate · 28 September 2026
293.3k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a modular deep learning framework that provides a unified, backend-agnostic interface for tensor operations, automatic differentiation, and neural network module construction. It supports diverse hardware execution through a dispatch layer routing to CPU, GPU, and remote backends, while offering comprehensive tools for data loading, model training, and reinforcement learning. The framework emphasizes performance via automatic operation fusion and memory optimization, alongside robust model persistence and distributed training capabilities.
How it got here
2022–2024 — modularization and remote execution
83 changes.
The project underwent a major architectural restructuring, introducing a unified Device API and splitting the codebase into modular extension crates like burn-core, burn-dataset, and burn-fusion. This period also established comprehensive documentation, expanded the example suite to cover diverse training workflows, and introduced a new remote backend for distributed tensor execution via Iroh.
2025 — comprehensive framework expansion
73 changes.
This period focused on significantly expanding the Burn framework by introducing new crates for neural network modules (burn-nn), vision operations (burn-vision), and intermediate representation (burn-ir). It also established robust multi-device and distributed training capabilities, including DDP support, while adding new backends like MLIR-based CPU and ROCm to broaden hardware compatibility.
2026 — Reinforcement learning and multi-backend dispatch
45 changes.
This period focused on introducing reinforcement learning capabilities through the new burn-rl crate and training infrastructure, alongside a major architectural shift to unified multi-backend execution via the burn-dispatch system. Significant work also went into expanding the CPU backend with optimized tensor operations, adding various vision and linear algebra metrics, and establishing robust remote training examples using Iroh networking.
Features
Add 1D and 2D interpolation modules with configurable resizing options
The \burn-nn\ crate now includes \Interpolate1d\ and \Interpolate2d\ modules for resizing tensors. These modules allow users to specify either an \output\_size\ or a \scale\_factor\ to determine the new dimensions, and support multiple interpolation modes including Nearest, NearestExact, Linear, Cubic, and Lanczos. Additionally, an \align\_corners\ option is available to control pixel alignment during the resizing process.
crates/burn-nn/src/modules/interpolate · high confidence
Add 1D, 2D, 3D, and Deformable Convolution modules with asymmetric padding support
The \burn-nn\ crate now includes new convolution modules (\Conv1d\, \Conv2d\, \Conv3d\, \ConvTranspose1d/2d/3d\, and \DeformConv2d\) that support asymmetric padding configurations. Users can now specify padding as explicit left/right (or top/bottom/left/right) pairs, allowing for precise control over output dimensions. The \Same\ padding mode is supported for odd kernel sizes, while even kernel sizes with \Same\ padding will panic to prevent ambiguous output sizes. These modules also include validation to ensure input/output channels are divisible by the number of groups.
crates/burn-nn/src/modules/conv · high confidence
Add 2D affine image transformation capabilities
The \burn-vision\ crate now includes a \Transform2D\ module that enables applying affine transformations to image tensors. Users can rotate, scale, translate, or shear images using dedicated constructors, and these transforms can be composed into a single operation. The implementation leverages the ONNX \GridSample\ operator with bilinear interpolation and configurable padding modes to perform the resampling.
crates/burn-vision/src/transform · high confidence
Add A-FINE image quality metric
Introduces the A-FINE (Adaptive Fidelity-Naturalness Evaluator) full-reference image quality metric in \crates/burn-train/src/metric/vision/afine\. This new capability combines a naturalness branch and a fidelity branch over CLIP ViT-B/32 features to produce a per-sample quality score in the range (0, 100). The implementation includes a custom CLIP visual encoder, naturalness and fidelity heads, logistic calibrators, and an adapter module, along with logic to automatically download and load pretrained weights from the PyIQA project on Hugging Face.
crates/burn-train/src/metric/vision/afine · high confidence
Add CIFAR-10 and CIFAR-100 dataset support
Users can now load the CIFAR-10 and CIFAR-100 image classification datasets directly via the new \CifarDataset\ accessor in \burn-dataset\. This feature automatically downloads, extracts, and caches the datasets (mirrored from fastai) into the system cache directory, providing convenient access to the 50,000-image training set and 10,000-image test set for both the 10-class and 100-class variants.
crates/burn-dataset/src/vision · high confidence
Add CPU-based morphological operations (erode/dilate) with SIMD acceleration
This change introduces basic morphological operations—specifically erosion and dilation—to the CPU backend in the burn-vision crate. The implementation supports various data types (including float, integer, and boolean) and utilizes SIMD vectorization for performance optimization. Users can now apply structural element filters to tensors with configurable kernel sizes, anchors, border types, and iteration counts directly on the CPU.
crates/burn-vision/src/backends/cpu/morphology · high confidence
Add Cube backend support for connected components operations
The Cube backend now implements the \BoolVisionOps\ trait, providing hardware-accelerated support for connected components analysis on boolean tensors. This includes \connected\_components\ and \connected\_components\_with\_stats\, which attempt to use a dedicated hardware-accelerated path and fall back to the CPU implementation if acceleration is unavailable. The \IntVisionOps\, \FloatVisionOps\, and \VisionBackend\ traits are also implemented but currently contain no specific logic.
crates/burn-vision/src/backends/cube · high confidence
Add DISTS perceptual image quality metric
Introduces the DISTS (Deep Image Structure and Texture Similarity) metric for full-reference image quality assessment. This new capability allows users to compute perceptual similarity between images by combining structure and texture similarity using deep features from a VGG16 backbone. The implementation includes a custom L2 pooling layer for smoother downsampling, an ImageNet-normalized VGG16 feature extractor, and support for loading pretrained weights (both the VGG16 backbone and DISTS-specific alpha/beta weights) automatically from remote URLs.
crates/burn-train/src/metric/vision/dists · high confidence
Add Frechet Inception Distance (FID) vision metric
Users can now evaluate the quality and diversity of generated images using the Frechet Inception Distance metric. This change introduces a new \Fid\ module in \crates/burn-train/src/metric/vision/fid\ that implements an InceptionV3 feature extractor and computes the FID score between real and generated image distributions. The implementation includes automatic downloading and caching of pretrained PyTorch InceptionV3 weights, normalization using ImageNet statistics, and computation of the Frechet distance via Newton-Schulz iteration for matrix square roots.
crates/burn-train/src/metric/vision/fid · high confidence
Add Gram Matrix Loss for Neural Style Transfer
The \burn-vision\ crate now includes a Gram Matrix Loss module, enabling users to compute style transfer losses by comparing texture correlations between images. This feature introduces a \GramMatrixLoss\ component that leverages a pretrained VGG19 feature extractor (with weights automatically downloaded and cached) to extract features from five specific convolutional layers. Users can configure layer weights and choose between average or max pooling for the feature extraction, then compute the loss between predicted and target images with configurable reduction (mean or sum).
crates/burn-vision/src/loss · high confidence
Add LPIPS perceptual similarity metric with pretrained backbone support
The LPIPS (Learned Perceptual Image Patch Similarity) metric is now available for evaluating image quality by measuring perceptual distance between images using deep features. This implementation supports three backbone networks—VGG16, AlexNet, and SqueezeNet—and includes logic to automatically download and cache official pretrained weights from the PerceptualSimilarity repository and PyTorch model hubs. Users can initialize the metric with pretrained weights via \LpipsConfig::init\_pretrained\ for accurate results or with random weights for experimentation, and can configure input normalization via the \normalize\ flag.
crates/burn-train/src/metric/vision/lpips · high confidence
Add LSTM and Bidirectional LSTM modules with configurable state management
This change introduces new \Lstm\ and \BiLstm\ modules in \burn-nn\, allowing users to implement unidirectional and bidirectional Long Short-Term Memory networks. The \Lstm\ module supports configurable input/output dimensions, bias, initialization strategies, and gate activations, with optional support for processing sequences in reverse order and clipping cell states to prevent explosion. The \BiLstm\ module combines forward and reverse LSTMs, handling dimension swapping for batch-first or sequence-first layouts. A new \LstmState\ type bundles cell and hidden states, providing methods for initialization, slicing, stacking, chunking, and dimension manipulation (squeeze/unsqueeze), including support for negative dimensions. These components enable more flexible and robust sequence modeling capabilities.
crates/burn-nn/src/modules/rnn/lstm · high confidence
Add MNIST guide example demonstrating training and inference workflow
A new example located at examples/guide provides a complete, runnable workflow for training and inferring a convolutional neural network on the MNIST dataset. The example includes binaries for training (train.rs), inference (infer.rs), and model inspection (print.rs), along with source modules for data batching, model definition, and training logic. It serves as a concrete reference implementation for the Burn book's basic workflow guide, showing how to configure datasets, optimizers, and learners, and how to save and load trained model artifacts.
examples/guide · high confidence
Add MNIST inference demo for the web
Introduces a new web-based inference example that runs a trained MNIST model directly in the browser using WebAssembly. The demo allows users to draw digits on a canvas, which are processed by a Web Worker to perform inference asynchronously without blocking the UI. It supports both the \flex\ (CPU) and \wgpu\ (WebGPU) backends, exposing the Rust model via \wasm-bindgen\ and providing a local HTTP server script to serve the static assets.
examples/mnist-inference-web · high confidence
Add P2P remote training example using Iroh networking
A new example demonstrates distributed training over a peer-to-peer network using the Iroh library. The \p2p-remote-training\ crate provides a server and client mode where participants connect via a shared topic string that acts as a stable identity. The client explicitly disables UDP segmentation offload (GSO) in the QUIC transport configuration to work around a known Iroh issue, ensuring stable connectivity for the demo.
examples/p2p-remote-training · high confidence
Add PyTorch model loading support via PytorchStore
Users can now load PyTorch checkpoint files (.pth, .pt) directly into Burn using the new \PytorchStore\. This feature automatically handles weight transformations, such as transposing Linear layer weights and renaming normalization parameters (gamma/beta to weight/bias). It also supports flexible loading options, including filtering specific layers via regex, remapping tensor keys, extracting nested state dicts, and allowing partial loads when some tensors are missing.
crates/burn-store/src/pytorch · high confidence
Add RNN, GRU, and LSTM modules to burn-nn
This change introduces new recurrent neural network modules to the \burn-nn\ crate, including a basic unidirectional RNN, a Gated Recurrent Unit (GRU), and a Long Short-Term Memory (LSTM) network. These modules provide configurable hidden state management, activation functions, and support for batch-first or sequence-first input layouts, enabling users to implement sequence modeling tasks directly within the framework.
crates/burn-nn/src/modules/rnn · high confidence
Add SafeTensors store for model persistence
Introduces the \SafetensorsStore\ in \crates/burn-store/src/safetensors\, enabling models to be saved to and loaded from the SafeTensors format. This store supports both file-based and in-memory operations, offering zero-copy tensor access, lazy loading, and efficient memory usage. Users can configure tensor filtering via regex or exact paths, remap tensor names for framework compatibility (e.g., PyTorch to Burn), and attach custom metadata. The implementation ensures atomic file writes and includes adapters for seamless integration with existing model architectures.
crates/burn-store/src/safetensors · high confidence
Add WGAN example for MNIST generation
A new Wasserstein GAN example has been added to the examples directory, providing a Burn implementation for generating MNIST digits. The example includes a generator and discriminator model, a training loop with weight clipping and RMSProp optimization, and utilities for saving generated images. Users can run the training or generation scripts using various backends (CUDA, WGPU, Tch, Flex) via Cargo.
examples/wgan · high confidence
Add advanced LSTM example with bidirectional and stacked variants
Introduces a new \modern-lstm\ example demonstrating an advanced Long Short-Term Memory (LSTM) implementation. This example features a custom \LstmNetwork\ that supports bidirectional and stacked configurations, utilizing combined weight matrices for input and hidden states, along with layer normalization and dropout for regularization. It includes complete training and inference pipelines (\lstm-train\ and \lstm-infer\) that support multiple backends (CUDA, WGPU, Tch, Flex) and use a synthetic sequence dataset to predict values based on previous inputs.
examples/modern-lstm · high confidence
Add configurable remote backend server example
The examples/server directory now includes a runnable server example that starts a WebSocket-based remote backend. Users can configure the listening port via the REMOTE\_BACKEND\_PORT environment variable (defaulting to 3000) and benefit from a new burn.toml configuration that enables compilation caching and sets specific logging levels for profiling, autotuning, and compilation.
examples/server · high confidence
Add custom CSV dataset example with in-memory and Polars backends
The \examples/custom-csv-dataset\ directory now provides a working example demonstrating how to implement the \Dataset\ trait for CSV data. It includes an in-memory approach using \InMemDataset\ with \serde\ for parsing, and an alternative \DataframeDataset\ implementation backed by Polars for efficient data manipulation. The example automatically downloads the diabetes dataset and shows how to load and access records using both methods.
examples/custom-csv-dataset · high confidence
Add custom CubeCL kernel example for fused matmul-add-relu
The \examples/custom-cubecl-kernel\ directory now contains a complete example demonstrating how to implement a custom backend extension in Burn. It includes a WGSL/CubeCL kernel for a fused matrix-multiply-add-relu operation, along with the necessary forward and backward (autodiff) implementations to support gradient computation. The example also provides a reference implementation and verification logic to ensure the custom kernel produces results consistent with standard tensor operations.
examples/custom-cubecl-kernel, examples/custom-wgpu-kernel · high confidence
Add custom image dataset training example with multi-backend support
A new example demonstrating how to train a model on a custom image dataset has been added. It configures an SGD optimizer with momentum and supports multiple backends: LibTorch (CUDA on Linux, MPS on macOS) and WGPU/Metal/Vulkan, allowing users to run the training on various hardware configurations.
examples/custom-image-dataset/examples · high confidence
Add custom learning strategy example
The \examples/custom-learning-strategy\ directory now contains a complete example demonstrating how to implement a custom learning strategy. It includes a CNN model for MNIST classification and a training script that uses a custom \SupervisedLearningStrategy\ to control the training loop, including epoch iteration, batch processing, and event processing.
examples/custom-learning-strategy · high confidence
Add custom training loop example for MNIST
The examples/custom-training-loop directory now contains a complete, runnable example demonstrating how to implement a custom training loop for the MNIST dataset. The example includes configuration management, data loading with a custom batcher, model initialization, and a manual training/validation loop that handles gradient computation and optimizer steps, serving as a reference for users building custom training workflows.
examples/custom-training-loop · high confidence
Add custom-renderer example demonstrating training progress logging
The examples/custom-renderer directory now includes a new example that demonstrates how to implement and use a custom metrics renderer for training. The example provides a \CustomRenderer\ struct that implements \MetricsRenderer\, \MetricsRendererTraining\, \MetricsRendererEvaluation\, \TrainingProgressLogger\, and \EvaluationProgressLogger\ traits, allowing users to customize how training and evaluation progress is logged (currently printing item counts and events to stderr via \dbg!\). It shows how to integrate this custom renderer into a supervised training loop using \SupervisedTraining::renderer\.
examples/custom-renderer · high confidence
Add median and variance statistics operations
The tensor statistics module now includes functions to compute the median of a tensor along a specified dimension, following PyTorch's behavior for both odd and even element counts. Additionally, variance calculation has been expanded with new functions (\var\, \var\_with\_mean\, \var\_bias\, \var\_with\_mean\_bias\, and \var\_with\_mean\_n\) that allow for more flexible computation of variance, including options for biased and unbiased estimators and pre-computed means.
crates/burn-tensor/src/tensor/stats · high confidence
Add multi-device training strategies with gradient accumulation and sharded optimization
Users can now train models across multiple devices using two distinct optimization strategies: \OptimMainDevice\, where gradients are accumulated and applied on a main device, and \OptimSharded\, where gradients are accumulated per device and applied via a multi-device optimizer step. The training loop supports gradient accumulation to simulate larger batch sizes, handles data splitting across devices, and ensures model parameters are synchronized back to the main device for validation after sharded training steps.
crates/burn-train/src/learner/supervised/strategies/multi · high confidence
Add simple regression example with California Housing dataset
A new example demonstrating how to define a custom dataset, create a data pipeline with min-max feature scaling, and train a simple neural network for regression using the California Housing dataset. The example includes source files for dataset loading (via HuggingFace datasets), model definition, training loop, and inference visualization, supporting CPU and GPU backends (Flex, Tch, Wgpu).
examples/simple-regression · high confidence
Add simple regression example with multi-backend support
A new regression example has been added to the examples directory, demonstrating how to train a model and perform inference using the \burn\ library. The example is designed to work across multiple hardware backends, including CPU and GPU options via \flex\, \tch\ (LibTorch), and \wgpu\, selected via compile-time features. It utilizes the high-level \Device\ struct to abstract backend-specific device selection, allowing users to easily switch between different execution environments.
examples/simple-regression/examples, examples/wgan/examples · high confidence
Add tensor grid utilities: meshgrid and affine\_grid\_2d
The \burn::grid\ module now exposes \meshgrid\ and \meshgrid\_stack\ for generating coordinate matrices from 1D tensors, supporting both dense and sparse output modes as well as matrix (ij) and Cartesian (xy) indexing conventions. Additionally, \affine\_grid\_2d\ has been added to generate transformed coordinate grids for 2D image processing, accepting a batched transformation matrix and output dimensions to produce a grid suitable for affine transformations.
crates/burn-tensor/src/tensor/grid · high confidence
Add text classification example with LoRA finetuning support
The \examples/text-classification\ directory now contains a complete example for training and inferring text classification models on the AG News and DbPedia datasets using the Burn deep learning library. The example supports multiple backends (Torch GPU/CPU, Flex, WGPU, CUDA, Metal) and includes a specific workflow for finetuning pre-trained models using Low-Rank Adaptation (LoRA). It demonstrates data loading via HuggingFace datasets, tokenization with BERT, and model training with metrics like accuracy and loss.
examples/text-classification · high confidence
Add text generation training example
A new example demonstrating text generation training has been added, allowing users to train a Transformer model on the DbPedia dataset. The example supports multiple hardware backends including CUDA, ROCm, wgpu, Metal, and CPU via LibTorch, as well as a remote server mode, and allows configuration of the floating-point precision (f32, f16, or flex32) through feature flags.
examples/text-generation/examples · high confidence
Add text-generation example with GPT-2 and DbPedia dataset
This change introduces a new text-generation example that demonstrates training a Transformer Encoder model on the DbPedia dataset using a GPT-2 tokenizer. The example includes a complete training pipeline with data batching, model definition (including token and position embeddings, attention masks, and cross-entropy loss), and a training loop configured with the Adam optimizer, Noam learning rate scheduler, and metrics such as accuracy, perplexity, and loss. It also provides configuration files and documentation for running the example on CUDA and Mac environments.
examples/text-generation · high confidence
Add utility to save tensors as images
A new \save\_tensor\_as\_image\ function has been added to the \burn-vision\ crate, allowing users to export tensor data directly to image files. The utility supports various dimension orders (HWC, CHW, NHWC, etc.) and offers configuration options for color mapping (RGB or monochrome with min/max scaling) and batch handling (tiled or aggregated layouts).
crates/burn-vision/src/utils · high confidence
Atomic file writes and streaming tensor serialization in burn-pack
The burn-pack crate now ensures that saving model files is atomic: writes occur to a temporary scratch file and are only renamed to the final destination once complete, preventing partial or corrupted checkpoints if the process fails mid-write. Additionally, the writer supports streaming tensors on demand via deferred providers, allowing models larger than available host memory to be saved by materializing tensor data one at a time rather than holding the entire model in memory. The reader also supports lazy, zero-copy loading of tensor data from files, reading bytes only when accessed.
crates/burn-pack/src · high confidence
CPU backend implementation for vision operations
The CPU backend for burn-vision now provides native implementations for key computer vision operations, including Non-Maximum Suppression (NMS) with SIMD acceleration, connected components analysis with optional statistics, and basic morphology operations. This allows users to run these vision-specific tasks directly on the CPU without relying on external backends like LibTorch or Flex, improving availability and performance for CPU-only environments.
crates/burn-vision/src/backends/cpu · high confidence
Complete autodiff operation implementations across tensor types
The autodiff engine now provides full backward-pass implementations for activation functions (GELU, ReLU, Sigmoid, LogSigmoid), boolean tensor operations (AND, OR, XOR, NOT, scatter, gather, mask operations), integer tensor operations (arithmetic, comparisons, cumulative operations, matmul), quantized tensor operations (quantize, dequantize, reshape, permute), and distributed operations (all-reduce). This change ensures that gradients are correctly computed for these previously unimplemented or partially implemented operations, enabling end-to-end training with boolean, integer, and quantized tensors, as well as distributed training scenarios.
crates/burn-autodiff/src/ops · high confidence
Dataset module restructured with Polars DataFrame support and Turso-backed SQLite
The dataset module has been reorganized into a modular structure (base, error, iterator, in\_memory, sqlite, dataframe) and now supports Polars DataFrames via the new \dataframe\ feature, allowing datasets to be backed by Polars DataFrames. Additionally, the SQLite dataset implementation has been migrated from rusqlite to Turso, providing a Rust-native SQLite engine while maintaining compatibility with existing database files.
crates/burn-dataset/src/dataset · high confidence
GPU-accelerated connected components for Cube backend
The Cube backend now includes a hardware-accelerated implementation for connected components analysis, utilizing custom CUDA kernels for 4-connected and 8-connected labeling. This change introduces GPU-based strip labeling and prefix sum operations to significantly speed up image processing tasks compared to CPU-only fallbacks.
_crates/burn-vision/src/backends/cube/connected\components · high confidence
Initial release of the Burn Book documentation site
The Burn Book documentation site is now available, providing a structured guide for users. This initial release includes configuration for the mdBook tool (book.toml), a .gitignore file to manage build artifacts and IDE files, and a Prettier configuration for consistent formatting. The site is authored by the core team and supports mathematical notation via MathJax.
burn-book · high confidence
Initial setup of the Burn Contributor Book
Added the foundational structure for the Burn Contributor Book, including configuration for mdBook (book.toml), formatting rules (Prettier), and license symlinks. This establishes the project as a standalone documentation site authored by the Burn community.
contributor-book · high confidence
Introduce Burnpack format with half-precision and adapter support
The \burn-store\ crate now uses the Burnpack format (\.bpk\) as its primary storage backend, replacing the previous \Record\ type. This new format supports atomic writes, streaming tensor writes, and includes a \HalfPrecisionAdapter\ that allows users to save models in F16 to reduce file size while automatically loading them back into F32. The system also introduces a \ModuleAdapter\ trait with chaining capabilities, enabling transformations like dtype casting and parameter renaming during save and load operations. An example tool \burnpack\_inspect\ is provided to help users examine the binary structure of these files.
crates/burn-store/src · high confidence
Introduce Iroh as a new transport backend for Burn Remote
Burn Remote now supports the Iroh networking stack alongside the existing WebSocket transport, enabling remote compute sessions and authenticated tensor transfers over QUIC. This change adds a new \iroh\ transport module that manages peer identities, multiplexes sessions over shared connections, and handles tensor movement with per-peer authorization. To work around a known kernel issue (iroh\#4555), the Iroh server explicitly disables Generic Segmentation Offload (GSO) in its QUIC transport configuration.
crates/burn-remote/src/transport · high confidence
Introduce \`burn-backend-extension\` crate for runtime backend dispatch and custom operations
The \burn-backend-extension\ crate is introduced to provide procedural macros (\\#\[backend\_dispatch\]\ and \\#\[backend\_extension\]\) that generate runtime dispatch for custom backend operations. This allows users to implement extension traits on specific backends (such as Cube, Flex, NdArray, LibTorch, and Remote) and expose them through a unified \Dispatch\ wrapper. The crate supports advanced features including automatic backend selection based on tensor inputs, autodiff context merging, and lazy fusion execution with metadata generation. It also enables struct and enum inputs/outputs via \ExtensionType\ derivation and supports async functions and futures for backend interop.
crates/burn-backend-extension · high confidence
Introduce burn-cuda crate as a wrapper for the CubeCL CUDA backend
The new burn-cuda crate provides a thin abstraction layer over the CubeCL CUDA runtime, re-exporting CudaDevice and defining a Cuda type alias for the Cube backend. It includes a test suite that verifies support for a comprehensive set of data types (including F32, F16, BF16, various integer types, Bool, and quantized floats) via the supports\_dtype API, while explicitly noting that F64 is not currently supported.
crates/burn-cuda/src · high confidence
Introduce burn-flex CPU backend
Adds the \burn-flex\ crate, a new pure-Rust CPU backend for Burn that enables tensor operations on the CPU without C dependencies. This backend introduces \FlexDevice\ (displayed as 'Cpu'), \FlexTensor\ for type-erased byte storage with zero-copy views, and \FlexQTensor\ for quantized data. It includes optimized implementations for strided iteration, layout management, and joint loop-nest collapsing for binary operations to support SIMD acceleration and efficient in-place mutations.
crates/burn-flex/src · high confidence
Introduce burn-flex CPU backend with optimized tensor operations
The new burn-flex CPU backend provides a comprehensive set of optimized tensor operations for CPU execution. This includes activation functions (ReLU, GELU, Sigmoid, Softmax) with correct NaN propagation, attention mechanisms with both naive and flash attention strategies for different sequence lengths, and convolution operations (1D, 2D, 3D) with tiled im2col + GEMM approach and various fast paths for 1x1 convolutions, depthwise convolutions, and small-channel convolutions. Binary operations benefit from SIMD acceleration and in-place optimizations, while boolean operations support broadcasting and efficient comparison operations. The backend also implements concatenation, cumulative operations, and other essential tensor operations with performance optimizations like layout-aware processing and parallel execution where applicable.
crates/burn-flex/src/ops · high confidence
Introduce burn-ir crate for tensor operation intermediate representation
The new \burn-ir\ crate defines an Intermediate Representation (IR) for tensors and operations, enabling execution across different targets (such as remote backends) and allowing optimization and transformation of tensor computations before execution (e.g., operator fusion). It provides core types like \TensorIr\, \ScalarIr\, and \OperationIr\, along with a \GraphIr\ for managing operation graphs and boundaries, and a \HandleContainer\ for managing tensor handles and error propagation.
crates/burn-ir · high confidence
Introduce burn-nn crate with neural network modules, activations, and initializers
The \burn-nn\ crate is introduced to centralize neural network building blocks. It provides a comprehensive set of activation functions including CELU, ELU, HardShrink, SoftShrink, Shrink, Softsign, ThresholdedReLU, and SELU, alongside standard ones like ReLU and Sigmoid. The crate also exposes a robust \Initializer\ enum supporting Constant, Ones, Zeros, Uniform, Normal, Kaiming (Uniform/Normal), Xavier (Uniform/Normal), and Orthogonal initialization strategies. Additionally, it includes \PaddingConfig\ types for 1D, 2D, and 3D operations, enabling explicit, valid, and 'same' padding modes with support for asymmetric padding calculations. Autoregressive caching mechanisms for tensor sequences are also provided to support efficient inference in recurrent or transformer-based models.
crates/burn-nn/src · high confidence
Introduce burn-vision backend module with CPU and CubeCl support
The burn-vision crate now includes a backend module that exposes CPU-based vision operations and conditionally includes CubeCl support when the 'cubecl-backend' feature is enabled. This module provides core utilities like KernelShape and create\_structuring\_element, laying the groundwork for morphology operations in vision tasks.
crates/burn-vision/src/backends · high confidence
Introduce checkpointing infrastructure with file and async backends
The training module now includes a new checkpointing system that allows saving, restoring, and deleting model states (module, optimizer, and learning rate scheduler records) to disk. This change introduces a \FileCheckpointer\ for standard synchronous file-based storage and an \AsyncCheckpointer\ that offloads checkpoint operations to a background thread to prevent blocking the training loop. Users can now persist training progress to specific directories using burnpack files, with support for interrupting training if a checkpoint operation fails.
crates/burn-train/src/checkpoint · high confidence
Introduce configurable remote server with WebSocket and Iroh transports
The remote server now supports two transport protocols: WebSocket (defaulting to port 3000) and Iroh peer-to-peer, selected via a new \RemoteServerBuilder\. This builder allows users to configure the transport, register custom operations, and start the server either synchronously or asynchronously. The server now manages sessions with dedicated OS threads to prevent blocking the async runtime, handles same-host tensor transfers via a local rendezvous service, and ensures proper cleanup of device memory and session state upon client disconnect.
crates/burn-remote/src/server · high confidence
Introduce execution plan store and index for fusion optimization caching
The \burn-fusion\ crate now includes a dedicated store and index system for managing fusion optimization plans. This change adds \ExecutionPlanStore\ to cache and retrieve execution strategies (such as block optimizations or individual operation orderings) based on operation sequences, and \ExecutionPlanIndex\ to efficiently look up these plans using hashed operation identifiers. This infrastructure allows the fusion compiler to reuse previously explored and scored optimization paths, improving performance by avoiding redundant searches for identical operation streams.
crates/burn-fusion/src/stream/store · high confidence
Introduce global Dispatch backend for unified multi-backend execution
The \burn-dispatch\ crate now provides a global \Dispatch\ backend that acts as a single entry point for executing tensor operations across multiple underlying backends (such as CubeCL, Flex, NdArray, LibTorch, Remote, and Capture). This change introduces \DispatchDevice\ and \DispatchTensor\ types that route operations to the appropriate backend based on the device variant, enabling seamless cross-backend tensor transfers and unified autodiff support. Users can now select a specific hardware device (e.g., CUDA, CPU, WebGPU) via \DispatchDevice\ variants, and the system automatically handles the necessary backend-specific execution and data movement.
crates/burn-dispatch/src · high confidence
Introduce minimal non-generic record system for saving and loading module parameters
A new \ModuleRecord\ system has been added to \burn-core\ to handle the serialization and deserialization of module parameters using the burnpack format. This feature allows users to save module states to disk or memory buffers and load them back, with configurable options for dtype casting (\DTypePolicy\), partial loading, and validation. The implementation provides a straightforward traversal of module parameters via \ModuleVisitor\ and \ModuleMapper\, serving as the core persistence mechanism within the \burn-core\ crate, while more advanced features like cross-framework adapters remain in the separate \burn-store\ crate.
crates/burn-core/src/store · high confidence
Introduce multi-backend router extension
Added the \burn-router\ crate, a new extension that enables tensor operations to be forwarded across multiple backends. It provides a \BackendRouter\ implementation that routes operations via a configurable \RouterChannel\, supporting tensor transfers between backends through a \MultiBackendBridge\, client-side graph caching for fused operations, and a server-side \TensorInterpreter\ with a registry for custom operations.
crates/burn-router · high confidence
Introduce multi-stream fusion execution with cross-stream tensor sharing
The fusion runtime now supports concurrent execution across multiple streams. This change introduces a \MultiStream\ manager that handles cross-stream tensor aliasing, ensuring that when a tensor is cloned or moved between threads, its backing buffer is correctly shared via reference counting and its producing operations are drained if necessary to materialize the handle. It also adds a \Context\ for managing relative graph tensor mappings and scalars, and includes optional memory leak detection via a background monitoring thread.
crates/burn-fusion/src/stream · high confidence
Introduce new CubeCl fusion engine with trace-based execution
The \burn-cubecl-fusion\ crate now provides a new fusion engine that generates and executes fused kernels via a trace-based approach. This engine supports both fuse-on-read and fuse-on-write strategies, allowing multiple compute-bound operations to be combined into a single kernel. It includes a new \TraceOperationFuser\ to translate operations into a \FuseTrace\, a \FuseTraceLauncher\ to manage execution, and comprehensive code generation for handling tensor I/O, quantization parameters, and dynamic layouts. This replaces the previous fusion implementation with a more flexible and performant architecture.
crates/burn-cubecl-fusion · high confidence
Introduce new Evaluator API for model evaluation
Added a new \Evaluator\ and \EvaluatorBuilder\ in \crates/burn-train/src/evaluator\ that provide a structured way to evaluate models on datasets. The \Evaluator\ handles running inference on one or multiple named data splits, processing items through the model, and managing interruptions. The \EvaluatorBuilder\ allows users to configure the evaluation by registering numeric and text metrics, setting up a custom renderer, enabling evaluation summaries, and attaching progress loggers. This new component integrates with the existing event processing and metric systems to provide detailed evaluation feedback.
crates/burn-train/src/evaluator · high confidence
Introduce reinforcement learning training infrastructure in burn-train
This change adds the core components for reinforcement learning training to the \burn-train\ crate. It introduces the \RLComponentsTypes\ trait to define the necessary types for environments, policies, and learning agents, along with the \RLCheckpointer\ for saving and loading policy and agent states. The \RLTraining\ struct provides a builder API to configure and launch RL experiments, supporting an off-policy learning strategy (\OffPolicyStrategy\) that handles multi-environment experience collection, replay buffer management, and evaluation loops. Additionally, it includes \EpisodeSummary\ for tracking cumulative rewards and episode lengths, enabling users to train and evaluate RL agents with integrated checkpointing and metric logging.
crates/burn-train/src/learner/rl · high confidence
Introduce remote backend with Iroh transport and op-graph caching
The \burn-remote\ crate now provides a peer-to-peer remote tensor execution backend. It uses Iroh (QUIC-based) as the primary transport, with an optional legacy WebSocket mode, allowing applications to run computations on remote devices. A key feature is client-side op-graph caching (fusion): recurring operation groups are cached on the server and invoked by ID, significantly reducing network traffic for repeated computations. The crate also includes built-in telemetry to log these network savings and supports custom operations hosted on the remote server.
crates/burn-remote/src · high confidence
Introduce standalone PyTorch checkpoint reader
A new \pytorch-reader\ crate provides a Burn-free, standalone library for reading PyTorch checkpoint files (\.pt\, \.pth\). It supports modern ZIP-based formats (PyTorch 1.6+), legacy pickle streams, early TAR archives, and plain pickles. The reader handles full-model saves by reconstructing the module's \state\_dict()\, supports lazy tensor loading to inspect metadata without reading bytes, and includes safety limits to prevent excessive memory allocation. It also features a nested value deserializer for extracting configuration data and an adapter system for module-specific transformations.
crates/pytorch-reader/src · high confidence
Introduce standalone burn-fusion crate for automatic operation fusion
The \burn-fusion\ crate is now a standalone library that provides automatic operation fusion for backends implementing the \FusionBackend\ trait. It introduces a \Fusion\ wrapper backend that intercepts tensor operations, queues them via a \GlobalFusionClient\, and executes them on a dedicated fusion server thread. This architecture enables deferred execution, allowing multiple operations to be combined into single kernels for improved performance. The crate includes support for custom operations via the \custom\ module, introspection capabilities through \FusionInspector\ for testing and leak detection, and an observer system (\FusionObserver\) to monitor registration and execution of fused blocks.
crates/burn-fusion/src · high confidence
Introduces CLI metrics renderer and automatic renderer selection
The training renderer system now includes a \CliMetricsRenderer\ that outputs plain-text progress logs to standard output, suitable for non-interactive environments. The \default\_renderer\ function automatically selects between the terminal UI renderer (when the \tui\ feature is enabled and stdout is a terminal) and the new CLI renderer (when the \tui\ feature is disabled or stdout is not a terminal), ensuring appropriate output formatting in all execution contexts.
crates/burn-train/src/renderer · high confidence
Introduces structured training and evaluation progress logging
The \burn-train\ logger module now includes dedicated traits (\TrainingProgressLogger\ and \EvaluationProgressLogger\) and a \ProgressSnapshot\ struct to track and report the lifecycle of training and evaluation runs. This allows users to monitor step-by-step progress within splits and epochs, as well as overall run completion, providing more granular visibility into the training process beyond simple metric values.
crates/burn-train/src/logger · high confidence
Introduces vision operations API with connected components, morphology, and NMS
The \burn-vision\ crate now exposes a public API for computer vision tasks, defining the core traits and options for connected components analysis, morphological operations (erosion/dilation with configurable border types and anchors), and non-maximum suppression (NMS) for object detection. This change establishes the \VisionBackend\ trait and associated extension points, allowing backends to implement these specific operations while maintaining compatibility with existing backend extension mechanisms like Fusion and Flex.
crates/burn-vision/src/ops · high confidence
Introduction of Burn Backend Dispatch crate with build-time backend detection
The new \burn-dispatch\ crate provides a multi-backend dispatch mechanism that forwards tensor operations to the appropriate backend. Its \build.rs\ script implements build-time detection of enabled backends (such as CUDA, CPU, WebGPU, etc.) via Cargo features, defining \backend\_enabled\ and \cube\_backend\ configuration flags to support both multi-backend and backend-free builds.
crates/burn-dispatch · high confidence
Introduction of Burn Remote shared protocol types and session management
The \burn-remote\ crate now includes a \shared\ module that defines the core application-layer protocol for remote backend communication. This introduces structured types for session lifecycle management (\SessionId\, \SessionInit\, \SessionInfo\), including versioning (\PROTOCOL\_VERSION\) to ensure handshake compatibility. It also defines the \Task\ enum, which specifies the units of work sent between client and server, such as registering operations (\RegisterOperation\), executing cached operation graphs (\RegisterAndExecuteGraph\, \ExecuteGraph\), and managing tensor data (\RegisterTensor\, \RegisterTensorRemote\). Additionally, it establishes security and routing mechanisms via \TransferCapability\ for tensor downloads and \RequestId\ for demultiplexing results back to the client.
crates/burn-remote/src/shared · high confidence
Introduction of Distributed Data Parallel (DDP) training strategy
Users can now train models across multiple devices using the Distributed Data Parallel (DDP) strategy. This new implementation in \burn-train\ spawns a dedicated thread for each device to run model replicas, synchronizing gradients between all peers via an \all-reduce\ operation after backward passes. The strategy designates the first device as the main worker responsible for validation and event processing (including UI updates), while secondary workers handle their assigned data shards. It supports gradient accumulation, learning rate scheduling per optimizer update, and integrates with existing checkpointing and early stopping mechanisms, requiring users to manage the collective configuration across nodes.
crates/burn-train/src/learner/supervised/strategies/ddp · high confidence
Introduction of MLIR-based CPU backend
The burn-cpu crate now provides a new CPU backend implementation based on MLIR (via CubeCL). This backend supports a wide range of data types including F64, F32, F16, various integer types (I64/I32/I16/I8, U64/U32/U16/U8), and quantized floats, while explicitly excluding Flex32, Bool, and BF16 due to LLVM dialect limitations. Users can now utilize this MLIR-driven CPU runtime for tensor operations within the Burn framework.
crates/burn-cpu · high confidence
Introduction of asynchronous environment runner for RL training
The RL training loop now supports an asynchronous agent-environment runner (\AgentEnvAsyncLoop\) alongside the existing synchronous base runner. This new component runs the environment interaction loop in a separate thread, communicating via channels to allow for non-blocking execution, which can improve performance when environment steps are slow or when parallelizing environment interactions. The runner handles state transitions, reward accumulation, and trajectory collection while respecting evaluation and deterministic mode configurations.
_crates/burn-train/src/learner/rl/env\runner · high confidence
Introduction of burn-core crate with documentation and license symlinks
A new burn-core crate has been added to the repository, providing the core traits and components for building and training deep learning models. The crate includes a README explaining its usage with the main burn library and noting support for no\_std environments via feature flags. License files (LICENSE-APACHE and LICENSE-MIT) are now symlinks pointing to the root directory licenses.
crates/burn-core · high confidence
Introduction of burn-rl crate for reinforcement learning
The new burn-rl crate provides reinforcement learning building blocks for Burn, accessible via the \burn::rl\ namespace. It introduces core abstractions such as the \Environment\ trait for defining RL environments and the \StepResult\ struct for handling environment steps, rewards, and terminal states. This functionality is opt-in and requires enabling the \rl\ feature flag alongside \train\ in the Burn configuration.
crates/burn-rl · high confidence
Introduction of burn-rl library for reinforcement learning
The new \burn-rl\ crate provides the foundational library for training reinforcement learning agents, exposing public modules for environments, policies, and transition buffers. This release establishes the core abstractions required for RL workflows within the Burn ecosystem.
crates/burn-rl/src · high confidence
Introduction of burn-std as the core shared library
The \burn-std\ crate is introduced to host the foundational types and utilities shared across the Burn ecosystem, including core definitions for shapes, indexing, and data types. This new location centralizes the global runtime configuration system (\BurnConfig\), providing structured settings for operation fusion, autodiff, and remote backends, along with a unified logging infrastructure. It also establishes the \TensorData\ API for host-side tensor representation, featuring typed views, robust conversion methods, and strict validation of byte lengths and element counts to ensure data integrity.
crates/burn-std · high confidence
Introduction of multi-device training step implementation
The supervised training step module now includes a new \MultiDevicesTrainStep\ implementation that enables training across multiple devices. This component manages a pool of worker threads, each assigned to a specific \Device\, to execute training steps in parallel. It handles the distribution of data loader batches to these workers, collects the resulting training outputs and progress metrics, and aggregates them into a unified result, thereby supporting multi-GPU or multi-device training workflows within the Burn framework.
crates/burn-train/src/learner/supervised/step · high confidence
Introduction of the activation module with core tensor activation functions
The \crates/burn-tensor/src/tensor/activation\ directory now contains the \base.rs\ module, which exposes a suite of element-wise activation functions for tensors, including ReLU, HardTanh, Tanhshrink, ReLU6, LeakyReLU, GELU (exact and approximate), PReLU, and Softmax. This change establishes the public API for these operations within the tensor crate, allowing users to apply these standard neural network activations directly to tensor data.
crates/burn-tensor/src/tensor/activation · high confidence
Introduction of the burn-cubecl JIT backend
The \burn-cubecl\ crate is introduced as a new generic backend that compiles operations just-in-time (JIT) to various shader language targets. This change adds the core backend implementation, including tensor primitives, device management, and kernel launches for operations like binary math, attention, and casting, effectively establishing the JIT execution path for the Burn framework.
crates/burn-cubecl · high confidence
Introduction of the burn-dataset crate
The \burn-dataset\ crate has been introduced to provide a dedicated library for creating and loading datasets. It exposes a modular structure including sources, transforms, and domain-specific modules for audio, vision, and NLP, gated by respective features. The crate re-exports network utilities from \burn-std\ and includes a HuggingFace source module when the \sqlite\ feature is enabled.
crates/burn-dataset/src · high confidence
Introduction of the burn-rocm backend for ROCm HIP runtime
A new \burn-rocm\ crate has been added, providing a backend that utilizes the ROCm HIP runtime. This allows users to execute computations on AMD GPUs by leveraging the \cubecl::hip::AmdDevice\. The backend requires ROCm version 6.2.2 or a compatible version, which can be specified via the \ROCM\_PATH\ or \CUBECL\_ROCM\_PATH\ environment variables.
crates/burn-rocm · high confidence
MNIST example now supports automatic backend selection via feature flags
The MNIST example now automatically selects the appropriate hardware backend (such as Flex, CUDA, Metal, Vulkan, WGPU, or remote) based on the enabled Cargo features, allowing users to run the example on different devices without modifying the source code.
examples/mnist/examples · high confidence
Major rewrite of derive macros for Modules, Configs, and Einsum
The \burn-derive\ crate has been completely rewritten to support new capabilities and improved code generation. The \Module\ derive macro now supports enum types alongside structs, allows generic fields to be explicitly marked with \\#\[module(skip)\]\ to exclude them from the module system, and generates \Display\ implementations that format module structures similarly to PyTorch. The \Config\ derive macro has been refactored to support both struct and enum configurations, automatically generating constructors, builder methods, and serialization logic. Additionally, a new \einsum\ macro has been added, which compiles Einstein summation equations into optimized tensor operations at compile time, and the \RecordState\ derive macro has been updated to handle new field types like \Vec\<Tensor\>\ and nested states for optimizer and scheduler serialization.
crates/burn-derive/src · high confidence
Module system overhaul with LoRA, QLoRA, and quantization support
The module system in \burn-core\ has been restructured to introduce first-class support for Low-Rank Adaptation (LoRA) and Quantized LoRA (QLoRA) fine-tuning, alongside general weight quantization capabilities. Users can now apply \Lora\ or \QLora\ adapters to freeze base weights and train only low-rank factors, with QLoRA specifically keeping base weights in a low-bit representation to reduce memory usage. The \Module\ trait now includes methods like \apply\_lora\, \apply\_qlora\, and \quantize\_weights\, allowing for targeted parameter group modifications. Additionally, the module display system has been enhanced with customizable formatting settings, enabling users to control the output of module debug prints, including parameter counts and attribute visibility.
crates/burn-core/src/module · high confidence
NdArray backend implements core tensor operations and new spatial layers
The NdArray backend now provides concrete implementations for a wide range of tensor operations, including activation functions (ReLU), boolean tensor logic (AND, OR, NOT), integer arithmetic, and basic math operations. It also introduces support for new spatial layers and transformations, specifically Adaptive Average Pooling (2D and 3D), standard Average Pooling (2D), Grid Sample (2D), and Nearest Neighbor Interpolation. These additions enable the backend to handle more complex neural network architectures and data augmentation pipelines previously unsupported or requiring external backends.
crates/burn-ndarray/src/ops · high confidence
New \#\[might\_panic\] test attribute for controlled panic handling
The backend test suite now includes a new procedural macro attribute, \#\[might\_panic\], which allows tests to declare expected panic messages. When a test marked with this attribute panics with a message matching the specified reason, the test is skipped rather than failing, enabling more robust testing of error conditions without breaking the test suite on expected failures.
crates/burn-backend-tests/src · high confidence
New DQN agent example for CartPole with flexible device selection
Added a new Deep Q-Network (DQN) agent example that trains on the CartPole environment. The example now includes a \burn.toml\ configuration file for CubeCL settings (such as autotuning and memory management) and a \select\_device\ function that automatically chooses the appropriate hardware backend (CUDA, ROCm, WGPU, Metal, LibTorch, or Flex) based on enabled features. The training loop utilizes the updated \RLTraining\ API with an explicit inference device and configurable off-policy parameters.
examples/dqn-agent · high confidence
New Hugging Face dataset loader with SQLite export
Users can now load datasets from Hugging Face directly into the Burn framework. This change introduces the \HuggingfaceDatasetLoader\, which downloads datasets using the Hugging Face \datasets\ library and exports them to a local SQLite database for efficient access. The loader supports configuration options for dataset subsets, authentication tokens, custom cache and data directories, and trusting remote code. It automatically handles Python environment setup (including optional virtual environments) and manages the conversion of Hugging Face dataset features into SQLite tables, preserving binary data for images and audio.
crates/burn-dataset/src/source/huggingface · high confidence
New Jupyter notebook examples for Burn
Added a new \examples/notebook\ directory containing Jupyter notebooks that demonstrate how to use the Burn deep learning framework in Rust via the Evcxr kernel. The included \README.md\ provides setup instructions for the Evcxr kernel and explains that the notebooks target the current workspace API, explicitly enabling the \flex\ backend and \autodiff\ feature where needed. The notebooks themselves cover basic tensor operations (creation, shape manipulation, indexing) and automatic differentiation (gradient computation, chain rule examples), serving as interactive documentation for users.
examples/notebook · high confidence
New MNIST example with data augmentation and training pipeline
The examples/mnist directory now contains a complete, runnable MNIST training example. It demonstrates defining a custom CNN model (ConvBlock with BatchNorm, MaxPool, and Gelu), building a data pipeline with a custom batcher and mapper, and applying random image augmentations (translation, shear, scale, rotation) via the Transform2D API. The training configuration uses AdamW with cautious weight decay, a composed learning rate scheduler (cosine annealing with linear warmup), and includes early stopping based on validation loss. The example also shows how to log metrics, save checkpoints, and evaluate the trained model.
examples/mnist · high confidence
New NLP dataset loaders: AG News and TextFolder
Added \AgNewsDataset\ and \TextFolderDataset\ to the \burn-dataset\ NLP module. \AgNewsDataset\ provides automatic download, extraction, and access to the AG News text classification dataset (120k training, 7.6k test samples) with thread-safe download locking. \TextFolderDataset\ enables loading text classification datasets from local directories, supporting automatic encoding detection (UTF-8, UTF-16, GB18030, GBK) and configurable file extensions. Both are exposed via the \burn-dataset::nlp\ module, with AG News gated behind the \builtin-sources\ feature.
crates/burn-dataset/src/nlp · high confidence
New Speech Commands dataset integration
Added a new \SpeechCommandsDataset\ in the audio module that loads the Google Speech Commands dataset (v0.02) from Hugging Face into a local SQLite database. This provides a ready-to-use dataset for audio classification tasks, exposing train, test, and validation splits with 37 distinct speech command classes (including silence and other words). The implementation handles WAV file decoding, normalizing audio samples to the \[-1.0, 1.0\] range, and mapping labels to a strongly-typed \SpeechCommandClass\ enum.
crates/burn-dataset/src/audio · high confidence
New WebSocket-based remote communication backend
The \burn-communication\ crate now provides a new WebSocket-based remote backend, enabling direct data transfer between distributed nodes without routing through a central client. This change introduces a canonical \Address\ type for network endpoints, a generic \Protocol\ trait, and a concrete \WebSocket\ implementation using \axum\ and \tokio-tungstenite\. Key behavioral improvements include robust connection handling that ignores keepalive frames, detection of vanished peers via TCP keep-alive probes, and an \external\_comm\ service that allows servers to expose and download tensors directly to one another.
crates/burn-communication/src · high confidence
New activation layers added to burn-nn
The \burn-nn\ crate now exposes a comprehensive suite of activation functions as configurable modules, including CELU, ELU, GELU (with tanh approximation), Hard Shrink, Hard Sigmoid, Hard Swish, HardTanh, Leaky ReLU, Log Sigmoid, Mish, PReLU, ReLU, ReLU6, RReLU, SELU, Shrink, Sigmoid, SiLU, Soft Shrink, Softplus, Softsign, SwiGLU, Tanh, Tanhshrink, Threshold, and Thresholded ReLU. These modules are unified under a single \Activation\ enum and \ActivationConfig\ for easy selection, and each includes configuration structs, forward-pass implementations, and unit tests.
crates/burn-nn/src/activation · high confidence
New and expanded training metrics for classification, NLP, and system monitoring
The \burn-train\ metric module now includes a comprehensive suite of new evaluation metrics and system monitors. For classification tasks, it adds Accuracy (with padding support), AUROC, AUC-PR, and F-beta scores, all supporting binary, multiclass, and multi-label configurations via macro/micro reduction. NLP and sequence tasks are covered by BLEU, Character Error Rate (CER), and Word Error Rate (WER) metrics. Additionally, system-level monitoring is introduced with metrics for CPU temperature, CPU usage, and CUDA device statistics (memory, utilization, and power draw).
crates/burn-train/src/metric · high confidence
New attention modules with configurable bias and masking utilities
The \burn-nn\ crate now exposes a complete attention module suite, including \MultiHeadAttention\, \CrossAttention\, and mask generation utilities. \MultiHeadAttention\ adds configuration flags to independently enable or disable bias for the query, key, value, and output linear layers, allowing users to match specific architectural requirements. \CrossAttention\ supports asymmetric input shapes (different model dimensions for query vs. context) and advanced attention mechanisms like Grouped Query Attention (GQA) and Multi-Query Attention (MQA) via the \n\_heads\_kv\ parameter. Additionally, new helper functions \generate\_autoregressive\_mask\ and \generate\_padding\_mask\ are provided to simplify sequence masking for training and inference.
crates/burn-nn/src/modules/attention · high confidence
New burn-linalg crate with linear algebra operations
The \burn-linalg\ extension crate has been introduced to host linear algebra operations previously located in \burn\_tensor::linalg\. This new crate provides implementations for matrix operations including LU, QR, and SVD decompositions, as well as utilities for determinants, traces, diagonal extraction, outer products, matrix-vector multiplication, and vector norms (including Lp norms with support for negative orders). It also includes a cosine similarity function that avoids denominator underflow by normalizing inputs separately. The crate enforces input validation through a dedicated check module and supports \no\_std\ builds.
crates/burn-linalg · high confidence
New burn-signal crate for signal processing operations
The \burn-signal\ extension crate has been introduced to provide signal processing capabilities, including real FFT (\rfft\), inverse real FFT (\irfft\), complex FFT (\cfft\), Short-Time Fourier Transform (\stft\), and inverse STFT (\istft\). It also includes windowing functions such as Blackman, Hamming, and Hann windows. The crate supports autodiff for FFT operations and provides backend implementations for Flex, CubeCL, and LibTorch, with custom operation support for remote execution and graph capture.
crates/burn-signal · high confidence
New burn-vision crate with color, filter, and morphology operations
The new \burn-vision\ crate introduces a suite of image processing operations, including color space conversions (RGB↔Grayscale, RGB↔HSV), 2D filtering (box blur, median blur), and morphological operations (erosion, dilation). It also provides utility structs for vision tasks like \Size\ and \Point\, and exposes tensor extension traits for connected components and non-maximum suppression (NMS). These operations are available for float, integer, and boolean tensors and require enabling a backend feature (e.g., \flex\, \wgpu\) alongside the \vision\ feature.
crates/burn-vision/src · high confidence
New capture backend for recording operation graphs
A new non-executing backend has been added to the \burn-capture\ crate that records Burn operation graphs into a \GraphIr\ structure instead of executing them. Users can now capture computational graphs by defining scopes via \CaptureDevice::capture\_scope\, which isolates operations and preserves initializers, returning a \CapturedGraph\ containing the ordered operations and concrete tensor values.
crates/burn-capture · high confidence
New checkpointing strategies: KeepLastN, Metric-based, and Composed
The checkpointing system now supports flexible, composable strategies for managing saved model states. Users can retain the last N checkpoints to minimize disk usage, automatically keep only the checkpoint with the best metric value (configurable by metric name, aggregation, direction, and data split), or combine multiple strategies using a builder pattern where deletions only occur when all constituent strategies agree. These strategies integrate with the event store to make decisions based on training progress and metrics.
crates/burn-train/src/checkpoint/strategy · high confidence
New custom-image-dataset example for CIFAR-10 training
Adds a new example demonstrating how to train a CNN on a custom image dataset using the ImageFolderDataset API. The example includes code to automatically download and unpack the CIFAR-10 dataset, defines a VGG-style CNN model, and implements a training loop with SGD, accuracy/loss metrics, and model checkpointing. It supports running on Torch GPU, WGPU, and Metal backends.
examples/custom-image-dataset · high confidence
New dataset transformation wrappers for composition, mapping, and windowing
The \burn-dataset\ transform module now exposes several new dataset wrapper types that allow users to compose, filter, and reshape data pipelines without modifying the underlying storage. \ComposedDataset\ enables concatenating multiple datasets into a single sequence, while \MapperDataset\ provides a lazy, type-safe way to transform items from an inner dataset. \PartialDataset\ allows slicing a dataset by index range and includes utilities to split datasets evenly or by batch chunks. \SelectionDataset\ and \ShuffledDataset\ offer index-based selection and shuffling capabilities, \SamplerDataset\ supports sampling with or without replacement using configurable random number sources, and \WindowsDataset\ creates overlapping sliding windows over data. These changes introduce new public APIs for dataset manipulation, enhancing flexibility in data loading and preprocessing.
crates/burn-dataset/src/transform · high confidence
New distributed training strategy API with DDP and multi-device support
The supervised learning strategies module has been refactored to introduce a new \ExecutionStrategy\ enum and \TrainingStrategy\ type, enabling users to configure single-device, multi-device, and Distributed Data Parallel (DDP) training modes. The new API exposes \MultiDeviceOptim\ to control whether optimization occurs on a main device or is sharded, and provides a \DistributedContext\ for DDP coordination. This change decouples the distributed backend requirements from the core training loop, allowing for more flexible configuration of distributed training workflows.
crates/burn-train/src/learner/supervised/strategies · high confidence
New einsum API with compile-time macro and runtime string parsing
Users can now perform Einstein summation operations using two new interfaces: the \einsum!\ macro for compile-time equation parsing and the \Tensor::einsum\ method for runtime string-based equations. This feature, implemented in the new \burn-einsum\ crate and demonstrated in \burn-tensor/examples/einsum.rs\, supports broadcasting, ellipses, diagonals, and multiple operands. The macro generates efficient left-to-right contraction stages at compile time, while the runtime API interprets the same equation plan, allowing for dynamic equation strings. Both interfaces share the same underlying parser and executor, ensuring consistent behavior across static and dynamic use cases.
crates/burn-einsum, crates/burn-tensor/examples · high confidence
New example for importing and converting model weights
The \examples/import-model-weights\ crate now provides a complete workflow for importing pre-trained model weights into Burn. Users can load weights directly from PyTorch (\.pt\) or Safetensors formats, or use the \convert\ binary to transform these formats into Burn's native Burnpack (\.bpk\) format for efficient loading. The example includes binaries for running inference on MNIST images using the imported weights and demonstrates the full pipeline from external format to native Burn model usage.
examples/import-model-weights · high confidence
New examples for HuggingFace and Speech Commands datasets
Added two new example programs to demonstrate dataset loading capabilities. The \hf\_dataset\ example shows how to load HuggingFace datasets into a SQLite store, including how to enable \trust\_remote\_code\ for datasets requiring custom scripts. The \speech\_commands\ example demonstrates loading and inspecting the Speech Commands dataset, provided the \audio\ feature is enabled.
crates/burn-dataset/examples · high confidence
New gradient checkpointing implementation for memory optimization
The autodiff module now includes a new gradient checkpointing system (in \crates/burn-autodiff/src/checkpoint\) that allows users to trade compute time for reduced memory usage during training. This system introduces a \CheckpointStrategy\ trait with implementations for \NoCheckpointing\ (default, no memory savings) and \BalancedCheckpointing\ (saves intermediate states for memory-bound operations). The implementation uses a \CheckpointerBuilder\ to record checkpointing actions during the forward pass and a \Checkpointer\ to efficiently retrieve or recompute node outputs during the backward pass using a topological sort that avoids redundant traversals. This enables more efficient training of large models by selectively recomputing intermediate activations instead of storing them all in memory.
crates/burn-autodiff/src/checkpoint · high confidence
New loss functions added to burn-nn
The \burn-nn\ crate now includes several new loss modules: BinaryCrossEntropyLoss (with label smoothing and class weights), CosineEmbeddingLoss, CrossEntropyLoss (with padding token exclusion, weights, and label smoothing), CTCLoss, GaussianNLLLoss, HingeEmbeddingLoss, HuberLoss, KLDivLoss, and LpLoss. These additions expand the available optimization criteria for classification, regression, and sequence modeling tasks.
crates/burn-nn/src/loss · high confidence
New neural network modules: CosineSimilarity, GaussianNoise, Dropout, Embedding, Fold4d, Identity, Linear, PairwiseDistance, PixelShuffle/Unshuffle, PositionalEncoding, and RotaryEncoding
The \burn-nn\ module collection now includes a comprehensive set of new layers and utilities. Users can compute similarity and distance metrics with \CosineSimilarity\ and \PairwiseDistance\, apply regularization and data augmentation via \Dropout\ and \GaussianNoise\, and handle sequence data with \Embedding\, \PositionalEncoding\, and \RotaryEncoding\. Standard transformations are available through \Linear\, \Identity\, \Fold4d\, and the \PixelShuffle\/\PixelUnshuffle\ pair for spatial rearrangement. All modules follow the standard configuration and initialization patterns, supporting features like negative dimensions and custom display formatting.
crates/burn-nn/src/modules · high confidence
New normalization modules and unified wrapper in burn-nn
The \burn-nn\ crate now includes dedicated modules for Batch, Group, Instance, Layer, RMS, and Local Response Normalization, each with configuration structs and forward passes. BatchNorm now supports a distinct training mode that updates running statistics only when the layer is not frozen, and Group/Layer norms widen internal accumulation to f32 when the input is f16 to prevent overflow. A new \Normalization\ enum and \NormalizationConfig\ provide a single abstraction to switch between these layers, and all modules include shape validation and custom display formatting.
crates/burn-nn/src/modules/norm · high confidence
New optimizer crate with gradient clipping and learning rate schedulers
The \burn-optim\ crate introduces core optimization building blocks, including gradient clipping (by value or norm, with F16/BF16 overflow protection) and a comprehensive set of learning rate schedulers: constant, linear, exponential, cosine annealing, Noam, step, sequential, and composed schedulers with configurable reduction strategies (average, sum, product). All schedulers support state serialization and deserialization via burnpack records, enabling checkpointing and resumption of training schedules.
crates/burn-optim · high confidence
New prelude module and backend extension support
Users can now import common types and macros via \burn::prelude::\*\ for a quicker start, and the \backend\ module exposes the \ExtensionType\ trait, enabling custom structs and enums to wrap tensor primitives in backend extensions with automatic dispatch routing and autodiff context merging.
crates/burn-core/src · high confidence
New remote MNIST inference example for the web
Added a new example that runs a MNIST classifier in the browser while offloading all tensor operations to a remote compute peer via Iroh. The browser holds only the model definition and weights, sending 28x28 inputs and receiving 10-class probabilities over an authenticated QUIC session, enabling GPU-backed inference without shipping native builds or exceeding browser memory limits.
examples/remote-inference-web · high confidence
New single-device supervised training strategy with gradient accumulation and checkpointing
The supervised learner now includes a new \SingleDeviceTrainingStrategy\ that manages the training and validation loop on a single device. This strategy supports gradient accumulation, allowing users to simulate larger batch sizes by accumulating gradients over multiple iterations before updating the optimizer and learning rate. It integrates with the event processing system to log training and validation progress, and ensures metrics are flushed before checkpoints or early stopping conditions are evaluated. The implementation handles dataset errors gracefully and supports interruption during training.
crates/burn-train/src/learner/supervised/strategies/single · high confidence
New tensor operations and distributed execution utilities
This change introduces several new capabilities to the tensor API. It adds a \DistributedContext\ and \all\_reduce\ function to manage multi-device synchronization and collective operations. New functional operations include \batch\_norm\ (with a training variant that returns statistics), \ctc\_loss\, \embedding\, \conv1d\, and \conv2d\. Additionally, a \check\_closeness\ utility is provided for debugging tensor comparisons, and a \cross\_entropy\_with\_logits\ loss function is added.
crates/burn-tensor/src/tensor · high confidence
New tensor shape and Einstein summation macros
The \burn-tensor\ crate now exposes \assert\_shape!\ and \debug\_assert\_shape!\ macros for compile-time rank and runtime dimension validation, alongside a new \einsum!\ macro for literal Einstein summation equations. These additions provide stronger type safety and performance for tensor shape checks and matrix operations directly within the tensor API.
crates/burn-tensor/src · high confidence
New vision metrics: Dice, SSIM, MS-SSIM, PSNR, and more
Added a suite of new image quality and segmentation metrics to the training loop. Users can now track the Dice-Sorenson coefficient for segmentation overlap, SSIM and multi-scale SSIM for structural similarity, PSNR for peak signal-to-noise ratio, as well as A-FINE, DISTS, LPIPS, and FID for advanced perceptual and generative quality assessment. These metrics are available in the \burn-train\ vision module and support configurable parameters such as pixel range, kernel size, and background inclusion.
crates/burn-train/src/metric/vision · high confidence
New xtask commands for books, remote backend validation, and build/test orchestration
The xtask tooling now includes several new subcommands to streamline development workflows. Developers can use the new \Books\ command to build or open the Burn Book and Contributor Book locally via mdbook. A \Remote\ command has been added to end-to-end validate the remote backend by spawning a server example and running backend tests against it. The \Build\ command now supports a \--ci\ flag to exclude unsupported crates during CI runs and handles complex no-std builds for various targets including ARM and WASM. The \Test\ command has been expanded to support CI-specific sharding (Backends, Crates, Examples) and runs tests for extension crates like burn-linalg and burn-signal alongside the main backend tests. Additionally, a \Validate\ command provides a fast pre-PR check sequence including formatting, linting, and no-std compatibility checks.
xtask · high confidence
Project governance and compliance documentation added
The repository now includes CITATION.cff for proper academic citation, a CODE-OF-CONDUCT.md establishing community standards, and a CONTRIBUTING.md detailing the contribution workflow and AI-assisted contribution policy. Additionally, license compliance is formalized with explicit LICENSE-MIT and LICENSE-APACHE files, a NOTICES.md file tracking third-party attributions, and a deny.toml configuration for dependency auditing.
(repo-wide) · high confidence
SIMD-accelerated operations for the ndarray backend
The ndarray backend now uses SIMD instructions to accelerate core tensor operations, including average pooling, max pooling, 2D convolution, binary element-wise operations, comparisons, and unary operations. This change introduces new optimized implementations in the \crates/burn-ndarray/src/ops/simd\ module, leveraging the \macerator\ library to dispatch vectorized code paths on supported architectures (x86, x86\_64, aarch64, wasm32, loongarch64). Users will see improved performance for these operations when running on compatible hardware, with the backend automatically selecting the accelerated path when conditions like tensor layout and data type allow.
crates/burn-ndarray/src/ops/simd · high confidence
Unified dispatch layer for backend operations
The \burn-dispatch\ crate now provides a centralized dispatch layer that routes tensor operations (activation, bool/int/float tensors, modules, quantized tensors, and distributed operations) to the appropriate backend implementation. This change introduces a new \Dispatch\ type that implements operation traits like \FloatTensorOps\, \ModuleOps\, and \DistributedOps\, allowing users to write backend-agnostic code that automatically selects the correct backend (e.g., Flex, NdArray, Cubecl) at runtime based on the device. The dispatch layer also handles cross-backend transfers and autodiff context promotion, ensuring that operations like \batch\_norm\, \conv2d\, and \all\_reduce\ work seamlessly across different backend configurations.
crates/burn-dispatch/src/ops · high confidence
Unified parameter system with lazy initialization, flags, and LoRA support
The module parameter system has been restructured to support lazy initialization, module-owned control flags, and low-rank adaptation (LoRA). Parameters now share a lazy initialization state across clones, ensuring the initialization function runs at most once and all clones resolve to the same value. A new \Flag\ type allows modules to hold runtime control state (like dropout training mode) that is preserved across validation and training transitions without being persisted in model records. LoRA adapters are now first-class citizens, attached as reparameterizations to frozen weights, with their trainable factors surfaced as regular parameters for optimizer and autodiff traversal. The system also introduces parameter groups for selecting parameters by path or regex, and running states for thread-safe statistics updates during the forward pass.
crates/burn-core/src/module/param · high confidence
Removals
NdArray backend deprecated in favor of burn-flex and CubeCL backends
The \burn-ndarray\ crate is now deprecated (since version 0.22.0) and will be removed in a future release. Users are advised to migrate to \burn-flex\ for pure-Rust CPU execution (supporting std, no\_std, and WebAssembly) or to CubeCL backends (such as \burn-cuda\, \burn-rocm\, \burn-wgpu\, or \burn-cpu\) for GPU acceleration. The codebase retains the implementation but marks it as legacy, with a migration guide available in the \burn-flex\ crate.
crates/burn-ndarray/src · high confidence
Architecture
Introduction of the burn-backend crate for core backend interfaces
The \burn-backend\ crate has been introduced to house the core backend interfaces and data structures for executing tensor operations. This new location defines the \Backend\ trait and its associated operations (such as \FloatTensorOps\, \ActivationOps\, and \DistributedOps\), along with the \BackendTypes\ associated types for device and tensor primitives. It also includes the \DeviceOps\ trait and a global registry for managing device settings, effectively centralizing the low-level backend abstractions previously scattered across other crates.
crates/burn-backend · high confidence
LibTorch backend tensor operations restructured into modular files
The tensor operations for the LibTorch (burn-tch) backend have been reorganized from a single monolithic file into distinct modules: \activation.rs\, \base.rs\, \bool\_tensor.rs\, \int\_tensor.rs\, \module.rs\, \qtensor.rs\, \tensor.rs\, and \transaction.rs\. This change improves code maintainability and separation of concerns by grouping operations by type (e.g., float, int, bool, quantized) and function (e.g., activations, modules like convolutions). The functionality itself remains consistent with the previous implementation, ensuring no behavioral changes for users.
crates/burn-tch/src/ops · high confidence
Tensor API restructured into modular source files
The tensor API implementation in \crates/burn-tensor/src/tensor/api\ has been reorganized from a single large file into distinct modules (e.g., \autodiff.rs\, \base.rs\, \bool.rs\, \cast.rs\, \einsum/\, \extension.rs\, \float.rs\). This change improves code maintainability and separation of concerns by grouping related functionality—such as automatic differentiation, type casting, boolean operations, and Einstein summation—into their own files, without altering the public API surface.
crates/burn-tensor/src/tensor/api · high confidence
Behavioural changes
Async metric processing for training and evaluation
The metric processor in burn-train now offloads event handling to dedicated background threads via \AsyncProcessorTraining\ and \AsyncProcessorEvaluation\. This ensures that metric computation, storage, and rendering do not block the main training or evaluation loops, improving responsiveness. The change introduces a new \ItemLazy\ trait to allow lazy synchronization of training and evaluation items before they are processed by the worker threads, and adds a \flush\ mechanism to guarantee that all queued events are processed before operations like checkpointing or early stopping proceed.
crates/burn-train/src/metric/processor · high confidence
Burn 0.22: Unified backend dispatch and modularized extension crates
This release introduces a new high-level \Device\ struct that replaces the previous generic \Tensor\ backend approach, allowing users to select hardware backends at runtime via \Device::default()\ or explicit selection. The framework has been restructured into modular extension crates, with linear algebra, signal processing, and neural network components now available under \burn-linalg\, \burn-signal\, and \burn-nn\ respectively, exposed through the main \burn\ crate. Several legacy backends have been deprecated or removed: \burn-ndarray\ and \burn-tch\ are deprecated in favor of \flex\ and other CubeCL backends, while \burn-candle\ has been removed entirely. A new \remote\ feature enables remote device execution over Iroh, and backend tracing is now opt-in via the \tracing\ feature flag. Users must explicitly enable backends (e.g., \wgpu\, \flex\) rather than relying on defaults, and \Device::default()\ will panic if no execution backend is available.
crates/burn/src · high confidence
Burn crate now uses symlinks for shared documentation and licenses
The burn crate directory now references the root-level README, Apache-2.0 license, and MIT license files via symlinks instead of containing local copies. This ensures that documentation and legal text remain consistent with the project root and reduces duplication.
crates/burn · high confidence
Deprecation of the burn-tch (LibTorch) backend
The burn-tch backend, which wraps the LibTorch library via the tch crate, is now deprecated as of version 0.22.0 and will be removed in a future release. Users are advised to migrate to actively maintained alternatives: CubeCL-based GPU backends (burn-cuda for NVIDIA, burn-rocm for AMD, or burn-wgpu for Metal, Vulkan, and WebGPU) or CPU backends (burn-cpu or burn-flex). The codebase now includes sample binaries (cpu.rs, cuda.rs, mps.rs) that demonstrate tensor operations using the deprecated LibTorch types, serving as validation examples while the backend remains available.
crates/burn-tch/src · high confidence
Fusion backend ops restructured into modular, IR-driven implementations
The \burn-fusion\ backend's operation implementations have been reorganized into distinct modules (activation, binary, bool\_tensor, distributed, int\_tensor, module, qtensor, tensor, transaction, unary) that execute operations via a unified IR-based client registration system. This change introduces a \NoOp\ placeholder for no-operation steps, implements distributed collective operations (like \all\_reduce\) within the fusion stream, and ensures that module operations (such as convolutions) are registered as IR nodes to preserve downstream fusion compatibility in backends like \burn-cubecl-fusion\. Users benefit from a more consistent and extensible fusion pipeline where tensor operations are explicitly described by IR and executed through a centralized client, improving maintainability and enabling more complex fused graphs.
crates/burn-fusion/src/ops · high confidence
Introduce new supervised training API with configurable strategies and checkpointing
The supervised learning module has been refactored to provide a new, builder-style API for configuring training runs. Users can now explicitly set the \TrainingStrategy\ (supporting both single-device and multi-device execution), customize checkpointing behavior via \CheckpointingStrategy\ (including keeping the last N checkpoints or saving based on metric thresholds), and register custom progress loggers. The default configuration retains backward-compatible behaviors such as saving the last two checkpoints and monitoring validation loss, but the new structure allows for greater flexibility in how training loops, metrics, and interruptions are handled.
crates/burn-train/src/learner/supervised · high confidence
Introduces a new bridge layer to decouple tensor kinds from backend primitives
A new \bridge\ module has been added to \burn-tensor\ to serve as an intermediary between the high-level tensor API and the backend dispatch system. This change introduces \BridgeTensor\ and \BridgeKind\ (supporting Float, Int, Bool, and QFloat variants) to wrap the underlying \DispatchTensor\, effectively keeping tensor kind tracking out of the backends and preventing the exposure of backend-level primitives in the public API. The module also defines a sealed \TensorKind\ trait with a closed set of types (Float, Int, Bool) and a \Kind\ enum, ensuring that the dispatch system operates on a fixed, predictable set of tensor kinds while allowing extension authors to construct bridge tensors via feature-gated methods.
crates/burn-tensor/src/bridge · high confidence
New adaptive pooling modules and enhanced average pooling behavior
The \burn-nn\ pooling module now includes new \AdaptiveAvgPool1d\, \AdaptiveAvgPool2d\, and \AdaptiveAvgPool3d\ layers for flexible output sizing. Additionally, \AvgPool1d\ and \AvgPool2d\ modules have been updated to support asymmetric padding and \ceil\_mode\, and they now correctly honor the \count\_include\_pad\ setting to match standard behavior where padding zeros are included in the average calculation.
crates/burn-nn/src/modules/pool · high confidence
New asynchronous policy inference with deterministic mode preservation
The \burn-rl\ crate now includes an asynchronous policy inference server (\async\_policy.rs\) that supports autobatching of agent requests. This implementation specifically preserves the deterministic mode for each individual request, ensuring that deterministic and stochastic inferences are processed correctly without cross-contamination, even when batched together. The policy traits in \base.rs\ define the interface for this behavior, including support for moving policies to specific devices and managing policy states.
crates/burn-rl/src/policy · high confidence
New autodiff graph structure with traversal and requirement tracking
The autodiff graph implementation has been restructured into a new modular design within \crates/burn-autodiff/src/graph\. This change introduces a \Node\ structure that tracks gradient requirements (\Grad\, \GradInBackward\, \None\) and computing properties, replacing the previous flat graph representation. A new \GraphTraversal\ module provides breadth-first search capabilities to validate ancestry and collect backward steps in the correct order, ensuring that gradients are computed only for nodes that require them. The \Step\ trait now includes methods for accessing distributed parameters and parent nodes, facilitating more complex distributed training scenarios. This refactoring supports better memory management and enables features like gradient checkpointing by explicitly managing node lifecycles and requirements.
crates/burn-autodiff/src/graph · high confidence
New dependency-graph algorithms for fusion block ordering
The fusion search now uses a dedicated dependency-graph toolkit to determine valid execution orders for fusion blocks. This change introduces a Directed Acyclic Graph (DAG) implementation that derives dependencies from data-flow hazards (read-after-write and write-after-read) rather than explicit declarations. It includes a topological sort that respects original program order for independent nodes, a reachability analysis to safely merge blocks without creating cycles, and a lifetime validator to ensure resources remain live throughout the execution sequence. These algorithms allow the fusion search to generate more efficient and correct execution plans by accurately modeling resource dependencies.
crates/burn-fusion/src/search/graph · high confidence
New fusion execution engine with policy-driven optimization and error-scoping
The \burn-fusion\ stream execution layer has been refactored into a new modular system located in \crates/burn-fusion/src/stream/execution\. This change introduces a \Policy\ component that manages execution plans and triggers, an \Explorer\ that handles beam-search optimization with configurable caps, and a \Processor\ that orchestrates the interaction between them. A new \WriteScope\ mechanism ensures that operation failures and panics correctly claim write sets, preventing data corruption in fused blocks. The system also includes a \Trace\ module for detailed fusion logging and a \Validator\ for checking operation compatibility, providing a more robust and observable fusion execution path.
crates/burn-fusion/src/stream/execution · high confidence
New fusion search algorithm for optimizing operation blocks
The fusion search logic in \burn-fusion\ has been refactored to introduce a new optimization strategy for grouping operations. This change implements a block-based search mechanism that evaluates dependency graphs to find the best fusion opportunities, settling unfusable heads of blocks when necessary to improve execution efficiency. Users will benefit from potentially faster inference and training times as the runtime more effectively combines operations into optimized execution strategies.
crates/burn-fusion/src/search · high confidence
New fusion search optimization engine with DAG-aware block merging
The fusion search system now uses a new optimization pipeline in \crates/burn-fusion/src/search/optimization\ that models operations as a dependency DAG. This allows the optimizer to safely reorder independent operations (e.g., interleaved blocks with no data dependencies) to find better fusion opportunities, merging them into composed strategies when possible. It also handles 'holes' in the execution stream by re-optimizing unresolved operations in a second pass, ensuring all operations are covered. The change includes logging for fusion levels and bounds the search complexity via a configurable max block limit.
crates/burn-fusion/src/search/optimization · high confidence
New modular training API with application logging and early stopping warmup
The \burn-train\ learner has been refactored into a new modular API. Users can now install a configurable application logger (e.g., to a file) via \ApplicationLoggerInstaller\ to capture training logs and panic information. The \MetricEarlyStoppingStrategy\ now supports a \warmup\_epochs\ configuration, allowing training to continue for a specified number of epochs before early stopping conditions are evaluated. Additionally, the crate introduces structured output types for classification, regression, and sequence tasks that adapt model outputs for metrics while keeping tensors on-device, and provides a \LearnerSummary\ to aggregate training, validation, and test metrics from artifact directories.
crates/burn-train/src/learner · high confidence
New thread-safe metric event store with aggregation capabilities
The metric store module has been refactored to introduce a new, thread-safe architecture for handling training metrics. This change adds a dedicated \EventStoreClient\ that communicates with a background worker thread, allowing metrics to be collected and queried concurrently without blocking the training loop. The store now supports aggregating numeric metrics (currently mean) across epochs and splits, and provides methods to find the best epoch based on metric direction (highest or lowest). This replaces the previous synchronous logging mechanism with a more robust, asynchronous event-driven system for metric management.
crates/burn-train/src/metric/store · high confidence
Optimized CPU connected components via Spaghetti algorithm
The CPU backend for connected component labeling now uses the Spaghetti algorithm, which generates optimized decision forests to process images in block-based passes. This change replaces the previous implementation with a faster, generated code approach that handles both 8-connectivity and 4-connectivity modes, significantly improving performance for vision tasks requiring component labeling.
_crates/burn-vision/src/backends/cpu/connected\components · high confidence
Redesigned TUI with mouse support and labeled metric groups
The terminal UI renderer has been rewritten to support mouse interactions for navigating metric tabs and selecting text metrics, and to display metrics grouped by user-provided labels. The layout now features a dedicated controls view, a status panel that shows training mode and event counters, and separate plots for recent and full history that handle numeric and text metrics with improved resolution and bounds calculation.
crates/burn-train/src/renderer/tui · high confidence
Redesigned autodiff graph and gradient management with distributed support
The autodiff engine has been refactored to use a new graph-based node system (\NodeRef\, \NodeId\) and a dedicated \Gradients\ container that supports hooks for distributed gradient synchronization. This change introduces \AutodiffTensor\ as the primary tensor type, managing node references and reference counts to ensure correct graph lifecycle, while the \backend.rs\ module now acts as a decorator implementing \AutodiffBackend\ to expose backward pass capabilities. The new architecture explicitly supports distributed training via \DistributedGradientRegistration\ and \GradSyncContext\, allowing for coordinated gradient reductions across devices during the backward pass.
crates/burn-autodiff/src · high confidence
Redesigned autodiff runtime with mutex-based graph management and improved memory cleanup
The autodiff runtime has been refactored to use a new \GraphMutexClient\ that manages computation graphs via mutex-protected synchronization, allowing multiple graphs to modify their data without blocking each other, which is essential for multi-device training. This change introduces a new memory management strategy (\GraphMemoryManagement\) that tracks node statuses and parent dependencies to more accurately identify and free unused roots and unavailable nodes, addressing previous memory leaks and graph cleanup issues. The runtime now explicitly rejects the reuse of consumed graphs to prevent undefined behavior, and includes better handling for distributed backward passes by validating that all distributed parameters use the same backend as the loss.
crates/burn-autodiff/src/runtime · high confidence
Refactored DataLoader architecture with persistent multithreaded workers
The data loading subsystem has been restructured to introduce a new \DataLoader\ trait and a \DataLoaderBuilder\ for constructing loaders, alongside a dedicated \BatchStrategy\ trait for flexible batching logic. A key behavioral improvement is in the multithreaded loader (\MultiThreadDataLoader\), which now maintains a persistent pool of worker threads across training epochs to prevent memory leaks and stabilize GPU streams, while also pre-shuffling datasets before splitting them among workers. The system now supports splitting a single loader across multiple devices via \split\_dataloader\ and provides a \Progress\ struct to track iteration status.
crates/burn-core/src/data/dataloader · high confidence
Refactored fusion operation queue with deferred memory management and safe execution fallbacks
The fusion operation queue in \burn-fusion\ has been restructured to improve safety and memory handling. The new \OperationQueue\ implementation introduces deferred tensor freeing, which delays the release of memory for tensors received from other threads until the next execution boundary, preventing interruptions during queue processing. Execution logic now includes robust fallbacks: if a cached optimization plan does not match the current stream's state (e.g., due to shape ID mismatches or operation index bounds), the system safely reverts to running operations in submission order rather than panicking. Additionally, unsafe code has been removed from this component, and the queue now explicitly tracks tensor references to ensure correct memory lifecycle management during fusion.
crates/burn-fusion/src/stream/queue · high confidence
Refactored tensor lifecycle and cross-thread synchronization in the fusion backend
The fusion backend's tensor handling has been restructured to improve determinism and safety during cross-thread operations. The \FusionTensor\ implementation now uses atomic reference counting to manage shared views across different execution streams, ensuring that tensor status (read-only vs. read-write) is accurately tracked based on usage. A new deferred drop mechanism prevents unsafe re-entries into the client during thread unwinding, ensuring that resource cleanup happens safely without causing deadlocks or aborts. Additionally, comprehensive tests have been added to verify the correctness of these synchronization primitives, particularly focusing on the timing of tensor releases and the behavior of fused kernels under various conditions.
crates/burn-fusion/src/stream/multi, crates/burn-fusion/src/tensor · high confidence
Remote backend client architecture refactored for lazy connection and transport abstraction
The remote backend client has been restructured to introduce a \RemoteClient\ handle that delegates to a \RemoteService\ running on a dedicated device-runner thread, enabling lazy connection establishment (the network handshake occurs on first use rather than at initialization) to avoid blocking on process-global locks. The transport layer is now abstracted via \RemoteEndpoint\, supporting both WebSocket and Iroh backends through a unified \SubmitChannel\/\ResponseChannel\ interface, while \OutgoingBatch\ implements configurable task-count and byte-threshold flushing to optimize data transfer performance.
crates/burn-remote/src/client · high confidence
Reorganized data module structure in burn-core
The data module in burn-core has been restructured to act as a central re-export hub. It now explicitly exposes the dataloader and dataset functionality (re-exported from burn\_dataset) under the 'dataset' feature, and network utilities (re-exported from burn\_std) under the 'network' feature. This change consolidates access to these components within the burn-core crate, aligning with the broader effort to move types and utilities to their respective dedicated crates.
crates/burn-core/src/data · high confidence
Restructure burn-train crate with new module organization and component traits
The burn-train crate has been reorganized into distinct modules (checkpoint, renderer, logger, metric, learner, evaluator) and introduces a new \components.rs\ file that defines the \LearnerModel\ trait, consolidating requirements for training and inference steps along with the core \Module\ trait. This change also exposes helper types for training and inference inputs/outputs and includes test utilities for classification metrics, reflecting a structural refactor to improve modularity and ease of use for defining learning models.
crates/burn-train/src · high confidence
Tensor operations are now implemented via a new internal bridge trait system
The tensor operation implementations in \crates/burn-tensor/src/bridge/ops\ have been restructured to use a new internal trait-based bridge (e.g., \BasicOps\, \BasicAutodiffOps\) that wraps backend dispatch calls. This change introduces a unified layer for tensor operations across different data types (float, int, bool) and separates autodiff logic, ensuring that all tensor math, comparison, and manipulation operations are routed through this new bridge interface rather than direct backend calls.
crates/burn-tensor/src/bridge/ops · high confidence
Tensor-backed replay buffer for RL transitions
The transition buffer in burn-rl now uses contiguous tensor storage instead of Vec-based collections, enabling efficient random sampling via tensor select operations. The new implementation lazily initializes storage on the first push, supports circular buffer behavior with a configurable capacity, and stores states, actions, rewards, and done flags as tensors on a specified device. This change improves performance for reinforcement learning workloads by leveraging tensor operations for batched data access.
_crates/burn-rl/src/transition\buffer · high confidence
Text classification examples updated to use new Device and ExecutionStrategy APIs
The text classification examples (ag-news and db-pedia) have been refactored to align with the framework's new high-level Device struct and ExecutionStrategy enum. Users can now select hardware backends (CUDA, ROCm, WGPU, Metal, LibTorch, Flex) via feature flags, and training jobs can be configured for single-device, multi-device sharded, or distributed data-parallel (DDP) execution. The examples also support remote inference and training via a WebSocket connection to a burn-remote server, allowing users to leverage external compute resources without local GPU requirements.
examples/text-classification/examples · high confidence
Transformer encoder and decoder modules now support configurable activation functions and layer normalization epsilon
The Transformer Encoder and Decoder modules in \burn-nn\ now allow users to configure the activation function used in the position-wise feed-forward network (defaulting to GELU) and the epsilon value for layer normalization (defaulting to 1e-5). This change introduces \ActivationConfig\ and \layer\_norm\_eps\ fields to the configuration structs for both encoder and decoder layers, enabling customization of these hyperparameters. Additionally, the position-wise feed-forward module marks its activation field with \\#\[module(skip)\]\ to maintain backward compatibility with saved records, noting that stateful activations like SwiGLU will not have their learnable parameters persisted.
crates/burn-nn/src/modules/transformer · high confidence
Wgpu backend now re-exports CubeCL types and exposes explicit dtype support checks
The \burn-wgpu\ crate has been refactored to serve as a thin wrapper around \burn\_cubecl\ and \cubecl\, re-exporting core types like \CubeBackend\, \CubeTensor\, and \WgpuRuntime\ directly from those dependencies. This change introduces explicit \supports\_dtype\ checks for the backend, allowing users to query which data types (such as F32, I64, or quantized formats) are supported on specific devices and graphics APIs (Vulkan, Metal, WebGPU) at runtime, rather than relying on compile-time assumptions. The module also clarifies that graphics API selection (e.g., Vulkan vs. Metal) is now a runtime dispatch decision handled by \AutoCompiler\ rather than a compile-time type distinction, simplifying the API surface for multi-platform GPU execution.
crates/burn-wgpu/src · high confidence
burn-tch backend deprecated with migration guidance
The burn-tch crate is now deprecated as of version 0.22.0 and will be removed in a future release. Users are advised to migrate to actively maintained backends: CubeCL GPU backends (burn-cuda, burn-rocm, burn-wgpu) for GPU acceleration, or CPU backends (burn-cpu, burn-flex) for portable pure-Rust execution. The README has been updated to reflect this deprecation and provide clear migration paths.
crates/burn-tch · high confidence
Fixes
Fix CUDA initialization issues in tch-rs by adding dependency stubs
This fix addresses CUDA initialization problems in the tch-rs backend by introducing two new C++ source files: \dummy\_cuda\_dependency.cpp\ and \fake\_cuda\_dependency.cpp\. The \dummy\_cuda\_dependency.cpp\ file explicitly links against PyTorch's CUDA components (such as cuBLAS and warp size) to ensure proper initialization, while the \fake\_cuda\_dependency.cpp\ file provides a no-op implementation for environments where CUDA is not required or available. This change ensures that the burn-tch crate can correctly handle CUDA dependencies during runtime.
_crates/burn-tch/src/cuda\hack · high confidence
Test coverage
Add PyTorch model compatibility tests for burn-store; Added CubeCL quantization backend tests; Added autodiff tests for tensor operations; Added backend tests for boolean tensor operations; Added backend tests for float tensor activation functions; Added backend tests for quantized tensor operations; Added backend tests for tensor statistics operations; Added backend-agnostic tests for integer tensor operations; Added common test utilities for vision transforms; Added integration tests for Iroh and WebSocket remote backends; Added integration tests for WebSocket server concurrency and disconnect handling; Added integration tests for multi-layer model serialization; Added integration tests for the burn-pack format; Added integration tests for training loop features; Added no-std integration tests for Burnpack, SafeTensors, and MNIST model; Added no-std tests for model storage and neural network modules; Added quantization accuracy and calibration tests; Added regression tests for burn-store save/load reliability and type safety; Added regression tests for deserializer robustness and memory safety; Added tensor backend tests for clone invariance, distributed operations, and multi-threading; Added test fixtures for new dataset formats and features; Added tests for PyTorch model loading and store configuration; Added tests for backend extension fusion, lazy parameter initialization, and module state transitions; Added tests for backend extension, profiling, and new extension crates; Added tests for einsum macros, shape assertion macros, and device settings isolation; Added tests for float and integer tensor primitive operations; Added tests for module quantization; Added tests for the SafeTensors model store; Added tests for vision operations; Backend test coverage for tensor modules; Expanded CubeCL backend test coverage; Expanded backend test coverage for autodiff, graph capture, and tensor operations; Expanded backend test coverage for float tensor operations; Migrate backend benchmarks to burn-backend-tests; New benchmarking suite for model loading and saving performance; Refactor backend tests to use centralized device configuration and reusable test modules.
Housekeeping
Added README and license symlinks for burn-wgpu crate; Burn Tensor crate structure and documentation updates; Documentation and license setup for burn-fusion; Standardize license and documentation for burn crates.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 72 → 67 (-5.1)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 83 → 83 (+0.0)
- Architecture 97 → 95 (-2.8)
- Maturity 69 → 69 (+0.0)
- Readiness 73 → 57 (-16.8)
- Security 68 → 74 (+6.5)
- Event Sourcing 100 → 100 (+0.0)
- Performance 100 (new)
Resolved (99)
- Change coupling: base.rs ↔ builder.rs (crates/burn-backend/src/backend/ops/modules/base.rs)
- Change coupling: builder.rs ↔ module.rs (crates/burn-ir/src/builder.rs)
- Change coupling: executor.rs ↔ optimization.rs (crates/burn-cubecl-fusion/src/engine/launch/executor.rs)
- Change coupling: module.rs ↔ module.rs (crates/burn-autodiff/src/ops/module.rs)
- Change coupling: module.rs ↔ module.rs (crates/burn-autodiff/src/ops/module.rs)
- Change coupling: optimization.rs ↔ base.rs (crates/burn-cubecl-fusion/src/optim/matmul/optimization.rs)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- Duplicated block (10 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (10 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (10 lines × 2) (xtask/src/commands/test.rs)
- Duplicated block (10–18 lines × 2) (crates/burn-cubecl/src/kernel/conv/conv_transpose2d/transpose_direct.rs)
- Duplicated block (11 lines × 2) (crates/burn-cubecl/src/kernel/conv/conv_transpose2d/transpose_direct.rs)
- Duplicated block (11–14 lines × 2) (crates/burn-cubecl/src/kernel/conv/conv_transpose2d/transpose_direct.rs)
- Duplicated block (13 lines × 2) (crates/burn-tensor/src/tensor/signal/stft.rs)
- Duplicated block (14 lines × 2) (crates/burn-train/src/learner/supervised/strategies/ddp/epoch.rs)
- Duplicated block (16 lines × 2) (crates/burn-train/src/learner/supervised/strategies/multi/epoch.rs)
- Duplicated block (17 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (23 lines × 2) (crates/burn-cubecl/src/kernel/conv/conv_transpose2d/transpose_direct.rs)
- Duplicated block (26 lines × 3) (crates/burn-ndarray/src/ops/base.rs)
- …and 79 more
New (114)
- Ambiguous retrieval semantics: consume takes a NodeRef while get/remove take AutodiffTensor. It is unclear if consume is a 'get and delete' operation or just a 'get' operation, and whether remove is 'get and delete' or 'delete only'. The naming convention for 'get vs remove vs consume' is inconsistent across the API surface.
- Change coupling: client.rs ↔ multi.rs (crates/burn-fusion/src/client.rs)
- Change coupling: float.rs ↔ int.rs (crates/burn-flex/src/ops/float.rs)
- Change coupling: module.rs ↔ module.rs (crates/burn-autodiff/src/ops/module.rs)
- Change coupling: module.rs ↔ module.rs (crates/burn-autodiff/src/ops/module.rs)
- ClassTooLong: NdArrayOps (crates/burn-ndarray/src/ops/base.rs)
- Duplicated block (11 lines × 2) (xtask/src/commands/test.rs)
- Duplicated block (11 lines × 3) (crates/burn-signal/src/functions/blackman_window.rs)
- Duplicated block (13 lines × 2) (crates/burn-signal/src/functions/stft.rs)
- Duplicated block (14 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (14 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (14 lines × 2) (crates/burn-train/src/learner/supervised/strategies/multi/epoch.rs)
- Duplicated block (15 lines × 2) (crates/burn-flex/src/ops/bool.rs)
- Duplicated block (18 lines × 4) (crates/burn-ndarray/src/ops/base.rs)
- Duplicated block (23 lines × 3) (crates/burn-autodiff/src/ops/tensor.rs)
- Duplicated block (27 lines × 2) (crates/burn-train/src/learner/supervised/strategies/ddp/epoch.rs)
- Duplicated block (37–38 lines × 2) (crates/burn-flex/src/ops/attention.rs)
- Duplicated block (5 lines × 2) (crates/burn-router/src/interpreter.rs)
- Duplicated block (5 lines × 4) (crates/burn-ndarray/src/ops/base.rs)
- Duplicated block (5 lines × 4) (crates/burn-ndarray/src/ops/base.rs)
- …and 94 more
Changes since last survey
- 145 commits — 71 feature/other, 74 fixes
By area
- (root) — 30 commits
- crates/burn-backend-tests — 27 commits
- burn-book/src — 10 commits
- crates/pytorch-reader — 10 commits
- crates/burn-flex — 7 commits
- crates/burn-remote — 7 commits
- crates/burn-core — 6 commits
- .github/workflows — 5 commits
- crates/burn-optim — 4 commits
- crates/burn-train — 4 commits
- crates/burn-autodiff — 3 commits
- crates/burn-backend — 3 commits
- crates/burn-cubecl-fusion — 3 commits
- crates/burn-pack — 3 commits
- crates/burn-store — 3 commits
- crates/burn-cubecl — 2 commits
- crates/burn-dataset — 2 commits
- crates/burn-linalg — 2 commits
- crates/burn-tensor — 2 commits
- examples/mnist-inference-web — 2 commits
Notable commits
- fix: chore(deps): trim zip default features and fix burn-dataset nlp feature (#5715)
- fix: fix(autodiff): correct repeat_dim gradient ordering (#5837)
- fix: fix(autodiff): preserve gradients for zero scalar exponents (#5692)
- fix: fix(backend)!: separate device settings queries from initialization (#5869)
- fix: fix(backend): match PyTorch pooling output size in ceil_mode (#5798)
- fix: fix(backends): correct average-pooling gradients with ceil mode (#5653)
- fix: fix(burn-dataset): cap MNIST item counts at the split size (#5667)
- fix: fix(burn-ndarray): sign(NaN) must be 0, not its hidden sign bit (#5665)
- fix: fix(burn-store): scope contiguous index mapping per prefix (#5750)
- fix: fix(ci): publish burn-einsum and include it in no-std checks (#5727)
- fix: fix(ci): use cargo info to check published crate versions (#5777)
- fix: fix(core): count a module's parameters without initializing them (#5746)
- fix: fix(core): let an init_mapper parameter train on an autodiff device (#5705)
- fix: fix(cubecl): handle broadcasting in mask_where and mask_fill (#5767)
- fix: fix(cubecl): propagate NaN through two-sided clamp (#5718)
- fix: fix(cubecl): reject reshape that splits a quantization block across rows (#5844)
- fix: fix(cubecl): same-runtime device moves without a peer transport, and quantized moves (#5678)
- fix: fix(cubecl): update dependencies to fix empty tensor readback (#5704)
- fix: fix(deps): avoid enabling CubeCL through linalg defaults and respect vision defaults (#5694)
- fix: fix(deps): restore gpu-allocator's windows version to match wgpu-hal (#5867)
- …and 125 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
tracel-ai/burn was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3bf6e039482de4e0cb68fc370d4c4e48b767a737 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.