huggingface/candle
49.7
Weak · 27 September 2026
154k
lines of production code
Rust
primary language
4
measurements over time
What this system is
Candle is a minimalist machine learning framework for Rust that provides a core tensor library with hardware-accelerated backends for CPU, CUDA, and Metal. It supports a wide range of neural network operations, including optimized attention mechanisms and quantized inference for efficient model execution. The system enables users to load, run, and train diverse models—such as LLMs, vision transformers, and diffusion models—directly from Hugging Face or ONNX formats.
How it got here
2023 — core library release and model expansion
84 changes.
This period focused on the initial release of the Candle framework's core crates, including candle-core, candle-nn, and candle-transformers, establishing the foundational tensor operations and neural network layers. The work significantly expanded the project's capabilities by adding support for a wide variety of models such as LLaMA, Whisper, Stable Diffusion, and BERT, alongside the introduction of Python bindings and WebAssembly examples.
2024 — multimodal expansion and backend refactoring
49 changes.
This period focused on significantly expanding the library's multimodal capabilities by adding support for a wide array of vision, audio, and text models, including Flux, Stable Diffusion 3, and various LLMs. Concurrently, the core tensor infrastructure underwent major refactoring across CPU, CUDA, and Metal backends to improve performance, correctness, and maintainability through modularization and optimized kernel integrations.
2025–2026 — MoE, Flash Attention v3, and multimodal expansion
32 changes.
This period focused on integrating advanced architectural features, specifically Mixture-of-Experts (MoE) support and Flash Attention v3 for CUDA, while significantly expanding the framework's multimodal capabilities with new models for speech, audio, and vision. The Metal backend underwent a major refactoring to improve concurrency and kernel performance, complementing the addition of optimized CPU flash attention. Concurrently, the project broadened its example coverage to include a wide array of modern LLMs, quantized models, and specialized vision-language applications.
Features
Add BERT model implementation to Python bindings
The Python bindings now include a native implementation of the BERT architecture in \candle/models/bert.py\. This adds support for BERT-specific components, including embeddings, self-attention mechanisms with attention mask handling, intermediate layers, and pooling, allowing users to run BERT models directly through the Python interface.
python · high confidence
Add BERT sentence embedding example for WebAssembly
Introduces a new Candle WASM example that runs BERT models in the browser via Web Workers. The package includes a Rust binary (m.rs) exposing a Model class for loading weights and computing normalized sentence embeddings, a JavaScript worker (bertWorker.js) to handle inference off the main thread, and a demo UI (lib-example.html) that allows users to input text, calculate embeddings for sentences, and visualize cosine similarity scores for search queries using pre-configured models like intfloat/e5-small-v2.
candle-wasm-examples/bert · high confidence
Add BLIP image captioning example for WebAssembly
This change introduces a new Web-based demo for the BLIP image captioning model, allowing users to generate text descriptions for images directly in the browser. The example includes a Rust backend compiled to WebAssembly that supports both standard (safetensors) and quantized (GGUF) model variants, and a Vanilla JavaScript frontend that runs inference inside a Web Worker to keep the UI responsive. Users can load models from Hugging Face, upload images via file picker or drag-and-drop, and view the generated captions in real-time.
candle-wasm-examples/blip · high confidence
Add BLIP image captioning example with quantized model support
Users can now generate captions for input images using the Salesforce BLIP model via the new \blip\ example. The example supports both standard and quantized (GGUF) model variants, allowing for CPU-only inference or GPU acceleration via CUDA. It automatically downloads the model and tokenizer from Hugging Face, processes the input image, and outputs a text description.
candle-examples/examples/blip · high confidence
Add Based LLM example from Hazy Research
Users can now run the Based small language model, an experimental non-instruction-tuned LLM from Hazy Research that combines local and linear attention layers. This new example allows generation with three model variants (360m, 1b, 1b-50b) and supports standard generation parameters like temperature, top-p, and repeat penalty.
candle-examples/examples/based · high confidence
Add BigCode (StarCoder) code generation example
Users can now run the StarCoder/BigCode model for code generation using the new \bigcode\ example. This addition includes the main executable (\main.rs\) which loads the model via safetensors, supports CPU/GPU execution, and allows configuration of generation parameters such as temperature, top\_p (nucleus sampling), and seed. A corresponding README provides usage instructions and a sample output demonstrating the model's ability to generate code snippets.
candle-examples/examples/bigcode · high confidence
Add CIFAR-10 and Fashion MNIST vision datasets
The \candle-datasets\ crate now includes support for the CIFAR-10 and Fashion MNIST datasets alongside the existing MNIST dataset. Users can load CIFAR-10 images (32x32 RGB) and Fashion MNIST images (28x28 grayscale) directly via the \candle\_datasets::vision::cifar::load()\ and \candle\_datasets::vision::fashion\_mnist::load()\ functions, which fetch data from Hugging Face in Parquet format and return standardized \Dataset\ structs containing training and test tensors.
candle-datasets/src/vision · high confidence
Add CLIP image-text matching example
Users can now run a CLIP (Contrastive Language-Image Pre-Training) example that matches input images against text sequences. The new \candle-examples/examples/clip\ directory includes \main.rs\ and \README.md\, demonstrating how to load the \openai/clip-vit-base-patch32\ model and tokenizer, process images, tokenize text, and compute similarity probabilities via softmax. The example supports CPU execution and Metal acceleration on macOS, with default sample images and text sequences provided for quick testing.
_candle-examples/examples/chinese\clip, candle-examples/examples/clip, candle-examples/examples/mobileclip · high confidence
Add CSM conversational speech generation example
A new example has been added to generate conversational speech using the Sesame CSM model. Users can now run the \csm\ example to produce audio output from text prompts, supporting multi-speaker dialogue separated by the \\|\ character. The implementation loads the CSM-1b model, a Llama-3.2 tokenizer, and Mimi audio weights, allowing users to specify custom voice embeddings and generation parameters via command-line arguments.
candle-examples/examples/csm · high confidence
Add Chinese CLIP model support
Introduces a new implementation of the Chinese Contrastive Language-Image Pre-Training (Chinese CLIP) architecture, enabling users to perform image-text matching and retrieval using models trained on Chinese data. The change adds the \ChineseClipModel\ along with its text and vision transformer components (\ChineseClipTextTransformer\, \ChineseClipVisionTransformer\) and configuration structs, allowing for the loading of weights from the \OFA-Sys/chinese-clip-vit-base-patch16\ checkpoint and providing feature extraction capabilities for both text and image inputs.
_candle-transformers/src/models/chinese\clip · high confidence
Add CodeGeeX4-9B example for code generation
Users can now run the THUDM CodeGeeX4-9B model via a new \codegeex4-9b\ example, supporting both CPU and CUDA execution. The example allows prompting the model for code generation tasks (e.g., writing Rust functions) with configurable parameters such as temperature, top-p, repeat penalty, and sample length, and includes documentation on how to run and cite the model.
candle-examples/examples/codegeex4-9b, candle-examples/examples/starcoder2 · high confidence
Add ColPali PDF retrieval example
A new example demonstrating ColPali-based retrieval from PDF documents has been added to the examples directory. Users can now run the \colpali\ example to query a PDF file with a text prompt and receive the top-k most relevant page numbers, leveraging the ColPali model for vision-language matching.
candle-examples/examples/colpali · high confidence
Add ConvNeXt and ConvNeXt-V2 image classification example
Users can now run image classification inference using ConvNeXt and ConvNeXt-V2 models. This new example supports a wide range of model variants (Atto through Huge, including V2 versions) and allows selecting the specific model via the \--which\ argument. It loads pre-trained weights from the Hugging Face Hub and outputs the top-5 predicted classes for a given input image.
candle-examples/examples/convnext · high confidence
Add DINOv2-reg4 plant species classification example
A new example demonstrates using the DINOv2-reg4 model to classify plant species from the PlantCLEF2024 dataset. Users can run inference on an image to get probability scores for the top 5 species among 7,806 categories, with the example automatically downloading the pre-trained weights and class mapping file.
candle-examples/examples/dinov2reg4 · high confidence
Add DeepSeek V2 example
Users can now run the DeepSeek V2 Mixture-of-Experts model with Multi-Latent Attention via a new example. The example supports Lite (16B) and full (236B) model variants, offering context lengths of 32k and 128k tokens respectively, and allows selection between standard, chat, and coder-lite-chat configurations.
candle-examples/examples/deepseekv2 · high confidence
Add Depth Anything V2 monocular depth estimation example
This location introduces a new example for running the Depth Anything V2 model to perform monocular depth estimation from a single image. The \main.rs\ file implements the inference pipeline, loading the underlying DINOv2 vision transformer and the Depth Anything V2 head from Hugging Face, processing the input image, and saving the resulting depth map. A new \color\_map.rs\ module provides a \SpectralRColormap\ utility that allows users to visualize the depth output as a color-coded image via the \--color-map\ CLI flag.
_candle-examples/examples/depth\_anything\v2 · high confidence
Add DistilBert example for sentence embeddings and masked token prediction
A new DistilBert example has been added to the Candle examples suite, enabling users to compute sentence embeddings and perform masked token prediction. The example supports two model variants—DistilBertModel for embeddings and DistilBertForMaskedLM for token prediction—and allows users to specify custom model IDs and revisions from the Hugging Face Hub. It defaults to downloading safetensors weights but also supports PyTorch weights, and includes options for CPU execution, tracing, and controlling the number of top predictions for masked tokens.
candle-examples/examples/distilbert · high confidence
Add EVA-02 ImageNet classification example
A new example has been added to run the EVA-02 computer vision model as an ImageNet classifier. Users can now classify images by providing an image path via the \--image\ argument, with the example loading the model weights from the Hugging Face Hub by default and outputting the top 5 predicted categories with their probabilities.
candle-examples/examples/eva2 · high confidence
Add Falcon text generation example with configurable sampling and precision
Introduces a new Falcon example that allows users to generate text from the Falcon-7B model. The example supports CPU and GPU execution, with a specific \--use-f32\ flag required for CPU inference due to the lack of BFloat16 support. It includes configurable sampling parameters such as temperature, top-p (nucleus sampling), and repeat penalty to control output diversity and reduce repetition.
candle-examples/examples/falcon · high confidence
Add Flux image generation example
A new example demonstrating image generation using the Flux 12B rectified flow transformer has been added. Users can now run the example to generate images from text prompts, supporting both the 'schnell' and 'dev' model variants. The implementation includes support for quantized models, configurable image dimensions, and a seed option for reproducible results.
candle-examples/examples/flux · high confidence
Add Flux image generation model support
Added the Flux 12B rectified flow transformer model for text-to-image generation, including the main model implementation, an autoencoder for latent space processing, a quantized model variant for reduced memory usage, and the necessary sampling logic to run inference.
candle-transformers/src/models/flux · high confidence
Add GLM-4 quantized model example
Introduces a new example for running quantized GLM-4 models (specifically the 0414 variants) in GGUF format. The \quantized-glm4\ example supports multiple quantization levels (Q2\_K and Q4\_K\_M) for both 9B and 32B model sizes, allowing users to load models from local files or automatically download them from Hugging Face repositories (THUDM for tokenizers, unsloth for weights). It includes command-line arguments for configuring generation parameters such as temperature, sampling strategy, and context size, and provides a README with usage instructions.
candle-examples/examples/quantized-glm4 · high confidence
Add Gemma 3 support to the Gemma example
The Gemma example now supports Google's Gemma 3 models (specifically the 1B variants) alongside existing Gemma 1 and 2 models. Users can run inference on these new models by selecting the appropriate model variant (e.g., \3-1b\ or \3-1b-it\) via the command-line interface, which handles the necessary model loading and text generation logic for the new architecture.
candle-examples/examples/gemma · high confidence
Add Gemma 4 multimodal model support
Added a new implementation for Google's Gemma 4 multimodal model in the \candle-transformers\ library. This addition introduces support for processing text, images, and audio inputs. The implementation includes a text decoder with sliding window and global attention layers, a vision tower for image processing, and a conformer-based audio encoder. It also features a multimodal embedding system to project vision and audio features into the language model's embedding space, allowing the model to handle mixed modality inputs.
candle-transformers/src/models/gemma4 · high confidence
Add Gemma 4 text and multimodal generation example
A new example script for the Gemma 4 model has been added, supporting both text-only and multimodal inference modes. Users can now run generation with configurable parameters including temperature, top-p, top-k, and repeat penalty, while the example handles tokenization and output streaming for both model variants.
candle-examples/examples/gemma4 · high confidence
Add Granite 7b Instruct model example
Users can now run the IBM Granite 7b Instruct large language model using the Candle framework. This new example demonstrates loading the model weights and tokenizer from Hugging Face, configuring inference parameters such as temperature and sampling strategy, and generating text completions. The implementation supports CPU and GPU acceleration (via Metal or other backends) and allows customization of data types (f16, bf16, f32) and attention mechanisms.
candle-examples/examples/granite · high confidence
Add Helium-1 LLM example with Gumbel-Softmax sampling
Users can now run the Helium-1 2B parameter language model via the new \helium\ example. This example supports both the standard V1 architecture and the Helium-1 preview model, allowing inference on CPU or GPU. It introduces Gumbel-Softmax as a sampling strategy for text generation, in addition to existing methods like Top-K and Top-P.
candle-examples/examples/helium · high confidence
Add Hiera vision model example
Users can now run the Hiera hierarchical vision transformer example to perform image classification. This new example loads pre-trained Hiera models (Tiny, Small, Base, Base Plus, Large, Huge) from the timm library and returns the top-5 class probabilities for a given input image, supporting both CPU and GPU execution.
candle-examples/examples/hiera · high confidence
Add LFM2.5 example for LiquidAI models
Users can now run text generation examples for the LFM2.5 (Liquid Foundation Model 2.5) architecture from LiquidAI. This new example supports the LFM2.5-1.2B-Instruct and LFM2.5-1.2B-Thinking model variants, allowing generation with standard prompts or chat templates, and includes support for CUDA and flash attention via optional features.
candle-examples/examples/lfm2 · high confidence
Add MMDiT model support for Stable Diffusion 3 and 3.5
This change introduces the Mix of Multi-scale Dilated and Traditional Convolutions (MMDiT) transformer architecture, enabling inference for Stable Diffusion 3 and Stable Diffusion 3.5 models. The implementation includes the core model logic, attention and MLP blocks with AdaLN modulation, and embedding layers for patches, positions, timesteps, and vectors. It supports both the standard MMDiT configuration used in SD3 and the MMDiT-X variant used in SD3.5 Medium, allowing users to run these specific diffusion models within the Candle framework.
candle-transformers/src/models/mmdit · high confidence
Add Mamba2 inference example
Users can now run a Mamba2 inference example that implements the State Space Duality (SSD) framework. The new \mamba2\ example supports models ranging from 130m to 2.7b parameters (hosted under the \AntonV\ Hugging Face namespace) and includes options for CPU/GPU execution, prefill mode, and standard text-generation parameters like temperature and repeat penalty.
candle-examples/examples/mamba2, candle-examples/examples/mobilenetv4 · high confidence
Add Marian MT neural machine translation example
A new \marian-mt\ example has been added to the Candle examples, enabling neural machine translation for language pairs such as French-to-English, English-to-Chinese, English-to-Hindi, English-to-Spanish, English-to-French, and English-to-Russian. The implementation supports both \base\ and \big\ model variants, automatically downloads model weights and tokenizer files from the Hugging Face Hub, and includes a Rust-based SentencePiece tokenizer converter to eliminate Python dependencies. Users can run translations via the command line, specifying the text, model variant, and target language pair.
candle-examples/examples/marian-mt · high confidence
Add MetaVoice text-to-speech example
Introduces a new Rust example for the MetaVoice-1B text-to-speech model, allowing users to generate audio from text prompts. The implementation supports both standard and quantized (Q4K) model variants, configurable data types (F32, F16, BF16), and includes specific handling for Metal devices by offloading the Encodec decoder to the CPU. Users can run the example via \cargo run --example metavoice\ to produce WAV output files.
candle-examples/examples/metavoice, candle-examples/examples/parler-tts · high confidence
Add Mimi audio compression model with streaming support
Introduces the Mimi audio compression model, a state-of-the-art encoder/decoder architecture using residual vector quantization. This addition enables low-latency audio processing by supporting streaming encoding and decoding of audio tokens on the fly. The implementation includes components for convolutional operations (with weight and spectral norm support), SeaNet encoder/decoder layers, residual vector quantization, and a transformer with rotary positional embeddings and rotating KV cache.
candle-transformers/src/models/mimi · high confidence
Add Mimi audio tokenizer example with streaming support
Users can now run the Mimi audio compression model example to encode audio files into tokens, decode tokens back into audio, or perform real-time audio-to-audio conversion. The example supports streaming mode for low-latency interaction, handles audio input/output via system devices, and automatically resamples input audio to the required 24kHz sample rate.
candle-examples/examples/mimi · high confidence
Add Mistral 7B LLM example with quantized and streaming support
Introduces a new example for running the Mistral-7B model, supporting both standard and quantized (GGUF) variants. Users can now generate text with streaming output, exit on EOS tokens, and utilize flash attention for faster inference. The example includes configuration for multiple Mistral variants (v0.1, v0.2, instruct, mathstral, nemo) and provides a README with usage instructions.
candle-examples/examples/mistral · high confidence
Add Mixtral-8x7B example
Users can now run the Mixtral-8x7B-v0.1 sparse mixture-of-experts LLM via a new \mixtral\ example. This addition introduces a new CLI tool that loads the model from the HuggingFace Hub, handles tokenization and generation with configurable parameters (temperature, top\_p, repeat penalty), and outputs the generated text along with performance metrics.
candle-examples/examples/mixtral · high confidence
Add MobileOne image classification example
Users can now run the MobileOne model for image classification via a new example in the candle-examples suite. This addition includes a README with usage instructions and a main.rs implementation that loads a pre-trained MobileOne backbone (supporting variants S0–S4), processes an input image, and outputs the top-5 predicted classes with probabilities.
(repo-wide) · high confidence
Add ModernBERT fill-mask example
Users can now run a new example demonstrating the ModernBERT bidirectional encoder for the fill-mask task. The example supports both ModernBERT-base and ModernBERT-large models, automatically downloading weights from Hugging Face, and allows users to provide custom prompts or use default sentences to see masked token predictions.
candle-examples/examples/modernbert · high confidence
Add Moondream 2 multimodal model demo for WebAssembly
This change introduces a new Candle WebAssembly example for the Moondream 2 model, enabling users to run a multimodal AI assistant directly in the browser. The entry includes the Rust source code (src/bin/m.rs, src/lib.rs) that compiles the model to WASM, a JavaScript worker (moondreamWorker.js) to handle model loading and inference off the main thread, and a vanilla JavaScript UI (index.html, code.js) that allows users to upload images, enter text prompts, and adjust generation parameters like temperature and top-p. The demo supports both standard and quantized model variants, fetching assets from Hugging Face and caching them in the browser.
candle-wasm-examples/moondream · high confidence
Add MusicGen example for text-to-music generation
A new example has been added that implements the MusicGen model, allowing users to generate music from text prompts. The entry point (\main.rs\) loads the \facebook/musicgen-small\ model and tokenizer, processes a text prompt (defaulting to a '90s rock song'), and runs the text encoder to produce embeddings. The implementation includes the model architecture (\musicgen\_model.rs\) with support for sinusoidal positional embeddings, attention mechanisms, and configuration for the small model variant, along with a README providing usage instructions.
candle-examples/examples/musicgen · high confidence
Add NV-Embed-v2 inference example
A new example has been added for the NV-Embed-v2 text embedding model, enabling users to generate sentence embeddings or compute retrieval similarity scores. The implementation supports both CPU and GPU execution, allows loading models from local files or the Hugging Face Hub, and includes options for L2 normalization and tracing.
_candle-examples/examples/nvembed\v2 · high confidence
Add OLMo and OLMo 2 text generation example
Users can now run inference on OLMo and OLMo 2 models using the new \candle-examples/examples/olmo\ example. The entry point supports selecting between original OLMo and OLMo 2 architectures via the \--model\ flag (e.g., \1b\, \7b\, \1.7-7b\, \2-1b\), allowing generation with configurable temperature, top-p, and repeat penalty settings.
candle-examples/examples/chatglm, candle-examples/examples/olmo · high confidence
Add ONNX example supporting SqueezeNet, EfficientNet, and Real-ESRGAN
The ONNX example now supports three models: SqueezeNet and EfficientNet for image classification, and Real-ESRGAN for image super-resolution. Users can select the desired model using the \--which\ flag (defaulting to SqueezeNet) and provide an input image via the \--image\ flag. For classification models, the example prints the top 5 predicted classes with probabilities; for Real-ESRGAN, it saves the enhanced image with a 'super\_' prefix.
candle-examples/examples/onnx · high confidence
Add Orpheus 3B Text-to-Speech example
Users can now generate speech from text using the Orpheus 3B model via the new \candle-examples/examples/orpheus\ example. This addition introduces a CLI tool that loads the Llama-based Orpheus model and the Snac audio codec, supporting multiple voice profiles (Tara, Leah, Jess, Leo, Dan, Mia, Zac, Zoe) and configurable generation parameters like temperature and top-p. The example allows users to run inference on CPU or GPU (with CUDA support) and save the output audio to a WAV file.
candle-examples/examples/orpheus · high confidence
Add PaddleOCR-VL example for document parsing
Added a new example demonstrating the PaddleOCR-VL vision-language model, enabling users to perform text, table, formula, and chart recognition on images, process multiple images in a single prompt or batch mode, and extract text from video files. The example supports multilingual input, dynamic resolution handling, and configurable tasks via command-line arguments.
candle-examples/examples/paddleocr-vl · high confidence
Add Phi 1.5 and Phi 2.0 WebAssembly inference example
This change introduces a new WebAssembly-based example for running Microsoft's Phi 1.5 and Phi 2.0 language models directly in the browser. It includes a Rust implementation (\src/bin/m.rs\) that loads quantized GGUF and standard safetensor weights, a Web Worker (\phiWorker.js\) to handle model initialization and token generation off the main thread, and a Vanilla JS UI (\index.html\) supporting multiple model variants (including Puffin Phi v2) with configurable generation parameters like temperature and top-p.
candle-wasm-examples/phi · high confidence
Add Pixtral 12B multimodal example
Added a new example for running the Pixtral-12B text-and-vision model. The \main.rs\ implementation handles image encoding, constructs the specific input embeddings required by the model (including image breaks and special tokens), and performs text generation conditioned on the provided image. The \README.md\ provides usage instructions and links to the model card.
candle-examples/examples/pixtral · high confidence
Add Pixtral multimodal model support
Introduces a new implementation of the Pixtral architecture, enabling the model to process both images and text. This change adds the core model components in \candle-transformers/src/models/pixtral\, including the vision encoder (\vision\_model.rs\) for processing image patches, the multi-modal projector (\llava.rs\) for aligning visual features with text embeddings, and the Mistral-based language model for text generation. Users can now run Pixtral models (such as Pixtral-12B) for image understanding tasks.
_candle-transformers/src/models/pixtral, candle-transformers/src/models/qwen3\vl, candle-transformers/src/models/voxtral · high confidence
Add RecurrentGemma example with quantized model support
A new example for the RecurrentGemma 2B model has been added, allowing users to run text generation with both the base and instruct variants. The implementation supports CPU and CUDA execution, includes a quantized model variant for reduced memory usage, and enforces top-k sampling for text generation.
candle-examples/examples/recurrent-gemma · high confidence
Add Replit Code completion example
Users can now run the \replit-code\ example to perform code completion using the \replit-code-v1\_5-3b\ model. This new example supports both standard and quantized (GGUF) model variants, allowing users to generate code snippets from text prompts via the command line.
candle-examples/examples/replit-code · high confidence
Add SNAC audio tokenizer example with multi-rate support
The \candle-examples/examples/snac\ directory now contains a new example demonstrating the SNAC audio tokenizer. This tool supports three sample rates (24kHz, 32kHz, and 44kHz) and enables three workflows: converting audio files to SNAC codes, reconstructing audio from SNAC codes, and real-time audio-to-audio processing via microphone input and speaker output. The implementation handles audio resampling, file I/O, and model loading from the Hugging Face Hub.
candle-examples/examples/snac · high confidence
Add SPLADE sparse retrieval example
A new example demonstrating SPLADE, a neural retrieval model that generates sparse embeddings via the BERT masked-language-model head. Users can now compute sparse embeddings for a single query or calculate cosine similarities between multiple sentences using the \prithivida/Splade\_PP\_en\_v1\ model by default.
candle-examples/examples/splade · high confidence
Add Segformer example for image classification and semantic segmentation
A new example has been added to run Hugging Face Segformer models for both image classification and semantic segmentation. Users can now execute the example to classify images (defaulting to a food classification model) or generate segmentation masks (defaulting to an ADE20K fine-tuned model) by providing an input image path. The example includes a label mapping file for segmentation tasks and supports CPU execution.
candle-examples/examples/segformer · high confidence
Add Segment Anything (SAM) example with prompting and tracing support
This location introduces the Segment Anything Model (SAM) example, enabling users to generate image masks via point prompts (including negative/background points) or automatic mask generation. The example supports the default ViT backbone and the smaller TinyViT-based MobileSAM variant, allows threshold tuning for mask selectivity, and includes optional Chrome tracing for performance profiling. It outputs a merged image with the mask overlay and prompt markers.
candle-examples/examples/segment-anything · high confidence
Add Segment Anything Model (SAM) implementation
The Segment Anything Model (SAM) is now available in candle-transformers, providing a robust image segmentation pipeline that supports interactive prompting via points and bounding boxes. This release includes the full model architecture (image encoder, prompt encoder, and mask decoder) with support for both the original Vision Transformer backbone and the faster TinyViT backbone (MobileSAM). Users can now segment objects in images by specifying foreground or background points, with the option to generate multiple mask candidates or a single best mask.
_candle-transformers/src/models/segment\anything · high confidence
Add SigLIP multimodal example
Users can now run the SigLIP (Sigmoid Loss for Language-Image Pre-training) model via a new example in \candle-examples/examples/siglip\. This addition includes the main Rust application (\main.rs\) which supports multiple model variants (v1 and v2, base and large, with various patch sizes and image dimensions) and a README with usage instructions. The example demonstrates loading models from Hugging Face, processing images and text sequences, and outputting similarity probabilities, effectively adding a new multimodal text-vision capability to the Candle examples suite.
candle-examples/examples/llava, candle-examples/examples/paligemma, candle-examples/examples/siglip · high confidence
Add Silero VAD v5 voice activity detection example
Users can now run a new example that performs voice activity detection using the Silero VAD v5 model. The example loads the ONNX model from Hugging Face, processes streaming audio input (supporting 8kHz and 16kHz sample rates), and outputs real-time predictions indicating whether speech is present. It includes documentation on how to run the example using standard audio tools like arecord or SoX.
candle-examples/examples/silero-vad · high confidence
Add SmolLM3 inference example with unified quantized and full-precision support
A new example (\candle-examples/examples/smollm3\) has been added to run the SmolLM3 model using the Candle framework. This implementation provides a unified interface for both quantized (GGUF) and full-precision (safetensors) models, supporting multiple quantization levels (Q4\_K\_M, Q8\_0, F16) and data types (F32, F16, BF16). It includes features such as automatic model downloading from HuggingFace Hub, chat template support for instruction-tuned variants, and a 'thinking mode' for reasoning traces. The example handles SmolLM3's specific NoPE/RoPE architecture configuration and offers command-line options for generation parameters like temperature, top-p, and repeat penalty.
candle-examples/examples/smollm3 · high confidence
Add Stable Diffusion 3 and 3.5 example with Skip Layer Guidance support
This change introduces a new example for running Stable Diffusion 3 Medium and Stable Diffusion 3.5 variants (Large, Large Turbo, Medium) using the Candle framework. The implementation supports the MMDiT architecture and includes a triple text-encoder setup (CLIP-L, CLIP-G, and T5-XXL) to handle the specific tokenization and embedding requirements of these models. Users can now generate images using these newer models via the \--which\ argument, with specific defaults for inference steps and CFG scale per variant. Additionally, the example supports Skip Layer Guidance (SLG) for Stable Diffusion 3.5 Medium to improve image quality, and offers Flash Attention optimization for faster inference on compatible GPUs.
candle-examples/examples/stable-diffusion-3 · high confidence
Add StableLM example with quantized and flash-attention support
Users can now run StableLM models (including StableLM-2, Stable-Code, and Zephyr variants) via a new CLI example. The implementation supports both full-precision and quantized GGUF model formats, and includes an option to enable flash attention for improved performance. The example handles gated model access and provides configuration for sampling parameters like temperature and repeat penalty.
candle-examples/examples/stable-lm · high confidence
Add Stella\_en\_v5 embedding model example
This location introduces a new example for the Stella\_en\_v5 embedding model, supporting both the 1.5B and 400M variants. The implementation allows users to generate text embeddings with configurable dimensions (256 to 8192) and supports both similarity (s2s) and retrieval (s2p) tasks via command-line flags.
candle-examples/examples/stella-en-v5 · high confidence
Add TinyStories NLP dataset support
Introduces a new \tinystories\ module within the \candle-datasets\ crate, providing a \Dataset\ struct and iterator for loading the pre-tokenized TinyStories dataset. This implementation uses memory-mapped files to efficiently handle binary token data generated by llama2.c tools, enabling users to iterate over training and validation sequences for NLP model training.
candle-datasets/src/nlp · high confidence
Add TrOCR image-to-text transcription example
Users can now run a TrOCR (Transformer OCR) example to transcribe text from images. This new example includes a custom image processor for resizing and normalizing inputs, supports four model variants (base/large for handwritten and printed text) via the \--which\ flag, and handles tokenization and generation to output the transcribed text.
candle-examples/examples/trocr · high confidence
Add Vision Transformer (ViT) example
Users can now run a Vision Transformer classification example that loads the google/vit-base-patch16-224 model, processes an input image, and prints the top-5 predicted classes with probabilities. The example supports specifying a custom model file, selecting CPU or GPU execution, and includes a README with usage instructions.
candle-examples/examples/vit · high confidence
Add Voxtral speech recognition example
A new Voxtral speech recognition example has been added to the examples directory. This example allows users to transcribe audio files using the Voxtral model (defaulting to \mistralai/Voxtral-Mini-3B-2507\), supporting both CPU and GPU execution via the \--cpu\ flag. It includes automatic downloading of model weights and sample audio files, and provides command-line options to specify the input audio file and the model ID.
candle-examples/examples/whisper · high confidence
Add Whisper ASR model implementation with audio processing and quantized support
The \candle-transformers/src/models/whisper\ module now provides a complete implementation of the Whisper automatic speech recognition model. This includes \audio.rs\ for multithreaded log-mel spectrogram computation, \model.rs\ for the standard transformer architecture with KV-cache support, and \quantized\_model.rs\ for a quantized variant to reduce memory usage. The module exposes the model configuration, tokenizer constants, and inference logic, enabling users to convert audio files to text with language detection and multilingual support.
candle-transformers/src/models/whisper · high confidence
Add Würstchen v2 text-to-image example
A new example has been added to generate images using the Würstchen v2 model, ported from the Hugging Face diffusers implementation. Users can now run the example via \cargo run --example wuerstchen\ to produce images from text prompts, with automatic weight downloading from the HuggingFace Hub and support for custom weight files, resolution settings, and Flash Attention.
candle-examples/examples/wuerstchen · high confidence
Add YOLOv3 object detection example
Users can now run a YOLOv3 object detection example that loads model weights and configuration from the Hugging Face Hub by default, parses Darknet-style network definitions, and performs inference on images to detect and visualize objects such as people, bicycles, and motorbikes.
candle-examples/examples/yolo-v3 · high confidence
Add YOLOv8 example with object detection and pose estimation
A new YOLOv8 example has been added to the Candle examples suite, supporting both object detection and human pose estimation tasks. The implementation includes a configurable CLI to select model variants (n, s, m, l, x) and tasks, with output images annotated with bounding boxes, confidence legends, and keypoint connections. The example downloads weights from the Hugging Face Hub by default and can run locally or in the browser via WebAssembly.
candle-examples/examples/yolo-v8 · high confidence
Add Yi-6B and Yi-34B bilingual LLM example
Users can now run inference on the Yi family of bilingual (English, Chinese) large language models using the new \yi\ example. The implementation supports both the 6B and 34B model variants, allowing users to generate text via command-line arguments for prompts, temperature, and other generation parameters.
candle-examples/examples/yi · high confidence
Add example for running quantized Qwen3 MoE models
A new example has been added to demonstrate inference with quantized Qwen3 Mixture-of-Experts (MoE) models in GGUF format. Users can now run the \quantized-qwen3-moe\ example to load models from local files or automatically download them from Hugging Face (specifically the \unsloth/Qwen3-16B-A3B-GGUF\ and \unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF\ repositories). The example supports various quantization levels (Q2\_K, Q4\_K\_M, Q6\_K, Q8\_0) for both the 16B and 32B model variants, allowing users to select the desired model size and precision via the \--which\ argument.
candle-examples/examples/quantized-qwen3-moe · high confidence
Add llama2-c example with inference, evaluation, and training
The candle-examples/examples/llama2-c directory now contains a complete example implementation of the llama2.c model, providing users with three distinct capabilities: inference, evaluation, and training. The main entry point uses subcommands to manage these tasks, supporting features like nucleus sampling (top\_p), repeat penalties, and configurable batch sizes. The inference mode allows loading models from the Hugging Face Hub or local config files, while the evaluation mode supports pre-tokenized datasets for performance assessment. Additionally, a new training module enables users to fine-tune the model using the AdamW optimizer on pre-tokenized data, with automatic checkpointing in safetensors format.
candle-examples/examples/llama2-c · high confidence
Add minimal Mamba language model example
A new minimal example for the Mamba architecture has been added to the Candle examples suite. This standalone implementation, derived from the mamba-minimal Python reference, allows users to run inference with various Mamba model sizes (130m through 2.8b) directly via the command line. It provides a simpler, more transparent codebase compared to the existing Mamba example, facilitating easier experimentation and understanding of the model's core components.
candle-examples/examples/mamba-minimal · high confidence
Add multiprocess LLaMA example with tensor parallelism
Introduces a new \llama\_multiprocess\ example that runs LLaMA models (Llama-2 and Llama-3) across multiple GPU processes using NCCL for tensor parallelism. The implementation shards model weights across ranks, handles inter-process communication via a shared NCCL communicator ID file, and supports configuration of model variants, data types (F16/BF16), and generation parameters like temperature and top\_p.
_candle-examples/examples/llama\multiprocess · high confidence
Add preliminary ONNX support to Candle
The candle-onnx crate has been added to provide preliminary support for ONNX models. This includes a build script that compiles the ONNX protobuf definitions and documentation explaining the requirement for the protoc installation to compile the crate.
candle-onnx · high confidence
Add quantized LFM2 model inference example
A new example has been added to run quantized LFM2 models (350M and 2.6B variants) using GGUF weights. Users can now download and infer from models like \lfm2-350m-q4\_k\_m\ or \lfm2-2.6b-q8\_0\ via the Hugging Face Hub, with support for CPU and GPU (Metal) execution, configurable sampling parameters, and token streaming output.
candle-examples/examples/quantized-lfm2 · high confidence
Add quantized Qwen3 example
A new example has been added to run quantized Qwen3 models (0.6B, 1.7B, 4B, 8B, 14B, and 32B variants) using GGUF files. Users can now download and run these models via the \quantized-qwen3\ example, with support for interactive and chat modes, configurable sampling parameters, and automatic model/tokenizer fetching from the Hugging Face Hub.
candle-examples/examples/quantized-gemma, candle-examples/examples/quantized-qwen3 · high confidence
Add reinforcement learning examples (DQN, DDPG, Policy Gradient)
New examples demonstrating reinforcement learning algorithms have been added to the candle-examples package. This includes implementations of Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), and Policy Gradient methods. The examples utilize a custom Rust-based Gymnasium wrapper (gym\_env.rs) to interface with Python environments, supporting both standard environments like CartPole-v1 and Atari games via vectorized wrappers (vec\_gym\_env.rs) and custom Atari preprocessing scripts (atari\_wrappers.py). Users can run these examples using the \--features=pyo3\ flag.
candle-examples/examples/reinforcement-learning · high confidence
Add support for Based, BEiT, BERT, BigCode, BLIP, ChatGLM, and CLIP models
The \candle-transformers/src/models\ module now includes implementations for several new model architectures, enabling users to run inference for Based (a linear attention language model), BEiT (a vision transformer), BERT (for sentence embeddings and similarity), BigCode/StarCoder (code generation), BLIP (image captioning), ChatGLM (multilingual conversational models), and CLIP (contrastive language-image pre-training). These additions expand the library's capabilities across text, code, and vision-language tasks.
candle-transformers/src/models · high confidence
Added Mamba inference example
A new example has been added to run Mamba State Space Model inference. The \main.rs\ file implements a text generation CLI that supports multiple model sizes (130m through 2.8b) and allows users to specify prompts, temperature, and sampling parameters via command-line arguments. A corresponding \README.md\ provides instructions on how to run the example.
candle-examples/examples/mamba · high confidence
Added Metal GEMM benchmark example
A new example file, \metal\_benchmarks.rs\, has been added to the \candle-metal-kernels\ package to allow users to benchmark Metal GEMM (General Matrix Multiply) performance. This tool measures throughput in GFLOPS for matrix multiplication operations using both f32 and f16 data types across various matrix dimensions (512, 1024, 2048, 4096), providing a way to evaluate the performance of the underlying Metal kernels.
candle-metal-kernels/examples · high confidence
Added Moondream image-question-answering example
Users can now run the Moondream computer-vision model via a new \moondream\ example in \candle-examples\. This example accepts an image and a text prompt, loads the model (supporting both standard and quantized variants), and generates a text answer describing the image content. It includes a README with usage instructions and supports command-line options for device selection (CPU/GPU), temperature, top-p, and repeat penalty.
candle-examples/examples/moondream · high confidence
Flash attention kernel integration with variable-length and paged KV cache support
The \candle-flash-attn/kernels\ directory now contains the CUDA implementation for Flash Attention, enabling efficient attention computation on NVIDIA GPUs. This update introduces support for variable-length sequences (varlen) via cumulative sequence length arrays (\cu\_seqlens\_q\, \cu\_seqlens\_k\), allowing models to process batches of sequences with different lengths without padding. It also adds support for paged key-value (KV) caches through a \block\_table\ parameter, which is essential for memory-efficient inference in long-context scenarios. The kernels support head dimensions of 32, 64, 96, 128, 160, 192, 224, 256, and 512, and include features such as ALiBi (Attention with Linear Biases) slopes, softcap scaling, local windowing, and dropout. The implementation is split into multiple files by head dimension and data type (FP16/BF16) to optimize compilation times, and includes specific optimizations for head dimension 512 (used in Gemma 4) by disabling unused features like ALiBi and softcap.
candle-flash-attn/kernels · high confidence
Initial ONNX model support in Candle
This change introduces the \candle-onnx\ crate, enabling users to load and execute ONNX models within the Candle framework. It includes the ONNX protocol buffer definitions, a parser for model files, and an evaluation engine (\eval.rs\) that maps ONNX operators to Candle tensor operations, allowing for inference of models exported in the ONNX format.
candle-onnx/src · high confidence
Initial implementation of FlashAttention v3 CUDA kernels
This change introduces the core CUDA kernel infrastructure for FlashAttention v3 within the \candle-flash-attn-v3/hkernel\ directory. It adds the host kernel headers (\flash.h\, \combine.h\, \epilogue\_fwd\_sm90\_tma.hpp\) defining the parameter structures, LSE combination logic, and SM90 TMA-based epilogue operations. It also includes the Python/C++ API bindings (\flash\_api.cpp\, \flash\_api.cu\) for launching the kernels from PyTorch, and version-specific SM90 TMA copy implementations (\copy\_paged\_sm90\_tma\_cutlass35.hpp\, \copy\_paged\_sm90\_tma\_cutlass36.hpp\) to support paged attention with different CUTLASS versions.
candle-flash-attn-v3/hkernel · high confidence
Initial implementation of Stable Diffusion model components
This change introduces the core building blocks for Stable Diffusion support, including the UNet 2D denoising model, CLIP text encoder, Variational Autoencoder (VAE), and multiple diffusion schedulers (DDIM, DDPM, Euler Ancestral Discrete). It also adds attention mechanisms, residual network blocks, and time-step embeddings required to run Stable Diffusion v1.5, v2.1, and XL models.
_candle-transformers/src/models/stable\diffusion · high confidence
Initial implementation of the Würstchen diffusion model
Adds the core components for the Würstchen image generation model, including the PaellaVQ tokenizer for encoding/decoding images, the WPrior network for generating latent representations from text conditions, and the WDiffNeXt decoder for reconstructing images from latents. The implementation includes specialized attention mechanisms with optional flash-attention support, custom layer normalization, and a DDPM scheduler for the denoising process.
candle-transformers/src/models/wuerstchen · high confidence
Initial integration of fused MoE kernel using cuTile
Added a new fused Mixture-of-Experts (MoE) kernel implementation in \candle-nn/src/moe/cutile.rs\ that leverages the cuTile library for route-aware grouped matrix multiplication. This change introduces a specialized CUDA kernel (\routed\_grouped\_matmul\) designed to optimize the performance of MoE layers by fusing routing and computation steps, supporting bfloat16 precision and configurable tile sizes.
candle-nn/src/moe · high confidence
Initial project scaffolding and documentation
This change establishes the foundational structure of the Candle project by adding the MIT and Apache 2.0 license files, a comprehensive README with usage examples and model listings, a CHANGELOG documenting versions up to v0.3.0, a Makefile for build tasks, a .gitignore for build artifacts, and a .pre-commit-config.yaml for Rust code formatting and linting.
(repo-wide) · high confidence
Initial release of the Candle Documentation Book
The Candle project now includes a comprehensive documentation book built with mdBook. This new resource provides users with a structured guide covering installation for various backends (CPU, CUDA, cuTile, MKL, Metal), a step-by-step MNIST tutorial (modeling, training, saving/loading), and reference sections on error management, tracing, and advanced CUDA kernel development. It also includes guides for creating applications in WASM, REST APIs, and Tauri desktop apps, along with a PyTorch cheatsheet and examples for loading models from the Hugging Face Hub.
candle-book · high confidence
Initial release of the Candle PyO3 Python bindings
This change introduces the \candle-pyo3\ package, providing Python bindings for the Candle machine learning library. It exposes core tensor operations, neural network modules (such as \Linear\, \LayerNorm\, and \Embedding\), and functional APIs (like \relu\, \softmax\, and pooling) to Python. The bindings include support for quantized tensors (\QTensor\), GGUF and safetensors model loading/saving, and optional hardware acceleration via CUDA, MKL, and Apple Accelerate. Additionally, it provides a \QuantizedLlama\ model implementation and comprehensive type hinting (\.pyi\ files) to support static analysis and IDE autocomplete.
candle-pyo3 · high confidence
Initial release of the candle-core tensor library
This change introduces the candle-core crate, providing the foundational tensor operations, backend abstractions, and data types for the Candle machine learning framework. It establishes the core \Tensor\ API, including support for multiple data types (such as F32, F16, BF16, and various quantized formats), device management (CPU, CUDA, and Metal), and a modular backend trait system that allows for hardware-specific implementations. The release also includes essential mathematical operations (convolutions, pooling, matrix multiplication), automatic differentiation (backpropagation), and utilities for loading and saving models in formats like PyTorch and Safetensors.
candle-core/src · high confidence
Initial release of the candle-nn neural network layer library
This change introduces the candle-nn crate, providing a comprehensive set of building blocks for constructing neural networks. It includes standard layers such as Linear, Conv1d, Conv2d, Embedding, and various normalization layers (BatchNorm, LayerNorm, GroupNorm, RmsNorm). The library also exposes a wide range of activation functions (including PReLU, Mish, SwiGLU, and LeakyReLU), loss functions (CrossEntropy, NLL, MSE, Huber, BCEWithLogit), and utility modules like VarBuilder for parameter management, Seq for sequential composition, and Func for custom closure-based layers. Additionally, it provides KV cache implementations for efficient autoregressive generation and encoding utilities like one-hot encoding.
candle-nn/src · high confidence
Initial support for Mixture-of-Experts (MoE) models with fused CUDA kernels
Added fused CUDA kernels for Mixture-of-Experts (MoE) architectures, enabling efficient inference for MoE models. The new implementation includes token alignment and expert routing logic, as well as GEMM kernels that support both standard and GGUF-quantized weights (Q8\_0, Q4K, Q2K, Q3K, Q5K, Q6K) using Tensor Cores (WMMA) for accelerated computation.
candle-kernels/src/moe · high confidence
Introduce Flash Attention v3 layer for Hopper GPUs
Adds a new Flash Attention v3 implementation for the Candle framework, targeting NVIDIA Hopper architecture (sm90a). This feature enables forward attention kernels with head dimensions of 64, 128, 256, and 512, supporting both FP16 and BF16 data types. It also includes support for Grouped Query Attention (GQA) with group counts of 2, 4, 8, 16, and 32. The build process uses cudaforge to compile the necessary CUDA kernels and links against the CUTLASS library (commit 4c42f73) to provide optimized attention mechanisms for compatible hardware.
candle-flash-attn-v3 · high confidence
Introduce FlashAttention v3 implementation for CUDA
Added the \candle-flash-attn-v3\ crate, providing a new FlashAttention v3 backend for CUDA devices. This feature enables faster attention computation with support for head dimensions of 64, 128, 256, and 512, and includes options for ALiBi slopes, causal masking, and GQA (Grouped Query Attention) packing.
candle-flash-attn-v3/src · high confidence
Introduce Llama2 C WASM example with nucleus sampling and performance metrics
Adds a new WebAssembly-based Llama2 C inference example that runs the model in a background worker thread. The UI allows users to adjust generation parameters including temperature and top\_p (nucleus sampling), and displays real-time performance metrics such as tokens per second and total generation time.
candle-wasm-examples/llama2-c/src · high confidence
Introduce Llama2 C WASM example with repeat penalty and nucleus sampling
The Llama2 C WASM example now includes a repeat penalty mechanism to reduce token repetition in generated text and supports top\_p (nucleus) sampling for more controlled output diversity. These features are implemented in the WASM module's model logic, allowing users to configure repeat\_penalty and top\_p parameters during initialization to influence the generation process.
candle-wasm-examples/llama2-c/src/bin · high confidence
Introduce T5 family example with translation, embedding, and model variants
Adds a new T5 example supporting encoder-decoder translation (including MADLAD-400 multilingual models), sentence embedding generation, and multiple model variants (T5-small/base/large, T5-3B, MT5 variants, and FLAN-T5/UL2 via revision flags). Users can enable/disable KV caching, adjust generation parameters (temperature, top\_p, repeat penalty), and run on CPU or GPU, with optional tracing support.
candle-examples/examples/t5 · high confidence
Introduce candle-datasets crate for dataset loading and batching
The new \candle-datasets\ crate provides utilities for handling datasets in Candle. It includes a \Batcher\ component that groups tensor iterators into batches, supporting configurable batch sizes and the option to return incomplete final batches. Additionally, it offers a \hub\ module to load Parquet files directly from the Hugging Face Hub, re-exporting the \FileReader\ trait for convenient metadata access without requiring users to add the \parquet\ crate as a direct dependency.
candle-datasets/src · high confidence
Introduction of Flash Attention integration
This change introduces the \candle-flash-attn\ crate, providing a Rust interface to Flash Attention kernels for accelerated transformer attention mechanisms. The implementation exposes a \FlashAttn\ struct that allows users to configure parameters such as \softmax\_scale\, \alibi\_slopes\, \window\_size\_left/right\, and \softcap\. It supports CUDA execution with BF16 and FP16 data types, validates tensor shapes (requiring rank 4 inputs and head dimensions that are multiples of 8 up to 512), and handles memory management via CUDA streams. The core logic bridges Rust tensors to C FFI bindings (\run\_mha\) to perform the attention computation efficiently on GPU.
candle-flash-attn/src · high confidence
New AVX-optimized quantized vector operations and CUDA fast paths for GGUF models
This change introduces new, highly optimized CPU and GPU backends for quantized inference within the \candle-core/src/quantized\ module. On x86\_64 CPUs, a new \avx.rs\ file provides AVX2/FMA-accelerated \vec\_dot\ implementations for various quantization types (Q2K through Q8K), significantly speeding up matrix-vector multiplications. On the GPU side, new \cuda.rs\, \fast\_mmq.rs\, and \fast\_mmvq.rs\ files implement fast CUDA kernels for GGUF quantized tensors, supporting batched matrix-matrix (MMQ) and matrix-vector (MMVQ) operations for quantization types including Q4\_0, Q4\_1, Q5\_0, Q5\_1, Q8\_0, and the K-quants (Q2K-Q6K). These additions are complemented by corresponding dummy implementations for non-CUDA/Metal builds and updated file loaders (\gguf\_file.rs\, \ggml\_file.rs\) to support the new storage types.
candle-core/src/quantized · high confidence
New BERT example for sentence embeddings and similarity
Added a new BERT example that demonstrates computing sentence embeddings and calculating cosine similarities between sentences. The example supports loading models from the Hugging Face Hub (defaulting to \sentence-transformers/all-MiniLM-L6-v2\), allows specifying custom model IDs, and includes options for CPU/GPU execution, PyTorch weight loading, L2 normalization of embeddings, and an approximate GELU activation for performance. It also handles attention masks during average pooling to align with standard sentence-transformer behavior.
candle-examples/examples/bert · high confidence
New CPU Flash Attention implementation with optimized single-batch and variable-length paths
A new CPU-optimized flash attention implementation has been added to \candle-nn\, providing significantly faster inference for large language models. The module introduces specialized kernels for single-batch (B=1) scenarios, which bypass batch-indexing overhead for direct slice access, and a variable-length path that packs sequences to avoid padding waste. It supports f32 and f16 data types, causal masking, ALiBi biases, and logit soft-capping, while falling back to an unfused standard implementation for unsupported configurations or as a correctness reference.
candle-nn/src/attention · high confidence
New CUDA kernel infrastructure and expanded operation support
The candle-kernels module has been restructured with a new set of CUDA source files (affine, binary, cast, conv, fill, indexing, etc.) and supporting headers (cuda\_utils, compatibility). This introduces native GPU implementations for a wide range of tensor operations, including binary math (add, sub, mul, div, min, max, comparisons), affine transforms (scale/shift), type casting (including fp8, bf16, f16, i64), convolution (conv1d/conv2d with im2col), indexing (index\_select), and data filling. It also adds specialized kernels for Mixture-of-Experts (MoE) GEMM and GGUF quantized matmul (MMVQ), along with utility functions for atomic operations and math across various data types and CUDA architectures.
candle-kernels/src · high confidence
New DeBERTaV2/V3 example for NER and text classification
Added a new \debertav2\ example that supports Named Entity Recognition (NER) and text classification tasks using DeBERTaV2 and DeBERTaV3 models. Users can run inference on models from the HuggingFace Hub or local directories, with support for both CPU and GPU execution, PyTorch weight loading, and basic benchmarking.
candle-examples/examples/debertav2 · high confidence
New EfficientViT image classification example
Added a new example demonstrating inference with the EfficientViT model from Microsoft Research Asia. Users can now run this example to classify images using pre-trained models (M0–M5) trained on ImageNet, with support for selecting specific model variants and running on CPU or GPU.
candle-examples/examples/convmixer, candle-examples/examples/efficientvit · high confidence
New Encodec audio compression example with live microphone support
The encodec example now supports three modes: converting audio files to compressed tokens, decoding tokens back to audio, and full audio-to-audio compression. A key behavioral change is the addition of live microphone input (via \cpal\) and real-time audio output, allowing users to record from their default input device and play back the processed audio directly to their default output device by specifying \-\ as the file path. The example also handles automatic sample-rate resampling to the model's required 24kHz and normalizes the loudness of the generated audio output.
candle-examples/examples/encodec · high confidence
New GGUF tokenizer and ONNX basics example programs
Two new example programs are now available in the examples directory. The \gguf-tokenizer\ example demonstrates how to load tokenizer metadata directly from GGUF files (supporting both local paths and Hugging Face Hub URLs) and tokenize/decode text prompts. The \onnx\_basics\ example provides a CLI tool to inspect ONNX model graphs and perform simple evaluations of ONNX models using Candle's ONNX runtime.
candle-examples/examples · high confidence
New GraniteMoeHybrid example for IBM Granite 4.0 Micro
Added a new example in \candle-examples/examples/granitemoehybrid\ that demonstrates text generation using IBM's Granite 4.0 Micro hybrid Mixture-of-Experts model. The example showcases the \GraniteMoeHybrid\ implementation, including specific features like embedding/logit scaling and hybrid attention, and supports CPU, CUDA, and Metal backends with configurable precision and sampling parameters.
candle-examples/examples/granitemoehybrid · high confidence
New Jina-Bert embedding example added
A new example demonstrating the Jina-Bert model (jina-embeddings-v2-base-en) has been added to the candle-examples suite. Users can now compute sentence embeddings for a single prompt or calculate cosine similarities between sets of sentences using average pooling and L2 normalization. The example supports CPU/GPU execution, configurable model/tokenizer paths, and tracing.
candle-examples/examples/jina-bert · high confidence
New Llama example supporting multiple model variants
The \candle-examples/examples/llama\ directory now contains a new \main.rs\ and \README.md\ that provide a unified CLI for running various Llama-based architectures. Users can select specific models via the \--which\ flag, including Llama v1, v2, v3, v3.1, v3.2 (1B and 3B), Solar-10.7B, TinyLlama, and SmolLM2 variants. The example supports configuration options for CPU/GPU execution, data types (f16, bf16, f32), sampling parameters (temperature, top\_p, top\_k, repeat penalty), and optional Flash Attention.
candle-examples/examples/llama, candle-examples/examples/resnet · high confidence
New Llama2.c WASM examples with model caching and UI improvements
The llama2-c WASM example location now provides two distinct ways to run the model: a Pure Rust UI (using Trunk) and a Vanilla JS/WebWorker interface. The WebWorker implementation introduces a browser Cache API layer to force-cache model weights and tokenizer files, improving load performance on subsequent visits. The UI has been updated to support nucleus sampling (top\_p), configurable repeat penalties, and seed selection, with visual feedback for loading, generating, and completion states.
candle-wasm-examples/llama2-c · high confidence
New MNIST training example with configurable options and checkpoint support
Added a new \mnist-training\ example that demonstrates training multi-layer perceptrons and convolutional networks on the MNIST dataset. The example supports loading pre-trained weights via a \--load\ flag, saving trained model checkpoints using a \--save\ flag, and configuring the number of training epochs. It also includes a README with usage instructions and sample output.
candle-examples/examples/mnist-training · high confidence
New NomicBert embedding example supporting task prefixes
Added a new example demonstrating the nomic-embed-text-v1.5 model, which uses the NomicBert architecture to generate 768-dimensional embeddings for semantic search and clustering. The example supports computing embeddings for a single prompt or calculating cosine similarities between multiple sentences, and includes optional task prefix support (e.g., 'search\_document:', 'search\_query:') to improve retrieval quality.
candle-examples/examples/nomic-bert · high confidence
New ONNX LLM example for SmolLM-135M
Added a new example demonstrating how to run ONNX-based Large Language Models in Candle, specifically supporting the SmolLM-135M model. Users can now execute this example via \cargo run --example onnx-llm --features onnx\ to generate text using an ONNX model loaded from Hugging Face, with support for CPU/GPU execution and various sampling strategies.
candle-examples/examples/onnx-llm · high confidence
New Phi example supporting multiple model versions and quantization
The \candle-examples/examples/phi\ directory now contains a new example application that supports running Microsoft Phi models (versions 1, 1.5, 2, 3, 3-Medium, 4-Mini, 2-Old) as well as the MixFormer and Puffin-Phi-v2 models. The implementation allows users to select specific model variants via command-line arguments, supports both standard and quantized (GGUF-style) model loading, and includes features like repeat penalty, configurable temperature/top-p sampling, and verbose token output. The example also provides documentation on how to run inference for these models.
candle-examples/examples/phi · high confidence
New Qwen3 WASM example with SIMD optimizations and shared chat template support
This change introduces a new WebAssembly-based Qwen3-0.6B language model example that runs entirely in the browser, featuring SIMD optimizations for faster inference and support for both Q8\_0 and Q4\_K\_M quantized models. It includes a new shared \candle-wasm-chat-template\ library providing Jinja-based chat template rendering compatible with HuggingFace's tokenizer, supporting multi-turn conversations and thinking mode for reasoning models. The Qwen3 example demonstrates this with a complete chat interface, performance profiling tools, and automatic model downloading via a Python server.
candle-wasm-examples/chat-template, candle-wasm-examples/quant-qwen3 · high confidence
New Stable Diffusion example with multi-version and inpainting support
A new \stable-diffusion\ example has been added, providing a Rust/Candle implementation of the Stable Diffusion pipeline. Users can now generate images using Stable Diffusion v1.5, v2.1, XL 1.0, and Turbo variants via the \--sd-version\ flag. The example supports automatic weight downloading from the HuggingFace Hub, image-to-image generation, and inpainting capabilities. It also includes performance optimizations such as Flash Attention support, F16 precision options, and batched generation, with output images saved as PNG files.
candle-examples/examples/stable-diffusion · high confidence
New T5 WebAssembly demo with quantized model support
This change introduces a new browser-based demo for running T5 transformer models (including t5-small, flan-t5-small, and quantized variants like t5-small-quantized and the CoEdIT text-rewrite model) directly in the browser using Candle and WebAssembly. The demo provides two main capabilities: conditional text generation (e.g., translation, summarization) and sentence embedding extraction. It supports both standard safetensors models and GGUF-quantized models for reduced size, with the UI allowing users to select models, tasks, and generation parameters like temperature and repeat penalty.
candle-wasm-examples/t5 · high confidence
New VGG example supporting VGG13, VGG16, and VGG19
A new example has been added to demonstrate the implementation of VGG models (VGG13, VGG16, and VGG19) using the Candle library. Users can now run the example to load an image and classify it using any of the three VGG variants by specifying the model via the \--which\ argument (e.g., \vgg13\, \vgg16\, or \vgg19\). The example handles device placement, loads pre-trained weights from the Hugging Face Hub, and outputs the top 5 classification predictions.
candle-examples/examples/vgg · high confidence
New Whisper WebAssembly examples with Rust and Vanilla JS UIs
This change introduces two new end-to-end examples for running OpenAI Whisper speech recognition in the browser using Candle-compiled WebAssembly. The first example features a Pure Rust UI built with Yew, allowing users to load and transcribe audio samples directly in the browser. The second example demonstrates a Vanilla JavaScript implementation using Web Workers, supporting drag-and-drop audio files and multiple model configurations (including quantized GGUF models and Distil-Whisper). Both examples include the necessary build scripts, HTML entry points, and Rust source code for the WASM decoder, audio processing, and worker communication.
candle-wasm-examples/whisper · high confidence
New XLM-RoBERTa example with fill-mask, reranking, and text classification tasks
A new example application for the XLM-RoBERTa model has been added, supporting three distinct tasks: fill-mask (using models like xlm-roberta-base), reranking (using BGE reranker models like bge-reranker-base), and text classification (specifically a formality classifier). The example allows users to run these tasks via command-line arguments, automatically downloading the necessary tokenizer, config, and weight files from Hugging Face, and supports both CPU and GPU execution.
candle-examples/examples/xlm-roberta · high confidence
New YOLOv8 WebAssembly example with Rust and Vanilla JS UIs
This change introduces a new YOLOv8 object detection and pose estimation example for the Candle WebAssembly runtime. It provides two distinct user interfaces: a Pure Rust UI built with Yew and Trunk, and a Vanilla JS UI that leverages Web Workers for background inference. The example supports multiple YOLOv8 model sizes (n, s, m, l, x) for both standard object detection and pose estimation, allowing users to run inference directly in the browser using WASM.
candle-wasm-examples/yolo · high confidence
New Z-Image text-to-image generation example
Added a new example demonstrating text-to-image generation using the Z-Image model (specifically the Turbo variant) from Alibaba. This example supports running on CUDA, Metal (macOS), and CPU, and includes features such as automatic weight downloading from HuggingFace, local weight loading, configurable image dimensions (must be divisible by 16), negative prompts for classifier-free guidance, and dynamic timestep shifting for the flow matching scheduler.
_candle-examples/examples/z\image · high confidence
New browser-based Segment Anything demo with WASM and Web Workers
This change introduces a complete, runnable example for the Segment Anything Model (SAM) in the browser. It includes a Rust/WASM module (m.rs) that loads SAM weights and computes image embeddings and point-based masks, exposed via wasm-bindgen. A Web Worker (samWorker.js) handles the heavy lifting, including model loading, image embedding caching, and segmentation, communicating with the main thread via postMessage. The UI (lib-example.html) provides a vanilla JS interface with drag-and-drop image support, point selection, undo functionality, and the ability to download the resulting mask as a PNG. A build script (build-lib.sh) automates the WASM compilation and bundling process.
candle-wasm-examples/segment-anything · high confidence
New candle-flash-attn crate with CUDA build infrastructure
This change introduces the \candle-flash-attn\ component, providing a Rust crate that compiles Flash Attention CUDA kernels. The implementation includes a \build.rs\ script that utilizes the \cudaforge\ library to fetch a specific commit of the CUTLASS library and compile 53 kernel files (supporting head dimensions 32–512, fp16/bf16, and causal variants) into a static library. It also adds a README file to document the crate.
candle-flash-attn · high confidence
New candle-transformers crate with quantized model support and object detection utilities
This change introduces the candle-transformers crate, providing foundational infrastructure for running quantized models and computer vision tasks. It includes a VarBuilder for loading GGUF files, quantized neural network layers (Embedding, Linear, RmsNorm), and a FusedMoe implementation for efficient Mixture-of-Experts models. Additionally, it adds object detection utilities for bounding box handling and non-maximum suppression, along with shared utilities for causal masking and repeat penalties.
candle-transformers/src · high confidence
New core examples for CPU, CUDA, Metal, and custom CUDA kernels
Added five new examples in candle-core/examples to demonstrate core tensor operations and device capabilities. basics.rs shows CPU tensor slicing; cuda\_basics.rs benchmarks 1D convolutions on CUDA; cuda\_sum\_benchmark.rs measures sum performance on CUDA; metal\_basics.rs demonstrates Metal device usage with GPU trace capture; and cutile.rs provides a complete example of implementing and running a custom CUDA kernel using the cutile library.
candle-core/examples · high confidence
New custom-ops example demonstrating RMS normalization on CPU and CUDA
The candle-examples/examples/custom-ops directory now contains a complete example illustrating how to implement custom operations with both forward and backward pass capabilities. This specific example implements RMS normalization for f32 tensors, providing a CPU implementation via the CustomOp1 trait and a CUDA implementation using custom PTX kernels. The example includes the necessary CUDA kernel source files (layernorm\_kernels.cu), reduction utilities, and a build configuration to compile and link the custom kernel, allowing users to run the operation on either CPU or GPU hardware.
candle-examples/examples/custom-ops · high confidence
New example embedding BERT model into a single binary
Added the \bert\_single\_file\_binary\ example, which demonstrates how to inline model configuration, tokenizer, and weights directly into the compiled binary using Rust's \include\_bytes!\ macro. This allows for a self-contained executable that does not require downloading model files at runtime. The build process requires setting the \CANDLE\_BUILDTIME\_MODEL\_REVISION\ environment variable and enabling the \buildtime-download\ feature flag to trigger the compile-time download step.
_candle-examples/examples/bert\_single\_file\binary · high confidence
New logits processing and sampling strategies for text generation
The \candle-transformers\ crate now includes a \LogitsProcessor\ in \src/generation/mod.rs\ that supports multiple sampling strategies for text generation, including ArgMax, standard temperature sampling, Top-K filtering, Top-P (nucleus) sampling, and combined Top-K then Top-P. It also introduces Gumbel-Softmax sampling and allows arbitrary temperature modifications, providing users with finer control over the randomness and diversity of generated text.
candle-transformers/src/generation · high confidence
New optimizer and CPU benchmark examples
Added a basic optimizer example demonstrating linear regression with the AdamW optimizer, and a CPU benchmarks example featuring im2col-based conv2d, standard conv1d/conv2d, and matmul performance tests.
candle-nn/examples · high confidence
New quantized LLaMA inference example with broad model support
A new \quantized\ example has been added to the \candle-examples\ crate, enabling fast inference of quantized LLaMA-style models using built-in quantization methods (2-bit through 8-bit integer) and GGUF/GGML file formats. The example automatically downloads weights from the HuggingFace Hub and supports a wide variety of pre-configured models via the \--which\ flag, including LLaMA variants (7b, 13b, 70b, code, chat), Mistral (including instruct v0.2), Zephyr, OpenChat 3.5, Starling, Mixtral (sparse mixture of experts), Llama3, Phi3, SmolLM2, and DeepseekR1. It features interactive and chat modes, SIMD optimizations for Apple Silicon and x86, and command-line flags for local model files, prompts, and sampling parameters.
candle-examples/examples/quantized · high confidence
New quantized Phi model example
Added a new example for running quantized Microsoft Phi models (Phi-2, Phi-3, Phi-3b, and Phi-4) using GGUF files. The example includes a README with usage instructions and a main.rs implementation that supports loading models from the Hugging Face Hub, interactive and chat modes, and various sampling parameters.
candle-examples/examples/quantized-phi · high confidence
New quantized Qwen2 and DeepSeek-R1 Distill Qwen examples
Added new \quantized-qwen2-instruct\ example supporting Qwen2 models (0.5B, 1.5B, 7B, 72B) and the DeepSeek-R1-Distill-Qwen-7B variant. Users can now run these quantized models via the \--which\ argument, with automatic downloading of the corresponding GGUF files and tokenizers from Hugging Face.
candle-examples/examples/quantized-qwen2-instruct · high confidence
New quantized T5 translation example
Added a new example demonstrating how to run quantized T5 models for sequence-to-sequence tasks like translation and text editing. The example supports multiple model variants (T5-small, Flan-T5 variants) and allows users to specify custom models, such as CoEdit or MADLAD-400, via command-line arguments. It includes documentation on generating quantized weights using the tensor-tools utility and running inference with configurable parameters like temperature and repeat penalty.
candle-examples/examples/quantized-t5 · high confidence
New shared utility library for examples
The \candle-examples\ crate now exposes a \src\ library module providing shared utilities for all examples. This includes audio processing helpers (loudness normalization via ITU-R BS.1770, PCM decoding with Symphonia, and resampling with Rubato), a Jinja-based chat template renderer compatible with Hugging Face tokenizers, and image handling functions (loading, resizing, and saving with ImageNet normalization). It also adds a streaming token output wrapper for incremental text generation, a Hub client wrapper for model downloads with progress reporting, and class label constants for COCO and ImageNet.
candle-examples/src · high confidence
New tensor-tools binary for inspecting and quantizing model files
A new standalone \tensor-tools\ binary has been introduced to allow users to inspect, print, and convert model weights. The tool supports listing contents (\ls\) and printing tensor data (\print\) for formats including Safetensors, Npz, Ggml, Gguf, and PyTorch Pth. It also provides \quantize\ and \dequantize\ commands to convert between GGUF and Safetensors formats, with specific quantization modes (e.g., Q4\_0, Q6K) and logic that mirrors llama.cpp's default quantization behavior for 2D weight tensors.
tensor-tools · high confidence
New whisper-microphone example for real-time speech transcription
Added a new example that transcribes audio in real-time using a microphone input. The implementation supports streaming audio chunks, handles non-aligned audio data, and includes language detection for 99 languages. It also supports the whisper large-v3 turbo model and provides options for quantized models.
candle-examples/examples/whisper-microphone · high confidence
RWKV example now supports v5, v6, and v7 models with custom tokenization
The RWKV example has been updated to support RWKV v5, v6, and v7 models (including v7a and v7b variants), replacing the previous single-version implementation. This change introduces a custom tokenizer for RWKV models and adds specific CLI options for v7 features such as prompt templates (chat, think, fake-think, fill-in-middle), sampling presets, and stop sequences. Users can now run inference on a wider range of model sizes and variants, with dedicated handling for v7's different configuration and state structures.
candle-examples/examples/rwkv · high confidence
Support for Qwen3 and GTE-Qwen embedding models
The Qwen example now supports the Qwen3 model family, including the 3-moe-a3b Mixture-of-Experts variant, allowing users to run inference on the latest Qwen architectures. Additionally, a new GTE-Qwen example has been added to support the gte-Qwen1.5-7B-instruct embedding model, enabling text embedding generation via the HuggingFace Hub or local repositories.
candle-examples/examples/qwen · high confidence
Architecture
Metal kernel implementation refactored into modular source files
The Metal kernel logic in \candle-metal-kernels/src/kernels\ has been reorganized from a single monolithic file into distinct modules (e.g., \affine.rs\, \binary.rs\, \cast.rs\, \convolution.rs\, \indexing.rs\, \mlx\_gemm.rs\, \quantized.rs\, \sdpa.rs\, \sort.rs\, \unary.rs\, etc.). This change improves code maintainability and readability by separating concerns for different kernel types while preserving the existing public API and functionality.
candle-metal-kernels/src/kernels · high confidence
Behavioural changes
Build-time model downloads and sandbox-safe CUDA kernel generation
The candle-examples now support downloading model files (config, tokenizer, weights) at build time via the new \buildtime-download\ feature, controlled by the \CANDLE\_BUILDTIME\_MODEL\_REVISION\ environment variable, and generate CUDA kernel bindings using \cudaforge\ in a sandbox-safe manner by writing outputs to \$OUT\_DIR\ instead of the source tree.
candle-examples · high confidence
CPU backend refactored with runtime feature detection and SIMD-optimized kernels
The CPU tensor backend has been restructured to use runtime CPU feature detection (via \candle-core/src/cpu/features.rs\) to automatically select the most efficient vectorization path at execution time. This change introduces dedicated SIMD kernel implementations for AVX (x86/x86\_64), NEON (ARM/aarch64), and SIMD128 (WebAssembly), enabling hardware-accelerated operations like vector addition, dot products, and reductions. It also adds optimized support for half-precision (f16) and brain-float (bf16) data types, including fallbacks for CPUs lacking specific instruction sets like f16c or fp16, and integrates the \libm\ library for accurate error function (erf) calculations.
candle-core/src/cpu · high confidence
CUDA backend refactored into modular components with new cuTile and cuDNN integrations
The CUDA backend implementation has been reorganized into distinct modules to improve maintainability and extend hardware support. A new \cudnn.rs\ module provides optimized 1D and 2D convolution operations using the cuDNN library, while a new \cutile.rs\ module introduces a context for launching custom cuTile kernels with proper CUDA stream interop. The core \device.rs\ module now includes a host-to-device (HTOD) cache specifically designed to optimize data transfers during CUDA graph captures, and the \error.rs\ module has been split out to centralize CUDA-specific error handling. These changes collectively restructure the backend to support more specialized operations and better performance on supported GPU architectures.
_candle-core/src/cuda\backend · high confidence
CUDA kernel build system refactored with \`cudaforge\` and conditional optimizations
The \candle-kernels\ build process has been rewritten to use the \cudaforge\ crate, replacing the previous build script. This change introduces a split build strategy: most kernels are compiled to PTX for broader compatibility, while specific high-performance kernels (MoE, MMQ, MMVQ) are statically compiled into a library. The build script now automatically detects the GPU's compute capability and disables bf16 WMMA kernels on pre-Ampere GPUs (compute capability \< 8.0) to prevent runtime errors. Additionally, it conditionally links \stdc++\ for non-MSVC targets and supports the \CUTILE\ feature to include alignment kernels.
candle-kernels · high confidence
Consolidate benchmark suite into a unified entry point
The benchmarking infrastructure has been reorganized to provide a single, unified entry point for running all performance tests. A new \bench\_main.rs\ file now aggregates and registers all existing benchmark modules—including vector dot products, affine operations, concatenation, contiguous memory handling, binary and broadcast operations, copying, 2D transposed convolutions, matrix multiplication, quantized matrix multiplication, random number generation, reduction, unary operations, and conditional selection—allowing users to execute the complete benchmark suite from one location.
candle-core/benches · high confidence
GLM4 example now supports both old and new model architectures
The GLM4 example has been updated to support the new GLM-4-0414 architecture alongside the original GLM-4 models. Because the two architectures are incompatible, users must now explicitly specify the model type using the \--which\ flag (e.g., \--which "glm4-new"\ or \--which "glm4-old"\) when running the example to avoid initialization errors.
candle-examples/examples/glm4 · high confidence
Metal backend kernel infrastructure and error handling
The \candle-metal-kernels\ crate has been restructured to provide a robust foundation for Metal GPU operations. A new \MetalKernelError\ enum in \err.rs\ introduces detailed error reporting, including specific variants for command buffer failures, library loading issues, and type mismatches, along with automatic backtrace capture for debugging. The \kernel.rs\ module implements a centralized \Kernels\ struct that manages Metal library and compute pipeline caching, ensuring efficient reuse of compiled shaders. Additionally, \source.rs\ embeds the Metal shader source code directly into the binary, and \utils.rs\ provides essential helpers for buffer management, command encoder parameter setting, and thread-group sizing.
candle-metal-kernels/src · high confidence
Metal backend refactored with MLX kernel integration and improved concurrency
The Metal backend implementation has been significantly refactored to integrate MLX matmul kernels by default, replacing previous GEMM implementations. This change includes new buffer management strategies, such as using private storage mode for intermediate compute buffers to reduce coherency overhead, and registering external buffers in the device residency set. Concurrency and stability are improved through thread-isolated command buffer maps and status semaphores to ensure the backend is Send/Sync, while also fixing CPU readback races under concurrent command submission. Additionally, support for new data types (BF16, FP8, I16, I32) and operations (scatter, gather, bilinear interpolation, affine) has been added, alongside optional debug labeling and deterministic random number generation.
_candle-core/src/metal\backend · high confidence
Metal backend refactored with new concurrency, synchronization, and residency management
The Metal backend in \candle-metal-kernels/src/metal\ has been refactored to improve performance and reliability. A new command queue and command buffer pool system enables concurrent dispatching and better multi-threaded performance. Cross-encoder synchronization is now handled via explicit fences and a global output map to manage memory visibility between encoders. Buffer residency is managed via a \ResidencySet\ to keep buffers in GPU memory, reducing overhead. The device module now uses \MTLCopyAllDevices()\ for reliable enumeration and guards against NULL architecture pointers on simulators. Additionally, command buffer errors are now captured and exposed to the user.
candle-metal-kernels/src/metal · high confidence
Metal kernel implementation overhaul
The Metal backend kernels in candle-metal-kernels have been completely rewritten and expanded. This change introduces new, optimized implementations for core operations including GEMM (via MLX Steel kernels), GEMV, binary/unary/affine/ternary element-wise ops, casting, quantized matmul (Q4/Q5/Q8), sorting, and indexing (gather/scatter). It also adds support for convolution, bilinear upsampling, and random number generation. The new kernels feature improved strided indexing, scalar broadcasting, and better handling of various data types (f32, f16, bf16, u8, i64, etc.), replacing the previous Metal kernel set.
_candle-metal-kernels/src/metal\src · high confidence
Refactored CPU backend with tiled Conv2d and NdIter-based operations
The CPU backend implementation has been restructured to improve performance and correctness. Conv2d operations now use a specialized tiled im2col kernel that processes input and output in parallel tiles, with specific fast paths for 1x1 convolutions, replacing the previous monolithic approach. Additionally, element-wise operations like \where\ and binary maps have been migrated to use the \NdIter\ iterator for more efficient handling of strided and non-contiguous layouts, ensuring correct behavior for complex tensor shapes.
_candle-core/src/cpu\backend · high confidence
Test coverage
Add WASM test suite for CPU and quantized operations; Add initial benchmark suite for candle-nn operations; Added tests for Flash Attention v3 integration; Added tests for Flash Attention variants; Added tests for text generation sampling and object detection NMS; Added unit tests for ONNX operators; Expanded test coverage for candle-core operations; Expanded test coverage for neural network components and operations; New structured benchmark suite for core tensor operations.
Dependencies
Candle workspace updated to version 0.11.0 with dependency upgrades
The Candle ML framework workspace has been bumped to version 0.11.0, updating all core crates (candle-core, candle-nn, candle-transformers, etc.) to this new version. This release includes significant dependency upgrades: cudarc is updated to 0.19.10, tokenizers to 0.23.1, safetensors to 0.8.0, gemm to 0.19.0, and hf-hub to 1.0.0. Other notable updates include pyo3 to 0.29, yew to 0.23.0, and web-sys to 0.3.104. The workspace now uses the Cargo resolver 2 and standardizes dependency versions across the monorepo.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 36 → 50 (+13.9)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 43 → 55 (+12.8)
- Architecture 97 → 72 (-24.5)
- Maturity 51 → 65 (+13.9)
- Readiness 21 → 100 (+79.1)
- Security 46 → 62 (+15.8)
- Accessibility 36 (new)
Resolved (12)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- LLM evaluation failed
- No exposed public API
- No tests found
- Test reliability not included
New (1715)
- Ambiguous method naming. clone_dtoh and clone_htod likely refer to memory copy operations (Device-to-Host, Host-to-Device), but the verb 'clone' is misleading for a data transfer operation. It suggests object duplication rather than memory movement. Furthermore, CudaStorage has transfer_to_device, which is clearer but less specific about direction.
- AttentionWeights::forward (cognitive 28) (candle-transformers/src/models/quantized_qwen3.rs)
- AvgPool2D::f (cognitive 21) (candle-core/src/cpu_backend/mod.rs)
- BPE::from_json (cognitive 16) (candle-transformers/src/models/metavoice.rs)
- BarrierPool::new (cognitive 27) (candle-core/src/utils.rs)
- BlockQ2K::from_float (cognitive 31) (candle-core/src/quantized/k_quants.rs)
- BlockQ2K::from_float_imatrix (cognitive 21) (candle-core/src/quantized/k_quants.rs)
- BlockQ2K::vec_dot_unopt (cognitive 16) (candle-core/src/quantized/k_quants.rs)
- BlockQ3K::from_float (cognitive 45) (candle-core/src/quantized/k_quants.rs)
- BlockQ3K::from_float (cyclomatic 17) (candle-core/src/quantized/k_quants.rs)
- BlockQ3K::from_float_imatrix (cognitive 38) (candle-core/src/quantized/k_quants.rs)
- BlockQ3K::to_float (cognitive 22) (candle-core/src/quantized/k_quants.rs)
- BlockQ3K::vec_dot_unopt (cognitive 42) (candle-core/src/quantized/k_quants.rs)
- BlockQ4K::from_float (cognitive 29) (candle-core/src/quantized/k_quants.rs)
- BlockQ4K::from_float_imatrix (cognitive 26) (candle-core/src/quantized/k_quants.rs)
- BlockQ4K::vec_dot_unopt (cognitive 26) (candle-core/src/quantized/k_quants.rs)
- BlockQ5K::from_float (cognitive 36) (candle-core/src/quantized/k_quants.rs)
- BlockQ5K::from_float_imatrix (cognitive 34) (candle-core/src/quantized/k_quants.rs)
- BlockQ5K::to_float (cognitive 19) (candle-core/src/quantized/k_quants.rs)
- BlockQ5K::vec_dot_unopt (cognitive 36) (candle-core/src/quantized/k_quants.rs)
- …and 1695 more
Changes since last survey
- 54 commits — 30 feature/other, 24 fixes
By area
- candle-core/src — 16 commits
- (root) — 7 commits
- .github/workflows — 6 commits
- candle-examples/examples — 5 commits
- candle-examples/Cargo.toml — 3 commits
- candle-kernels/src — 2 commits
- candle-onnx/src — 2 commits
- candle-transformers/src — 2 commits
- candle-book/Cargo.toml — 1 commit
- candle-book/src — 1 commit
- candle-core/tests — 1 commit
- candle-flash-attn-v3/src — 1 commit
- candle-metal-kernels/src — 1 commit
- candle-nn/src — 1 commit
- candle-pyo3/Cargo.toml — 1 commit
- candle-pyo3/src — 1 commit
- candle-wasm-examples/llama2-c — 1 commit
- candle-wasm-examples/yolo — 1 commit
- candle-wasm-tests/tests — 1 commit
Notable commits
- fix: Fix CPU argsort for offset tensor views (#3875)
- fix: Fix CUDA copy_strided_src over-copying from an offset source view (#3944)
- fix: Fix FlashAttention v3 null tile count semaphore, stream, and LSE sizing (#3892)
- fix: Fix Metal rank-5+ strided reduce reading out of bounds (#3873)
- fix: Fix SIMD128 half-precision CPU kernels (#3904)
- fix: Fix broken examples that were hidden behind feature flags (#3945)
- fix: Fix concrete shape element count validation (#3813)
- fix: Fix conv2d CUDA im2col offset bug for non-contiguous kernels (#3894)
- fix: Fix decode regression from aaarch64/x86 repack kernels (#4000)
- fix: Fix index_select backward with non-contiguous ids (#3997)
- fix: Fix repeat with zero factors (#3877)
- fix: Fix silently incorrect CPU matmul for a stride-zero broadcast batch (#3960)
- fix: Fix the rotary emb convention in mimi transformer (#3949)
- fix: Fix the symphonia-gated examples build against the 0.6 audio API (#3931)
- fix: Fix the ug+metal build against the guard-based encoder API (#3855)
- fix: Fix unused warnings in cande-wasm-tests (#3946)
- fix: fix(cpu): avoid very recently stable aarch64 fp16 vector type (#3845)
- fix: fix(kernels): gate the __hmax_nan/__hmin_nan shim (#3909)
- fix: fix(metal): surface command buffer errors that occur during execution (#3862)
- fix: fix(models): centralize the additive causal mask, build it batch-independently (#3879)
- …and 34 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- python
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
huggingface/candle was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit aebc405d2b4bf42808387e0ca597bf7dad9b565f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.