Skip to content
CAI
Software that uses CAICheck a score

EricLBuehler/mistral.rs

66.1

Adequate · 29 September 2026

487.6k

lines of production code

Rust

with Python

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a high-performance, Rust-based inference engine for large language and multimodal models, optimized for local deployment on CPU, NVIDIA GPU, and Apple Silicon hardware. It provides a comprehensive suite of capabilities including agentic tool use with sandboxed code execution, support for diverse model architectures (such as MoE, diffusion, and speech models), and extensive quantization methods for efficient resource usage. The system exposes these features through a Python SDK, a CLI, and a server offering OpenAI and Anthropic API compatibility.

How it got here

2024 — multimodal expansion and PagedAttention

61 changes.

This period focused on expanding the inference engine to support a wide range of new multimodal, MoE, and diffusion models, including LLaVA, FLUX, and Qwen3-VL. It also introduced a comprehensive PagedAttention scheduler with FP8 quantization and MLA support to significantly improve memory efficiency and throughput. Additionally, the project released dedicated Rust and Python SDKs with agentic tool-use capabilities and revamped its quantization infrastructure with new CUDA and Metal kernels.

2025–2026 — multimodal expansion and quantization acceleration

77 changes.

This period focused on integrating a wide array of new multimodal models, including vision, audio, and speech capabilities for architectures like Gemma 4, Qwen 3.5, and Phi-4. It simultaneously advanced high-performance inference through extensive CUDA kernel development for FP8, MXFP4, and NVFP4 quantization, alongside the introduction of agentic tooling, sandboxed code execution, and distributed tensor parallelism infrastructure.

Features

Add CUDA dequantization kernels for bitsandbytes quantization formats

This change introduces a new CUDA kernel file (\dequant.cu\) that implements blockwise dequantization for bitsandbytes quantization schemes, specifically supporting General 8-bit, FP4, and NF4 data types. The implementation provides optimized device-side functions to convert quantized weights back to floating-point (f32) or half-precision (f16) formats, enabling the use of these specific quantization methods on CUDA hardware.

mistralrs-quant/kernels/bitsandbytes · high confidence

Add CUDA kernel for indexed Mixture-of-Experts (MoE) inference

This change introduces a new CUDA implementation for the indexed Mixture-of-Experts (MoE) forward pass, specifically designed to handle quantized weights. By adding the \indexed\_moe.cu\ file, the system now supports efficient GPU-accelerated inference for MoE architectures using various quantization formats (such as Q4\_K, Q8\_0, etc.), enabling faster processing for models that utilize this specific routing mechanism.

_mistralrs-quant/kernels/indexed\moe · high confidence

Add CUDA kernels for Affine Fast Quantization (AFQ)

New CUDA implementation files (afq.cu, afq\_gemm.cu, afq\_utils.cuh) have been added to the AFQ kernel directory, introducing GPU-accelerated dequantization, quantization, and fused GEMM operations. These kernels support 2, 3, 4, 6, and 8-bit quantization with group sizes of 32, 64, and 128, and include specialized handling for non-power-of-2 bit widths (3-bit and 6-bit) as well as embedding lookups. The addition enables on-the-fly dequantization during matrix multiplication to reduce memory bandwidth usage, supporting float, half, and bfloat16 data types.

mistralrs-quant/kernels/afq · high confidence

Add Conformer encoder for vision models

Introduces a new Conformer encoder implementation within the vision\_models module, including configuration structures, attention and feed-forward layers, Nemo-style convolutional subsampling, and positional encoding (absolute and T5-relative). This enables the runtime to process models that rely on the Conformer architecture for visual feature extraction.

_mistralrs-core/src/vision\models/conformer · high confidence

Add Dia 1.6b TTS model support with loudness normalization

Users can now generate speech using the Dia 1.6b text-to-speech model. This change introduces the core infrastructure for speech models, including a new \SpeechLoaderType\ enum and auto-detection logic to identify Dia models from their configuration. It also adds ITU-R BS.1770-4 compliant loudness analysis and normalization utilities to ensure consistent audio output levels, along with helpers to write generated PCM data as WAV files.

_mistralrs-core/src/speech\models · high confidence

Add Dia 1.6b text-to-speech model support

Users can now generate speech from text using the Dia 1.6b model. This change introduces the core inference pipeline in \mistralrs-core/src/speech\_models/dia\, including the \DiaPipeline\ for orchestration, \DiaKvCache\ for managing key-value states, and configuration structures (\DiaConfig\) for model parameters. It also implements the Descript Audio Codec (DAC) for audio encoding/decoding and specific audio processing utilities like delay pattern application and reverting, enabling the model to process text prompts and produce audio output.

_mistralrs-core/src/speech\models/dia · high confidence

Add FP8 linear layer support with quantization and UQFF serialization

Introduces a new FP8 quantization method for linear layers, enabling models to be loaded and run using FP8 weights. The implementation includes the \FP8Linear\ struct which handles quantization, dequantization, and forward passes, including optimized Metal-specific embedding gathering. It integrates with the existing UQFF format for serialization and deserialization, allowing FP8 models to be saved and loaded alongside other quantization types. The change also adds tests for FP8 roundtrip accuracy and embedding operations.

mistralrs-quant/src/fp8 · high confidence

Add GPTQ quantization support for CUDA

Users can now load and run models quantized with the GPTQ (and AWQ) format on CUDA devices. This change introduces the GPTQ backend implementation in \mistralrs-quant/src/gptq\, including CUDA kernels for matrix multiplication (with optional Marlin acceleration), FFI bindings for the underlying compute routines, and a CPU stub that explicitly rejects GPTQ loads on non-CUDA hardware.

mistralrs-quant/src/gptq · high confidence

Add Gemma 3n multimodal model support

This change introduces full support for the Gemma 3n model, enabling the processing of text, images, and audio inputs. The implementation includes a new audio tower with cumulative group normalization and relative position embeddings, an audio processor for resampling and mel-spectrogram generation, and a vision tower with specific block types like edge residuals and multi-query attention. A dedicated inputs processor handles multimodal layouts, mapping image and audio tokens to their respective embeddings, while the core text model supports sliding window attention and KV-cache sharing for efficient generation.

_mistralrs-core/src/vision\models/gemma3n · high confidence

Add Gemma 4 multimodal model support

Introduces first-class support for the Gemma 4 model family, including a new vision model implementation with both tower and unified vision paths, audio processing and encoding capabilities, and MTP (Multi-Token Prediction) speculative decoding. The change adds dedicated modules for audio handling (audio.rs, audio\_processing.rs), vision processing (vision.rs), text model logic (text.rs), and multimodal embedding (multimodal\_embedding.rs), along with configuration (config.rs) and input processing (inputs\_processor.rs) to handle images, audio, and video inputs alongside text.

_mistralrs-core/src/vision\models/gemma4 · high confidence

Add HQQ quantization support

Introduces High Quantization Quality (HQQ) as a new quantization method, allowing models to be compressed using 1, 2, 3, 4, or 8-bit precision. This implementation includes weight optimization to improve accuracy, dequantization kernels for CPU and Metal backends, and CUDA bit-packing kernels for faster processing on NVIDIA GPUs.

mistralrs-quant/src/hqq · high confidence

Add LLaVA and LLaVA-Next vision model support

Users can now run LLaVA and LLaVA-Next multimodal models, enabling image understanding capabilities alongside text generation. This change introduces the core configuration structures, vision tower (CLIP) integration, multi-modal projectors, and input processors required to handle image tokens and layout for both the original LLaVA-1.5 and the newer LLaVA-Next architectures.

_mistralrs-core/src/vision\models/llava · high confidence

Add LLaVA model support with LLaMA and Mistral backends

This change introduces the core LLaVA (Large Language and Vision Assistant) implementation within the vision models module, adding new \llama.rs\ and \mistral.rs\ files that provide LLaMA and Mistral-based LLM backends specifically tailored for LLaVA's multimodal requirements. The implementation includes custom attention mechanisms, RoPE position embedding handling, and integration with PagedAttention and quantization layers, enabling the system to process visual inputs alongside text using these specific model architectures.

_mistralrs-core/src/vision\_models/llava/llava\llm · high confidence

Add Llama 4 multimodal model support

Users can now load and run the Llama 4 vision-language model. This change introduces the model configuration, a processor for handling image inputs and special tokens (such as \\<\|image\|\>\), and the core inference logic including the vision encoder, text language model, and multimodal projector. The implementation supports features like tiled image processing, encoder caching for repeated images, and quantization configuration propagation.

_mistralrs-core/src/vision\models/llama4 · high confidence

Add MLlama (Llama 3.2 Vision) model support

This change introduces the MLlama architecture, enabling the use of Llama 3.2 Vision models. The implementation includes the full model structure (vision encoder, text decoder, and multimodal projector), a dedicated input processor to handle image tokens and cross-attention masks, and configuration parsing for vision-specific parameters like aspect ratios and rope scaling. Users can now load and run MLlama-based multimodal models.

_mistralrs-core/src/vision\models/mllama · high confidence

Add MXFP4 quantization support with optimized CUDA and Metal kernels

Users can now load and run models quantized with the MXFP4 format, which packs weights into 4-bit values to reduce memory usage. This change introduces the MXFP4 quantization method, including optimized matrix multiplication kernels for both CUDA (using standard and WMMA instructions) and Apple Metal, as well as specialized kernels for Mixture-of-Experts (MoE) models. The implementation supports both f16 and bf16 input precisions and includes optional bias handling, enabling faster inference for compatible models on supported hardware.

mistralrs-quant/src/mxfp4 · high confidence

Add Marlin CUDA kernels for GGUF affine and quantized matrix multiplication

This change introduces a new set of Marlin CUDA kernels in the \mistralrs-quant/kernels/marlin\ directory to accelerate quantized matrix operations. The update adds support for GGUF affine repacking (handling formats Q4\_0 through Q8\_K) and provides Marlin matmul implementations for GPTQ 4-bit and AWQ 4-bit quantization schemes, with kernels for both FP16 and BF16 data types. These additions enable more efficient inference for models using these specific quantization formats.

mistralrs-quant/kernels/marlin · high confidence

Add Marlin kernel support for GPTQ and AWQ quantization formats

The Marlin kernel implementation in this location now supports GPTQ (4-bit and 8-bit) and AWQ model formats. This is achieved by introducing new CUDA header files that define the necessary data types, scalar types for half and bfloat16 precision, and low-level asynchronous memory copy primitives required to handle these specific quantization schemes on NVIDIA GPUs.

mistralrs-quant/kernels/marlin/marlin · high confidence

Add NVFP4 Blackwell (SM121) acceleration via CUTLASS kernels

Users with Blackwell GPUs (SM121) and CUDA 13.3+ can now leverage native NVFP4 GEMM kernels for significantly faster inference. This change introduces a new CUDA kernel implementation in \mistralrs-quant/kernels/nvfp4\_cutlass\ and a Rust FFI binding in \mistralrs-quant/src/nvfp4/cutlass.rs\ that wraps CUTLASS collective operations for both prefill and decode phases. The implementation supports BF16 and F16 output types, handles NVFP4 weight packing with E2M1 precision and F8E4M3 block scales, and includes comprehensive random and deterministic tests to verify numerical correctness against reference implementations.

_mistralrs-quant/kernels/nvfp4\cutlass, mistralrs-quant/src/nvfp4 · high confidence

Add Phi-4 Multimodal support with audio and image processing

The Phi-4 vision model implementation now supports multimodal inputs, including audio and images. This change introduces new modules for audio embedding (using a Conformer encoder) and image embedding (using a Siglip vision transformer), along with configuration structures and input processors to handle multimodal layouts and special tokens for both media types.

_mistralrs-core/src/vision\models/phi4 · high confidence

Add Qwen2-VL multimodal model support

Users can now load and run the Qwen2-VL vision-language model. This adds a new model type that processes both text and images/videos, featuring a dedicated vision encoder, text decoder, and an input processor that handles media placeholders and grid-based image/video tokenization. The implementation supports configuration options for sliding window attention, rotary embeddings (MRoPE), and quantized weights, enabling multimodal inference capabilities within the existing MistralRS pipeline.

_mistralrs-core/src/vision\models/qwen2vl · high confidence

Add Qwen3.5 vision-language model support with MTP and packed multimodal inference

Users can now run the Qwen3.5 vision-language model family. This change introduces the model implementation in mistralrs-core, including configuration parsing (config.rs), the main multimodal model structure (mod.rs), and the text decoder (text.rs). It adds support for Multi-Token Prediction (MTP) speculative decoding via a dedicated drafter head (mtp.rs) and implements optimized packed multimodal inference for images and videos (packed\_visual.rs, packed\_gdn.rs), allowing multiple visual inputs to be processed in a single forward pass.

_mistralrs-core/src/vision\_models/qwen3\5 · high confidence

Add Vector FP8 quantization support with CUDA kernels

This change introduces a new Vector FP8 quantization method (VectorFP8Linear) in the mistralrs-quant module. It adds CPU and CUDA implementations for vector-based FP8 (F8E4M3) quantization and dequantization operations, including optimized CUDA kernels for dequantizing to f32, f16, and bf16. The module also integrates with the existing ISQ (In-Session Quantization) pipeline, allowing Vector FP8 weights to be further quantized using HQQ, AFQ, or GGUF formats.

_mistralrs-quant/src/vector\fp8 · high confidence

Add Voxtral Mini 4B multimodal model support

Introduces the Voxtral Mini 4B model, enabling real-time speech recognition and multimodal (audio-text) capabilities. This change adds the core implementation files for the Voxtral architecture, including configuration parsing, a Whisper-style audio encoder with mel spectrogram processing, a temporal adapter for audio downsampling, and the necessary input processors to handle audio prompts alongside text.

_mistralrs-core/src/vision\models/voxtral · high confidence

Add bitsandbytes quantization support (NF4, FP4, Int8)

Users can now load and run models quantized with the bitsandbytes library, specifically supporting NF4, FP4, and Int8 data types. This change introduces the necessary FFI bindings and CPU dequantization logic to handle these quantized weights, enabling inference with models that use these specific compression formats.

mistralrs-quant/src/bitsandbytes · high confidence

Add fast CUDA MMQ GGUF kernels

This change introduces a new set of high-performance CUDA kernels for GGUF quantized matrix multiplication in the \mistralrs-quant/kernels/mmq\_gguf\ directory. The implementation includes a self-contained common header (\mmq\_common.cuh\) defining GGML quantization types (Q2\_K through Q8\_K, IQ series, MXFP4, NVFP4) and block structures, along with instance files for specific quantization types (Q2\_K, Q3\_K, Q4\_0, Q4\_1, Q4\_K, Q5\_0, etc.). These kernels utilize MMQ (Matrix Multiply Quantized) techniques with stream-k launch policies and dynamic tiling to optimize performance on NVIDIA GPUs, replacing or supplementing previous quantization implementations for faster inference with GGUF models.

_mistralrs-quant/kernels/mmq\gguf · high confidence

Add mistral.rs-owned Flash Attention 2 CUDA kernels

This change introduces a new \mistralrs-flash-attn\ component containing a fork of Hugging Face Candle's Flash Attention 2 (FA2) CUDA kernels, specifically tailored for mistral.rs attention paths. The package includes a build script (\build.rs\) that compiles 53 pre-generated CUDA kernel variants (supporting FP16 and BF16 data types, causal and non-causal modes, and head dimensions from 32 to 512) against a pinned CUTLASS revision. It also provides the necessary C++ glue code, parameter structures, and device-side utilities (such as ALiBi, dropout, and block info handling) to enable high-performance attention inference on supported NVIDIA architectures.

mistralrs-flash-attn · high confidence

Add per-tensor FP8 linear layer implementation

Introduces the \PerTensorFP8Linear\ struct in the \mistralrs-quant\ crate, enabling support for FP8 quantization with per-tensor weight scales. This implementation validates that weights use the E4M3 format and scales use F32, supports various scale layouts (tensor, channel, or block), and integrates with the \cutile\ library for optimized CUDA execution of W8A16 and W8A8 matrix multiplications when available.

_mistralrs-quant/src/pertensor\fp8 · high confidence

Add project scaffolding and developer guidance files

The repository now includes foundational configuration and documentation files to standardize development workflows. This includes \.gitattributes\ to enforce consistent line endings, \.typos.toml\ to configure spell-checking for code identifiers, and \SECURITY.md\ to define the vulnerability reporting policy. Additionally, \AGENTS.md\ and \CLAUDE.md\ provide structured guidance for AI coding assistants on repository structure, build commands, and code style conventions. Standard Dockerfiles for CPU and CUDA environments are also added to support containerized builds and runtime.

(repo-wide) · high confidence

Add support for DeepSeek V2, V3, GLM-4 MoE, GPT-OSS, and Granite models

The model inference engine now supports several new architectures: DeepSeek V2 and V3 (including their MoE variants with MLA attention and LoRA projections), GLM-4 MoE and GLM-4 MoE Lite, GPT-OSS (with YARN rope scaling and custom SwiGLU), and IBM Granite (supporting hybrid Mamba/Attention layers and rope scaling). These additions expand the range of models that can be loaded and run, particularly for complex MoE and hybrid state-space architectures.

mistralrs-core/src/models · high confidence

Add support for Gemma and Qwen 3 embedding models

Users can now generate embeddings using Gemma and Qwen 3 architectures. This change introduces the core model implementations (\EmbeddingGemma\ and \Qwen3Embedding\), a dedicated input processor for handling embedding-specific tokenization and batching, and shared pooling/normalization layers required to produce final vector outputs.

_mistralrs-core/src/embedding\models · high confidence

Add support for Google's Gemma 3 model (text and vision variants)

Users can now load and run Google's Gemma 3 models in MistralRS. This change introduces the core implementation for both text-only and multimodal (vision) variants, including the necessary configuration parsing, input processing for image tokens, and the multi-modal projector to bridge vision and text embeddings. The implementation supports features like sliding window attention, rotary embeddings, and paged attention for efficient inference.

_mistralrs-core/src/vision\models/gemma3 · high confidence

Add support for Idefics 3 multimodal models

Users can now load and run Idefics 3 (including SmolVLM-Instruct) models. This change introduces the model architecture, configuration, and input processing logic required to handle image-text inputs, enabling multimodal capabilities for this specific model family.

_mistralrs-core/src/vision\models/idefics3 · high confidence

Add support for MiniCPM-O multimodal model

Users can now run the MiniCPM-O vision-language model. This change introduces the model architecture, which combines a Qwen2 text backbone with a SigLIP vision encoder and a custom resampler, along with the necessary input processing to handle image tokens and slicing within prompts.

_mistralrs-core/src/vision\models/minicpmo · high confidence

Add support for Mistral 3 multimodal model

Users can now load and run the Mistral 3 vision-language model. This change introduces the model configuration, a processor for handling multimodal inputs (images and text), and the core model implementation including the vision encoder, multimodal projector, and patch merging logic.

_mistralrs-core/src/vision\models/mistral3 · high confidence

Add support for Phi-3 Vision models

Introduces the core implementation for Microsoft's Phi-3 Vision models, including the model architecture in \mod.rs\ and the multimodal input processing logic in \phi3\_inputs\_processor.rs\. This enables users to load and run Phi-3 Vision models with support for image inputs, PagedAttention, and device mapping.

_mistralrs-core/src/vision\models/phi3 · high confidence

Add support for Qwen3 VL MoE multimodal model

Users can now load and run the Qwen3 VL MoE model, a vision-language variant that combines the Qwen3 VL vision encoder with a Mixture-of-Experts (MoE) text decoder. This change introduces the model configuration, vision processing logic, and MoE-specific text layers (including expert routing and MLP handling) within the \qwen3\_vl\_moe\ module, enabling inference for this specific architecture.

(repo-wide) · high confidence

Add support for the FLUX diffusion model

Users can now generate images using the FLUX diffusion model architecture. This change introduces the core model implementation (including the transformer blocks, attention mechanisms, and autoencoder), sampling logic with configurable guidance and timestep shifting, and the necessary text encoders (T5-XXL and CLIP) to process prompts.

_mistralrs-core/src/diffusion\models/flux · high confidence

Add support for the Muse Glimmer vision-language model

Users can now run the Muse Glimmer model, a multimodal architecture capable of processing both images and video. This change introduces the core implementation files in \mistralrs-core/src/vision\_models/muse\_glimmer\, including the model configuration (\config.rs\), input processing for media tokens (\inputs\_processor.rs\), and the text and vision encoder components (\text.rs\, \vision.rs\). The model supports specific attention mechanisms like sliding window attention for text and handles visual inputs through a dedicated vision encoder with caching.

_mistralrs-core/src/vision\_models/muse\glimmer · high confidence

Add support for the Qwen 3.5 MoE vision-language model

Users can now load and run the Qwen 3.5 MoE model, a multimodal architecture combining a vision encoder with a text model featuring Mixture-of-Experts (MoE) layers and hybrid attention (full and linear/GDN). This change introduces the model configuration, text backbone implementation, and vision integration logic, enabling inference for images and videos using the existing Qwen3-VL processor and caching infrastructure.

_mistralrs-core/src/vision\_models/qwen3\_5\moe · high confidence

Added CUDA kernels for MoE activation and alignment

New CUDA kernels have been added to the MoE module to support efficient activation functions and token routing. This includes \gelu\_tanh\_and\_mul\ and \silu\_and\_mul\ kernels for BF16 and FP16 data types, as well as a \moe\_sum\ kernel for aggregating expert outputs. Additionally, \moe\_align\ and \hunyuan\_moe\_capacity\_mask\ kernels are now available to handle token sorting, expert alignment, and capacity masking for MoE models.

mistralrs-quant/kernels/moe · high confidence

Added CUDA kernels for counting and indexing non-zero elements

This change introduces new CUDA kernel implementations for the \nonzero\ operation, allowing users to efficiently count and retrieve the indices of non-zero elements in tensors. The implementation supports multiple data types including float, double, int32, int64, uint8, uint32, int16, and half-precision formats (f16/bf16) where hardware architecture permits, leveraging CUB for device-side reduction and selection operations.

mistralrs-quant/kernels/ops · high confidence

Added FP8 vector quantization and dequantization CUDA kernels

New CUDA kernels have been added to support FP8 (E4M3) vector quantization and dequantization for 1xK vectors. This includes implementations for quantizing inputs from float32, float16, and bfloat16 into FP8, as well as dequantizing FP8 weights back to float32, float16, and bfloat16. Dummy fallback implementations are also provided for environments where FP8 is not supported, ensuring the build remains functional while logging a warning.

_mistralrs-quant/kernels/vector\fp8 · high confidence

Added GPTQ quantization kernel support for older CUDA architectures

The GPTQ quantization implementation in mistralrs-quant now supports CUDA compute capability 5.3 and lower. This is achieved by introducing custom CUDA kernel files (compat.cuh, matrix\_view.cuh, q\gemm.cu, and qdq\\*.cuh) that provide atomic operations and matrix view utilities for half and half2 types on architectures where native support is missing, enabling GPTQ-quantized models to run on older GPUs.

mistralrs-quant/kernels/gptq · high confidence

Added Gemma 3n model configuration presets

Users can now select from a variety of pre-configured Gemma 3n model variants, including the main 3.98B parameter model and several optimized configurations ranging from 1.91B to 3.79B parameters. These presets, defined in the new \gemma3n.csv\ config file, offer different trade-offs between model size, MMLU accuracy, and layer skipping strategies (both layer-level and block-level) to suit different performance and resource requirements.

_matformer\configs · high confidence

Added HQQ CUDA kernels for quantized weight packing and dequantization

This change introduces new CUDA kernels in the \mistralrs-quant/kernels/hqq\ directory to support the HQQ quantization scheme. The implementation includes \hqq.cu\, which provides dequantization logic for 8-bit, 4-bit, and 2-bit weights across float32, float16, and bfloat16 precisions, and \hqq\_bitpack.cu\, which handles the packing of 1-bit through 4-bit values into dense storage formats. These low-level operations enable efficient memory usage and inference speed for models utilizing HQQ quantization.

mistralrs-quant/kernels/hqq · high confidence

Added custom CUDA GEMV kernels for decode-phase inference

The \mistralrs-quant\ crate now includes a new \gemv\ module that provides optimized General Matrix-Vector multiplication kernels for CUDA. These custom kernels replace cuBLAS for small output workloads during the decode phase, offering performance improvements through vectorized loads, read-only cache paths, and warp-level reductions. The implementation supports batch sizes of 1–8 and data types BF16, F16, and F32, with an automatic selection mechanism that considers device compute capability and tensor shapes to decide when to use the optimized path.

mistralrs-quant/src/gemv · high confidence

Added fast CUDA MMVQ GGUF kernels

This change introduces a new CUDA kernel implementation (\mmvq\_gguf.cu\) for GGUF quantized matrix-vector operations. The file provides optimized device-side routines for various quantization types (Q8\_0, Q8\_1, Q2\_K through Q6\_K) and supports fused activation functions (SiLU, GELU, ReLU, GELU-ERF, Sigmoid) within the computation path, enabling faster inference for models using these specific GGUF formats on NVIDIA GPUs.

_mistralrs-quant/kernels/mmvq\gguf · high confidence

Added specialized CUDA GEMV kernels for small batch sizes

A new custom CUDA kernel implementation for General Matrix-Vector multiplication (GEMV) has been added to the \mistralrs-quant/kernels/gemv\ directory. This kernel is optimized for LLM decode-phase inference with small batch sizes (1-8), utilizing vectorized loads, read-only cache paths, and warp-level reductions to improve performance. It supports float, half, and bfloat16 data types and includes logic for optional bias application.

mistralrs-quant/kernels/gemv · high confidence

Expanded Python SDK examples for agentic, multimodal, and specialized model capabilities

The Python examples directory has been significantly expanded to demonstrate new SDK capabilities. Key additions include agentic tool usage with \agentic\_tools.py\ (showing local callbacks and external dispatch), secure code execution with sandboxing and approval workflows (\code\_execution.py\, \code\_execution\_approval.py\), and Model Context Protocol (MCP) integration (\mcp\_client.py\). The examples also cover a wide range of new model architectures and features: AnyMoE configuration and LoRA adapters (\anymoe.py\, \anymoe\_inference.py\, \anymoe\_lora.py\, \lora.py\), diffusion models like FLUX and DiffusionGemma (\flux.py\, \diffusion\_gemma.py\), embedding models (\embedding\_gemma.py\), and speech generation with Dia (\dia.py\). Support for various multimodal models (Gemma 3/3n/4, Llama 4, Idefics 2/3, LLaVA Next, LFM2.5-VL, Llama 3.2 Vision) is demonstrated, along with specialized models like DeepSeek R1/V2, GPT-OSS, IBM Granite, GLM-4 MoE, and LFM2.5. Additional examples show file input handling (\file\_inputs.py\), custom search callbacks (\custom\_search.py\), and grammar-based output constraints using Lark and llguidance (\json\_schema.py\, \lark.py\, \lark\_llg.py\, \llguidance.py\).

examples/python · high confidence

Initial AnyMoE training infrastructure and input handling

This change introduces the foundational components for AnyMoE (Mixture of Experts) training within the core library. It adds new modules to parse training inputs from structured CSV and JSON files, supporting prompts, expert assignments, and optional image URLs. It also defines the core traits and structs for managing MoE layers, including gating mechanisms, LoRA adapter support, and the ability to save/load gating layer weights after training.

mistralrs-core/src/amoe · high confidence

Initial X-LoRA support for Gemma, Gemma 2, Llama, Mistral, Mixtral, Phi-2, Phi-3, and Starcoder 2

This change introduces the core implementation for X-LoRA (X-Linear LoRA) adapters in the \mistralrs-core\ library. It adds a new \xlora\_models\ module containing model-specific implementations for Gemma, Gemma 2, Llama, Mistral, Mixtral, Phi-2, Phi-3, and Starcoder 2, along with a dedicated \XLoraClassifier\ for determining adapter scalings and an \XLoraConfig\ for configuration. These components enable the runtime loading, activation, and inference of X-LoRA adapters for the specified model families, integrating with the existing LoRA and device mapping infrastructure.

_mistralrs-core/src/xlora\models · high confidence

Initial implementation of CLIP text encoder for diffusion models

Added the core CLIP text encoder module (mistralrs-core/src/diffusion\_models/clip) to support text-conditioned diffusion models. This new module implements the necessary components for processing text inputs, including token and position embeddings, multi-head self-attention with causal masking, and a multi-layer perceptron (MLP) with configurable activation functions (e.g., QuickGELU). This provides the foundational text-processing capability required for models like FLUX that rely on CLIP-style text encoders.

_mistralrs-core/src/diffusion\models/clip · high confidence

Initial release of the mistralrs Python SDK

Introduces the \mistralrs\ Python package, providing a Python SDK for the mistral.rs inference engine. The release includes type stubs (\mistralrs.pyi\) and a \py.typed\ marker for PEP 561 compliance, enabling static type checking. The API exposes agentic capabilities, including tool use, web search, code execution, and shell access, with configurable permissions (Auto, Ask, Deny) and approval callbacks. It also supports LoRA adapter management, grammar-constrained generation, and various sampling parameters.

mistralrs-pyo3 · high confidence

Introduce AFQ (Adaptive Fast Quantization) quantization method

Users can now apply the new AFQ quantization method to models, supporting 2, 3, 4, 6, and 8-bit precision with configurable group sizes (32, 64, 128). This implementation provides optimized GPU acceleration for both Metal and CUDA backends, while maintaining a CPU fallback for other devices. The change also adds support for loading and serializing AFQ-quantized models via the UQFF format.

mistralrs-quant/src/afq · high confidence

Introduce AnyMoE pipeline and auto-loader support for Mixture-of-Experts models

The pipeline now supports AnyMoE models, which allow building a Mixture-of-Experts architecture from existing base models. This includes a new \AnyMoeLoader\ and \AnyMoePipeline\ that load a target model and apply a gating layer trained on a pre-specified dataset. The auto-loader has been updated to automatically detect and route to the correct loader for these models. Note that AnyMoE currently does not support PagedAttention.

mistralrs-core/src/pipeline · high confidence

Introduce Metal PagedAttention with FP8 KV-cache quantization

This change adds the core Metal GPU kernels for PagedAttention, enabling efficient, block-based key-value cache management on Apple Silicon. The implementation includes optimized kernels for copying, gathering, and reshaping cache blocks, along with new FP8 (E4M3/E5M2) support that allows quantizing the KV cache to reduce memory usage and improve throughput. Additionally, the Rust-side kernel loader has been updated to load precompiled Metal libraries directly from memory, ensuring compatibility with sandboxed environments like macOS apps distributed via TestFlight.

mistralrs-paged-attn/src/metal/kernels · high confidence

Introduce Model Context Protocol (MCP) client integration

The \mistralrs-mcp\ crate now provides a client implementation for connecting to external Model Context Protocol (MCP) servers. Users can configure connections via HTTP, WebSocket, or local processes, enabling the automatic discovery and registration of external tools and resources. The client supports bearer token authentication, concurrent tool execution, and customizable tool name prefixes to prevent conflicts with built-in capabilities.

mistralrs-mcp/src · high confidence

Introduce OS-level sandboxing for LLM-generated subprocesses

A new \mistralrs-sandbox\ library provides OS-level isolation for subprocesses spawned by the model, ensuring safer code execution. On Linux, it enforces resource limits (memory, CPU, processes, file descriptors) via cgroups v2 and rlimits, isolates the filesystem using Landlock, restricts network access via seccomp-bpf and network namespaces, and scrubs environment variables to a strict allowlist. On macOS, it leverages Seatbelt (\sandbox-exec\) with generated profiles to restrict file and network access. The system supports two configuration profiles—\Restricted\ (default, with loopback-only networking) and \Developer\ (with full networking and access to common development toolchains)—and can be toggled at runtime via the \MISTRALRS\_SANDBOX\ environment variable.

mistralrs-sandbox · high confidence

Introduce PagedAttention backend for CUDA and Metal

Users can now leverage PagedAttention for significantly improved memory efficiency and inference performance on CUDA and Metal devices. This change introduces a new \PagedAttention\ layer implementation that supports advanced features such as FP8 KV-cache quantization, FlashAttention V3 integration, and block-level prefix caching. The implementation is gated behind the \cuda\ and \metal\ feature flags, ensuring that users without these hardware backends continue to use the standard attention mechanism without disruption.

_mistralrs-core/src/paged\attention/layers · high confidence

Introduce UnquantLinear for unquantized layers with optimized CUDA wide GEMV support

A new \UnquantLinear\ implementation is added to handle unquantized linear layers, enabling the system to load and run models that contain layers without quantization (e.g., specific embeddings or heads). This change introduces optimized CUDA execution paths, including a 'wide GEMV' kernel for single-token decoding on specific hardware (SM121) and improved batch matrix multiplication logic via cuBLASLt. The implementation also includes logic to correctly handle bias addition in these unquantized layers, ensuring numerical stability and correctness during inference.

mistralrs-quant/src/unquantized · high confidence

Introduce \`mistralrs-server-core\` as the unified server implementation

The \mistralrs-server-core\ crate is introduced as the foundational library powering the Mistral.rs server. This new module consolidates the server's API surface, providing OpenAI-compatible endpoints for chat completions, completions, embeddings, and file uploads, alongside a dedicated implementation for the Anthropic Messages API. It also adds support for the OpenAI Responses API with background task management and response caching, and introduces an approval broker to handle app-driven tool approvals for agentic workflows.

mistralrs-server-core · high confidence

Introduce dedicated paged-attention library with FP8 and Metal support

This change introduces the \mistralrs-paged-attn\ crate, a new standalone library for PagedAttention kernels. It adds support for FP8 KV-cache quantization on NVIDIA GPUs (specifically targeting compute capability 9.0/Blackwell with Flash Attention 3) and provides a new Metal backend for Apple Silicon devices. The library includes build scripts for compiling CUDA and Metal kernels, along with Rust bindings for operations like block copying, cache gathering, and scale updates.

mistralrs-paged-attn · high confidence

Introduce dedicated vision preprocessing crate

A new \mistralrs-vision\ crate has been added to provide reusable image preprocessing utilities. It exposes a composable transform pipeline (via \Transforms\ and \ApplyTransforms\) that converts \DynamicImage\ inputs into normalized \Tensor\ outputs. The crate includes specific operations for resizing (nearest interpolation), padding to a maximum edge or image size, normalization by mean and standard deviation, and scaling, supporting RGB and RGBA image formats.

mistralrs-vision/src · high confidence

Introduce in-process file store for agentic tool outputs

Added a new \mistralrs-core/src/files\ module that provides an in-memory \FileStore\ with TTL-based expiration and a hard cap of 4096 entries to manage files generated by agentic tools. This module defines the \File\ and \FileContent\ types and includes helpers to convert tool outputs into structured file records, inject input file metadata into tool prompts, and compose tool responses with file summaries. This change establishes the internal mechanism for surfacing generated files from shell and code execution tools to the model and clients.

mistralrs-core/src/files · high confidence

Introduce sandboxed code and shell execution with automatic file output surfacing

The \mistralrs-code-exec\ module now provides a complete, sandboxed environment for executing Python code and shell commands. This change introduces persistent sessions for both Python and shell execution, enforcing configurable timeouts and OS-level sandbox policies (including network isolation and filesystem restrictions) to ensure safe execution. A key behavioral improvement is the automatic detection and surfacing of generated files: the system captures new or modified files created during execution and makes them available as downloadable artifacts, while also supporting explicit output specifications and a dedicated tool to surface previously generated files. Input files are safely mounted into the session working directory, and all execution results—including stdout, stderr, exceptions, and generated images or video frames—are returned in a structured format.

mistralrs-code-exec/src · high confidence

Introduce v1-style paged attention with block-level prefix caching and multimodal support

The paged attention subsystem has been rewritten to follow vLLM's v1 architecture, introducing a new block-based KV cache management system. This change adds block-level prefix caching, which uses content-addressable hashing to detect and reuse identical KV cache blocks across requests, significantly improving throughput for repeated prompts. The new system includes a dedicated block pool for O(1) allocation and free operations, a KV cache manager for high-level block tracking, and an encoder output cache to skip redundant vision/audio encoder passes for identical media. It also supports multimodal inputs by incorporating content hashes for images, audio, and video into the prefix cache keys, and allows configuring FP8 KV cache quantization for supported CUDA and Metal devices.

_mistralrs-core/src/paged\attention · high confidence

Introduce web search and content extraction tools with RAG capabilities

Users can now equip models with two new tools: \mistralrs\_search\_the\_web\ for querying the web and \mistralrs\_website\_content\_extractor\ for retrieving page content. The search module includes a retrieval-augmented generation (RAG) pipeline that uses a dedicated embedding model to chunk, embed, and re-rank search results, ensuring that retrieved content is capped to the tokenizer's context window. Configuration is handled via \WebSearchOptions\, which supports user location details and custom search descriptions.

mistralrs-core/src/search · high confidence

Introduction of PagedAttention scheduler with prefix caching support

The core scheduling engine now supports PagedAttention, introducing a new \PagedAttentionScheduler\ alongside the existing \DefaultScheduler\. This change enables block-level prefix caching, allowing the system to reuse previously computed KV cache blocks for identical input prefixes, which improves performance for repeated prompts. The scheduler module now exposes configuration for PagedAttention-specific parameters (such as max batched tokens and cache configuration) and includes validation logic (\PagedPrefixCacheValidator\) to manage cache hits and commits during sequence scheduling.

mistralrs-core/src/scheduler · high confidence

Introduction of core Mixture-of-Experts (MoE) module

A new core module for Mixture-of-Experts (MoE) models has been added to the library. This change introduces the foundational infrastructure for handling MoE architectures, including the definition of expert configurations, expert projection logic, and utilities for sharding experts across distributed environments. This module serves as the basis for supporting MoE-based models within the framework.

mistralrs-core/src/moe · high confidence

Introduction of first-class LoRA support with dynamic and static adapter implementations

The library now includes a dedicated LoRA module (\mistralrs-quant/src/lora\) that enables loading and applying Low-Rank Adaptation adapters. This change introduces two distinct execution modes: dynamic LoRA, which supports runtime switching between multiple adapters and expert routing via new kernel launches, and static LoRA, which merges adapter weights into the base layer at load time (similar to Phi-4 multimodal style). The implementation provides comprehensive configuration parsing for PEFT-compatible adapters, including support for rank, alpha, target/exclude modules, bias handling, and various advanced features like DoRA, QLoRA, and layer replication.

mistralrs-quant/src/lora · high confidence

Launch of the mistral.rs landing page with platform-specific install commands and static blog

The website directory now hosts the static landing page for mistralrs.dev, built with Vite for Cloudflare Pages. Users can view and copy platform-specific installation commands (macOS/Linux and Windows) via an accessible tabbed interface that includes a robust copy-to-clipboard feature with fallbacks. The site also features a statically generated blog section that renders release posts from the repository root, complete with Open Graph metadata, image handling, and SEO-friendly URLs.

website · high confidence

Metal GPU support for Gated Delta Net and Selective State Space Model kernels

Users running on Apple Silicon can now leverage optimized Metal GPU kernels for Gated Delta Net (GDN) and Selective State Space Model (SSM) operations, replacing previous CPU or unoptimized paths. This change introduces new Rust bindings and Metal shader implementations for the gated delta rule recurrence and Mamba-2 style selective scan, enabling faster inference for models utilizing these architectures on Metal devices.

mistralrs-core/src/metal · high confidence

Metal backend for PagedAttention with FP8 KV-cache quantization support

This change introduces the Metal-specific implementation for PagedAttention, enabling efficient key-value cache management on Apple Silicon hardware. The new backend includes optimized kernels for gathering KV cache blocks, copying and swapping cache blocks between memory locations, and updating KV scales. Crucially, it adds support for FP8 (F8E4M3) KV-cache quantization, requiring explicit F32 scale tensors for key and value caches, which allows for reduced memory usage and potentially higher throughput during inference on Metal devices.

mistralrs-paged-attn/src/metal/backend · high confidence

Native GGUF support for Gemma 3, Gemma 3n, Llama 4, and other multimodal architectures

The \mistralrs-core/src/gguf\ module now includes native GGUF loading support for several new model families, including Gemma 3, Gemma 3n, Llama 4, Idefics3/SmolVLM, and LFM2-VL. This change introduces dedicated configuration and tensor-binding logic for these architectures, enabling the system to correctly map GGUF tensor layouts to internal model structures. It also adds robust handling for multimodal components (vision, audio, and text) and improves tokenizer validation to prevent out-of-bounds panics from malformed GGUF metadata.

mistralrs-core/src/gguf · high confidence

New CUDA MoE kernels for Hunyuan and CUTLASS 2.x

Added new CUDA-based Mixture-of-Experts (MoE) operations in \mistralrs-quant/src/moe\, including token alignment, fused activation functions, and expert routing. This introduces support for Hunyuan v1's per-expert capacity masking and provides a BF16-only CUTLASS 2.x grouped-GEMM path for expert-sorted forward passes on Ampere (sm\_80+) or newer GPUs, enabling faster inference for MoE models on compatible hardware.

mistralrs-quant/src/moe · high confidence

New CUDA kernels for MLA and FP8-quantized PagedAttention

This update introduces new CUDA kernels and supporting infrastructure for Multi-Head Latent Attention (MLA) and FP8-quantized KV-cache in PagedAttention. The diff adds device kernels for concatenating and caching MLA-specific KV and PE tensors (\concat\_and\_cache\_mla\_kernel.cu\), block-copying operations for FP8 and standard types (\copy\_blocks\_kernel.cu\), and tiled FlashAttention-2 kernels with per-head attention sinks for prefill (\flash\_attn\_sinks.cu\). It also integrates FlashInfer's cascade and decode kernels for state merging and attention computation, alongside new FFI bindings in Rust to expose these operations. Supporting headers define vector types and arithmetic for BF16, FP16, FP32, and FP8 data types, enabling lower-precision inference and optimized memory handling for MLA architectures.

mistralrs-paged-attn/src/cuda · high confidence

New CUDA kernels for NVFP4, MoE, and dynamic LoRA

The \mistralrs-quant\ crate now includes new CUDA kernels that enable NVFP4 checkpoint loading with cuTile Blackwell acceleration, optimized Mixture-of-Experts (MoE) grouped GEMM operations, and true dynamic LoRA adapter support. These additions are accompanied by new benchmarking examples (\nvfp4\_bench.rs\, \routed\_lora\_bench.rs\) and build-system updates to support the new kernel compilation and header bundling.

mistralrs-quant · high confidence

New CUDA kernels for dynamic convolution and DFlash context operations

This change introduces new CUDA kernels and their Rust bindings in the \mistralrs-core/src/cuda\ directory to support dynamic convolution and DFlash context operations. The \dynamic\_conv\ implementation provides a fused kernel for applying dynamic weights to hidden states, supporting BF16, F16, and F32 data types with configurable kernel and group sizes. The \dflash\_context\ module adds kernels for packing taps and computing context keys with RMS normalization and RoPE, while \dflash\_selector\ implements greedy and sampling selection logic for DFlash decoding. These additions expand the available CUDA primitives for specific model architectures and decoding strategies.

mistralrs-core/src/cuda · high confidence

New CUDA paged attention backend with FP8 and MLA support

A new CUDA backend for paged attention has been introduced, adding support for FP8 KV-cache quantization (F8E4M3) and Multi-Head Latent Attention (MLA) patterns. This implementation includes optimized block-copying logic for prefix caching, FlashAttention-3 integration for FP8 decode, and specific context attention kernels for MLA models, enabling more efficient memory usage and inference performance for supported architectures.

mistralrs-paged-attn/src/cuda/backend · high confidence

New F8Q8 quantization method with CPU SIMD optimizations

Added the F8Q8 quantization method, enabling 8-bit floating-point (F8E4M3) weights with 8-bit integer activations for CPU inference. This new capability includes optimized vector dot-product kernels for AVX (x86), NEON (ARM), and SIMD128 (WebAssembly) architectures to accelerate matrix multiplications, alongside the necessary serialization and deserialization logic to support the format in the UQFF file standard.

mistralrs-quant/src · high confidence

New FP8 CUDA kernels for quantization, dequantization, and GEMM

This change introduces a new set of CUDA kernels in the \mistralrs-quant/kernels\ directory to support FP8 (E4M3) precision. The \blockwise\_fp8\ module adds blockwise quantization and dequantization kernels, optimized GEMM implementations (including a CUTLASS-based path for SM90 hardware and a tiled path for earlier architectures), and a specialized kernel for indexed Mixture-of-Experts (MoE) GEMM. Additionally, the \scalar\_fp8\ module provides element-wise conversion kernels between FP8 and standard floating-point types (F16, BF16, F32), while \fused\_rms\_norm\_fp8\ adds fused normalization and quantization steps. Dummy stub implementations are also included to ensure compatibility on GPUs that do not support FP8 compute capabilities.

_mistralrs-quant/kernels/blockwise\fp8 · high confidence

New GGUF backend with optimized CUDA kernels and indexed MoE support

A new GGUF weight source and backend has been added to mistralrs-quant, enabling first-class loading and execution of GGUF-quantized models. This includes a full GGUF archive parser, CPU/Metal fallback paths for indexed Mixture-of-Experts (MoE) layers, and high-performance CUDA kernels for fast matmul (fast\_mmq/fast\_mmvq) and indexed MoE forward passes. The implementation also introduces a Marlin-based packed affine backend for specific quantization types on CUDA, significantly expanding the range of supported quantized formats and accelerating inference for MoE architectures.

mistralrs-quant/src/gguf · high confidence

New Gated Delta Net (GDN) backend for hybrid models

This change introduces a new Gated Delta Net (GDN) implementation within the core engine, enabling support for hybrid models that utilize this architecture. The new module includes a backend with optimized recurrence kernels (including fused operations for CUDA and Metal), a caching system for recurrent states, and configuration handling for state data types (F16, BF16, F32) and value-head layouts (Grouped, Tiled). It also provides weight loading and projection logic compatible with quantization and LoRA, allowing users to run models that depend on GDN layers.

mistralrs-core/src/gdn · high confidence

New MTP and DFlash speculative decoding engine

The engine now supports Multi-Token Prediction (MTP) and DFlash block-diffusion draft models for speculative decoding. This introduces a new \mistralrs-core/src/speculative\ module that manages the full lifecycle: \config.rs\ defines the MTP and DFlash settings (including external assistant models and draft sampling methods), \cache.rs\ handles paged-attention KV cache reservation and rollback for speculative tokens, \proposer.rs\ drives the draft generation using target model hidden states, \verifier.rs\ validates drafts against the target model (supporting greedy and sparse rejection on CUDA), and \dflash.rs\ implements the block-diffusion architecture. Users can now enable speculative decoding to accelerate inference with these advanced drafting techniques.

mistralrs-core/src/speculative · high confidence

New Metal compilation subsystem for precompiled and runtime kernels

The \mistralrs-metal-compile\ crate introduces a dedicated system for managing Metal kernel compilation, supporting both Ahead-of-Time (AOT) precompilation for macOS, iOS, and tvOS, and runtime compilation. This change adds the infrastructure to embed Metal source files, configure compilation options (such as language version 3.1 and fast math mode), and handle platform-specific SDKs, enabling more efficient and flexible Metal backend support.

mistralrs-metal-compile · high confidence

New Metal kernels for quantization, data movement, and attention

This change adds a suite of new Metal GPU kernels to the \mistralrs-quant\ library, expanding hardware-accelerated capabilities for quantized inference and model operations. The new code includes dequantization kernels for BitsAndBytes formats (NF4, FP4, INT8), blockwise FP8, and F8Q8, alongside bit-packing utilities for HQQ (1-bit to 8-bit). It also introduces optimized kernels for fused GLU activations (SiLU, GELU variants, ReLU, Sigmoid), comprehensive data copy and transposition operations across various types, bitwise logical operations, and specialized Flash Attention routines for padding and block masking. These additions enable more efficient execution of quantized models and complex neural network layers on Apple Silicon hardware.

_mistralrs-quant/src/metal\kernels · high confidence

New Python SDK bindings for advanced model and agent capabilities

The Python bindings in \mistralrs-pyo3/src\ have been expanded to expose several new capabilities to Python users. You can now configure and train Any-Mixture-of-Experts (AnyMoE) models using the new \AnyMoeConfig\ and \AnyMoeExpertType\ classes. The SDK also introduces a sandboxed code-execution environment, allowing you to define \SandboxPolicy\ limits and manage \AgentPermission\ settings for agentic tool approvals. Additionally, you can now handle file inputs and outputs via the \InputFile\ and \RequestedFile\ classes, and the streaming interface (\ChatCompletionStreamer\) has been updated to support agentic tool call progress and approval events.

mistralrs-pyo3/src · high confidence

New Python code execution engine with rich output capture

A new Python executor script has been added to the mistralrs-code-exec module, enabling the runtime execution of user-provided code within a persistent namespace. This change introduces support for capturing rich outputs from executed code, including standard output and error streams, Matplotlib figures (as base64-encoded PNGs), PIL images, and Pandas DataFrame representations. The executor also enforces security by blocking stdin access and handles clean interruption via SIGINT.

mistralrs-code-exec/python · high confidence

New Rust SDK with unified model builders and agentic capabilities

The \mistralrs/src\ crate now provides a comprehensive Rust SDK featuring a unified \ModelBuilder\ for automatic model type detection (text, multimodal, embedding), alongside specialized builders for GGUF, diffusion, and AnyMoE models. This release introduces a built-in \Agent\ for executing agentic loops with tool calling, supports both synchronous (\BlockingModel\) and asynchronous APIs, and includes new configuration options for MCP client integration, Python code execution, and shell execution.

mistralrs/src · high confidence

New \`\#\[tool\]\` macro for defining agentic tools

The \mistralrs-macros\ crate now provides a \\#\[tool\]\ attribute macro that allows developers to define tools for the mistral.rs agentic loop by simply annotating Rust functions. This macro automatically generates the necessary \Tool\ definitions, callbacks, and argument structs, supporting features like parameter descriptions, default values, and optional parameters, thereby simplifying the creation of integrable AI tools.

mistralrs-macros · high confidence

New advanced examples for agentic tool use, code execution, and constrained generation

The \mistralrs/examples/advanced\ directory now includes a comprehensive suite of demonstration programs. These cover the agentic loop with both non-streaming and streaming tool calling (including parallel execution), strict tool choice restrictions, and app-driven approval callbacks for code execution. Additional examples demonstrate sandboxed Python code execution with first-class file outputs, Model Context Protocol (MCP) client integration, and various constrained generation techniques using GBNF regex, JSON schemas, and llguidance grammars. The collection also features examples for LoRA adapter loading, AnyMoE mixture-of-experts configurations, automatic device mapping, concurrent request batching, embedding generation, and custom logits processing.

mistralrs/examples/advanced · high confidence

New agentic loop with tool approvals, file outputs, and session management

The engine now supports an agentic loop that enables multi-turn tool use, including code execution, shell commands, and web search. This change introduces app-driven tool approvals, allowing applications to control when tools are executed. It also adds support for file outputs from shell commands and general execution, with new session management capabilities for storing and retrieving agentic conversations. The engine now handles input files and integrates them into the conversation context. Additionally, the admission queue and CUDA decode completion mechanisms have been updated to support these new agentic workflows.

mistralrs-core/src/engine · high confidence

New benchmarking, conversion, and build utility scripts

Added a suite of new scripts to support model workflows and release processes. The \bench.py\ and \bench\_nvfp4\_serving.py\ scripts provide tools for benchmarking server performance and throughput under various concurrency and quantization settings (including NVFP4). Conversion utilities (\convert\_to\_gptq.py\, \convert\_awq\_marlin.py\) allow users to transform models into GPTQ and Marlin quantized formats. The \build\_wheels.py\ script automates the creation of Python wheels for different platforms and accelerators (CUDA, Metal, CPU). Additional utilities include \production\_soak.py\ for long-duration server stability testing, \fetch\_nvfp4\_test\_fixtures.py\ for downloading test data, and \generate\_readme\_banner.py\ for creating documentation assets.

scripts · high confidence

New blockwise FP8 inference providers for CUDA

Added new CUDA execution paths for blockwise FP8 quantized models, introducing support for DeepGEMM SM90, CUTLASS SM90, and tensor-core GEMV providers. This change adds the \deepgemm.rs\, \mma.rs\, and \ffi.rs\ modules to handle JIT compilation, kernel preparation, and FFI bindings for these high-performance FP8 operations, while also adding scalar FP8 conversion kernels for CPU, CUDA, and Metal backends to support broader FP8 data type handling.

_mistralrs-quant/src/blockwise\fp8 · high confidence

New chat templates for tool calling, reasoning, and multimodal models

The chat\_templates directory now includes dedicated templates for several new model families and capabilities. Tool-calling support is added for DeepSeek, Mistral Nemo, Mistral Small 3, and Hermes 2 Pro/3, each with specific formatting for function signatures and calls. Reasoning capabilities are supported via the SmolLM3 template, which handles thinking modes. Multimodal support is introduced with templates for Gemma 3n (handling text, image, and audio tokens) and Idefics3 (handling text and image tokens). Standard templates are also provided for ChatML, LLaMA 2, LLaMA 3, Mistral, Phi 3, Phi 3.5, and Vicuna.

_chat\templates · high confidence

New cuBLASLt FP8 batch matrix multiplication with fused activations

The \mistralrs-quant\ module now includes a new cuBLASLt-based implementation for FP8 (F8E4M3) batch matrix multiplications. This change introduces fused kernels that combine matrix multiplication with optional bias addition and activation functions (ReLU, Gelu, SwiGLU), supporting both scalar and batched scaling factors. The implementation includes device-specific handle management to ensure correct operation on multi-GPU setups and adds corresponding unit tests to verify numerical accuracy against standard floating-point operations.

mistralrs-quant/src/cublaslt · high confidence

New cuTile CUDA kernels for FP8, MoE, and GDN prefill

Added a new cuTile backend module in mistralrs-quant/src/cutile that introduces fused CUDA kernels for FP8 matrix multiplication (W8A8 and W8A16), FP8 Mixture-of-Experts (MoE) grouped GEMM, and Gated Delta Network (GDN) prefill. These kernels leverage the cuTile library for tile-based execution on NVIDIA Blackwell and Ada GPUs, featuring JIT autotuning to optimize launch configurations based on tensor shapes and device capabilities. The implementation includes context bridging to integrate with the Candle CUDA stream, persistent blockwise GEMM logic with scale handling, and specialized routing for MoE experts, providing accelerated inference paths for models utilizing these quantization and architectural patterns.

mistralrs-quant/src/cutile · high confidence

New distributed backend infrastructure for tensor parallelism

The \mistralrs-quant\ crate now includes a new distributed module (\mistralrs-quant/src/distributed\) that provides the foundational infrastructure for multi-node and multi-GPU tensor parallelism. This change introduces a unified communication abstraction (\Comm\) supporting both NCCL (for CUDA) and Ring backends, along with socket-based synchronization primitives (\Server\ and \Client\) for cross-node coordination. It also adds \layers.rs\ to handle sharded linear layers and packed output layouts, enabling the framework to distribute model weights and computations across multiple devices or nodes.

mistralrs-quant/src/distributed · high confidence

New dynamic LoRA adapter support with CUDA kernels

The system now supports dynamic loading and application of LoRA adapters at runtime, including specialized kernels for Mixture-of-Experts (MoE) models. This change introduces a new execution arena and loader that handle adapter weights, allowing adapters to be applied on-the-fly without recompiling the model. For CUDA users, this includes new optimized kernels (via FFI bindings) for both standard linear layers and routed MoE expert layers, supporting FP32, FP16, and BF16 data types to accelerate inference with active adapters.

mistralrs-quant/src/lora/dynamic · high confidence

New example scripts for multimodal, diffusion, and speech models

Added example scripts demonstrating how to use the library with new model capabilities: audio transcription (Voxtral Mini), image generation (FLUX.1-schnell), block-diffusion text generation (DiffusionGemma), text-to-speech synthesis (Dia-1.6B), and combined audio/image multimodal interactions (Phi-4-multimodal-instruct). The examples also include unified templates for text and multimodal models, showcasing features like streaming, multi-turn conversations, and quantization.

mistralrs/examples/models · high confidence

New getting-started examples for embeddings, GGUF, multimodal, and streaming

The \mistralrs/examples/getting\_started\ directory now includes dedicated example scripts demonstrating key capabilities: \embedding/main.rs\ shows how to generate text embeddings using \EmbeddingModelBuilder\; \gguf/main.rs\ and \gguf\_locally/main.rs\ illustrate loading GGUF models from Hugging Face or local paths, including configuration of chat templates and paged attention; \multimodal/main.rs\ demonstrates handling image inputs with multimodal models; \streaming/main.rs\ provides a token-by-token streaming chat example; and \text\_generation/main.rs\ covers basic chat with ISQ quantization and logprobs. These examples serve as practical entry points for integrating these specific features.

_mistralrs/examples/getting\started · high confidence

New grouped MoE GEMM kernel for optimized prefill performance

A new CUDA kernel (\moe\_grouped.cu\) has been added to the \mistralrs-quant/kernels/moe\_grouped\ directory to optimize the prefill phase for Mixture-of-Experts (MoE) models. This implementation groups tokens by expert rather than launching one block per token/expert combination, which significantly reduces the number of CUDA blocks (e.g., from \~44.8M to \~90K for Gemma4 with 8K tokens) and improves weight data reuse via L1 cache. The kernel supports various quantization formats (Q4\_0, Q5\_0, Q8\_0, etc.) and includes helper functions for vector dot products and data type conversions, aiming to deliver substantial speedups for quantized MoE prefill operations.

_mistralrs-quant/kernels/moe\grouped · high confidence

New mistralrs-audio crate for audio processing utilities

The new \mistralrs-audio\ crate provides audio utilities for \mistral.rs\, mirroring the functionality of \mistralrs-vision\. It includes structures and functions for reading audio data (via WAV files or raw bytes), resampling, channel handling, and computing mel spectrogram features. Key features include bounds checking to prevent memory exhaustion during decoding (limiting to 30 minutes of 48kHz stereo audio), audio normalization to prevent clipping, fade-in/fade-out application to reduce artifacts, and DC offset removal. The crate also includes tests for WAV reading and PCM16 normalization.

mistralrs-audio · high confidence

New model-specific tool call parsers for ATEM, DeepSeek, Gemma 4, Harmony, Hunyuan, Liquid, Llama, Mistral Nemo, and Qwen

The tool-calling system now supports a wider range of model-specific output formats through new parsers in \mistralrs-core/src/tools/parsers\. This adds native parsing and grammar generation for Muse Glimmer (ATEM), DeepSeek, Gemma 4 (including strict schema-constrained mode), Harmony (GPT-OSS), Hunyuan, Liquid (LFM 2.5), Llama, Mistral Nemo, and Qwen. These parsers enable correct detection, extraction, and mid-stream grammar enforcement for tool calls across these model families, ensuring that tool arguments are validated against their schemas and that the model's output is constrained to the expected format.

mistralrs-core/src/tools/parsers · high confidence

New offline and online calibration infrastructure for In-Situ Quantization (ISQ)

The \isq\_flow\ module now provides the orchestration layer for ISQ, introducing support for both offline calibration (using a provided calibration file to generate imatrix statistics) and online calibration (collecting activation statistics from live model traffic). This change adds the logic to plan quantization strategies, drive calibration data through normal, multimodal, and embedding models, and apply the resulting imatrix data to requantize and swap model layers at runtime or load time.

_mistralrs-core/src/pipeline/isq\flow · high confidence

New optimized MXFP4 GEMM CUDA kernels

Added two new CUDA kernel implementations for MXFP4 matrix multiplication in the quantization module: a tiled GEMM kernel (mxfp4\_gemm.cu) and a WMMA tensor-core accelerated kernel (mxfp4\_gemm\_wmma.cu). These kernels enable efficient inference with MXFP4 quantized weights on NVIDIA GPUs by using LUT-based dequantization and vectorized loads, providing a new high-performance path for models using this quantization format.

mistralrs-quant/kernels/mxfp4 · high confidence

New quantization examples for In-Situ Quantization, NVFP4, and UQFF formats

Added demonstration scripts in the quantization examples directory showcasing advanced quantization capabilities. Users can now see how to perform In-Situ Quantization (ISQ) with automatic type selection, runtime re-ISQ, and online calibration that collects activation statistics from live traffic to hot-swap layers without restarting. The examples also cover NVFP4 checkpoint loading with cuTile Blackwell acceleration, per-layer quantization control using Topology, Mixture of Quant Experts (MoQE), and loading pre-quantized UQFF models for both text and multimodal inputs.

mistralrs/examples/quantization · high confidence

New server examples for agentic workflows, Anthropic API compatibility, and model-specific capabilities

The examples/server directory now includes a comprehensive suite of new client scripts demonstrating advanced server features. Agentic capabilities are covered by examples for tool-round loops (agentic\_tool\_rounds.py), strict tool selection (allowed\_tools.py), and interactive code-execution approval flows (code\_execution\_approval.py). Anthropic API compatibility is demonstrated through dedicated scripts for chat, streaming, tool calling, and skills (including file upload/download), alongside a Claude Code settings configuration. The collection also provides model-specific usage examples for Gemma 3/4 (vision, audio, video), LFM 2.5 (dense and MoE), GPT-OSS, Idefics 2/3, and Dia (TTS), as well as examples for LoRA adapter management, JSON schema enforcement, Lark grammar constraints, and file inputs.

examples/server · high confidence

New utility modules for quantization, FP8, and bitwise operations

The \mistralrs-quant/src/utils\ directory has been restructured into dedicated modules (\ffi\, \fp8\, \isq\, \log\, \ops\, \uqff\) to support advanced quantization and performance features. This includes CUDA-accelerated contiguous byte-copy for FP8 tensors, immediate and parallel In-Source Quantization (ISQ) execution, fused GPT-OSS SwiGLU kernels, and bitwise operations (left-shift, non-zero) for both CUDA and Metal backends. Additionally, UQFF serialization utilities and deduplicated logging helpers are now available to support the UQFF file format and reduce log noise.

mistralrs-quant/src/utils · high confidence

New vision model infrastructure and CLIP/SigLIP encoder support

This change introduces the foundational infrastructure for vision models in the core library, including a new \vision\_models\ module with dedicated files for configuration, preprocessing, and multimodal layout management. It adds initial implementations for CLIP and SigLIP vision encoders, enabling the system to process and embed visual inputs for multimodal models. The update also includes a generic image preprocessor trait and configuration parsers to handle diverse image and audio preprocessing requirements across different model architectures.

_mistralrs-core/src/vision\models · high confidence

Behavioural changes

CLI restructured with dedicated argument modules and new management commands

The CLI argument definitions have been reorganized into dedicated modules (model, paged\_attn, quantize, sandbox, server) to improve discoverability and reduce duplication. This change introduces new subcommands for managing the installation (\update\, \uninstall\), inspecting the environment (\doctor\), and managing the Hugging Face model cache (\cache list\, \cache delete\). It also adds a \bench\ command for performance benchmarking and a \login\ command for Hugging Face authentication, while consolidating server, sandbox, and quantization options into clearly structured argument structs.

mistralrs-cli · high confidence

Integrated FlashInfer CUDA primitives for optimized attention kernels

Added a suite of low-level CUDA header files (cp\_async, exception, fastdiv, fp16, frag\_layout\_swizzle, layout, math, mma, page, permuted\_smem) to the paged attention implementation. These headers provide PTX wrappers for asynchronous memory copies, matrix multiply-accumulate (MMA) instructions, half-precision conversions, and paged key-value cache management, enabling the underlying attention kernels to leverage hardware-accelerated operations for improved performance.

mistralrs-paged-attn/src/cuda/flashinfer · high confidence

Introduce mid-stream grammar enforcement and strict tool call modes

The tool-calling subsystem has been refactored to support mid-stream constrained decoding via new grammar helpers in \grammar.rs\ and a strategy-based architecture in \strategy.rs\. This enables the engine to enforce tool-call syntax dynamically during generation, including a new 'strict' mode that validates tool arguments against their JSON schemas using \anyOf\ variants. The system now supports multiple model-specific formats (Llama, Mistral Nemo, Harmony, etc.) through dedicated parsers and introduces a \ToolChoice::Required\ mode that tracks obligations and forces a tool call if the generation nears the token deadline without one.

mistralrs-core/src/tools · high confidence

Introduction of DummyLayer for UQFF placeholder handling

A new \DummyLayer\ implementation has been added to the quantization module to serve as a temporary placeholder for UQFF (Unified Quantization Format) files. This layer implements the \QuantMethod\ interface but is designed to fail with descriptive errors if used for actual inference operations like dequantization, forward passes, or LoRA delta application, ensuring that these temporary placeholders are correctly replaced before model execution.

mistralrs-quant/src/dummy · high confidence

Introduction of UQFF v1.2 format for unquantized layer serialization

The UQFF (Unquantized Quantization Format) module has been updated to version 1.2.0, introducing improved serialization and deserialization capabilities for unquantized layers and handling of AFQ fallbacks. This change adds a new \UqffReader\ for loading artifacts, a \Tracker\ system to monitor module quantization types and promotion policies, and a comprehensive reporting system (\report.rs\) that generates detailed JSON reports on quantization issues, layer details, and fallbacks. The format now enforces version tags, rejecting pre-1.0 artifacts, and supports sharding logic for weights and biases.

mistralrs-quant/src/uqff · high confidence

New README and CUDA build system using cudaforge

The mistralrs-core crate now includes a README.md and a comprehensive build.rs script that manages CUDA kernel compilation. The build system uses the cudaforge library to compile CUDA sources, supports custom NVCC flags via environment variables, and handles platform-specific linking (e.g., .lib on Windows, .a on Linux). It also introduces conditional compilation for specific CUDA features like FP8 producers and SM90 kernels based on detected CUDA toolkit versions and compute capabilities.

mistralrs-core · high confidence

New core utility modules for logging, memory, and model configuration

The \mistralrs-core/src/utils\ directory has been restructured into distinct modules to improve code organization and functionality. \debug.rs\ introduces a new verbosity-controlled logging system (\LogVerbosity\) and device representation helpers, allowing users to control log detail levels via the \MISTRALRS\_DEBUG\ environment variable. \memory\_usage.rs\ provides a unified \DeviceMemory\ abstraction that accurately queries memory for both discrete (CUDA) and unified (Metal, CPU) architectures, including specific handling for integrated GPUs. \normal.rs\ implements the \ModelDType\ enum with an \Auto\ mode that automatically selects the best data type (BF16, F16, or F32) based on hardware capabilities (e.g., CUDA compute capability) and runtime probes. \model\_config.rs\ refactors model loading parameters into a structured \ModelParams\ type system, supporting both quantized and adapter-based (LoRA/XLora) configurations. Additionally, \gguf\_metadata.rs\ adds robust parsing for GGUF model metadata, \progress.rs\ enhances loading feedback with parallel-aware progress bars, and \tokenizer.rs\/\tiktoken.rs\ improve tokenization support including Tekken and tiktoken formats.

mistralrs-core/src/utils · high confidence

New topology system for layer-specific device and quantization mapping

The core topology module has been replaced with a new system that allows users to specify per-layer device placement (CPU, CUDA, or Metal) and integer quantization (ISQ) settings via a structured configuration format. This change introduces support for defining mappings by specific layer indices, ranges (e.g., 0-10), or regex patterns, enabling fine-grained control over how model layers are distributed across hardware and quantized, replacing the previous generic topology handling.

mistralrs-core/src/topology · high confidence

New web UI with agent tool approvals and code execution support

The web UI has been rebuilt using Svelte, introducing a complete interface overhaul that includes an AgentApproval component for user-controlled approval of tool executions (such as shell commands and code), dedicated display components for code execution results and file outputs, and updated chat input handling that sends images as data URLs.

mistralrs-cli/webui · high confidence

Optimized CPU attention with SIMD kernels and new LoRA adapter management

The core inference engine now features significantly faster CPU attention paths, introducing AVX512 and NEON SIMD micro-kernels for dot products and online-softmax operations to accelerate x86 and ARM processors. Additionally, a new LoRA adapter management system has been introduced, providing runtime budgeting for resident adapters and a flexible selection mechanism that supports aliasing, generation-based targeting, and pinning for efficient adapter handling during inference.

mistralrs-core/src · high confidence

Optimized CUDA RoPE kernels

The rotary embedding implementation in the quantization module has been optimized for CUDA, introducing new FFI bindings and a custom operator to improve performance. This change enhances the efficiency of RoPE calculations on supported hardware.

mistralrs-quant/src/rotary · medium confidence

Optimized CUDA RoPE kernels with ROCm compatibility

The rotary embedding (RoPE) kernels in the mistralrs-quant module have been optimized for CUDA execution and extended to support ROCm. A new compatibility header introduces abstractions for CUDA intrinsics (such as load and shuffle operations) that map to their HIP equivalents, enabling the same kernel logic to run on AMD hardware. The implementation includes specialized kernels for both standard and GPT-NeoX style rotary embeddings, supporting FP16, BF16, and FP32 data types, which improves performance for attention mechanisms on supported accelerators.

mistralrs-quant/kernels/rotary · high confidence

Optimized cross-GPU tensor transfers via CUDA Peer Access

The device mapping layer now utilizes CUDA Peer-to-Peer (P2P) access to accelerate tensor movement between GPUs. When multiple CUDA devices are detected, the system attempts to enable direct P2P communication; if successful, tensors are transferred directly between GPUs, bypassing the CPU. If P2P is unavailable or unsupported, transfers fall back to staging through system memory. This optimization also includes pre-creating attention masks for each unique device to avoid repeated allocations during inference loops.

_mistralrs-core/src/device\map · high confidence

Refactored KV cache system with new cache types and hybrid support

The KV cache implementation has been restructured into distinct modules (single, rotating, full, and hybrid) to support more flexible caching strategies. A new HybridCache type enables vLLM-style continuous batching for models mixing attention and recurrent layers (like GraniteMoeHybrid and Qwen3 Next) using pool-based state with indexed access. The system now includes a unified EitherCache enum to handle Normal, Full, and Hybrid cache variants, and introduces snapshot/restore capabilities for better state management during sequence cloning and rollback operations.

_mistralrs-core/src/kv\cache · high confidence

Refactored attention backend architecture with new FlashAttention and Sinks implementations

The attention computation logic has been reorganized into a modular backend system located in \mistralrs-core/src/attention/backends\. This change introduces dedicated modules for FlashAttention (\flash.rs\), Metal-optimized FlashAttention (\metal\_flash\_attn.rs\), naive SDPA (\naive.rs\), and a new fused 'Sinks' attention mechanism (\sinks.rs\) that supports per-head sinks for improved performance on CUDA and Metal devices. The refactoring centralizes dispatch logic, adds support for variable-length sequences and sliding windows in FlashAttention, and ensures correct memory synchronization for low-memory devices in the naive path.

mistralrs-core/src/attention/backends · high confidence

Refactored attention module with new mask abstraction and chunked computation

The attention module has been restructured to introduce a new \AttentionMask\ enum that explicitly distinguishes between no masking, flash attention (causal), and custom tensor masks, replacing the previous implicit handling. This change is accompanied by the addition of \chunked\_attention\ and \chunked\_attention\_with\_offset\ functions, which split long sequence attention computations into 1024-token chunks to prevent out-of-memory errors during inference. These changes improve memory stability for long contexts and provide a clearer interface for different attention backends.

mistralrs-core/src/attention · high confidence

Release v0.9.0 benchmark data and v0.8.2 performance report added

The releases directory now includes a comprehensive benchmark report for v0.8.2, documenting mistral.rs performance against llama.cpp and vLLM on NVIDIA GB10 and B200 hardware for the Gemma 4 E4B model. Additionally, raw benchmark metadata and detailed CPU affinity test results for the Qwen3-4B model have been added under the v0.9.0 release artifacts, providing granular throughput data across various quantization levels and thread configurations.

releases · high confidence

Revamped LoRA support with new internal architecture

The LoRA implementation in \mistralrs-core/src/lora\ has been completely rewritten, introducing new \LoraLinear\ and \QLoraLinear\ structs to handle adapter application and weight merging. This change replaces the previous ordering system with a new internal structure that manages adapter scaling, stacking, and merging logic directly within these modules, fundamentally altering how LoRA adapters are processed and applied to linear layers.

mistralrs-core/src/lora · high confidence

Unified MoE experts layer with multiple backend support

The MoE experts implementation has been refactored into a unified module that automatically selects the optimal execution backend based on the model configuration, device, and quantization settings. Users benefit from improved performance and compatibility as the system now supports specialized backends including fused CUDA kernels, cuTile blockwise FP8 grouped GEMM, CUTLASS grouped GEMM, and gather-based execution for CPU/Metal/ISQ. The new checkpoint loader automatically detects and handles various on-disk weight layouts (such as fused gate/up projections or per-expert naming conventions like Mixtral's w1/w3/w2), ensuring seamless loading of diverse MoE model architectures without manual intervention.

mistralrs-core/src/moe/experts · high confidence

Unified automatic device mapping and loader architecture

The model loading system has been refactored to support automatic device mapping across all model types, allowing the library to automatically distribute model layers across available GPUs and CPU based on memory availability. This change introduces a new \auto\_device\_map\ module that calculates memory requirements and assigns layers, a \checkpoint\_inventory\ module for efficient safetensor size analysis, and unified loader traits (\DeviceMappedModelLoader\) for normal, multimodal, embedding, and diffusion models. Users benefit from improved memory management and the ability to load larger models that exceed single-GPU memory without manual configuration.

mistralrs-core/src/pipeline/loaders · high confidence

Vendored FlashAttention SM90 subset for FP8 paged decode

The repository now includes a vendored subset of FlashAttention (derived from vLLM's fork at commit f3e1a4f) and a pinned version of NVIDIA CUTLASS (commit 62750a2) in the third\_party/flash-attention directory. This source closure provides the transitive dependencies required by the new FP8 paged decode provider, specifically targeting SM90 (Hopper) architectures. The vendored code has been modified to remove Torch header dependencies, handle CUDA errors via the C ABI, and retain only the noncausal forward and BF16 hdim-256 combine instantiations used by this provider.

_mistralrs-paged-attn/third\party · high confidence

Test coverage

Added CUDA integration tests for paged attention and FlashInfer GQA; Added CUDA test coverage for quantized MoE, LoRA, and NVFP4 kernels; Added integration tests for image normalization and resizing transforms; Added tests for sandboxed code execution sessions.

Dependencies

Mistral.rs v0.9.4 release with new workspace structure and dependency updates

This release updates the project to version 0.9.4 and restructures the codebase into a Cargo workspace containing multiple crates (mistralrs-core, mistralrs-cli, mistralrs-server-core, etc.). It upgrades the Rust minimum version to 1.94, updates the Candle backend to commit 66a8cf1, and refreshes various dependencies including axum, serde, and clap. The web UI and documentation sites are also updated with new dependencies like Svelte 5, Tailwind CSS 4, and Astro 7.

(dependencies) · high confidence

Housekeeping

Added third-party attribution and license files for FlashInfer and DeepGEMM SM90 kernels

This change adds the Apache 2.0, BSD-3-Clause, and MIT license texts, along with NOTICE and README attribution files, for the FlashInfer GDN, FlashInfer radix top-k, and DeepGEMM SM90 providers. These files document the upstream sources (FlashInfer, DeepGEMM, CUTLASS, vLLM) and their pinned revisions, clarifying the licensing and provenance of the adapted CUDA kernels in mistralrs-core and mistralrs-quant.

_mistralrs-core/third\_party, mistralrs-quant/third\party · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 63 → 66 (+2.9)
  • Rubric changed (rubric-2026.09.9 → rubric-2026.09.17) — scores are not directly comparable.

Lenses

  • Code Health 78 → 77 (-0.3)
  • Architecture 99 → 95 (-3.8)
  • Maturity 66 → 68 (+2.1)
  • Readiness 77 → 64 (-13.4)
  • Security 56 → 68 (+12.0)
  • Event Sourcing 100 → 100 (+0.0)
  • Accessibility 63 → 63 (+0.0)
  • Performance 100 (new)

Resolved (60)

  • Duplicated block (12 lines × 2) (mistralrs-quant/src/cutile/nvfp4_glu.rs)
  • High CVE: [GHSA redacted] (mistralrs-cli/webui/package-lock.json)
  • High CVE: [GHSA redacted] (Cargo.lock)
  • High CVE: [GHSA redacted] (mistralrs-cli/webui/package-lock.json)
  • Hotspot: mistralrs-cli/src/commands/doctor.rs (mistralrs-cli/src/commands/doctor.rs)
  • Hotspot: mistralrs-cli/src/commands/quant/gguf_discovery.rs (mistralrs-cli/src/commands/quant/gguf_discovery.rs)
  • Hotspot: mistralrs-cli/src/commands/quantize.rs (mistralrs-cli/src/commands/quantize.rs)
  • Hotspot: mistralrs-cli/src/commands/tune.rs (mistralrs-cli/src/commands/tune.rs)
  • Hotspot: mistralrs-cli/webui/src/lib/services/streaming.ts (mistralrs-cli/webui/src/lib/services/streaming.ts)
  • Hotspot: mistralrs-core/src/attention/backends/cpu/mask.rs (mistralrs-core/src/attention/backends/cpu/mask.rs)
  • Hotspot: mistralrs-core/src/cuda/speculative_rejection.rs (mistralrs-core/src/cuda/speculative_rejection.rs)
  • Hotspot: mistralrs-core/src/diagnostics.rs (mistralrs-core/src/diagnostics.rs)
  • Hotspot: mistralrs-core/src/engine/cuda_memory.rs (mistralrs-core/src/engine/cuda_memory.rs)
  • Hotspot: mistralrs-core/src/layers.rs (mistralrs-core/src/layers.rs)
  • Hotspot: mistralrs-core/src/models/hunyuan_v1_moe.rs (mistralrs-core/src/models/hunyuan_v1_moe.rs)
  • Hotspot: mistralrs-core/src/moe/experts/backends.rs (mistralrs-core/src/moe/experts/backends.rs)
  • Hotspot: mistralrs-core/src/moe/experts/config.rs (mistralrs-core/src/moe/experts/config.rs)
  • Hotspot: mistralrs-core/src/pipeline/auto.rs (mistralrs-core/src/pipeline/auto.rs)
  • Hotspot: mistralrs-core/src/pipeline/paths.rs (mistralrs-core/src/pipeline/paths.rs)
  • Hotspot: mistralrs-core/src/reasoning_parsers/harmony.rs (mistralrs-core/src/reasoning_parsers/harmony.rs)
  • …and 40 more

New (195)

  • Duplicated block (11 lines × 2) (scripts/production_soak.py)
  • Duplicated block (12–13 lines × 3) (examples/server/anthropic_chat.py)
  • Duplicated block (13 lines × 2) (examples/server/anthropic_agentic.py)
  • Duplicated block (14 lines × 2) (scripts/production_soak.py)
  • Duplicated block (14 lines × 2) (scripts/production_soak.py)
  • Duplicated block (15 lines × 2) (scripts/production_soak.py)
  • Duplicated block (15 lines × 2) (scripts/production_soak.py)
  • Duplicated block (16 lines × 2) (scripts/production_soak.py)
  • Duplicated block (16 lines × 2) (scripts/production_soak.py)
  • Duplicated block (17 lines × 2) (scripts/production_soak.py)
  • Duplicated block (18–21 lines × 3) (scripts/production_soak.py)
  • Duplicated block (19 lines × 2) (examples/python/custom_search.py)
  • Duplicated block (20 lines × 2) (scripts/production_soak.py)
  • Duplicated block (21 lines × 2) (scripts/production_soak.py)
  • Duplicated block (21 lines × 24) (examples/server/chat.py)
  • Duplicated block (22–25 lines × 2) (scripts/production_soak.py)
  • Duplicated block (23 lines × 2) (scripts/production_soak.py)
  • Duplicated block (23 lines × 2) (scripts/production_soak.py)
  • Duplicated block (25 lines × 3) (examples/python/shell_skills.py)
  • Duplicated block (26 lines × 2) (scripts/production_soak.py)
  • …and 175 more

Changes since last survey

  • 12 commits — 3 feature/other, 9 fixes

By area

  • (root) — 3 commits
  • mistralrs-core/src — 3 commits
  • mistralrs-quant/src — 2 commits
  • docs/openapi.json — 1 commit
  • mistralrs-audio/src — 1 commit
  • mistralrs-cli/src — 1 commit
  • mistralrs-server-core/src — 1 commit

Notable commits

  • fix: chore(deps): fix open Dependabot security alerts (#2444)
  • fix: fix(core): only honor tool-call and reasoning delimiters emitted as special tokens (#2455)
  • fix: fix(core): read the CUDA graph instantiate flag through the cudarc 0.19.10 newtype (#2454)
  • fix: fix(quant): restore CUDA builds and inference after the candle bump (#2451)
  • fix: fix(search): share the remote fetch guard with the web search and extraction tools (#2449)
  • fix: fix(server): accept only a bare file name for calibration save_cimatrix (#2445)
  • fix: fix(server): bound decoded GIF frames and audio samples (#2448)
  • fix: fix(server): judge IPv6 transition addresses by their embedded IPv4 in the media guard (#2446)
  • fix: fix(ui): validate chat ids before building chat file paths (#2447)
  • change: chore(deps): bump candle to e65eb1de and cutile to 0.3.1 (#2450)
  • change: chore(release): 0.9.4 finalization (#2452)
  • change: chore: add SECURITY.md

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

EricLBuehler/mistral.rs was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 2370966bb91e2e3dafa0b1521b87c50fd5c01244 — the exact code this score is about.
  • Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-705631bb727e.