vllm-project/vllm
61.9
Adequate · 11 October 2026
1075.5k
lines of production code
Python
with Rust
3
measurements over time
What this system is
This system is vLLM, a high-throughput inference engine for large language models that supports serving, offline inference, and specialized pooling tasks like embedding and reranking. It provides extensive hardware acceleration across NVIDIA, AMD, Intel, and ARM platforms through custom CUDA/HIP kernels, quantization support, and disaggregated serving capabilities. The project also includes a Rust-based frontend for low-latency request handling and a comprehensive benchmarking suite for performance profiling.
How it got here
2023–2025 — Hardware acceleration and quantization expansion
42 changes.
This period focused on expanding vLLM's hardware support and performance through extensive low-level kernel development for CUDA, ROCm, and various CPU architectures. Significant work was done to introduce native support for FP8, BF16, and 4-bit quantization formats, alongside new C++ infrastructure for memory management and KV caching. The effort also included major overhauls to the build system, CI pipelines, and benchmarking suites to ensure stability and efficiency across diverse deployment environments.
2026 — Rust frontend and stable ABI migration
111 changes.
This period focused on introducing a comprehensive Rust-based inference frontend, including chat, tokenizer, and parser components, while migrating core CUDA kernels to the libtorch stable ABI for improved binary compatibility. The work also expanded the example suite with extensive demonstrations for multimodal, tool-calling, and disaggregated serving, alongside significant additions to benchmarking and CI infrastructure.
Features
Add BFCL tool-call correctness evaluation script
A new CI script (.buildkite/scripts/tool\_call/run-bfcl-eval.sh) has been added to run Berkeley Function Call Leaderboard (BFCL) evaluations against a local vLLM server. The script automates starting a vLLM instance with specific tool-calling configurations (such as prefix caching and auto-tool-choice), patches the BFCL evaluation library to target the local server, and executes correctness tests for tool calls. It supports configurable parameters including model selection, API type (chat completions or responses), test categories, and temperature, defaulting to greedy decoding (temperature 0.0).
_.buildkite/scripts/tool\call · high confidence
Add CPU attention and fused MoE micro-benchmarks
New benchmark scripts have been added to the CPU kernels directory to measure performance of specific inference components. The new \benchmark\_cpu\_attn.py\ script allows users to profile CPU attention kernels, supporting features like FP8 KV cache, sliding windows, and ISA-specific optimizations (AMX, RISC-V RVV). Additionally, \benchmark\_cpu\_fused\_moe.py\ provides a way to benchmark the fused Mixture-of-Experts kernel, reporting metrics such as throughput (TFLOP/s) and latency for various expert configurations and activation functions.
benchmarks/kernels/cpu · high confidence
Add CUTLASS extensions for vLLM custom quantized types and utilities
This change introduces new header and Python files in the \csrc/cutlass\_extensions\ directory to support vLLM-specific quantization formats. It defines custom biased integer subbyte types (\vllm\_uint4b8\_t\ and \vllm\_uint8b128\_t\) in C++, provides corresponding Python enumerations and type mappings in \vllm\_cutlass\_library\_extension.py\, and adds utility functions for layout permutation and pointer handling in \cute\_utils.cuh\ and \vllm\_type\_utils.cuh\. These additions enable the underlying CUTLASS kernels to correctly handle and map these specific 4-bit and 8-bit quantized data types.
_csrc/cutlass\extensions · high confidence
Add DeepGEMM FP8 Block Dense GEMM benchmark for Hopper GPUs
A new benchmark script and documentation have been added to compare DeepSeek's DeepGEMM block FP8 kernels against vLLM's existing Triton and CUTLASS implementations. This tool allows users to evaluate performance (TFLOPS, bandwidth) and accuracy on Hopper architecture GPUs, providing a side-by-side analysis of the new DeepGEMM integration.
benchmarks/kernels/deepgemm · high confidence
Add DeepSeek V3.2, V4, and V4.1 DSML tool-call parsers
The Rust frontend now includes dedicated tool parsers for DeepSeek V3.2, V4, and V4.1 models that recognize the DSML markup format. V3.2 uses \<|DSML|function\_calls\> framing, V4 uses \<|DSML|tool\_calls\>, and V4.1 uses spaced tags (e.g., \<|DSML| calls\>). The parsers preserve special tokens, emit tool calls only after a complete invoke block, and handle parameter typing via the string attribute (string="true" forces strings; string="false" or missing lets schema-based conversion apply, including preserving the literal "null" string). Streaming is supported, including handling markers split across chunks, and whitespace framing is preserved.
_rust/src/parser/src/tool/deepseek\dsml · high confidence
Add DeepSeek V4.1 chat renderer to Rust frontend
The Rust chat frontend now includes a dedicated renderer for the DeepSeek V4.1 model. This renderer handles the model's specific prompt formatting, including reasoning effort controls (mapped to a numeric budget of 1–100), tool invocation schemas, and multimodal image placeholders. It also ensures that media order is correctly tracked even when tool results are reordered or historical developer content is dropped.
_rust/src/chat/src/renderer/deepseek\v41 · high confidence
Add GLM-4.5/4.6 and GLM-4.7 MoE XML tool parsers
Users can now parse XML-style tool calls for GLM-4.5/4.6 and GLM-4.7 Mixture-of-Experts models. The new parsers handle argument extraction, whitespace preservation, and schema-based type conversion (e.g., integers, booleans, JSON objects) from the model's output. GLM-4.7 specifically supports a more flexible function-name separator, allowing the name to be followed by whitespace, a newline, or directly by the first argument tag.
_rust/src/parser/src/tool/glm\xml · high confidence
Add GPTQ quantization CUDA kernels to libtorch\_stable
New CUDA source files (compat.cuh, matrix\_view.cuh, q\_gemm.cu, qdq\_2/3/4/8.cuh, qdq\_util.cuh) are added to csrc/libtorch\_stable/quantization/gptq, providing device-side utilities and kernels for GPTQ quantization (2/3/4/8-bit dequantization, matrix views, and GEMM) adapted from exllamav2/exllama. This extends the libtorch\_stable ABI with GPTQ activation kernels, enabling GPTQ quantized inference within the stable libtorch build.
_csrc/libtorch\stable/quantization/gptq · high confidence
Add Hugging Face response template parser for structured output
The unified parser now supports Hugging Face \response\_template\ definitions, enabling structured parsing of model outputs (such as tool calls, reasoning, and content) based on declarative templates from \tokenizer\_config.json\. This implementation mirrors the Transformers \chat\_parsing\ protocol, streaming events for text, reasoning, and tool calls while handling complex patterns like XML-inline and key-value line parsing.
rust/src/parser/src/unified/hf · high confidence
Add Kimi K3 XTML chat renderer
A new native Rust renderer for the Kimi K3 model has been added to the chat frontend. It implements the model's specific XTML prompt format, including support for reasoning controls (thinking effort), dynamic tool declarations from developer messages, and media placeholders. The renderer includes comprehensive golden-fixture tests to ensure the generated prompts match the expected token sequences and structure.
_rust/src/chat/src/renderer/kimi\k3 · high confidence
Add Prithvi/Terratorch geospatial pooling examples
Added three new example scripts in the pooling plugin directory demonstrating offline and online inference with the Prithvi-EO model and Terratorch integration. The examples show how to process geospatial TIFF images using the io\_processor\_plugin, including handling base64-encoded outputs and configuring the LLM for multimodal embedding tasks.
examples/pooling/plugin · high confidence
Add QuickReduce all-reduce implementation for ROCm
Introduces a new QuickReduce all-reduce communication primitive for ROCm devices, supporting full-precision (FP16/BF16) and symmetric integer quantization (INT8, INT6, INT4, and INT3). The implementation includes device-side kernels and codec logic for optimized collective communication, with INT3 quantization restricted to two-way tensor parallelism (world\_size == 2) due to performance overhead on larger groups. It also enables support for the gfx1250 ROCm architecture and includes fixes for handling different input shapes and accumulation errors in CUDA graph mode.
csrc/quickreduce · high confidence
Add Rust chat smoke test example for external vLLM engines
A new Rust example (\external\_engine\_chat\_qwen\) and its accompanying README have been added to demonstrate how to smoke-test the Rust chat facade against an external vLLM engine. This example connects to a headless vLLM instance via a handshake protocol, sends a text prompt, and streams structured assistant events (including reasoning blocks and text), while also reporting output token counts and finish reasons.
rust/src/chat/examples · high confidence
Add Rust multimodal preprocessing for audio, video, and images
The Rust chat frontend now supports full multimodal input handling. It can preprocess audio clips, video clips, and image frames using the \llm-multimodal\ library, converting them into batched or flat tensor features for the engine. The implementation includes prompt placeholder expansion to replace media markers with model-specific tokens, validation of preprocessed features, and zero-copy tensor slicing for efficient data transfer. This enables users to send audio, video, and image content in chat requests via the Rust frontend.
rust/src/chat/src/multimodal · high confidence
Add SGLang CPU kernels for GDN, MLA, and attention operations
Introduces a new set of CPU-optimized kernels adapted from the SGLang project, including fused GDN (Gated Delta Networks) support with AMX acceleration, MLA (Multi-Latent Attention) backends for DeepSeek models, and extend/decode attention kernels. These additions enable high-performance speculative decoding, prefix caching, and efficient handling of FP8/INT4 quantized weights on CPU architectures.
csrc/cpu/sgl-kernels · high confidence
Add bfloat16 and FP8 support to attention kernels
The attention kernels in csrc/attention now support bfloat16 and FP8 data types in addition to the existing float16 and float32. This includes new device-side vector types and arithmetic operations for bfloat16 (bf16\_4\_t, bf16\_8\_t) and FP8 (uint8\_t based vectors), as well as an enum and helper function to select between FP8 E4M3 and E5M2 formats for the KV cache. These changes enable users to leverage lower-precision data types for improved performance and memory efficiency in attention computations.
csrc/attention · high confidence
Add dynamic 4-bit MoE path for ARM64
A new CPU implementation for dynamic 4-bit MoE inference has been added, specifically targeting AArch64 (ARM64) architectures. This change introduces a new source file that leverages the ARM-specific \\_dyn\_quant\_matmul\_4bit\ kernel to perform efficient matrix multiplications. The implementation handles token bucketing, expert routing, and activation functions (SwiGLU variants), enabling dynamic 4-bit quantized MoE operations on ARM hardware where this specific kernel is available.
csrc/moe · high confidence
Add fused CUDA kernel for GDN post-conv MTP decode
A new fused CUDA kernel (\fused\_gdn\_decode\_kernel.cu\) has been added to the GDN module to accelerate the post-convolution multi-token prediction (MTP) decode step. This implementation optimizes memory access patterns using shared memory and asynchronous copies, supporting specific data types (bfloat16, float16, float) and bias configurations to improve inference performance for compatible models.
_csrc/libtorch\stable/gdn · high confidence
Add kernel microbenchmark skill for GPU performance analysis
A new agent skill named 'kernel-microbenchmark' has been added to help build, debug, and interpret vLLM GPU kernel microbenchmarks. It provides a structured workflow and reference numbers for CUDA, ROCm/HIP, Triton, and CuteDSL, including specific guidance on using FlashInfer CUPTI timing for CUDA and HIP graph replay for ROCm to avoid cache-related measurement errors. The skill includes example benchmark scripts for single-GPU CUPTI timing, platform-agnostic graph replay, and multi-GPU GEMM with reduce-scatter, ensuring users can accurately measure throughput (TFLOPS/GB/s) and verify correctness before timing.
.agents/skills/kernel-microbenchmark · high confidence
Add native Harmony output processor for GPT-OSS models
Introduces a new \HarmonyChatOutputProcessor\ in the chat output pipeline that consumes token IDs directly and uses the official \openai-harmony\ parser to reconstruct structured assistant messages at the token level. This processor replaces the generic text-first parsing approach for \gpt\_oss\ models, enabling support for features like parallel tool calls and dynamic tool definitions while ensuring that interrupted or incomplete messages are preserved correctly in the final output.
rust/src/chat/src/output/harmony · high confidence
Add native Harmony renderer for GPT-OSS chat
The Rust chat frontend now includes a native Harmony renderer specifically for the GPT-OSS model. This renderer handles prompt formatting and tokenization using the official Harmony encoding, supporting features such as reasoning effort controls, tool usage (including round-trips), and developer-defined tools. It also manages system message handling, including the ability to move leading instructions into the system identity via environment configuration, and ensures proper message history handling like dropping stale analysis messages.
rust/src/chat/src/renderer/harmony · high confidence
Add native Inkling chat renderer
Introduces the InklingChatRenderer, a new native component that formats chat interactions into token sequences using Inkling-specific special tokens. This renderer supports multimodal inputs (text, images, and audio), reasoning content, and tool usage (declarations and invocations), ensuring that user messages and model responses are correctly structured for the Inkling model architecture.
rust/src/chat/src/renderer/inkling · high confidence
Add structural-tag grammar support for HY tool calls
Introduces a new structural-tag builder for the HY dialect, enabling the parser to generate and enforce XML-style structural markers for tool calls. This change adds the \HyStructuralTagBuilder\ implementation which handles argument key separation (required vs optional), tool call formatting with specific start/end markers, and supports different HY dialect versions (V3 and V4) to ensure consistent tool invocation syntax.
rust/src/parser/src/tool/hy · high confidence
Add structural-tag grammar support for Kimi K3 model
Users can now generate responses from the Kimi K3 model using structural tags. This change introduces a new Rust frontend implementation for Kimi K3 that defines a specific XTML-based grammar for tool calls and arguments. The implementation ensures that reserved markers are excluded from response text to prevent parsing conflicts and supports both reasoning modes and forced/automatic tool choices via the XGrammar structural-tag builder.
_rust/src/parser/src/unified/kimi\k3 · high confidence
Add tool parsers for DeepSeek V3/V3.1, HY, Kimi K2, MiMo, and MiniMax M2/M3
The Rust tool parser module now includes dedicated parsers for several new model-specific tool-call formats. DeepSeek V3 and V3.1 parsers handle JSON-fenced and raw JSON arguments respectively, streaming them as raw deltas without schema conversion. HY parsers support V3 and V4 XML-style dialects with schema-based parameter conversion. Kimi K2 parser handles token-delimited calls with raw JSON arguments. MiMo parser uses Qwen Coder's XML grammar with compact tags. MiniMax M2 and M3 parsers handle XML-style calls, with M3 supporting recursive namespace-delimited parameters and schema conversion. All parsers implement the unified ToolParser interface and provide structural tag builders for grammar generation.
rust/src/parser/src/tool · high confidence
Add vLLM IR op benchmarking infrastructure
A new benchmarking harness has been added at benchmarks/kernels/ir to measure the performance of vLLM IR operations. The tool allows users to run benchmarks for specific ops (such as rms\_norm and fused\_add\_rms\_norm) across various shape configurations and data types, supporting both CUDA graph and eager execution modes. Results are collected and analyzed to provide timing metrics and speedup comparisons against the native implementation.
benchmarks/kernels/ir · high confidence
Added AWQ 4-bit GEMM kernel and dequantization utilities to libtorch\_stable
This change introduces the CUDA implementation for Activation-aware Weight Quantization (AWQ) into the stable Torch C++ API. It adds a new \dequantize.cuh\ header providing a device-side function to convert 4-bit integers to FP16, and a \gemm\_kernels.cu\ file containing the \gemm\_forward\_4bit\_cuda\ kernel. This kernel performs matrix multiplication on 4-bit quantized weights by dequantizing them on-the-fly using scaling factors and zero-points, enabling efficient inference for models using AWQ quantization within the stable ABI.
_csrc/libtorch\stable/quantization/awq · high confidence
Added OpenAI batch inference example and documentation
The \examples/features/openai\_batch\ directory now includes a README guide and a sample \openai\_example\_batch.jsonl\ file. This addition provides users with concrete examples of how to perform offline batch inference using the OpenAI batch file format, covering local file execution, remote URL inputs, and AWS S3 integration via presigned URLs.
_examples/features/openai\batch · high confidence
Added ReplaySSM end-to-end decode benchmark
A new benchmark script (\e2e\_decode\_speedup.py\) has been added to the \benchmarks/replayssm\ directory to measure the performance of the ReplaySSM implementation against the standard SSM kernel. This tool runs end-to-end autoregressive decode tests on hybrid Mamba2 models, allowing users to compare per-step latency and throughput between the two modes via separate subprocesses with clean CUDA contexts.
benchmarks/replayssm · high confidence
Added automated skill for upgrading Transformers and removing vendored code
A new agent skill has been added to automate the process of upgrading the pinned Transformers version and performing the subsequent 'unvendoring pass' to remove redundant vLLM code. This skill includes a playbook and helper scripts (find\_upstreamed, dump\_configs, dump\_model, diff\_dumps) that inventory vendored configs and processors now provided by upstream, identify version-gated code, and verify model behavior changes via config and model dumps, streamlining the maintenance of vLLM's compatibility with newer Transformers releases.
.agents/skills/transformers-upgrade · high confidence
Added chat templates for new reranking models
Added Jinja chat templates for several reranking models to support specific formatting requirements. This includes templates for BGE Reranker v2 Gemma, mxbai\_rerank\_v2, Nemotron rerank, and Qwen3 reranker. Additionally, templates for Qwen3-VL and Nemotron VL rerankers were added to handle multimodal inputs, including logic to manage image tokens and clean content. A template for zerank-2 was also introduced, supporting optional instructions and specific system/user turn formatting.
examples/pooling/score/template · high confidence
Added multimodal embedding templates for Qwen2-VL, Nemotron, and Phi-3V
New Jinja templates have been added to the pooling/embedding examples to support multimodal (text and image) embedding models. These templates (dse\_qwen2\_vl, nemotron\_embed\_vl, vlm2vec\_phi3v, vlm2vec\_qwen2vl) handle the formatting of mixed content types, including image placeholders and role prefixes, ensuring that embedding models receive correctly structured inputs for vision-language tasks.
examples/pooling/embed/template · high confidence
Adds fused Kimi-K3 KDA and AttnRes kernels for CUDA and ROCm
This change introduces new CUDA and ROCm GPU kernels to accelerate the Kimi-K3 model. For NVIDIA (CUDA) hardware, it adds the \attn\_res\_kernel\ for Blackwell (SM100) architectures, implementing a warp-specialized online softmax with residual and RMSNorm, and replaces the previous Triton-based dispatch with fused KDA (Kron-Diagonal Attention) kernels for both prefill and decode phases. For AMD (ROCm) hardware, it adds fused KDA kernels for prefill and decode that replace the prior Triton implementation, consolidating causal conv1d updates, gated delta-rule recurrences, and output RMSNorm into single launches to reduce HBM traffic and kernel launch overhead.
_csrc/libtorch\_stable/kimi\k3 · high confidence
Automated XPU Docker image and Triton shim publishing
New Buildkite scripts have been added to automate the release and nightly publishing of Intel XPU artifacts. The pipeline now creates multi-architecture ECR manifests for the main vLLM XPU Docker image, publishes the XPU Triton shim wheel (v3.8.0) to an S3-hosted index with integrity verification, and pushes nightly Docker tags to the public ECR repository.
.buildkite/scripts/xpu · high confidence
Benchmark for fused FP8 output quantization in merge\_attn\_states
Added a new benchmark script (\merge\_attn\_states\_benchmarks.py\) that compares the performance of fused versus unfused implementations for merging attention states with FP8 output quantization. The benchmark evaluates four approaches—fused CUDA, fused Triton, unfused CUDA, and unfused Triton—across various tensor parallelism sizes, token counts, and input data types to help users understand the performance trade-offs of these specific kernel configurations.
_benchmarks/fused\kernels · high confidence
CI pipeline now records GPU utilization and memory telemetry alongside test timelines
The CI pipeline now automatically samples NVIDIA GPU utilization and memory usage for every test command, providing visibility into hardware resource consumption during test runs. This is achieved through new OpenTelemetry helpers in \.buildkite/scripts/ci-otel\ that inject a pytest timing plugin and run a background GPU sampler (\ci\_gpu.py\) which queries device metrics via \nvidia-smi\ or NVML bindings without importing heavy frameworks like PyTorch. The telemetry is attached to command spans and uploaded via OTLP, allowing the Builds timeline dashboard to display GPU controls alongside job and command rows. The implementation includes safeguards such as dropping root privileges for the sampler, handling MIG (Multi-Instance GPU) memory metrics, and ensuring sampling can be disabled via the \CI\_INFRA\_GPU\_SAMPLING\ environment variable.
.buildkite/scripts/ci-otel · high confidence
DeepSeek V3.2 renderer now supports developer messages with tools and respects generation prompt options
The new DeepSeek V3.2 prompt renderer (rust/src/chat/src/renderer/deepseek\_v32) implements the official Hugging Face encoding logic, adding support for developer-role messages that carry tool definitions. It also honors the continue\_final\_message and add\_generation\_prompt chat options, ensuring the final assistant turn is rendered correctly and generation prompts are applied as configured.
_rust/src/chat/src/renderer/deepseek\v32 · high confidence
DeepSeek V4 chat renderer with tool-use and reasoning support
The Rust frontend now includes a dedicated DeepSeek V4 chat renderer that formats prompts with DeepSeek-specific tokens, supporting tool-use schemas, reasoning content, and developer messages. This renderer is accompanied by test fixtures and assertions that verify correct prompt generation for tool calls, multi-turn conversations, and developer tool definitions.
_rust/src/chat/src/renderer/deepseek\v4 · high confidence
Hugging Face chat template renderer now supports multimodal content, structured output, and custom roles
The Rust frontend's Hugging Face chat template renderer has been rewritten to support advanced template features and content formats. Users can now render multimodal content (images, video, audio) via dedicated placeholder tokens, and the renderer respects the \{% generation %}\ block syntax used by Hugging Face Transformers for assistant-token masking. Content format selection is configurable (auto, string, or OpenAI-compatible structured form), and templates can now define custom chat roles, handle \continue\_final\_message\ logic, and use the \raise\_exception\ function for explicit error signaling. Additionally, a custom \tojson\ filter ensures Python-compatible JSON serialization (no HTML escaping, configurable indent/separators/sort\_keys), and a fuel budget prevents resource-exhaustion denial-of-service during template rendering.
rust/src/chat/src/renderer/hf · high confidence
Initial project scaffolding and repository structure
Establishes the foundational structure for the vLLM project, including the initial README, documentation configuration (MkDocs, ReadTheDocs), build system (CMake, setup.py), and development tooling (pre-commit, .gitignore, clang-format). This commit also introduces the AGENTS.md file to define contribution policies and coding standards for AI-assisted development, alongside the Developer Certificate of Origin (DCO) and security policy.
(repo-wide) · high confidence
Introduce C++ KV cache and memory management infrastructure
This change adds the core C++ backend components for managing KV cache operations and GPU memory allocation. It introduces \csrc/cache.h\ to define the C++ interfaces for cache operations such as swapping, reshaping, and gathering KV cache blocks, including support for quantized caches (e.g., FP8, indexer K quantization) and MLA (Multi-Head Latent Attention) fusion. It also adds \cumem\_allocator.cpp\ and \cumem\_allocator\_compat.h\, which implement a CUDA PluggableAllocator based on the \cuMem\ APIs to manage GPU memory, with specific support for ROCm's sleep mode memory chunking and address reservation. Additionally, it includes \cuda\_compat.h\ and \cuda\_utils.h\ to provide cross-platform (CUDA/ROCm) compatibility macros and utilities for device attributes and error checking.
csrc · high confidence
Introduce Hugging Face model backend for text generation
The Rust frontend now includes a new Hugging Face (HF) text backend that loads and parses HF model configurations (config.json, tokenizer\_config.json, generation\_config.json) to handle EOS tokens, sampling parameters, and MoE metadata. It supports resolving model files from local directories, the HF Hub cache, or remote downloads, and allows applying JSON Merge Patch overrides to the model config at runtime.
rust/src/text/src/backend/hf · high confidence
Introduce Machete mixed-precision GEMM kernels for Hopper architectures
This change adds the Machete kernel implementation to the stable Torch ABI, providing a new optimized path for mixed-precision matrix multiplication on NVIDIA Hopper (Sm90) hardware. The Machete kernels are based on CUTLASS and feature a weight prepacking strategy that aligns data with tensor core layouts to enable wider shared memory loads, improving performance for quantized workloads. The implementation includes the core GEMM kernels, prepacking logic, and dispatchers, exposing \machete\_mm\, \machete\_prepack\_B\, and \machete\_supported\_schedules\ as new operations in the \\_C\ library.
_csrc/libtorch\stable/quantization/machete · high confidence
Introduce Rust LLM facade for external engine integration
Adds a new Rust \llm\ crate that provides a thin, high-level generate-and-abort facade over the \EngineCoreClient\. This library enables external applications to connect to a vLLM engine via a handshake protocol, submit tokenized generation requests, and consume streaming outputs. It handles critical frontend responsibilities including external-to-internal request ID mapping for reliable aborts, configurable stream intervals, request metrics tracking, and periodic stats logging. A smoke-test example is included to demonstrate connecting to a headless vLLM instance.
rust/src/llm/src · high confidence
Introduce Rust chat frontend with structured reasoning and multimodal support
The Rust chat frontend is now available, providing a new path for chat requests that supports structured assistant outputs including extracted reasoning blocks and tool calls. It handles multimodal inputs (images, video, audio) with configurable per-prompt limits, and exposes server-side controls for tool-call strictness and reasoning effort. The implementation includes a dedicated error model for request validation, a streaming event interface for collecting final assistant messages, and integration with the engine's token usage reporting.
rust/src/chat/src · high confidence
Introduce Rust text-generation frontend with prompt truncation and sampling validation
The Rust text frontend now handles text-generation requests, including prompt tokenization, truncation (with configurable side and limit), and sampling parameter validation. Users can now control prompt truncation via \truncate\_prompt\_tokens\ and \truncation\_side\, and the frontend enforces bounds on sampling parameters like \min\_tokens\ vs \max\_tokens\, \thinking\_token\_budget\, and logprobs limits based on model and tokenizer vocabulary sizes.
rust/src/text/src · high confidence
Introduce Rust-based Hugging Face chat backend with multi-renderer support
The chat frontend now uses a new Rust implementation (\HfChatBackend\) to load and manage Hugging Face models, replacing previous approaches. This change introduces support for multiple chat renderers (Hugging Face native, DeepSeek V3.2/V4/V4.1, Harmony, Inkling, and Kimi K3) selected via the \--renderer\ flag. It also adds configuration options for Hugging Face model revisions (\--hf-revision\), JSON config overrides (\--hf-overrides\), generation config modes (\--generation-config\), and multimodal limits (\--limit-mm-per-prompt\). The backend automatically loads multimodal processors unless \--language-model-only\ is specified, and supports the \hf\ parser for response templates derived from tokenizer configs.
rust/src/chat/src/backend · high confidence
Introduce Rust-based OpenAI-compatible inference server
This change adds a new Rust implementation of the vLLM frontend server, providing an OpenAI-compatible HTTP API and a gRPC interface for model inference. The server supports streaming chat completions, multimodal inputs (images, video, audio), LoRA adapter management, and structured output tool calling. It includes configuration for TLS/mTLS, CORS, API key authentication, and health reporting, along with a smoke-test example demonstrating integration with an external vLLM engine.
rust/src/server · high confidence
Introduce Rust-based metrics library with lock-free histograms
The \rust/src/metrics\ crate now provides the core metrics infrastructure for the vLLM frontend, replacing the previous Python-based definitions. It introduces a custom lock-free \Histogram\ implementation to prevent serialization of Tokio workers during high-frequency observations, ensuring better performance for latency-sensitive metrics. The library registers a comprehensive set of Prometheus families, including HTTP request durations, weight operation and sleep-mode operation tracking, request lifecycle metrics (such as time-to-first-token and inter-token latency), and scheduler statistics (including KV-cache residency and Mooncake/NIXL connector telemetry).
rust/src/metrics · high confidence
Introduce Rust-based tokenizer frontend with incremental decoding and multi-backend support
A new Rust tokenizer implementation is added to the \rust/src/tokenizer\ module, providing a high-performance alternative for tokenization and detokenization. The frontend supports multiple backends, including HuggingFace \tokenizers\, \fasttokens\, \tiktoken\, and Mistral's Tekken, automatically selecting the fastest available option. It introduces an incremental \DecodeStream\ that emits text chunks token-by-token with precise token attribution and anchoring, improving streaming latency. The implementation also handles edge cases such as out-of-vocabulary token IDs, UTF-8 boundary panics, and raw character preservation in byte-level decoders.
rust/src/tokenizer/src · high confidence
Introduce ScalarType abstraction and batch-invariant environment caching
The core library now includes a new \ScalarType\ class that allows representing sub-byte and custom quantized data types (such as MXFP4) beyond standard PyTorch dtypes, enabling more flexible kernel dispatch for quantization. Additionally, a new \batch\_invariant.hpp\ header caches the \VLLM\_BATCH\_INVARIANT\ environment variable check to improve performance, and utility headers for exception macros and Python extension registration have been added to support these internal abstractions.
csrc/core · high confidence
Introduce \`vllm-rs\` CLI with managed engine and standalone modes
Users can now launch the Rust frontend via the new \vllm-rs\ binary, which supports three modes: \serve\ (manages a headless Python vLLM engine and the Rust OpenAI-compatible frontend in one process), \frontend\ (connects to an existing Python engine via a handshake address), and \render\ (runs an engine-free text renderer). The CLI handles argument partitioning, forwarding Python-specific flags after \--\ to the managed engine while parsing Rust-owned options (such as TLS/mTLS via \--ssl-\*\, model length, and parser selections) on the Rust side. It also includes a benchmark subcommand (\bench serve\) and explicit handling for unsupported or no-op arguments to prevent silent misconfiguration.
rust/src/cmd · high confidence
Introduce extensible IO Processor plugin system for pooling models
Users can now extend vLLM's handling of pooling models by installing custom IO Processor plugins. This change adds a new plugin group (\vllm.io\_processor\_plugins\) and an abstract \IOProcessor\ interface that allows third-party code to define custom pre-processing and post-processing logic for model inputs and outputs. The system automatically discovers and loads these plugins, enabling users to inject custom behavior into the pooling pipeline without modifying core vLLM code.
_vllm/plugins/io\processors · high confidence
Introduce mock engine for frontend stress testing
Added a standalone \vllm-mock-engine\ process that emulates the engine-core protocol to allow frontend stress testing without a real model. The mock engine joins a frontend-owned handshake, treats prefill as instant, and emits random decode tokens until each request reaches its \max\_tokens\. It supports configurable parameters such as \--engine-count\ for data-parallel simulation, \--output-token-chunk-size\ for spec-decode shaped tests, and deterministic token generation via \--seed\. The component includes a CLI entry point, IO loop for ZMQ message handling, and integration tests verifying connection, multi-identity registration, token chunking, and abort behavior.
rust/src/mock-engine · high confidence
Introduce multi-turn conversation benchmarking tool
Adds a new benchmark suite in \benchmarks/multi\_turn\ for evaluating model performance over multi-turn conversations. The tool supports generating synthetic conversations using configurable random distributions (uniform, constant, lognormal, zipf, poisson) via a JSON configuration file, and includes a utility to convert ShareGPT datasets into the required OpenAI format. It provides detailed metrics including throughput, latency, and token counts, with optional warmup-inclusive runtime reporting.
_benchmarks/multi\turn · high confidence
Introduce standalone Rust benchmark client for vLLM endpoints
Adds a new high-performance Rust benchmark client (\vllm-bench\) as a standalone binary with no Python runtime dependency, serving as a drop-in replacement for the Python \vllm bench serve\ command. This tool supports multiple backend types (OpenAI completions/chat, embeddings, pooling, and rerank) and diverse dataset sources (random, ShareGPT, HuggingFace, SPEED-Bench, and Mooncake-style timed traces). It features advanced benchmarking capabilities including concurrency and rate parameter sweeps, multi-run statistical aggregation, multi-turn conversation simulation, and side-by-side result comparison. The implementation prioritizes performance and memory efficiency through features like \Arc\<str\>\ prompt sharing, parallel dataset generation via \rayon\, and a \mimalloc\ allocator, while ensuring output compatibility with the existing Python benchmarking tool.
rust/src/bench · high confidence
Introduce unified parser interface with model-specific implementations
The parser module now provides a single \UnifiedParser\ interface that combines reasoning and tool-call parsing into one stream, replacing the previous separate mechanisms. This location introduces the core \CombinedParser\ adapter and adds dedicated unified parsers for Google Gemma4, HY3, HY4, Inkling, and Kimi K3 models, enabling consistent handling of reasoning text, visible text, and tool calls across these architectures.
rust/src/parser/src/unified · high confidence
Introduce vLLM Attention Benchmarking Suite
A new, unified benchmarking suite is added to benchmarks/attention\_benchmarks, providing a single \benchmark.py\ entry point for both standard attention (Flash/Triton/FlashInfer) and MLA backends. It features a simplified batch specification grammar (e.g., \q2k\, \8q1s1k\) to concisely define prefill, decode, and speculative workloads, alongside YAML configurations for common scenarios like MLA decode, mixed batches, and speculative decoding. The suite also supports CLI-driven parameter sweeps (e.g., for \num\_kv\_splits\ or \reorder\_batch\_threshold\) and includes specific benchmarks for sparse MLA backends and FA4 FP8 output performance.
_benchmarks/attention\benchmarks · high confidence
Introduce vllm-proto Rust crate for gRPC schema and bindings
The \rust/proto\ directory now serves as the canonical source for vLLM's gRPC schema, providing a dedicated Rust crate (\vllm-proto\) that exposes generated Prost message types and Tonic client/server modules for the Inference and Control APIs. Consumers can add \vllm-proto = "0.5"\ to their dependencies to access these bindings without needing Buf credentials or a separate \protoc\ installation. The crate includes the \control.proto\ and \inference.proto\ files, which define services for server/model discovery, request management, LoRA lifecycle, RL weight updates, and KV event discovery, as well as request/response structures for text and multimodal generation. Schema updates are now published to crates.io via a dedicated workflow rather than the Buf Schema Registry.
rust/proto · high confidence
Native Rust chat renderers for DeepSeek, Kimi K3, and Inkling models
The Rust frontend now includes dedicated prompt renderers for DeepSeek V3.2, V4, V4.1, Kimi K3, and Inkling models, replacing generic template rendering with model-specific logic. These renderers handle native features such as DeepSeek's reasoning effort controls, tool-call argument alignment, and media placeholder ordering, as well as Kimi K3's XTML format and Inkling's native tokenization. The \RendererSelection\ enum allows users to explicitly choose or auto-detect these implementations, ensuring accurate prompt construction for these model families.
rust/src/chat/src/renderer · high confidence
New Buildkite CI pipeline for AMD disaggregated vLLM inference
This change introduces a complete CI infrastructure for testing vLLM's disaggregated (prefill/decode) inference on AMD hardware. It adds a configuration script (cluster.sh) for cluster topology and networking, a model catalog (models.yaml) defining performance flags for models like DeepSeek-V3, Kimi-K2.5, and GLM-5.2-FP8, and a launcher script (vllm\_disagg.sh) to orchestrate the disaggregated servers. The pipeline (pipeline-disagg.yaml) and submission script (run-slurm-disagg-test.sh) enable automated testing of these configurations on SLURM clusters, supporting both standard tensor parallelism and wide expert parallelism modes.
.buildkite/amd-disagg · high confidence
New Buildkite CI scripts for artifact annotation, macOS wheel building, and ROCm caching
The \.buildkite/scripts\ directory now includes several new helper scripts to improve the CI pipeline. \annotate-build-artifact.sh\ and \annotate-image-build.sh\ append build metadata and Docker image tags to Buildkite annotations for better visibility. \build-macos-wheel.sh\ enables native arm64 CPU wheel builds on macOS agents. \cache-rocm-base-wheels.sh\ manages S3-based caching of ROCm base wheels to speed up builds. \ci-clean-log.sh\ and \ci-fetch-log.sh\ provide utilities for cleaning and fetching CI logs. \cleanup-nightly-builds.sh\ automates the removal of old nightly Docker tags from DockerHub. \crcr-report.sh\ and \crcr\_report.py\ handle reporting nightly build results to the PyTorch Cross-Repo CI Relay. \check-ray-compatibility.sh\ verifies dependency compatibility with Ray. \cherry-pick-from-milestone.sh\ assists in identifying missing commits for cherry-picking. \ci-bake-rocm.sh\ wraps Docker buildx for ROCm CI builds. \detect-manylinux-tag.py\ ensures correct manylinux platform tags for wheels.
.buildkite/scripts · high confidence
New CMake-based build system for CPU and GPU extensions
The project introduces a new CMake-based build system for CPU and GPU extensions, replacing the previous build mechanism. This change enforces C++20 as the standard and adds support for various CPU architectures including x86\_64, ARM64, PowerPC, and RISC-V, with specific optimizations for features like AMX, AVX2, and BF16. It also integrates ROCm/HIP support through a new hipify tool and handles dependencies like OpenMP and PyTorch's libgomp more robustly across different platforms.
cmake · high confidence
New CPU attention backend with AMX-FP8 and vectorized activation kernels
This change introduces a new CPU attention implementation that adds native AMX-FP8 support for Intel Diamond Rapids processors, enabling high-performance FP8 key-value cache operations. It also includes new vectorized activation kernels (silu, gelu variants) and a lookup-table-based BF16 activation path to accelerate common operations on x86, ARM, RISC-V, and Power architectures.
csrc/cpu · high confidence
New CUDA kernels for Direct DCP attention and SM100 MLA support
This change introduces new CUDA kernels in the attention module to support Direct Data-Parallel Communication (DCP) and NVIDIA Blackwell (SM100) architectures. It adds utilities for direct symmetric-memory DCP operations, including KV and query gathering, LSE reduction, and attention state merging, enabling efficient multi-GPU attention. Additionally, it integrates Cutlass 3.x-style MLA kernels optimized for SM100, featuring TMA-based warp-specialized execution and reduction logic to improve performance on newer hardware.
_csrc/libtorch\stable/attention · high confidence
New CUTLASS epilogue support for device-side quantization scales
Added new \broadcast\_load\_epilogue\ header files that modify the CUTLASS epilogue logic to support row, column, or scalar broadcasting where scale tensors are passed via device pointers. This change allows a single compiled kernel to handle per-tensor, per-channel, or per-token quantization cases and ensures scales remain on the device, preventing performance hazards and torch.compile graph breaks associated with moving scalar values to the CPU.
_csrc/cutlass\extensions/epilogue · high confidence
New CUTLASS-based FP8 MoE kernels for SM90 and SM100
Added new CUDA kernels for w8a8 (FP8) quantized Mixture-of-Experts (MoE) grouped matrix multiplications, targeting NVIDIA Ada (SM90) and Blackwell (SM100) architectures. The implementation introduces \grouped\_mm\_c3x.cuh\ and architecture-specific dispatchers (\grouped\_mm\_c3x\_sm90.cu\, \grouped\_mm\_c3x\_sm100.cu\) that utilize CUTLASS 3.x collective APIs and TMA warp-specialized schedules. These kernels support Float8\_e4m3fn inputs with BFloat16 or Half outputs, featuring optimized tile shapes and cluster configurations for varying problem dimensions (M, N, K) and expert counts. Supporting utilities in \get\_group\_starts.cuh\ and \moe\_data.cu\ handle expert offset calculations, input/output permutation, and problem size computation for both gated and non-gated activation paths.
_csrc/libtorch\stable/quantization/w8a8/cutlass/moe · high confidence
New CUTLASS-based FP8 and INT8 scaled\_mm kernels for Blackwell (SM90/SM100/SM120)
This change introduces a new set of CUDA kernels in the c3x directory for w8a8 quantized matrix multiplication, leveraging the CUTLlass library to support SM90, SM100, and SM120 architectures. It adds specialized implementations for FP8 (blockwise and standard) and INT8 (with zero-point and bias support), including performance optimizations like CTA raster swizzling for SM120 when weights exceed L2 cache, and batch-invariant epilogues for SM100/SM120. The code also includes a dispatcher that routes between these new kernels and existing ones based on the tensor data types and scaling configurations, ensuring compatibility with the stable Torch ABI.
_csrc/libtorch\stable/quantization/w8a8/cutlass/c3x · high confidence
New Claude agent skills for API compatibility, CI debugging, and model development
The .claude directory now exposes several new skills to the agent via symlinks, enabling specialized assistance for specific workflows. These include skills for checking API compatibility, handling Buildkite CI failures, debugging CUDA IMA issues, writing Triton kernels, running kernel microbenchmarks, managing PR checklists, upgrading the Transformers library, and following model recipes. This expands the agent's capabilities in areas such as kernel development, CI troubleshooting, and library migration support.
.claude · high confidence
New Jinja chat templates for Alpaca, ChatGLM, Falcon, and tool-calling models
The examples directory now includes a collection of new Jinja chat templates. Standard conversational formats are added for Alpaca, ChatGLM, ChatGLM2, ChatML, Falcon, Falcon 180B, Inkbot, and Tele-FLM. Additionally, tool-calling templates are provided for Apertus, DeepSeek R1, DeepSeek V3, DeepSeek V3.1, Function Gemma, Gemma 3 (Pythonic), Gemma 4, GLM 4, Granite, Granite 20B Function Calling, and Hermes, enabling structured function calling and tool-use workflows for these models.
examples · high confidence
New KV events subscriber example with replay support
Added a new Python example script (\kv\_events\_subscriber.py\) that demonstrates how to subscribe to and process KV cache events. The subscriber connects to a ZMQ topic to receive event batches and includes logic to detect missed messages, automatically requesting a replay of missing events from a separate replay socket to ensure continuity.
_examples/features/kv\events · high confidence
New LoRA examples for quantization and multi-adapter usage
Added two new example scripts in the features/lora directory: \lora\_with\_quantization\_offline.py\ demonstrates offline inference with LoRA combined with bitsandbytes, AWQ, and GPTQ quantization techniques, while \multilora\_offline.py\ illustrates how to manage and run multiple distinct LoRA adapters within a single engine instance.
examples/features/lora · high confidence
New MoE permute/unpermute CUDA kernels with FP8 and padded-layout support
This change introduces a new set of CUDA kernels for Mixture-of-Experts (MoE) routing operations, located in csrc/libtorch\_stable/moe/permute\_unpermute\_kernels. The implementation provides the core logic for permuting and unpermuting token data across experts, including expert offset computation, inverse index mapping, and top-k ID preprocessing. It adds support for Float8 data types (e5m2 and e4m3fn) alongside existing Float, Half, BFloat16, and Byte types via a new dispatch mechanism. The kernels also handle padded layouts and support skipping non-local expert tokens, enabling more efficient MoE inference and training workloads on CUDA 12.0+.
_csrc/libtorch\_stable/moe/permute\_unpermute\kernels · high confidence
New Model Recipes skill for vLLM deployment guidance
Added a new agent skill that helps users find vLLM deployment recipes and configurations for LLMs and Hugging Face models based on their specific hardware. The skill provides instructions for searching the published catalog at recipes.vllm.ai, selecting appropriate hardware and strategies, and retrieving deployment commands and caveats. It includes a helper script to parse recipe JSON and summarize deployment guidance, enabling users to quickly identify verified configurations for their models without manually constructing commands.
.agents/skills/model-recipes · high confidence
New RL examples for batch invariance, pause/resume, and weight transfer
Added new demonstration scripts in examples/rl to showcase vLLM's reinforcement learning capabilities. The batch\_invariance/reproducibility\_offline.py example shows how to achieve deterministic generation using the VLLM\_BATCH\_INVARIANT flag. The pause\_resume/ directory includes data\_parallel\_pause\_resume.py and pause\_resume\_offline.py, which demonstrate coordinated pause and resume of generation across data-parallel ranks and the async engine, respectively. Additionally, several new weight-transfer examples were added: rdt\_vllm\_serve.py and rdt\_weight\_source.py for sharded RDT-based sync; rlhf\_async\_new\_apis.py for async RL with NCCL weight syncing; rlhf\_http\_ipc.py, rlhf\_http\_nccl.py, rlhf\_ipc\_fsdp\_ep.py, rlhf\_m2n.py, and rlhf\_nccl\_fsdp\_ep.py, which illustrate various weight transfer backends (IPC, NCCL, M2N) and configurations (FSDP, expert parallelism) for syncing training model weights to the inference engine.
examples/rl · high confidence
New ROCm custom kernels for GEMM, MoE, and Paged Attention
This release introduces a suite of new, hardware-optimized CUDA/HIP kernels for AMD ROCm devices, replacing generic fallbacks with specialized implementations. It adds native W4A16 GPTQ kernels for RDNA3 (gfx1100) that use WMMA instructions for larger batch sizes and a deterministic split-K epilogue to fix prior accuracy and nondeterminism issues. A fused MoE W4A16 kernel is added for RDNA3, combining expert routing with dequantization and GEMM in a single launch. The update also brings custom skinny GEMM kernels (wvSplitK/wvSplitKrc) for unquantized linear layers, supporting fp16, bf16, and fp8, with bias and padding support. Finally, a custom paged attention kernel is introduced, featuring FP8 quantization support and optimized V-cache padding masking to prevent NaN propagation.
csrc/rocm · high confidence
New Ray Serving examples for batch inference, multi-node clusters, and elastic serving
Added several new example scripts in the examples/ray\_serving directory to demonstrate advanced Ray integration patterns. These include batch\_llm\_inference.py for data-parallel batch inference using Ray Data, multi-node-serving.sh and run\_cluster.sh for orchestrating multi-node Ray clusters (with Docker support), and a new elastic\_ep/ subdirectory containing serve\_deepseek\_v2.sh, scale.py, and bench.sh to demonstrate elastic expert parallelism and dynamic scaling for DeepSeek models. Additionally, ray\_serve\_deepseek.py provides an example of deploying DeepSeek models via Ray Serve LLM with an OpenAI-compatible API.
_examples/ray\serving · high confidence
New Rust JSON tool-call parsers for Granite 4, Llama, and other models
The JSON tool-parsing module in the Rust frontend now includes dedicated parsers for Granite 4, Llama 3 JSON, Hermes, InternLM2, Mistral, Phi-4 Mini, and Qwen XML formats. These parsers extract tool-call names and arguments from model output using model-specific markers (such as \\<tool\_call\>…\</tool\_call\>\ for Granite 4 and Hermes, \\<\|action\_start\|\>\<\|plugin\|\>…\<\|action\_end\|\>\ for InternLM2, \\[TOOL\_CALLS\] \[…\]\ for Mistral, \functools\[…\]\ for Phi-4 Mini, and \\<tool\_call\>…\</tool\_call\>\ with newline framing for Qwen). They stream argument deltas as raw JSON text without schema conversion or normalization, support parallel tool calls where applicable, and preserve whitespace and special tokens according to each model's format.
rust/src/parser/src/tool/json · high confidence
New Rust text output processing layer with incremental decoding and logprob support
This change introduces a new Rust module (\rust/src/text/src/output\) that handles the conversion of raw LLM engine outputs into incrementally decoded text events. It provides \TextDecodeOptions\ for configuring stop-string handling and special token skipping, and emits \DecodedTextEvent\ streams that separate prompt metadata from generated token deltas. The implementation includes detailed logprob decoding for both prompt and generated tokens, supports sampling masks, and exposes connector-specific transfer parameters (KV and encoder cache) for disaggregated serving scenarios. It also provides a \collect\_output\ helper to aggregate the stream into a final \CollectedTextOutput\ containing the full text, token IDs, and usage statistics.
rust/src/text/src/output · high confidence
New automated server parameter tuning and batched benchmarking scripts
The benchmarks/auto\_tune directory now includes auto\_tune.sh, a script that automatically finds the optimal vLLM server parameters (max-num-seqs and max-num-batched-tokens) to maximize throughput while respecting latency and prefix cache constraints, and batch\_auto\_tune.sh, which allows running multiple tuning experiments sequentially from a JSON configuration file with optional Google Cloud Storage upload. These tools provide a structured way to benchmark and profile vLLM server performance across different hardware configurations (GPU/TPU) and model settings.
_benchmarks/auto\tune · high confidence
New benchmark library utilities for readiness checks and result handling
The benchmark suite now includes a new \vllm/benchmarks/lib\ package providing core utilities for benchmark execution and reporting. This adds a \ready\_checker\ module that allows users to configure a timeout and retry interval when waiting for the serving endpoint to become available, improving reliability during startup. It also introduces \utils\ functions to redact sensitive credentials from CLI argument logs, indicate the compilation mode (e.g., torch.compile) in benchmark results, and write results to JSON with support for PyTorch OSS benchmark format and handling of infinite values.
vllm/benchmarks · high confidence
New benchmarking suite and CLI integration
The benchmarks directory has been reorganized to support the new vLLM CLI. Legacy standalone scripts (benchmark\_latency.py, benchmark\_serving.py, benchmark\_throughput.py) now display deprecation notices and direct users to the corresponding vLLM CLI commands (vllm bench latency, serve, throughput). A comprehensive set of new specialized benchmarks has been added, including tools for measuring batch invariance overhead, block pool performance, hash function efficiency (xxHash vs SHA-256), hidden state extraction throughput, long document QA, N-gram proposer speed, pinned memory impact, prefix block hashing, request prioritization, and top-k/top-p sampling performance. A new README provides an overview of these categories and links to the official CLI documentation.
benchmarks · high confidence
New build and CI tooling for DeepGEMM, NIXL, and Triton
The \tools/\ directory now includes dedicated scripts to streamline building and verifying third-party kernel dependencies. \build\_deepgemm\_C.py\ and \check\_wheel\_deepgemm.py\ handle building and validating the vendored DeepGEMM TORCH\_LIBRARY extension, while \build\_nixl\_with\_torch215\_for\_rubin.py\ provides a temporary build path for NIXL against specific Torch versions. Additionally, \build\_triton\_from\_source.sh\ and \flashinfer-build.sh\ automate the compilation of Triton and FlashInfer wheels, and \kernel\_symbol\_map.py\ generates a mapping of GPU kernel symbols to source files for CI test selection.
tools · high confidence
New classification example scripts for text, vision, and LoRA
Added three new example scripts in the pooling classification directory: \classification\_online.py\ demonstrates online text classification via the \/classify\ API endpoint using both raw strings and token IDs; \vision\_classification\_online.py\ extends this to multimodal inputs, showing how to classify text, images (via URL and base64), and videos through the same API; and \classification\_with\_lora\_offline.py\ provides an offline example of performing classification with a LoRA adapter loaded via the vLLM Python API.
examples/pooling/classify · high confidence
New deployment examples and Helm chart for vLLM
Added new example scripts for offline inference, including \async\_llm\_streaming.py\ which demonstrates token-by-token streaming with the AsyncLLM engine, and \llm\_engine\_example.py\ for direct LLMEngine usage. Introduced a new Helm chart (\chart-helm\) for deploying vLLM on Kubernetes, featuring templates for Deployment, Service, HPA, PVC, and Jobs, along with a schema for configuration validation and unit tests. Also added a SageMaker entrypoint script (\sagemaker-entrypoint.sh\) to handle environment variable-based configuration for AWS SageMaker deployments.
examples/deployment · high confidence
New disaggregated serving and encoder examples
Added new example scripts and documentation for disaggregated serving and encoder (EPD) configurations. This includes a proxy demo for XpYd prefill/decode separation, a push-mode proxy demo using the NixlPushConnector, and a MoRI-IO toy proxy server for single-node prefill/decode disaggregation. Additionally, examples for the Disaggregated Encoder feature are provided, demonstrating 1e1p1d and 1e1pd topologies with support for local media inputs, EC connectors (Example and Mooncake), and dynamic instance registration. A KV cache event publishing example and a FlexKV prefix caching example are also included.
examples/disaggregated · high confidence
New example applications for API server, chatbots, and RAG
Added example scripts in the examples/applications directory to demonstrate vLLM usage. This includes an API server (server.py) and client (client.py) for simple performance benchmarks, Gradio and Streamlit-based chatbot web interfaces (including a Streamlit interface with multi-session management and optional reasoning display), and Retrieval Augmented Generation (RAG) examples using LangChain and LlamaIndex with Milvus vector storage.
examples/applications · high confidence
New example for long text embedding with chunked processing
Added a new example directory (\examples/pooling/embed/openai\_embedding\_long\_text\) demonstrating how to use vLLM's chunked processing feature to handle text inputs that exceed the model's maximum context length. The example includes a server startup script (\service.sh\) that enables chunked processing with configurable pooling types (MEAN, CLS, LAST) and a maximum embedding length up to 3 million tokens, along with a test client (\client.py\) that validates short, medium, long, and extreme-length text embeddings, including batch processing scenarios.
_examples/pooling/embed/openai\_embedding\_long\text · high confidence
New example scripts for batched completions, trace replay, and multimodal/reasoning models
Added several new example scripts in the \examples/generate\ and \examples/reasoning\ directories. \batched\_chat\_completions\_online.py\ demonstrates using the \/v1/chat/completions/batch\ endpoint with structured outputs (regex and JSON schema). \trace\_replay\_offline.py\ shows how to use deterministic decode replay for benchmarking and logprob extraction. New multimodal examples include \mistral-small\_offline.py\ for offline inference with Mistral-Small-3.1, \openai\_chat\_completion\_client\_for\_multimodal.py\ for online multimodal serving, and a \qwen2\_5\_omni\ folder with offline inference scripts. Additionally, \qwen\_1m\_offline.py\ provides a demo for processing 600k-token prompts with Qwen2.5-1M, and reasoning examples (\openai\_chat\_completion\_with\_reasoning.py\, \openai\_chat\_completion\_tool\_calls\_with\_reasoning.py\, etc.) illustrate using reasoning parsers and tool calling with models like DeepSeek-R1 and QwQ.
examples/generate · high confidence
New examples for custom logits processors and offline profiling
Added example scripts in \examples/features/logits\_processor/\ demonstrating how to implement and use custom logits processors with vLLM's offline inference API. The examples cover batch-level processors (\custom.py\), wrapping request-level processors for batch compatibility (\custom\_req.py\), and handling engine configuration during initialization (\custom\_req\_init.py\). Additionally, a DRY (Don't Repeat Yourself) repetition penalty example (\dry.py\) is provided, ported from llama.cpp for the Model Runner V2 interface. New profiling examples (\examples/features/profiling/\) demonstrate how to use vLLM's offline profiling capabilities with PyTorch profiler, including simple offline profiling and batch-specific profiling with configurable parameters.
_examples/features/logits\processor · high confidence
New examples for prompt embedding inference
Added two Python scripts to the examples/features/prompt\_embed directory demonstrating how to use pre-computed prompt embeddings with vLLM. The new prompt\_embed\_inference\_with\_openai\_client.py script shows how to generate embeddings using Hugging Face Transformers and send them to a vLLM server via both the OpenAI-compatible Chat Completions API (where the server applies the chat template) and the Completions API (where the caller must apply the template beforehand). The prompt\_embed\_offline.py script demonstrates single and batch inference using the vLLM Python API directly with the enable\_prompt\_embeds flag.
_examples/features/prompt\embed · high confidence
New examples for reranking and multimodal scoring
Added a suite of new example scripts in examples/pooling/score to demonstrate vLLM's reranking and scoring capabilities. These include online and offline clients for the Cohere and Jina rerank APIs, dedicated examples for ColBERT late-interaction models (including ColQwen3, ColQwen3.5, and ColModernVBERT) with multimodal support, and scripts for the Qwen3-Reranker model. The collection also provides a utility to convert causal language models to sequence classification models and demonstrates how to use chat templates for various rerankers.
examples/pooling/score · high confidence
New examples for sequence and token reward models
Added four new example scripts in the \examples/pooling/reward\ directory demonstrating how to use reward models via the pooling runner. The \sequence\_reward\_offline.py\ and \sequence\_reward\_online.py\ files show how to generate single rewards for entire input sequences using \pooling\_task="classify"\, while \token\_reward\_offline.py\ and \token\_reward\_online.py\ demonstrate token-level reward generation using \pooling\_task="token\_classify"\. These examples cover both offline inference via the Python API and online inference via the HTTP pooling endpoint.
examples/pooling/reward · high confidence
New fused RMSNorm and SiLU+Mul quantization kernels for vLLM
Added new CUDA kernels for fused RMSNorm and SiLU+Mul operations with dynamic per-token and per-block quantization (FP8 and INT8). These kernels, located in the vLLM \libtorch\_stable\ quantization module, implement vectorized RMS computation, dynamic scale calculation with upper-bound clamping, and quantization logic, supporting both residual addition and transposed scale layouts to improve performance and compatibility with the stable ABI.
_csrc/libtorch\_stable/quantization/fused\kernels · high confidence
New hardware CI test scripts for AMD, CPU, HPU, Intel, and NPU platforms
The \.buildkite/scripts/hardware\_ci\ directory now includes dedicated test runner scripts for various hardware backends, replacing or supplementing previous ad-hoc CI configurations. \run-amd-test.sh\ provides a robust framework for ROCm tests with improved quoting, diagnostics, and multi-node detection. New scripts \run-cpu-test.sh\, \run-cpu-test-arm.sh\, \run-cpu-test-ppc64le.sh\, and \run-cpu-test-s390x.sh\ standardize CPU testing across architectures, including disk hygiene, retry logic, and specific kernel/model tests for Arm and PowerPC. \run-hpu-test.sh\ and \run-npu-test.sh\ introduce structured testing for Intel Gaudi and Ascend NPU, respectively, with compatibility pinning and device configuration. \run-intel-test.sh\ and \run-intel-ci-test.sh\ handle Intel XPU testing with parallelism support and specific test suites. \run-cpu-distributed-smoke-test.sh\ adds distributed CPU smoke tests. These scripts collectively enhance CI reliability, diagnostics, and coverage for hardware-specific features.
_.buildkite/scripts/hardware\ci · high confidence
New kernel micro-benchmarking infrastructure
The \benchmarks/kernels\ directory now provides a dedicated suite of micro-benchmarks for measuring the performance of specific vLLM CUDA and Triton kernels. This includes scripts for comparing the new \concat\_mla\_q\ kernel against \torch.cat\ for MLA queries, benchmarking FP8 and BF16 cache gathering and upconversion operations (\cp\_gather\), and evaluating CPU sampling kernels (fused Gumbel-max and greedy argmax) against standard PyTorch baselines. Additionally, it introduces performance comparisons for W8A8 block FP8 GEMMs (Triton vs. CUTLASS), CUTLASS MoE kernels for FP8 and NVFP4 quantization, and the new FlashMLA mega attention kernel for DeepSeek V4.1. These tools allow users to profile and validate the throughput and latency of these specific low-level operations across various batch sizes, sequence lengths, and hardware configurations.
benchmarks/kernels · high confidence
New logging configuration documentation and tensorization example script
The examples/features directory now includes a new documentation file explaining how to configure vLLM logging via CLI arguments (such as \--logging-config\ and \--log-level\) and custom JSON configuration files, including support for built-in JSON formatters and environment variable overrides. Additionally, a new Python script example demonstrates how to serialize and deserialize vLLM models using Tensorizer, supporting local, S3, and HTTP/HTTPS endpoints with optional encryption.
examples/features · high confidence
New observability examples for dashboards, metrics, and tracing
Added new example files in examples/observability to help users set up monitoring for vLLM. This includes native dashboard configurations for Grafana (JSON) and Perses (YAML) to track performance and query statistics, a Python script to dump vLLM metrics locally, and a complete OpenTelemetry proof-of-concept with a dummy client and Jaeger setup for distributed tracing.
examples/observability · high confidence
New offline examples for saving and loading sharded model checkpoints
Added two new example scripts in the sharded state feature directory: \save\_sharded\_state\_offline.py\ and \load\_sharded\_state\_offline.py\. The save script demonstrates how to export a model's state dict into a sharded checkpoint format optimized for tensor-parallel loading, while the load script shows how to restore such a checkpoint and run inference. These examples provide a concrete workflow for users to save and reload models using the \sharded\_state\ load format, particularly useful for large-scale deployments where efficient checkpoint management is required.
_examples/features/sharded\state · high confidence
New offline inference and online serving examples with watermark detection
The examples/basic directory now includes a comprehensive set of scripts for offline inference (basic, chat, generate, classify, embed, score) and online serving (OpenAI chat and completion clients). Additionally, a new watermark detection server example has been added, demonstrating the use of the Gumbel-max algorithm for detecting watermarked text via a FastAPI endpoint.
examples/basic · high confidence
New pooling embedding examples for Matryoshka, multimodal, and binary encoding formats
The examples/pooling/embed directory now includes new scripts demonstrating advanced embedding capabilities. Users can generate variable-dimension embeddings using Matryoshka parameters (embed\_matryoshka\_fy\_offline.py, openai\_embedding\_matryoshka\_fy\_client.py), run multimodal text-and-image embedding inference with models like CLIP, E5-V, Qwen3-VL, and SigLIP (vision\_embedding\_offline.py, vision\_embedding\_online.py), and utilize the new binary encoding formats (bytes, bytes\_only) with specific dtype and endianness settings (embedding\_requests\_bytes\_online.py).
examples/pooling/embed · high confidence
New pooling examples for ColQwen3, Jina Reranker v3, and multi-vector retrieval
Added new example scripts in the pooling/token\_embed directory to demonstrate the Pooling API for several models. This includes online and offline examples for the ColQwen3 multi-vector retrieval model (supporting text and image inputs with MaxSim scoring), the Jina Reranker v3 model (supporting task instructions for scoring and reranking), and the BAAI/bge-m3 model for multi-vector retrieval. These examples show how to use the pooling runner to generate token-level embeddings and compute similarity scores.
_examples/pooling/token\embed · high confidence
New scale-out examples for disaggregated multimodal and token-based serving
Added example scripts in the \examples/scale\_out\ directory demonstrating how to use the new \--enable-scale-out\ flag for disaggregated serving. The \example\_mm\_serve.py\ script illustrates a two-phase multimodal workflow where a chat completion is first rendered into token IDs and features, which are then passed directly to a separate generate endpoint for inference. The \token\_generation\_client.py\ script provides a simpler example of sending pre-tokenized inputs to the generate endpoint. These examples require the server to be launched with the \--enable-scale-out\ argument to expose the necessary endpoints.
_examples/scale\out · high confidence
New shared tracing library for vLLM Rust binaries
A new \vllm-tracing\ crate has been introduced in \rust/src/tracing\ to provide a unified, process-wide tracing subscriber for vLLM Rust components. This library standardizes log formatting with vLLM-style colored output (including process labels, timestamps, and location info) and establishes a consistent log-level filtering strategy that respects both \VLLM\_LOGGING\_LEVEL\ and \RUST\_LOG\ environment variables. Additionally, it includes a \RequestTimingLayer\ that automatically aggregates wall-clock timings for request stages (identified by \stage\ fields) per request ID, allowing consumers to easily extract performance metrics without manual instrumentation.
rust/src/tracing · high confidence
New speculative decoding examples for hidden state extraction and offline benchmarking
Added \extract\_hidden\_states\_offline.py\ to demonstrate using the \extract\_hidden\_states\ speculative method with the \ExampleHiddenStatesConnector\ for saving and loading hidden states, and \mlpspeculator\_offline.py\ to provide a performance comparison between generation with and without speculative decoding. The existing \spec\_decode\_offline.py\ benchmark script was updated to support additional speculative methods (eagle3, mtp, draft\_model) and configuration options like heterogeneous vocabularies and parallel drafting.
_examples/features/speculative\decoding · high confidence
New speech-to-text examples for transcription, translation, and real-time audio
Added four new example scripts in the \examples/speech\_to\_text\ directory to demonstrate vLLM's audio capabilities. The \openai\_transcription\_client.py\ and \openai\_translation\_client.py\ scripts show how to use the OpenAI-compatible API for synchronous and streaming transcription and translation, with the transcription client now supporting an optional \prompt\ parameter for style or vocabulary hints. Additionally, two new scripts in the \realtime\ subdirectory (\openai\_realtime\_client.py\ and \openai\_realtime\_microphone\_client.py\) demonstrate real-time audio transcription via WebSocket, including a Gradio-based demo that streams audio from a microphone for live transcription.
_examples/speech\_to\text · high confidence
New streaming parser library for chat completions
A new streaming parser library has been introduced in \rust/src/parser\ to handle chat completion outputs. It provides token-aware marker parsing that distinguishes structural special tokens from model-generated text, supports recursive argument parsing with depth limits, and constructs output grammars for tool calls and reasoning phases. The library includes utilities for streaming-safe text scanning, JSON schema resolution for tool parameters, and readable grammar outlines for testing.
rust/src/parser/src · high confidence
New structured diffusion example for DiffusionGemma with Jev-style decision API
Adds a new example in \examples/features/structured\_diffusion\ that demonstrates structured generation using the DiffusionGemma model. The \structured\server.py\ script provides a local HTTP server exposing two endpoints: \/v1/chat/completions\ (OpenAI-compatible) and \/v1/systemone\ (Jev-style decision API). This allows users to define question schemas (yes/no, choice, score) with dependencies and conditional logic, which the server translates into specific diffusion canvas arguments (seeding, pinning, constrained reads) to return probability distributions over answers. The included README documents the required vLLM server configuration, the specific \diffusion\\*\ extra arguments, and supported question types.
_examples/features/structured\diffusion · high confidence
New token classification pooling examples for forced alignment and NER
Added four new example scripts in the token classification pooling directory: offline and online implementations for Qwen3 forced alignment (producing word-level timestamps from audio) and offline and online implementations for Named Entity Recognition (NER) using NeuroBERT. These examples demonstrate how to use the pooling runner with the token\_classify task for both batch inference and HTTP API usage.
_examples/pooling/token\classify · high confidence
New tool-calling examples for vLLM
Added a suite of new Python examples in the \examples/tool\_calling\ directory demonstrating how to use vLLM's tool-calling capabilities. These include offline function calling with the native vLLM API, OpenAI-compatible client examples for standard tool calls, streaming tool calls, and required tool choices. The examples also cover advanced scenarios such as using the xLAM tool-call parser, leveraging the OpenAI Responses API with strict tool definitions, and integrating Model Context Protocol (MCP) tools with various filtering options.
_examples/tool\calling · high confidence
New vLLM feature demonstration scripts
Added example scripts demonstrating several vLLM capabilities: automatic prefix caching (with offline comparison), context extension using the YARN method, data parallel inference (including multi-node and multi-instance setups), KV cache reset, and distributed inference using torchrun for both tensor and data parallelism.
(repo-wide) · high confidence
New vLLM performance benchmarking infrastructure
This change introduces a new set of scripts for running and reporting vLLM performance benchmarks. The \run-performance-benchmarks.sh\ script now supports benchmarking on both GPU and CPU (including aarch64) architectures, with adaptive concurrency controls and SLA checks for Time To First Token (TTFT) and Time Per Output Token (TPOT). The \launch-server.sh\ script adds support for launching benchmark servers using multiple backends, including vLLM, TGI, SGLang, LMDeploy, and TensorRT-LLM, with specific handling for FP8 quantization. Additionally, a new Python utility \convert-results-json-to-markdown.py\ is added to parse benchmark results and generate human-readable markdown reports for Buildkite.
.buildkite/performance-benchmarks/scripts · high confidence
Opt-in non-root support for vLLM OpenAI image
The vLLM OpenAI Docker image now includes an opt-in \vllm-openai-nonroot\ target that allows the container to run as an arbitrary non-root user (such as those required by OpenShift Restricted Pod Security Standards). A new entrypoint wrapper (\vllm-nonroot-entrypoint.sh\) ensures compatibility by dynamically setting the \HOME\ directory to a writable location (preferring \/home/vllm\ or falling back to \/tmp\), defaulting the \USER\ and \LOGNAME\ environment variables to \vllm\, and appending a synthetic entry to \/etc/passwd\ when the running UID is not present. This enables relative path workflows and cache directory resolution to function correctly without requiring root privileges.
docker/entrypoints · high confidence
Token attribution for decoded text
The incremental tokenizer now tracks which generated tokens correspond to specific positions in the decoded output text. This change introduces a \TokenAttribution\ structure that records whether a token produced visible characters or was zero-width, along with its byte offset in the decoded string. Users can now access this mapping to understand exactly which tokens contributed to the final text output, enabling more precise debugging and analysis of the tokenization process.
rust/src/tokenizer/src/incremental · high confidence
Unified parser registry with model-specific auto-selection
The chat frontend now uses a centralized parser registry that automatically selects the correct output parser based on the model ID. Supported models include Gemma 4, Hy3/Hy4, Inkling, and Kimi K3, with an opt-in Hugging Face template parser. Users can override this auto-detection via the \ParserSelection\ enum (e.g., \auto\, \none\, or explicit names) and validate overrides before request processing.
rust/src/chat/src/parser · high confidence
Unified reasoning parser factory with multi-model support
The chat frontend now uses a centralized \ReasoningParserFactory\ to manage reasoning output parsing. This change introduces a registry that automatically selects the correct parser based on the model name, supporting a wide range of models including DeepSeek (R1, V3, V4, V4.1), GLM (4.5, 4.7), Kimi (K2), MiniMax (M2, M3), Qwen3, Seed-OSS, Step3/Step3p5, and others. The factory handles pattern matching to distinguish between similar model families (e.g., QwQ vs Qwen3, GLM 4.5 vs 4.7) and ensures specific parsers like Step3p5 are routed correctly before generic substrings. This provides a consistent interface for reasoning parsing across different model providers.
rust/src/chat/src/parser/reasoning · high confidence
Unified reasoning parser with model-specific implementations
The reasoning parser module has been rewritten to use a shared \DelimitedReasoningParser\ state machine that initializes its state from the prompt token IDs rather than hardcoding model-family conventions. This allows multiple model families with the same delimiters to share the same implementation. New model-specific parsers have been added for Cohere Command, DeepSeek R1/V3/V4/V4.1, GLM-4.5/4.6/4.7/5, HY, Kimi (including K2), MiniMax M2/M3, Qwen3/Qwen3.5, SeedOSS, and Step3/Step3p5/Nemotron V3, each configuring the shared parser with their specific start/end tags and framing. The \ReasoningParser\ trait now includes an \initialize\ method to handle this prompt-aware setup, and \ReasoningDelta\ correctly attributes tokens to reasoning or content while dropping delimiter marker tokens.
rust/src/parser/src/reasoning · high confidence
Unified tool parser registry with automatic model-based selection
The Rust chat frontend now uses a centralized \ToolParserFactory\ to manage tool parsing. This registry automatically selects the correct parser based on the model ID (e.g., \deepseek-v4\ maps to \DeepSeekV4ToolParser\, \qwen3-coder\ maps to \Qwen3CoderToolParser\) or allows explicit selection by name. It supports a wide range of models including DeepSeek V3/V4, GLM-4/5, Granite 4, Hermes, Kimi K2, MiniMax M2/M3, Mistral, Phi-4 Mini, Qwen 3, and Seed-OSS, ensuring consistent tool call handling across different model architectures.
rust/src/chat/src/parser/tool · high confidence
Behavioural changes
14191 commits (5052 fixes) modifying vllm
A change to existing behaviour in vllm — 14191 commits (5052 fixs), 3337 files.
vllm · medium confidence · unverified
Added quantization utility headers with ROCm-specific Float8 adjustments
A new header file, csrc/quantization/utils.cuh, has been introduced to provide quantization utilities, including adjusted maximum values and minimum scaling factors for quantization types. This change specifically addresses accuracy issues on ROCm platforms by overriding the default maximum value for Float8\_e4m3fnuz (using 0x7E instead of 0x7F) and adjusting host/device macro definitions to accommodate ROCm's requirements.
csrc/quantization · high confidence
Align Rust sampling and logprobs validation with Python
The Rust frontend now validates sampling parameters (temperature, top\_p, min\_p, frequency/presence penalties, repetition\_penalty) and logprobs settings (logprobs counts, logprob\_token\_ids, prompt\_logprob\_token\_ids) against the same rules and vocabulary ranges as the Python implementation. This ensures consistent error reporting for out-of-range values, empty sequences, and out-of-vocabulary token IDs before requests reach the engine core.
rust/src/text/src/lower · high confidence
FP8 quantization kernels migrated to the stable libtorch ABI
The FP8 weight-8-activation-8 (w8a8) quantization kernels in the \csrc/libtorch\_stable/quantization/w8a8/fp8\ directory have been migrated to use the stable libtorch ABI. This change updates the implementation of per-token-group quantization and common FP8 utilities to rely on \torch::stable::Tensor\ and stable headers, ensuring compatibility with the stable library interface. The update includes new CUDA kernel implementations for scaled FP8 quantization and abs-max reduction, supporting both contiguous and strided memory layouts, and introduces a register-resident path for optimized quantization on Blackwell architectures.
_csrc/libtorch\stable/quantization/w8a8/fp8 · high confidence
FP8 quantization utilities restructured for ROCm and CUDA
The FP8 quantization implementation in the w8a8 module has been reorganized to support both NVIDIA CUDA and AMD ROCm backends with platform-specific optimizations. New header files (\amd/quant\_utils.cuh\ and \nvidia/quant\_utils.cuh\) provide hardware-accelerated conversion functions for FP8 types, including vectorized conversions between FP8 and higher-precision formats (float, half, bfloat16). The ROCm implementation uses hardware conversion instructions where available (ROCm 6.3+) and falls back to software bit-manipulation for older versions. A common header (\fp8/common.cuh\) abstracts device property checks to select the appropriate FP8 type (OCP vs FNUZ) based on the platform. This change ensures correct and efficient FP8 quantization/dequantization across both GPU architectures.
csrc/quantization/w8a8 · high confidence
Introduce pytest-based LM Eval Harness for CI correctness testing
Replaces the previous bash-based evaluation scripts with a structured pytest harness in \.buildkite/lm-eval-harness\. This new setup allows CI to run model evaluations against HuggingFace baselines using YAML configuration files, supporting features like tensor parallelism, ROCm-specific engine shutdowns, and GPU architecture checks. It standardizes how vLLM's accuracy is verified in continuous integration by parameterizing tests via a config list file.
.buildkite/lm-eval-harness · high confidence
Introduce unified output processing and structured event assembly
The chat module now uses a new \ChatOutputProcessor\ trait and \AssistantEvent\ stream to unify how raw token decoding is converted into structured chat events. This change introduces support for parallel tool calls (controlled by a \parallel\_tool\calls\ flag), ensures tool call IDs follow the OpenAI-style \call\\<id\>\ format, and applies parser-owned structural tags (via xgrammar) to override user-supplied structured output constraints. It also adds handling for connector-specific transfer parameters (KV and EC) in the final \Done\ event and improves internal memory efficiency by boxing rarely-set fields in output types.
rust/src/chat/src/output · high confidence
Major Dockerfile overhaul: Ubuntu 24.04, CUDA 13, and uv adoption
The Docker build infrastructure has been significantly restructured. The default base OS is now Ubuntu 24.04 and the default CUDA version is 13.0.3, replacing previous Ubuntu 22.04 and CUDA 12 defaults. The build process has migrated from pip to the \uv\ package manager for faster, more reliable dependency resolution. Additionally, the Dockerfiles now support building for ARM64 (AArch64) CPU platforms alongside x86\_64, and include optimizations such as sccache integration for Rust/CUDA builds and consolidated apt-get commands to reduce image size.
docker · high confidence
Marlin quantization kernels migrated to the torch stable ABI
The Marlin quantization kernels in \csrc/libtorch\_stable/quantization/marlin\ have been refactored to use the \torch::stable\ C++ API (e.g., \torch::stable::Tensor\, \torch::stable::accelerator::DeviceGuard\) instead of the legacy PyTorch internals. This change updates the kernel entry points, tensor handling, and device management to align with the stable ABI, ensuring compatibility with the current PyTorch build system while preserving the existing quantization logic for formats like AWQ and GPTQ.
_csrc/libtorch\_stable/quantization/fp4, csrc/libtorch\stable/quantization/marlin · high confidence
Migrate CI configuration to a new directory structure and add ABI/wheel-size checks
The monolithic test-pipeline.yaml has been deprecated and its content migrated into a new directory structure under .buildkite/test\_areas, .buildkite/image\_build, and .buildkite/hardware\_tests, with corresponding configuration files (ci\_config.yaml, ci\_config\_rocm.yaml, ci\_config\_intel.yaml) now defining job triggers and exclusions. Additionally, new scripts check-torch-abi.py and check-wheel-size.py have been added to enforce PyTorch stable ABI compliance and limit wheel sizes to 500 MB during the release pipeline.
.buildkite · high confidence
Migrate CUDA kernels to the libtorch stable ABI
The CUDA kernels in this directory have been migrated to use the libtorch stable ABI, replacing legacy PyTorch tensor and operator interfaces with the stable equivalents (e.g., torch::stable::Tensor, torch::stable::ops). This change ensures binary compatibility across PyTorch versions for these specific kernel implementations, including activation, cache, and communication primitives, without altering their functional behavior.
_csrc/libtorch\stable · high confidence
Migrate MoE kernels to the torch stable ABI
The MoE (Mixture of Experts) CUDA kernels in \csrc/libtorch\_stable/moe\ have been migrated to use the \torch::stable\ tensor and accelerator APIs, replacing the legacy internal tensor interfaces. This change updates the core routing, sorting, and GEMM implementations (including \topk\_softmax\, \moe\_permute\, and \moe\_wna16\) to align with the stable ABI, ensuring compatibility with the new torch stable layer while preserving existing MoE functionality.
_csrc/libtorch\stable/moe · high confidence
Migrated w8a8 quantized GEMM kernels to the Torch Stable ABI
The quantized matrix multiplication kernels in the w8a8 Cutlass backend have been migrated from the legacy Torch C++ API to the new Torch Stable ABI. This change updates the internal tensor handling and kernel dispatch logic for SM75, SM80, SM89, SM90, SM100, and SM120 architectures to use the stable tensor types, ensuring binary compatibility and future-proofing the quantization operations.
_csrc/libtorch\stable/quantization/w8a8/cutlass · high confidence
Refactored CI image build pipeline with dedicated scripts and non-root support
The CI image build process has been restructured into a modular pipeline defined in \image\_build.yaml\, utilizing dedicated shell scripts for each target (e.g., \image\_build\_cpu.sh\, \image\_build\_arm64.sh\, \image\_build\_torch\_nightly.sh\). This change introduces non-root support for the \vllm-openai\ image, validated by new smoke tests that verify filesystem permissions and user configuration. It also adds specific build targets for PyTorch nightly (using CUDA 13.0 and Ubuntu 24.04), AMD Zen CPU, Intel XPU, and Habana Gaudi (HPU), while enabling zstd compression for Docker image layers to optimize storage and transfer.
_.buildkite/image\build · high confidence
Stabilize ROCm CI image builds with content-addressed base handoffs
The ROCm CI pipeline now uses a new set of build scripts to ensure reproducible and efficient image construction. A content-hash mechanism determines when the ROCm base image needs refreshing, and digest-pinned image references are passed between pipeline stages via Buildkite metadata to prevent redundant pulls and ensure downstream tests validate the exact base used. The pipeline also includes a structural smoke test that verifies the presence of key directories, executables, and Python libraries within the final ROCm CI image.
.buildkite/scripts/rocm · high confidence
Stabilized quantization kernels and alignment-safe vectorization utilities
The quantization implementation in csrc/libtorch\_stable/quantization has been migrated to the stable ABI, introducing new header-only vectorization utilities (vectorization.cuh, vectorization\_utils.cuh) and a new activation kernel source (activation\_kernels.cu). The vectorization utilities now explicitly check output alignment in vectorize\_with\_alignment, preventing misaligned-address crashes when head sizes are not multiples of 8. The new activation kernels provide optimized, vectorized FP8 quantization paths for activation functions like SiLU, supporting both CUDA and ROCm backends with appropriate type mappings and async copy optimizations.
_csrc/libtorch\stable/quantization · high confidence
Unified parser integration for reasoning and tool calls
The default chat output processor now uses a unified parser interface to handle reasoning text and tool calls within a single parsing pipeline. This change introduces support for the \hf\ (Hugging Face) parser for response templates, allows server-side control of structural tag strictness via the \--tool-strict-level\ flag, and ensures that reasoning token counts are accurately reported in chat completion usage statistics. Users benefit from more consistent parser selection logic and better handling of structured outputs like tool calls and reasoning blocks.
rust/src/chat/src/output/default · high confidence
Fixes
Fix memory leak in stable string conversion
Applied a patch to the stable string conversion logic to prevent a memory leak. The change modifies the return path to construct a std::string directly on the stack instead of allocating a new one on the heap, ensuring the original string memory is properly handled without leaving dangling allocations.
cmake/patches · high confidence
Fix missing added tokens in Hugging Face tokenizer
The Hugging Face tokenizer now correctly merges extra added tokens defined in \tokenizer\_config.json\ into the final \tokenizer.json\ structure. Previously, tokens specified in the \added\_tokens\_decoder\ map of the config file were ignored, leading to incomplete tokenizers. This change ensures that all custom tokens are properly included, preserving existing tokenizer definitions while adding the missing entries.
rust/src/tokenizer/src/hf · high confidence
Fixed garbage outputs in QuIP transforms by respecting the inplace parameter
The hadamard transform kernel in the hadacore CUDA implementation now correctly respects the \inplace\ parameter. Previously, ignoring this flag could result in garbage outputs when performing QuIP transforms. This fix ensures that the transformation behaves as expected whether it modifies the input tensor directly or writes to a separate output buffer.
_csrc/libtorch\stable/quantization/hadamard · high confidence
Test coverage
6785 commits adding/updating tests in tests; Add Rust engine-core client smoke tests and examples; Add roundtrip and grammar-replay tests for the Rust chat frontend; Add scheduled integration tests for DeepSeek and Qwen models with EPLB and prefetch offloading; Added benchmark suites for multiple tool parsers; Added dummy platform plugin for testing; Added integration tests for the Rust LLM generate API; Added test plugin for BGE-M3 sparse embeddings processing; Added test plugin for registering dummy models; Added tests for DeepSeek V4.1 tool parsing behavior; Added tokenizer benchmark suites for HuggingFace and Tiktoken backends; Added vLLM chat template fixtures for Rust test alignment.
Dependencies
Introduce structured dependency manifests for benchmarks, examples, and Rust workspace
The project now uses dedicated dependency files to isolate requirements for specific areas: \benchmarks/kernels/requirements.txt\ and \benchmarks/multi\_turn/requirements.txt\ pin libraries like pandas, numpy, and transformers for benchmarking; \examples/features/structured\_outputs/pyproject.toml\ defines the OpenAI and Pydantic dependencies for that example; and the Rust frontend is fully managed via a \Cargo.toml\ workspace with a \Cargo.lock\ file, specifying versions for core crates such as \tokio\, \axum\, \tonic\, and \xgrammar-structural-tag\.
(dependencies) · high confidence
Reorganize and pin dependency requirements
The project has reorganized its dependency management by moving all requirements into a structured \requirements/\ directory, introducing a \common.txt\ file for shared dependencies and platform-specific files (e.g., \cpu.txt\, \cuda.txt\, \rocm.txt\, \tpu.txt\, \xpu.txt\) for hardware-specific needs. This change pins critical libraries to specific versions to ensure stability and security, including upgrading PyTorch to 2.13.0 for CUDA and CPU backends, updating FlashInfer to 0.7.0.post1, and bumping Starlette to \>= 1.0.1 to address [CVE redacted]. It also introduces new dependencies for structured output (xgrammar 0.2.8, llguidance 1.7.0), observability (OpenTelemetry SDK/API \>= 1.27.0), and model loading (fastsafetensors \>= 0.3.3, instanttensor \>= 0.1.9), while removing the default Ray dependency to make it optional.
requirements · high confidence
Restructure test requirements into platform-specific source and lock files
The test dependency management has been reorganized to support distinct hardware backends. New \.in\ source files (cuda, rocm, xpu, cpu) define the specific packages required for each platform's test suite, while the corresponding \.txt\ files contain the fully resolved, pinned dependency trees generated by \uv pip compile\. This change replaces the previous monolithic or less-structured approach, ensuring that CUDA, ROCm, XPU, and CPU tests install the correct, non-conflicting versions of libraries such as PyTorch, transformers, and platform-specific accelerators.
requirements/test · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 62.
Lenses
- Code Health 72
- Architecture 92
- Maturity 81
- Readiness 54
- Security 57
- Performance 91
Changes since last survey
- 300 commits — 178 feature/other, 122 fixes
By area
- tests/v1 — 37 commits
- vllm/model_executor — 34 commits
- vllm/v1 — 24 commits
- vllm/models — 20 commits
- tests/entrypoints — 17 commits
- tests/kernels — 17 commits
- vllm/entrypoints — 15 commits
- rust/src — 14 commits
- tests/models — 12 commits
- .buildkite/test_areas — 10 commits
- vllm/distributed — 6 commits
- (root) — 5 commits
- .buildkite/intel_jobs — 5 commits
- docs/features — 5 commits
- tests/model_executor — 5 commits
- .buildkite/scripts — 4 commits
- .buildkite/hardware_tests — 3 commits
- tests/compile — 3 commits
- tests/distributed — 3 commits
- tests/test_config.py — 3 commits
Notable commits
- fix: [BUG] Fix KV offload shared-memory slot selection using worker rank (#58497)
- fix: [BugFix] Pick the DP world-group port at bind time via the coordination store (#51018)
- fix: [BugFix][DiffusionGemma][CPU] Copy async output snapshots instead of aliasing reused buffers (#59829)
- fix: [Bugfix] Bound GDN speculative decode query widths (#55504)
- fix: [Bugfix] Check batch-invariance support for every attention backend (#60271)
- fix: [Bugfix] Clean up async client sockets after event loop closure (#57757)
- fix: [Bugfix] Close MessageQueue zmq resources via finalizer to avoid exit (#60472)
- fix: [Bugfix] Don't int-sort attention-type names in --kv-cache-dtype-skip-layers for packed KV dtypes (#60364)
- fix: [Bugfix] Don't warn about missing watermark config for default requests (#60412)
- fix: [Bugfix] Eliminate frontend ZMQ port TOCTOU with inherited listeners (#54113)
- fix: [Bugfix] Fix InternLM2 tool parser dropping characters from streamed arguments (#60709)
- fix: [Bugfix] Fix ZeroDivisionError in Qwen3-VL video sampling when the source fps is unknown (#60501)
- fix: [Bugfix] Fix pre-commit check (#60840)
- fix: [Bugfix] Followup to #59299: preserve Unicode in structured decision state prompts (#60651)
- fix: [Bugfix] Index Mamba align state copies by request slot in MRV2 (#60210)
- fix: [Bugfix] Keep the GLM-5.3 kpool tail and Qwen4 QSA ring out of the null block (#59528)
- fix: [Bugfix] Limit GlmMoeDsaForCausalLM fp8 KV cache default to SM100 (#60281)
- fix: [Bugfix] Make suppress_stdout redirect fd 1 instead of sys.stdout.fileno() (#59336)
- fix: [Bugfix] Make disable_any_whitespace actually disable whitespace on xgrammar (#58067)
- fix: [Bugfix] Persist FlashInfer autotune cache per rank to fix the rank-0-only cache-hit deadlock (#57635)
- …and 280 more
Architecture
- 0 containers · 1 bounded contexts · 0 dependency edges (baseline)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
vllm-project/vllm was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 11 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3c062d391d15cfa41d3f321cad24f3030c892a97 — the exact code this score is about.
- Scored under rubric-2026.10.5 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fe8540b5da9b.