verl-project/verl
52.5
Weak · 19 September 2026
102.9k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a distributed reinforcement learning framework for training large language models, supporting algorithms like PPO, GRPO, and DAPO across NVIDIA, AMD, and Huawei Ascend hardware. It provides a unified engine architecture that integrates with inference backends such as vLLM, SGLang, and TensorRT-LLM to manage rollout generation and weight synchronization. The platform offers comprehensive tooling for data preprocessing, multi-turn agent loops, and performance profiling, enabling scalable training on diverse cluster topologies.
How it got here
2024–2025 — Unified engine architecture and backend expansion
107 changes.
The project underwent a major architectural shift by replacing legacy, monolithic training modules with a unified, pluggable model engine interface supporting FSDP, Megatron, VeOmni, and TorchTitan. This period also focused on expanding hardware and inference backend support, notably adding native SGLang integration, Ascend NPU compatibility, and modular reward management systems for advanced RL algorithms like GRPO and GMPO.
2026 — asynchronous training and hardware abstraction
43 changes.
This period focused on introducing asynchronous training architectures, including fully-async policies, one-step-off-policy trainers, and a unified V1 PPO trainer with TransferQueue-backed replay buffers. It also established a hardware-agnostic platform abstraction layer to support NVIDIA, Ascend, and AMD accelerators, alongside a unified checkpoint engine for efficient weight synchronization across diverse backends.
Features
Add FP8 refit support for TensorRT-LLM rollout
Users can now utilize FP8 precision during TensorRT-LLM rollouts via a new quantization helper. This change introduces the \verl/utils/trtllm/trtllm\_fp8\utils.py\ module, which provides the \TRTLLMFP8QuantizerHelper\ class to manage FP8 quantization configurations specifically for the TRT-LLM backend, alongside the necessary \\\init\\_.py\ to expose the package.
verl/utils/trtllm · high confidence
Add Megatron-specific distillation loss implementation
The \verl/trainer/distillation/megatron\ package is introduced, providing the core components for on-policy distillation within a Megatron-LM distributed training context. This includes a new \\_\init\\_.py\ to expose the module and a \losses.py\ file implementing vocab-parallel KL divergence calculations. The loss implementation handles tensor model parallelism by correctly sharding vocabulary logits, aggregating statistics across parallel ranks, and managing global-to-local index mapping for teacher-student probability comparisons.
verl/trainer/distillation/megatron · high confidence
Add MindSpeed engine support for NPU training
This change introduces a new MindSpeed backend engine for NPU devices, providing \MindspeedEngineWithLMHead\ and \MindspeedEngineWithValueHead\ classes that extend the existing Megatron engine. The implementation ensures correct context parallelism by applying necessary MindSpeed patches before device mesh initialization and handles FP8 weight management on NPU by resetting quantized weight caches during device transfers.
verl/workers/engine/mindspeed · high confidence
Add NVFP4 QAT training and vLLM inference support via ModelOpt integration
This change introduces a new ModelOpt integration module in \verl/utils/modelopt\ that enables Quantization-Aware Training (QAT) with NVFP4 (W4A16) quantization for Megatron backends and provides the necessary patches for vLLM inference. Users can now apply fake quantization to Megatron model modules, export QAT-trained weights in the NVFP4 format, and utilize specific vLLM patches to handle dynamic weight updates and weight reloading for the Marlin backend, ensuring compatibility between the training and inference stages for this quantization scheme.
verl/utils/modelopt · high confidence
Add NVFP4 Quantization-Aware Training (QAT) support for FSDP
Users can now perform Quantization-Aware Training with NVFP4 (W4A4 and W4A16) quantization modes in FSDP training. This change introduces a new \verl.utils.qat\ module that provides configuration (\QATConfig\), a custom \QATLinear\ layer with Triton-based fake quantization kernels, and utilities for weight scale computation and fusion. It also includes patches for vLLM to enable dynamic weight reloading for NVFP4 quantized models, allowing seamless integration with vLLM's inference engine.
verl/utils/qat · high confidence
Add NeMo Automodel as an alternative training engine
Users can now select the NeMo Automodel engine as an alternative backend for supervised fine-tuning (SFT) training. This new engine, implemented in \verl/workers/engine/automodel\, delegates model building, parallelization (including FSDP2, Megatron FSDP, and DDP strategies), optimizer sharding, and checkpointing to the NeMo Automodel infrastructure while retaining verl's training loop, data pipeline, and loss functions. It supports features such as parameter/optimizer offloading, FP8 quantization, torch compile, and specific optimizations like chunk entropy calculation.
verl/workers/engine/automodel · high confidence
Add PRIME code reward scoring utility
Added a new \verl/utils/reward\_score/prime\_code\ module that implements code correctness evaluation for the PRIME reward system. This includes utilities to execute generated Python code in a sandboxed environment with timeouts, parse test cases, and return pass/fail results along with detailed metadata, supporting both call-based and standard input code styles.
_verl/utils/reward\_score/prime\code · high confidence
Add SGLang FP8 quantization utilities
Added \verl/utils/sglang/\_\init\\_.py\ and \verl/utils/sglang/sglang\_fp8\_utils.py\ to provide utilities for SGLang block-wise FP8 quantization. The new \SGLangFP8QuantizerHelper\ class and \build\_sglang\_fp8\_quant\_config\ function enable users to configure FP8 quantization for SGLang, including support for ignoring specific layers via configuration or environment variables.
verl/utils/sglang · high confidence
Added GSPO training script for Qwen3-32B on Ascend NPU
A new shell script has been added to the Ascend NPU examples directory to facilitate training the Qwen3-32B model using the GSPO algorithm. This script configures the environment for FSDP2 training with vLLM rollouts, setting specific parameters for data handling, model architecture, actor and reference strategies, and rollout configurations tailored for Ascend hardware.
_examples/ascend\_extras/gspo\trainer · high confidence
Added MoE weight merging and splitting utilities for VeOmni
New scripts (moe\_merge.py and moe\_split.py) have been added to the VeOmni toolset to handle expert weight transformations. The merge script converts HuggingFace checkpoints with individual expert weights into stacked tensors for faster loading and better memory efficiency, supporting both Qwen3-MoE and DeepSeek formats. The split script performs the reverse operation, unstacking merged weights back to individual expert formats.
scripts/veomni · high confidence
Introduce Continuous Token support for multi-turn rollouts
The \verl/utils/tokenizer\ package now includes a Continuous Token system that enables incremental, multi-turn prompt construction for agent rollouts. This adds a \ContinuousTokenBuilder\ framework with model-specific implementations (e.g., for Qwen, DeepSeek, GLM, and MiniMax families) to handle complex chat templates, tool calls, and multimodal inputs. It also introduces utilities to automatically detect and restore turn separators that are often dropped during incremental generation, ensuring that the token sequence remains consistent with full-conversation templates.
verl/utils/tokenizer · high confidence
Introduce FSDP-Turbo backend and refactor FSDP engine location
The FSDP engine implementation has been relocated from verl/trainer/ppo/hybrid\_engine to verl/workers/engine/fsdp and refactored to support a new FSDP-Turbo backend. This change adds FSDPTurboEngineWithLMHead, which integrates the fsdp\_turbo library for parallel state management, CPU offloading, and Ulysses sequence parallelism, while also updating the optimizer step logic to handle gradient clipping and mixed-precision scaling via the FSDP-Turbo API. The module now exposes FSDPEngine, FSDPEngineWithLMHead, and FSDPTurboEngineWithLMHead, consolidating the FSDP-specific engine logic into a dedicated location.
verl/workers/engine/fsdp · high confidence
Introduce MCore backend integration via Megatron-Bridge
The \verl/models/mcore\ module has been added to provide a new backend for running models using NVIDIA's Megatron Core. This change introduces a bridge layer (\bridge.py\, \mbridge.py\) that adapts Megatron-Bridge APIs for model initialization and weight conversion, along with a comprehensive set of forward pass implementations (\model\_forward.py\, \model\_forward\_fused.py\, \model\_forward\_1f1b\_overlap.py\) that support sequence packing (THD/BSHD formats), Multi-Token Prediction (MTP), and fused log-probability/entropy calculations. The module also includes configuration converters (\config\_converter.py\) to map HuggingFace model configs to Megatron Core transformer configs and specific initializers for architectures like Qwen2-MoE, DeepSeek-V3, and Mixtral.
verl/models/mcore · high confidence
Introduce One Step Off Policy Async Trainer for improved RL efficiency
This change introduces a new 'One Step Off Policy' asynchronous training mode within the \verl/experimental/one\_step\_off\_policy\ directory. Unlike the standard synchronous PPO workflow, this new trainer parallelizes the generation and training processes by using samples generated in the previous step for the current training update. This approach reduces GPU idle time during long-tail sample generation, offering significant performance improvements (up to 40% faster in benchmarks) for large-scale reinforcement learning tasks. The implementation includes a new \OneStepOffRayTrainer\, a dedicated entry point (\main\_ppo.py\), and specific configuration for async rollout and NCCL-based parameter synchronization.
_verl/experimental/one\_step\_off\policy · high confidence
Introduce Ray-based single controller for distributed worker management
The \verl/single\_controller/ray\ module is introduced, providing the core infrastructure for managing distributed workers via Ray. This includes \RayResourcePool\ and \ResourcePoolManager\ for defining and allocating compute resources (including GPU/NPU placement groups), and \RayWorkerGroup\ for executing operations across these workers. The implementation supports colocated workers, custom master port ranges, and platform-specific device handling (e.g., Ascend NPU), enabling the framework to orchestrate distributed training and rollout tasks on Ray clusters.
_verl/single\controller/ray · high confidence
Introduce SGLang as a native rollout engine with disaggregated and HTTP-server support
This change adds a new SGLang-based rollout engine to verl, replacing the legacy vLLM rollout module in the \verl/workers/rollout/sglang\_rollout\ package. It introduces an asynchronous HTTP server adapter (\AsyncHttpServerAdapter\) and a dedicated \SGLangPDReplica\ to support prefill-decode disaggregated rollouts, allowing separate scaling of prefill and decode stages. The implementation includes a custom sparse delta weight loader (\delta\_loader.py\) for efficient, in-place weight updates without staging full-model mirrors, and handles FP8 quantization, LoRA adapter management, and multi-turn message tokenization. Users can now configure SGLang as the backend for rollout via \engine\_kwargs\, enabling features like disaggregated deployment, async multi-turn generation, and improved memory efficiency during weight synchronization.
_verl/workers/rollout/sglang\rollout · high confidence
Introduce Sandbox Fusion integration for code execution scoring
Added the \verl/utils/reward\_score/sandbox\_fusion\ module to enable code correctness verification via the Sandbox Fusion service. The new \compute\_score\ function extracts code from completions, parses test cases (including optional \assert\_case\ definitions), and executes them against a remote Sandbox Fusion API endpoint. The implementation supports configurable memory limits, concurrent execution via semaphores, and continuous scoring modes, allowing users to evaluate generated code against test inputs using either a self-hosted service or public cloud FaaS.
_verl/utils/reward\_score/sandbox\fusion · high confidence
Introduce SkipManager to cache and reuse rollout generation steps
A new SkipManager utility has been added to uniformly manage skipping schemes during rollout. This feature allows the system to skip sequence generation for specified steps by loading previously dumped data (cached results) or repeating the latest available results, thereby avoiding redundant computation. The implementation includes configuration classes for standard, async, and V1 TransferQueue-based rollout skips, a base skip class with a registration mechanism, and a manager that integrates these skip instances into the training loop via decorators to intercept and optimize generation calls.
verl/utils/skip · high confidence
Introduce TensorRT-LLM as an asynchronous rollout engine
Adds initial support for TensorRT-LLM as an asynchronous rollout engine in VERL's reinforcement learning pipeline. This feature enables distributed inference with Ray-based orchestration, dynamic weight updates via Inter-Process Communication (IPC), and efficient GPU memory management for GRPO training. The implementation uses a hybrid engine colocate mode where training and inference workers share GPUs, managed through resume/release APIs. It includes an HTTP server wrapper for OpenAI-compatible API communication, KV cache management, and CUDA Graph optimization. The rollout supports multi-node deployments with placement group-based GPU assignment and is currently focused on FSDP and Megatron backend support for Qwen model variants.
_verl/workers/rollout/trtllm\rollout · high confidence
Introduce TorchTitan as an alternative training engine
Users can now select TorchTitan as a training backend for the PPO critic, enabling distributed training with FSDP2, tensor parallelism, and activation checkpointing. This new engine automatically maps HuggingFace model configurations to TorchTitan model flavors and handles data loading integration, providing an alternative to the existing DataParallelPPOCritic implementation.
verl/workers/engine/torchtitan · high confidence
Introduce V1 PPO trainer with unified sync/async modes and TransferQueue-backed replay buffer
The \verl/trainer/ppo/v1\ package introduces a new PPO trainer implementation that unifies synchronous and asynchronous training modes under a single abstraction. It provides three concrete trainer classes—\PPOTrainerSync\, \PPOTrainerColocateAsync\, and \PPOTrainerSeparateAsync\—registered via a central \register\_trainer\ mechanism. The core replay buffer (\ReplayBuffer\ and \ReplayBufferAsync\) replaces the previous storage approach with TransferQueue, enabling asynchronous trajectory collection with status tracking (pending, running, finished, failure) and support for DAPO group filtering, staleness control, and failed-group refilling. The separate async trainer adds hybrid rollout switching, allowing idle training GPUs to serve rollout requests when configured, while the colocate async trainer manages weight updates and generation pausing for colocated setups. Metric aggregation is handled by a new \MetricsAggregator\ that supports weighted averages, sums, and special derivations like sequence length min-max differences, and advantage computation is adapted to handle multi-trajectory sessions from agent loops.
verl/trainer/ppo/v1 · high confidence
Introduce VeOmni engine as an alternative training backend
Adds a new VeOmni-based training engine (\VeOmniEngine\, \VeOmniEngineWithLMHead\, \VeOmniEngineWithValueHead\) to the \verl/workers/engine/veomni\ module, providing an alternative to the existing FSDP engine. This implementation integrates VeOmni's native distributed parallelism (FSDP2, Ulysses sequence parallelism, expert parallelism) and includes utilities for handling vision-language model masks, MoE parameter mapping, and model/optimizer offloading to CPU.
verl/workers/engine/veomni · high confidence
Introduce chunked top-K log-prob computation for long-context distillation
The FSDP distillation module now includes a new \losses.py\ implementation that supports a chunked top-K log-probability calculation. This feature allows users to enable \use\_chunked\_topk\ in their \DistillationLossConfig\ to process logits in smaller chunks (default 4096 tokens) rather than materializing the full \[B, T, V\] log-softmax buffer. This significantly reduces memory usage, preventing out-of-memory errors during on-policy distillation with long contexts (e.g., \>=64K tokens). The module also computes diagnostic metrics for teacher-student top-K overlap, including overlap count and token advantage, to help monitor distillation quality.
verl/trainer/distillation/fsdp · high confidence
Introduce driver-side checkpoint callback hook and refactor PPO trainer utilities
Users can now register custom logic to run after each checkpoint save by setting the \trainer.checkpoint\_callback\_class\ configuration option, which instantiates a \CheckpointCallback\ subclass on the driver process to handle side effects like shard uploading or model version registration. This new hook is part of a broader refactoring of the \verl/trainer/ppo\ module that consolidates previously scattered logic into dedicated files: \metric\_utils.py\ for PPO metrics, \padding\_utils.py\ for multi-trajectory batch padding, \prefix\_grouper\_utils.py\ for shared-prefix optimization, \reward.py\ for reward manager loading, and \rollout\_corr\_helper.py\ for off-policy rollout correction. The refactoring also removes the legacy \BasePPOActor\ and \dp\_actor\ classes, replacing them with the new modular structure to support the v1 PPO trainer architecture.
verl/trainer/ppo · high confidence
Introduce dynamic resource scheduling for fully-async training
The fully-async policy now includes a dynamic resource scheduling module that allows Trainer-node GPUs to participate in rollout generation during idle periods, improving overall GPU utilization. This feature introduces a dual-mode inference resource design with Standalone replicas (dedicated rollout nodes) and Hybrid replicas (sharing Trainer-node GPUs). A pluggable policy system (with default, static, and fixed\_ratio strategies) manages the lifecycle of Hybrid replicas, activating them when Standalone capacity is insufficient and deactivating them to free resources for training. Configuration parameters such as \use\_dynamic\_resource\_scheduling\, \dynamic\_schedule\_policy\, and \dynamic\_schedule\_deactivate\_ratio\ control this behavior, with corresponding monitoring metrics provided to track resource utilization.
_verl/experimental/fully\_async\policy · high confidence
Introduce experimental Agent Loop framework for agentic RL rollouts
This release adds a new experimental agent loop module (\verl/experimental/agent\_loop\) that provides a coroutine-based framework for multi-turn rollout and agentic reinforcement learning. It introduces \AgentLoopBase\, \AgentLoopManager\, and \AgentLoopWorker\ classes to manage parallel agent execution, along with two concrete implementations: \SingleTurnAgentLoop\ for standard single-turn completions and \ToolAgentLoop\ for ReAct-style multi-turn interactions with tool calling. The module includes a pluggable \ToolParser\ registry (supporting Hermes, GPT-OSS, and Qwen3 formats) to extract function calls from model outputs, and handles multimodal inputs (images, video, audio) within the rollout generation process.
_verl/experimental/agent\loop · high confidence
Introduce function-based tool registration and unified tool loading
The \verl/tools\ package now supports a simpler, decorator-based way to register tools via the \@function\_tool\ decorator in \function\_tool.py\, which automatically infers OpenAI-compatible schemas from Python function signatures and docstrings. A new \ToolRegistry\ in \tool\_registry.py\ unifies loading both these function-based tools and traditional native tools (defined via YAML config and inheriting from \BaseTool\), while enforcing unique tool names across both sources. New schema models in \schemas.py\ and a base class in \base\_tool.py\ provide the underlying structure for tool definitions, execution, and responses.
verl/tools · high confidence
Introduce model merger for FSDP and Megatron checkpoints
Added a new \verl.model\_merger\ module that converts distributed checkpoints from FSDP and Megatron backends into standard Hugging Face format. The module provides a CLI entry point (\python -m verl.model\_merger merge\) supporting both backends, with specific handling for FSDP DTensor sharding and Megatron tensor/pipeline parallelism, including options to tie word embeddings and upload directly to Hugging Face. It also includes an output validation step to ensure the merged model is complete and a test operation for verifying merged checkpoints against a reference model.
_verl/model\merger · high confidence
Introduce modular reward manager architecture with multiple implementations
The reward computation logic has been restructured into a modular \verl.workers.reward\_manager\ package, replacing the previous monolithic approach. This change introduces an \AbstractRewardManager\ base class and a registry system (\register\/\get\_reward\_manager\_cls\) to manage different reward strategies. Four specific implementations are now available: \NaiveRewardManager\ for standard single-process scoring (with optional per-sample timeouts), \BatchRewardManager\ for batched scoring, \DAPORewardManager\ for DAPO-specific logic including overlong response penalties, and \PrimeRewardManager\ for parallel, multi-process scoring. This allows users to select the appropriate reward computation strategy via configuration rather than relying on a single hardcoded method.
_verl/workers/reward\manager · high confidence
Introduce new MegatronEngine and delta export machinery
The \verl/workers/engine/megatron\ package now provides a new \MegatronEngine\ implementation (along with \MegatronEngineWithLMHead\ and \MegatronEngineWithValueHead\) that serves as the primary interface for running training and inference with Megatron-Core. This engine integrates with Megatron-Bridge for parameter mapping and includes a new \delta\_export\ module that enables exporting model weights from Megatron's parallelized format to standard Hugging Face coordinates. The export logic uses a specialized probing mechanism to safely convert sharded parameters across Tensor, Expert, and Pipeline parallelism without requiring full collective communication during the conversion step. Additionally, the engine supports dynamic context parallel scheduling and includes utilities for handling fused kernels and random seeding across different hardware backends.
verl/workers/engine/megatron · high confidence
Introduce on-policy distillation loss module
Added a new \verl/trainer/distillation\ package that provides the core loss computation logic for on-policy distillation. This includes a registry for distillation loss functions, support for top-k loss calculations across FSDP, FSDP2, VeOmni, and Megatron strategies, and integration with the existing PPO policy loss pipeline.
verl/trainer/distillation · high confidence
Introduce platform abstraction layer and new utility modules for device management and performance tracking
The \verl/utils\ package has been expanded with a new platform abstraction layer (\verl.plugin.platform\) that unifies hardware detection and configuration across CUDA, Ascend NPU, and other accelerators. This is exposed through updated device utilities (\verl/utils/device.py\) that automatically detect the hardware vendor and set appropriate environment variables (e.g., \ASCEND\_RT\_VISIBLE\_DEVICES\). New modules include \activation\_offload.py\ for CPU offloading of backward-pass tensors to reduce GPU memory usage, \attention\_utils.py\ for unified attention padding/rearranging across backends, \dynamic\_cp\_scheduler.py\ for Megatron-Core dynamic context parallel scheduling, \flops\_counter.py\ for hardware-specific MFU/FLOPS estimation, \fp8\_utils.py\ for async FP8 weight quantization, and \groupwise.py\ for robust group-wise statistical operations. These changes provide the foundational utilities required for multi-hardware support, memory optimization, and performance monitoring in the training loop.
verl/utils · high confidence
Introduce resource-separated PPO trainer with model detach support
This change introduces a new experimental PPO trainer architecture in \verl/experimental/separation\ designed for resource-separated scenarios, such as fully async and one-step-off-policy training. The core addition is the \DetachActorWorker\ engine worker, which enables saving and restoring the actor model state to and from CPU memory, supporting FSDP, FSDP2, VeOmni, and Megatron strategies to optimize resource management during training. A new \SeparateRayPPOTrainer\ orchestrates this workflow using dedicated Ray resource pools for training and rollout roles, while utility functions handle the creation of resource pools and role-to-worker mappings. The previous \BasePPOCritic\ abstract class has been removed from this location as it is no longer part of this new separation-focused implementation.
verl/experimental/separation · high confidence
Introduce structured, modular PPO trainer configuration system
The \verl/trainer/config\ package now provides a comprehensive, dataclass-based configuration system for PPO training, replacing the previous flat or implicit structures. This change introduces modular YAML definitions for actor strategies (FSDP, Megatron, Torchtitan, VeOmni) and algorithmic components (KL control, rollout correction, group filtering), which are automatically flattened into reference YAML files (\\_generated\ppo\\*\_trainer.yaml\) for easy inspection and override. Users can now configure training parameters through a unified, hierarchical structure that supports OmegaConf variable interpolation, allowing for more maintainable and extensible training setups across different backend engines.
verl/trainer/config · high confidence
Introduce unified Model Engine interface with pluggable backends
The training infrastructure now uses a new \BaseEngine\ interface in \verl/workers/engine\ to standardize model training, replacing the previous worker-based approach. This change introduces a pluggable backend system that supports multiple training engines, including FSDP, TorchTitan, VeOmni, Automodel, Mindspeed, and Megatron, which are conditionally imported with warnings if unavailable. The new engine architecture provides a consistent API for training and inference batches, handles parameter offloading, and supports sharded delta weight synchronization for efficient checkpointing and export across different distributed strategies.
verl/workers/engine · high confidence
Introduce vLLM colocated training-inference rollout with process separation
This change introduces a new vLLM rollout implementation located in \verl/workers/rollout/vllm\_rollout\ that supports colocated training and inference with process separation. It adds a \ServerAdapter\ client to communicate with a co-located vLLM HTTP server, enabling weight updates via bucketed ZMQ/IPC or shared memory transfer (\bucketed\_weight\_transfer.py\), and integrates with vLLM's delta-sharded weight transfer API (\delta\_weight\_transfer.py\) for efficient checkpoint patching. The module also includes utilities for handling LoRA, online FP8 quantization, and QAT patches within the vLLM worker subprocess, along with support for prefill-decode disaggregated rollouts (\vllm\_pd\_replica.py\) and robust server lifecycle management (\vllm\_async\_server.py\).
_verl/workers/rollout/vllm\rollout · high confidence
MoE router replay support for VeOmni
Added a new router replay controller (verl/utils/veomni/router\_replay.py) that enables recording and replaying of Mixture-of-Experts routing decisions within the VeOmni engine. This feature allows the system to capture expert selection patterns during training and replay them during inference or subsequent training steps, which can help stabilize training and improve performance by ensuring consistent expert usage across different phases.
verl/utils/veomni · high confidence
New Ascend Docker images for SGLang, vLLM, and Qwen 3.5
Added a comprehensive set of new Dockerfiles in the \docker/ascend\ directory to support building Ascend NPU images for multiple hardware generations (A2, A3, 910b) and software stacks. The new images cover SGLang (versions 8.3.rc1 and 8.5.0), vLLM (versions 8.2.rc1, 8.3.rc1, 8.5.0, and 9.0.0), and Qwen 3.5 (version 8.5.2). These images integrate specific versions of PyTorch, torch\_npu, MindSpeed, Megatron-LM, and the verl framework, providing pre-configured environments for training and inference on Huawei Ascend hardware.
docker/ascend · high confidence
New DPPO and GSPO example scripts for Qwen3 models
Added new example directories for Divergence Proximal Policy Optimization (DPPO) and Group Sequence Policy Optimization (GSPO). The \examples/dppo\_trainer\ directory provides scripts to run DPPO (using TV or KL divergence clips) on Qwen3-30B-A3B with Megatron, while \examples/gspo\_trainer\ provides scripts for GSPO on Qwen3-30B-A3B (Megatron) and Qwen3-8B (FSDP), supporting both NVIDIA GPUs and Ascend NPUs. These examples demonstrate specific policy-loss modes and configuration overrides for large MoE models.
_examples/dppo\_trainer, examples/gspo\trainer · high confidence
New Docker images for Isaac Lab, SGLang, vLLM, TRT-LLM, and AWS platforms
This release introduces a new set of Dockerfiles to support multiple inference and training backends. A new image for Isaac Lab (Dockerfile.isaaclab230) provides a desktop environment with VNC access, pre-installed LIBERO, and specific Python dependencies. Stable images for SGLang (Dockerfile.stable.sglang), vLLM (Dockerfile.stable.vllm), and TensorRT-LLM (Dockerfile.stable.trtllm) are added, featuring CUDA 13 support, cuDNN 9, and integration with Megatron-Bridge, DeepEP, and FlashMLA. A new uv-managed image (Dockerfile.uv.cu130) allows runtime selection of backends from a baked package cache. Additionally, AWS-specific images are provided for EFA support (Dockerfile.extention.awsefa) and SageMaker compatibility (Dockerfile.ngc.vllm0.8.sagemaker).
docker · high confidence
New Docker images for verl v0.4 and v0.5 with updated base stacks and DeepEP support
This change introduces new Dockerfile definitions for two distinct runtime environments: a v0.4 stack (CUDA 12.4, PyTorch 2.6.0, Flash Attention 2.7.4) and a v0.5 stack (CUDA 12.6, PyTorch 2.7.1, Flash Attention 2.8.0). The v0.4 images provide application layers for SGLang and vLLM, including specific configurations for Megatron-Core 0.12 and a preview for 0.13, with TransformerEngine pinned to v2.2.1 (or v2.5 in preview) to ensure compatibility with Megatron's RoPE fusion. The v0.5 images focus on SGLang 0.4.8 and include DeepEP support directly in the base image, while the v0.4 DeepEP variants add DeepEP as an optional application layer. These images allow users to select a base environment based on their CUDA/PyTorch version requirements and choose between SGLang or vLLM inference backends.
docker/verl0.4-cu124-torch2.6-fa2.7.4, docker/verl0.5-cu126-torch2.7.1-fa2.8.0 · high confidence
New Docker images for verl v0.5 with CUDA 12.8 and PyTorch 2.7.1
This change introduces new Docker images (base and app variants) for the verl v0.5 preview release, upgrading the underlying stack to CUDA 12.8, cuDNN 9.8, and PyTorch 2.7.1. The base image includes Flash Attention 2.8.0, FlashInfer 0.2.6, and TransformerEngine 2.5, while the app image adds support for sglang 0.4.8 and Megatron-LM core\_r0.13.0. These images are published under the \verlai/verl\ namespace and replace previous configurations to support newer hardware and framework versions.
docker/verl0.5-preview-cu128-torch2.7.1-fa2.8.0 · high confidence
New Docker images for verl v0.5 with PyTorch 2.7 and updated inference engines
This change introduces a new set of Docker images under the \verl0.5-cu126-torch2.7-fa2.7.4\ directory, upgrading the base framework to PyTorch 2.7.1 and CUDA 12.6. It provides specific application images for SGLang (0.4.9.post6 and 0.4.10.post2) and vLLM (0.10.0), paired with TransformerEngine 2.2.1 and Megatron-LM 0.13.0, as well as a preview image for TransformerEngine 2.7 and Megatron-LM 0.15.0. These images also include updated dependencies such as Transformers 4.55.4 and FlashInfer 0.2.9rc1 to support the new PyTorch version and fix known issues.
docker/verl0.5-cu126-torch2.7-fa2.7.4 · high confidence
New FP8 and Triton-accelerated linear cross-entropy kernels
The \verl/utils/kernel\ module now includes new high-performance kernels: an FP8 blockwise quantization implementation in \fp8\_kernel.py\ for memory-efficient weight storage, and Triton-based fused linear cross-entropy kernels in \kernels.py\ and \linear\_cross\_entropy.py\ that accelerate training by fusing the forward and backward passes. These changes introduce new capabilities for FP8 quantization and optimized cross-entropy computation, improving training throughput and reducing memory usage for large vocabulary models.
verl/utils/kernel · high confidence
New GMPO trainer example with geometric-mean policy loss and uv integration
The examples/gmpo\_trainer directory now includes a new training example for Geometric-Mean Policy Optimization (GMPO), which improves stability over standard GRPO by maximizing the geometric mean of token-level rewards rather than the arithmetic mean. Users can run this example via the provided run\_qwen3\_8b\_fsdp.sh script, which configures the actor to use loss\_mode=geo\_mean and supports vLLM rollouts with FSDP training. The example also integrates uv for dependency management, allowing users to launch the trainer via uv run when the VERL\_USE\_UV environment variable is set. Additionally, a custom reward function for AIME validation has been added to ensure proper scoring normalization.
_examples/gmpo\trainer · high confidence
New GRPO training examples for DeepSeek V3/V4, GPT-OSS, GLM-4.1V, and Hy V3
The \examples/grpo\_trainer\ directory now includes a comprehensive set of new launch scripts and documentation for Group Relative Policy Optimization (GRPO). Users can now train DeepSeek-V3 (671B), DeepSeek-V4-Flash (on NVIDIA and AMD), GPT-OSS-20B, GLM-4.1V-9B, and Hy V3 using various backends including FSDP, Megatron, Megatron Lite, and VeOmni. The new README documents the supported model matrix, configuration knobs (such as rollout count and batch sizes), and hardware-specific requirements like FP8 KV cache or expert parallelism.
_examples/grpo\trainer · high confidence
New GRPO training examples for GLM-5.2 and Qwen3 models on Ascend NPUs
Added new launch scripts in the GRPO trainer examples directory to support Reinforcement Learning with Generalized Reward Proximal Optimization (GRPO) for several large language models on Ascend NPU hardware. The new examples include \run\_glm5\_2\_megatron.sh\ for GLM-5.2, \run\_qwen3\_235b\_256k\_megatron.sh\ for Qwen3-235B with 256k context, \run\_qwen3\_30b\_a3b\_megatron.sh\ for Qwen3-30B-A3B, \run\_qwen3\_32b\_fsdp.sh\ for Qwen3-32B using FSDP, \run\_qwen3\_5\_122b\_a10b\_32k\_megatron.sh\ for Qwen3.5-122B, and \run\_qwen3\_next\_80b\_fsdp.sh\ for Qwen3-Next-80B. These scripts provide pre-configured parameters for data, model parallelism (Megatron or FSDP), rollout engines (vLLM or SGLang), and algorithm settings tailored for these specific model architectures and hardware constraints.
_examples/ascend\_extras/grpo\trainer · high confidence
New Multi-Token-Prediction (MTP) training examples for MiMo-7B-RL
Added a new \examples/mtp\_trainer\ directory containing documentation and shell scripts to demonstrate Multi-Token-Prediction (MTP) training using the MiMo-7B-RL model with the Megatron backend. The entry includes a README explaining MTP usage and key configuration flags, alongside three run scripts: a standard synchronous hybrid-engine setup (\run\_mimo\_7b\_mtp\_megatron.sh\), a variant supporting both SGLang and vLLM rollouts with speculative decoding (\run\_mimo\_7b\_mtp\_rl\_vllm\_sgl\_megatron.sh\), and a fully-async multi-node layout (\run\_mimo\_7b\_mtp\_fully\_async\_megatron\_multinode.sh\). All scripts are configured to use the \uv\ package manager for dependency resolution and include specific instructions to adjust \max\_position\_embeddings\ in the model config.
_examples/mtp\trainer · high confidence
New PRIME math reward scoring module
Added a new math reward scoring implementation in \verl/utils/reward\_score/prime\_math\ that uses SymPy for symbolic and numerical equality checking. This module includes utilities for normalizing mathematical expressions (handling LaTeX, units, and mixed numbers) and a grader that compares model predictions against ground truth answers with configurable tolerance and timeout protection to prevent hangs during verification.
_verl/utils/reward\_score/prime\math · high confidence
New PrefixGrouper example for accelerated GRPO training
Added a new example directory demonstrating how to use PrefixGrouper to optimize GRPO training by grouping samples with shared prompts, reducing redundant computations. The entry includes a README explaining the installation, configuration (setting \use\_prefix\_grouper=True\), and performance benefits, along with a shell script (\run\_qwen3\_8b\_fsdp.sh\) that configures a Qwen3-8B model on GSM8K using FSDP and vLLM, including support for \uv\ integration.
_examples/prefix\grouper · high confidence
New ReMax trainer example with synchronous training support
Added a new example directory for the ReMax policy-gradient method, providing scripts to train Qwen models using vLLM for rollouts and FSDP for training. The example includes a new synchronous training script (\run\_qwen2.5\_math\_7b\_sync\_fsdp.sh\) that leverages the TransferQueue trainer, alongside a standard asynchronous script (\run\_qwen3\_8b\_fsdp.sh\). Both scripts are configured to use the \remax\ advantage estimator and support execution via \uv\ on GPU environments.
_examples/remax\trainer · high confidence
New Tinker-style TrainingWorker with per-step optimizer overrides
The \verl/workers\ module now includes a new \TrainingWorker\ and \TinkerTrainingWorker\ that expose a Tinker-like API for split training primitives (explicit gradient clearing, forward/backward, and optimizer stepping). This allows callers to retain model outputs after backward passes and apply per-step optimizer parameter overrides (such as learning rate or weight decay) directly to all optimizer param groups before stepping. The worker also integrates with the new engine registry and profiler infrastructure, providing a more granular control surface for training loops compared to the previous coarse-grained batch training API.
verl/workers · high confidence
New dataset preprocessing scripts for multimodal and tool-use scenarios
Added preprocessing scripts in examples/data\_preprocess to convert various datasets into the required parquet format for training. This includes support for the AIME 2024 and DAPO-Math-17k datasets with tool-use capabilities, the Geometry3k dataset (both standard and multiturn tool-use variants), GSM8k (standard, multiturn SFT, and multiturn tool-use), and the MATH-lighteval dataset. Additionally, new scripts handle multimodal data from the Open-R1 and TinyLLaVA-Video-R1 datasets, as well as the SearchR1 dataset. All scripts support loading from local paths or HuggingFace, saving to local directories or HDFS, and include specific formatting for prompts, images, videos, and tool arguments.
_examples/data\preprocess · high confidence
New diagnostic and validation tooling for scripts
The scripts directory now includes several new utilities to improve environment setup and model integrity. A new \diagnose.py\ script provides a comprehensive check of the OS, hardware, Python/pip versions, network connectivity, and installed packages (vLLM, SGLang, Ray, Torch) to aid in troubleshooting. A \chat\_template\_checker.py\ script and its associated \chat\_template\_mock\_trajectories.py\ data have been added to validate chat-template append-only behavior and Continuous Token builder correctness against various mock trajectories. Additionally, \init\_random\_model.py\ allows users to create small, debug-friendly models with custom configurations and random weights, while \generate\_trainer\_config.sh\ automates the flattening of Hydra trainer configurations into reference YAML files. Finally, installation scripts for NPU backends (\install\_sglang\_mcore\_npu.sh\, \install\_vllm\_mcore\_npu.sh\) and a unified \install\_vllm\_sglang\_mcore.sh\ have been updated or added to support specific hardware stacks and dependency versions.
scripts · high confidence
New experimental Dockerfiles for SGLang and vLLM backends
Added Dockerfiles for experimental builds using SGLang (v0.5.6) and vLLM (v0.12.0) as inference backends. These images include specific dependencies for Qwen3-Next training, such as causal-conv1d and flash-linear-attention, alongside HybridEP support and NVIDIA profiling tools.
docker/verl0.6.1-experimental · high confidence
New experimental fused linear PPO kernel with memory optimization
Added a new experimental module (\verl/utils/experimental/torch\_functional.py\) providing a fused linear operation for PPO training that computes log-probabilities and entropy in a single pass. This implementation supports chunked processing to reduce memory usage and integrates optional acceleration via the \liger\_kernel\ (LigerFusedLinearScaledCrossEntropyFunction) or \flash\_attn\ (cross\_entropy\_loss) libraries, falling back to standard PyTorch operations if these dependencies are unavailable. The module also includes the corresponding backward pass logic for gradient computation.
verl/utils/experimental · high confidence
New hardware-agnostic platform abstraction layer with multi-chip support
The \verl/plugin\ module now introduces a unified platform abstraction layer that routes all device-specific logic through a \PlatformBase\ singleton, preventing direct calls to vendor-specific libraries like \torch.cuda\ or \torch.npu\. This system supports auto-detection of available hardware (via the \VERL\_PLATFORM\ environment variable or system probes) and includes built-in backends for NVIDIA CUDA, Huawei Ascend NPU, and AMD ROCm/HIP. Users can now run verl on these diverse accelerators without modifying core code, and external plugins can register new hardware backends via the \PlatformRegistry\ decorator.
verl/plugin · high confidence
New learning-oriented tutorials for agent loops, Ray, and cloud deployment
Added a new \examples/tutorial\ directory containing learning-oriented content that serves as starting points for using the platform. This includes a notebook walkthrough for the agent-loop API with code sandbox execution, a Ray API crash-course notebook, a SLURM job template for running Ray-on-SLURM with configurable network interfaces, and SkyPilot task specifications for running PPO and GRPO training on Kubernetes or cloud platforms.
examples/tutorial · high confidence
New loss and padding utilities for nested-tensor training
The \verl/workers/utils\ package now includes dedicated modules for handling loss computation and tensor padding in no-padding (nested-tensor) scenarios. \losses.py\ introduces \sft\_loss\ and \ppo\_loss\ functions that support both standard and no-padding data modes, normalizing PPO and SFT losses over the global mini-batch to ensure gradient invariance across micro-batch splits. \padding.py\ provides utilities to convert between left-right padded and no-padded formats, including \left\_right\_2\_no\_padding\ for creating nested tensors from input IDs and attention masks, \no\_padding\_2\_padding\ for slicing model outputs back to batched tensors, and \build\_attention\_mask\_from\_nested\ for generating attention masks from nested inputs. These changes enable more efficient training by avoiding padding overhead while maintaining correct loss aggregation.
verl/workers/utils · high confidence
New metric aggregation utilities for distributed training
Added a new \verl.utils.metric\ module providing a \Metric\ class and \reduce\_metrics\ function to standardize how training metrics are collected and aggregated. The \Metric\ class supports configurable aggregation strategies (mean, sum, min, max) and includes a dedicated \aggregate\_dp\ method to correctly handle data-parallel synchronization across ranks. The \reduce\_metrics\ helper automatically selects the appropriate reduction operation (mean, max, or min) based on the metric key name, simplifying the consolidation of metrics from distributed workers.
verl/utils/metric · high confidence
New on-policy distillation trainer examples for Qwen3 and Qwen3.5
The examples/on\_policy\_distillation\_trainer directory now provides a comprehensive set of launch scripts for on-policy distillation, covering single-teacher and multi-teacher (MOPD) configurations for Qwen3 and Qwen3.5 models. These scripts support training on NVIDIA GPUs and Huawei Ascend NPUs using FSDP2 or Megatron backends, with vLLM for rollouts. Key capabilities include VeOmni engine support with fused top-K distillation kernels to reduce GPU memory usage, multi-modal (text and vision) distillation, and integration with the uv package manager for reproducible environments.
_examples/on\_policy\_distillation\trainer · high confidence
New profiling examples for GRPO runs with torch, torch\_memory, and NPU profilers
The \examples/profile\ directory now provides ready-to-run scripts for capturing performance traces during GRPO training. Users can profile PyTorch execution and vLLM rollout inference using the torch profiler, analyze CUDA memory allocation with torch memory profiling, or capture NPU traces on Ascend hardware. These examples demonstrate how to configure the profiler via environment variables and Hydra overrides to capture training and inference timelines, with options for discrete or end-to-end tracing modes.
examples/profile · high confidence
New router implementations for the reward loop
The reward loop router module has been expanded with two new router implementations: an inner SGLang router that launches a dedicated SGLang router process with health checks, and a NaiveRouter that provides async load-balancing with configurable retries and optional deterministic routing based on request body hashing. These routers are now part of the \verl.experimental.reward\_loop.router\ package, replacing the previous location under \verl.third\_party.vllm\.
_verl/experimental/reward\loop/router · high confidence
New router\_replay example for deterministic MoE training
Added a new example directory (examples/router\_replay) demonstrating how to use the Router Replay feature for Mixture of Experts (MoE) models to ensure deterministic training by recording and replaying routing decisions. The entry includes a README explaining the R2 (standard) and R3 (rollout-specific) modes, along with shell scripts for running Qwen3-30B-A3B with Megatron and Qwen3.5-35B-A3B with VeOmni. It highlights configuration requirements, such as using vLLM \>= 0.22.0 for certain hybrid-attention models in R3 mode, and shows how to enable replay via YAML or command-line arguments.
_examples/router\replay · high confidence
New shell scripts for DAPO and GRPO one-step-off-policy training
Added a suite of new shell scripts in the one-step-off-policy directory to support the DAPO and GRPO algorithms. These scripts provide ready-to-run configurations for training Qwen2.5-7B on math datasets (MATH, AIME) and Qwen3-0.6B on GSM8K. The configurations cover various distributed training strategies, including FSDP2, Megatron, and disaggregated rollout setups using vLLM and SGLang. Key features include support for dynamic batch sizing, overlong response buffering, and delta-sharded weight synchronization for efficient disaggregated training.
_verl/experimental/one\_step\_off\policy/shell · high confidence
New shell scripts for fully async DAPO training on Qwen3 and Qwen2.5 models
Added a set of new shell scripts in \verl/experimental/fully\_async\_policy/shell\ to launch fully async DAPO (Direct Preference Optimization) training jobs. These scripts cover configurations for Qwen3-30B-A3B (on both CPU/GPU via FSDP and NPU) and Qwen2.5-7B across various node and GPU topologies (e.g., 4-12, 16-16, 32-32, 64-64). They configure the \verl.experimental.fully\_async\_policy.fully\_async\_main\ entry point with specific parameters for async rollout (using vLLM), FSDP2 training, dynamic batch sizing, and staleness thresholds, enabling users to run large-scale math reasoning training on these models.
_verl/experimental/fully\_async\policy/shell · high confidence
New training examples and uv integration for advanced RL algorithms
This update adds new example scripts and documentation for several reinforcement learning algorithms, including CISPO, GDPO, GPG, OTB, REINFORCE++, RLOO, and SAPO, providing ready-to-run configurations for vLLM and FSDP/Megatron backends. It also introduces uv integration for dependency management on GPU environments, allowing users to launch training via \uv run\ while retaining system Python as a fallback.
(repo-wide) · high confidence
New transformer model adapters and Ulysses sequence parallelism support
This change introduces a new \verl/models/transformers\ package that provides model-specific forward implementations and monkey patches for several architectures, including Apertus, GLM-4V, Kimi VL, LLaMA, Qwen2, and Qwen3 (via NPU patches). These adapters integrate Ulysses sequence parallelism by inserting all-to-all communication steps into the attention and forward passes, enabling efficient distributed training. The package also includes generic dense model utilities for PPO training, such as fused linear layers for log-probability and entropy calculation, and patches to support the PrefixGrouper optimization.
verl/models/transformers · high confidence
New unified Checkpoint Engine with multiple backend support
The \verl/checkpoint\_engine\ module introduces a unified abstraction layer for synchronizing weights between training and inference backends, replacing ad-hoc mechanisms with a pluggable architecture. It provides three core streaming APIs (\send\_weights\, \receive\_weights\, \get\_weights\) and includes a registry for multiple communication backends: NCCL (NVIDIA GPU), HCCL (Ascend NPU), NIXL (various transports), Mooncake (P2P/TransferEngine), Kimi (Mooncake+NCCL/HCCL), and a DeltaSharded engine for efficient sparse weight updates. This allows users to select the optimal backend for their hardware and topology (e.g., disaggregated rollout, elastic clusters) while maintaining a consistent interface.
_verl/checkpoint\engine · high confidence
New unified profiler module with multi-backend support and memory debugging
The \verl/utils/profiler\ package has been introduced to centralize performance profiling and memory debugging. It provides a \DistProfiler\ dispatcher that supports multiple backends: Nsight Systems (via NVTX) for GPU, MSTX for Ascend NPU, and PyTorch's native profiler, with automatic selection based on hardware availability. The module includes configuration dataclasses for fine-grained control, such as \TorchProfilerScheduleConfig\ to sub-sample update mini-batches and \PrecisionDebuggerToolConfig\ for msprobe integration. Additionally, it adds \TorchMemoryProfiler\ for automatic CUDA/NPU out-of-memory snapshotting and history tracking, along with utility functions like \GPUMemoryLogger\ and \marked\_timer\ for logging and timing.
verl/utils/profiler · high confidence
New vLLM integration utilities and patches
The \verl/utils/vllm\ package has been introduced to centralize vLLM-specific utilities and patches. This includes \TensorLoRARequest\ and \VLLMHijack\ to enable loading LoRA adapters directly from tensors rather than file paths, and \resolve\_weight\_name\ to correctly handle weight synchronization across different vLLM versions and LoRA configurations. The package also provides \apply\_npu\_vllm\_patches\ to automatically apply necessary compatibility fixes for Ascend NPU hardware, including specific patches for GLM-5.2 and general vLLM-Ascend integration. Additionally, it introduces \restore\_moe\_expert\_maps\ to fix expert parallel routing on ROCm, and dedicated modules (\vllm\_fp8\_utils\, \vllm\_fp4\_utils\) to support weight refitting for FP8 and MXFP4 quantized models during RL rollout, ensuring CUDA graph validity and correct parameter metadata preservation.
verl/utils/vllm · high confidence
Router replay support and distributed checkpointing for Megatron
This update introduces router replay capabilities for MoE models, allowing the system to record and replay expert routing decisions during training and inference, which is essential for algorithms like DPO. It also adds distributed checkpointing support to save and load model states efficiently across parallel ranks. Additionally, the optimizer module has been updated to support the Muon optimizer algorithm and align optimizer states with model precision, while deprecated optimizer arguments have been removed.
verl/utils/megatron · high confidence
Removals
Removal of DataParallel and Megatron PPO Critic implementations
The \DataParallelPPOCritic\ and \MegatronPPOCritic\ classes have been removed from the \verl/trainer/ppo/critic\ module. This change eliminates the specific implementations for handling PPO critic updates via standard data parallelism and Megatron-LM pipeline parallelism, respectively.
verl/trainer/ppo/critic · high confidence
Removal of FSDP and Megatron PPO worker implementations
The \fsdp\_workers.py\ and \megatron\_workers.py\ files in the PPO workers directory have been deleted. This removes the specific worker classes (\ActorRolloutRefWorker\) and their associated model initialization, FSDP/Megatron wrapping, and offloading logic that previously enabled PPO training using these parallel strategies.
verl/trainer/ppo/workers · high confidence
Removal of FSDP and Megatron vLLM hybrid engine implementations
The FSDP-based sharding manager (\fsdp\_vllm.py\) and the Megatron-based hybrid engine (\megatron\_vllm.py\) have been removed from the PPO trainer. This eliminates the specific weight synchronization, offloading, and parallel state management logic previously provided for integrating FSDP and Megatron models with the vLLM inference engine.
_verl/trainer/ppo/hybrid\engine · high confidence
Removal of Megatron-LM v4 compatibility patches
The \megatron\_v4.patch\ file has been deleted, removing custom modifications previously applied to the Megatron-LM codebase. These changes included updates to argument validation logic in \megatron/arguments.py\, the removal of the \group\ argument from pipeline parallel P2P communication calls in \megatron/core/pipeline\_parallel/p2p\_communication.py\, and the introduction of explicit \hidden\_size\ parameters in pipeline parallel scheduling functions (\megatron/core/pipeline\_parallel/schedules.py\).
patches · high confidence
Removal of Megatron-specific LLaMA model implementation
The Megatron-specific implementation for the LLaMA model (modeling\_llama\_megatron.py) has been removed from the codebase. This change eliminates the custom Megatron Core integration for LLaMA, meaning users can no longer run LLaMA models using this specific Megatron backend within this module.
verl/models/llama/megatron · high confidence
Removal of custom Megatron Llama parallel layers
The custom parallel implementation for Llama models in \verl/models/llama/megatron/layers\ has been removed. This change deletes the files \parallel\_attention.py\, \parallel\_decoder.py\, \parallel\_linear.py\, \parallel\_mlp.py\, and \parallel\_rmsnorm.py\, which previously provided tensor-parallel attention, MLP, decoder layer, and RMSNorm components specifically tailored for the Megatron backend. Users relying on these specific internal layer definitions for Llama model construction within this directory will need to adjust their dependencies or switch to alternative implementations.
verl/models/llama/megatron/layers · high confidence
Removal of initial single\_controller base implementation
The initial version of the single\_controller module has been removed, deleting the base worker definitions, decorator-based dispatch logic, Ray integration classes, and Megatron worker group implementations. This cleanup removes the foundational components for distributed worker management and execution that were part of the first open-source release.
_single\controller · high confidence
Removal of vLLM rollout implementation
The vLLM-specific rollout module (\vllm\_rollout.py\) has been removed from the codebase. This change eliminates the legacy implementation that handled model inference via vLLM, including its specific initialization logic for tensor parallelism, weight loading, and sampling parameter management.
verl/trainer/ppo/rollout · high confidence
Removed VeRL Ray API tutorial notebook
The \examples/ray/tutorial.ipynb\ file, which previously provided a step-by-step guide for using the VeRL Ray API (including Ray initialization, GPU accumulation, and Megatron-LM integration), has been deleted. Users can no longer follow this specific tutorial to set up or understand the Ray-based distributed training workflow within the examples directory.
examples/ray · high confidence
Removed vLLM 0.4.2 compatibility layer
The vLLM 0.4.2 adapter module has been removed from the codebase. This change deletes the entire \verl/third\_party/vllm/vllm\_v\_0\_4\_2\ directory, including files such as \arg\_utils.py\, \config.py\, \llm.py\, \llm\_engine\_sp.py\, and various weight loaders (\hf\_weight\_loader.py\, \megatron\_weight\_loaders.py\, \dtensor\_weight\_loaders.py\). Users relying on vLLM version 0.4.2 will no longer have this compatibility layer available.
_verl/third\_party/vllm/vllm\_v\_0\_5\4 · high confidence
Architecture
Relocate vLLM 0.3.1 integration to experimental module
The vLLM 0.3.1 integration code has been moved from the third-party directory to the experimental module. This change reorganizes the codebase structure without altering the functionality of the integration itself.
verl/experimental · high confidence
Rollout logic reorganized into a modular workers package
The rollout subsystem has been moved from the PPO trainer directory to a dedicated \verl.workers.rollout\ package, introducing a cleaner separation of concerns. This change adds a \BaseRollout\ abstract class and a registry-based factory (\get\_rollout\_class\) to standardize backend integration, while extracting server lifecycle management into \LLMServerManager\ and \LLMServerClient\. It also introduces a \GlobalRequestLoadBalancer\ for sticky-session routing and a \RolloutReplica\ abstraction to support hybrid, colocated, and standalone deployment modes. Additionally, the HuggingFace rollout implementation now explicitly passes \position\_ids\ to the model generator and refines sampling logic to correctly handle \num\_return\_sequences \> 1\.
verl/workers/rollout · high confidence
Behavioural changes
Checkpoint logic refactored into a reusable handler and managers
The checkpointing logic previously embedded in the PPO actor module has been extracted into a dedicated \verl/utils/checkpoint\ package. This introduces a \CheckpointHandler\ that manages orchestration (SPMD vs Ray), dataloader state, and LoRA metadata, alongside specialized \BaseCheckpointManager\, \FSDPCheckpointManager\, and \MegatronCheckpointManager\ classes to handle model, optimizer, and scheduler persistence. This change decouples checkpointing from the PPO-specific actor implementation, making the functionality reusable across different training modes and backends.
verl/utils/checkpoint · high confidence
Drop support for vLLM versions prior to 0.18.0 and add SGLang fallback
The vLLM integration in verl now requires vLLM version 0.18.0 or higher, removing support for older versions (0.3.1, 0.4.2, 0.5.4). If the installed vLLM version is unsupported or missing, the system will now check for SGLang availability as a fallback backend; if neither is available, it raises an error. Additionally, the default sleep level is set to 2 for supported vLLM versions, but is forced to 1 when running on NPU hardware (AMD/Ascend) due to compatibility constraints.
_verl/third\party/vllm · high confidence
Enable FULL\_DECODE\_ONLY cudagraph mode in Ascend PPO example
The run script for the Qwen3-8B FSDP PPO example on Ascend NPU now enables the FULL\_DECODE\_ONLY cudagraph mode via the vLLM engine configuration. This change optimizes the rollout phase for decode-only generation, potentially improving inference throughput during the PPO training loop.
_examples/ascend\_extras/ppo\trainer · high confidence
Expanded Megatron model support and unified weight loading
The Megatron backend now supports additional model architectures, including Qwen2, Apertus, Mixtral, Qwen2-MoE, Qwen2.5-VL, DeepSeek-V3, and Qwen3 variants. Weight loading has been unified to use the Mcore GPTModel loader for Llama and Qwen2, while a new weight-saving registry handles checkpoint merging for the expanded set of models, enabling broader compatibility for training and inference workflows using the Megatron engine.
verl/models · high confidence
Expanded math and QA reward scoring with performance optimizations
The reward scoring module now supports a wider range of datasets including Geo3K, Search-R1 QA, and Math-DAPO, alongside a new batched reward computation utility. To improve performance, GSM8K solution extraction has been optimized by clipping the input string to the last 300 characters before regex matching. Additionally, the legacy \math.py\ module has been renamed to \math\_reward.py\ and the internal \\_default\_compute\_score\ API is now deprecated in favor of the new \default\_compute\_score\ function.
_verl/utils/reward\score · high confidence
Introduce experimental teacher loop for distillation logprob computation
The teacher loop logic for on-policy distillation has been reorganized into a new \verl/experimental/teacher\_loop\ module, introducing \AsyncTeacherLLMServerManager\ and \TeacherModelManager\ to handle asynchronous logprob computation and teacher model lifecycle management. This change enforces a fixed temperature of 1.0 for teacher prompt logprobs (ignoring the student's rollout temperature) and disables detokenization to prevent crashes and improve efficiency. It also adds support for multi-teacher routing via configuration keys and validates that teacher GPU resource pools align correctly with node boundaries.
_verl/experimental/teacher\loop · high confidence
Introduce new asynchronous reward loop with legacy configuration migration
The \verl/experimental/reward\_loop\ module has been replaced with a new implementation that uses an asynchronous \RewardLoopManager\ and \RewardLoopWorker\ for computing rewards, supporting rule-based, discriminative, and generative reward models via a new \RewardModelManager\. To ensure backward compatibility, a \migrate\_legacy\_reward\_impl\ function is now provided to automatically transform old \reward\_model\ configuration keys into the new \reward\ structure, allowing existing setups to continue working without manual config updates.
_verl/experimental/reward\loop · high confidence
Introduce new reward manager architecture with DAPO, GDPO, and rate-limiting support
The reward loop now uses a new, extensible reward manager system located in \verl/experimental/reward\_loop/reward\_manager\. This change introduces a registry-based architecture (\base.py\, \registry.py\) and several new manager implementations: \DAPORewardManager\ (with configurable overlong-response penalties), \GDPORewardManager\ (for Group reward-Decoupled Normalization Policy Optimization), \NaiveRewardManager\, \RemoteRewardManager\ (offloading score computation to Ray workers to avoid thread-pool issues), and \RateLimitedRewardManager\ (with concurrency, RPM, and TPM limits for API-based rewards). Existing code should migrate to this new module structure.
_verl/experimental/reward\_loop/reward\manager · high confidence
Introduce structured worker configuration dataclasses
The configuration system for training workers has been reorganized into a modular package under \verl/workers/config\. This change introduces dedicated dataclasses for Actor, Critic, Engine, Model, Optimizer, Checkpoint, Rollout, Distillation, and Disaggregation settings, replacing the previous flat or monolithic configuration approach. Users can now configure specific components like router replay, policy loss modes, engine parallelism (FSDP, Megatron, VeOmni, Torchtitan), and disaggregated prefill-decode workflows through clearly defined, validated config objects.
verl/workers/config · high confidence
Introduce unified model engine-based trainers and standalone generation server
The \verl/trainer\ module is restructured around a new 'model engine' abstraction. A standalone generation server (\main\_generation\_server.py\) is added for launching rollout servers and generating responses via HTTP. The previous FSDP-specific SFT trainer (\fsdp\_sft\_trainer.py\) and legacy generation script (\main\_generation.py\) are removed and replaced by \sft\_trainer.py\ and \sft\_trainer\_ray.py\, which utilize the new \TrainingWorker\ engine to support multiple backends (FSDP, Megatron, Veomni, Torchtitan) in both SPMD and single-controller Ray modes. The PPO entry point (\main\_ppo.py\) is updated to use a modular \BaseTaskRunner\ that registers workers via the unified engine, and a new \constants\_ppo.py\ centralizes Ray runtime environment configuration with conditional logic for specific hardware (e.g., GB200 NCCL workarounds, ROCm visibility fixes).
verl/trainer · high confidence
Introduces BaseConfig, removes legacy reward model, and adds auto-padding and NPU compatibility to DataProto
Users can now use the new BaseConfig class, which provides a frozen, dictionary-like interface for dataclass configurations. The legacy Megatron-based PPO reward model and its base class have been removed from the codebase. DataProto gains automatic padding capabilities (controlled by the VERL\_AUTO\_PADDING environment variable) to ensure batch sizes are divisible by a specified size, and its union operations now correctly handle N-D arrays, complex objects, and NaN values. Additionally, the library now includes a workaround for Ascend NPU precision issues during tensordict transfers and disables a conflicting Ray runtime-env hook when running under uv.
verl · high confidence
Migrate to uv for dependency management and pre-commit for code quality
verl has replaced the legacy setup.py and yapf/pylint tooling with a modern, unified development workflow. The project now uses \uv\ for environment and dependency management, featuring a single \uv.lock\ file that resolves mutually exclusive backend extras (vLLM, SGLang, FSDP, Megatron) into one lockfile, simplifying installation and CI. Code formatting and linting have migrated from yapf and pylint to \ruff\ and \mypy\ via pre-commit hooks, ensuring consistent style and type checking. Additionally, the \recipe\ directory has been moved to a dedicated external repository (\verl-recipe\) and integrated as a git submodule, decoupling core library development from recipe maintenance.
(repo-wide) · high confidence
Move single\_controller to verl and introduce extensible dispatch/execute decorators
The single\_controller module has been relocated to the verl package, updating import paths for Worker, WorkerGroup, and related classes. This change introduces a new decorator system in verl/single\_controller/base/decorator.py that defines extensible Dispatch and Execute modes (such as RANK\_ZERO, ONE\_TO\_ALL, DP\_COMPUTE) using a DynamicEnum registry, enabling more flexible data distribution strategies. The WorkerGroup class now supports explicit dispatch and collect info registration, and the ResourcePool class has been updated with corrected spelling (max\_colocate\_count) and improved documentation. These changes provide a more structured foundation for distributed worker management within the verl framework.
_verl/single\controller/base · high confidence
New debug metrics and profiler migration in verl.utils.debug
The \verl/utils/debug\ module now includes a new \metrics.py\ file that calculates debug metrics comparing rollout and actor log-probabilities, including max, mean, and standard deviation of differences, as well as a Pearson correlation coefficient. Additionally, the module has been refactored to migrate functionality to \verl/utils/profiler\; \performance.py\ and \\_\init\\_.py\ now re-export symbols from the profiler package for backward compatibility, while \trajectory\_tracker.py\ received minor formatting and import cleanup.
verl/utils/debug · high confidence
New streamlined ROCm Docker image with DeepSeek-V4-Flash GRPO support
The \docker/rocm\ directory now provides a new primary \Dockerfile.rocm\ (based on \rocm/primus:v26.4\) that replaces older historical variants like \Dockerfile.rocm6\ and \Dockerfile.rocm702\. This updated image targets ROCm 7.14, Python 3.12, and GPU architectures gfx942/gfx950, and includes pre-built and source-compiled dependencies for vLLM, SGLang, TransformerEngine, and aiter. It also adds support for DeepSeek-V4-Flash GRPO by installing \mbridge\, \megatron-core\, and \mathruler\, and includes a patch to fix a non-leaf tensor issue in Megatron-LM's distributed optimizer.
docker/rocm · high confidence
PPO trainer examples updated to canonical scripts with uv integration and deprecated legacy runs
The PPO trainer examples have been restructured to use a new set of canonical, user-adjustable shell scripts (e.g., \run\_qwen3\_8b\_fsdp.sh\, \run\_qwen3\_8b\_megatron.sh\) that replace the previous legacy scripts (such as \run\_deepseek7b\_llm.sh\ and \run\_qwen2-7b.sh\). These new scripts standardize the configuration interface, support both FSDP and Megatron training backends, and integrate \uv\ for dependency management. The update also includes a comprehensive README documenting the new configuration options, including advanced features like KL divergence control and dual-clip PPO.
_examples/ppo\trainer · high confidence
Refactor dataset utilities and introduce multi-turn SFT support
The \verl/utils/dataset\ module has been restructured to support more flexible and multimodal training workflows. The legacy \SFTDataset\ has been removed and replaced with a new \MultiTurnSFTDataset\ that natively handles conversation data, tool configurations, and multimodal inputs (images, videos) via a \ProcessorMixin\. A new \SFTTensorCollator\ has been added to manage batching of variable-length sequences using NestedTensors and non-tensor data. Additionally, the \RLHFDataset\ and \RMDataset\ have been updated to support sample limiting, shuffling, and improved multimodal processing, while the \\_\init\\_.py\ exports have been adjusted to reflect these structural changes.
verl/utils/dataset · high confidence
Refactored logging utilities with new rank-aware functions and simplified LocalLogger
The logging module in \verl/utils/logger\ has been significantly refactored to improve clarity and distributed computing support. The \LocalLogger\ class has been simplified to remove deprecated remote logger and Weights & Biases integration, now focusing solely on console output with automatic flushing. New utility functions have been added: \print\_rank\_0\ for printing only from the primary process in distributed setups, \print\_with\_rank\ and \print\_with\_rank\_and\_timer\ for formatted output including rank and optional timestamps, and \log\_with\_rank\ for structured logging via Python's \logging\ module. Additionally, the \concat\_dict\_to\str\ helper now uses \pprint\ for better number formatting, and the module's public API is explicitly defined in \\\init\\_.py\.
verl/utils/logger · high confidence
SFT examples migrated to verl.trainer.sft\_trainer with modern engine support
The GSM8K, multiturn, and VLM SFT examples have been updated to use the unified verl.trainer.sft\_trainer entry point, replacing the legacy fsdp\_sft\_trainer. This change introduces support for multiple training backends including FSDP, Megatron, Megatron Lite, and the new Nemo-Automodel engine, while also adding examples for specific models like DeepSeek-V4, Qwen3, and Qwen3-VL. The old shell scripts referencing the deprecated trainer and HDFS paths have been removed.
examples/sft · high confidence
Updated DeepSeek generation example with vLLM 0.8.2 rollout configuration
The DeepSeek LLM generation example has been updated to use the vLLM rollout engine with explicit configuration for multi-node support and correct load formatting. The new script (run\_deepseek\_llm\_7b.sh) sets the rollout load format to 'auto' to ensure compatibility in generation-only runs without a trainer, whereas the previous example (run\_deepseek\_v2\_lite\_math.sh) used a generic rollout configuration. The updated example also introduces user-adjustable parameters for node count, tensor parallel size, and GPU memory utilization, and defaults to the DeepSeek-LLM-7B-Chat model.
examples/generation · high confidence
single\_controller package relocated to verl namespace with updated version path
The single\_controller module has been moved into the verl package (verl/single\_controller), which is a breaking change for existing imports that referenced the previous top-level single\controller path. The package's \\init\\_.py now explicitly re-exports symbols from the base module and updates the logic for reading the version file to account for the new directory structure, ensuring the version string is still correctly resolved from the relocated location.
_verl/single\controller · high confidence
Fixes
Fix FSDP OOM by vendoring PyTorch 2.7 state dict logic
The FSDP integration now uses a vendored copy of PyTorch 2.7's distributed checkpoint and state dict utilities to replace the PyTorch 2.6 implementation, which caused out-of-memory errors. This change adds local copies of the state dict handling code (including \state\_dict.py\ and \\_state\_dict\_utils.py\) to ensure stable memory usage during model state operations.
_verl/third\party/torch · high confidence
Fixes NaiveRollout class name and attention mask logic
The NaiveRollout class has been renamed from NativeRollout to match its file name, and the module now explicitly exports this class. Additionally, a bug in the response attention mask generation has been fixed: the code now iterates over individual token IDs to correctly handle cases where eos\_token\_id is a list, ensuring that attention masks are properly updated for each token during sequence generation.
verl/workers/rollout/naive · high confidence
Shared skills now use directory symlinks for .claude and .codex
The .claude and .codex directories now use symlinks to point to the shared skills located in .agent/skills. This change ensures that both .claude and .codex reference the same underlying skill definitions, maintaining consistency across different agent configurations.
.claude, .codex · high confidence
Test coverage
Added Ascend NPU quick-start test scripts for Qwen3-0.6B; Added Ascend NPU test scripts for Qwen3 and Qwen3.5 models; Added CPU tests for reward manager configurations and registry; Added CPU-based tests for dataset utilities; Added CPU-based unit tests for PPO core algorithms and metrics; Added CPU-only unit tests for V1 PPO trainer components; Added SFT engine accuracy alignment tests; Added comprehensive test suite for the checkpoint engine; Added detached worker test suite for Ray-based training; Added end-to-end test for PPO trainer with custom function rewards; Added integration tests for sandbox fusion correctness checks; Added multi-GPU unit tests for distributed training stability and correctness; Added performance benchmarking script for vLLM async rollout backends; Added regression tests for worker engine behavior and correctness; Added test suite for single controller worker groups; Added tests for FSDP model merger and output validation; Added tests for Megatron router replay utilities; Added tests for Muon optimizer support and precision-aware optimizer behavior; Added tests for PRIME reward scoring and sandbox fusion; Added tests for SGLang rollout HTTP server and LoRA handling; Added tests for debug metrics calculation; Added tests for experimental VLA simulation environments; Added tests for experimental reward loop and async policy configurations; Added tests for function-based tool registration and mixed tool loading; Added tests for single controller dispatch mode registration and updates; Added tests for the experimental agent loop; Added tests for the platform abstraction layer and plugin system; Added tests for trainer constants and multi-trajectory advantage computation; Added tests for vLLM rollout stability and determinism on Ascend NPU and CPU; Added unit and integration tests for TRT-LLM rollout workers; Added unit tests for VeOmni MoE parameter export and router replay utilities; Added unit tests for checkpoint manager logic and configurations; Added unit tests for worker configuration dataclasses; Expanded model engine and training regression tests; Expanded test coverage for protocol, configuration, and distributed worker health; Expanded test coverage for rollout workers; Expanded test coverage for utility modules; New CI sanity checks for documentation, code style, and configuration; New end-to-end test suite for multi-GPU and specialized training configurations.
Dependencies
Migrate to pyproject.toml with uv dependency management and updated core libraries
The project has replaced the legacy requirements.txt-based setup with a structured pyproject.toml configuration, introducing a unified dependency resolution system via uv. This change introduces new optional dependency groups (extras) for specific backends—vllm, sglang, fsdp, and megatron—allowing users to install only the necessary components (e.g., \uv sync --extra vllm\). Core dependencies have been updated, notably upgrading numpy to \>=2.0.0, tensordict to the 0.8–0.10 range, and pyarrow to \>=19.0.0, while adding support for newer versions of transformers (\>=5.5.3, excluding 5.6.0) and vllm (0.24.0). The migration also includes new requirements files for NPU (requirements-npu.txt) and testing (requirements-test.txt), and updates documentation dependencies to include myst\_parser and tokenizers.
(dependencies) · high confidence
New Docker images for v0.6 with SGLang 0.5.2 and updated base dependencies
This change introduces new Docker image definitions for the v0.6 release (CUDA 12.8, PyTorch 2.8.0). The \Dockerfile.app.sglang\ adds support for SGLang version 0.5.2 and includes the \torch-memory-saver\ package. The \Dockerfile.base\ establishes a new foundation using NVIDIA PyTorch 25.03, upgrading core dependencies such as \transformers\ to 4.55.4 and \flash\_attn\ to 2.7.4.post1, while also integrating DeepEP and Megatron-LM. Additionally, \Dockerfile.vllm011.mcore\_gpt-oss\ provides a specialized image based on NVIDIA NeMo for GPT-OSS, installing vLLM 0.11.0 and related utilities.
docker/verl0.6-cu128-torch2.8.0-fa2.7.4 · high confidence
Housekeeping
Updated package version to 0.10.0.dev
The version identifier in the package has been updated from 0.0.2 to 0.10.0.dev, reflecting the current development state of the software.
verl/version · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 53.
Lenses
- Code Health 77
- Architecture 98
- Maturity 73
- Readiness 44
- Security 43
Changes since last survey
- 300 commits — 127 feature/other, 173 fixes
By area
- verl/workers — 47 commits
- .github/workflows — 38 commits
- verl/trainer — 36 commits
- verl/utils — 24 commits
- tests/utils — 23 commits
- docs/ascend_tutorial — 20 commits
- examples/grpo_trainer — 14 commits
- docker/ascend — 12 commits
- verl/models — 12 commits
- tests/workers — 10 commits
- verl/checkpoint_engine — 8 commits
- (root) — 7 commits
- verl/experimental — 7 commits
- docs/advance — 6 commits
- tests/special_npu — 6 commits
- examples/ascend_extras — 3 commits
- tests/experimental — 3 commits
- tests/models — 2 commits
- tests/special_distributed — 2 commits
- tests/trainer — 2 commits
Notable commits
- fix: Revert "[megatron] fix: Qwen3.5 LoRA & MTP support (with Megatron-Bridge) (#5599)" (#7173)
- fix: Revert "[megatron] fix: preserve R2 router replay for THD-packed batches" (#7786)
- fix: [BREAKING][megatron, cfg] fix: remove unused grad_offload (#7544)
- fix: [BREAKING][trainer] fix: separate_async should use the same step granularity with other trainers (#6977)
- fix: [algo] fix: carry running_return through observation spans in REINFORCE++ (#7278) (#7300)
- fix: [algo] fix: micro-batch normalization for distillation loss (#7225)
- fix: [algo] fix: normalize critic value loss over the global mini-batch, not per micro-batch (#6957)
- fix: [cfg, megatron, doc] fix: drop unused actor.router_replay in favor of engine config (#7466)
- fix: [cfg] fix: drop unused ref router replay config (#7536)
- fix: [ci, hardware] fix: stabilize ROCm PPO trainer CI (#7377)
- fix: [ci, trainer] fix: Fix ci AssertionError for the environment variable CUDA_DEVICE_MAX_CONNECTIONS (#7040)
- fix: [ci] chore: Fix docker image upload (#7311)
- fix: [ci] chore: Fix npu nightly ci (#7345)
- fix: [ci] chore: Fix transformers version in Ascend docker images (#7867)
- fix: [ci] chore: fix ci failure (#7629)
- fix: [ci] chore: fix ci failure with transformers==5.9.0 (#7654)
- fix: [ci] chore: fix nightly ci of npu (#7081)
- fix: [ci] chore: fix npu and gpu ci (#7066)
- fix: [ci] chore: fix npu ci (#7354)
- fix: [ci] chore: remove vllm_ascend patch and fix ci (#7042)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
verl-project/verl was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 3efe38c759c14622fd1b2c9e3679f2d02f86bdac — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.