ray-project/ray
47.8
Weak · 26 September 2026
676.6k
lines of production code
Python
with C++, C
4
measurements over time
What this system is
This system is a distributed computing framework that enables scalable task execution and model serving across heterogeneous clusters. It provides a unified API for Python, C++, and Java to define remote functions, actors, and data pipelines, with specialized support for machine learning workloads including distributed training, hyperparameter tuning, and LLM inference. The codebase also includes a comprehensive CI/CD infrastructure for building, testing, and releasing artifacts, alongside extensive documentation and interactive examples for users.
How it got here
2016–2020 — Build system modernization and C++ API expansion
37 changes.
This period focused on migrating the project's build infrastructure to Bazel and standardizing Docker images to improve reproducibility and cross-platform support. Simultaneously, it introduced a comprehensive new C++ API with local and cluster execution modes, enabling native C++ workers and cross-language integration alongside significant overhauls to the documentation site and CI pipelines.
2021–2022 — CI modernization and documentation expansion
34 changes.
This period focused on overhauling the continuous integration infrastructure by migrating to the rayci orchestration and Wanda build system, while introducing comprehensive local build tooling and Kubernetes-based test runners. Simultaneously, the project significantly expanded its documentation and SDK coverage, adding extensive code examples for Ray Core, Serve, Train, and RLlib, alongside new Java and C++ API implementations and performance benchmarks.
2023–2024 — Dashboard modularization and CI v2 infrastructure
40 changes.
This period focused on restructuring the Ray Dashboard into isolated subprocess modules to improve stability and performance, while simultaneously launching the Ray CI v2 infrastructure to modernize build and testing pipelines. Significant effort was also directed toward expanding documentation with new templates for LLM fine-tuning, batch inference, and serving, alongside comprehensive updates to support PyTorch Lightning 2.0 and new Dask APIs.
2025–2026 — LLM serving documentation and CI infrastructure
23 changes.
This period focused on expanding Ray's LLM serving capabilities through extensive documentation examples for vLLM, SGLang, and custom model integration, alongside new tutorials for reinforcement learning and video analysis. Concurrently, the project strengthened its CI and deployment foundations by introducing the raydepsets tool for dependency management, new Docker base images, and a centralized event aggregator for the dashboard.
Features
Add C++ default worker entry point
A new default worker executable is introduced at \cpp/src/ray/worker/default\_worker.cc\. This entry point initializes the Ray C++ API using \ray::RayConfig\ and \ray::Init\, then starts the task execution loop via \ray::RunTaskExecutionLoop\, providing a standard way to run C++ Ray tasks.
cpp/src/ray/worker · high confidence
Add C++ example project with Bazel build and dynamic library support
A new template project has been added to cpp/example to help users get started with the Ray C++ API. The entry includes a Bazel build configuration (\BUILD.bazel, \.bazelrc) that compiles the example and a shared library (example.so) using C++17, along with a run script (run.sh) that handles dynamic library loading via LD\_LIBRARY\_PATH/DYLD\_LIBRARY\_PATH. The example code (example.cc) demonstrates core Ray C++ features including remote functions, actors, and object storage.
cpp/example · high confidence
Add DeepSpeed ZeRO-3 and GPT-J fine-tuning examples to Ray Train documentation
New documentation examples have been added to the Ray Train guide for distributed training with DeepSpeed. The \deepspeed\_example.rst\ page provides an intermediate tutorial on using Ray Data with DeepSpeed ZeRO-3, while the \gptj\_deepspeed\_fine\_tuning.ipynb\ notebook demonstrates an advanced workflow for fine-tuning the GPT-J-6B model. Both examples include explicit runtime environment configurations for required dependencies such as \transformers\, \accelerate\, and \deepspeed\, and feature direct links to the Anyscale console for one-click execution.
doc/source/train/examples/deepspeed · high confidence
Add Docker retag Lambda function for automated image tagging
A new AWS Lambda function has been added to the \docker/retag-lambda\ directory to automate the process of retagging Docker images. The function reads supported Python and CUDA versions from configuration files (\python\_versions.txt\ and \cuda\_versions.txt\) and uses the \docker-retag\ utility to update image tags (e.g., from a specific version to \latest\) for the \ray\, \ray-ml\, \ray-deps\, and \base-deps\ repositories. It retrieves Docker credentials securely from AWS Secrets Manager and provides instructions in a new \README.md\ for packaging and deploying the function.
docker/retag-lambda · high confidence
Add GPU-to-GPU RDT and GRPO reinforcement learning example
A new documentation example has been added at \doc/source/ray-core/examples/rdt/grpo\_contextual\_bandits.py\ that demonstrates reinforcement learning using the Group Relative Policy Optimization (GRPO) algorithm. The example illustrates how to leverage Ray Direct Transport (RDT) for GPU-to-GPU data transfer, featuring a custom contextual bandit environment, a residual MLP model, and Ray actors for the replay buffer and scorer.
doc/source/ray-core/examples/rdt · high confidence
Add Java actor performance benchmark suite
Added a new Java-based performance testing tool in the \java/performance\_test\ module to measure Ray actor communication throughput. The suite introduces \Receiver\ and \Source\ actors to simulate task execution and data transfer, supporting configurable parameters such as argument size, return values, and direct byte buffer usage. An \ActorPerformanceTestBase\ orchestrates the test across multiple actor layers, while \ActorPerformanceTestCase1\ provides a default 1-to-1 benchmark scenario to evaluate engine performance.
_java/performance\test · high confidence
Add PBT visualization example and utilities
Added a new documentation example in \doc/source/tune/examples/pbt\_visualization\ that demonstrates how to visualize Population Based Training (PBT) hyperparameter optimization. The entry includes a Jupyter notebook (\pbt\_visualization.ipynb\) showing how to set up checkpointing, configure the PBT scheduler, and visualize algorithm behavior using a toy optimization problem, along with helper Python utilities (\pbt\_visualization\_utils.py\) for plotting parameter history, true reward values, and animations.
_doc/source/tune/examples/pbt\visualization · high confidence
Add Ray Serve LLM documentation examples for Qwen models
Added new documentation examples in the \doc/source/llm/doc\_code/serve/qwen\ directory that demonstrate how to deploy Qwen models (specifically Qwen2.5-0.5B and Qwen2.5-1.5B-Instruct) using Ray Serve. The changes include a Python script (\qwen\_example.py\) showing programmatic configuration via \LLMConfig\ and a YAML configuration file (\llm\_config\_example.yaml\) with a corresponding loader script (\llm\_yaml\_config\_example.py\). These files serve as both user-facing documentation and CI test fixtures, ensuring the examples are executable and validating deployment status.
_doc/source/llm/doc\code/serve/qwen · high confidence
Add Ray Serve fault tolerance documentation and examples
This change introduces documentation and code examples for Ray Serve's fault tolerance capabilities. It includes a Kubernetes configuration file (k8s\_config.yaml) demonstrating a RayService deployment with Redis for state management and fault tolerance enabled via annotations. Additionally, it provides Python code snippets: sleepy\_pid.py, which illustrates a deployment with a slow initialization, and replica\_health\_check.py, which shows how to implement a custom health check method to detect and recover from issues like broken database connections.
_doc/source/serve/doc\_code/fault\tolerance · high confidence
Add Ray Train examples for Hugging Face Transformers
New documentation and build configuration have been added for Hugging Face Transformers integration with Ray Train. This includes a new BUILD.bazel file to manage notebook tests and CI configurations, a detailed Jupyter notebook demonstrating how to fine-tune a text classifier on the GLUE benchmark using Ray Data and Ray Train, and an RST page providing a runnable code example and a quick-start link to Anyscale.
doc/source/train/examples/transformers · high confidence
Add SGLang integration examples for Ray Serve
Added four new documentation examples demonstrating how to use the SGLang engine with Ray Serve: a single-node serving example, a multi-node serving example with specific placement group configurations, a batch processing example using the Ray Data API, and a client-side query example showing OpenAI-compatible API usage.
_doc/source/llm/doc\code/serve/sglang · high confidence
Add documentation and example scripts for deploying Ray on YARN via Skein
New files have been added to the documentation directory to guide users in deploying Ray clusters on YARN using the Skein framework. This includes a YAML configuration file (ray-skein.yaml) defining head and worker services, a Python script (dashboard.py) to register the Ray dashboard with Skein's UI, and an example script (example.py) demonstrating a basic Ray workload. These resources provide a concrete reference for setting up and running Ray applications in a YARN environment.
doc/yarn · high confidence
Add documentation and test example for vLLM-based audio transcription
This change introduces a new documentation example and CI test script for the Ray Serve vLLM engine backend, demonstrating how to deploy and query a Whisper model for audio transcription. The file \transcription\_example.py\ configures an \LLMConfig\ with the \whisper-small\ model, sets up an OpenAI-compatible client to send audio data, and validates that the deployment reaches a running status. This serves as both a user-facing reference for implementing transcription features and an automated test to ensure the vLLM integration works correctly in CI environments.
_doc/source/llm/doc\code/serve/transcription · high confidence
Add documentation examples for custom datasources and resource allocation
Added two new code snippets to the documentation: \custom\_datasource\_example.py\ demonstrates how to implement and use custom \FileBasedDatasource\ and \RowBasedFileDatasink\ classes for reading and writing image data, and \key\_concepts.py\ illustrates how to configure Ray Tune to reserve cluster resources for Dataset execution by setting \max\_concurrent\_trials\.
_doc/source/data/doc\code · high confidence
Add gRPC proxy documentation and example code
The documentation now includes a complete example for the gRPC proxy feature, demonstrating how to configure gRPC options, define a deployment with unary, streaming, and multiplexed RPCs, and interact with the service using generated Protocol Buffer client code.
_doc/source/serve/doc\_code/grpc\proxy · high confidence
Add gap-filling scheduler for post-merge CI pipeline
A new gap-filling scheduler has been added to the CI infrastructure to automatically re-trigger blocked builds when the latest pipeline run fails. This tool identifies commits between the most recent passing and failing builds, locates the corresponding blocked jobs in Buildkite, and unblocks them to ensure coverage is maintained. It is implemented as a Python module with a CLI entry point, includes unit tests, and is integrated into the Bazel build system for execution on Python 3.10.
_ci/ray\ci/pipeline · high confidence
Add production guide example for Ray Serve with Hugging Face models
The documentation now includes a concrete code example (text\_ml.py) demonstrating how to deploy a summarization and translation pipeline using Ray Serve and Hugging Face Transformers. The example defines two Ray Serve deployments—a Translator and a Summarizer—showing how to bind them together, handle HTTP requests, and perform runtime reconfiguration of model parameters like target language and summary length.
_doc/source/serve/doc\_code/production\guide · high confidence
Added Bazel action listener to generate compile commands database
A new tool has been added to the CI lint infrastructure that listens to Bazel C++ compile actions and generates a compile\_commands.json database. This is implemented via a new Bazel target in ci/lint/generate\_compile\_commands that uses an action\_listener and extra\_action to invoke a C++ binary (extract\_compile\_command.cc), which parses Bazel's extra action protobuf data and outputs JSON entries for each compilation. This enables better integration with IDEs and static analysis tools that rely on compile commands.
_ci/lint/generate\_compile\commands · high confidence
Added documentation code examples for Ray observability and debugging
New Python code snippets have been added to the documentation to demonstrate Ray's observability and debugging capabilities. These examples cover application logging, common environment variable pitfalls and their fixes, memory profiling with Memray for both actors and tasks, exporting custom metrics (counters, gauges, histograms) via the Ray metrics API, and using the distributed debugger with breakpoints and post-mortem debugging.
_doc/source/ray-observability/doc\code · high confidence
Added documentation style guide for AI agents and contributors
A new style guide rule file (doc/.cursor/rules/ray-docs-style.mdc) has been added to enforce consistent documentation standards. This guide defines specific requirements for voice (active, present tense), word choice (plain language, contractions), heading structure, and formatting (MyST admonitions, code block tagging). It also establishes strict rules for cross-references and prohibits the fabrication of technical details, ensuring that all generated or edited documentation under the doc/ directory adheres to the project's quality standards.
doc/.cursor · high confidence
Added example code snippet for documentation contribution guide
A new Python file, example\_module.py, has been added to the documentation source directory to serve as a reference for writing code snippets. This file contains a simple is\even function wrapped with specific comment markers (\\_is\_even\begin\\_ and \_\_is\_even\end\\_), which are used to define the boundaries of code blocks in the documentation.
_doc/source/ray-contribute/doc\code · high confidence
Added installation scripts for Prometheus server and GDRCopy
New shell scripts have been added to the documentation tools directory to automate the setup of specific dependencies. The install-prometheus-server.sh script downloads and extracts the Prometheus server (version 2.8.0) for Linux and macOS systems. The install\_gdrcopy.sh script automates the installation of the GDRCopy library (default version 2.5.1-1) on Debian-based systems by downloading the appropriate package for the specified OS version, CUDA version, and architecture (x64 or aarch64).
doc/tools · high confidence
C++ Worker metric API implementation
The C++ runtime now exposes a metric API for workers, aligning it with the Java and Python implementations. This change introduces the \cpp/src/ray/runtime/metric/metric.cc\ implementation file, which provides concrete classes for Gauge, Histogram, Counter, and Sum metrics, allowing C++ workers to record and manage metrics in a manner consistent with other language runtimes.
cpp/src/ray/runtime/metric · high confidence
C++ worker now supports dynamic library loading and cross-language calls
Users can now load C++ dynamic libraries (e.g., .so files) at runtime via the \code\_search\_path\ configuration, enabling Python drivers to invoke C++ functions and actors using \ray.cross\_language.cpp\_function\ and \ray.cross\_language.cpp\_actor\_class\. This change introduces a new C++ worker build structure in \cpp/BUILD.bazel\ that compiles \libray\_api.so\ and a \default\_worker\, and adds test coverage for these cross-language interactions and C++ job submission scenarios.
cpp · high confidence
C++ worker now supports loading dynamic libraries for remote functions
The C++ worker runtime can now dynamically load shared libraries (e.g., .so, .dll, .dylib) from specified search paths to discover and execute remote functions and actors. This change introduces a \FunctionHelper\ that scans directories or files for dynamic libraries, initializes the Ray runtime within them, and registers their \RAY\_REMOTE\-decorated functions for execution. Additionally, the \ProcessHelper\ has been updated to configure the worker with these code search paths and runtime environment settings, enabling users to extend the C++ worker's capabilities by loading external code at runtime.
cpp/src/ray/util · high confidence
Dashboard module for ingesting and caching Kubernetes platform events
The Ray Dashboard now includes a new platform events module that ingests, caches, and exposes Kubernetes events (such as Pod lifecycle changes) via a new REST API endpoint at /api/v0/platform\_events. This feature is opt-in, controlled by the RAY\_DASHBOARD\_INGEST\_PLATFORM\_EVENTS environment variable, and relies on the KubernetesEventProvider to watch relevant K8s resources and map them to standardized RayEvent protobufs for caching and display.
_python/ray/dashboard/modules/platform\events · high confidence
Documentation example for serving custom vLLM reward models via Ray Serve
This location adds a complete, runnable example demonstrating how to serve a custom vLLM model architecture—specifically a Qwen3-based reward model with a scalar output head—using Ray Serve LLM. The provided code includes a Python plugin that registers the custom model class with vLLM, a Ray Serve server hook, and example scripts (both programmatic and YAML-based) that deploy the model and verify the \/classify\ endpoint returns a single scalar reward score.
_doc/source/llm/doc\_code/serve/custom\vllm · high confidence
Documentation example for the prefix cache-aware router
Added a new documentation example (\prefix\_aware\_example.py\) demonstrating how to configure and use the \PrefixCacheAffinityRouter\ within a Ray Serve LLM application. The example shows setting up an \LLMConfig\ with specific router parameters such as \imbalanced\_threshold\, \match\_rate\_threshold\, and eviction settings, and includes test logic to verify the deployment reaches a running state.
_doc/source/llm/doc\_code/serve/prefix\_aware\router · high confidence
Introduce C++ local-mode object store implementation
Added a new local-mode object store implementation for the C++ worker, consisting of the \LocalModeObjectStore\ class and its supporting \ObjectStore\ interface. This change enables object storage operations (Put, Get, Wait) to function correctly when running Ray in local mode by delegating to the \CoreWorkerMemoryStore\ instead of the distributed Plasma store, ensuring that C++ tasks can serialize, store, and retrieve objects locally without requiring a cluster.
cpp/src/ray/runtime/object · high confidence
Introduce C++ local-mode task execution infrastructure
The C++ runtime now includes a new local-mode task submission path, enabling users to run Ray tasks and actors in a single-process mode without a cluster. This change adds the \LocalModeTaskSubmitter\ component (along with \InvocationSpec\, \TaskExecutor\, and the \TaskSubmitter\ interface) which handles task submission, actor creation, and execution within the local runtime, laying the groundwork for local-mode C++ workflows.
cpp/src/ray/runtime/task · high confidence
Introduce C++ runtime environment support
The C++ API now supports runtime environments, allowing users to configure and manage execution contexts (such as dependencies and environment variables) for C++ tasks and actors. This change adds the \RuntimeEnv\ class and its serialization logic in \cpp/src/ray/runtime/runtime\_env.cc\, enabling runtime environment configuration to be passed alongside task submissions in the C++ worker.
cpp/src/ray/runtime · high confidence
Introduce Ray CI v2 infrastructure for building and testing
The \ci/ray\_ci\ directory now contains the core implementation for the Ray CI v2 system, replacing the previous CI setup. This includes a new Bazel build file (\BUILD.bazel\) defining the test and library targets, and a suite of Python modules for managing Docker containers (e.g., \builder\_container.py\, \linux\_container.py\, \docker\_container.py\) used to build Ray wheels and images. The new infrastructure supports building artifacts for multiple Python versions (3.9–3.14) and architectures (x86\_64, aarch64), and handles publishing to AWS ECR and GCP registries (with Azure temporarily disabled). It also introduces a new sharding mechanism (\bazel\_sharding.py\) for distributing Bazel tests and provides configuration files (\configs.py\, \oss\_config.yaml\) to manage pipeline settings and supported image variants.
_ci/ray\ci · high confidence
Introduce Ray Event Aggregator Agent for centralized event collection and publishing
The dashboard now includes a new Event Aggregator Agent module that collects Ray events via gRPC, buffers them, and publishes them to configurable destinations including GCS, the dashboard head, and external HTTP services. This agent supports multiple consumers with independent cursors, configurable event filtering (including an 'ALL' option), and robust retry logic with exponential backoff. It also provides OpenTelemetry metrics for monitoring publish performance and failures, and includes a dedicated buffer for task event metadata to track dropped attempts.
python/ray/dashboard/modules/aggregator · high confidence
Introduce automated release pipeline with pre-release checks and wheel validation
The release automation pipeline now includes a dedicated pre-release check stage that validates the commit hash, builds the update-version binary, and triggers downstream nightly, weekly, and macOS test suites. Additionally, the pipeline adds a wheel upload and validation stage that publishes wheels to TestPyPI, verifies them across Linux (x86\_64 and arm64) and macOS for Python 3.10–3.14, and finally uploads them to PyPI.
.buildkite/release-automation · high confidence
Introduce automated test bisecting for Linux/Windows and macOS
Adds a new CI tooling module in \ci/ray\_ci/bisect\ that automatically identifies the specific commit causing a test failure. The implementation includes a \Bisector\ class that performs binary search over a range of revisions, supported by \GenericValidator\ (which triggers and monitors Buildkite builds for Linux/Windows tests) and \MacOSValidator\ (which runs local Bazel tests for macOS). The entry point \bisect\_test.py\ orchestrates the process, updates the test state machine, and comments the blamed revision on the associated GitHub issue.
_ci/ray\ci/bisect · high confidence
Introduce base-deps Docker image with unified build infrastructure
This change introduces a new \docker/base-deps\ image that consolidates system-level dependencies (such as \libjemalloc-dev\, \openssh-client\, and \netbase\) and Python packages (including \setuptools\, \numpy\, \psutil\, and \smart\_open\) into a single foundational layer. The build process now uses Miniforge for Conda management, installs the \uv\ package manager, and leverages Wanda configuration files (\cpu.wanda.yaml\, \cuda.wanda.yaml\, \tpu.wanda.yaml\) to generate images for CPU, GPU (CUDA), and TPU platforms across x86\_64 and aarch64 architectures. This unifies the dependency resolution and image creation workflow, replacing previous fragmented approaches with a standardized, architecture-aware base image that supports Python 3.10+ and includes specific TPU ecosystem dependencies like JAX and Flax.
docker/base-deps · high confidence
Introduce centralized CI linting infrastructure and pre-commit hooks
This change establishes a new \ci/lint\ directory containing a comprehensive suite of linting and formatting scripts, including checks for Bazel buildifier, C++ clang-format/clang-tidy, Python docstrings, API stability annotations, and documentation style (Vale). It also adds Git hooks (\prepare-commit-msg\ for signoff, \pre-push\ for pre-commit validation) and a \lint.sh\ orchestrator that integrates these tools into the CI pipeline, ensuring consistent code quality and style enforcement across the repository.
ci/lint · high confidence
Introduce initial Java API for Ray Serve
This change adds the foundational Java SDK for Ray Serve, enabling Java applications to define, configure, and deploy services. The new \Serve\ class provides entry points to start and shut down the Serve instance, while \Deployment\ and \DeploymentCreator\ classes allow users to define deployment logic, set configuration options (such as replica counts, autoscaling, and health checks), and bind arguments. The implementation includes a \ServeControllerClient\ to communicate with the Python-based Serve controller, a DAG-based model for composing deployments, and support for cross-language invocation via \DeploymentHandle\.
java/serve · high confidence
Introduce new CI automation scripts and libraries for image tagging, microcheck test selection, and artifact extraction
This change adds a new Bazel build file and a suite of Python scripts to the CI automation layer. It introduces \determine\_microcheck\_tests.py\ and \determine\_microcheck\_step\_ids.py\ to automatically select high-impact tests for the microcheck pipeline based on historical failure data. It also adds \copy\_wanda\_image.py\ and \extract\_wanda\_artifact.py\ to manage Wanda-cached container images using the \crane\ tool, and provides supporting libraries (\crane\_lib\, \image\_tags\_lib\, \docker\_tags\_lib\) for generating Docker image tags and handling registry operations. Additionally, \check\_nightly\_ray\_commit.py\ is added to validate nightly image contents, and \get\_contributors.py\ is refactored to use a new \GitHubClient\ for release announcements.
_ci/ray\ci/automation · high confidence
Introduce new slim base Docker image for Ray
A new 'slim' base image variant is now available, targeting environments like Anyscale that require a lighter footprint. This image is built on Ubuntu 22.04 (with CUDA support) and includes Miniforge for Python 3.10, the \uv\ package manager, and essential system tools such as \awscli\, \azcopy\, and \dynolog\. It pre-installs a specific set of Python dependencies defined in \requirements.in\ (including \anyscale\, \adlfs\, and \boto3\) and configures the \ray\ user with sudo privileges and a standard working directory.
docker/base-slim · high confidence
Introduce ray-ml Docker image with consolidated ML dependencies
Adds the \ray-ml\ Docker image, which extends the base \rayproject/ray\ image by installing a comprehensive set of machine learning dependencies for RLlib, Serve, and Tune. The build process copies specific requirement files (including core, data, rllib, train, and tune requirements) and executes a script that installs system libraries (such as GCC, CMake, and Mesa GL) and Python packages, ensuring TensorFlow Probability is correctly installed alongside TensorFlow.
docker/ray-ml · high confidence
Introduce raydepsets tool for managing CI dependency lock files
Adds a new Python-based tool, raydepsets, to the CI infrastructure for managing and validating Python dependency lock files. The tool reads declarative YAML configurations to model relationships between lock files as a dependency graph, supporting operations like compile, subset, expand, and relax to ensure consistency across different Python versions, platforms, and CUDA variants. It leverages \uv pip compile\ for reproducible resolution and includes a \--check\ mode to validate that committed lock files are up-to-date and mutually consistent in CI pipelines.
ci/raydepsets · high confidence
Introduction of the C++ Ray API header
The file \cpp/include/ray/api.h\ has been added, providing the main entry point for the C++ Ray API. This header exposes functions for initializing and shutting down the Ray runtime, storing and retrieving objects via \Put\ and \Get\ (including timeout support), and executing remote tasks and actors using \Task\ and \Actor\ builders. It also includes APIs for managing named actors, placement groups, and retrieving the current job namespace.
cpp/include/ray · high confidence
New Azure deployment guide with ARM template and init script
Added documentation and supporting files for deploying Ray on Azure, including an Azure Resource Manager (ARM) template (azure-ray-template.json) and an initialization shell script (azure-init.sh). The ARM template defines the infrastructure for head and worker nodes, supporting configuration for VM sizes, spot instance priorities, and network security rules for services like JupyterLab, Ray Web UI, and TensorBoard. The init script automates the setup of the Conda environment, installation of Ray, and configuration of systemd services for Ray and TensorBoard to ensure they start at boot.
doc/azure · high confidence
New Bazel build infrastructure and dependency definitions
This change introduces a comprehensive set of Bazel build files and helper scripts to the bazel/ directory. It defines build targets for key third-party dependencies including hiredis, jemalloc, msgpack, nlohmann\_json, rapidjson, rocksdb, and spdlog, along with their specific build configurations (e.g., platform-specific compiler options, CMake/Make generators). It also establishes a hermetic Python 3.10 toolchain and runtime environment for CI and internal scripts. Additionally, it provides utility macros for running Python doctests and tests via pytest, a script for deterministic zip file creation, and a dependency resolution helper for the CI driver closure.
bazel · high confidence
New Bazel export options script for CI logging and artifacts
A new \ci/run/bazel\_export\_options\ script has been added to standardize Bazel CI execution. It configures environment variables to direct build event JSON logs to a temporary directory and sets up paths for archiving test failure logs and generating test summaries, with specific handling for Windows (msys) environments to ensure artifacts are correctly mounted and uploaded.
ci/run · high confidence
New C++ API headers for actors, object references, and task execution
The \cpp/include/ray/api\ directory now includes a comprehensive set of new header files that define the core C++ API surface. This introduces \ActorHandle\ and \ActorCreator\ for managing actor lifecycles and remote method calls, \ObjectRef\ with reference counting for distributed object storage, and \ActorTaskCaller\ for invoking actor tasks. Additionally, it adds \Arguments\ for serializing task inputs, \RayConfig\ for cluster and runtime configuration, \RayRemote\ macros for registering remote functions, and \Metric\ classes (Gauge, Histogram, Counter) for observability, establishing the foundational interface for C++ workers to interact with the Ray cluster.
cpp/include/ray/api · high confidence
New CI checks for API documentation consistency and parameter coverage
The \ci/ray\_ci/doc\ package introduces automated checks to ensure the public API surface remains consistent with documentation. The \cmd\_check\_api\_discrepancy\ command compares the \@PublicAPI\ symbols found by importing Ray against those documented in Sphinx reference pages, flagging mismatches, duplicates, and undocumented symbols. Additionally, \cmd\_check\_api\_param\_coverage\ performs a static, diff-scoped analysis to fail pull requests that introduce new \@PublicAPI\ callables or parameters without corresponding \Args:\ entries in their docstrings. These tools help maintain documentation quality by catching drift and missing details early in the development cycle.
_ci/ray\ci/doc · high confidence
New CI infrastructure for package mirroring and build reproducibility
This change introduces new CI tooling to improve build reliability and debugging. It adds a PyPI proxy configuration system (pypi\_proxy\_profile.sh, pypi\_proxy\_agent.sh) that automatically routes pip, uv, and Bazel downloads to an internal CI mirror when available, falling back to public PyPI if the mirror is unreachable. It also includes a bazel mirror downloader (bazel\_mirror\_downloader.sh) to rewrite Bazel's direct HTTP archive downloads to the same mirror. Additionally, a new repro-ci.py script is added, allowing developers to automatically provision an AWS EC2 instance that mirrors a specific Buildkite CI job's environment, enabling local reproduction of CI failures.
ci · high confidence
New CI pipeline infrastructure for conditional test execution and API rule synchronization
This change introduces the core logic for the new RayCI conditional testing system. It adds a Python script to determine which tests to run based on file changes using a configurable rule file (test.rules.txt) and a Bazel build target to execute this logic. Additionally, it includes a new test that ensures the CI rules for API documentation checks remain synchronized with the actual API surface defined in the documentation source code, preventing drift between the two.
ci/pipeline · high confidence
New Data Head module exposes dataset and operator metrics via the Ray Dashboard
The Ray Dashboard now includes a new \DataHead\ module (running as a subprocess) that exposes detailed metrics for Ray Data datasets and their operators. This change adds a new API endpoint \/api/data/datasets/{job\_id}\ which returns dataset state, progress, total rows, and specific Prometheus-backed metrics such as output rows, spilled bytes, current bytes, and CPU/GPU/memory usage. The response structure includes both dataset-level and operator-level metrics, allowing users to monitor data pipeline performance directly from the dashboard. Tests verify the correct schema and unique operator identification in the returned data.
python/ray/dashboard/modules/data · high confidence
New DreamBooth fine-tuning workspace template with LoRA support
A new workspace template for fine-tuning Stable Diffusion models using DreamBooth and Ray Train has been added. This template enables users to generate images of custom subjects by training on a small set of photos, supporting both standard full fine-tuning and parameter-efficient LoRA fine-tuning. The package includes a run script for automated data preparation and training, a Jupyter notebook for interactive image generation, and documentation detailing configuration options such as worker count, learning rates, and dataset paths.
_doc/source/templates/05\_dreambooth\finetuning · high confidence
New DreamBooth finetuning template with LoRA support
The DreamBooth finetuning example now includes a complete, runnable template that supports both full model finetuning and parameter-efficient finetuning using LoRA. The new scripts handle downloading example datasets, caching base models, and building Ray Data pipelines for image preprocessing and tokenization. Users can now train models with LoRA adapters to reduce memory usage and training time, and generate images using the trained models with optional Ray Data-based batch inference for parallelization across multiple GPUs.
_doc/source/templates/05\_dreambooth\finetuning/dreambooth · high confidence
New Intel Gaudi training examples for BERT, Llama, ResNet, and Stable Diffusion
Added Jupyter notebook examples in the Intel Gaudi documentation directory demonstrating how to run distributed training on Intel Gaudi (HPU) accelerators using Ray Train. The new notebooks cover fine-tuning BERT for sequence classification, fine-tuning Llama-2 with LoRA and DeepSpeed, pre-training Llama models, training ResNet-50 for image classification, and fine-tuning Stable Diffusion with textual inversion. Each example includes setup instructions for the required Gaudi Docker container and shows how to configure Ray TorchTrainer with HPU resources.
_doc/source/train/examples/intel\gaudi · high confidence
New JAX training template for Ray Train
Added a new documentation example in \doc/source/train/examples/jax/intro\_to\_jax\_trainer\ that demonstrates how to distribute a JAX/Flax training loop using Ray Train's \JaxTrainer\. The template provides a complete guide for training a GPT-2-style Transformer on the OpenWebText dataset, covering dataset preparation, Ray Data integration, and distributed execution on GPUs or TPUs.
doc/source/train/examples/jax · high confidence
New K8s CI infrastructure and test runners
Added new scripts and configuration files to the ci/k8s directory to support Kubernetes-based testing. This includes install-k8s-tools.sh to provision specific versions of kind (v0.22.0), kubectl (v1.28.4), kustomize (v5.2.1), and helm (v3.12.2); kind.config.yaml for cluster networking setup; prep-k8s-environment.sh to initialize the local cluster; and dedicated runners for chaos tests (run-chaos-test.sh) and KubeRay operator tests (run-operator-tests.sh), which now explicitly use Python 3.10 and support the autoscaler v2 environment.
ci/k8s · high confidence
New Llama-2 finetuning template with LoRA support
This location introduces a new documentation template and accompanying scripts for fine-tuning Llama-2 series models (7B, 13B, and 70B) using Ray Train, DeepSpeed, and Hugging Face Accelerate. The template now supports both full-parameter fine-tuning and parameter-efficient LoRA fine-tuning, including a new \merge\_lora\_weights.py\ script to combine LoRA adapters with base models for evaluation. It provides cluster configuration guidelines for AWS instances, automated dataset creation for GSM8K, and instructions for syncing checkpoints to AWS S3.
_doc/source/templates/04\_finetuning\_llms\_with\deepspeed · high confidence
New PyTorch training examples and CI configuration
The documentation now includes new PyTorch training examples, specifically a ResNet finetuning notebook, a DreamBooth finetuning guide, and a tutorial on converting existing PyTorch code to Ray Train. These examples are supported by a new Bazel build file that configures continuous integration tests, separating GPU-accelerated notebook runs from CPU-only tests to ensure reliable validation of the new content.
doc/source/train/examples/pytorch · high confidence
New Python-based Grafana dashboard panel definitions for Ray Data, Serve, and LLM
The Ray Dashboard now includes a new Python module (\python/ray/dashboard/modules/metrics/dashboards\) that defines Grafana dashboard panels programmatically for Ray Data, Ray Serve, and Ray Data LLM (vLLM). This change introduces structured Python classes (e.g., \Panel\, \Target\) and corresponding JSON base templates to replace or supplement previous static dashboard configurations. Users will see updated visualizations for Ray Data metrics (such as bytes spilled, object store memory, and input/output rates), Ray Serve metrics (including QPS, error rates, and latency percentiles), and new vLLM-specific metrics (like token throughput, cache utilization, and TTFT/TPOT latency). The default dashboard panels are also updated to use this new structure, ensuring consistent rendering of cluster utilization, node counts, and task states.
python/ray/dashboard/modules/metrics · high confidence
New Ray Train V2 documentation code examples
Added a comprehensive set of new code examples in the Ray Train documentation (\doc/source/train/doc\_code/\) demonstrating the V2 API. These include guides for asynchronous validation, checkpointing, fault tolerance, data ingestion, and integration with Ray Tune, alongside quickstarts for PyTorch, TensorFlow, and LightGBM. The examples also cover key concepts like result handling, metric logging, and custom data configurations, providing users with updated patterns for building and managing distributed training jobs.
_doc/source/train/doc\code · high confidence
New Sphinx extensions for API navigation, code callouts, and AI agent documentation
The documentation build now includes four new Sphinx extensions in the \\_ext\ directory. The \api\_sidebar\ extension replaces the previous sidebar approach by capturing the full API toctree once and rendering it as a single, client-side loaded fragment (\api-nav.html\), which prevents build memory issues and keeps individual pages lightweight. A new \callouts\ extension introduces \callout\ and \annotations\ directives, allowing authors to embed code blocks with numbered, circled-line annotations for clearer explanations. The \llms\_txt\ extension replaces the third-party \sphinx-llms-txt\ package with an in-repo generator that produces \llms.txt\ and \llms-full.txt\ files structured by site navigation, optimized for AI coding assistants with features like sectioned output, description extraction, and token-based sharding. Finally, the \queryparamrefs\ extension adds a \query-param-ref\ directive to support documentation links that include URL query parameters.
_doc/source/\ext · high confidence
New Tune documentation code examples
Added a comprehensive set of Python code snippets in the \doc/source/tune/doc\_code\ directory to serve as executable examples for Ray Tune documentation. These files cover key concepts such as defining search spaces, configuring resources, integrating with HyperOpt and Optuna, managing checkpoints, handling fault tolerance and restoration, and implementing custom stopping criteria.
_doc/source/tune/doc\code · high confidence
New base-extra Docker image with cloud and GPU tooling
A new \base-extra\ Docker image layer has been introduced, extending the standard Ray base images with additional system packages and tools. This layer installs the Google Cloud CLI, Azure Storage AzCopy (v10.30.0), and AWS CLI v2, alongside utilities like GDB, NFS client, and SSH. It also includes GPU-specific optimizations for AWS environments, automatically detecting CUDA versions (11, 12, and 13) to install corresponding EFA, GDRCopy, and AWS OFI NCCL components. Python dependencies are managed via a dedicated lockfile, including the Anyscale CLI (\>=0.26.105) and Jupyter ecosystem packages.
docker/base-extra · high confidence
New batch inference examples for LoRA adapters and structural outputs
Added two new Jupyter notebook examples in the documentation for Ray Data LLM batch inference: one demonstrating how to use LoRA adapters with vLLM (including configuration for \enable\_lora\, \max\_lora\_rank\, and dynamic loading paths), and another showing how to enforce structural outputs (guided decoding) using JSON schemas via the \structured\_outputs\_config\ and \sampling\_params\. A corresponding Bazel build file (\BUILD.bazel\) was also added to define and run these notebook examples as GPU tests.
doc/source/llm/examples · high confidence
New batch inference template for GPU-accelerated image classification
Added a new documentation template demonstrating large-scale batch inference using Ray Data and PyTorch. The example walks users through loading the Imagenette dataset from S3, preprocessing images, and running parallel inference across multiple GPUs using a pretrained ResNet model, with results saved to S3 or local disk. The template specifies a minimum compute requirement of 4 nodes with NVIDIA T4 GPUs and utilizes the Anyscale-provided Ray ML image (Python 3.9).
_doc/source/templates/01\_batch\inference · high confidence
New dashboard UI with token authentication and dark mode support
The Ray Dashboard client has been replaced with a new implementation built on MUI and React Router. This update introduces token-based authentication, requiring users to enter a token via a dedicated dialog when enabled, and adds support for dark mode with theme persistence in local storage. The new interface also includes a global context for managing application state, lazy-loaded routes for performance, and various UI components like collapsible sections and duration text formatting.
python/ray/dashboard/client · high confidence
New documentation build infrastructure and LLM serving examples
This change introduces the core tooling for the Ray documentation build system, including a new Makefile, Bazel build rules, and Python scripts (load\_doc\_cache.py, rtd\_doctor.py, update\_cache\_env.py) to support incremental builds, Read the Docs fidelity checks, and cache management. It also adds executable notebook examples and CI tests for Ray Serve LLM direct streaming and multi-host TPU serving, demonstrating how to configure and deploy LLM applications using the build\_openai\_app API.
doc · high confidence
New documentation code examples for Ray Core patterns and features
Added a comprehensive set of Python code examples in the \doc\_code\ directory to illustrate Ray core concepts. These include anti-patterns (e.g., nested \ray.get\, global variables, fine-grained tasks), actor management (synchronization, checkpointing, cancellation, restarts), and advanced execution models like Compiled Graphs (C-Graphs) with NCCL/Gloo tensor transports, cross-language calls, and custom direct transports.
_doc/source/ray-core/doc\code · high confidence
New documentation code examples for Slurm, YARN, and XGBoost clusters
Added new code snippets in the cluster documentation to demonstrate deploying Ray on Slurm (including symmetric-run and launch scripts), YARN (using Skein with dashboard registration), and submitting XGBoost jobs via the JobSubmissionClient. These files provide concrete, runnable examples for users setting up these specific cluster environments.
_doc/source/cluster/doc\code · high confidence
New documentation examples for Ray Data LLM batch inference
Added a comprehensive suite of code examples in the 'Working with LLMs' guide, covering basic vLLM batch inference, OpenAI API integration, and multimodal tasks including image, video, and audio processing. The examples also demonstrate advanced configurations such as classification and embedding model support, custom tokenizer pipelines, disaggregated tokenization, and checkpointing for fault tolerance.
_doc/source/data/doc\code/working-with-llms · high confidence
New documentation examples for multi-GPU LLM serving configurations
Added four new code examples in the multi-GPU serving documentation to demonstrate specific deployment patterns: basic data parallel attention, data parallel autoscaling, prefill-decode disaggregation with data parallel attention, and simplified bundle-per-worker placement group configuration. These files serve as both Sphinx documentation snippets and CI tests to validate the correct usage of the new LLM serving APIs for these advanced multi-GPU scenarios.
_doc/source/llm/doc\_code/serve/multi\gpu · high confidence
New documentation skills for RST-to-MyST conversion and Markdown soft-wrapping
Added two new skills to the Ray documentation tooling: \rst-to-myst\ and \ray-soft-wrap\. The \rst-to-myst\ skill provides a comprehensive guide and mapping for converting existing reStructuredText documentation pages to MyST Markdown, including handling of directives, cross-references, and verification steps to ensure rendered output remains identical. The \ray-soft-wrap\ skill introduces deterministic Python scripts (\softwrap.py\ and \verify.py\) to normalize line wrapping in Markdown files, ensuring each paragraph and list item is on a single line while preserving content integrity through content-invariant, render-equality, and idempotency checks.
doc/.claude/skills · high confidence
New event reporting and export infrastructure in the Ray Dashboard
The Ray Dashboard now includes a dedicated event module that introduces an event agent to monitor local event log files and report them to the dashboard head via an HTTP API, with support for token-based authentication and configurable retry logic. The dashboard head exposes endpoints to receive these events and list cluster events via a new API, while also supporting the generation and persistence of export events (such as task and submission job events) for observability. This change adds comprehensive tests to verify event monitoring, API responses, and export event generation.
python/ray/dashboard/modules/event · high confidence
New example for distributed language model training with Ray and Fairseq
Added a new documentation example in the \lm\ directory that demonstrates how to use Ray to perform fault-tolerant, distributed training of language models (specifically RoBERTa) using the Fairseq library. The example includes a cluster configuration file (\lm-cluster.yaml\) for provisioning AWS infrastructure, a preprocessing script (\preprocess.sh\) to prepare the WikiText-103 dataset, and a Python training script (\ray\_train.py\) that manages distributed actors, handles checkpointing, and automatically restarts training on new resources or failures.
doc/source/ray-core/examples/lm · high confidence
New interactive documentation features: AI assistant, CSAT feedback, and example gallery filtering
The documentation site now includes several new interactive capabilities. An 'Ask AI' chat widget has been added, allowing users to ask questions and receive streamed, syntax-highlighted responses with copy-to-clipboard functionality. A Customer Satisfaction (CSAT) widget enables users to provide feedback on documentation pages, with votes and comments tracked via Google Analytics. The example gallery now supports filtering by use case, library, framework, and contributor, with search and URL-based filtering. Additionally, interactive terminal simulations (Termynal) are now supported for dynamic code examples, and announcement banners can be dismissed and remembered via local storage.
_doc/source/\static/js · high confidence
New local build tooling for wheels and Docker images
The \ci/build\ directory now provides a new set of Python scripts and shell utilities to build Ray wheels and Docker images locally. \build\_wheel.py\ and \build\_image.py\ act as CLIs that invoke the \raymake\ tool, while \build\_common.py\ supplies shared helpers for platform detection and environment setup. This is accompanied by a new \container\_resource\_utils.py\ script that calculates optimal Bazel resource flags based on container cgroup limits, and a \BUILD.bazel\ file that registers these new Python modules and their corresponding unit tests.
ci/build · high confidence
New log streaming module for the Ray Dashboard
The Ray Dashboard now includes a new log module (python/ray/dashboard/modules/log) that provides backend support for streaming and listing logs. This module introduces a LogsManager that resolves log file locations for jobs, actors, and workers, and streams log content via gRPC. It includes utilities for efficient file offset calculation and MIME type registration, enabling users to view and stream logs directly from the dashboard interface.
python/ray/dashboard/modules/log · high confidence
New many-model training template for parallel time-series forecasting
A new documentation template has been added to demonstrate how to parallelize the training of hundreds of time-series forecasting models using Ray Tune. The template utilizes the \statsforecast\ library to fit models on partitions of the M4 forecasting competition dataset, allowing users to run multiple training jobs concurrently across a distributed cluster. It includes a Jupyter notebook (\start.ipynb\) and a README that guide users through setting up dependencies, defining custom training functions, and evaluating cross-validation metrics to select the best model for generating forecasts.
_doc/source/templates/02\_many\_model\training · high confidence
New modular job submission system with CLI, SDK, and dashboard integration
The job submission functionality has been restructured into a new modular package under \python/ray/dashboard/modules/job\. This introduces a dedicated CLI (\ray job\) for submitting, stopping, and listing jobs, alongside a new \JobSubmissionClient\ SDK that communicates with the cluster via HTTP. The dashboard now exposes these capabilities through new \JobHead\ and \JobAgent\ modules, which handle job lifecycle management, log streaming, and status reporting. The system also includes new Pydantic models for job details and a structured logging approach for job supervisors, providing a more robust and consistent interface for managing Ray jobs.
python/ray/dashboard/modules/job · high confidence
New modular reporter subsystem with unified health checks and profiling capabilities
The Ray Dashboard's reporter functionality has been restructured into a new, modular subsystem under \python/ray/dashboard/modules/reporter\. This change introduces dedicated modules for GPU metrics (\gpu\_providers.py\), GPU profiling (\gpu\_profile\_manager.py\), JAX profiling (\jax\_profile\_manager.py\), and CPU/memory profiling (\profile\_manager.py\). It also adds a unified health check endpoint (\/api/healthz\) via \healthz\_agent.py\ and refactors the head and agent reporters (\reporter\_head.py\, \reporter\_agent.py\) to use these new components. The subsystem includes Pydantic models (\reporter\_models.py\) for consistent data serialization and comprehensive unit tests for the new profiling and provider logic.
python/ray/dashboard/modules/reporter · high confidence
New monitoring code examples for Ray Serve
Added new Python code snippets in the Serve documentation to demonstrate key monitoring capabilities. These include examples for creating and using custom metrics (Counters), configuring logging (including JSON encoding, log levels, and access logs), tracking request IDs via the X-Request-ID header, and querying the current Serve status to monitor application health.
_doc/source/serve/doc\code/monitoring · high confidence
New object detection batch inference example and Bazel build configuration
Added a new Jupyter notebook example demonstrating object detection batch inference using PyTorch and Ray Data, including a 'Run on Anyscale' quickstart button. Introduced a Bazel BUILD file to define test targets for running all notebooks in the directory and to discover CI configuration files.
doc/source/data/examples · high confidence
New release pipeline configuration and build orchestration
The release pipeline now uses a new configuration structure in .buildkite/release, including a config.yml defining artifact buckets and queues, and a build.rayci.yml that orchestrates image builds for Python 3.10–3.14 and CUDA 12.3/13.0 variants using Wanda definitions. A new custom-image-build-and-test-init.sh script handles environment setup, including Bazel and uv installation, and generates custom build and test job definitions. Symlinks in the release directory point to shared CI configuration files for forge, images, and wheel builds.
.buildkite/release · high confidence
New template for serving Stable Diffusion models with Ray Serve
Added a new template in \doc/source/templates/03\_serving\_stable\_diffusion\ that demonstrates how to serve a Stable Diffusion model using Ray Serve and FastAPI. The template includes an application entry point (\app.py\) defining an \/imagine\ endpoint, a cluster environment configuration (\cluster\_env.yaml\) based on Ray 2.9.0 with Python 3.9, and instructions for local development and production deployment via Anyscale Services. It also provides a query script (\query.py\) and a Jupyter notebook (\start.ipynb\) to guide users through setup and inference.
_doc/source/templates/03\_serving\_stable\diffusion · high confidence
New tutorial for RL post-training with TRL and Ray Train
Added a new documentation example demonstrating how to fine-tune a Qwen2.5 0.5B model using Hugging Face TRL's Group Relative Policy Optimization (GRPO) algorithm and Ray Train. The notebook and accompanying scripts show how to scale training across multiple GPUs, including specific configurations for AWS and GCP, and provide a custom reward function that disables timeouts to ensure compatibility with Ray's execution environment.
_doc/source/train/examples/transformers/transformer\_reinforcement\learning · high confidence
New video analysis tutorial with Ray Serve
Added a new tutorial demonstrating a production-grade video analysis pipeline using Ray Serve. The example includes a FastAPI-based application (\app.py\) that orchestrates three distinct deployments: a GPU-bound \VideoEncoder\ using SigLIP for embeddings, a CPU-bound \MultiDecoder\ for tag classification and caption retrieval, and a CPU-heavy \VideoAnalyzer\ ingress for S3 downloads and FFmpeg chunking. The tutorial provides a Jupyter notebook (\README.ipynb\) explaining the architecture, setup prerequisites (S3, Pexels API, GPU), and code for generating text embeddings. It also includes supporting scripts for load testing (\client/load\_test.py\), performance benchmarking (\chunk\_video\_benchmark.ipynb\), and a custom coordinated autoscaling policy (\autoscaling\_policy.py\) to manage replica counts across the heterogeneous workloads.
doc/source/serve/tutorials/video-analysis · high confidence
Serve HTTP guide now includes examples for WebSocket, streaming, and FastAPI patterns
The Serve HTTP guide documentation has been expanded with new code examples demonstrating advanced HTTP capabilities. Users can now see how to handle client disconnections and cancellation in streaming responses, implement WebSocket echo servers using FastAPI, and utilize the FastAPI factory pattern to avoid serialization issues with instrumentations. These additions provide concrete patterns for building more robust and interactive web services with Ray Serve.
_doc/source/serve/doc\_code/http\guide · high confidence
Serve documentation code examples updated for new APIs and features
The doc\_code directory has been refreshed with new and updated Python and YAML examples demonstrating Ray Serve capabilities. Key additions include application-level autoscaling with coordinated policies, async inference autoscaling using Celery adapters, and custom autoscaling policies (including class-based and stateful variants). The examples also cover new request routing strategies such as capacity queues, round-robin, and consistent hashing, alongside deployment-scoped actors and cross-node parallelism (tensor/pipeline) for LLMs. Existing code samples have been migrated to the new Application builder API and DeploymentHandle interface, and best practices for async I/O and batching are now explicitly documented.
_doc/source/serve/doc\code · high confidence
Architecture
Bazel build system configuration
The repository now uses a \.bazelrc\ file to configure the Bazel build system, enabling platform-specific configurations and strict action environments.
(repo-wide) · high confidence
Java build system migrated to Bazel with modular dependency management
The Java build infrastructure has been replaced with a Bazel-based system, introducing modular definitions for the \api\, \runtime\, and \serve\ components in \java/BUILD.bazel\. This change centralizes dependency management in \java/dependencies.bzl\, specifying versions such as Jackson 2.18.8, GSON 2.11.0, and Log4j 2.25.4. The new build process includes automated generation of Maven POM files and Java proto sources via dedicated Python scripts, and enforces code quality through Google-style Checkstyle configurations. Additionally, the \java/runtime\ module now uses the Maven Shade plugin to relocate Guava and Protobuf packages to prevent conflicts, and the release workflow is streamlined through \build-jar-multiplatform.sh\ for cross-platform JAR assembly and Maven Central deployment.
java · high confidence
StateHead module converted to run as a subprocess
The StateHead component in the Ray Dashboard has been refactored to run as a separate subprocess module rather than within the main dashboard process. This architectural change isolates state API handling (such as listing actors, jobs, nodes, and tasks) to reduce GIL contention and improve stability. The module now uses a dedicated ThreadPoolExecutor with a constrained number of workers (defaulting to 1) to manage concurrency, and it integrates with the new subprocess routing infrastructure while maintaining existing REST API endpoints for state observability.
python/ray/dashboard/modules/state · high confidence
Behavioural changes
1 commit (0 fixes) modifying doc/data
A change to existing behaviour in doc/data — 1 commit, 1 file.
doc/data · low confidence · unverified
Docker image build documentation and release tagging automation
The docker directory now includes a README.md that documents the image hierarchy, clarifying that images without CPU/GPU suffixes are aliases for the CPU variant built on Ubuntu 22.04, and introduces a new fix-docker-latest.sh script to automate retagging release images to 'latest' via an AWS Lambda function.
docker · high confidence
Docs site overhaul: new templates, AI assistant, and analytics
The documentation site has been significantly updated with a new set of Jinja2 templates in doc/source/\_templates. A custom 404 page now uses a \<base\> tag to ensure relative links resolve correctly on missing pages, and a new API sidebar template loads navigation client-side to improve performance. The homepage (index.html) has been revamped with a new layout and code examples for Ray Data, Train, Tune, and Serve. A new 'RunLLM' AI chat widget named 'Ray Docs' has been added via extrahead.html, replacing previous integrations. Additionally, the site now includes a Customer Satisfaction (CSAT) feedback widget, Fathom analytics, and a pre-paint script to prevent announcement banner flashing. Navigation templates have been updated to include a 'Try Managed Ray' link pointing to console.anyscale.com, and template files for Jupyter Notebooks and Markdown have been added to support documentation authoring.
_doc/source/\templates · high confidence
Documentation build system overhaul and example gallery redesign
The documentation build infrastructure has been significantly restructured to improve reliability, speed, and maintainability. A new standalone API stub generator (api\_autogen.py) allows the API-doc consistency check to run without a full Sphinx build, while a centralized mock configuration (api\_mock\_imports.py) ensures consistent handling of optional dependencies across the build, stub generation, and CI checks. The example gallery has been redesigned with a new template (examples.html) featuring client-side search and filtering by use case, library, framework, and contributor. External example templates are now fetched at build time from templates.ci.ray.io with pinned build IDs (template\_pins.json) and robust retry logic to prevent build stalls. Additionally, a custom 404 page has been added, and the llms.txt generation has been migrated to an in-repo extension with stricter exclusion rules for low-signal pages.
doc/source · high confidence
Documentation examples updated for the new DeploymentHandle API
The documentation code samples in the model composition section have been updated to demonstrate the new \DeploymentHandle\ API. The new examples illustrate how to chain deployments by passing \DeploymentResponse\ objects directly between steps, how to integrate Ray tasks with Serve responses using \\_to\_object\_ref()\, and how to handle streaming responses via \DeploymentResponseGenerator\. These changes reflect the shift to the new handle API, which is now the default behavior.
_doc/source/serve/doc\_code/model\composition · high confidence
Explicit symbol visibility control for libray\_api.so on Linux
A new linker version script (ray\_api\_exported\_symbols\_linux.lds) has been introduced to explicitly define the public API surface of the libray\_api.so library on Linux. This change restricts exported symbols to those starting with 'ray' (e.g., TaskExecutionHandler, InitRayRuntime) and necessary absl::flags symbols, while hiding all other internal symbols. This ensures a stable ABI by preventing accidental exposure of internal implementation details and resolves potential conflicts with user-linked absl::flags libraries.
cpp/symbols · high confidence
Java API refactored with new type-safe call interfaces and placement group management
The Java API in the \java/api\ module has been restructured to provide a more type-safe and fluent interface for remote execution. This includes the introduction of dedicated caller classes (such as \ActorTaskCaller\, \TaskCaller\, and cross-language variants like \PyActorTaskCaller\) that replace previous invocation patterns, and the addition of a \PlacementGroups\ utility class to manage placement group creation, retrieval, and removal. The \Ray\ entry point now exposes \getActor\ for named actor lookup and \wait\ with a \fetchLocal\ parameter, while \ObjectRef\ and \BaseActorHandle\ define the core interfaces for object retrieval and actor lifecycle management (including \kill\ with restart options).
java/api · high confidence
Java runtime refactored into new local and native execution modes
The Java runtime implementation has been restructured to separate local (development) and cluster (native) execution paths. A new \AbstractRayRuntime\ base class now defines the core API surface, with \RayDevRuntime\ handling local mode via a \LocalModeTaskExecutor\ and \LocalModeObjectStore\, while \RayNativeRuntime\ manages cluster mode by initializing a GCS client and delegating to native C++ components through \NativeTaskExecutor\ and \NativeObjectStore\. Actor handles are now language-agnostic wrappers (\NativeActorHandle\) that support serialization and reference counting for Java, Python, and C++ actors, and configuration is centralized in \RayConfig\ to support both run modes.
java/runtime · high confidence
Migrate CI Docker infrastructure to Wanda pipeline with Ubuntu 22.04 base
The CI Docker build system has been restructured to use the Wanda pipeline, replacing the previous build mechanism. All base images now target Ubuntu 22.04 (Jammy), and the toolchain has been updated to use Clang 14 and the uv package manager. The new structure introduces a hierarchical image dependency chain: \base.test\ (Ubuntu 22.04 with Python, Bazel, and Miniforge) serves as the foundation for \base.build\ (which adds Java), which in turn supports specialized images like \base.ml\ (ML dependencies), \base.gpu\ (CUDA 12.8/13.0 support), and \forge\ (CI orchestration). This change also adds support for ARM64 builds and introduces new depset-based dependency management for specific test environments like Data, LLM, and Doc.
ci/docker · high confidence
Migrate Ray CI to rayci orchestration and Wanda build system
The Buildkite configuration has been replaced with a new structure driven by the rayci tool and Wanda container images. Pipeline definitions are now organized into modular YAML files (e.g., \core.rayci.yml\, \build.rayci.yml\) that define steps using Wanda images for base environments, wheel builds, and Docker image construction. Conditional test selection is handled by \test.rules.txt\ and \always.rules.txt\, which map file changes to tags to determine which tests run. This change introduces support for Python versions 3.10 through 3.14 and CUDA versions up to 13.0.0 in the build matrix, and moves the CI infrastructure itself to run on the new 'forge' base image.
.buildkite · high confidence
New Buildkite hooks for artifact cleanup and test summary uploads
Added three new Buildkite hooks (pre-command, post-command, post-artifact) to manage the /tmp/artifacts directory on the host machine. The pre-command hook cleans up stale artifacts and sets up permissions before jobs run, while the post-artifact hook ensures cleanup after uploads to prevent node contamination. The post-command hook reads condensed test summaries from /tmp/artifacts/test-summaries and uploads them as Buildkite annotations for easier visibility of failures, then removes the summary files. These hooks also handle platform-specific cleanup logic for macOS and Linux, using Docker for permission-safe operations on Linux.
.buildkite/hooks · high confidence
New C++ API initialization and configuration subsystem
The C++ API now introduces a dedicated initialization flow via \ray::Init()\ and \ray::IsInitialized()\, replacing previous ad-hoc startup patterns. This change adds a new internal configuration layer (\ConfigInternal\) that consolidates runtime settings, supporting both programmatic \RayConfig\ objects and command-line arguments parsed via \absl::flags\. Users can now configure cluster connectivity (\ray\_address\), Redis authentication (\ray\_redis\_username\, \ray\_redis\_password\), code search paths, and runtime environment details through a unified interface, with the system automatically detecting whether to run in single-process or cluster mode based on the provided configuration.
cpp/src/ray · high confidence
New CI environment management scripts for build, test, and credential setup
The \ci/env\ directory now contains a suite of new shell and Python scripts to standardize the CI environment. \install-miniforge.sh\ replaces Miniconda with Miniforge for Linux, macOS, and Windows, supporting Python 3.14 environments. \install-bazel.sh\ installs Bazelisk and configures CI-specific build settings, including remote caching and symlink validation. \install-dependencies.sh\ orchestrates the installation of base system packages, Node.js (upgraded to v20 on macOS), and Python dependencies. \check\_minimal\_install.py\ validates that minimal installation tests are not tainted by extra packages. \setup\_credentials.py\ securely retrieves and exports API keys for external services like Weights & Biases, Comet ML, and Snowflake from AWS Secrets Manager, while \cleanup\_test\_state.py\ provides utilities to remove test artifacts from these services after runs.
ci/env · high confidence
New documentation stylesheets for custom UI components
The documentation site now includes dedicated CSS files to style several new or updated interface elements: a custom 404 error page with a search bar, an AI assistant chat widget (Kapa), a customer satisfaction (CSAT) feedback widget, dismissable announcement banners, an interactive example gallery with search and filtering, a landing page with tabs and community cards, and terminal-style code blocks (termynal). These styles ensure consistent theming across light and dark modes and isolate component-specific rules to prevent leakage into other pages.
_doc/source/\static/css · high confidence
New pre-hooks for building platform-independent wheels and managing dependency constraints
Added three new pre-hook scripts to the CI dependency set workflow: \build-placeholder-wheel.sh\ generates a platform-independent wheel using \uv\ with \RAY\_DEBUG\_BUILD=deps-only\; \remove-compiled-headers.sh\ strips GPU-specific index URLs and find-links from compiled requirements files to ensure platform independence; and \strip-pytorch-lightning-constraint.sh\ removes \pytorch-lightning\ and \numpy\ constraints from Python 3.13 requirements and removes the \lightning\ package from tune test requirements to maintain backward compatibility with \pytorch-lightning\ v1.
_ci/raydepsets/pre\hooks · high confidence
New unified macOS CI scripts for testing and wheel building
This change introduces new shell scripts (macos\_ci.sh, macos\_ci\_build.sh, pypi\_proxy.sh) that consolidate and standardize the macOS continuous integration workflow. The test script (macos\_ci.sh) sets up a Python 3.10 environment with specific Torch versions, filters flaky tests, and runs various test suites (smoke, small, medium, large, core dashboard, and C++ tests) using Bazel, while also handling result persistence and log uploads. The build script (macos\_ci\_build.sh) handles the creation of macOS wheels and JARs, including specific setup for Apple Silicon (arm64) and Java, and manages uploading artifacts to branch/latest directories based on the build branch. Both scripts utilize a new pypi\_proxy.sh helper to resolve Python packages through a CI mirror, improving reliability against public PyPI outages.
_ci/ray\ci/macos · high confidence
Pre-baked intersphinx inventory snapshots for faster, more reliable documentation builds
Documentation builds now use committed snapshots of third-party Sphinx inventories (such as NumPy, PyTorch, and pandas) stored in \doc/source/\_intersphinx\, eliminating the need to fetch these files over the network during every build. This change significantly reduces build startup time (by 20–60 seconds) and prevents intermittent failures caused by network flakiness or expired signed URLs. A new \refresh.py\ script and a monthly scheduled job keep these snapshots up to date, ensuring that cross-references remain valid while maintaining the ability to fall back to upstream sources if a snapshot is missing.
_doc/source/\intersphinx · high confidence
RLlib documentation updates for new API stack and evaluation workflows
The RLlib documentation code examples have been updated to reflect the new API stack, which is now the default for PPO, SAC, and DQN. The examples demonstrate the new \AlgorithmConfig\ builder pattern for configuring environments, exploration, and evaluation (including parallel evaluation and duration settings). New code snippets illustrate how to perform inference with DreamerV3 using the \RLModule\ directly, how to use the \SingleAgentEpisode\ class for episode data handling, and how to configure and customize replay buffers. The examples also show the migration to the Gymnasium API for custom environments.
_doc/source/rllib/doc\code · high confidence
RLlib new API stack enabled by default for BC, MARWIL, and CQL
The new RLlib API stack (based on RLModules) is now enabled by default for the Behavioral Cloning (BC), MARWIL, and Conservative Q-Learning (CQL) algorithms. Users running these algorithms will automatically benefit from the new API's features and performance improvements without needing to explicitly configure the legacy API.
rllib · high confidence
Ray Core documentation examples migrated to Jupyter notebooks and MyST Markdown
The Ray Core examples in the documentation have been converted from RST to Jupyter notebooks (.ipynb) and MyST Markdown (.md). This change introduces new notebook-based examples for batch prediction, gentle walkthrough, highly parallel tasks, map-reduce, Monte Carlo Pi estimation, hyperparameter tuning, parameter servers, and Pong. The build system has been updated to use BUILD.bazel files to manage these examples, with explicit test targets defined in the parent doc/BUILD.bazel to avoid duplicate test runs and allow specific configuration (size, team, tags) for each notebook.
doc/source/ray-core/examples · high confidence
Ray Tune examples migrated to Jupyter notebooks with Bazel test coverage
The Ray Tune documentation examples in \doc/source/tune/examples\ have been converted to Jupyter notebooks (\.ipynb\) and are now automatically tested via Bazel. A new \BUILD.bazel\ file defines a test suite that runs these notebooks, with the \RAY\_TRAIN\_V2\_ENABLED\ environment variable set to \1\ to ensure compatibility with the new training API. Several notebooks are excluded from the standard test run due to high resource requirements (\pbt\_ppo\_example.ipynb\), missing CI dependencies (\tune-aim.ipynb\, \bohb\_example.ipynb\), or use of legacy APIs (\pbt\_transformers.ipynb\), while \tune-xgboost.ipynb\ is reserved for separate GPU-based testing.
doc/source/tune/examples · high confidence
Refactored Node Head into a subprocess module with new data organization and caching
The Node Head module has been restructured to run as a subprocess module, introducing a new \DataOrganizer\ and \DataSource\ architecture to manage node, actor, and worker statistics. This change implements a configurable dead-node cache (controlled by \RAY\_maximum\_gcs\_dead\_node\_cached\_count\) to prevent GCS RPC errors from overwhelming the dashboard, and optimizes performance by offloading blocking operations to a constrained thread pool executor. The update also adds comprehensive tests for actor states (alive, dead, infeasible, placement group) and node info endpoints, ensuring the dashboard correctly displays updated node summaries and actor metadata.
python/ray/dashboard/modules/node · high confidence
Release test infrastructure migration to BYOD and Python 3.10
The release test infrastructure has been modernized by migrating tests to use the Bring Your Own Docker (BYOD) image system and standardizing on Python 3.10. This change replaces legacy SDK runners and custom environment configurations with a unified, reproducible build process, ensuring that release tests run against the exact wheel versions and dependencies intended for the release. Additionally, support for older Python versions (3.7, 3.9) has been removed from the release test suite to align with the project's current runtime requirements.
release · high confidence
Serve dashboard module refactored into a subprocess with SDK and API updates
The Serve dashboard module has been restructured: the ServeHead component now runs as a subprocess module, and the REST API endpoint for listing applications now supports an optional 'api\_type' filter. The Serve SDK has been updated to require the Ray dashboard's HTTP(S) address (rejecting ray:// URLs) and enforces a minimum cluster version of 1.12. Additionally, runtime environment values are now redacted in the dashboard response when appropriate.
python/ray/dashboard/modules/serve · high confidence
Third-party library patches for build stability and new capabilities
This update applies a comprehensive set of patches to vendored third-party libraries to improve build compatibility, fix platform-specific issues, and enable new features. Key changes include: enabling configurable gRPC thread counts via the RAY\_num\_grpc\_internal\_threads environment variable; adding explicit shutdown and immediate export (ExportNow) APIs to the OpenCensus stats exporter to prevent hangs during worker death; enabling mTLS support for the OpenTelemetry OTLP gRPC exporter; fixing Cython compilation by ensuring it uses the project's COPTS; resolving build failures on Windows (MSVC) for hiredis, msgpack, and prometheus-cpp by addressing header conflicts and missing definitions; patching Boost to use BoringSSL instead of OpenSSL and exporting headers for external use; updating Bazel rules for protobuf (exec\_tools to tools) and Apple support (Xcode version check); and fixing spdlog rotation filename formatting. These patches ensure the build system works correctly across Windows, macOS (including Apple Silicon), and Linux, while providing better control over telemetry and network resources.
thirdparty · high confidence
Unified API navigation and enhanced documentation styling
The documentation now features a unified API navigation section that dynamically loads a shared sidebar fragment, automatically highlighting the current page and expanding its parent sections for easier browsing. Additionally, new CSS styles have been added to improve the visual presentation of dataframes within the docs and to support Binder integration badges.
_doc/source/\static · high confidence
Updated Dask-on-Ray documentation examples
The documentation code samples for Dask-on-Ray have been refreshed to reflect the new Dask task class API. The examples now demonstrate modern usage patterns, including the use of \dask.annotate\ for resource constraints, the \RayDaskCallback\ interface for task hooks and caching, and updated graph inspection methods (\collections\_to\_expr\) compatible with recent Dask versions.
_doc/source/ray-more-libs/doc\code · high confidence
Updated PyTorch Lightning examples to support Lightning 2.0 and new import paths
The Lightning example notebooks in the documentation have been updated to support PyTorch Lightning 2.0. This includes adding fallback logic to import from \lightning.pytorch\ (the new namespace) while maintaining compatibility with older \pytorch\_lightning\ installations, and updating dependency requirements in the examples to specify \lightning\>=2.0\.
doc/source/train/examples/lightning · high confidence
Updated Ray Serve getting-started code samples to use the new DeploymentHandle API
The code samples in the Ray Serve documentation have been updated to reflect the new \DeploymentHandle\ API. The examples now demonstrate how to compose deployments by passing a \DeploymentHandle\ (imported from \ray.serve.handle\) as a dependency to other deployments, enabling inter-deployment communication via methods like \.remote()\. The samples also show the explicit \Application\ concept using \.bind()\ and \.options()\ before running the application with \serve.run()\.
_doc/source/serve/doc\_code/getting\started · high confidence
Updated Serve deployment configuration example
The documentation example for configuring a Serve deployment has been updated to include specific deployment options such as replica count, resource allocation (CPUs/GPUs), concurrency limits, and health check settings. The code now demonstrates how to bind these options to a deployment class and run the application, providing a more complete reference for users setting up model deployments.
_doc/source/serve/doc\_code/configure\_serve\deployment · high confidence
Windows CI environment rebuilt with Bazelisk, uv, and pinned dependencies
The Windows CI infrastructure has been restructured to use Bazelisk instead of the standard Bazel binary, with a shortened output base path to avoid Windows file-length limits. The build environment now installs the \uv\ package manager and pins specific versions for Python, \pyopenssl\, and CA certificates to resolve compatibility and trust issues. Additionally, the configuration disables Java and extra C++ API installations, restricts remote cache uploads for premerge pipelines, and cleans up system caches to reduce image size and improve build speed.
_ci/ray\ci/windows · high confidence
Fixes
Fix performance regression in GPU object handling
Resolves a performance regression by ensuring that small and non-GPU objects are correctly cleared from the object reference tracking, preventing unnecessary overhead during object management.
src/ray · high confidence
Test coverage
Added C++ cluster mode and cross-language integration tests; Added C++ test infrastructure with example unit tests; Added C++ worker API and serialization tests; Added C++ worker example applications for metrics, jobs, and KV store; Added CI verification for external PyTorch tutorial examples; Added Java benchmarking, documentation demos, and actor lifecycle tests; Added mock implementations for core worker and GCS components; Added test harness and validation tests for Ray Dashboard modules; Added tests for the raydepsets CLI and workspace configuration; Refactor object spilling test.
Dependencies
Updated dependency lockfiles and build manifests
This change updates the compiled dependency lockfiles (requirements\compiled\\*.txt) and build manifests (pyproject.toml, java/pom.xml, package-lock.json) to reflect the latest resolved versions for the project's Python, Java, and JavaScript dependencies. It also introduces new test data files for the dependency set tooling and updates documentation build requirements, including a pin on setuptools to 84.0.0 and updates to Sphinx and related extensions, ensuring deterministic builds and compatibility with Python 3.12+.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 43 → 48 (+4.9)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 75 → 82 (+7.0)
- Architecture 63 → 75 (+12.1)
- Maturity 73 → 77 (+3.6)
- Readiness 20 → 40 (+20.4)
- Security 60 → 45 (-15.5)
- Domain Modelling 100 (new)
- Accessibility 46 (new)
Resolved (124)
- (anonymous) (cognitive 17) (doc/source/_static/api-nav-loader.js)
- Coverage not measured — test suite did not build
- Critical CVE: [GHSA redacted] (python/ray/dashboard/client/package-lock.json)
- Critical CVE: [GHSA redacted] (release/air_examples/dreambooth/dreambooth/requirements.txt)
- Critical CVE: [GHSA redacted] (release/ray_release/byod/requirements_ml_byod_3.10.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/tune-test-requirements.txt)
- Critical CVE: [GHSA redacted] (release/air_examples/dreambooth/dreambooth/requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/test-requirements.txt)
- Critical CVE: [GHSA redacted] (release/ray_release/byod/requirements_ml_byod_3.10.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/core-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/core-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/train-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements_compiled.txt)
- Critical CVE: [GHSA redacted] (python/ray/dashboard/client/package-lock.json)
- Critical CVE: [GHSA redacted] (python/requirements/ml/core-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/core-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/test-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/tune-test-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/test-requirements.txt)
- Critical CVE: [GHSA redacted] (python/requirements/ml/core-requirements.txt)
- …and 104 more
New (4199)
- APPOConfig.training (cognitive 16) (rllib/algorithms/appo/appo.py)
- APPOConfig.training (cyclomatic 17) (rllib/algorithms/appo/appo.py)
- APPOTorchPolicy.loss (cognitive 17) (rllib/algorithms/appo/appo_torch_policy.py)
- AWSNodeProvider._create_node (cognitive 24) (python/ray/autoscaler/_private/aws/node_provider.py)
- AWSNodeProvider._merge_tag_specs (cognitive 20) (python/ray/autoscaler/_private/aws/node_provider.py)
- AWSNodeProvider.create_node (cognitive 16) (python/ray/autoscaler/_private/aws/node_provider.py)
- AWSNodeProvider.fillout_available_node_types_resources (cognitive 23) (python/ray/autoscaler/_private/aws/node_provider.py)
- AWSNodeProvider.terminate_nodes (cognitive 17) (python/ray/autoscaler/_private/aws/node_provider.py)
- ActorClass._remote (cognitive 64) (python/ray/actor.py)
- ActorClass._remote (cyclomatic 55) (python/ray/actor.py)
- ActorDetail.ActorDetailPage (cognitive 21) (python/ray/dashboard/client/src/pages/actor/ActorDetail.tsx)
- ActorDetail.ActorDetailPage (cyclomatic 21) (python/ray/dashboard/client/src/pages/actor/ActorDetail.tsx)
- ActorHandle._actor_method_call (cyclomatic 16) (python/ray/actor.py)
- ActorMethod._remote (cognitive 20) (python/ray/actor.py)
- ActorMethod._remote (cyclomatic 18) (python/ray/actor.py)
- ActorPoolStrategy.init (cognitive 26) (python/ray/data/_internal/compute.py)
- ActorPoolStrategy.init (cyclomatic 21) (python/ray/data/_internal/compute.py)
- ActorReplicaWrapper.check_ready (cognitive 30) (python/ray/serve/_private/deployment_state.py)
- ActorReplicaWrapper.check_ready (cyclomatic 18) (python/ray/serve/_private/deployment_state.py)
- ActorTable.ActorTable (cognitive 34) (python/ray/dashboard/client/src/components/ActorTable.tsx)
- …and 4179 more
Changes since last survey
- 300 commits — 267 feature/other, 33 fixes
By area
- python/ray — 186 commits
- doc/source — 36 commits
- src/ray — 14 commits
- release/nightly_tests — 13 commits
- (root) — 7 commits
- release/ray_release — 7 commits
- python/deplocks — 6 commits
- release/release_data_tests.yaml — 6 commits
- doc/requirements-doc.txt — 3 commits
- .buildkite/dependencies.rayci.yml — 2 commits
- .buildkite/doc.rayci.yml — 2 commits
- .buildkite/test.rules.test.txt — 2 commits
- .github/CODEOWNERS — 2 commits
- .vale/styles — 2 commits
- doc/redirects — 2 commits
- .buildkite/release-automation — 1 commit
- ci/build — 1 commit
- ci/docker — 1 commit
- ci/raydepsets — 1 commit
- doc/.claude — 1 commit
Notable commits
- fix: [Data] Fix Multitenancy Placement Check Failure (#66268)
- fix: Revert "[Data] Add autoscaler_scales_up_when_compute_bound release test (#65772)" (#66376)
- fix: Revert "[Data] [1/4] Add generated_id_column to CheckpointConfig with V2 Parquet row-group IDs" (#66183)
- fix: [Autoscaling] Fix request_remaining cases that #66001 missed (#66370)
- fix: [CI] Fix deepspeed astral's cu128 index incompatible with Torch 2.9 (#65936)
- fix: [Core] Fix test_runtime_env_working_dir failures on Windows (#66365)
- fix: [Data] Fix ResourceBudget backpressure causing pipeline stall (#64601)
- fix: [Data] Fix WebDataset docs and encoder typing (#65278)
- fix: [Data] Fix crash when partition pruning removes every file in a manifest (#65955)
- fix: [Data] Fix empty global aggregation and duplicate aggregation column names (#66218)
- fix: [Data] Fix flaky CI timeouts in test_join and test_projection_fusion (#66223)
- fix: [Data] Fix hash shuffle after workers scale to zero (#66043)
- fix: [Data] Fix silent data loss when a file read retries mid-file (#66011)
- fix: [Fix][Core/Client] Skip replayed chunks of completed requests after reconnect (#66476)
- fix: [RLlib] Fix DQN exploration device mismatch (#66198)
- fix: [RLlib] Fix test_offline_env_runner (#66212)
- fix: [RLlib][docs] Fix broken tuned-example links on the algorithms page (#66338)
- fix: [Train] Fix _WrappedDataLoader record_stream for nested Tensors (#66344)
- fix: [Tune] Fix Repeater crash when search finishes (#65960)
- fix: [Tune] Fix qrandint and qlograndint to sample values uniformly (#65864)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ray-project/ray was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit c1cd00a387f9b9f3c04a05bb1b5a93443ed19aa0 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-9984f8053b7b.