Skip to content
CAI
Software that uses CAICheck a score

Rust-GPU/rust-cuda

66.6

Adequate · 30 September 2026

55.7k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a Rust-based toolkit for developing and executing GPU-accelerated applications on NVIDIA hardware. It provides a comprehensive stack that includes a custom NVVM code generation backend to compile Rust code for CUDA, alongside low-level bindings for the CUDA Driver API, cuBLAS, cuDNN, and OptiX ray tracing. The project also supplies a curated standard library for GPU kernels, memory management utilities, and various examples demonstrating linear algebra, hashing, and hardware-accelerated rendering.

How it got here

2021 — Initial project scaffolding and core CUDA bindings

26 changes.

This period established the foundational infrastructure for a Rust-based CUDA toolkit, including the initial repository setup, workspace configuration, and comprehensive documentation. It introduced core crates for low-level CUDA driver access (cust), GPU kernel compilation (rustc\_codegen\_nvvm, nvvm), and a curated standard library for GPU programming (cuda\_std). Additionally, it laid the groundwork for high-level libraries by adding initial implementations for cuBLAS (blastoff), cuDNN, OptiX ray tracing, and GPU-compatible random number generation.

2022–2026 — OptiX integration and compiler testing

29 changes.

This period focused on integrating NVIDIA OptiX hardware ray tracing through new device-side bindings, examples, and version-aware build configurations. It also established a comprehensive compile-time test suite to validate CUDA kernel generation, atomic operations, and language feature support across various compute architectures.

Features

Add CUDA atomic intrinsics and mid-level emulation for fences, loads, and stores

The \cuda\_std\ crate now includes low-level PTX atomic intrinsics (for loads, stores, fences, and CAS) and mid-level wrappers that emulate atomic operations on older GPU architectures (pre-Compute Capability 7.0) using memory barriers. This enables correct atomic behavior across a wider range of CUDA devices by falling back to volatile loads/stores surrounded by \membar\ instructions when hardware atomic support is unavailable.

_crates/cuda\std/src/atomic · high confidence

Add DeviceCopy derive macros for cust types

The \cust\_derive\ crate now provides \DeviceCopy\ and \DeviceCopyCore\ derive macros. These macros automatically implement the \cust::memory::DeviceCopy\ or \cust\_core::DeviceCopy\ traits for structs, enums, and unions, while generating compile-time checks to ensure all nested fields also implement the respective trait.

_crates/cust\derive · high confidence

Add OptiX hardware ray-tracing window example

The \ex03\_window\ example now demonstrates hardware-accelerated ray tracing using the OptiX API. It compiles a CUDA kernel (\ex03\_window.cu\) into PTX via a new \build.rs\ script and uses the \optix\ and \cust\ crates to launch a ray-generation program that renders a dynamic color pattern to a window. The example manages the full OptiX pipeline lifecycle, including module loading, shader binding table setup, and CUDA context integration, while using \gl\ and \glfw\ for windowing and display.

_crates/optix/examples/ex02\_pipeline, crates/optix/examples/ex03\window · high confidence

Add OptiX hardware raytracing example for mesh rendering

The ex04\_mesh example now uses OptiX for hardware-accelerated raytracing instead of software or previous methods. This includes a CUDA kernel (kernels/src/lib.rs) defining raygen, closest-hit, any-hit, and miss programs, a build script (build.rs) to compile these kernels to PTX, and a Rust renderer (src/renderer.rs) that sets up the OptiX pipeline, acceleration structure, and shader binding table to render a mesh onto a GL window.

_crates/optix/examples/ex04\mesh · high confidence

Add async API kernel sample for element-wise increment

A new kernel sample has been added to the introduction area that demonstrates an asynchronous API pattern for performing an element-wise increment on GPU data. The provided Rust code defines a CUDA kernel that calculates global thread indices and applies an addition operation to the target memory buffer, including safety documentation regarding thread count limits relative to data size.

_samples/introduction/async\api/kernels · high confidence

Added GPU Developer Toolbox (GDT) common library for OptiX examples

The \crates/optix/examples/common/gdt\ directory now includes the source code and build configuration for the GPU Developer Toolbox (GDT) library. This addition provides shared mathematical utilities (linear and affine spaces, quaternions) and CMake modules to locate and configure dependencies such as OptiX 7.0, CUDA, Intel TBB, and GLUT, enabling the OptiX hardware ray-tracing examples to build and run.

crates/optix/examples/common · high confidence

Added Rust/CUDA matrix multiplication sample with Kahan summation

A new sample demonstrating matrix multiplication in Rust using the CUDA ecosystem has been added to the introduction samples. The implementation includes a CUDA kernel (\matrix\_mul\_cuda\) that utilizes shared memory for performance and employs the Kahan summation algorithm to improve numerical stability and reduce floating-point accumulation errors. The host-side application manages device memory allocation, kernel launching via the \cust\ crate, and validates the results against a CPU reference calculation.

samples/introduction/matmul, samples/introduction/matmul/kernels · high confidence

Added sample project structure and build scripts

The samples directory now includes a README.md providing an overview and navigation to the Introduction chapter. Additionally, build.rs scripts have been added for the async\_api and matmul samples, utilizing the cuda\_builder crate to compile CUDA kernels into PTX format during the build process.

samples · high confidence

Added scripts to generate CUDA libdevice intrinsics and download OptiX headers for CI

This change introduces a new toolchain in the \scripts\ directory to automate the generation of Rust bindings for NVIDIA libdevice math intrinsics. It adds \gen\_libdevice\_json.py\ to parse the libdevice PDF documentation into a structured JSON file, and \gen\_intrinsics.py\ to convert that JSON into \std\intrinsics.rs\, exposing raw \\\nv\\ functions for use in the codebase. Additionally, a new \download\_ci\_optix.bash\ script is added to fetch OptiX 7.0 headers into the CI environment, and a \vast-ai.sh\ script is included to facilitate building and running binaries on vast.ai infrastructure.

scripts · high confidence

Added xtask utility to extract individual LLVM functions

A new xtask command \extract\_llfns\ has been added, allowing users to take an LLVM IR file and split it into separate standalone LLVM IR files, one for each function defined within. The tool uses \llvm-extract\ to isolate each function and outputs them to a specified directory, sorting the results by function length.

xtask · high confidence

Initial NVVM LLVM wrapper implementation

The NVVM codegen backend now includes a C++ wrapper library (rustc\_llvm\_wrapper) that bridges Rust and LLVM. This wrapper provides FFI bindings for LLVM operations, including memory management, pass manager configuration, and intrinsic handling. It supports LLVM versions 4 through 21, with conditional compilation ensuring compatibility across major LLVM API changes (e.g., pass manager migration in LLVM 19+). The implementation includes specific handling for NVVM target features, such as address spaces and function attributes, enabling the backend to generate correct NVPTX code.

_crates/rustc\_codegen\_nvvm/rustc\_llvm\wrapper · high confidence

Initial PTX lexer and type definitions

Added a new lexer for the PTX ISA in \crates/ptx/src/lexer.rs\ that tokenizes ASCII source files using the \ascii\ crate for performance, along with corresponding type definitions in \types.rs\ for directives (such as version, target, kernel, and function declarations) and token kinds. Included unit tests in \lexer\_tests.rs\ to verify tokenization of punctuation, identifiers, directives, and instructions, and exposed the lexer and types via \lib.rs\.

crates/ptx · high confidence

Initial debug info generation support for NVVM codegen

The NVVM codegen backend now includes a complete debug information subsystem, enabling the generation of DWARF debug metadata for compiled code. This change introduces the core infrastructure for mapping MIR source scopes to LLVM debug scopes, handling variable locations, and managing namespace hierarchies. Users will now have access to basic debugging capabilities, including source-level variable inspection and stack traces, when building with debug info enabled.

_crates/rustc\_codegen\_nvvm/src/debug\info · high confidence

Initial implementation of the NVVM codegen backend

This change introduces the \rustc\_codegen\_nvvm\ crate, providing a new code generation backend for NVIDIA GPU targets (NVVM/PTX). The implementation covers core compiler infrastructure including ABI handling (with specific adjustments for passing structs and arrays by value), LLVM IR generation for functions, constants, and statics, inline assembly support, and attribute management. It also includes the allocator shim, target machine factory configuration, and initial support for atomic operations and 128-bit integer emulation, enabling Rust code to be compiled for CUDA devices.

_crates/rustc\_codegen\nvvm/src · high confidence

Initial project setup and documentation

The repository is initialized with core configuration files including a Rust toolchain pinned to nightly-2026-04-02, a Nix flake for reproducible development environments, and a VS Code dev container configuration. It also introduces comprehensive project documentation, including a CONTRIBUTING guide, GitHub issue and PR templates, a CODEOWNERS file, and dual Apache-2.0/MIT licensing.

(repo-wide) · high confidence

Initial release of the NVVM crate for safe CUDA kernel compilation

This change introduces the \crates/nvvm\ crate, providing high-level safe Rust bindings to the NVVM compiler (libnvvm) for writing CUDA GPU kernels using a subset of LLVM IR. The crate exposes functions to query NVVM, IR, and debug metadata versions, defines a comprehensive \NvvmError\ enum for handling compiler errors, and implements \NvvmOption\ to manage compilation settings such as architecture (\NvvmArch\), debug info generation, and floating-point optimizations. It serves as the foundational layer for NVVM-based code generation within the project.

crates/nvvm · high confidence

Initial release of the cust CUDA driver API wrapper

This change introduces the \cust\ crate, a Rust wrapper around the CUDA Driver API. It provides core abstractions for CUDA programming, including device enumeration and attribute retrieval (\device.rs\), error handling (\error.rs\), asynchronous event tracking (\event.rs\), kernel function launching with grid/block sizing (\function.rs\), and CUDA Graph management for capturing and replaying execution sequences (\graph.rs\). The crate also includes utilities for linking PTX code (\link.rs\), managing legacy and primary contexts (\context.rs\, \legacy.rs\), and handling external memory resources (\external.rs\). Additionally, it ships with a sample CUDA kernel source (\add.cu\) and its compiled PTX representation (\add.ptx\) for testing and demonstration purposes.

crates/cust/src · high confidence

Introduce CUDA Array and Unified Memory APIs

The memory module now includes support for CUDA Array objects via the new \array\ submodule, providing \ArrayFormat\ and \ArrayObjectFlags\ for creating and managing texture/surface arrays. Additionally, unified memory capabilities are expanded with the introduction of \UnifiedBox\ and \UnifiedPointer\ types, allowing for heap-allocated memory that is seamlessly accessible by both host and device, alongside existing device memory allocation functions.

crates/cust/src/memory · high confidence

Introduce CUDA GPU programming macros

Added the \cuda\_std\_macros\ crate, which provides procedural macros for GPU kernel development. The \\#\[kernel\]\ attribute registers functions as GPU kernels, enforcing safety constraints such as requiring \unsafe\ blocks, \Copy\ parameters, and no return values. The \\#\[gpu\_only\]\ attribute allows defining CPU fallbacks that panic when executed outside the GPU target. The \\#\[externally\_visible\]\ attribute ensures functions are preserved in the PTX output for linking. Additionally, the \\#\[address\_space\]\ macro allows users to explicitly place static variables into specific GPU memory spaces (global, shared, constant, or local), with detailed documentation on the safety requirements for shared memory usage.

_crates/cuda\_std\macros · high confidence

Introduce CUDA builder crate with configurable architecture and optimization options

A new \cuda\_builder\ crate is added to provide a builder API for compiling Rust GPU crates using \rustc\_codegen\_nvvm\. This allows users to configure the target CUDA architecture (defaulting to \compute\_75\ with support for CUDA 12.9 suffixes), enable or disable debug line info, control NVVM optimization passes, and toggle specific floating-point behaviors like fast sqrt/division and FMA contraction. The builder also supports exporting LLVM IR for debugging, overriding \libm\ calls with faster \libdevice\ intrinsics, and targeting OptiX hardware raytracing with immediate abort on panic.

_crates/cuda\builder · high confidence

Introduce CUDA runtime bindings and kernel launch support

This change adds the foundational runtime layer for CUDA operations in the \cuda\_std\ crate. It introduces raw FFI bindings to the CUDA Driver API (including device management, stream handling, and IPC memory functions) and provides a safe Rust wrapper for creating and managing CUDA streams. Additionally, it implements a \launch!\ macro that allows users to launch GPU kernels with specified grid and block dimensions, parameter buffers, and stream targets, while also integrating the \glam\ library for vector type conversions.

_crates/cuda\std/src/rt · high confidence

Introduce OptiX 9 bindings and runtime stubs

The \crates/optix-sys\ crate now provides Rust bindings for the OptiX 9 SDK. The build script locates the SDK via environment variables, parses the version from headers, and uses bindgen to generate bindings for OptiX types and functions. It also compiles a C stub file to handle OptiX's runtime function table loading mechanism, exposing an \optixInit\ function for initialization.

crates/optix-sys · high confidence

Introduce OptiX device-side Rust bindings for hardware ray tracing

This change adds the \optix\_device\ crate, providing a Rust interface for NVIDIA OptiX hardware ray tracing on the GPU. It exposes device-side functions for tracing rays, handling intersections, and managing payloads via inline assembly, organized by OptiX program types (raygen, miss, anyhit, closesthit). The library supports custom hit kinds, triangle barycentrics, ray attributes, and motion transforms, enabling users to implement hardware-accelerated rendering logic directly in Rust.

_crates/optix\device · high confidence

Introduce OptiX hardware ray tracing support

This change adds the \optix\ crate, providing a Rust wrapper for NVIDIA OptiX 7 to enable hardware-accelerated ray tracing. It introduces core components including \DeviceContext\ for GPU context management, \Accel\ for building geometry and instance acceleration structures, and a GPU-accelerated AI \Denoiser\ to improve image quality with fewer path tracing iterations. The implementation includes safe APIs for building acceleration structures, managing compilation caching, and handling OptiX-specific errors, along with documentation and Glam math integration.

crates/optix/src · high confidence

Introduce blastoff crate for high-level cuBLAS bindings

The new \blastoff\ crate provides a high-level Rust interface to the NVIDIA cuBLAS library, enabling GPU-accelerated linear algebra operations. It exposes a \CublasContext\ for managing the cuBLAS handle and supports Level 1 (scalar/vector) operations such as \amin\, \amax\, \axpy\, \copy\, \dot\, \nrm2\, \rot\, \scal\, and \swap\, as well as Level 3 (matrix) operations including generic matrix multiplication (\gemm\). The API supports single precision (\f32\), double precision (\f64\), and complex types (\Complex32\, \Complex64\), with optional half-precision (\f16\) support via the \half\ feature. It also includes utilities for stream execution, precision configuration (e.g., TF32 tensor cores), and atomic mode settings.

crates/blastoff · high confidence

Introduce cuda\_std as a GPU-targeted standard library

This change adds the \cuda\_std\ crate, providing a curated set of abstractions for writing GPU kernels in Rust. It introduces modules for thread and block indexing (\thread\), warp-level operations including shuffles, reductions, and votes (\warp\), and safe atomic types with block, device, and system scopes (\atomic\). It also provides GPU-compatible floating-point math via the \GpuFloat\ trait and \FloatExt\ extension, raw libdevice intrinsics (\intrinsics\), CUDA-specific pointer handling (\ptr\), dynamic shared memory access (\shared\), and safe printing/assertion macros (\io\).

_crates/cuda\std/src · high confidence

Introduce gpu\_rand crate with GPU-compatible random number generators

The new gpu\rand crate provides a suite of fast, non-cryptographic random number generators (including Xoroshiro and Xoshiro variants) designed to run on both CPU and GPU. It exposes a default 64-bit generator (Xoroshiro128\\*) and a GpuRand trait for generating uniform and normal floating-point numbers, enabling users to reuse the same random states across CPU and GPU environments.

_crates/gpu\rand · high confidence

Introduce high-level Rust bindings for the CUDA NVPTX compiler

Adds the \ptx\_compiler\ crate, providing a safe Rust interface to the CUDA NVPTX compiler. Users can now compile PTX source strings into cubin binaries via the \NvptxCompiler\ struct, with automatic resource management through \Drop\ implementations. The API exposes detailed error handling via the \NvptxError\ enum and allows retrieval of compilation logs (errors and info) as strings, ensuring proper cleanup of compiler handles even when compilation fails.

_crates/ptx\compiler · high confidence

Introduction of DeviceCopy trait for CUDA-safe type markers

The \cust\_core\ crate now exposes a \DeviceCopy\ marker trait, allowing users to explicitly mark types that can be safely copied to or from a CUDA device. This trait ensures types contain no references to non-device-accessible memory and does not imply Rust's \Copy\ semantics. The crate provides a \DeviceCopy\ derive macro for automatic implementation and includes built-in implementations for primitive types, pointers, \MaybeUninit\, \Option\, \Result\, tuples, arrays, and specific vector/matrix types from the \glam\ and \vek\ libraries when their respective features are enabled.

_crates/cust\core · high confidence

New CUDA examples for GEMM, SHA-2, and vector addition

Added new example applications demonstrating CUDA capabilities: a GEMM benchmark comparing naive and tiled matrix multiplication kernels against cuBLAS, SHA-256/SHA-512 hashing examples using both one-shot and incremental APIs, and a vector addition example. These examples showcase the use of the \cuda\_std\ library for kernel development and the \cust\ crate for CUDA runtime management.

(repo-wide) · high confidence

New async API sample demonstrating CUDA event timing and concurrency

A new Rust sample has been added at samples/introduction/async\_api that demonstrates how to use CUDA events to measure GPU execution time and enable concurrent CPU-GPU operations. The example shows recording events in a non-blocking stream, performing asynchronous memory copies and kernel launches, and querying event status to coordinate work without blocking the CPU, while also printing timing metrics for both GPU and CPU activity.

_samples/introduction/async\api · high confidence

New container images for Rocky Linux 9 and Ubuntu 24 with CUDA 13 and LLVM 21 support

The container directory now includes new Dockerfiles for Rocky Linux 9 (with CUDA 12 and 13) and Ubuntu 24.04 (with CUDA 12, 13, and a debug variant). These images provide updated base environments, including support for CUDA 13.3.1 and the addition of an LLVM 21.1.8 build path for the \ubuntu24-cuda13-llvm21\ variant, which sets \LLVM\_CONFIG\_21\ to enable the \llvm21\ cargo feature. The containers also introduce helper scripts (\dcr\ and \dex\) to manage persistent development containers with GPU access and volume mounts for Cargo and Rustup data.

container · high confidence

New cuDNN backend API and activation/attention operations

The \cudnn\ crate introduces a new backend API layer (using \cudnnBackend\*\ functions) that provides builder-based construction for convolution operations (forward, backward data, and backward filter) and their configuration. Additionally, the crate adds support for neuron activation functions (forward and backward passes for modes like ReLU, Sigmoid, Swish, etc.) and multi-head attention operations (including descriptors, sequence data handling, and forward/backward passes), enabling these specific neural network primitives to be executed via the GPU.

crates/cudnn · high confidence

New device memory allocation types with async support

The \cust\ crate introduces a new \memory::device\ module containing \DeviceBox\, \DeviceBuffer\, \DeviceSlice\, and \DeviceVariable\. These types provide heap-allocation and buffer management on the CUDA device, featuring both synchronous and asynchronous (stream-based) allocation, copying, and deallocation. \DeviceBox\ and \DeviceBuffer\ support async operations via \new\_async\ and \drop\_async\, while \DeviceSlice\ offers a zero-sized type representation for device memory slices. All types implement \Send\ and \Sync\ where appropriate and integrate with the \bytemuck\ feature for zero-initialization.

crates/cust/src/memory/device · high confidence

New i128 arithmetic kernel for CUDA

Added a new CUDA kernel in the i128 demo example that performs 128-bit integer operations (addition, subtraction, multiplication, bitwise logic, shifts, and division/rem) on parallel arrays. This allows users to execute and benchmark these specific 128-bit arithmetic routines on the GPU.

_examples/i128\demo/kernels · high confidence

Behavioural changes

Consolidated CUDA binding modules with feature-gated access

The cust\_raw crate now exposes a unified set of modules for CUDA, cuDNN, and OptiX bindings (including driver, runtime, cuBLAS, cuBLASLt, cuBLASXt, NVPTX compiler, and NVVM) that are conditionally compiled based on specific features. Each module uses build-time generated code from the OUT\_DIR, ensuring that only the requested bindings are included in the final binary, which simplifies the API surface and reduces compilation overhead for users who do not need all CUDA components.

_crates/cust\raw/src · high confidence

Denoiser example updated to use glam and modern image-rs APIs

The denoiser example in crates/optix/examples/denoiser has been refactored to replace the vek and image crates with glam and the updated image-rs API. This change updates image loading to use ImageReader, simplifies vector math operations using glam's Vec3 (such as direct component access and splatting), and adjusts the RGB-to-linear conversion and clamping logic to align with these library changes.

crates/optix/examples/denoiser · high confidence

Improved Windows cuDNN discovery and Linux CUDA 13 support

The build system now automatically discovers cuDNN installations on Windows by scanning for vX.Y versioned directories under standard NVIDIA paths, rather than relying on hardcoded legacy paths. On Linux, the build script has been updated to include architecture-specific include directories (e.g., x86\_64-linux-gnu, aarch64-linux-gnu) to support CUDA 13.0, which relocates headers into these arch-specific folders. Users benefit from more robust automatic detection of cuDNN versions and successful compilation against newer CUDA 13 toolchains without manual environment variable overrides.

crates/cudnn-sys · high confidence

New CUDA examples and improved vecadd build reliability

The examples directory now includes a README and dedicated build scripts for the vecadd, gemm, i128\_demo, and sha2\_crates\_io examples, showcasing GPU and CPU usage with the cuda\_builder. The vecadd example has been updated to explicitly link advapi32 on Windows to resolve nanorand symbol errors, and its build script now supports environment variables to dump the final LLVM module or emit LLVM IR for debugging.

examples · high confidence

OptiX version-aware build configuration

The build process now automatically adjusts Rust configuration flags based on the installed OptiX version to ensure compatibility. Specifically, it enables \optix\_build\_input\_instance\_array\_aabbs\ for versions prior to 7.2.0, \optix\_module\_compile\_options\_bound\_values\ for versions 7.2.0 and above, and \optix\_pipeline\_compile\_options\_reserved\ and \optix\_program\_group\_options\_reserved\ for versions 7.3.0 and above.

crates/optix · high confidence

Path tracer example re-enabled with GPU math library migration and OptiX disabled

The path tracer example has been moved to \crates/optix/examples/path\_tracer\ and re-enabled for use. The implementation has been migrated from the \vek\ math library to \glam\, updating vector operations (e.g., \normalized\ to \normalize\, \magnitude\ to \length\) and type definitions (e.g., \Vec3\<f32\>\ to \Vec3\). Additionally, OptiX hardware raytracing support is currently disabled; the \optix\ module and related CUDA denoiser/renderer logic are commented out, leaving the example to rely solely on CPU and standard CUDA kernel rendering.

_crates/optix/examples/path\tracer · high confidence

Primary context handling is now the default for CUDA context management

The library has shifted its default context management strategy from the legacy thread-local stack model to a primary context model that aligns with the CUDA Runtime API. This change improves compatibility with runtime-based libraries like cuBLAS and cuFFT, which previously caused issues when interacting with the explicit context stack. Users now benefit from a simpler, reference-counted context model where creating a context for a device returns the same underlying instance, eliminating the need for manual unowned context handling in multi-threaded scenarios. The legacy context handling remains available in the \legacy\ module for users who require the old behavior.

crates/cust/src/context · high confidence

Rebuilt CUDA bindings with improved macro handling and SDK detection

The build script for the CUDA bindings has been restructured to improve reliability and maintainability. It now features a centralized CUDA SDK detection module that locates the toolkit via environment variables or platform-specific defaults and exposes version metadata to dependent crates. The binding generation process has been consolidated into a unified callback system that correctly handles function renames through C macro expansion, ensuring accurate symbol mapping. Additionally, the build script now generates bindings for the CUDA driver, runtime, cuBLAS, cuBLASLt, cuBLASXt, NVPTX compiler, and NVVM libraries, while disabling comment generation to prevent documentation warnings and excluding specific CUDA node parameter structs from derived traits.

_crates/cust\raw/build · high confidence

Updated changelog and build configuration for CUDA 13 compatibility

The changelog has been updated to document recent API changes, including the deprecation of \vek\ in favor of \glam\, the implementation of \memcpy\_dtoh\, and the restoration of \Stream::add\_callback\ behavior to execute on context errors. Additionally, the build script now detects CUDA driver versions to conditionally enable configuration flags (such as \cuMemAdvise\_v2\ and \reserved\_graph\_node\_16\), ensuring compatibility with CUDA 13 features like merged function signatures and anonymous union field layouts.

crates/cust · high confidence

Upgrade to LLVM 21 and CUDA 13.3 with i128 emulation and warp intrinsic fixes

The NVPTX codegen backend now supports LLVM 21 (and CUDA 13.3) alongside the legacy LLVM 7 path, allowing users to build with modern toolchains via the \llvm21\ feature. To ensure compatibility with LLVM 21's stricter verifier, warp shuffle and match intrinsics now pack their return values into a single \i64\ instead of a struct, avoiding invalid alignment attributes. The backend also adds emulation for 128-bit integer operations (add, sub, mul, bitwise) and provides symbols for \cuda\_std\ warp intrinsics, while completely removing support for 32-bit CUDA.

_crates/rustc\_codegen\nvvm · high confidence

Test coverage

Added UI compile tests for language features and standard library intrinsics; Added compile tests for CUDA atomic operations; Added compile tests for CUDA float extensions and thread indexing/synchronization; Added compile tests for CUDA shared memory allocations; Added compile-time test suite for Rust CUDA; Added compile-time tests for CUDA warp-level functions; Added compile-time tests for kernel operations; Added compiletests for NVVM intrinsic truncation and target feature inheritance; Added tests for LLVM wrapper attribute handling.

Dependencies

Initial Cargo workspace and dependency lockfile

The project now includes a \Cargo.lock\ file and a root \Cargo.toml\ defining a workspace that aggregates multiple crates (such as \cust\, \cuda\_builder\, \cuda\_std\, \cudnn\, \gpu\_rand\, \nvvm\, and \optix\) along with various examples and test harnesses. This establishes the dependency graph and version constraints for the Rust CUDA toolkit.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 64 → 67 (+3.0)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 86 → 86 (+0.0)
  • Architecture 100 → 98 (-1.6)
  • Maturity 61 → 65 (+4.5)
  • Readiness 68 → 64 (-4.2)
  • Security 55 → 62 (+7.5)

Resolved (11)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • High IaC: DS-0029 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • High IaC: DS-0029 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • High IaC: WD-DOCKER-0001 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Medium IaC: DS-0013 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • Medium IaC: WD-DOCKER-0003 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • Medium IaC: WD-DOCKER-0003 (container/ubuntu24-cuda13-llvm19/Dockerfile)
  • Off-boarding risk: anonymized user #1

New (42)

  • High IaC: DS-0029 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • High IaC: DS-0029 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • High IaC: WD-DOCKER-0001 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Inconsistent API for retrieving version information. CublasContext and CudnnContext return a tuple (u32, u32, u32) directly. Context and UnownedContext return a CudaResult wrapping an opaque or unspecified type (likely requiring further calls). CudaApiVersion uses a builder-like or accessor pattern (get, major, minor). This forces users to handle versioning differently depending on which context or API they are querying.
  • Inconsistent handling of stream context. CublasContext has a with_stream method that takes a closure, suggesting a scoped stream execution, but also has set_stream (implied by Cudnn having it, though Cublas doesn't explicitly list set_stream in the provided snippet, it has with_stream). More critically, CudnnContext has set_stream but no with_stream. This creates an inconsistency in how users manage stream scope: one library encourages scoped execution via closures, the other uses imperative state setting.
  • Inconsistent resource management pattern: Some types use a static method drop(ctx: Self) (Cublas, Cudnn, Context, Event, ArrayObject) while others rely on standard Rust Drop trait or do not expose a manual drop method in the signature list. While drop as a static method is a valid pattern for explicit resource release, the inconsistency lies in the fact that CudaBuilder and GpuBuffer/GpuBox do not show a drop method, implying reliance on RAII, whereas the context/event types require explicit manual cleanup. This mixes RAII and manual memory management paradigms without a clear, consistent boundary.
  • Low cohesion: Builder (LCOM4 4) (crates/rustc_codegen_nvvm/src/builder.rs)
  • Low cohesion: CodegenCx (LCOM4 14) (crates/rustc_codegen_nvvm/src/context.rs)
  • Low cohesion: CudaBuilder (LCOM4 15) (crates/cuda_builder/src/lib.rs)
  • Low cohesion: CudnnContext (LCOM4 4) (crates/cudnn/src/context.rs)
  • Medium IaC: DS-0013 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • Medium IaC: WD-DOCKER-0003 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • Medium IaC: WD-DOCKER-0003 (container/ubuntu24-cuda13-llvm21/Dockerfile)
  • Off the main sequence: blastoff
  • Off the main sequence: cust_raw
  • Off-boarding risk: anonymized user #1
  • Outdated: anyhow
  • …and 22 more

Changes since last survey

  • 9 commits — 5 feature/other, 4 fixes

By area

  • .github/workflows — 3 commits
  • crates/cuda_std — 3 commits
  • crates/rustc_codegen_nvvm — 1 commit
  • guide/src — 1 commit
  • tests/compiletests — 1 commit

Notable commits

  • fix: ci: fix digest artifact pattern matching llvm21 image digests
  • fix: fix(atomic): use system memory scope for SystemAtomic types
  • fix: fix(nvvm): cast bit-counting intrinsic results back to the return width
  • fix: fix(warp): fix float casting in warp_match_any/warp_match_all
  • change: Grant Pages deploy artifact read permission
  • change: Point LLVM 21 prebuilt URLs at the upstream release
  • change: Upgrade modern backend to LLVM 21.1.8 and CUDA 13.3
  • change: docs(guide): warn that the default NvvmArch fails silently on pre-Turing GPUs
  • change: mark warp_shuffle_64/128/16/8 as gpu_only

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

Rust-GPU/rust-cuda was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 3a82c35a7196731fb60eebdac2a40d2b1bf0ba95 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-505904ce13c1.