vosen/ZLUDA
55.0
Adequate · 28 September 2026
83.8k
lines of production code
Rust
primary language
2
measurements over time
What this system is
ZLUDA is a compatibility layer that enables CUDA applications to run on AMD hardware by intercepting CUDA API calls and translating them into equivalent AMD HIP operations. It provides comprehensive support for core CUDA libraries, including cuBLAS, cuFFT, cuSPARSE, and cuDNN, by mapping them to their ROCm counterparts like rocBLAS and hipFFT. The system also features a PTX-to-LLVM compiler for offline code translation, a persistent kernel cache for performance, and tools for DLL injection and API tracing to facilitate debugging and deployment.
How it got here
2020–2021 — Workspace restructuring and HIP backend migration
18 changes.
The project was reorganized into a Cargo workspace with over 40 member crates, establishing the foundational structure for the offline compiler, injection tools, and performance bindings. Significant work focused on rewriting the CUDA driver API implementation to delegate to the HIP runtime, enabling execution on AMD hardware, while introducing Windows DLL injection capabilities via Microsoft Detours.
2022–2025 — AMD backend and library integration
44 changes.
This period focused on establishing a comprehensive AMD GPU backend by integrating LLVM for code generation and binding to ROCm libraries such as rocBLAS, hipBLASLt, MIOpen, and hipFFT. The project significantly expanded its compatibility layer to support NVIDIA's core performance libraries (cuBLAS, cuDNN, cuFFT, cuSPARSE) by mapping them to their AMD equivalents, while simultaneously rewriting the PTX parser and compiler pipeline to improve robustness and maintainability.
2026 — 32-bit CUDA and AMD GPU support
6 changes.
This period focused on extending compatibility to 32-bit CUDA applications on Windows through a new inter-process communication architecture and dedicated 32-bit metadata handling. It also expanded hardware support by introducing Rust bindings for AMD's rocSPARSE and hipFFT libraries, enabling Fourier transform and sparse matrix operations on ROCm hardware.
Features
32-bit CUDA support via inter-process communication
Users can now run 32-bit CUDA applications on Windows. The \zluda32\ library acts as a client that spawns the 64-bit \zluda64\_server\ process, communicating with it via Windows shared memory and events to forward CUDA API calls. This architecture allows 32-bit processes to leverage the 64-bit server's capabilities, with the server handling the actual CUDA operations and returning results to the 32-bit client.
_zluda32, zluda64\_server, zluda\_server\common · high confidence
Add 32-bit CUDA support via nvapi and nvapi\_trace
This change introduces support for 32-bit CUDA applications by adding the \nvapi\ and \nvapi\_trace\ modules. The \nvapi\ module provides a Rust binding to the NVIDIA API, exposing functions for GPU enumeration, memory info, and CUDA compute capabilities. The \nvapi\_trace\ module implements a dynamic loader that intercepts calls to \nvapi.dll\ (or \nvapi64.dll\ on 64-bit systems), allowing 32-bit applications to interact with NVIDIA drivers through a traced interface. This enables compatibility for 32-bit CUDA workloads that previously relied on 64-bit specific paths.
_nvapi, nvapi\trace · high confidence
Add AMD device library bitcode and license
The device-libs directory now includes the AMD device libraries in bitcode form (ockl.bc) along with the corresponding University of Illinois/NCSA Open Source License (LICENSE.TXT) and a README. This provides the necessary runtime library components for device-side execution.
_llvm\zluda/src/device-libs · high confidence
Add MIOpen integration to zluda\_dnn
The zluda\_dnn module now integrates AMD's MIOpen library to provide backend support for cuDNN operations. This change introduces the infrastructure to load and dispatch MIOpen functions (via \MIOpenVtable\), manages context handles and search caches for convolution algorithms, and exposes a set of implemented cuDNN API functions (such as convolution forward/backward data and filter operations) that are routed through this new MIOpen backend.
_zluda\dnn · high confidence
Add MIOpen system bindings and build configuration
The \ext/miopen-sys\ crate has been added to provide low-level bindings for the MIOpen library. This includes a build script that links against the \MIOpen\ dynamic library on non-Windows platforms (searching \/opt/rocm/lib/\) and a generated Rust source file exposing MIOpen constants, types, and C API functions (such as handle creation and error handling) for use by higher-level components like \zluda\_dnn\.
ext/miopen-sys · high confidence
Add NVML device management stubs for AMD GPUs
The zluda\_ml library now includes platform-specific implementations (via impl\_unix.rs and impl\_win.rs) and common utilities (impl\_common.rs) to expose a subset of the NVIDIA Management Library (NVML) API. On Unix systems, this enables basic device discovery, PCI bus ID resolution, and memory usage reporting by mapping NVML calls to the ROCm SMI library (rocm\_smi\_sys). On Windows, these functions are stubbed as unimplemented. This change allows applications relying on NVML for AMD GPU information to proceed without immediate crashes, though full feature parity is not yet achieved.
_zluda\ml · high confidence
Add PTX intrinsic implementation library
The ptx/lib directory now includes a new C++ source file (zluda\_ptx\_impl.cpp) and its compiled bitcode artifacts (zluda\_ptx\_impl.bc, zluda\_ptx\_impl\_constrained.bc). This library provides device-side implementations for various PTX intrinsics and system registers (such as activemask, thread/cluster IDs, lane masks, and bit-field extract operations) required for CUDA code execution, replacing previous no-op or missing stubs with functional logic.
ptx/lib · high confidence
Add cuDNN 8 support with delayed loading of AMD HIP runtime on Windows
This change introduces a new zluda\_dnn8 module that provides bindings for cuDNN 8. On Windows, the build configuration links against delayimp.lib and configures the loader to delay-load amdhip64\_7.dll, ensuring the AMD HIP runtime is loaded only when needed. The library exports unmangled cuDNN 8 functions and registers a Windows-specific hook to open the AMD HIP library if it is already loaded in the process.
_zluda\dnn8 · high confidence
Add cuDNN 9 compatibility layer for ROCm 7
This change introduces a new \zluda\_dnn9\ crate that provides a compatibility layer for the cuDNN 9 API, enabling support for ROCm 7. The implementation exposes unmangled cuDNN 9 functions that delegate to the underlying \zluda\_dnn\ library, with specific Windows loader hooks to ensure correct dynamic linking of \amdhip64\_7.dll\. It also includes a comprehensive test suite that validates core operations, such as convolution backward data passes, by comparing GPU results against CPU references.
_zluda\dnn9 · high confidence
Add hipBLASLt system bindings for ROCm
This change introduces the \ext/hipblaslt-sys\ crate, providing Rust bindings for the hipBLASLt library. The \build.rs\ script configures the linker to search \/opt/rocm/lib/\ and link against \hipblaslt\ on non-Windows platforms, while \src/lib.rs\ exposes generated constants and types (such as \hipblasOperation\_t\, \hipblasComputeType\_t\, and \hipDataType\) required by higher-level components like PyTorch.
ext/hipblaslt-sys · high confidence
Add hipFFT system bindings for AMD GPU FFT operations
The ext/hipfft-sys crate now provides the low-level Rust bindings for the hipFFT library, enabling Fourier transform computations on AMD ROCm hardware. This addition includes the build configuration to link against the hipfft dynamic library on non-Windows platforms and exposes core API functions such as hipfftPlan1d, hipfftPlan2d, hipfftPlan3d, and hipfftPlanMany, along with necessary type definitions and constants for single and double precision complex transforms.
ext/hipfft-sys · high confidence
Add ptxas stub binary for CUDA compilation
Introduces a new \ptxas\ executable that mimics the command-line interface of the NVIDIA CUDA compiler (reporting version 12.8). It accepts standard arguments such as input/output files, optimization levels, and GPU architecture flags, but functions as a stub by simply copying the input file to the output location without performing actual compilation.
ptxas · high confidence
Add rocSPARSE system bindings
Added a new \ext/rocsparse-sys\ crate that provides Rust bindings for the rocSPARSE library. This includes a build script to link against the \rocsparse\ dynamic library on non-Windows platforms and generated source code defining the necessary types, handles, and descriptors for interacting with the rocSPARSE API.
ext/rocsparse-sys · high confidence
Add rocblas-sys bindings and build configuration
This change introduces the \ext/rocblas-sys\ crate, providing Rust bindings for the AMD rocBLAS library. The \build.rs\ script configures the linker to search \/opt/rocm/lib/\ and link against \rocblas\ on non-Windows platforms, while a \.rustfmt.toml\ file disables formatting for the generated code. The \src/lib.rs\ file contains the auto-generated bindings (produced by \zluda\_bindgen\), exposing rocBLAS data types (such as \rocblas\_half\, \rocblas\_float\_complex\), operation enums, and handle structures required for linear algebra operations.
ext/rocblas-sys · high confidence
Add zluda\_precompile tool for batch CUDA binary precompilation
A new command-line utility has been added to scan directories for CUDA binaries (ELF and PE formats) and extract their PTX code for precompilation. The tool supports parallel processing via Rayon, displays progress bars for file scanning and compilation, and allows users to specify the CUDA device ID and whether to follow symbolic links.
_zluda\precompile · high confidence
Added CUDA library bindings for cuBLAS, cuBLASLt, cuDNN, cuFFT, and cuSPARSE
The \cuda\_macros\ crate now includes generated FFI bindings for several core CUDA libraries, enabling their use in host-side code. This adds support for the cuBLAS linear algebra library, the lightweight cuBLASLt API, the cuFFT transform library, and the cuSPARSE sparse matrix library. It also introduces separate bindings for both cuDNN version 8 and version 9, allowing applications to target specific major versions of the deep neural network library.
_cuda\macros/src · high confidence
Added NVML function tracing capability
The zluda\_trace\_nvml library now provides a mechanism to trace calls to NVIDIA Management Library (NVML) functions. It dynamically loads the libnvidia-ml library and wraps its exported functions to log calls, allowing users to monitor NVML interactions.
_zluda\_trace\nvml · high confidence
Added Rust type bindings for CUDA libraries and infrastructure
New Rust type definitions have been added to the \cuda\_types\ crate to support CUDA performance libraries and core infrastructure. This includes bindings for cuBLAS (version 13.0.2.14), cuBLASLt, cuFFT (version 12.1.0.31), cuSPARSE (version 12.6.3.3), and two versions of cuDNN (v8.9.7 and v9.13.1). Additionally, the commit introduces the \dark\_api\ module, which defines structures for handling CUDA fat binary formats (FatbinHeader, FatbincWrapper), and updates the core \cuda.rs\ module with new types and constants for CUDA 13.0.
_cuda\types/src · high confidence
Added Windows CUDA library loading verification tool
A new diagnostic tool for Windows has been added to verify the loading and availability of CUDA libraries (such as cuBLAS, cuSPARSE, cuFFT, and cuDNN). The tool attempts to load specified DLLs, checks for specific function symbols to validate functionality, and reports whether the libraries are found and working, helping users troubleshoot CUDA environment issues on Windows.
_cuda\check · high confidence
Added generated CUDA library bindings for cuBLASLt, cuDNN v8, and cuFFT
The build directory now includes automatically generated header files for several NVIDIA libraries, enabling better support for performance libraries and PyTorch. This change adds \cublasLt\_internal.h\ (generated via a new Ghidra-based Python script), wrapper headers for cuDNN v8 (covering inference, training, backend, and operations), and updated wrappers for cuBLAS and cuFFT. These bindings expose the necessary function signatures for the host code to interface with the underlying CUDA runtime components.
_zluda\bindgen/build · high confidence
Added generated formatting implementations for CUDA libraries
The format module now includes auto-generated display and logging implementations for CUDA types and functions across cuBLAS, cuBLASLt, cuDNN (v8 and v9), and cuFFT. This enables human-readable logging of API calls and internal state for these libraries, supporting better debugging and traceability when using the translation layer.
format · high confidence
Added tracing support for cuBLAS, cuBLAS-LT, cuFFT, and cuSPARSE libraries
The ZLUDA trace component now includes dedicated modules for intercepting and logging calls to key CUDA libraries: cuBLAS, cuBLAS-LT, cuFFT, and cuSPARSE. These new modules dynamically load the respective shared libraries (libcublas.so, libcublasLt.so, libcufft.so, and libcusparse.so) and wrap their functions to log API usage, enabling users to trace these specific computational backends alongside existing CUDA operations.
(repo-wide) · high confidence
Added tracing support for cuDNN 8 and cuDNN 9
New trace modules have been added for cuDNN 8 and cuDNN 9, enabling the interception and logging of calls to these specific deep neural network libraries. The cuDNN 8 module loads \libcudnn.so.8\ (configurable via \ZLUDA\_DNN8\_LIB\), while the cuDNN 9 module loads \libcudnn.so.9\ (configurable via \ZLUDA\_DNN9\_LIB\). Both modules dynamically resolve function pointers from the respective shared libraries and route calls through the common tracing infrastructure to log execution details.
_zluda\_trace\_dnn8, zluda\_trace\dnn9 · high confidence
Initial HIP runtime bindings and build configuration
This change introduces the \hip\_runtime-sys\ crate, providing the foundational Rust bindings for the AMD HIP runtime. It includes a build script that configures linking to \amdhip64\ on Linux (searching \/opt/rocm/lib/\) and handles delayed loading of \amdhip64\_7.dll\ on Windows. The library exposes generated constants and types for HIP operations, such as memory allocation flags, stream behaviors, and JIT options, enabling Rust applications to interface with AMD GPU hardware.
_ext/hip\runtime-sys · high confidence
Initial PTX source library structure
The ptx/src/lib.rs file introduces the foundational module structure for the PTX source component, exposing the \to\_llvm\_module\ translation function, \Attributes\, and \TranslateError\ types from the \pass\ module, while also including internal \pass\ and test modules.
ptx/src · high confidence
Initial cuFFT implementation via hipFFT
The zluda\_fft module now provides a functional implementation of the CUDA cuFFT library by translating calls to the AMD hipFFT backend. This change introduces core FFT capabilities, including plan creation (1D, 2D, 3D, and many), execution (C2C, R2C, C2R, Z2Z, D2Z), and resource management (create, destroy, stream setting). The implementation handles handle mapping between CUDA and HIP contexts and includes a build script for Windows resource embedding.
_zluda\fft · high confidence
Initial cuSPARSE support via rocsparse
This change introduces the zluda\_sparse module, providing a minimal implementation of the cuSPARSE API by mapping CUDA sparse operations to the rocsparse library. It exposes functions for creating and destroying sparse matrix descriptors (COO, dense matrix), managing handles, and performing sparse matrix-matrix multiplication (SpMM) operations. On Windows, it dynamically loads the rocsparse DLL, while on other platforms it links directly to the system library. Unimplemented functions return an error status, ensuring compatibility with applications expecting cuSPARSE interfaces.
_zluda\sparse · high confidence
Initial implementation of cuBLAS support via rocBLAS
This change introduces the \zluda\_blas\ crate, providing the first layer of cuBLAS compatibility by mapping CUDA BLAS calls to the rocBLAS library on AMD ROCm platforms. It implements core functions such as \cublasCreate\_v2\, \cublasDestroy\_v2\, and various GEMM operations (e.g., \cublasSgemm\_v2\, \cublasSgemmStridedBatched\), along with handle management and status reporting. The implementation includes platform-specific logic for loading \rocblas.dll\ on Windows and \librocblas.so\ on Unix, and adds integration tests to verify the behavior of these BLAS functions against both the Zluda implementation and native CUDA libraries.
_zluda\blas · high confidence
Initial import of ROCm SMI system bindings
This change introduces the \rocm\_smi-sys\ crate, providing Rust bindings for the AMD ROCm System Management Interface (RSMI). The diff shows the addition of a build script that links against the \rocm\_smi64\ dynamic library and a generated source file (\src/lib.rs\) exposing RSMI constants, enums (such as performance levels and event groups), and initialization flags. This enables Rust applications to interact with AMD GPU hardware monitoring and management features.
_ext/rocm\smi-sys · high confidence
Initial import of the highs-sys Rust binding for HiGHS
This change introduces the \highs-sys\ crate, providing Rust bindings for the HiGHS linear programming solver. It includes a build script that can either compile HiGHS from source (with optional features for libz and Ninja) or discover a pre-installed system version via pkg-config. The crate exposes the HiGHS C API, including functions for solving linear programs (\Highs\_lpCall\), mixed-integer programs (\Highs\_mipCall\), and quadratic programs (\Highs\_qpCall\), along with comprehensive test coverage for these solver interfaces.
ext/highs-sys · high confidence
Initial release of detours-sys Rust bindings
This change introduces the \detours-sys\ crate, providing Rust bindings to Microsoft Detours for Windows. It includes a build script that compiles the underlying C++ Detours sources using Clang or MSVC, pre-generated bindings via rust-bindgen, and dual Apache-2.0/MIT licensing. The crate exposes the Detours API for function hooking on Windows, enabling users to intercept and replace system or library calls.
detours-sys · high confidence
Introduce cuBLASLt support via hipBLASLt
This change adds a new \zluda\_blaslt\ module that implements the NVIDIA cuBLASLt API by delegating calls to the AMD hipBLASLt library. Users gain access to lightweight BLAS operations (such as matrix multiplication, handle creation/destruction, and descriptor management) on AMD hardware, with the implementation handling the translation between CUDA and HIP types and loading the appropriate native library (\hipblaslt.dll\ or \libhipblaslt.dll\).
_zluda\blaslt · high confidence
Introduce dark API fatbin parsing and unified module iteration
The dark API layer now includes a new \fatbin\ module that provides a higher-level interface for parsing CUDA fatbin binaries, supporting both V1 and V2 wrapper formats. This change introduces a unified iterator (\FatbinIter\) that abstracts over the different internal structures of fatbin versions, allowing the runtime to consistently traverse submodules and files (PTX, ELF, or binary payloads) regardless of the specific fatbin version encountered. This enables more robust handling of compiled CUDA modules within the dark API.
_dark\api · high confidence
Introduce persistent SQLite-based kernel cache
The zluda\_cache module now uses a SQLite database to store compiled kernel binaries, replacing any previous in-memory or file-based caching. The cache is stored in the system's cache directory (zluda/ComputeCache/zluda2.db) and persists across sessions. It tracks module metadata (hash, compiler version, device, etc.) and automatically manages total cache size via database triggers. This enables faster subsequent runs by reusing previously compiled kernels.
_zluda\cache · high confidence
Introduce xtask build automation tool
Added a new Rust-based build automation tool (xtask) to streamline the compilation and packaging process for ZLUDA. This tool provides CLI commands to build ZLUDA and create distribution packages, handling platform-specific output paths, symlinks, and file copying for both Linux and Windows targets.
xtask · high confidence
Introduce zluda\_common crate with shared constants, error handling, and test utilities
The new zluda\_common crate centralizes shared infrastructure for the project. It provides a compute capability helper that respects the ZLUDA\_CC and ZLUDA\_SM environment variables, allowing users to override the default GPU capability (8.5). It defines a unified CudaErrorType trait and FromCuda conversion logic to standardize error handling across CUDA, HIP, cuBLAS, cuSPARSE, cuFFT, cuDNN, MIOpen, and NVML bindings. Platform-specific self-path detection is now handled via os\_win.rs and os\_unix.rs. Additionally, a test Runtime utility is included to load CUDA or HIP libraries and perform basic memory and stream operations for testing purposes.
_zluda\common · high confidence
Introduce zluda\_trace\_common library for unified tracing infrastructure
The zluda\_trace\_common crate has been added to provide shared utilities for the tracing system. It implements platform-specific logic for loading the CUDA driver (libcuda.so on Unix, nvcuda.dll on Windows) and handles DLL loading without redirection on Windows to ensure accurate trace data. The library also defines the ReprUsize trait to standardize the serialization and formatting of CUDA API status codes (such as CUresult, cublasStatus\_t, and others) into a common format for the trace output.
_zluda\_trace\common · high confidence
Introduction of 32-bit CUDA kernel metadata support
The kernel metadata module now includes support for 32-bit CUDA architectures. A new \ModuleMetadata32Bit\ struct and associated \Global32Bit\ definitions have been added to handle 32-bit specific global variables and argument sizes. The implementation introduces a dedicated ELF section named \zluda32\ for storing this metadata, allowing the system to read, write, and copy 32-bit kernel metadata alongside the existing 64-bit support.
_kernel\metadata · high confidence
Introduction of the ZLUDA offline compiler with error-tolerant PTX compilation
The compiler now includes a new offline tool (zoc) that translates PTX files into LLVM IR and compiles them for AMD GPUs. A key behavioral change is the addition of an \--ignore-errors\ flag, which allows the compilation process to proceed and produce output even when parsing errors occur, rather than failing immediately. The compiler automatically detects the target GPU architecture via the HIP runtime if not explicitly specified, defaults to \gfx1100\, and outputs intermediate LLVM IR files alongside the final compiled artifacts.
compiler · high confidence
New LLVM-based PTX code generation backend
The compiler now includes a new LLVM-based backend for generating PTX code, located in \ptx/src/pass/llvm\. This addition introduces a new code generation path that utilizes raw LLVM-C bindings to emit optimized machine code, handling module creation, function attributes, and type mappings for various scalar and vector types. This backend supports AMDGPU-specific features such as constrained floating-point operations, specific calling conventions, and address space management, providing an alternative to existing emission methods.
ptx/src/pass/llvm · high confidence
PTX parser macro implementation
The \ptx\_parser\_macros\ crate has been added, providing the procedural macro infrastructure for parsing PTX instructions. This new component introduces logic for handling opcode definitions, including support for postfix modifiers (such as vector and type suffixes like \.v2\, \.f32\, \.pred\), automatic pattern selection for disambiguating instruction variants, and generation of enum types for parser rules. This enables the parser to correctly identify and process complex PTX instruction syntaxes.
_ptx\_parser\macros · high confidence
Removals
Removal of CUDA stub library and CLI injector
The CUDA stub library (src/lib.rs) and the command-line injector (src/bin.rs) have been removed from the source tree. This eliminates the local implementation of CUDA API functions (such as cuInit and cuDeviceGetCount) that previously returned static success values, as well as the executable used to inject the nvcuda\_redirect.dll into target processes.
src · high confidence
Behavioural changes
Automated compilation of test helpers during build
The build process now automatically compiles Rust source files located in the \tests/helpers\ directory into executable binaries during the build phase. This ensures that the necessary helper tools for integration and unit tests are always up-to-date and available without requiring manual pre-compilation steps.
_zluda\inject · high confidence
Build system now compiles and links LLVM from source
The build process has been updated to automatically compile LLVM and LLD from the local source repository (ext/llvm-project) during the build phase. This change introduces a build script that configures the LLVM build with specific components (such as AMDGPU, BitWriter, and LLD) and links the resulting static libraries directly into the application, removing the dependency on pre-built external LLVM binaries.
_llvm\zluda · high confidence
Implementation of CUDA driver API via HIP runtime
The driver implementation layer has been rewritten to translate CUDA driver API calls into HIP runtime equivalents, replacing the previous execution model. This change introduces new modules for managing device contexts, memory allocations, kernel modules, and streams, all of which now delegate to the \hip\_runtime\_sys\ library to execute workloads on AMD hardware.
zluda/src/impl · high confidence
Improved handling of floating-point denormal and rounding modes
The compiler now more accurately propagates and optimizes denormal (flush-to-zero) and rounding modes across control-flow paths. This ensures that floating-point operations behave consistently with the specified modes throughout the kernel, reducing potential inaccuracies in complex kernels where modes might change between basic blocks.
_ptx/src/pass/instruction\_mode\_to\_global\mode · high confidence
Introduce LLVM-based compiler backend for GPU code generation
The \llvm\_zluda/src\ module now provides a new compiler backend that replaces previous code generation methods with an LLVM-based pipeline. This change introduces support for atomic load/store operations and 32-bit compilation targets, while also enabling optimized i8 matrix multiply-accumulate (MMA) instructions. The backend handles the full compilation flow from linking LLVM bitcode modules (including OCKL/OCML constants) to generating AMDGPU object files via LLD, ensuring improved floating-point accuracy and broader hardware support.
_llvm\zluda/src · high confidence
Introduce runtime initialization guard and platform-specific library loading
The library now prevents CUDA driver calls during shutdown by checking an initialization flag, returning a DEINITIALIZED error if operations are attempted after deinitialization. Platform-specific logic has been added to handle library loading: on Windows, it includes a DllMain entry point for cleanup and hooks to downgrade AMD HIP library loads, while on Unix it uses standard dynamic loading. Additionally, a test harness is provided to verify API compatibility between the implementation and the native CUDA driver.
zluda/src · high confidence
Introduce zluda\_trace module for CUDA API interception and logging
The zluda\_trace module has been introduced to replace the previous zluda\_dump functionality, providing a comprehensive system for intercepting CUDA API calls and logging them. This change adds a new dark\_api layer that overrides CUDA export tables to redirect function calls through a tracing mechanism, enabling detailed logging of CUDA operations. The implementation includes platform-specific OS abstraction layers for both Unix (using dlsym and dynamic assembly thunks) and Windows (using GetProcAddress and OutputDebugString), along with a state tracker that records CUDA module loading and saves opaque ELF binaries. Users will now have access to more structured tracing capabilities with improved error handling and logging infrastructure.
_zluda\trace · high confidence
PTX parser macro implementation rewritten
The \ptx\_parser\_macros\_impl\ crate has been completely rewritten to support a new, more robust PTX instruction parsing architecture. This change introduces a new macro-based system for generating instruction types, including support for generic type parameters, variant-specific display implementations, and visitor patterns (visit, visit\_mut, visit\_map) for traversing parsed PTX structures. The new parser module handles opcode definitions, patterns, and rules with greater flexibility, enabling better support for complex PTX instruction syntaxes and improving the overall maintainability of the parser.
_ptx\_parser\_macros\impl · high confidence
PTX parser rewritten with new AST and macro-driven instruction definitions
The PTX parser has been completely rewritten, introducing a new Abstract Syntax Tree (AST) in \ast.rs\ and a macro-based instruction definition system (\ptx\_parser\_macros::generate\_instruction\_type!\) that simplifies adding new instructions and supports generic argument visiting for compilation passes. The parser implementation in \lib.rs\ now uses the \winnow\ and \logos\ crates for tokenization and parsing, replacing the previous approach. This change provides a more robust foundation for supporting a wider range of PTX instructions and features, such as extended precision arithmetic and various memory operations, by making the instruction set easier to extend and maintain.
_ptx\parser · high confidence
Redesigned Windows DLL injection with configurable CUDA redirection modes
The Windows injector has been rewritten to support flexible CUDA library redirection via command-line arguments. Users can now choose between redirecting calls to ZLUDA (default), enabling ZLUDA tracing to a temporary directory, intercepting calls for NVIDIA tracing, or providing a custom set of DLL paths. The injection mechanism now uses Detours to create suspended processes, inject the specified libraries, and manage child process lifecycle, replacing the previous implementation.
_zluda\inject/src · high confidence
Refactored PTX compilation pipeline with new intermediate passes
The PTX compilation pipeline in the \ptx/src/pass\ module has been restructured into a series of distinct, modular passes to improve code clarity and maintainability. New passes introduced include \convert\_32bit\_to\_64bit\ for handling 32-bit module compatibility, \deparamize\_functions\ to normalize function parameter handling, \expand\_operands\ for operand flattening, \fix\_special\_registers\ to resolve special register usage, \hoist\_globals\ to move global variable declarations, \insert\_explicit\_load\_store\ to manage memory state, \insert\_implicit\_conversions\ to handle type conversions, and \insert\_post\_saturation\ for floating-point saturation logic. These changes refactor the internal compilation process without altering the external API or user-facing behavior.
ptx/src/pass · high confidence
Switches library interception from LD\_PRELOAD to LD\_AUDIT
The zluda\_ld component now uses the LD\_AUDIT mechanism instead of LD\_PRELOAD to intercept and redirect specific NVIDIA libraries (such as libcublas, libcudnn, and libcuda). This change, which includes a fix for segmentation faults observed in certain configurations, allows the library to hook into the dynamic linker's object loading process more robustly, ensuring that requests for these specific shared objects are redirected to the appropriate replacement paths while preventing self-redirect loops.
_zluda\ld · high confidence
Updated CUDA API version support and binding generation infrastructure
The \zluda\_bindgen\ tool now generates bindings for a broader range of CUDA versions, explicitly adding support for CUDA 13.0.0 and 12.9.x in its known version list, while also including version-specific logic for newer APIs like \cuCheckpointProcessCheckpoint\ (CUDA 12.8) and \cuArrayGetPlane\ (CUDA 11.2). This update ensures that the generated bindings correctly resolve function pointers and handle version-dependent behaviors for these newer CUDA toolkit releases.
_zluda\bindgen/src · high confidence
Updated build process to support ROCm 7 and CUDA 12.4
The build system has been updated to support ROCm 7, requiring the linking of the delayed-load library amdhip64\_7.dll on Windows, and has been upgraded to target CUDA 12.4. The build script now includes logic to verify the integrity of the nvcudart\_hybrid64.dll file to ensure it is not a Git LFS stub, and integrates git commit information into the build.
zluda · medium confidence
Windows DLL injection and CUDA library redirection logic
The zluda\_redirect component on Windows now implements a comprehensive mechanism to intercept and redirect CUDA-related DLL loads (such as nvcuda) and process creation calls. By hooking functions like LoadLibrary and CreateProcess, the library allows applications to transparently use ZLUDA's implementation instead of the native NVIDIA drivers, supporting both ASCII and UTF-16 path lookups to ensure compatibility with various application loading behaviors.
_zluda\redirect · high confidence
Windows library loading and manifest configuration
The Windows build now includes an application manifest to enable visual styles and proper isolation awareness, alongside a Rust source module that defines the specific CUDA libraries (such as nvcuda, cudnn, cublas, and cusparse) and their version-specific DLL names required for loading on Windows.
_zluda\windows · high confidence
Fixes
Build script validates PTX implementation binary integrity
The build process now includes a check to ensure that the embedded PTX implementation binaries (zluda\_ptx\_impl.bc and zluda\_ptx\_impl\_constrained.bc) are actual LLVM bitcode files rather than Git LFS stubs. This prevents build failures or runtime errors caused by missing binary assets by verifying the file magic bytes during compilation.
ptx · high confidence
Test coverage
Add PTX test suite for vector operations and CUDA runtime stubs; Added LLVM IR golden tests for PTX instruction translation; Added integration tests for CUDA injection scenarios; Added pass-level unit tests for PTX IR transformations; Added test helpers for CUDA initialization injection scenarios; Expanded PTX test coverage for atomic, synchronization, and conversion instructions.
Dependencies
Add HiGHS and LLVM project submodules
The repository now includes two new external submodules: HiGHS (commit 364c83a) and the LLVM project (commit ff4dc1f). These additions provide the underlying optimization and compilation infrastructure required by the project, enabling features such as MMA support and improved floating-point accuracy described in the associated commit history.
ext · high confidence
Project reorganized into a Cargo workspace with new tooling and dependency updates
The project has been restructured from a single crate into a Cargo workspace containing over 40 member crates, including the offline compiler (\compiler\/\zoc\), CUDA injection tools (\zluda\_inject\, \zluda\_redirect\), and various performance library bindings (BLAS, DNN, FFT, Sparse). This change introduces new dependencies such as \bpaf\ for command-line parsing, \highs\ for linear programming, and \vergen-gix\ for build-time versioning, while updating core libraries like \syn\, \bindgen\, and \windows\ to newer versions.
(dependencies) · high confidence
Housekeeping
Project initialization with licensing, submodules, and repository configuration
The repository is initialized with the Apache 2.0 and MIT license files, a README describing ZLUDA as a CUDA replacement, and a Geekbench performance graph. Repository configuration is established via .gitattributes (marking .dll and .bc files for LFS, third-party code as vendored), .gitmodules (adding submodules for a custom LLVM fork and HiGHS), .git-blame-ignore-revs, and .rustfmt.toml (enforcing Unix line endings).
(repo-wide) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 62 → 55 (-7.3)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 81 → 46 (-35.0)
- Architecture 99 → 98 (-1.6)
- Maturity 54 → 55 (+0.4)
- Readiness 64 → 62 (-2.1)
- Security 60 → 60 (+0.0)
- Domain Modelling 100 → 100 (+0.0)
Resolved (7)
- Documentation: no installation or build instructions (README.md)
- Hotspot: ptx_parser/src/ast.rs (ptx_parser/src/ast.rs)
- Hotspot: ptx_parser/src/lib.rs (ptx_parser/src/lib.rs)
- Hotspot: zluda64_server/src/main.rs (zluda64_server/src/main.rs)
- Hotspot: zluda_common/src/lib.rs (zluda_common/src/lib.rs)
- Members sharing a duplicated core (4 members, 50+ identical tokens) (zluda_bindgen/src/main.rs)
- Off-boarding risk: anonymized user #1
New (13)
- Documentation: no architecture or design documentation (README.md)
- Members sharing a duplicated core (5 members, 50+ identical tokens) (zluda_bindgen/src/main.rs)
- Off the main sequence: cuda_types
- Off the main sequence: detours-sys
- Off the main sequence: kernel_metadata
- Off the main sequence: llvm_zluda
- Off the main sequence: ptx_parser
- Off the main sequence: ptx_parser_macros_impl
- Off the main sequence: zluda_cache
- Off the main sequence: zluda_windows
- Off-boarding risk: anonymized user #1
- Projects may be oversized for their cohesion
- TooManyFunctions: zluda_fft::impl (zluda_fft/src/impl.rs)
Changes since last survey
- 1 commits — 1 feature/other, 0 fixes
By area
- ext/hipfft-sys — 1 commit
Notable commits
- change: Implement core cuFFT APIs with hipFFT (#672)
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
vosen/ZLUDA was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit ee2f25a180099fa42f36b2346732e1f2470a03ad — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.