Skip to content
CAI
Software that uses CAICheck a score

ArroyoSystems/arroyo

48.8

Weak · 30 September 2026

89.5k

lines of production code

Rust

with TypeScript

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a distributed stream processing platform that executes SQL-defined data pipelines with stateful windowing, aggregations, and joins. It manages the lifecycle of these jobs through a control plane that handles scheduling, checkpointing to object storage, and resource allocation across worker nodes. The platform supports real-time data ingestion and output via various connectors, including Kafka and cloud storage, while allowing users to extend processing logic with Python or Rust user-defined functions.

How it got here

2023 — gRPC removal and Rust 2024 migration

30 changes.

This period focused on removing the legacy gRPC API and associated client code in favor of a REST-based architecture across the console and backend services. It also involved stripping out numerous legacy Rust modules, including old SQL compilation, state management, and worker operators, while upgrading the project to the Rust 2024 edition. Concurrently, the team established initial Kubernetes deployment support via a new Helm chart and modernized the Docker build infrastructure.

2024–2026 — v2 architecture and UDF support

52 changes.

This period focused on a comprehensive architectural overhaul to version 2, introducing a new RPC protocol, Parquet-based state backend, and unified CLI. It significantly expanded functionality with support for Python and asynchronous User-Defined Functions, alongside new connectors for Delta Lake, Iceberg, and Protobuf. The work also included a major redesign of the web UI and controller infrastructure to improve observability and deployment flexibility.

Features

Add Azure and Cloudflare R2 object storage support

The storage layer now supports Azure Blob Storage and Cloudflare R2 in addition to existing backends. This is implemented by adding new regex matchers for Azure (abfs/https) and R2 (r2/https) URL schemes in the storage provider, and introducing a dedicated AWS credential provider in a new \aws.rs\ module to handle token caching and refresh for S3-compatible stores.

crates/arroyo-storage · high confidence

Add Protobuf format support with Confluent Schema Registry and length-delimited message handling

The \arroyo-formats\ crate now includes a new \proto\ module that enables deserializing Protobuf messages into JSON and converting Protobuf schemas to Arrow schemas. This implementation supports Confluent Schema Registry wire formats (automatically skipping the magic byte and schema ID header) and length-delimited Protobuf messages (reading the varint length prefix). It handles basic types, nested messages, repeated fields, enums, and maps (mapped to JSON strings), allowing users to ingest and process Protobuf data sources directly.

crates/arroyo-formats/src/proto · high confidence

Automated generation of the Arroyo API client

The arroyo-openapi crate now automatically generates a Rust API client from the OpenAPI specification during the build process. This change introduces a build script that extracts the API documentation, converts it into code using the progenitor library, and includes the resulting client implementation in the library, ensuring the client stays synchronized with the API definition without manual maintenance.

crates/arroyo-openapi · high confidence

Generated TypeScript API type definitions for the web UI

The web UI now includes an auto-generated TypeScript file (\webui/src/gen/api-types.ts\) that defines the types for the backend API endpoints. This file, created by openapi-typescript, provides strict type checking and autocompletion for API interactions such as managing connection profiles, tables, connectors, and jobs, ensuring the frontend code aligns with the current API contract.

webui/src/gen · high confidence

Initial Helm chart release for Arroyo

This change introduces the initial Helm chart for deploying Arroyo on Kubernetes, providing a structured way to install and manage the platform. The chart includes templates for the controller deployment, worker configuration via ConfigMap, RBAC roles, and service accounts, along with helper templates for consistent naming and labeling. It supports configurable PostgreSQL backends (deployed via chart or external), Prometheus metrics integration, and custom environment variables. Post-installation, users can access the web UI via the exposed HTTP service or port-forwarding, and the chart includes test hooks to verify HTTP and gRPC connectivity.

k8s/arroyo/templates · high confidence

Initial release of Arroyo RPC and API protocol definitions

This change introduces the foundational gRPC protocol definitions for the Arroyo streaming platform, establishing the contract between clients and the control plane. The new \api.proto\ defines the schema for pipeline operators, including windowing (tumbling, sliding, session), aggregations, joins (including lookup joins), and user-defined functions (UDFs) such as async UDFs and Wasm functions. The \rpc.proto\ file defines the service interfaces for worker registration, heartbeats, task lifecycle management, and checkpointing events, enabling the distributed execution of data pipelines.

crates/arroyo-rpc/proto · high confidence

Initial release of the Arroyo Helm chart for Kubernetes deployment

This change introduces the official Helm chart for deploying the Arroyo stream processing engine on Kubernetes, establishing the baseline for version 0.17.0-dev. The chart configures the Arroyo control plane and worker pods, including resource limits, service ports, and support for node selectors and tolerations. It bundles dependencies for PostgreSQL (using a legacy Bitnami image) and Prometheus, allowing users to deploy in-cluster instances or connect to external databases and metrics servers. Configuration is managed via standard Helm values, supporting S3 object storage for checkpoints and artifacts, or local volume mounts for development environments.

k8s/arroyo · high confidence

Initial support for Python UDFs

Users can now define and execute user-defined functions in Python within Arroyo. This change introduces a new Python UDF module that allows developers to annotate Python functions with a \@udf\ decorator, automatically inferring input and return types from Python type hints (such as \int\, \float\, \str\, and \bool\). The implementation uses isolated Python sub-interpreters to run these functions safely and efficiently, converting data between Arrow arrays and Python objects to integrate with the existing DataFusion query engine.

crates/arroyo-udf/arroyo-udf-python · high confidence

Introduce Parquet-based state backend with new checkpoint protocol

The state management layer has been replaced with a new Parquet-based backend that stores operator state in Parquet files and manages checkpoint metadata via a new protocol. This change introduces a \BackingStore\ trait and \ParquetBackend\ implementation to handle loading, writing, and cleaning up checkpoint metadata, as well as compacting operator data. Users benefit from a more robust and scalable state storage mechanism that integrates with the new checkpoint system, improving reliability and performance for stateful stream processing.

crates/arroyo-state/src · high confidence

Introduce centralized metrics library for task counters and queue gauges

The \arroyo-metrics\ crate now provides a centralized library for managing observability data, introducing typed counters for messages, bytes, batches, and deserialization errors per subtask, along with helper functions to register queue gauges. This change establishes the foundational metric infrastructure (counters and gauges) used by the system, replacing ad-hoc metric handling with a structured, label-based approach.

crates/arroyo-metrics · high confidence

Introduce connection table versioning

The API now supports versioned connection tables, allowing users to create and retrieve specific versions of table configurations. The \get\_connection\_tables\ and \get\_connection\_table\ endpoints return the latest version by default, while \get\_connection\_table\_version\ allows fetching a specific historical version. This change enables tracking changes to table schemas and configurations over time.

crates/arroyo-api/queries · high confidence

Introduce core type definitions and serialization traits for Arroyo

This change adds the \arroyo-types\ crate, establishing foundational types for the system. It defines identifier wrappers (\WorkerId\, \MachineId\, \PipelineId\, \JobId\) and utility functions for time conversion. Crucially, it introduces the \Key\ and \Data\ traits, which enforce \bincode\ encoding/decoding capabilities, enabling consistent serialization for state and data processing. It also defines the \Watermark\ and \ArrowMessage\ enums, which structure how data and control signals (like barriers and stop commands) are passed through the pipeline.

crates/arroyo-types · high confidence

Introduces Arrow-based operator implementations for streaming computations

The worker now executes streaming operators using the Arrow data format and DataFusion physical plans, replacing the previous implementation. This change adds new operator modules for asynchronous UDFs, incremental aggregations, various windowing strategies (tumbling, sliding, session), and join types (instant, lookup, and joins with expiration). Users benefit from improved performance and consistency as these core data processing components are now handled by the optimized Arrow execution engine.

crates/arroyo-worker/src/arrow · high confidence

Introduces a worker-side job controller for leader-mode checkpointing and metrics

The worker now includes a new \job\_controller\ module that manages job execution state, checkpoint lifecycle, and metrics collection when operating in leader mode. This change adds structured state tracking for checkpoints (including operator and table-level details, start/finish times, and event spans), implements a two-phase commit protocol for transactional connectors, and introduces a metrics subsystem that collects and aggregates task-level performance data (such as bytes/messages sent/received and backpressure) for the UI. The controller also handles checkpoint history pruning and cleanup to prevent resource starvation.

_crates/arroyo-worker/src/job\controller · high confidence

Introduces common infrastructure for async and Python UDFs

This change adds the \arroyo-udf-common\ crate, providing the foundational types and parsing logic required to support new UDF capabilities. It introduces \async\_udf.rs\ to handle asynchronous UDF execution via FFI handles and result queues, and \parse.rs\ to define \UdfDef\ structures and map Rust types (including primitives, strings, bytes, and timestamps) to Arrow data types. These components enable the system to parse, validate, and execute both async and Python-based user-defined functions within the data pipeline.

crates/arroyo-udf/arroyo-udf-common · high confidence

Introduces server-common crate with TLS, logging, and profiling infrastructure

The new \arroyo-server-common\ crate centralizes shared server infrastructure, providing configurable logging (plaintext, logfmt, or JSON with static fields like service name and pipeline ID), CPU profiling endpoints, and a structured shutdown mechanism. It also adds comprehensive TLS support, including mTLS for HTTP/gRPC/TCP services and specific TLS configuration handling for Postgres connections.

crates/arroyo-server-common · high confidence

Introduction of embedded and process-based worker schedulers

The controller now supports running worker nodes directly within the controller process (EmbeddedScheduler) or as separate child processes (ProcessScheduler), replacing the previous external scheduling mechanism. This change allows for tighter integration and simplified deployment for local or embedded use cases, while the ProcessScheduler manages worker lifecycle via spawned processes with specific environment variables and shutdown handling.

crates/arroyo-controller/src/schedulers · high confidence

New Dockerfile consolidates build and runtime images with Python and Protobuf support

The Docker build process has been restructured into a new Dockerfile that consolidates the builder and runtime stages. This change introduces support for Python 3.12 (installed via standalone binaries) and Protobuf (protoc), enabling UDF compilation and schema registry features. It also adds mold as a linker for faster builds and exposes the API on port 5115 by default. The build now supports multi-architecture (x86\_64 and aarch64) and allows feature flags (like 'python') to be passed via build arguments.

docker · high confidence

New JSON format encoder and schema conversion support

The JSON format module now includes a custom encoder factory that serializes binary fields as base64 strings, supports configurable timestamp formats (converting various time units to Unix milliseconds), and handles Decimal128 values via string or base64 encoding. It also introduces schema conversion utilities that translate Arrow schemas to JSON Schema and Kafka Connect JSON formats, enabling better interoperability with external systems.

crates/arroyo-formats/src/json · high confidence

New RPC types and modules for checkpoints, configuration, and worker identity

The \arroyo-rpc\ crate introduces several new modules that define the data structures and validation logic used by the control plane and workers. \checkpoints.rs\ adds the \CheckpointMetadataStore\ trait and request types (\CreateCheckpointReq\, \UpdateCheckpointReq\, \FinishCheckpointReq\) to support the new checkpoint protocol. \config.rs\ implements a new configuration system using \figment\, supporting TOML, YAML, and JSON sources with environment variable overrides and legacy config migration. \identity.rs\ introduces gRPC interceptors (\InjectWorkerId\, \VerifyWorkerId\) to validate worker identity on RPC calls, preventing requests from being delivered to the wrong worker. \errors.rs\ defines a structured \DataflowError\ enum with \ErrorDomain\ and \RetryHint\ for better error classification. \df.rs\ adds \ArroyoSchema\ for managing Arrow schemas with timestamp and key indices. \formats.rs\ defines data format configurations (JSON, Avro, etc.) and \public\_ids.rs\ provides ID generation and validation logic.

crates/arroyo-rpc/src · high confidence

New SQL planner architecture with DataFusion extensions

The SQL planner has been rewritten to use a new extension-based architecture built on DataFusion. This introduces a \Planner\ and \PlanToGraphVisitor\ that convert logical plans into a logical execution graph, utilizing custom extensions for core operations: \AggregateExtension\ for windowed aggregations, \JoinExtension\ for standard joins, \LookupJoin\ for connector-based lookups, \DebeziumUnrollingExtension\ for CDC data handling, and \KeyCalculationExtension\ for partitioning. The build system now uses a \build.rs\ script to track test query files for incremental compilation.

crates/arroyo-planner · high confidence

New compiler service for compiling User-Defined Functions (UDFs)

A new \arroyo-compiler-service\ has been introduced to handle the compilation of User-Defined Functions (UDFs). This service exposes a gRPC interface (\CompilerGrpc\) that accepts UDF definitions and dependencies, generates a Rust crate (using edition 2024), and compiles it into a dynamic library. It supports automatic installation of the Rust toolchain if not present, allows environment variable substitution in dependency strings, and can optionally secure its gRPC endpoint with TLS. The compiled artifacts are stored via a configurable storage provider.

crates/arroyo-compiler-service · high confidence

New filesystem sink v2, Delta Lake/Iceberg connectors, and Confluent Cloud support

The filesystem connector now defaults to a new v2 sink implementation that supports configurable file rolling policies (size, interval, inactivity, watermark expiration), multipart upload tuning, and optional partition shuffling. New connectors have been added for Delta Lake and Iceberg tables, enabling writes to these formats with specific metadata and validation logic. Additionally, a dedicated Confluent Cloud connector is introduced, wrapping the existing Kafka connector to provide seamless integration with Confluent's managed Kafka and Schema Registry services. A new 'Blackhole' no-op sink connector is also added for testing purposes.

crates/arroyo-connectors · high confidence

New home dashboard displaying job statistics

The home page now features a dashboard that displays real-time statistics for jobs, including counts for running, all, finished, and failed jobs. This new view provides users with an immediate overview of their pipeline activity directly on the landing page.

webui/src/routes/home · high confidence

New leader management and job configuration structures in the controller

The controller now includes a dedicated \LeaderManager\ component that handles RPC communication with job leaders, including status polling, stopping, and state waiting with retry logic and timeouts. Additionally, the controller's core library introduces \JobConfig\ and \PipelineInfo\ structures to support per-job environment variables, scheduler configuration overlays, and pipeline-level metadata such as state URLs and tags, enabling more granular control and observability for individual pipeline jobs.

crates/arroyo-controller/src · high confidence

New operator execution framework with batch-sized queues and connector abstractions

The \arroyo-operator\ crate now provides the core runtime for executing dataflow operators, introducing a new \Connector\ trait for defining data sources and sinks, and a \BatchSender\/\BatchReceiver\ system that limits queue depth by row count rather than batch count to improve backpressure handling. The module also includes the \ArroyoUdaf\ accumulator wrapper to support vector-argument User Defined Aggregate Functions in DataFusion, and restructures the operator lifecycle through \ConstructedOperator\ and \OperatorNode\ types to support operator chaining and unified source/processing execution.

crates/arroyo-operator · high confidence

New state table implementations with version-aware migration support

The state backend now includes new table implementations (\ExpiringTimeKeyTable\ and \GlobalKeyedTable\) that store state in Parquet format and support version-aware migration. The \GlobalKeyedTable\ specifically checks the state version stored in checkpoint metadata and automatically migrates data from the previous version if a single-step version bump is detected, ensuring compatibility during operator upgrades.

crates/arroyo-state/src/tables · high confidence

Redesigned connection creation and management interface

The web UI now features a complete overhaul of the connections experience, replacing the legacy source management view with a new, multi-step wizard for creating connections. Users can now select from a grid of available connectors, configure connection profiles (reusable connection settings), define table configurations, and specify schemas (including support for JSON, Avro, and Protobuf with Confluent Schema Registry integration) before testing and finalizing the connection. The existing connections list has been updated to display connection tables with details like connector type, format, and associated pipelines, and allows for easier deletion and viewing.

webui/src/routes/connections · high confidence

Redesigned pipeline creation and editing experience with Python UDF support

The pipeline editor UI has been redesigned to support creating and editing pipelines with Python User-Defined Functions (UDFs). Users can now write and validate Python code alongside SQL queries in a unified editor, manage local and global UDFs, and configure pipeline settings such as parallelism and environment variables. The interface includes a catalog tab for browsing sources and sinks, a resource panel for managing UDFs, and improved output visualization with real-time data streaming.

webui/src/routes/pipelines · high confidence

Redesigned pipeline creation experience with interactive tour and new checkpoint visualization

The web UI now features a guided interactive tour for new users, including a Welcome modal, a Pipeline Builder introduction, and an Example Queries drawer that allows users to copy SQL snippets directly into the editor. The pipeline creation workflow has been updated with a Start Pipeline modal that collects the pipeline name and parallelism settings. Additionally, a new Checkpoint Details component provides a visual timeline of operator and subtask execution spans, while the Checkpoints list now displays checkpoint types and supports caching with clear status alerts when data is stale or unavailable.

webui/src/components · high confidence

Support for async User-Defined Functions in DataFusion

The UDF plugin now allows users to define asynchronous functions (using the \async\ keyword) alongside synchronous ones. This change introduces the necessary runtime infrastructure to execute these async UDFs within the DataFusion pipeline, enabling non-blocking operations in user-defined logic.

crates/arroyo-udf/arroyo-udf-plugin · high confidence

Support for asynchronous User-Defined Functions (UDFs)

The UDF host now supports asynchronous UDFs, allowing users to define functions that perform non-blocking operations. This change introduces the necessary infrastructure to manage async execution lifecycles, including starting the runtime, sending input data via FFI, and draining results. The implementation handles both ordered and unordered async execution modes, enabling more complex data processing workflows within the Arroyo stream processing engine.

crates/arroyo-udf/arroyo-udf-host · high confidence

Support for raw bytes and Vec\<&str\> arguments in UDFs

The UDF macro now supports User Defined Functions that accept raw binary data (Binary type) and aggregate functions that take vector arguments of string slices (Vec\<&str\>). This allows users to process binary payloads directly and handle list-based string inputs in aggregations without manual conversion.

crates/arroyo-udf/arroyo-udf-macros · high confidence

Removals

Removal of CloudComponents and utility formatting libraries

The CloudComponents module, which previously handled gRPC-based API client creation and cloud-specific UI routing, has been removed from the console library. Additionally, utility functions for formatting durations, data rates, and byte sizes, as well as a BigInt stringifier, have been deleted. This indicates a shift away from the previous gRPC-centric architecture and cloud-specific component structure in this part of the codebase.

arroyo-console/src/lib · high confidence

Removal of arroyo-controller legacy components

The arroyo-controller module has been removed from the codebase. This change deletes the controller's build script, SQL query definitions, compiler logic, job controller, scheduler implementations, and the entire state machine (including states like Compiling, Running, Recovering, and CheckpointStopping).

arroyo-controller · high confidence

Removal of arroyo-macro proc-macro crate

The \arroyo-macro\ crate, which previously provided the \wasm\_fn\ proc-macro for defining WebAssembly functions and the \arroyo\_data\ attribute for deriving serialization traits, has been removed from the codebase. This eliminates the ability to define custom WASM UDFs and auto-derive data structures using these specific macros in this module.

arroyo-macro · high confidence

Removal of arroyo-metrics library

The arroyo-metrics library has been removed from the codebase. This deletion eliminates the previously available helper functions for creating Prometheus counters, gauges, and histograms with automatic task-based labeling, meaning any components relying on this module for metric registration must now use an alternative approach.

arroyo-metrics, arroyo-server-common · high confidence

Removal of arroyo-types crate

The \arroyo-types\ crate has been deleted from the codebase. This removal eliminates the legacy type definitions, environment variable constants, and utility functions that were previously housed in this module, indicating a structural shift in how core types and configuration are managed within the system.

arroyo-types · high confidence

Removal of built-in sink operators

The built-in implementations for File, Null, Console, and gRPC sinks have been removed from the \arroyo-worker\ module. This change eliminates the default sink operators that previously allowed users to write data to local files, discard records, log to the console, or send data to the controller via gRPC, indicating that sink functionality is now handled externally or through a different mechanism.

arroyo-worker/src/operators/sinks · high confidence

Removal of legacy Arroyo Console build artifacts and configuration

The \arroyo-console\ directory has removed its legacy build configuration and static assets, specifically deleting \index.html\, \vite.config.ts\, and the \pre\ and \post\ query definition files. This cleanup eliminates the previous Vite-based React build setup and associated sample query definitions from the console component.

arroyo-console · high confidence

Removal of legacy Kafka sink implementation

The legacy Kafka sink operator implementation located in \arroyo-worker/src/operators/sinks/kafka/mod.rs\ has been removed from the codebase. This deletion eliminates the previous \KafkaSinkFunc\ component that handled data serialization and producer creation for Kafka topics, indicating a structural change to how data is written to Kafka within the Arroyo worker.

arroyo-worker/src/operators/sinks/kafka · high confidence

Removal of legacy Kafka source operator

The legacy Kafka source implementation, including its state management, serialization logic, and associated unit tests, has been removed from the worker. This deletion eliminates the previous JSON and JSON Schema Registry parsing modes and the manual offset checkpointing mechanism previously handled by \KafkaSourceFunc\.

arroyo-worker/src/operators/sources/kafka · high confidence

Removal of legacy arroyo-datastream library

The \arroyo-datastream/src/lib.rs\ file has been deleted, removing the legacy datastream library and its associated components (such as Wasm UDFs, windowing types, and aggregators) from the codebase.

arroyo-datastream · high confidence

Removal of legacy gRPC API definitions

The \api.proto\ and \rpc.proto\ files defining the legacy gRPC API and controller communication have been deleted. This removes the previous interface for job management, worker registration, and checkpointing, indicating a migration away from this specific gRPC-based control plane.

arroyo-rpc/proto · high confidence

Removal of legacy state management modules

The \arroyo-state\ module has removed its legacy implementation files (\lib.rs\, \parquet.rs\, and \tables.rs\), which previously defined the \StateStore\, \ParquetBackend\, and time-keyed table structures. This deletion indicates that the state management logic in this location has been refactored or replaced by newer implementations, likely as part of the ongoing migration to RecordBatch and DataFusion compute.

arroyo-state · high confidence

Removal of legacy windowing and join operators

The \arroyo-worker/src/operators\ module has removed the previous implementations for windowed aggregations, top-N windows, and windowed joins. Specifically, the files \aggregating\_window.rs\, \joins.rs\, \sliding\_top\_n\_aggregating\_window.rs\, \tumbling\_aggregating\_window.rs\, \tumbling\_top\_n\_window.rs\, and \windows.rs\ have been deleted, along with their exports in \mod.rs\. This eliminates the old \KeyedWindowFunc\, \WindowedHashJoin\, and specialized window operator structs, indicating a shift away from this code path in the dataflow engine.

arroyo-worker/src/operators · high confidence

Removal of single-node Docker image build files

The Dockerfile and supervisord configuration for the single-node deployment have been removed. This eliminates the ability to build and run the Arroyo system (including the API, controller, and PostgreSQL database) as a single Docker container using the previous setup.

docker/single · high confidence

Removed legacy pipeline and job management UI components

The legacy pipeline creation and job management views have been removed from the console. Specifically, the \CreatePipeline\, \JobDetail\, \JobsIndex\, \OldCreatePipeline\, and \SqlEditor\ components in the \arroyo-console/src/routes/pipelines\ directory have been deleted. This cleanup eliminates the old gRPC-based UI implementations for creating pipelines, viewing job details, managing job lists, and editing SQL queries, likely as part of a broader migration to a new REST API and UI architecture.

arroyo-console/src/routes/pipelines · high confidence

arroyo-sql: Removal of legacy SQL compilation modules

The \arroyo-sql\ crate has removed its legacy SQL compilation implementation, specifically deleting the \expressions.rs\, \lib.rs\, \operators.rs\, \pipeline.rs\, \schemas.rs\, \test.rs\, and \types.rs\ source files. This change eliminates the previous internal representation and code generation logic for SQL queries, marking a structural shift in how SQL is processed within the system.

arroyo-sql · high confidence

Architecture

Refactored controller state machine to use JobContext and Scheduler

The controller's state management logic has been refactored to replace the legacy \Context\ and \job\_controller\ abstractions with a new \JobContext\ and \Scheduler\ interface. This change updates the state machine implementation (including states like \Stopping\, \Recovering\, and \Scheduling\) to interact with the scheduler for worker lifecycle management, such as stopping workers and handling task assignments, rather than relying on the previous job controller. This structural shift aligns the state machine with the decoupled scheduling architecture.

crates/arroyo-controller/src/states · high confidence

Behavioural changes

Build system now generates database access code with SQLite support

The build process for the API and controller crates has been updated to automatically generate SQL query bindings using the Cornucopia library. This new build script connects to a PostgreSQL database (using the DATABASE\_URL environment variable or local defaults) and an in-memory SQLite instance to generate type-safe, async-compatible code from SQL queries. The generation is configured to include SQLite support, enabling the application to operate with SQLite as a database backend in addition to PostgreSQL.

crates/arroyo-api · high confidence

Connections and Sinks UI migrated from gRPC to REST API

The Connections and Sinks management pages in the Arroyo Console have been updated to communicate with the backend via a REST API instead of the previous gRPC implementation. This change affects the user experience for listing, creating, editing, and deleting connections and sinks, as the underlying client calls in the UI components have been switched to use the new REST endpoints.

arroyo-console/src/routes/connections · high confidence

Consolidated arroyo binary with unified CLI and local execution support

The \arroyo\ crate now provides a single, unified binary that replaces the previous separate \arroyo-bin\ and \arroyo-df\ components. This new entry point exposes a comprehensive CLI with subcommands to start individual control plane services (API, Controller, Compiler, Node, Worker), manage database migrations, and run local pipeline clusters or visualize query plans directly. The change consolidates Docker and Kubernetes configurations to use this single image and integrates TLS support via Rustls for all service communications.

crates/arroyo · high confidence

Database schema migrations for API control plane

This update applies a series of database migrations to the Arroyo API control plane, introducing structural changes to support new features and improve data integrity. Key changes include merging the \pipelines\ and \pipeline\_definitions\ tables into a single entity, adding support for User-Defined Functions (UDFs) with a dedicated table and language tracking, and introducing per-job configuration overlays for scheduler and pipeline settings. The schema also adds public IDs (\pub\_id\) to most tables for stable referencing, implements cascade deletes for better data consistency, and introduces new fields for checkpoint events, error handling context, and connection table versioning.

crates/arroyo-api/migrations · high confidence

Introduce connection table versioning

The API now tracks and persists a version number for each connection table. When compiling SQL pipelines, the API resolves and stores the specific version of each connection used, ensuring that pipeline execution is tied to a consistent snapshot of the connection definitions rather than potentially drifting configurations.

crates/arroyo-api/src · high confidence

Introduce structured API types for checkpoints, connections, metrics, pipelines, and UDFs

The API surface is now defined by explicit, serializable types in \crates/arroyo-rpc/src/api\_types\. Checkpoints now include a \checkpoint\_type\ (Scheduled or Stopping) and detailed event spans for diagnostics. Connections support a richer schema with \Decimal128\, \Struct\, \List\, and \Timestamp\ field types, along with validation to prevent duplicate fields. Pipelines expose \state\_url\, \tags\, per-job \env\_vars\, \scheduler\_config\, and \pipeline\_config\ overlays, and jobs now expose the \pipeline\_id\. UDFs now support Python in addition to Rust, and metrics are exposed via structured \MetricGroup\ and \OperatorMetricGroup\ types. All JSON responses use snake\_case for consistency.

_crates/arroyo-rpc/src/api\types · high confidence

Migration from gRPC to REST API for console connections

The Arroyo console has switched its connection endpoints from the gRPC API to the REST API. This change updates the UI to use the new connections API, ensuring that connection management and related interactions now rely on REST endpoints instead of the previous gRPC-based implementation.

arroyo-console/src · high confidence

New Avro serialization and deserialization implementation

The Avro format module has been replaced with a new implementation in \crates/arroyo-formats/src/avro\. This change introduces a complete rewrite of Avro handling, including deserialization (\de.rs\) that supports Confluent Schema Registry integration and raw datum modes, serialization (\ser.rs\) that maps Arrow data types to Avro values, and schema conversion utilities (\schema.rs\) for translating between Avro and Arrow schemas. Users will now interact with this updated Avro processing logic, which handles type mappings for decimals, timestamps, and nested structures differently than the previous version.

crates/arroyo-formats/src/avro · high confidence

New default configuration and updated protobuf build process

The crate now includes a default configuration file (default.toml) that defines settings for pipeline parameters, API and controller ports, Kubernetes scheduler resources, and logging formats. Additionally, the build script has been updated to use the experimental proto3 optional feature and to compile both rpc.proto and api.proto from the proto directory, ensuring generated code includes necessary serialization attributes.

crates/arroyo-rpc · high confidence

New object-store checkpoint protocol with leader-mode garbage collection

Arroyo introduces a new checkpoint protocol for object storage, implemented in the \arroyo-state-protocol\ crate. This change adds a dedicated garbage collection mechanism (\gc.rs\) that safely cleans up old checkpoints and their data files in leader mode, preventing storage bloat and checkpoint starvation. The protocol also defines a robust state machine for checkpoint resolution and recovery, ensuring that workers can safely restore from the correct generation even during failures or restarts. This new protocol replaces previous state management logic with a more reliable, object-store-native approach for handling checkpoint lifecycle and consistency.

crates/arroyo-state-protocol · high confidence

New source batching and serialization infrastructure

The \arroyo-formats\ crate now includes a new deserialization module (\de.rs\) and serialization module (\ser.rs\) that introduce configurable, bytes-based flushing for source collectors. Users benefit from more granular control over data ingestion performance through pipeline-level configuration of batch size, linger time, and maximum buffer bytes, as well as improved structured error handling that redacts raw values from deserialization error logs to prevent sensitive data exposure.

crates/arroyo-formats/src · high confidence

Node worker execution shifts from binary files to in-process RPC with new configuration system

The node service no longer writes pipeline binaries or WASM files to disk to start workers; instead, it spawns the worker process in-process via RPC, passing configuration through a new config system and environment variables. This change replaces the old file-based execution model with a more secure and efficient in-process approach, leveraging the new configuration system for settings like config paths and directories.

crates/arroyo-node · high confidence

Operator chaining preserves edge schemas

The datastream library now includes a chaining optimizer that merges adjacent operators into single execution units to improve performance. Crucially, this optimization preserves the schema definitions on the edges between operators within the chain, ensuring that type information is not lost during the merge. This change is implemented in the \crates/arroyo-datastream\ module, which defines the logical graph structures and the optimization logic.

crates/arroyo-datastream · high confidence

Per-slot Kubernetes resource management and per-job scheduler configuration

The Kubernetes scheduler now supports allocating CPU and memory resources per task slot rather than per pod, allowing finer-grained resource control when using the \PerSlot\ resource mode. Additionally, users can now provide a per-job \scheduler\_config\ overlay via the \StartPipelineReq\ to override global Kubernetes scheduler settings, and worker environment variables can be specified directly in the pipeline request.

crates/arroyo-controller/src/schedulers/kubernetes · high confidence

Refactor controller queries to use inner joins and new schema columns

The controller's database interaction layer has been updated to use an inner join between \job\_configs\ and \job\_statuses\ when retrieving jobs, ensuring that only fully constructed jobs are loaded and preventing partial states from being processed. The query schema now includes new columns such as \env\_vars\, \scheduler\_config\, and \pipeline\_config\ from the job configurations, as well as \state\_url\ and \tags\ from the pipelines table. Additionally, a new \clean\_preview\_pipelines\ query has been added to remove finished, stopped, or failed pipelines that have exceeded their TTL.

crates/arroyo-controller/queries · high confidence

Removal of gRPC API client and types from RPC library

The \arroyo-rpc/src/lib.rs\ file has been deleted, removing the gRPC-based API client implementation, authentication interceptor, and associated control message types (such as \ControlMessage\, \CheckpointCompleted\, and \CheckpointEvent\) from the RPC library. This change eliminates the gRPC transport layer from this module, aligning with the broader migration to a REST API.

arroyo-rpc/src · high confidence

Removal of generated gRPC API client code

The generated TypeScript files for the gRPC API client (api\_connectweb.ts and api\_pb.ts) have been removed from the console source. This eliminates the client-side bindings for the gRPC service, indicating that the console no longer communicates with the backend via this specific gRPC interface.

arroyo-console/src/gen · high confidence

Removal of legacy SQL query definitions

The \api\_queries.sql\ file, which contained the raw SQL definitions for API endpoints (such as connection, source, sink, pipeline, and job management), has been deleted. This indicates that the underlying data access layer for these resources has been refactored or migrated away from this specific SQL file, likely to a new implementation or ORM structure.

arroyo-api/queries · high confidence

Removal of legacy source operator implementations

The \mod.rs\ file in the sources operator module has been deleted, removing the previous implementations for the ImpulseSource and FileSource operators. This change eliminates the legacy source style, aligning with the migration to a new source architecture.

arroyo-worker/src/operators/sources · high confidence

Removal of legacy worker engine and network manager modules

The \arroyo-worker\ crate has removed its legacy execution engine and networking infrastructure, specifically deleting \engine.rs\, \lib.rs\, \network\_manager.rs\, and \process\_fn.rs\. This eliminates the previous custom dataflow runtime, TCP-based network links, and associated context management, indicating a shift away from the internal binary serialization and manual task scheduling logic previously handled in this location.

arroyo-worker/src · high confidence

Removal of the gRPC API service in arroyo-api

The gRPC API implementation has been removed from the arroyo-api service. This change deletes the build script that generated gRPC query bindings, the main entry point that started the gRPC server, and all associated handler modules (connections, jobs, metrics, pipelines, sinks, sources, testers, and optimizations). The API service no longer exposes the gRPC interface, shifting the system to rely on the REST API for client interactions.

arroyo-api/src · high confidence

Removed legacy gRPC-based home dashboard component

The \Home.tsx\ component, which previously rendered the main dashboard using gRPC calls to fetch jobs, sources, and sinks, has been removed from the codebase. This change eliminates the direct dependency on the gRPC API client for the home route, aligning with the broader shift toward using the REST API for console data fetching.

arroyo-console/src/routes/home · high confidence

Schema enrichment with hash and operation metadata for state management

The state schema now automatically appends \\_generation\, \\_key\_hash\, and \\_operation\ columns to the underlying memory schema. This structural change enables the system to compute and store routing key hashes for efficient filtering and statistics extraction, while also tracking the specific data operation type associated with each record batch.

crates/arroyo-state/src/schemas · high confidence

Source creation UI refactored to use new connections API

The source creation workflow in the console has been updated to align with the new connections API. The previous multi-step wizard components (ConfigureSource, CreateSource, DefineSchema, TestSource) and their associated styles have been removed, indicating a shift in how source configuration and validation are handled within the application.

arroyo-console/src/routes/sources · high confidence

Web UI library codebase reorganization and new utility components

The web UI library has been restructured with the introduction of new core modules: CloudComponents.tsx provides the root rendering logic and placeholder components for cloud-specific features, while data\_fetching.ts centralizes API client creation and SWR-based hooks for fetching pipelines, jobs, metrics, and checkpoints. New utility modules include util.ts for formatting durations, data rates, and timestamps; types.ts for shared SQL options; and example\_queries.ts providing sample SQL queries for tutorials. Additionally, the RadioGroup component has been migrated from the arroyo-console package to this library with minor refactoring.

webui/src/lib · high confidence

Web UI migration to Vite and new error pages

The web UI has been migrated to use Vite as the build tool, replacing the previous setup. This includes a new \index.html\ that supports a configurable base path via the \\_\_ARROYO\_BASENAME\ variable and integrates PostHog telemetry. Additionally, dedicated error pages have been added for API unavailability and 404 not-found scenarios, and the AG Grid theme has been updated to version 31.3.2 with a custom dark theme.

webui · high confidence

Web UI restructured with new routing, theming, and UDF state management

The web UI source has been reorganized to support a more modular architecture. A new React Router setup in router.tsx now handles navigation for Home, Connections, and Pipelines, respecting a configurable base path via window.\_\_ARCYO\_BASENAME. The application shell (App.tsx) features a collapsible sidebar with navigation buttons and integrates a new tour system (tour.ts) for user onboarding. Additionally, a dedicated state management layer (udf\_state.ts) has been introduced to handle local and global User Defined Functions (UDFs) for both Python and Rust, allowing users to create, edit, and manage UDFs directly within the interface. Theming (theming.ts) has been customized for Chakra UI components like modals and tabs to support the new tour and UI density requirements.

webui/src · high confidence

Worker introduces TLS networking, Python UDF support, and program version validation

The worker now supports TLS-encrypted communication between nodes via a new \NetworkStream\ abstraction that handles both plain TCP and TLS connections, and it adds support for executing Python UDFs by loading them into the operator registry. Additionally, the worker validates incoming execution requests against a supported program version (v2) to ensure compatibility with the controller, and it includes logic to truncate oversized error messages sent over gRPC to prevent excessive network payload sizes.

crates/arroyo-worker/src · high confidence

Test coverage

Added integration tests for API endpoints; Expanded SQL test coverage for windowing, aggregations, and joins; New SQL smoke test framework with .sql-driven queries and local UDFs.

Dependencies

Upgrade to Rust 2024 edition and modernize dependencies

The project has migrated to the Rust 2024 edition across all crates, updating the \edition\ field in \Cargo.toml\ files from 2021 to 2024. This change is accompanied by a comprehensive update of dependencies, including upgrading \axum\ to 0.9, \axum-server\ to 0.7, \tonic\ to the workspace version, and \rustls\ to 0.22. The web UI (\webui\) has also been updated, moving \vite\ to 8.0.16, \react\ to 18.3.1, and \framer-motion\ to 10.18.0, while removing older protocol buffer dependencies like \@bufbuild/connect-web\ in favor of the new REST API approach.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 47 → 49 (+1.7)
  • Rubric changed (rubric-2026.09.11 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 82 → 84 (+2.2)
  • Architecture 58 → 65 (+6.7)
  • Maturity 54 → 54 (+0.2)
  • Readiness 64 → 58 (-6.9)
  • Security 43 → 50 (+7.1)
  • Event Sourcing 100 → 100 (+0.0)
  • Accessibility 40 → 39 (-0.4)
  • Performance 100 (new)

Resolved (75)

  • Change coupling: lib.rs ↔ lib.rs (crates/arroyo-controller/src/lib.rs)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [CVE redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • …and 55 more

New (52)

  • FunctionTooLong: PipelineConfigs.PipelineConfigs (webui/src/routes/pipelines/PipelineConfigs.tsx)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • High CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • Inconsistent parameter naming for identical signatures. Most connectors use options (plural), while ConfluentConnector uses opts (abbreviated).
  • Inconsistent parameter naming for identical signatures. Most connectors use options (plural), while ConfluentConnector uses opts (abbreviated). Note: ConfluentConnector does not expose table_from_options, but the inconsistency in connection_from_options is sufficient to flag the naming convention drift.
  • Medium CVE: [GHSA redacted] (webui/pnpm-lock.yaml)
  • Medium vulnerability: RUSTSEC-2026-0285 (Cargo.lock)
  • Off the main sequence: arroyo-datastream
  • Off the main sequence: arroyo-rpc
  • Off the main sequence: arroyo-storage
  • Off the main sequence: arroyo-types
  • Off the main sequence: arroyo-udf-common
  • Off the main sequence: arroyo-udf-host
  • Off the main sequence: arroyo-udf-python
  • Off-boarding risk: anonymized user #1
  • Orphaned knowledge (crates/arroyo-connectors/src/nats/source/mod.rs)
  • Outdated: arc-swap
  • …and 32 more

Changes since last survey

  • 11 commits — 8 feature/other, 3 fixes

By area

  • crates/arroyo-connectors — 2 commits
  • crates/arroyo-controller — 2 commits
  • crates/arroyo-sql-testing — 2 commits
  • crates/arroyo-udf — 2 commits
  • crates/arroyo-api — 1 commit
  • docker/Dockerfile — 1 commit
  • webui/pnpm-lock.yaml — 1 commit

Notable commits

  • fix: Fix udf macro by tagging no_mangle unsafe (#1157)
  • fix: fix(filesystem): classify recovery auth errors as user failures (#1168)
  • fix: fix(worker): drain and restore async UDF calls correctly (#1150)
  • change: Add support for pipeline-level configs (#1151)
  • change: Decouple Controller scheduling from DataFusion (#1148)
  • change: Introduce connection table versioning (#1147)
  • change: Run the Iceberg commit in the background (#1170)
  • change: Test string-numeric comparison coercion (#1169)
  • change: Update version of pnpm in docker (#1166)
  • change: Upgrade js-yaml to 4.3.2 (from 4.3.1) (#1152)
  • change: test(sql): cover null aggregate inputs (#1153)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

ArroyoSystems/arroyo was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 30 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 0630bece757a867844f0056dd4fccb1a56cd9830 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-cb25ca4feafa.