Skip to content
CAI
Software that uses CAICheck a score

risingwavelabs/risingwave

61.9

Adequate · 29 September 2026

773.4k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a distributed streaming database engine that processes real-time data through materialized views, supporting both batch and streaming SQL execution. It ingests data from diverse sources like Kafka, CDC connectors, and object storage, and sinks results to destinations including Iceberg, Elasticsearch, and Kafka. The platform features a comprehensive observability dashboard for monitoring cluster health and query plans, alongside extensive tooling for development, testing, and CI/CD integration.

How it got here

2021–2022 — repository scaffolding and connector node introduction

77 changes.

This period established the project's foundational development infrastructure, including repository scaffolding, CI/CD tooling, and a new Java-based Connector Node service. It simultaneously removed legacy Java catalog and planner implementations while introducing a standalone SQL parser and comprehensive end-to-end test suites for batch and streaming engines.

2023 — Java connector node and integration testing

77 changes.

This period focused on establishing the Java connector node infrastructure, including the connector API, embedded gRPC service, and Java bindings for Hummock storage and stream processing. It also involved significant expansion of the integration test suite, adding comprehensive coverage for diverse sinks, CDC sources, and complex streaming scenarios alongside a new synthetic data generation tool.

2024–2025 — connector expansion and testing infrastructure

89 changes.

This period focused on expanding the platform's connectivity by adding support for numerous new sources and sinks, including Iceberg, MongoDB, SQL Server, and MQTT, alongside significant enhancements to existing connectors like Kafka and PostgreSQL CDC. A major effort was dedicated to building a comprehensive end-to-end testing infrastructure, introducing extensive test suites for backward compatibility, schema evolution, and complex streaming behaviors across these new integrations. The work also included foundational architectural changes such as a new DML crate for transactional writes, a license management system for feature gating, and a Nix-based development environment.

2026 — Backwards compatibility and test expansion

11 changes.

This period focused on expanding end-to-end test coverage for backwards compatibility, particularly for streaming parallelism, watermark handling, and AsOf joins. It also introduced new AI agent skills for development workflows and enforced stricter restrictions on CDC table creation and rate limits.

Features

Add Cassandra sink connector

Users can now sink data to Apache Cassandra. This change introduces the Cassandra sink implementation, including configuration for connection details (URL, keyspace, table, datacenter, credentials), batch size, and request timeout. It supports both append-only and upsert sink types, validates the target table schema and primary keys, and maps RisingWave data types to Cassandra CQL types.

java/connector-node/risingwave-sink-cassandra · high confidence

Add Feature Store integration test demo

A new integration test environment has been added at \integration\_tests/feature-store\ to demonstrate and validate RisingWave's feature store capabilities. This includes a Docker Compose setup with Kafka, RisingWave, and a custom feature-store service, along with two demo scenarios: an NYC taxi fare prediction case and a multi-factor authentication (MFA) change tracking case. The entry provides the necessary SQL scripts, Python generators, and shell scripts to simulate data ingestion, feature aggregation, and model inference within the test environment.

(repo-wide) · high confidence

Add Feature Store integration test simulator

A new Rust-based simulator has been added to the integration tests to exercise the Feature Store. It connects to the server at localhost:2666 and supports two modes: 'taxi' (default), which processes offline and online taxi trip data to report actions and retrieve predicted fare amounts, and 'mfa', which simulates user activity to report authentication events and retrieve feature counts. The simulator uses gRPC (tonic) for communication and reads input data from JSON or CSV sources.

_integration\_tests/feature-store/server, integration\tests/feature-store/simulator · high confidence

Add Java binding for Hummock storage reads and demo data generation

The \src/java\_binding\ directory now includes the core implementation for reading data from the Hummock storage engine via Java, featuring a new \HummockIterator\ native method that decodes a \ReadPlan\ protobuf to iterate over table data across virtual nodes. To support testing and demonstration, the change adds scripts (\run\_demo.sh\, \make-java-binding.toml\) to build the Java binding, start a demo cluster, and ingest data, alongside a Python script (\gen-demo-insert-data.py\) for generating sample SQL insert statements and a Rust binary (\data-chunk-payload-generator.rs\) for producing binary stream chunk payloads.

_src/java\binding · high confidence

Add Presto/Trino integration demo for querying RisingWave

Added a new integration test and demo setup in \integration\_tests/presto-trino\ that allows users to query RisingWave data using the PostgreSQL connectors in Presto and Trino. The change introduces a \docker-compose.yml\ configuration to spin up a RisingWave standalone cluster alongside Trino (v418) and Presto (v0.284) instances, along with necessary catalog properties (\risingwave.properties\) and client entrypoints. It also includes SQL scripts (\create\_source.sql\) to populate test data with various data types and a Python script (\sink\_check.py\) to verify data consistency between RisingWave and the query engines.

_integration\tests/presto-trino · high confidence

Add Prometheus integration test with Kafka source and materialized view

Added a new integration test suite for Prometheus monitoring, including SQL scripts to create a Kafka source (\prometheus\ topic), a materialized view (\metric\_avg\_30s\) that aggregates metrics into 30-second tumbling windows, and a read-only user (\grafanareader\) for Grafana access. The test also provisions a Docker Compose environment with Prometheus, a Kafka adapter, and RisingWave, along with a Prometheus configuration file to scrape metrics from RisingWave components and write them to Kafka.

_integration\tests/prometheus · high confidence

Add Pulsar source and materialized view for Twitter demo

The integration test environment for the Twitter-Pulsar demo now includes SQL definitions to set up the data pipeline. A new source named 'twitter' is created to consume messages from a Pulsar topic using Protobuf schema ('twitter.schema.Event') with a configurable subscription name prefix. Additionally, a materialized view 'hot\_hashtags' is introduced to calculate the top 10 hashtags by occurrence within daily time windows, enabling users to test real-time aggregation logic against the demo data.

_integration\tests/twitter-pulsar/pb · high confidence

Add SLT coverage check tool and ignore configuration

Introduces a new \check\_slt\_coverage.py\ script and a \.coverageignore\ file to the \e2e\_test\ directory. The script allows developers to verify that all SQL Logic Test (SLT) files are referenced by CI scripts, helping to identify orphaned or untested test cases. The ignore file specifies patterns for SLT files that should be excluded from this coverage analysis, such as those manually executed in backward-compatibility tests, Iceberg tests, or template files.

_e2e\test · high confidence

Add Twitter integration test scenario

Added a new integration test suite for the Twitter demo, including Docker Compose configuration to spin up the necessary infrastructure (RisingWave, Kafka, Postgres, etc.), SQL scripts to define a Kafka source and a materialized view for top hashtags, and schema definitions (Avro, Protocol Buffers) to validate data structures.

_integration\tests/twitter · high confidence

Add Twitter-Pulsar integration test with Protobuf support

The integration test suite now includes a new Twitter-Pulsar demo scenario. This addition introduces a Docker Compose configuration that spins up a Pulsar message queue alongside the standard RisingWave, Postgres, and monitoring services, along with a data generator to simulate Twitter events. The test validates end-to-end functionality by defining a Pulsar source, creating materialized views to identify influencers and trending hashtags, and verifying results via SQL queries. It also incorporates Protobuf schema definitions and testing to ensure compatibility with Protobuf-encoded data streams.

_integration\tests/twitter-pulsar · high confidence

Add Vector integration test for log ingestion

Added a new integration test in the \integration\_tests/vector\ directory that demonstrates ingesting logs from Vector into RisingWave. The setup uses a Docker Compose configuration to run a RisingWave standalone instance alongside a Vector agent configured to generate fake JSON logs via the \demo\_logs\ source and sink them to the \event\ table. The test validates the end-to-end flow by creating the source table, materialized view, and verifying the data presence.

_integration\tests/vector · high confidence

Added Delta Lake sink integration demo

A new demo in the integration\_tests/deltalake-sink directory allows users to sink data to a Delta Lake table stored on MinIO. The demo includes SQL scripts to create an append-only source and sinks, along with a README explaining how to launch the cluster via Docker Compose, create the Delta Lake table, and query the results.

_integration\tests/deltalake-sink · high confidence

Added Java binding integration demos for Hummock and StreamChunk

The java-binding-integration-test module now includes new demo applications: HummockReadDemo, which reads data from Hummock storage using metadata and vnode configuration, and StreamChunkDemo, which processes stream chunk payloads. A shared Utils class validates row data, including support for timestamp and decimal types, ensuring the Java bindings correctly handle these data formats during integration testing.

java/java-binding-integration-test · high confidence

Added Snowflake sink integration test and demo examples

Added a new integration test directory for the Snowflake sink connector, including a README tutorial and SQL scripts to demonstrate setting up the required Snowflake objects (table, stage, pipe) and RisingWave sources, materialized views, and sinks. The examples cover both append-only and upsert modes, providing users with concrete templates for configuring S3 credentials and matching column names between RisingWave and Snowflake.

_integration\tests/snowflake-sink · high confidence

Added mock data fetching script for dashboard development

A new shell script (fetch.sh) and a .gitignore file have been added to the dashboard/mock directory. The script automates the retrieval of mock data from a local development server (localhost:5691) by curling various API endpoints (actors, clusters, fragments, tables, sinks, sources, metrics, etc.) and saving the responses as JSON files, facilitating local dashboard testing and development.

dashboard/mock · high confidence

Automated TypeScript generation for dashboard protobuf definitions

A new shell script, generate\_proto.sh, has been added to the dashboard/scripts directory to automate the generation of TypeScript types from Protocol Buffer definitions. This script copies .proto files to a temporary directory, applies platform-specific sed commands to rename conflicting JavaScript keywords (Array, Object) to avoid type collisions, and then invokes protoc with ts-proto to generate TypeScript code into the proto/gen directory. This ensures that the dashboard's TypeScript interfaces stay synchronized with the underlying protobuf schema.

dashboard/scripts · high confidence

Expanded PostgreSQL system catalog and information\_schema support for tool compatibility

This release adds extensive support for PostgreSQL system catalogs and the information\_schema to improve compatibility with third-party database tools. New system tables and views are now available, including pg\_settings, pg\_am, pg\_collation, pg\_constraint, pg\_conversion, pg\_database, pg\_description, pg\_enum, pg\_index, pg\_indexes, pg\_locks, pg\_namespace, pg\_opclass, pg\_operator, pg\_postmaster\_start\_time, pg\_roles, pg\_sequence, pg\_shadow, pg\_size functions (pg\_table\_size, pg\_relation\_size, pg\_indexes\_size), pg\_stat\_activity, pg\_tables, pg\_type, pg\_user, and rw\_depend. Additionally, the information\_schema now supports views, tables, table\_constraints, schemata, key\_column\_usage, and columns. These additions enable better integration with tools like DBeaver, Metabase, and Atlas, allowing them to correctly query metadata, schema definitions, privileges, and dependencies.

_e2e\test/batch/catalog · high confidence

Expanded Postgres CDC type support and schema change reliability

The Postgres CDC connector now supports additional data types including PostGIS geography and geometry (mapped to bytea), pgvector vectors, and PostgreSQL point types (mapped to struct\<x double precision, y double precision\>). It also handles arrays of user-defined composite types and enums by deriving them as varchar arrays during auto schema changes, and correctly processes quoted table names. Additionally, the default time precision mode for CDC sources is now microseconds, and the system ensures that auto schema changes on tables with complex user-defined array types do not break unrelated column additions.

_e2e\_test/source\inline/cdc/postgres · high confidence

Expanded batch SQL function coverage with new tests

The batch execution engine now supports a significantly wider range of SQL functions, validated by new end-to-end tests in the \e2e\_test/batch/functions\ directory. Users can now rely on built-in support for array operations (including \array\_sort\, \array\_transform\, \array\_join\, \array\_max\, \array\_min\, \array\_sum\, \array\_reverse\, and overlap predicates), advanced date/time handling (\date\_trunc\, \date\_bin\, \AT TIME ZONE\, \now()\), mathematical functions (\gamma\, \lgamma\, \pow\, \pi\, \round\ with ties-to-even), and utility functions (\format\, \map\_filter\, \crc32\, \crc32c\, \overlay\, \convert\_from\, \convert\_to\). Additionally, scalar functions are now supported in the FROM clause, and session-timezone awareness has been applied to batch plan generation.

_e2e\test/batch/functions · high confidence

Iceberg sink now supports AWS IAM role assumption for Glue and S3 catalogs

The Iceberg sink connector now supports authenticating to AWS Glue catalogs and S3 storage using IAM role assumption. Users can configure \glue.iam-role-arn\ and \s3.iam-role-arn\ to allow the connector to assume a specific IAM role, which is particularly useful for cross-account access or when using temporary credentials. The implementation introduces dedicated credential providers (\GlueCredentialProvider\ and \S3FileIOAssumeRoleAwsClientFactory\) that use refreshable STS credentials and ensure that each catalog instance manages its own credentials provider lifecycle, avoiding issues with shared singleton providers. Additionally, the connector includes a JNI-based catalog wrapper (\JniCatalogWrapper\) to support various Iceberg catalog implementations (like JDBC) through the Java Native Interface, along with necessary dynamic class loading utilities (\DynClasses\, \DynConstructors\, etc.) to handle class resolution correctly in JNI contexts.

java/connector-node/risingwave-sink-iceberg · high confidence

Initial repository scaffolding and developer tooling configuration

The repository has been initialized with essential configuration files to support development workflows. This includes a pre-commit hook configuration (\.pre-commit-config.yaml\) that enforces code formatting with Rust edition 2024, runs \cargo sort\ on \Cargo.toml\, and checks for typos using \typos-cli\. A typo dictionary (\.typos.toml\) is provided to handle valid technical terms and exclude specific test data directories. Additionally, a \CODEOWNERS\ file establishes ownership for \Cargo.lock\ and license source files, while \AGENTS.md\ (linked via \CLAUDE.md\) provides structured instructions for coding agents and developers on building, testing, and contributing to the project.

(repo-wide) · high confidence

Introduce Python-based Grafana dashboard generation with multi-cluster support

The Grafana dashboard definition has been migrated from static JSON files to a Python-based generation pipeline using \grafanalib\. This change introduces a \generate.sh\ script to build the dev and user dashboards, allowing for dynamic configuration via environment variables. Users can now enable multi-cluster deployment features, such as namespace and RisingWave instance filtering, and support for dynamic Prometheus data sources, directly through the dashboard generation process.

grafana · high confidence

Introduce RisingWave SQL parser as a fork of sqlparser-rs

The \src/sqlparser\ directory now contains a standalone SQL parser implementation, forked from \sqlparser-rs\. This change adds the parser's source code, including AST definitions for data types (such as \VECTOR\ and \JSONB\), DDL operations (like \ALTER ... SET RESOURCE\_GROUP\ and \ALTER ... SET PARALLELISM\), and query structures, along with supporting files like an Apache 2.0 license, a README, a fuzzing target, and a benchmark suite. This establishes the foundational parsing logic for RisingWave's SQL dialect.

src/sqlparser · high confidence

Introduce dedicated DML crate with transactional write handles and back-pressured channels

The \src/dml\ crate has been introduced to centralize data manipulation logic, replacing the previous \source\ crate. It provides a \DmlManager\ to coordinate table readers and a \TableDmlHandle\ that exposes \WriteHandle\ for executing SQL insert, delete, and update statements. These writes are transmitted via a new \txn\_channel\ implementation that uses semaphore-based permits to enforce back-pressure and prevent out-of-memory errors during high-throughput inserts. The crate also defines specific error types (\DmlError\) for schema changes and missing readers, ensuring robust transactional behavior.

src/dml · high confidence

Introduce embedded connector node service

This change introduces the \risingwave-connector-service\, a new Java-based gRPC service that acts as an embedded connector node. It provides the implementation for sink validation, sink writer streaming, sink coordinator management, and source validation via JNI bindings. The service includes a main entry point to start the gRPC server and Prometheus metrics endpoint, along with handlers for file, JDBC, Elasticsearch, OpenSearch, and Cassandra sinks, enabling the system to offload connector execution and validation to a dedicated Java process.

java/connector-node/risingwave-connector-service · high confidence

Introduce jni\_core crate for embedded JVM and Java bindings

The new \jni\_core\ crate centralizes the embedded JVM runtime, Java binding macros, and schema history storage. It now manages JVM initialization (including heap sizing and classpath loading), registers native methods for Java interaction, and bridges Java logging to Rust's tracing system. Additionally, it implements schema history persistence using the object store with configurable timeouts and security checks for file extensions.

_src/jni\core · high confidence

Introduce license manager and feature tiers

The system now enforces feature availability based on license tiers. A new license manager validates JWT-based license keys, supporting Free, Paid (AllAsOf2\_5), All, and Custom tiers. Enterprise features such as TimeTravel, GlueSchemaRegistry, SnowflakeSink, RedisSinkStream, DynamoDbSink, OpenSearchSink, BigQuerySink, ClickHouseSharedEngine, SecretManagement, SqlServerSink, SqlServerCdcSource, CdcAutoSchemaChange, IcebergSinkWithGlue, ElasticDiskCache, ResourceGroup, DatabaseFailureIsolation, IcebergCompaction, SinkAutoSchemaChange, StateTableMemoryPreload, LocalityBackfill, SinkSinceTimestamp, and WebSocketIngest are now gated by license checks. The default production license enables all features on clusters with up to 4 CPU cores and 16 GiB memory; exceeding these limits invalidates the license. License keys are redacted in logs and diagnostics, and telemetry events are reported for license checks.

src/license · high confidence

Introduce risectl command modules for cluster diagnostics and storage management

The risectl CLI now exposes a structured set of subcommands for inspecting and managing the cluster's state. Users can dump and analyze async stack traces (await trees) to detect bottlenecks, run benchmarks against state tables, and view compute node configurations. Storage operations are now accessible via dedicated modules, allowing users to list and filter Hummock versions, inspect key-value pairs at specific epochs, dump SSTable contents with schema-aware formatting, manage compaction groups (including splitting and updating mutable configs), and perform maintenance tasks like migrating legacy object prefixes, resizing caches, and triggering full garbage collection or manual compaction.

src/ctl · high confidence

Introduce single-node and standalone modes in the all-in-one binary

The \risingwave\ binary now supports two new deployment modes: \single-node\ (the new default) and \standalone\. The single-node mode simplifies local development and testing by automatically configuring in-memory or file-based storage backends and hiding low-level node options, while the standalone mode allows users to start multiple services (meta, frontend, compute, compactor) within a single process with explicit, node-level configuration. This change also adds a build script to embed the Git SHA into the binary version string and includes demo scripts for running full standalone clusters with peripherals like Kafka and MinIO.

_src/cmd\all · high confidence

Introduce strong-typed IDs and prost helper macros

The \src/prost\ crate now uses strong-typed ID wrappers (e.g., \TableId\, \SourceId\) instead of raw integers, which prevents accidental mixing of different ID types at compile time. Additionally, the new \AnyPB\ derive macro automatically generates safe getter methods for protobuf messages, simplifying optional field handling and enum casting, while \Pb\-prefixed type aliases are generated for all message types.

src/prost · high confidence

Introduction of the Connector Node service and Java bindings

This change introduces the Connector Node, a new Java-based service that acts as a bridge for external sinks (JDBC, Iceberg, Delta Lake) and sources (CDC), replacing previous RPC mechanisms with JNI. It includes the \com\_risingwave\_java\_binding\_Binding.h\ header for Java-to-native interactions, S3 configuration utilities, and a Python-based integration test client. The build system has migrated from Gradle to Maven, and the project now enforces code formatting via Spotless and Checkstyle.

java · high confidence

Java binding now supports StreamChunk iteration and CDC source messaging

The Java binding API has been extended to allow iterating over StreamChunk objects directly via the new StreamChunk and StreamChunkIterator classes, enabling users to process data in Java without intermediate serialization. Additionally, a new CdcSourceChannel class has been added to allow sending CDC source messages and errors back to the connector, and the BaseRow class now exposes getters for a wider range of data types including timestamp, timestamptz, time, date, decimal, interval, jsonb, bytea, and one-dimensional arrays.

java/java-binding · high confidence

New AWS Docker build configuration and script

Added a new Dockerfile and build script for AWS deployments. The Dockerfile now uses Ubuntu 24.04 as the base image and bundles ca-certificates. The aws-build.sh script facilitates building the RisingWave binary with static linking features and pushing the resulting Docker image to a registry.

docker/aws · high confidence

New CI build environment image and tooling

The CI infrastructure now uses a dedicated Docker image (\ci/Dockerfile\) based on Ubuntu 24.04 to standardize the build environment. This image installs Java 21, Python 3.12, Node.js 20.11.1, and the Rust toolchain, along with essential build dependencies like OpenSSL, CMake, and Protobuf. It also pre-installs key development and testing tools, including \sqllogictest-bin\ v0.29.1, \cargo-nextest\, \sccache\, and \risedev\, ensuring consistent and faster builds across the CI pipeline.

ci · high confidence

New Elasticsearch 7 and OpenSearch Java connector implementation

This change introduces a new Java-based sink connector for Elasticsearch 7 and OpenSearch, located in the \risingwave-sink-es-7\ module. It provides a dedicated implementation using the REST High Level Client to handle bulk writes, supporting features such as configurable batch sizes, concurrent requests, retry-on-conflict policies, and dynamic index routing. The connector includes factory validation to verify connectivity and schema compatibility, along with unit tests using Testcontainers to ensure correct data ingestion for both Elasticsearch and OpenSearch backends.

java/connector-node/risingwave-sink-es-7 · high confidence

New Iceberg sink demo with Airflow-based compaction

Added a new integration test demo in \integration\_tests/iceberg-sink\ that showcases sinking data from RisingWave to an Apache Iceberg data lake. The demo includes a Docker Compose setup with RisingWave, MySQL CDC, Spark, Presto, and Airflow. It provides SQL scripts to create sources, materialized views, and upsert sinks to Iceberg, along with Airflow DAGs to automate Iceberg compaction tasks (rewriting small files and removing orphan files) to maintain query performance.

_integration\tests/iceberg-sink · high confidence

New Java common utilities for metadata, serialization, and vnode management

This update introduces a new \java/common-utils\ module containing core Java utilities for RisingWave clients. It adds \MetaClient\ to handle gRPC communication with the metadata service (including worker registration, heartbeat loops, and version pinning), \ObjectSerde\ for Java object serialization, \UrlParser\ for validating storage location URIs, and \VnodeHelper\ for distributing vnode IDs across groups. A test suite is also included to verify the vnode splitting logic.

java/common-utils · high confidence

New Nix-based development environment and SQL benchmarking framework

Developers can now use a Nix flake in \develop/nix\ to automatically provision a consistent development shell with all necessary dependencies (Rust toolchain, Java, PostgreSQL, etc.) across Linux and macOS. Additionally, a new SQL benchmarking framework (\develop/sql\_bench\) allows developers to define and run lightweight performance benchmarks using YAML configurations and \hyperfine\, facilitating rapid iteration on SQL features like \GAP\_FILL\.

develop · high confidence

New User Dashboard with modular monitoring sections

The user-facing Grafana dashboard has been restructured into a modular Python package under \grafana/dashboard/user\, introducing distinct sections for Overview, CPU, Memory, Network, Storage, Streaming, Actor Info, and Batch metrics. This change replaces the previous single-file definition with a component-based architecture where each module registers its panels via a decorator, allowing for clearer organization and easier maintenance of the monitoring views for RisingWave users.

grafana/dashboard/user · high confidence

New agent skills for PR labeling, CI triage, Rust analysis, and SLT authoring

The .agents directory now includes four new skills that guide AI agents in specific development workflows. The PR labeling skill provides a decision framework for assigning CI coverage and documentation labels based on PR content and repository rules. The fix-buildkite-ci skill enables agents to programmatically diagnose failing Buildkite checks by extracting logs and artifacts to apply focused fixes. The risingwave-rust-analyzer skill offers guidance on using the rust-analyzer CLI for semantic analysis, including handling workspace feature flags. Finally, the risingwave-slt-authoring skill defines conventions for writing and validating SQLLogicTest files, covering test structure, system command rules, and CI integration.

.agents · high confidence

New built-in system commands and validation scripts for e2e tests

The \e2e\_test/commands\ directory now provides a set of 'built-in' system commands available to sqllogictests, including wrappers for ClickHouse, MongoDB, MySQL, Redis, and SparkSQL, as well as utilities for Pulsar, Schema Registry, and search sinks. It also introduces Python-based validation scripts (\psql\_validate.py\, \psql\_refill\_validate.py\) for checking query outputs and table-refill state, alongside a script to prepare Delta Lake tables on MinIO. These additions enable more robust and localized end-to-end testing for connectors and internal state behaviors.

_e2e\test/commands · high confidence

New connector API interfaces and supporting classes

The connector-api module now includes a comprehensive set of new interfaces and classes that define the contract for connector implementations. This includes the sink API (SinkWriter, SinkFactory, SinkCoordinator, SinkRow, and related utilities like ArraySinkRow and CommonSinkConfig) and the source API (CdcEngine, CdcEngineRunner, SourceHandler, and SourceTypeE). It also introduces core data structures like TableSchema and ColumnDesc, along with helper classes such as TestUtils, Monitor, and PkComparator. These additions establish the foundational types and behaviors required for the connector node to manage data ingestion and export.

java/connector-node/connector-api · high confidence

New dashboard API client and streaming metrics support

The dashboard now includes a dedicated API layer in \dashboard/lib/api\ that centralizes HTTP requests, supports configurable endpoints (production, mock, and external meta-node), and provides a React hook (\useFetch\) for data fetching with optional polling. New API modules expose cluster metrics and version info, while the streaming module adds support for querying streaming jobs, relation dependencies, and fragment-to-relation mappings. Additionally, a new streaming stats module introduces Prometheus-based metrics collection with backpressure and throughput data, including a fallback to embedded dashboard stats if Prometheus is unavailable. Utility functions for time parsing and formatting are also added to support time-based metric queries.

dashboard/lib/api · high confidence

New dashboard utility components for backpressure visualization and iconography

The dashboard now includes new utility files in the components/utils directory to support UI enhancements. backPressure.tsx provides functions to calculate edge colors and widths based on backpressure values (normalized to 0-1) and to convert epochs to Unix milliseconds, while also offering latency-to-color mapping for performance monitoring. Additionally, icons.tsx and stroke-icons.tsx introduce a suite of SVG icons (such as server, database, and workflow icons) using a consistent 24px stroke design system, enabling more expressive visualizations in the relation dependency graph and other dashboard views.

dashboard/components/utils · high confidence

New dashboard utility libraries for layout, algorithms, and design tokens

The dashboard now includes a set of new core libraries in \dashboard/lib\ to support advanced visualization and UI consistency. \algo.ts\ provides graph traversal and connected-component detection utilities, while \layout.ts\ defines the type system for fragment and relation boxes used in the stream plan overview. \design-tokens.ts\ centralizes the visual style guide (colors, fonts, shadows, and motion) to ensure a consistent look across the dashboard. Additionally, \util.js\ offers basic array helpers, and \extractInfo.ts\ standardizes column information extraction from plan data.

dashboard/lib · high confidence

New developer utility scripts for installation, code quality, and debugging

This change introduces several new scripts to the \scripts/\ directory to improve the developer experience. The \install/install-risingwave.sh\ script now automates the installation of RisingWave, supporting both Homebrew on macOS and direct binary downloads for Linux (x86\_64 and aarch64), while also warning users if Java is missing for connector support. A new \check/check-trailing-spaces.sh\ tool allows developers to detect and automatically fix trailing whitespace in the codebase. Additionally, \coredump/coredump.entitlements\ adds macOS-specific entitlements to facilitate core dump generation during development panics, and \source/README.md\ documents the data preparation scripts used for testing Kafka sources.

scripts · high confidence

New error handling utilities and gRPC error serialization in the error crate

The \src/error\ crate now provides a comprehensive set of error-handling utilities. It introduces \def\_anyhow\_newtype\ and \def\_anyhow\_variant\ macros to create type-safe newtype wrappers around \anyhow::Error\, enforcing explicit context attachment during conversion. A complete \PostgresErrorCode\ enum is added for type-safe handling of SQLSTATE codes. For gRPC communication, the crate now serializes the full error source chain and includes the service name and call path in metadata, ensuring clients receive detailed error context. Additionally, it supports forwarding PostgreSQL error codes and error scores through gRPC boundaries to aid in root-cause analysis, and wraps \iceberg::Error\ to expose backtraces.

src/error · high confidence

New integration test datagen tool with multi-mode data generation and sink support

A new \datagen\ tool has been added to \integration\_tests/datagen\ to generate synthetic load for integration tests. It supports multiple data generation modes including ad-click, ad-CTR, CDN metrics, clickstream, ecommerce, delivery, livestream, Nexmark, and compatible data types. The tool can output generated events to various sinks such as PostgreSQL, MySQL, Kafka, Pulsar, Kinesis, S3, and NATS, with configurable QPS throttling, topic filtering, and record formatting (JSON, Protobuf).

_integration\tests/datagen · high confidence

New lint warns against direct error formatting

A new \rw::FORMAT\_ERROR\ lint has been added to the project's custom lint suite (powered by \cargo dylint\). It detects when errors are formatted directly using \format!\, \println!\, \tracing::event\, or \anyhow::anyhow\, which can cause loss of error context. The lint suggests using \thiserror\_ext::AsReport\ (e.g., \err.as\_report()\) to preserve the full error chain.

lints · high confidence

New sink benchmarking tool

A new \sink\_bench\ tool has been added to the \src/bench\ directory to evaluate sink performance. The tool includes a main implementation that simulates data generation and throughput measurement, along with configuration files defining a test schema and connection options for various sinks including Clickhouse, Redis, Kafka, Pulsar, Iceberg, MySQL, PostgreSQL, Delta Lake, Elasticsearch, Cassandra, Doris, StarRocks, and BigQuery.

src/bench · high confidence

New useErrorToast hook for standardized error notifications

The dashboard now includes a new \useErrorToast\ hook that provides a consistent way to display error messages to users. This hook leverages Chakra UI's toast component to show alerts with a 5-second duration and a close button. It intelligently formats error details, using the error message as the title and the cause as the description when available, ensuring users receive clear and actionable feedback when issues occur.

dashboard/hook · high confidence

New utility crates for async streams, iterators, and pgwire infrastructure

This change introduces several new utility crates under src/utils to support core system capabilities. The futures\_util crate adds stream control primitives, including a pausable stream with a Valve handle and a buffered stream that respects 'fence' futures to prevent out-of-order processing. The iter\_util crate provides optimized iterator adapters like zip\_eq\_fast and zip\_eq\_debug to improve performance in hot paths. The delta\_btree\_map crate introduces a new data structure that wraps a snapshot and delta BTreeMap, providing efficient cursor-based iteration over updated versions. Additionally, the pgwire crate is established as a dedicated module for PostgreSQL wire protocol handling, containing error definitions, LDAP authentication logic, and memory management for frontend messages.

src/utils · high confidence

PostgreSQL CDC now supports vector and composite types via custom converters

The PostgreSQL CDC connector now handles \pgvector\ and user-defined composite types by introducing \PgVectorToStringConverter\ and \PgCompositeToStringConverter\. These converters map vector columns and composite arrays to plain strings, allowing the connector to correctly process TOAST unchanged-value placeholders and avoid streaming crashes. This change enables reliable replication of tables containing vector embeddings and complex composite data structures.

java/connector-node/risingwave-source-cdc · high confidence

Redis sink integration test demo added

A new integration test demo for the Redis sink connector has been added to the \integration\_tests/redis-sink\ directory. This includes a Docker Compose setup to launch a RisingWave cluster with a Redis instance and a data generator, along with SQL scripts to create a source, a materialized view, and sinks using both JSON and Template encoding formats, allowing users to verify data sinking to Redis.

_integration\tests/redis-sink · high confidence

RiseDev tooling and configuration infrastructure restructured

The RiseDev developer tooling has been reorganized to improve configuration management and component handling. A new interactive configuration wizard (\risedev configure\) allows developers to enable or disable specific components (such as MinIO, Prometheus, Grafana, Lakekeeper, and ADBC drivers) via a CLI interface, replacing the previous manual environment file setup. The tool now uses a structured YAML configuration format (\risedev.yml\) with JSON schemas for validation, supporting template-based component definitions and profile expansion. Additionally, RiseDev now includes built-in support for downloading and managing external dependencies like the ADBC Snowflake driver, Lakekeeper, and Maven, with version-specific download scripts and task definitions integrated into the cargo-make workflow.

src/risedevtool · high confidence

Unify command entry points and add ctl binary

The \src/cmd\ crate now provides a unified entry point for all RisingWave components (compute, meta, frontend, compactor, and ctl) via a shared \main!\ macro and explicit entry functions. This change introduces a new \ctl\ binary (\src/cmd/src/bin/ctl.rs\) that delegates to the \risingwave\_ctl\ crate, enabling the \ctl\ command-line tool to be built and run as a standalone component alongside the other nodes. The entry functions handle common initialization tasks such as logging setup and graceful shutdown coordination.

src/cmd · high confidence

Removals

Removal of custom Calcite type system components

The custom \RisingWaveDataType\ interface, \RisingWaveDataTypeSystem\, and \RisingWaveTypeFactory\ classes have been removed from the common module. This eliminates the project's specific extensions to Apache Calcite's type system, likely simplifying type handling or preparing for a different integration approach.

java/common · high confidence

Removal of legacy Java catalog implementation

The legacy Java-based catalog service has been removed from the \java/catalog\ module. This change deletes the entire \com.risingwave.catalog\ package, including the \CatalogService\ interface, \SimpleCatalogService\ implementation, and the entity hierarchy (\BaseEntity\, \AbstractNonLeafEntity\, \DatabaseCatalog\, \SchemaCatalog\, \TableCatalog\, \ColumnCatalog\, \ColumnDesc\, \ColumnEncoding\, \DataDistributionType\). This eliminates the previous in-memory catalog structure that exposed tables and schemas via Apache Calcite interfaces.

java/catalog · high confidence

Architecture

Dev dashboard restructured into modular Python components

The development Grafana dashboard definition has been refactored from a single monolithic file into a modular Python package structure. The new layout organizes panels into distinct sections—Cluster, Metadata, Streaming, Source, Sink, Iceberg, Batch, Storage, Meta, and Misc—each implemented in its own module (e.g., \cluster\_alerts.py\, \streaming\_relations.py\, \hummock\_manager.py\). This change improves maintainability by separating concerns and allows for easier addition or modification of specific dashboard sections without touching unrelated code.

grafana/dashboard/dev · high confidence

Behavioural changes

Added global styles for text selection and loading animations

The dashboard now includes a new global stylesheet that disables text selection across the interface to improve the app-like feel, and defines a 'rw-spin' keyframe animation for loading indicators.

dashboard/styles · high confidence

CI now uploads Docker logs on integration test failure

The CI pipeline now automatically captures and uploads Docker container logs when an integration test fails. A new post-command hook checks the exit status of the test run; if it fails, it identifies the running containers (either a standalone Risingwave node or specific components like compactor, compute, frontend, and meta nodes) and uploads their logs as Buildkite artifacts for debugging.

ci/plugins/docker-compose-logs · high confidence

Compute node memory management and configuration refactoring

The compute node now uses a dedicated memory management module to control memory usage via a configurable total memory limit and a gradient-based reserved memory policy (30% for the first 16GB, 20% for the rest). This module includes an LRU watermark controller that adjusts eviction based on jemalloc and JVM memory statistics. Configuration for memory limits, parallelism, and other runtime options is now handled through command-line arguments and environment variables, with the deprecated \connector\_rpc\_endpoint\ option removed.

src/compute · high confidence

Concurrency protection for ALTER MATERIALIZED VIEW during sink creation

The system now prevents altering a materialized view if a sink job that depends on it is currently being created. This change ensures that DDL operations do not interfere with ongoing background sink initialization, avoiding potential state inconsistencies. The behavior is verified by new end-to-end tests in the \alter\_mv\ suite, which confirm that \ALTER MATERIALIZED VIEW\ statements are rejected with an error while a \CREATE SINK\ operation is in progress, and succeed once the sink is fully created.

_e2e\_test/streaming/alter\mv · high confidence

Dashboard UI redesign with new layout and catalog navigation

The dashboard components have been refactored to implement a new UI layout, featuring a fixed 216px left sidebar for navigation and a fluid main scroll region. This update introduces a comprehensive catalog navigation section that provides direct links to Sources, Tables, Materialized Views, Indexes, Internal Tables, Sinks, Views, Subscriptions, and Functions. The visualizations have been updated to use d3-dag for the relation graph and dagre for the fragment graph, replacing previous rendering approaches. Supporting components such as a new CatalogModal for viewing relation details, a GraphvizComponent for DOT output visualization, and utility components for time controls and metrics display have been added to support the new interface.

dashboard/components · high confidence

Dashboard pages rewritten with Next.js and Chakra UI

The dashboard pages have been completely rewritten as a Next.js application using the Chakra UI component library. This change introduces a new app shell (\\_app.tsx\) that manages global layout, loading states, and integrates the Monaco editor via vendored assets. Existing pages such as Cluster, Fragment Graph, Relation Graph, and Catalog views (Tables, Views, Sources, Sinks, etc.) have been migrated to this new framework, providing a modernized user interface and improved routing capabilities.

dashboard/pages · high confidence

Dashboard restructured as a Next.js application with new tooling and design system

The dashboard has been refactored to use the Next.js framework, enabling both standalone static HTML deployment and server-side rendering. This change introduces a new design system defined in \design-dna.json\ (featuring a warm-neutral palette and Linear-style aesthetics), updates the Node.js runtime to version 20, and migrates the package manager from npm to pnpm. Development and build tooling are now standardized with ESLint, Prettier, and TypeScript configurations, and a mock server is provided for local development.

dashboard · high confidence

Deprecate parallel unit mapping in favor of worker slot mapping

The system has replaced the parallel unit mapping mechanism with a worker slot mapping approach. This change simplifies the internal scheduling and resource allocation logic by aligning the frontend's view of compute resources with the actual worker slots, removing the abstraction layer of parallel units.

src/common, src/meta · high confidence

Deprecate unused MergeNode fields in streaming protocol

The proto definitions for the streaming engine have been cleaned up by removing unused fields from the MergeNode message. This change reduces the size of the serialized protocol messages and removes legacy data that is no longer referenced by the streaming execution logic.

proto · high confidence

Docker images and compose files now default to RisingWave v3.0.0 with Java 21 and Ubuntu 24.04

The Docker build environment has been upgraded to use Ubuntu 24.04 as the base OS and OpenJDK 21 for Java components, replacing previous versions. The default image tag in all Docker Compose configurations (including standalone and distributed setups for S3, GCS, Azure Blob, HDFS, and LakeKeeper) has been updated to v3.0.0. Additionally, the main Dockerfile now bundles the connector node and ADBC Snowflake driver, and the distributed compose file explicitly configures the meta node to use a PostgreSQL SQL backend.

docker · high confidence

Extended TTL retention support to non-append-only tables

Users can now configure storage retention (TTL) on non-append-only tables, a capability previously restricted to append-only tables. This feature is opt-in and requires enabling the \unsafe\_enable\_storage\_retention\_for\_non\_append\_only\_tables\ setting, reflecting the associated risks. The change also confirms that append-only tables with retention settings continue to function correctly with their downstream indexes.

_e2e\test/ttl · high confidence

Frontend query handling refactoring and error message refinement

The frontend's query handling logic has been refactored to reduce redundant code, particularly in the extended query protocol, and to improve the clarity of user-facing error messages for invalid operations. This change ensures that error reporting is more precise and helpful when users encounter issues with their SQL statements.

src/frontend · high confidence

JDBC sink connector rewritten with dialect-based writers and Snowflake key-pair auth

The JDBC sink connector has been refactored to use a new dialect-based architecture that automatically selects the correct SQL generation logic for PostgreSQL, MySQL, MariaDB, Redshift, Snowflake, and SQL Server. This change introduces dedicated sink writers (such as BatchAppendOnlyJDBCSink for Redshift and Snowflake) and enforces stricter validation, including mandatory primary key definitions for upsert sinks and pre-flight checks that the target table and columns exist. Snowflake sinks now support key-pair authentication via private key files or PEM content, and TCP keep-alive can be configured to prevent idle connection drops. The connector also improves reliability with automatic connection recreation on failure, batched statement execution, and explicit transaction isolation settings.

java/connector-node/risingwave-sink-jdbc · high confidence

Java connector logs are now routed to the unified tracing system

The Java connector-node now uses a custom SLF4J provider (TracingSlf4jAdapter) to forward log events to the Rust-based tracing infrastructure instead of using a standalone Java logging backend. This change ensures that Java logs, including exception stack traces and parameterized messages, are emitted through the same tracing channels as the rest of the system. It also introduces support for the RW\_JAVA\_LOG environment variable, allowing users to control log granularity for specific Java packages (e.g., io.debezium, org.apache.kafka) independently of the global Rust log level.

java/connector-node/tracing · high confidence

Kinesis connector restored and stabilized with improved error handling

The Kinesis connector has been fixed to work reliably again after previous regressions. This change resolves startup and throughput issues by adding proper timeout and retry logic to the Kinesis client, handling provisioned throughput exceeded exceptions, and correcting the NextToken/StreamName parameter conflict. It also suppresses expected Kafka enumerator errors and ensures the connector correctly handles timestamp-based startup modes and scale-in scenarios.

src/connector · high confidence

Moved legacy Kafka source test scripts to a dedicated location

The scripts and test data used to prepare Kafka topics for legacy source end-to-end tests have been relocated from \scripts/source\ to \e2e\_test/source\_legacy/basic/scripts\. This directory now contains the \prepare\_ci\_kafka.sh\ script for setting up topics and data, the \schema\_registry\_producer.py\ utility for producing Avro and JSON Schema messages, and the associated test data files, ensuring the test environment is correctly initialized for these legacy source tests.

_e2e\_test/source\legacy/basic/scripts · high confidence

New expression evaluation architecture with strict and non-strict modes

The expression evaluation engine has been restructured to support both strict and non-strict evaluation modes. Strict mode (the default) preserves existing behavior by wrapping expressions to report errors. Non-strict mode allows batch queries to continue processing and return null values instead of failing when an expression error occurs, controlled by the \batch\_expr\_strict\_mode\ setting. This change introduces new core modules for expression definitions (\def.rs\), the \AggregateFunction\ trait (\mod.rs\), and specific implementations for scalar wrappers and user-defined aggregates (\user\_defined.rs\), alongside the error handling (\error.rs\) and expression building logic (\build.rs\) that enforce these modes.

src/expr · high confidence

RPC client layer refactored with distributed tracing and observability

The RPC client implementation in \src/rpc\_client\ has been refactored to enhance observability and reliability. A new \WrappedChannel\ service wrapper now automatically injects distributed tracing context into gRPC request headers and attaches the gRPC call path to response headers, enabling end-to-end trace correlation and more precise error reporting. Additionally, the client layer now enforces a strict error handling policy where gRPC statuses must be converted via service-specific methods to ensure the service name is always included in error messages, and connection pooling is standardized across all client types.

_src/rpc\client · high confidence

Refactor dashboard definition into modular Python components

The Grafana dashboard definition logic has been restructured from a single file into multiple Python modules within the \grafana/dashboard\ package. This change introduces a \common.py\ module containing shared configuration, environment variable handling (such as namespace filtering and dynamic datasource selection), and reusable layout and panel helper classes. This modularization simplifies the maintenance of dashboard definitions and allows for more granular control over common dashboard elements.

grafana/dashboard · high confidence

Refactor object store to use OpenDAL as the default S3 backend

The object store implementation has been refactored to use OpenDAL as the default backend for S3 and S3-compatible storage (including MinIO), replacing the previous direct AWS SDK integration. This change introduces a unified \OpendalObjectStore\ engine that supports multiple storage backends (S3, MinIO, GCS, Azure Blob, HDFS, OSS, OBS, WebHDFS, and local FS) through a consistent interface. The new implementation includes configurable HTTP transport, retry and timeout layers, concurrency limits, and streaming upload capabilities. The old S3 implementation remains available but is now secondary to the OpenDAL-based approach.

_src/object\store · high confidence

Removal of legacy planner interface definitions

The \java/planner\ module has removed several core interface definitions that were part of the previous planner architecture. Specifically, the \Planner\ interface, along with the \RisingWaveRel\, \RisingWaveLogicalRel\, and \RisingWavePhyRel\ interfaces (which extended Calcite's \RelNode\ and \PhysicalNode\), have been deleted. This cleanup removes the old abstraction layer for logical and physical relations, likely as part of the ongoing implementation of the new RisingWave planner framework.

java/planner · high confidence

Restrictions on CDC table creation and rate limits

Users can no longer directly create CDC tables using connectors like mysql-cdc, postgres-cdc, or sqlserver-cdc; they must instead use CREATE SOURCE followed by CREATE TABLE FROM SOURCE. Additionally, creating CDC sources with a source\_rate\_limit of 0 is now rejected for mysql-cdc and sqlserver-cdc, as these connectors require reading an initial offset before creation completes.

_e2e\_test/source\inline/cdc · high confidence

Support for altering struct column types with downstream dependencies and indexes

Users can now alter the type of columns containing structs (including empty structs) while maintaining compatibility with existing materialized views and indexes. When a struct column is modified, existing downstream materialized views continue to function using the previous schema definition, with dropped fields appearing as NULLs, while new downstream objects reflect the updated structure. Additionally, internal indexes referencing struct fields are automatically rewritten to account for field shifts or removals, ensuring query correctness without manual intervention.

_e2e\_test/ddl/alter\_table\_column\type · high confidence

Fixes

Fix compaction task memory estimation

Corrects the calculation of memory usage for compaction tasks, ensuring that the system accurately accounts for the memory required during the compaction process. This prevents potential issues related to memory overcommitment or underestimation during compaction operations.

src/storage · high confidence

Fix orphaned pending sink state rows during upgrade

The system now correctly handles the upgrade migration for exactly-once Iceberg sinks by cleaning up orphaned \pending\_sink\_state\ rows that were left behind in older versions when a sink was dropped. This ensures that the new \ON DELETE CASCADE\ foreign key constraint can be added without violating referential integrity, preventing upgrade failures on Postgres, MySQL, and SQLite backends.

_e2e\test/backwards-compat-tests/slt/pending-sink-state · high confidence

Fixes for temporal join cache and watermark handling in hash join

Resolves a bug where the temporal join cache was not functioning correctly, ensuring accurate results for time-based joins. Additionally, improves watermark handling for hash join operations, leading to more efficient state management and better performance when processing streaming data with watermarks.

src/stream · high confidence

Upload CI failure logs as artifacts

When a CI build step fails, the system now automatically packages and uploads relevant diagnostic logs to Buildkite artifacts. This includes zipped RisingWave logs, regression test output results (if present), and the connector node log, making it easier to debug failures without needing to access the full build environment.

ci/plugins/upload-failure-logs-zipped · high confidence

Test coverage

Add BigQuery sink integration test demo; Add Citus CDC integration test environment; Add CockroachDB sink integration test demo; Add Debezium-PostgreSQL integration test environment; Add Doris sink integration test demo; Add Iceberg sink2 integration test for REST, storage, and JDBC catalogs; Add Iceberg sink2 integration tests for Hive, JDBC, REST, and storage catalogs; Add JMH benchmarks for Java binding row iteration; Add Kafka CDC Sink integration test; Add Kafka CDC compatibility integration test; Add MQTT integration tests for source and sink connectors; Add MindsDB integration test with home rentals prediction; Add MongoDB CDC integration test infrastructure; Add MongoDB CDC integration test with Debezium; Add MySQL CDC integration test suite; Add Nexmark benchmark materialized view definitions; Add PostgreSQL sink integration test demo; Add SQL Server + Debezium integration test; Add client-library integration tests for Go, Python, Java, Node.js, PHP, Ruby, and C\#; Add comprehensive PostgreSQL CDC data compatibility test; Add comprehensive end-to-end tests for struct types and operations; Add end-to-end test for Iceberg CDC with MySQL source; Add end-to-end test for PostgreSQL wire protocol extended mode; Add end-to-end tests for error UI formatting and license behavior; Add extended mode end-to-end tests; Add integration test environments for Iceberg Hive, JDBC, REST, and Storage catalogs; Add integration test for Avro upsert source with Kafka; Add integration test for Kinesis and S3 sources with timestamp support; Add integration tests for Iceberg source with REST catalog location derivation; Add legacy end-to-end tests for Kafka and Pulsar sources; Add local execution script for backward compatibility tests; Add streaming ClickHouse benchmark test suite; Add streaming TPC-H end-to-end test suite; Add streaming TPC-H view definitions for queries 1–22; Add streaming demo test cases for ad CTR, ecommerce, metric analysis, and Twitter; Added CDC source integration tests for MongoDB, MySQL, PostgreSQL, and Oracle; Added CH-Benchmark source definitions and cleanup scripts; Added CH-benchmark batch query test suite; Added DDL creation race-condition tests; Added DataFusion engine test suite for Iceberg tables; Added DuckDB compatibility tests for inner and left outer joins; Added DuckDB compatibility tests for schema-qualified column references; Added Iceberg predicate pushdown benchmarks; Added Iceberg source integration test for REST table location derivation; Added Java client integration tests; Added Kafka backwards-compatibility tests for invalid options and upsert formats; Added MongoDB CDC end-to-end tests for basic data types and TLS/mTLS connectivity; Added NATS integration test demo; Added NexMark end-to-end test suite with Kafka source support; Added Protobuf-based Twitter integration test; Added Pub/Sub integration test suite; Added SQL Server CDC end-to-end tests and regression suite; Added TPC-H end-to-end test fixtures for Kafka sources and data insertion; Added TiDB CDC sink integration test configuration and SQL definitions; Added backward compatibility tests for MySQL and PostgreSQL CDC sources; Added backward compatibility tests for SST filter metadata and data integrity; Added backward-compatibility test for stale table IDs in Hummock; Added backward-compatibility tests for AsOf join with materialized views; Added backward-compatibility tests for sink-into-table functionality; Added backward-compatibility tests for streaming parallelism configuration migration; Added backwards-compatibility tests for hash join watermark handling; Added compatibility test scripts for MySQL and PostgreSQL CDC connectors; Added direct CDC offset injection test suite for MySQL; Added e2e extended mode test and compaction test tool; Added e2e test for PostgreSQL hstore type array support; Added e2e tests for Parquet source decoding and rate-limiting; Added e2e tests for UDF retry and graceful shutdown behavior; Added e2e tests for streaming over window state cleaning and watermark forwarding; Added end-to-end test for Kafka background DDL; Added end-to-end test for debug splits configuration; Added end-to-end test for user keyword handling; Added end-to-end tests for CASCADE drop operations and background DDL; Added end-to-end tests for CDC auto schema change and PostgreSQL TOAST handling; Added end-to-end tests for CDC table alterations and backfill rate limiting; Added end-to-end tests for CUBE, ROLLUP, and GROUPING SETS aggregation; Added end-to-end tests for DML rate limiting and persistence behavior; Added end-to-end tests for DuckDB Common Table Expression (CTE) behavior; Added end-to-end tests for DuckDB join edge cases; Added end-to-end tests for DuckDB limit execution; Added end-to-end tests for EMIT ON WINDOW CLOSE streaming features; Added end-to-end tests for Group TopN and index-accelerated TopN; Added end-to-end tests for HashiCorp Vault secret backend integration; Added end-to-end tests for IEJoin operator; Added end-to-end tests for Kafka Protobuf source schema evolution and field presence; Added end-to-end tests for Kafka SASL authentication and connection alteration; Added end-to-end tests for LDAP authentication; Added end-to-end tests for MQTT source with RabbitMQ and Protobuf support; Added end-to-end tests for MySQL CDC table sources and schema handling; Added end-to-end tests for Nexmark streaming sink queries; Added end-to-end tests for ORDER BY and LIMIT/OFFSET behavior; Added end-to-end tests for Pulsar source connector capabilities; Added end-to-end tests for arrangement backfill runtime behavior; Added end-to-end tests for backfill order control; Added end-to-end tests for batch join operations; Added end-to-end tests for batch session window functionality; Added end-to-end tests for batch transaction features; Added end-to-end tests for bit aggregate functions and list min/max; Added end-to-end tests for cross products and right outer joins; Added end-to-end tests for dashboard graph rendering and backpressure stats; Added end-to-end tests for generated columns, default values, and watermark restrictions; Added end-to-end tests for list type casting and storage; Added end-to-end tests for locality backfill scenarios; Added end-to-end tests for materialized view backfill progress tracking; Added end-to-end tests for meta store snapshot creation and cleanup; Added end-to-end tests for mysql\_query and postgres\_query TVFs; Added end-to-end tests for null-safe join operations; Added end-to-end tests for refreshable table behavior; Added end-to-end tests for shared CDC sources and validation rules; Added end-to-end tests for sink backfill behavior and validation; Added end-to-end tests for snapshot backfill; Added end-to-end tests for streaming GroupTopN and WindowTopN; Added end-to-end tests for subscription cursor functionality; Added end-to-end tests for time travel functionality; Added end-to-end tests for timezone handling and user documentation examples; Added end-to-end tests for vector search capabilities; Added end-to-end tests for webhook source and WebSocket ingest features; Added integration test for Iceberg sink2; Added integration test setup for Protobuf-based live stream metrics; Added regression tests for streaming bug fixes; Added sink test definitions for Nexmark queries; Added slow tests for streaming backfill rate limiting and arrangement backfill behavior; Added test for backfill rate limiting with slow UDFs; Added test for cross-database materialized view backfill and recovery; Added test for group aggregation internal state consistency; Backwards compatibility tests for multiple version columns support; Comprehensive e2e test suite for batch query operations; Consolidated DuckDB end-to-end test suite; E2E test coverage for session initialization defaults; E2E tests for Kafka source alteration capabilities; End-to-end test coverage for COPY query output; End-to-end tests for NATS source with parallelism, subject inclusion, and stream creation controls; End-to-end tests for background DDL operations on materialized views, indexes, and sinks; End-to-end tests for connection lifecycle, ownership, and schema visibility; End-to-end tests for streaming temporal joins; End-to-end tests for visibility mode barriers and checkpoints; Expanded JDBC sink e2e test coverage for MySQL and advanced data types; Expanded S3 file source/sink testing for Parquet and refresh behaviors; Expanded TPC-H batch query test coverage; Expanded batch aggregation test coverage; Expanded end-to-end Nexmark test coverage; Expanded end-to-end test coverage for Iceberg engine tables and sinks; Expanded end-to-end test coverage for Iceberg sinks and sources; Expanded end-to-end test coverage for Kafka sink configurations and formats; Expanded end-to-end test coverage for Kafka source connectors; Expanded end-to-end test coverage for User-Defined Functions; Expanded end-to-end test coverage for over window functions; Expanded streaming aggregate test coverage; Expanded support for complex subquery patterns in batch queries; Iceberg V3 sink and engine test coverage expansion; Inline end-to-end tests for Google Pub/Sub source connector; Moved NEXMark and TPC-H backward compatibility tests to e2e\_test; New Iceberg end-to-end test suite with Spark 4.0 support; New Iceberg end-to-end test utilities for test discovery and execution; New MySQL sink integration tests added; New UDF implementations for end-to-end testing; New batch e2e test suite for local execution, configuration, and query planning; New end-to-end tests for Avro schema evolution and source/table alteration; New end-to-end tests for DDL operations; New end-to-end tests for core data types; New end-to-end tests for streaming materialized views; New integration test automation scripts for demo validation; New integration tests for ad analytics, CDN metrics, clickstream, and live streaming; New micro-benchmarking framework for batch executors; New streaming join tests and configurable join encoding; Nexmark endless materialized view tests use retry-based assertions; Regression test for force-append-only sink panic on non-unique primary keys; Sink-into-table e2e tests reorganized and expanded.

Dependencies

Routine dependency updates across Rust, Java, and JavaScript stacks

This release includes routine updates to dependencies across the Rust, Java, and JavaScript stacks. Key Rust updates include bumping tokio to v1.44.0, syn from 1.0.109 to 2.0.66, and tower from 0.4.13 to 0.5.0. Java dependencies have been upgraded, including mongodb from 3.5.1 to 3.9.1 and org.apache.hive:hive-metastore from 4.1.0 to 4.2.0. JavaScript dependencies in the dashboard have been updated, such as next from 14.2.22 to 14.2.25. These updates ensure the project uses the latest stable versions of its core libraries.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 60 → 62 (+1.9)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.

Lenses

  • Code Health 82 → 81 (-0.9)
  • Architecture 56 → 60 (+4.0)
  • Maturity 74 → 74 (+0.0)
  • Readiness 69 → 59 (-10.0)
  • Security 55 → 64 (+9.2)
  • Domain Modelling 92 → 68 (-24.5)
  • Event Sourcing 100 → 100 (+0.0)
  • Accessibility 66 → 66 (+0.0)
  • Performance 100 (new)

Resolved (210)

  • Boundary-crossing change coupling: create_table.rs ↔ stream_cdc_table_scan.rs (src/frontend/src/handler/create_table.rs)
  • Boundary-crossing change coupling: function_catalog.rs ↔ mod.rs (src/frontend/src/catalog/function_catalog.rs)
  • Boundary-crossing change coupling: function_catalog.rs ↔ mod.rs (src/frontend/src/catalog/function_catalog.rs)
  • Boundary-crossing change coupling: hummock_service.rs ↔ hummock_meta_client.rs (src/meta/service/src/hummock_service.rs)
  • Boundary-crossing change coupling: hummock_service.rs ↔ hummock_meta_client.rs (src/meta/service/src/hummock_service.rs)
  • Boundary-crossing change coupling: notification_service.rs ↔ observer_manager.rs (src/meta/service/src/notification_service.rs)
  • Boundary-crossing change coupling: observer_manager.rs ↔ notification_service.rs (src/common/common_service/src/observer_manager.rs)
  • Boundary-crossing change coupling: observer_manager.rs ↔ observer_manager.rs (src/frontend/src/observer/observer_manager.rs)
  • Boundary-crossing change coupling: relabeled_metric.rs ↔ streaming_stats.rs (src/common/metrics/src/relabeled_metric.rs)
  • Boundary-crossing change coupling: schedule.rs ↔ iceberg_compactor_runner.rs (src/meta/src/manager/iceberg_compaction/schedule.rs)
  • CatalogController::update_source_props_by_source_id (cognitive 16) (src/meta/src/controller/streaming_job.rs)
  • CatalogController::update_source_props_by_source_id (cyclomatic 16) (src/meta/src/controller/streaming_job.rs)
  • CatalogController::update_source_rate_limit_by_source_id (cognitive 17) (src/meta/src/controller/streaming_job.rs)
  • Change coupling: batch_hash_join.rs ↔ batch_lookup_join.rs (src/frontend/src/optimizer/plan_node/batch_hash_join.rs)
  • Change coupling: batch_hash_join.rs ↔ stream_hash_join.rs (src/frontend/src/optimizer/plan_node/batch_hash_join.rs)
  • Change coupling: config.rs ↔ mod.rs (src/risedevtool/src/config.rs)
  • Change coupling: create_source.rs ↔ create_table_as.rs (src/frontend/src/handler/create_source.rs)
  • Change coupling: kinesis.rs ↔ utils.rs (src/connector/src/sink/kinesis.rs)
  • Change coupling: local_hummock_storage.rs ↔ traced_store.rs (src/storage/src/hummock/store/local_hummock_storage.rs)
  • Change coupling: main.rs ↔ store_impl.rs (src/storage/hummock_test/src/bin/replay/main.rs)
  • …and 190 more

New (222)

  • Boundary-crossing change coupling: distributed_lookup_join.rs ↔ scan.rs (src/batch/executors/src/executor/join/distributed_lookup_join.rs)
  • Boundary-crossing change coupling: function_catalog.rs ↔ mod.rs (src/frontend/src/catalog/function_catalog.rs)
  • Boundary-crossing change coupling: function_catalog.rs ↔ mod.rs (src/frontend/src/catalog/function_catalog.rs)
  • Boundary-crossing change coupling: handle_privilege.rs ↔ utils.rs (src/frontend/src/handler/handle_privilege.rs)
  • Boundary-crossing change coupling: hummock_service.rs ↔ hummock_meta_client.rs (src/meta/service/src/hummock_service.rs)
  • Boundary-crossing change coupling: hummock_service.rs ↔ hummock_meta_client.rs (src/meta/service/src/hummock_service.rs)
  • Change coupling: hash_join.rs ↔ tests.rs (src/stream/src/executor/hash_join.rs)
  • Change coupling: logical_join.rs ↔ stream_delta_join.rs (src/frontend/src/optimizer/plan_node/logical_join.rs)
  • Change coupling: logical_over_window.rs ↔ over_window_to_topn_rule.rs (src/frontend/src/optimizer/plan_node/logical_over_window.rs)
  • Change coupling: replay_impl.rs ↔ monitored_store.rs (src/storage/hummock_test/src/bin/replay/replay_impl.rs)
  • Change coupling: state_table.rs ↔ sort.rs (src/stream/src/common/table/state_table.rs)
  • Change coupling: stream_delta_join.rs ↔ stream_temporal_join.rs (src/frontend/src/optimizer/plan_node/stream_delta_join.rs)
  • Change coupling: stream_hash_agg.rs ↔ stream_simple_agg.rs (src/frontend/src/optimizer/plan_node/stream_hash_agg.rs)
  • Change coupling: stream_hash_agg.rs ↔ stream_stateless_simple_agg.rs (src/frontend/src/optimizer/plan_node/stream_hash_agg.rs)
  • Change-coupling hub: batch_hash_join.rs → batch_lookup_join.rs, stream_delta_join.rs, stream_hash_join.rs (src/frontend/src/optimizer/plan_node/batch_hash_join.rs)
  • Change-coupling hub: mod.rs → config.rs, risedev_env.rs, service_config.rs (src/risedevtool/src/task/mod.rs)
  • Change-coupling hub: reader.rs → reader.rs, reader.rs, reader.rs (src/connector/src/source/nexmark/source/reader.rs)
  • Change-coupling hub: store_impl.rs → main.rs, replay_impl.rs, test_utils.rs, test_utils.rs (src/storage/src/store_impl.rs)
  • Change-coupling hub: stream_hop_window.rs → stream_dynamic_filter.rs, stream_hash_agg.rs, stream_hash_join.rs, stream_project.rs, stream_project_set.rs, stream_union.rs (src/frontend/src/optimizer/plan_node/stream_hop_window.rs)
  • Change-coupling hub: stream_union.rs → stream_dynamic_filter.rs, stream_hash_agg.rs, stream_project.rs, stream_stateless_simple_agg.rs (src/frontend/src/optimizer/plan_node/stream_union.rs)
  • …and 202 more

Changes since last survey

  • 94 commits — 63 feature/other, 31 fixes

By area

  • (root) — 35 commits
  • src/meta — 13 commits
  • src/stream — 8 commits
  • src/frontend — 6 commits
  • src/storage — 6 commits
  • java/connector-node — 3 commits
  • src/connector — 3 commits
  • src/tests — 3 commits
  • docker/dashboards — 2 commits
  • e2e_test/source_inline — 2 commits
  • e2e_test/source_legacy — 2 commits
  • ci/scripts — 1 commit
  • dashboard/package.json — 1 commit
  • docs/dev — 1 commit
  • e2e_test/iceberg — 1 commit
  • e2e_test/streaming — 1 commit
  • integration_tests/feature-store — 1 commit
  • src/batch — 1 commit
  • src/common — 1 commit
  • src/config — 1 commit

Notable commits

  • fix: fix(cdc): align heartbeat interval validation across Rust and Java (#27017)
  • fix: fix(cdc): enhance string decimal handling mode validation (#27143)
  • fix: fix(cdc): force-close SQL Server connections on shutdown (#27186)
  • fix: fix(ci): download minio and mc from RisingWave CI mirror (#27097)
  • fix: fix(common): prevent memory accounting drift (#27069)
  • fix: fix(connector): preserve SQL Server composite primary key order (#27160)
  • fix: fix(dashboard): normalize output blocking ratio in user dashboard (#27003)
  • fix: fix(grafana): use container_memory_rss for Node Memory relative (#27142)
  • fix: fix(iceberg): derive REST table location from namespace (#27104)
  • fix: fix(iceberg): merge manifests during COW overwrite (#26992)
  • fix: fix(iceberg): reject row lineage column names for V3 tables (#27223)
  • fix: fix(iceberg): reject unsupported primary keys in engine tables (#27105)
  • fix: fix(meta): fence sink coordinators during recovery (#27180)
  • fix: fix(meta): preserve compaction candidates across topology changes (#27042)
  • fix: fix(meta): preserve connection and secret refs during connector alters (#26923)
  • fix: fix(meta): recover source splits for snapshot backfill jobs (#27208)
  • fix: fix(meta): stabilize automatic compaction group split and merge (#27044)
  • fix: fix(object-store): propagate metadata errors during listing (#27266)
  • fix: fix(optimizer): avoid unsafe correlated aggregate decorrelation (#27065)
  • fix: fix(refresh): finish a table refresh only after all materialize actors and abandon it on recovery (#27041)
  • …and 74 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

risingwavelabs/risingwave was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 5149d7b2cbc1fa8b6d66f56e782f02965b8554a2 — the exact code this score is about.
  • Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-fbec9b1e08c2.