Skip to content
CAI
Software that uses CAICheck a score

GreptimeTeam/greptimedb

60.1

Adequate · 29 September 2026

786.6k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a distributed time-series database that manages data ingestion, storage, and querying across a cluster of frontend, datanode, and metasrv components. It provides a SQL-compatible interface with support for complex data types, full-text and inverted indexing, and flexible write-ahead log backends like Kafka and Raft Engine. The platform includes robust operational tooling for metadata management, data export/import, and performance benchmarking, alongside a comprehensive authentication and authorization subsystem for secure access control.

How it got here

2022 — Core infrastructure and distributed query foundation

51 changes.

This period focused on establishing the foundational architecture for a distributed time-series database, introducing core crates for storage, logging, and data types. Significant work included building the distributed query engine with DataFusion integration, implementing region management, and creating unified APIs for gRPC and SQL. The effort also covered essential operational capabilities such as cluster metadata management, telemetry, and comprehensive testing infrastructure.

2023 — infrastructure modernization and extensibility

56 changes.

This period focused on rebuilding core infrastructure with a modular, procedure-driven architecture, introducing a common framework for background tasks, metadata caching, and key-value backends. Significant work was done to enhance extensibility through new plugins, authentication, and telemetry systems, while expanding data handling capabilities with support for diverse file formats, object storage, and distributed query testing.

2024–2025 — Flownode, WAL, and Metadata Infrastructure

77 changes.

This period focused on establishing the Flownode service for dual-engine flow execution and introducing a comprehensive, configurable Write-Ahead Log (WAL) system supporting Kafka, Raft Engine, and Object Store backends. Significant work was also done to refactor the metadata layer, implementing robust reconciliation procedures, soft-drop capabilities, and extensive CLI tools for metadata inspection and repair.

2026 — CLI v2 and observability infrastructure

17 changes.

This period focused on introducing the Import/Export V2 CLI workflows and restructuring OTLP trace ingestion to support advanced schema reconciliation and v2 pipelines. Concurrently, the codebase expanded its internal observability by implementing structured DDL and maintenance event recording, while establishing foundational batching libraries and object-store-backed WAL capabilities.

Features

Add CLI commands for logically deleting metadata keys and tables

New CLI commands have been added to the metadata control module to allow users to logically delete metadata entries. The \del key\ command removes specific key-value pairs or all keys matching a given prefix from the metadata store by creating tombstones, while the \del table\ command logically deletes a table's metadata by resolving the table ID and invoking the table metadata manager to create corresponding tombstones. Both commands support configuration via store backend options and include validation to ensure correct usage, such as requiring a single target for table deletion.

src/cli/src/metadata/control/del · high confidence

Add Decimal128 type support

A new \Decimal128\ type has been added to the common decimal library, providing a 128-bit decimal representation backed by \i128\ with configurable precision (up to 38 digits) and scale. This implementation includes parsing from strings (leveraging \rust\_decimal\ for shorter inputs and \bigdecimal\ for longer ones to avoid underflow), serialization/deserialization support, and standard traits like \Display\, \Hash\, and comparison operators, enabling precise decimal arithmetic in downstream components.

src/common/decimal · high confidence

Add MySQL and PostgreSQL RDS key-value backends with TLS support

Users can now store metadata in MySQL or PostgreSQL databases instead of only embedded or other backends. The new MySQL and PostgreSQL implementations in the RDS key-value backend support configurable TLS modes (Disable, Prefer, Require, VerifyCa, VerifyFull) for secure connections, and the MySQL backend automatically creates the required schema and uses MEDIUMBLOB for value storage to handle larger metadata entries.

_src/common/meta/src/kv\backend/rds · high confidence

Add Substrait serialization support for DataFusion logical plans

The \src/common/substrait\ module now provides the infrastructure to serialize and deserialize DataFusion logical plans using the Substrait protocol. This includes a \DFLogicalSubstraitConvertor\ for encoding/decoding plans, an \ExtensionSerializer\ that handles custom PromQL-related logical nodes (such as \Absent\, \HistogramFold\, and \UnionDistinctOn\), and the associated error definitions. This enables the system to exchange query plans in a standardized format.

src/common/substrait · high confidence

Add TLS support for MySQL and Postgres metadata backends

The meta-srv now supports TLS-encrypted connections for its MySQL and Postgres metadata backends, allowing users to secure metadata storage in transit. New utility modules in \src/meta-srv/src/utils\ provide functions to build MySQL and Postgres key-value backends and election implementations with configurable TLS modes (Disable, Prefer, Require, VerifyCa, VerifyFull) and mutual TLS certificates. For Postgres, Unix domain socket connections automatically fall back to non-TLS when TLS is enabled, as TLS is not supported over Unix sockets. The etcd backend also gains TLS configuration support with a fallback to non-TLS if the preferred TLS mode fails.

src/meta-srv/src/utils · high confidence

Add in-memory catalog manager for testing

Introduces a new \MemoryCatalogManager\ implementation in the catalog module, providing a simple in-memory list of catalogs, schemas, and tables. This component supports standard catalog operations such as listing and checking existence of catalogs, schemas, and tables, as well as retrieving table metadata by name or ID. It is designed primarily for use in test environments and internal query engine construction, offering a lightweight alternative to persistent catalog backends.

src/catalog/src/memory · high confidence

Add object store backend implementations for S3, GCS, Azure Blob, OSS, and local filesystem

New backend builder functions have been added for Amazon S3, Google Cloud Storage, Azure Blob Storage, Alibaba Cloud OSS, and the local filesystem within the object store module. These implementations allow users to configure and connect to these storage providers using specific connection parameters (such as endpoints, credentials, and buckets) and automatically apply retry and instrumentation layers to the underlying object store operations.

_src/common/datasource/src/object\store · high confidence

Add static and file-watching user providers for authentication

Users can now authenticate against a static set of credentials or a credential file. The new \StaticUserProvider\ supports inline configuration via a \cmd:\ mode (e.g., \cmd:user=pwd\) or loading from a file via \file:\<path\>\, with support for PBKDF2-SHA256 password hashes. The new \WatchFileUserProvider\ reads credentials from a specified file and automatically reloads them when the file changes, allowing runtime credential updates without restarting the service; it fails initialization if the file is missing but retains the last known good credentials if the file is deleted at runtime.

_src/auth/src/user\provider · high confidence

Add timestamp filter expression builder for DataFusion

The logical plan layer now includes a new \expr.rs\ module that provides utilities for constructing DataFusion filter expressions based on timestamp ranges. This enables the query engine to automatically generate efficient time-range filters (e.g., \\>= start\ and \\< end\) for timestamp columns, supporting various time units (second, millisecond, microsecond, nanosecond) and integrating with the DataFusion expression system.

_src/common/query/src/logical\plan · high confidence

Add trigger DDL management with interval, for, and keep\_firing\_for support

This change introduces the core data structures and serialization logic for managing triggers within the DDL subsystem. Users can now define triggers with specific execution intervals, as well as optional \for\ and \keep\_firing\_for\ duration parameters to control alert firing behavior. The implementation includes support for webhook notification channels and handles the bidirectional conversion between internal task representations and protocol buffer messages for cluster communication.

src/common/meta/src/rpc/ddl · high confidence

Added NoopLogStore as a placeholder WAL implementation

A new NoopLogStore implementation has been added to the log-store module, providing a non-functional Write-Ahead Log (WAL) provider. This implementation returns empty results for reads and namespaces, and assigns entry ID 0 for all appended batches, effectively acting as a stub or placeholder for the WAL layer. This change supports the preparation for retiring the previous soft-drop WAL mechanism by introducing a no-op alternative that can be used where a real WAL is not required or during transitional phases.

src/log-store/src/noop · high confidence

Added cgroup-based CPU and memory limit and usage metrics

The \src/common/stat\ module now includes a new \cgroups.rs\ implementation that reads CPU and memory limits and usage directly from the cgroups filesystem (supporting both v1 and v2). This enables the system to accurately report resource constraints in containerized environments, addressing a previous issue where CPU cores were incorrectly calculated as zero. The module exposes functions to retrieve total CPU (in millicores) and memory (in bytes) limits, falling back to host system values when cgroup limits are unset, and provides a \ResourceStat\ trait with an implementation that periodically collects CPU usage metrics for cgroup v2 environments.

src/common/stat · high confidence

Added dummy catalog implementation for region server queries

A new \DummyCatalog\ implementation has been added to the catalog module to handle table resolution for the region server. This component delegates catalog, schema, and table lookups directly to the central \CatalogManager\, ensuring that queries executed in the region server context can correctly resolve tables through the existing catalog infrastructure rather than relying on local static definitions.

_src/catalog/src/table\source · high confidence

Added meta-client example demonstrating heartbeat and KV operations

A new example file \src/meta-client/examples/meta\_client.rs\ has been added to demonstrate how to initialize a MetaClient for a datanode, configure channel settings, and perform core operations including sending heartbeats, retrieving leader information, and executing key-value range and batch-get requests.

src/meta-client/examples · high confidence

Added protobuf definitions and build script for log-store

The log-store module now includes a build script and proto definitions to generate Rust bindings for log-store data structures. This introduces the \logstore.proto\ file, which defines the \EntryImpl\, \LogStoreState\, and \NamespaceImpl\ messages, ensuring that the log-store component has the necessary protocol buffer infrastructure to serialize and deserialize its internal state.

src/log-store · high confidence

Added table metadata benchmarking tool

A new \TableMetadataBencher\ component has been introduced in the CLI benchmarking suite to measure the performance of table metadata operations. This tool allows users to benchmark the latency of creating, retrieving, deleting, and renaming table metadata by interacting with the \TableMetadataManager\.

src/cli/src/bench · high confidence

Added utility functions for system schema column definitions

A new \tables.rs\ module was added to \src/catalog/src/system\_schema/utils\ providing helper functions (\string\_columns\, \string\_column\, \bigint\_column\, \timestamp\_micro\_column\) to generate \ColumnSchema\ definitions for system catalog tables. This simplifies the creation of standard column schemas (string, bigint, timestamp) used in system schema definitions and includes unit tests for the string column generation logic.

_src/catalog/src/system\schema/utils · high confidence

Automated Grafana dashboard generation and validation

Added scripts to automatically generate Grafana dashboard JSON files from a source definition using the \dac\ tool and to validate the resulting dashboards. The generation script (\gen-dashboards.sh\) creates intermediate YAML/Markdown representations and produces standalone dashboards by removing instance-specific filters. The check script (\check.sh\) ensures that generated dashboards have valid descriptions for stats and timeseries panels and that datasource UIDs are correctly configured as variables.

grafana/scripts · high confidence

CLI metadata control commands for inspecting and modifying store data

The CLI now includes a new set of metadata control subcommands that allow users to directly interact with the underlying metadata store. Users can retrieve key-value pairs or specific table metadata (info and routes) using the new \get\ commands, which support prefix queries, limits, and pretty-printing. Additionally, new \put\ commands enable writing raw key-value pairs or table metadata into the store, while \del\ commands allow for the removal of specific keys or table entries, including support for tombstone prefixes. These changes provide granular, low-level control over metadata storage for debugging and maintenance tasks.

src/cli/src/metadata/control · high confidence

CLI metadata put commands for keys and table metadata

The CLI now includes new \metadata put key\ and \metadata put table\ commands to allow direct updates to the metadata store. The \put key\ command writes arbitrary key-value pairs, enforcing validation for supported metadata types (such as flow state, views, catalogs, schemas, and topics) while blocking unsafe direct writes to critical structures like table routes and info unless the \--no-validate\ flag is used. The \put table\ command provides subcommands for updating table info and table route metadata, reading JSON values from standard input and leveraging the \TableMetadataManager\ to ensure consistency during the update process.

src/cli/src/metadata/control/put · high confidence

Derive macros for automatic Row and Schema generation

New \IntoRow\ and \Schema\ derive macros are now available for structs in the row macro crate. By applying these derives, users can automatically generate \into\_row()\ and \schema()\ methods, which convert struct instances into row data and define column schemas respectively. The macros support configuration via a \col\ attribute to specify column names, data types (including JSON), and semantic types (field, tag, timestamp), significantly reducing boilerplate for data model definitions.

src/common/macro/src/row · high confidence

Expanded information\_schema virtual tables for cluster, flow, and metadata visibility

The information\_schema now exposes several new virtual tables to improve observability and metadata introspection. The new cluster\_info table provides topology details for each peer (frontend, datanode, metasrv), including CPU and memory usage, version, and uptime. The flows table now includes the full flow\_definition, source and sink table names, and flownode addresses, while a dedicated flow\_statistics table adds runtime metrics such as uptime and state size. Additionally, the key\_column\_usage table now reports Greptime-specific index types (inverted, fulltext, skipping), and the partitions table includes a greptime\_partition\_id column to map partitions to regions.

_src/catalog/src/system\_schema/information\schema · high confidence

Implement drop database procedure

The drop database procedure is now implemented in the meta service, allowing users to remove a database and all its contents. The procedure follows a multi-step state machine: it first validates the schema's existence, then iterates through all tables to drop them logically and physically (including handling views and invalidating caches), and finally removes the schema metadata and invalidates the schema cache.

_src/common/meta/src/ddl/drop\database · high confidence

Introduce Bloom Filter as a Full-Text Search Backend

The index module now supports Bloom Filters as a backend for full-text search, providing an alternative to the existing Tantivy-based implementation. This change adds the core components for Bloom Filter management, including a creator for building filters with configurable false-positive rates and memory control, a reader for loading and probing filters, and an applier for executing search predicates against row groups. The implementation includes support for intermediate file storage to handle memory constraints during creation and benchmarks to validate performance for both English and Chinese tokenization scenarios.

src/index · high confidence

Introduce Cyborg automation for CI, documentation, and version management

This change introduces the 'cyborg' tool, a centralized automation service for the GreptimeTeam repositories. It adds several new scripts: 'bump-versions' to trigger downstream version updates for the website, demo-scene, and docs repositories (skipping nightly versions for the website and demo); 'check-pull-request' to automatically label PRs with 'breaking-change' based on conventional commit parsing; 'follow-up-docs-issue' to manage documentation requirements by adding/removing labels and creating follow-up issues in the docs repo upon PR merge; 'report-ci-failure' to create or update GitHub issues tracking CI workflow failures; and 'schedule' to automatically unassign stale issues from non-members. It also includes the core 'docs-version' logic for determining the correct version bump workflow and corresponding unit tests.

cyborg · high confidence

Introduce DDL event recording for database, table, view, and flow operations

The metadata layer now emits structured lifecycle events for Data Definition Language (DDL) operations. New event types have been added for creating, altering, and dropping databases, tables, views, and flows. These events capture key details such as operation intent (e.g., \create\_if\_not\_exists\, \or\_replace\), specific options (like TTL for databases), and resource identifiers (table IDs, flow IDs, view IDs) to provide a comprehensive audit trail of schema changes.

src/common/meta/src/ddl/event · high confidence

Introduce Datanode Region Server and Query Runtime Stream

The datanode now exposes a dedicated Region Server component that handles region-level requests and a new Query Runtime Stream to manage query result batches produced on the datanode's query runtime. This change adds the core infrastructure for region lifecycle events (registration/deregistration), query result streaming with metrics and ordering support, and partitions expression fetching from the table route, laying the groundwork for distributed query execution and region management.

src/datanode/src · high confidence

Introduce DatanodeWorkloadType enum for workload classification

Added a new \DatanodeWorkloadType\ enum in the \src/common/workload\ library to define datanode workload capabilities, currently supporting only the \Hybrid\ mode. This change introduces serialization/deserialization support, conversion methods between the enum and integer representations, and a \sanitize\_workload\_types\ helper that defaults to \Hybrid\ if no workload types are specified. Tests verify the correct mapping between the enum and its integer values.

src/common/workload · high confidence

Introduce DistAnalyzeExec for distributed query diagnostics

The query engine now includes a new \DistAnalyzeExec\ physical plan node to support distributed \EXPLAIN ANALYZE\ operations. This node collects and exposes execution metrics from the underlying plan, including specific metrics from \MergeScanExec\ stages, and outputs them in a structured format (JSON) for HTTP streaming and verbose analysis. This enables users to debug and understand the performance characteristics of distributed queries by seeing per-stage and per-node timing and resource usage.

src/query · high confidence

Introduce Flownode with dual-engine flow execution

Users can now run flows using a dedicated Flownode service that supports both streaming and batching execution modes. This change adds a new gRPC service for flow management, including creating, removing, and inserting data into flows, along with a heartbeat mechanism to report flow statistics to the metasrv. The batching mode allows flows to be executed as time-window-aware queries triggered by new data, with configurable options for query timeouts, incremental reads, and read preferences. A DataFusion-based optimizer is integrated to improve query performance through rules like group-by validation and expression simplification.

src/flow/src · high confidence

Introduce GreptimeDB SQL parser and dialect

The \src/sql\ module now provides a dedicated SQL parser and dialect implementation for GreptimeDB. This includes a custom \GreptimeDbDialect\ that supports specific identifier characters (like \\#\ and \@\), enables trailing commas, and allows struct literals. The parser exposes a comprehensive set of statement parsers (CREATE, ALTER, DROP, INSERT, SELECT, etc.) and re-exports core AST types from \sqlparser\ to standardize the SQL interface.

src/sql · high confidence

Introduce Import V2 with resume, parallelism, and progress reporting

The CLI now supports a new Import V2 workflow that allows users to import data from packed snapshots with resume capability, parallel task execution, and progress reporting. Users can specify the number of concurrent import tasks (1–64) via the \--task-parallelism\ flag to speed up large imports, and the import process can be resumed from a previous state if interrupted, with the state file path optionally overridden using \--state-path\. Progress is reported in real-time, with the mode controllable via the \--progress\ flag. The import command also supports dry-run verification, schema filtering, and proxy configuration.

_src/cli/src/data/import\v2 · high confidence

Introduce JSON2 type with configurable type hints and auto-expansion limits

The datatypes crate now supports the JSON2 data type, which allows users to define column-level type hints for specific JSON subpaths to improve query performance and storage layout. It also introduces a configurable limit on the number of unhinted JSON paths that are automatically expanded into Arrow fields, defaulting to 100 to prevent unbounded schema growth. These settings are stored in column schema metadata and enforced during JSON encoding and decoding.

src/datatypes/src · high confidence

Introduce JSON2 vector storage and builder implementation

This change adds the core implementation for the new JSON2 data type within the datatypes library. It introduces a new physical storage layout (v2) that supports explicit type hints, auto-expansion of paths, and a remainder field for unstructured data, replacing the previous legacy merging approach. The diff provides the \JsonVectorBuilder\ to construct these vectors, \JsonArray\ to read and project them, and \variant.rs\ to handle encoding/decoding of JSON values into the Parquet Variant format, enabling more efficient querying and storage of JSON data.

src/datatypes/src/vectors/json · high confidence

Introduce Kafka-backed remote WAL log store

This change adds a new Kafka-based implementation of the remote Write-Ahead Log (WAL) log store. It introduces a \ClientManager\ to handle Kafka client lifecycle and SASL/TLS configuration, an \OrderedBatchProducer\ for batching and compressing WAL entries before sending them to Kafka topics, and a \Consumer\ that fetches records while respecting the WAL index for replay. The \KafkaLogStore\ coordinates these components, including a \PeriodicOffsetFetcher\ to track topic stats and a \PeriodicTopicStatsReporter\ to expose metrics, enabling the system to use Kafka as a durable, remote storage backend for WAL data.

src/log-store/src/kafka · high confidence

The \src/puffin\ module now implements the Puffin file format, providing asynchronous readers and writers for managing blobs and directories. A key capability is the support for LZ4 compression on the file footer payload, which reduces I/O overhead when reading metadata. The implementation includes a \PuffinManager\ trait for unified access, a \PuffinMetadataCache\ for caching file metadata, and a \PartialReader\ for efficient range-based blob retrieval.

src/puffin · high confidence

Introduce RaftEngine-based log store and KV backend

The log store and metadata KV backend are now implemented using the RaftEngine library, replacing the previous implementation. This change introduces a new \RaftEngineLogStore\ that manages WAL operations (append, replay, GC, and periodic sync) via RaftEngine's engine, and a \RaftEngineBackend\ that implements the \KvBackend\ and \TxnService\ traits for metadata storage. Users benefit from RaftEngine's durability guarantees, configurable recovery modes, and background garbage collection tasks.

_src/log-store/src/raft\engine · high confidence

Introduce TableFlownodeSetCache for flow node mapping

Added a new \TableFlownodeSetCache\ component that caches the mapping between source tables and the specific flow nodes (peers) responsible for processing their partitions. This cache automatically initializes from the key-value backend and stays consistent by handling \CreateFlow\, \DropFlow\, and \FlowNodeAddressChange\ events, ensuring that flow execution routes are resolved locally without repeated remote lookups.

src/common/meta/src/cache/flow · high confidence

Introduce batching mode engine for scheduled flow queries

A new batching mode engine has been added to the flow system, enabling scheduled, time-window-aware flow queries that execute as discrete tasks rather than continuous streams. This implementation introduces a dedicated execution engine, a frontend client for offloading computation, and a checkpointing system that supports incremental reads with automatic fallback to full snapshots when necessary. Users can now define flows with an \EVAL INTERVAL\ to control execution frequency, and the engine handles dirty time window tracking, state persistence, and query plan rewriting to ensure correctness across scheduled runs.

_src/flow/src/batching\mode · high confidence

Introduce common WAL configuration and options module

The \src/common/wal/src\ module now centralizes WAL configuration and region-level options. It defines \MetasrvWalConfig\ and \DatanodeWalConfig\ enums supporting \raft\_engine\, \kafka\, \noop\, and \experimental\_object\_store\ providers, with the datanode defaulting to \raft\_engine\. Region creation now carries \WalOptions\ (serialized as JSON with keys like \wal.provider\ and \wal.kafka.topic\) so the metasrv can pass provider-specific details—such as the Kafka topic and an initial pruned entry ID—to the datanode. The module also adds Kafka client error-to-retry-hint mapping, IPv4-only broker endpoint resolution, and test utilities for Kafka integration tests.

src/common/wal/src · high confidence

Introduce common error handling with retry hints and masked internal messages

The \src/common/error\ crate now provides a unified error framework that decouples retryability from status codes via a new \RetryHint\ enum and \ErrorExt\ trait, allowing callers to determine if an operation should be retried based on transient conditions (e.g., I/O errors) rather than just the error category. Additionally, the \output\_msg\ implementation in \ErrorExt\ now masks \Internal\ and \Unknown\ status codes from end-user output to prevent leaking sensitive internal details, while still exposing root causes for other error types. The crate also defines a comprehensive \StatusCode\ enum covering SQL, catalog, storage, server, auth, flow, and trigger domains, and exposes error details (code, message, and retry hint) via HTTP headers for gRPC/HTTP interoperability.

src/common/error/src · high confidence

Introduce common event recorder for unified event persistence

A new common event recorder library has been added to \src/common/event-recorder\, providing a unified mechanism to record and persist system events (such as procedure triggers, DDL operations, and admin function executions) into a shared \events\ system table. This library defines the canonical schema for the event table—including columns for event type, payload, timestamp, and various contextual dimensions like catalog, schema, and procedure state—and implements the core recording logic with batched, asynchronous flushing. It also introduces a configurable \EventTypeFilter\ to allow users to restrict which event types are recorded, and provides a \PersistentEventContext\ to store structured metadata about why an event was triggered.

src/common/event-recorder · high confidence

Introduce common-telemetry crate for unified logging, metrics, and distributed tracing

The new \common-telemetry\ crate centralizes observability infrastructure, providing a unified API for logging, metrics, and distributed tracing. It introduces configurable logging with file retention, dynamic log level reloading, and a \slow!\ macro for performance monitoring. Metrics are now exposed via Prometheus-compatible formats, including support for remote write. Distributed tracing is enabled through OpenTelemetry integration, featuring W3C trace context propagation across services, configurable sampling rules based on protocol and request type, and a runtime switch to enable or disable trace collection without restarting the application.

src/common/telemetry/src · high confidence

Introduce distributed table support and refined filter pushdown semantics

The table module now includes a \DistTable\ implementation that wraps tables with a \DummyDataSource\, enabling the system to handle distributed table references (though currently returning an unsupported error for stream retrieval). Additionally, the module introduces a \FilterPushDownType\ enum (with \Unsupported\, \Inexact\, and \Exact\ states) to replace the previous boolean filter pushdown flag, allowing the query engine to distinguish between filters that can be fully enforced by the table and those that only help minimize data retrieval. This change also adds a \DfTableProviderAdapter\ to bridge the internal table representation with DataFusion's \TableProvider\ interface, including support for dictionary encoding hints on primary key columns.

src/table · high confidence

Introduce dual-engine flow execution with DataFusion streaming adapter

The flow adapter now manages a dual-engine architecture that supports both streaming and batching execution modes. A new \FlowDualEngine\ coordinates between the \StreamingEngine\ and \BatchingEngine\, handling flow creation, dropping, flushing, and state reporting. The streaming engine has been refactored to use a stateless DataFusion execution model, where flow plans are executed against transient \MemTable\ providers for each input batch. This includes logic to rewrite logical plans to carry source timestamps through filters and projections, ensuring correct data lineage. The adapter also introduces a \ManagedTableSource\ to resolve table schemas and metadata from the catalog manager, and provides utilities to convert between internal column schemas and the API protobuf format.

src/flow/src/adapter · high confidence

Introduce event recording and configurable memory limits in the frontend

The frontend now supports recording lifecycle events (such as admin actions) by ingesting them into a dedicated table via a new \EventRecorder\ and \EventHandlerImpl\. Additionally, the frontend instance and server handlers now enforce a configurable memory limit for concurrent write requests (\max\_in\_flight\_write\_bytes\), applying a specified policy (wait or fail) when the limit is exhausted to prevent resource exhaustion.

src/frontend/src · high confidence

Introduce fuzz testing framework for GreptimeDB

Added a new \tests-fuzz\ crate that provides a fuzz testing infrastructure for GreptimeDB, including generators for SQL operations like CREATE TABLE, ALTER TABLE, INSERT, and SELECT, as well as partition management. The framework uses \cargo-fuzz\ with a nightly Rust toolchain and includes configuration templates, error handling, and intermediate representation (IR) definitions to drive randomized SQL generation and validation against the database.

tests-fuzz · high confidence

Introduce global runtime management and weighted workload scheduler

The \src/common/runtime\ crate now provides a centralized runtime infrastructure, including global runtimes for query, ingestion, and compact operations, along with a new weighted workload scheduler that allows configuring relative poll shares for query versus write tasks. It also introduces CPU resource constraints via a \ThrottleableRuntime\ that rate-limits task polls based on priority, and exposes runtime metrics (thread counts, scheduler stats) via Prometheus.

src/common/runtime · high confidence

Introduce incremental read checkpoints and hardened state management for batching flows

This change adds new source files (ckpt.rs, inc.rs, test.rs) to the batching mode task module, implementing support for incremental read checkpoints and exact sequence-range scans. Users will see improved correctness and performance for flows using incremental mode, as the system now classifies query failures (such as stale snapshot fences or incremental query failures) to decide whether to fall back to full snapshots or continue incremental processing. The implementation also introduces logic to verify source table capabilities (like \preserve\_row\_sequence\) before enabling incremental reads, and provides hooks for rewriting completed plans to ensure data consistency during repair and checkpoint advancement.

_src/flow/src/batching\mode/task · high confidence

Introduce logical table reconciliation procedure

Adds a new reconciliation procedure for logical tables, implemented as a state machine in the meta service. The procedure validates that logical tables share the same physical table and schema, checks region metadata consistency, and then reconciles the system by creating missing regions and updating table metadata to match the physical table's column definitions. It also includes cache invalidation to ensure consistency after updates.

_src/common/meta/src/reconciliation/reconcile\_logical\tables · high confidence

Introduce memory profiling library with flamegraph support

Added a new \mem-prof\ crate that enables memory profiling via jemalloc, including the ability to dump profiles, generate pprof data, and directly produce flamegraph visualizations. The library exposes functions to activate/deactivate heap profiling and control the \prof.gdump\ setting, while providing a consistent error model. Note that memory profiling is currently unsupported on Windows, where all profiling functions return a 'not supported' error.

src/common/mem-prof/src · high confidence

Introduce native histogram support and query engine refactoring

The common-query module now includes a runtime model for Prometheus native histograms, providing schema definitions, validation, and encoding logic for the new histogram data type. Query results are restructured into an Output object that separates data from metadata (such as execution cost and plan details). Additionally, the module introduces a ColumnarValue abstraction to bridge scalar and vector representations, and adds a PromqlAnnotationCollector to track and return warnings and informational messages during query execution.

src/common/query/src · high confidence

Introduce new MetaClient library for cluster metadata operations

A new \meta-client\ library has been added to provide a unified interface for interacting with the metasrv cluster. This client supports multiple node roles (Frontend, Datanode, Flownode) and exposes capabilities for leader discovery, heartbeat management, distributed key-value storage, and executing distributed procedures (such as DDL tasks). It includes configurable gRPC channel settings, heartbeat keep-alive parameters, and a builder pattern for flexible initialization, enabling components to reliably manage cluster state and coordinate distributed operations.

src/meta-client/src · high confidence

Introduce new data types and type conversion logic

The datatypes module now includes implementations for Binary, Boolean, Date, Decimal128, Dictionary, Duration, Interval, JSON, List, and Null types, each providing specific vector builders and Arrow type mappings. Additionally, a new cast module has been added to handle value conversions between these types, supporting both strict and non-strict casting modes to ensure data integrity during type transitions.

src/datatypes/src/types · high confidence

Introduce new log-store crate with Kafka, Raft Engine, and Object Store WAL implementations

The \src/log-store\ crate has been introduced to centralize log storage logic, providing three distinct backend implementations: a Kafka-based log store (\kafka.rs\), a Raft Engine-based log store (\raft\_engine.rs\), and a durable Object Store WAL (\object\_store\_wal.rs\) that stores entries as immutable objects with chain-based recovery. The crate exposes a unified \ObjectStoreLogStore\ type, defines specific error handling in \error.rs\, and tracks operational metrics in \metrics.rs\.

src/log-store/src · high confidence

Introduce object store-backed Write-Ahead Log (WAL) with durable and enqueued modes

The log store now supports a new object store-based WAL implementation, adding new files for batch accumulation, cataloging, object formatting, and I/O operations. This change introduces two acknowledgement modes: 'durable', where an append returns only after the object is persisted and indexed, and 'enqueued', where it returns upon admission and the object is created in the background. The implementation includes recovery logic to rebuild the object catalog from existing objects, ensuring that entry IDs remain monotonic and unique across object sequences.

_src/log-store/src/object\_store\wal · high confidence

Introduce object-store library with OpenDAL-backed storage and caching

The new \src/object-store\ library provides a unified configuration and factory for object storage backends (S3, GCS, OSS, Azure Blob, HDFS, MySQL, and local filesystem) powered by OpenDAL. It introduces configurable read and write caching for remote stores, a secure filesystem sandbox for untrusted paths, and an \ObjectStoreManager\ to support per-table storage selection. The library also includes Prometheus metrics, retry logic, and compatibility layers for atomic writes on HDFS.

src/object-store · high confidence

Introduce relation type representation for dataflow

Added a new \relation.rs\ module in the flow representation layer that defines \Key\ and \RelationType\ structs. These structures model the schema of data relations by tracking column types, primary keys, time indices, and automatically added columns, providing the foundational type system for the new dataflow framework.

src/flow/src/repr · high confidence

Introduce remote WAL logical pruning capability

MetaSrv now supports pruning remote WAL data, allowing users to reclaim storage by deleting obsolete records from Kafka. This feature introduces a new \WalPruneManager\ that periodically triggers pruning procedures, tracks running tasks to prevent duplicates, and limits concurrency via a semaphore. The implementation supports a 'logical delete' mode that updates metadata without invoking Kafka's \DeleteRecords\ API, and includes utilities to determine prunable entry IDs based on region heartbeats and to update topic metadata accordingly.

_src/meta-srv/src/procedure/wal\prune · high confidence

Introduce servers module documentation and benchmarking infrastructure

The \src/servers\ crate now includes an \AGENTS.md\ guide detailing the module's architecture, including protocol translation, handler boundaries, and testing procedures. Additionally, a new benchmarking suite has been added under \src/servers/benches/\ to measure performance for Prometheus decoding, Loki label parsing, and HTTP output conversion.

src/servers · high confidence

Introduce store-api crate with storage engine and WAL abstractions

The \src/store-api\ crate is introduced to provide a unified API layer for the storage engine. It defines the \LogStore\ trait for Write-Ahead Log operations, including support for multiple providers such as RaftEngine, Kafka, and Object Store. The crate also introduces the \DataSource\ trait for retrieving data as record batch streams, the \RegionEngine\ trait for managing region lifecycle and role states (including staging and downgrading), and constants for the metric engine. This establishes the foundational interfaces for region handling, WAL management, and data access within the store layer.

src/store-api · high confidence

Introduce structured log query API with flexible filtering and expression support

This change introduces the \log-query\ crate, providing a new structured API for querying logs. Users can now define queries using a \LogQuery\ struct that supports time range filtering via \TimeFilter\ (handling dates, timestamps, and spans), pagination with \limit\ and \offset\, and column selection. The API enables complex filtering through a \Filters\ enum that supports AND, OR, and NOT logical combinations, as well as column-level filters. Additionally, it allows post-filter processing via \LogExpr\, supporting scalar functions, aggregation functions (\AggFunc\), binary operations, aliases, and decomposition of values (e.g., JSON/CSV). Error handling is standardized with specific error types for invalid time filters, date formats, and span formats.

src/log-query · high confidence

Introduce structured reconciliation procedures and event recording

The meta service now uses a new \ReconciliationManager\ to coordinate background reconciliation for catalogs, databases, physical tables, and logical tables. Each scope has its own procedure (e.g., ReconcileCatalogProcedure, ReconcileDatabaseProcedure, ReconcileTableProcedure, ReconcileLogicalTablesProcedure) that runs as a persistent, resumible procedure with state serialization and lock keys. The manager registers these procedure loaders and exposes entry points such as reconcile\_table, reconcile\_database, reconcile\_physical\_table, and reconcile\_logical\_tables, which dispatch to the appropriate procedure based on table type. Reconciliation events are now recorded via a new event module (event.rs) with stable event types (reconcile\_catalog, reconcile\_database, reconcile\_logical\_tables, reconcile\_table) and structured payloads that capture submission details (resolve strategy, fail-fast, parallelism) and results (success/failure counts). The manager also integrates with cache invalidation and node management to ensure metadata consistency across the cluster.

src/common/meta/src/reconciliation · high confidence

Introduce the File Engine for querying external files

Adds a new file engine that enables querying external data files (such as CSV) stored in object storage or local filesystems. This engine implements the region lifecycle (create, open, drop) and query execution, allowing users to scan external files as if they were database regions. It includes configuration for file options, manifest management for region metadata, and query capabilities with projection and filter pushdown to the underlying file format.

src/file-engine/src · high confidence

Introduce the common procedure framework

Adds the core \procedure\ library in \src/common/procedure/src\, providing the foundational traits, structures, and execution engine for managing background tasks. This includes the \LocalManager\ for executing procedures, a \ProcedureStore\ for persisting state, a \PoisonStore\ for handling failure isolation, and a \KeyRwLock\ for fine-grained concurrency control. The framework supports procedure states (running, done, failed, poisoned), retry logic with configurable limits, and an event system (\ProcedureEvent\) for recording lifecycle changes and errors.

src/common/procedure/src · high confidence

Introduce triggers and flow flushing capabilities in the operator

The operator crate now supports creating triggers with webhook notification channels and flushing flow state to flownodes. A new \PendingRowsBatcher\ trait decouples prepared table writes from the inserter, and the \FlowServiceOperator\ implements the \flush\ handler to coordinate flow updates across nodes.

src/operator · high confidence

Introduces Arrow extension types for histograms and JSON2 data

The datatypes module now includes dedicated Arrow extension types to identify native-histogram struct columns (tagged as 'greptime.histogram') and JSON2 columns. This enables precise schema alignment for JSON2 arrays and supports the JSON2 v2 physical layout, allowing readers to distinguish these specific column types by extension metadata rather than field names.

src/datatypes/src/extension · high confidence

Introduces WAL index collection, encoding, and iteration infrastructure

This change adds the core components for managing Write-Ahead Log (WAL) indexes within the Kafka log store. It introduces the \IndexCollector\ trait and \GlobalIndexCollector\ implementation, which manages index entries across multiple Kafka providers and runs a background \CollectionTask\ to periodically persist indexes to object storage. It also adds \IndexEncoder\ and \JsonIndexEncoder\ to serialize these indexes using delta encoding for efficient storage, and \RegionWalIndexIterator\ implementations (\RegionWalRange\, \RegionWalVecIndex\, \MultipleRegionWalIndexIterator\) to iterate over WAL entry IDs for replay or consumption. These files provide the specific mechanics for collecting, encoding, and iterating over WAL indexes, supporting the broader feature of WAL index management.

src/log-store/src/kafka/index · high confidence

Introduces a new cache module with registry builders and error handling

The \src/cache\ directory now contains the core implementation for the system's caching layer, including a new \Error\ enum for handling cache retrieval failures and a \lib.rs\ that defines cache registry builders. These builders construct specific caches for datanodes (schema, table ID-to-schema) and fundamental components used by both frontend and datanode (table info, name, route, flownode set, view info, and schema caches), as well as a composite registry that adds table and partition info caches. This change establishes the foundational caching infrastructure and configuration constants (capacity, TTL, TTI) used across the application.

src/cache · high confidence

Introduces a structured heartbeat response handler with parallel region operations

The datanode now uses a centralized \RegionHeartbeatResponseHandler\ to process batch instructions for regions, including opening, closing, flushing, upgrading, downgrading, and entering staging states. This handler supports configurable parallelism for opening and upgrading regions to improve performance and includes a \TaskTracker\ to manage and monitor the lifecycle of long-running asynchronous tasks associated with these region operations.

src/datanode/src/heartbeat · high confidence

Introduces common CPU profiling library with Unix support

Adds a new \src/common/pprof\ crate that provides a unified CPU profiling utility. On Unix systems, it leverages the \pprof\ crate to capture performance data, supporting output in text, flamegraph, and protobuf formats. For non-Unix platforms, it provides a dummy implementation that returns an unsupported error, ensuring cross-platform compatibility without runtime failures.

src/common/pprof · high confidence

Introduces dedicated caches for table metadata, schema, routes, and views

The system now uses specialized caches to store and invalidate table metadata, schema names, table routes, and view information. This change improves performance and consistency by caching frequently accessed metadata and ensuring that stale data is cleared when the underlying information changes.

src/common/meta/src/cache/table · high confidence

Introduces group-level repartition procedure states

The meta-srv repartition group procedure now implements a structured state machine to manage the lifecycle of a repartition group. New states include RepartitionStart to validate and capture source/target region routes, SyncRegion to synchronize new regions from a central region, EnterStagingRegion to transition regions into a staging state (including flushing pending deallocations), RemapManifest to update manifest paths, ApplyStagingManifest to apply the new partition expressions, and UpdateMetadata to persist changes and exit staging. This adds the concrete orchestration logic for group repartitioning within the meta-srv.

src/meta-srv/src/procedure/repartition/group · high confidence

Introduces region migration procedure in meta-srv

Adds a new region migration procedure framework to the meta-srv service, including a manager to orchestrate migration tasks and a state machine with steps for opening candidate regions, flushing and downgrading leader regions, upgrading candidates, and cleaning up old regions. This enables the system to move region leaders between datanodes while maintaining metadata consistency and handling abort/rollback scenarios.

_src/meta-srv/src/procedure/region\migration · high confidence

Introduction of JSON2 data type implementation

This change introduces the internal implementation for the new 'JSON2' data type, adding the \src/datatypes/src/json/value.rs\ file which defines core structures like \JsonNumber\ and \JsonVariant\. This new type provides a distinct storage and processing path for JSON data, separate from the existing JSON handling, enabling features such as optimized writes, nested path fallback reads, and specific encoding of variant payloads as jsonb.

src/datatypes/src/json · high confidence

Introduction of ReadPreference enum for frontend routing

The session module now defines a ReadPreference enum, currently supporting only the Leader mode, which determines whether frontend route operations read from the region leader or follower. This change introduces the foundational type for controlling read routing behavior, with the Leader mode set as the default.

src/common/session · high confidence

Introduction of common plugin constants for execution metrics

A new \plugins\ crate has been added to the common layer to centralize shared utilities and constants. This change introduces specific constants for Greptime execution metrics, including a preserved prefix (\greptime\exec\\) and keys for read and write costs, making these values available for use across other components like the frontend and datanode.

src/common/plugins/src · high confidence

Leader node metadata caching for improved read performance

The meta-srv service now includes a dedicated in-memory cache for metadata on the leader node. This new \LeaderCachedKvBackend\ intercepts read requests (specifically exact key lookups) and serves them from the local cache if available, falling back to the underlying store only on misses or for range queries. The cache is automatically populated during initialization and invalidated on mutations, reducing latency for metadata reads on the leader while ensuring followers continue to read directly from the persistent store.

src/meta-srv/src/service/store · high confidence

Log directory size retention policy

Added a new file retention mechanism for the telemetry logging system that manages disk usage by enforcing limits on the total size of log files and the maximum number of log files kept. This feature automatically prunes older log files when the configured size or file count thresholds are exceeded, helping to prevent unbounded growth of the log directory.

src/common/telemetry/src/logging · high confidence

New CI and quality-assurance scripts for Rust toolchain, licensing, and code style

The repository now includes several new shell and Python scripts to enforce build consistency and code quality. \check-builder-rust-version.sh\ validates that the Rust toolchain in the dev-builder Docker image matches the pinned version in \rust-toolchain.toml\, supporting both stable and nightly channels. \check-enterprise-license.py\ (with its test suite) ensures that files gated by \\#\[cfg(feature = "enterprise")\]\ are correctly listed in the enterprise license configuration, preventing license mismatches. Additional scripts \check-snafu.py\ and \check-super-imports.py\ detect unused SNAFU error variants and improper \use super::\ imports, respectively. Certificate generation helpers (\generate-etcd-tls-certs.sh\, \generate\_certs.sh\) and a dashboard asset fetcher (\fetch-dashboard-assets.sh\) are also added to support testing and build workflows.

scripts · high confidence

New CLI binaries for query performance and regression testing

The \greptime\ binary now includes built-in support for query performance and regression testing via two new executable targets: \query\_perf\_fixture\ and \query\_regression\_runner\. These tools allow users to generate performance fixtures and run regression scenarios directly from the command line, facilitating easier validation of query behavior and performance characteristics.

src/cmd/src/bin · high confidence

New CLI commands for metadata inspection, modification, and repair

The CLI now includes new subcommands to manage metadata directly: \metadata get\, \put\, and \del\ allow users to retrieve, set, and remove metadata keys and tables, while \metadata repair\ provides tools to fix logical table metadata and resolve partition column mismatches. Additionally, \metadata snapshot\ enables saving, restoring, and inspecting metadata snapshots, supporting both local file systems and object stores (such as S3) for storage backends.

src/cli/src/metadata · high confidence

New CLI export and import commands with V2 snapshot support

The CLI now includes new data export and import capabilities, introducing a V2 snapshot-based workflow alongside the existing export/import commands. The new \export-v2\ and \import-v2\ commands support JSON-based schema exports, manifest-based snapshot management, and resume capabilities for interrupted operations. Users can export and import data to/from multiple storage backends including S3, OSS, GCS, Azure Blob, and local filesystems. The V2 export supports parallel database and table operations, time-range filtering, and experimental metric export with packed Parquet objects. Progress reporting is available with interactive bars on TTYs and log-based output for non-interactive runs. The existing export and import commands continue to support schema, data, or both targets with configurable parallelism and retry settings.

src/cli/src/data · high confidence

New CLI export-v2 snapshot commands and chunked export engine

The CLI now includes a new export-v2 feature set with commands to create, list, verify, and delete data snapshots. Snapshots are stored as structured directories containing a manifest, schema definitions, and chunked data files (Parquet or CSV) written to local or remote object stores (S3, GCS, OSS, Azure Blob). The export engine splits data into time-bounded chunks for parallel export, supports resuming interrupted exports, and includes integrity verification to detect missing or corrupted files.

_src/cli/src/data/export\v2 · high confidence

New CLI tool for table metadata benchmarking

The CLI now includes a \bench\ command that allows users to benchmark table metadata operations (create, get, rename, delete) against various key-value backends. The tool supports connecting to etcd, PostgreSQL, or MySQL (when respective features are enabled) to measure average operation costs, helping users evaluate metadata performance under different storage configurations.

src/cli/src · high confidence

New CLI tools for repairing logical table metadata

A new CLI tool has been added to the \src/cli/src/metadata/repair\ directory to help administrators fix inconsistencies in logical table metadata. This includes a \RepairPartitionColumnCommand\ that detects and reconciles mismatches between partition columns defined in table info and those stored in region routes, with support for dry-run mode and update limits. The tool also introduces helper modules (\alter\_table.rs\ and \create\_table.rs\) to generate and execute alter and create table expressions for repairing table structures across regions.

src/cli/src/metadata/repair · high confidence

New DDL utility modules for table and region metadata management

The \src/common/meta/src/ddl/utils\ directory now includes new utility modules (\raw\_table\_info.rs\, \region\_metadata\_lister.rs\, \table\_id.rs\, and \table\_info.rs\) that provide core infrastructure for Data Definition Language operations. These additions introduce functions to build and update physical table metadata, batch-fetch table IDs and info values by name or ID, and asynchronously collect region metadata from datanodes, laying the groundwork for table reconciliation and logical table procedures.

src/common/meta/src/ddl/utils · high confidence

New Dockerfiles for CI, fuzz testing, and Kafka WAL helper

This change introduces three new Dockerfiles in the docker/ci/ubuntu directory to support different CI and operational needs. The main Dockerfile builds a GreptimeDB image based on Ubuntu 22.04, allowing the binary name to be configured via the TARGET\_BIN argument and enabling heap profiling by default through the MALLOC\_CONF environment variable. A second Dockerfile, Dockerfile.fuzztests, provides a base for running fuzz tests using a specified binary path. A third Dockerfile, Dockerfile.kafka-wal-helper, creates a lightweight image for the Kafka WAL helper service based on Ubuntu 24.04, exposing port 8080 and running the helper as the entrypoint.

docker/ci/ubuntu · high confidence

New Frontend Instance Builder and Protocol Handlers

The frontend instance is now constructed via a dedicated \FrontendBuilder\ that wires up core services like the catalog manager, node manager, and procedure executor. This change introduces new protocol handlers for InfluxDB line writes (with configurable merge modes and timestamp alignment), Jaeger trace queries (with permission checks), and database exports (including an experimental metric export feature). It also adds support for a dashboard table schema, entity graph derivation from table options, and packed database imports with robust cancellation handling.

src/frontend/src/instance · high confidence

New HTTP admin API endpoints for maintenance, procedure control, and system status

The meta-srv admin service now exposes a suite of new HTTP endpoints for operational control and monitoring. Operators can manage maintenance and recovery modes via dedicated status, enable, and disable endpoints, and can pause or resume the procedure manager to control background tasks. Additional endpoints provide visibility into the cluster leader, datanode heartbeats (with optional address filtering), and node leases, while the sequencer API allows setting the next table ID (restricted to recovery mode) and peeking the current sequence value. A standard health check endpoint is also available to verify service readiness.

src/meta-srv/src/service/admin · high confidence

New KV-backed catalog manager with configurable builder and caching

The catalog subsystem now uses a new \KvBackendCatalogManager\ backed by a \KvBackendRef\ and a \LayeredCacheRegistry\. A new builder (\KvBackendCatalogManagerBuilder\) allows configuring the manager with optional procedure and process managers, and extra information schema table factories. The manager initializes a \PartitionRuleManager\, a \TableMetadataManager\, and a \SystemCatalog\ that includes caches for catalogs and pg\_catalog, plus providers for information\_schema, pg\_catalog, and a numbers table. A new \CachedKvBackend\ wraps the underlying KV store with configurable capacity, TTL, and TTI, invalidating cache entries on writes. Table lookups are cached via a new \TableCache\ (a \CacheContainer\ for \TableName\ to \TableRef\) that initializes from \TableInfoCache\ and \TableNameCache\ and invalidates on \CacheIdent::TableName\. The manager also exposes an injectable entity-graph provider for computed semantic tables.

src/catalog/src/kvbackend · high confidence

New Kafka record serialization and range utility helpers

Added \range.rs\ and \record.rs\ to the Kafka log-store utilities. \record.rs\ introduces the \Record\ struct and \RecordType\ enum to handle serialization of log entries into Kafka records, including support for splitting large entries into multipart records (First, Middle, Last) and converting them to/from \KafkaRecord\ objects. \range.rs\ provides utility iterators \ConvertIndexToRange\ and \MergeRange\ to convert entry ID sequences into size-based ranges and merge overlapping or close ranges within a window size, likely to optimize batch processing or offset calculations.

src/log-store/src/kafka/util · high confidence

New administration functions for table and region lifecycle management

The \src/common/function\ module introduces a suite of new SQL administration functions to manage the database's internal storage lifecycle. Users can now invoke \flush\_table\, \compact\_table\, \flush\_region\, and \compact\_region\ to manually control data persistence and compaction. Additional functions like \gc\_table\ and \gc\_regions\ allow for garbage collection of unused data, while \migrate\_region\ enables moving data between cluster nodes. The module also adds \build\_index\ and \build\_series\_index\ for manual index creation, \discard\_unflushed\ to clear pending writes, and \purge\_table\ (an enterprise feature) to permanently remove dropped tables. These functions are registered in the function registry, with sensitive operations like \discard\_unflushed\ and \purge\_table\ restricted to ADMIN-only execution.

src/common/function · high confidence

New agent skills for GreptimeDB development Docker image builds and fuzz CI failure investigation

Added two new agent skills to the \.agents/skills\ directory. The \greptimedb-development-docker-image\ skill provides a guided workflow for building development-only GreptimeDB Docker images from local debug binaries, including scripts for context collection, binary platform validation, image building (via Docker or Podman), and safe \.env\ management. The \greptimedb-fuzz-ci-failure-investigation\ skill enables agents to diagnose failed GreptimeDB fuzz CI jobs by downloading GitHub Actions logs and artifacts, then correlating them with local source code.

.agents/skills · high confidence

New authentication and permission subsystem

The \src/auth\ crate has been introduced to centralize authentication and authorization logic. It adds support for multiple authentication methods, including MySQL native password, PostgreSQL SCRAM-SHA-256, and bearer token authentication. The system also introduces a granular permission model with named actions (e.g., \promql.query\, \log.write\) and table-level access control, allowing users to define read/write permissions for specific catalogs and schemas. User providers can now be configured via static configuration or file watching, and the subsystem includes robust error handling for authentication failures and permission denials.

src/auth/src · high confidence

New catalog management and system schema infrastructure

The \src/catalog\ crate now provides the core catalog management layer, introducing a \CatalogManager\ trait and a \KvBackendCatalogManager\ implementation for persistent storage, alongside a \MemoryCatalogManager\ for testing. This change adds a \ProcessManager\ to track and kill running queries across catalogs, exposes catalog and process metrics via Prometheus, and establishes the \SystemSchemaProvider\ and \SystemTable\ traits to power \information\_schema\ and \pg\_catalog\ system tables. Additionally, it includes a \DistributedInformationExtension\ to support distributed inspection of datanodes and a \DfTableSourceProvider\ to resolve table references and handle view execution within the DataFusion query engine.

src/catalog/src · high confidence

New common base library with foundational types and utilities

The \src/common/base\ module introduces a new set of core types and utilities for the application. This includes \Bytes\ and \StringBytes\ for efficient, type-safe byte and string handling with serialization support, and a \BitVec\ type alias for optimized bit manipulation. It adds a \CancellationHandle\ and \CancellableFuture\ to enable explicit cancellation of asynchronous operations. A \Plugins\ system provides a thread-safe, type-erased registry for storing and retrieving application plugins. The library also defines a \Channel\ enum to identify the protocol or subsystem (e.g., MySQL, gRPC, HTTP) through which a query is received. For resource management, it provides \ReadableSize\ for parsing and displaying human-readable byte sizes, \MemoryLimit\ for configuring memory limits by absolute size or percentage, and \SecretBox\/\SecretString\ for securely handling sensitive data with automatic memory zeroing. Finally, it includes a \RangeReader\ trait and an \AsyncReadAdapter\ for abstracting and adapting byte-range reads from various sources.

src/common/base · high confidence

New common batching library for protocol-independent batch management

A new \src/common/batcher\ library has been added, providing reusable, protocol-independent components for managing timed and size-based batching. It introduces a \PendingWorker\ that groups submissions by key and applies a \FlushPolicy\ (such as the new \TimingFlushPolicy\ which triggers on deadlines or row thresholds) to decide when to flush. The library also includes a \FlushLimiter\ to cap concurrent flushes, a \RequestLimiter\ to limit in-flight original requests, a \WorkerRegistry\ for managing worker senders by key, and a \Notifier\ for best-effort delivery of completion events. This refactors ingestion logic to reuse these shared building blocks.

src/common/batcher · high confidence

New common configuration and file-watching infrastructure

The \src/common/config\ module now provides a unified \Configurable\ trait for loading settings from config files, environment variables, and defaults with a defined precedence, alongside a robust file-watcher that monitors configuration paths (including symlink chains) to support hot-reloading. This change also introduces a \KvBackendConfig\ with sensible defaults for metadata store log size and purge behavior, and establishes \./greptimedb\_data\ as the default data home directory.

src/common/config · high confidence

New common datasource library for file formats and object storage

The \src/common/datasource\ module has been introduced to centralize data source handling. It provides a unified \Format\ enum supporting CSV, JSON, Parquet, and ORC file formats, along with a \CompressionType\ enum for GZIP, BZIP2, XZ, and ZSTD. The library includes a \Lister\ for object store directory and file enumeration, and a \LocalFileAccess\ system that enforces sandboxed access to the local filesystem for SQL operations. Additionally, it introduces \PackedWriter\ and \PackedSnapshot\ components to support incremental database backups and the export of metric snapshots as packed Parquet objects.

src/common/datasource/src · high confidence

New common frontend library for process management and slow query events

This change introduces a new \common/frontend\ library that provides core infrastructure for frontend-related operations. It defines a \FrontendClient\ trait and \MetaClientSelector\ to discover active frontend nodes and execute process management commands (listing and killing processes) via gRPC. Additionally, it introduces a \SlowQueryEvent\ type that structures slow query data—including cost, query text, schema, and Prometheus query details—for recording into the \slow\_queries\ table.

src/common/frontend · high confidence

New common gRPC library with TLS reloading and Arrow Flight support

A new \src/common/grpc\ crate has been introduced to centralize gRPC client infrastructure. It provides a \ChannelManager\ that manages connection pooling and supports dynamic TLS certificate reloading via file watching, ensuring connections update automatically when certificates change. The library includes an Arrow Flight encoder/decoder for serializing RecordBatches (with LZ4 compression support) and a \select\ module that handles the conversion of various data types (such as timestamps, intervals, and decimals) into gRPC values.

src/common/grpc/src · high confidence

New common memory manager with configurable granularity and policies

The \src/common/memory-manager\ crate introduces a generic, semaphore-based memory quota system that subsystems (such as compaction, flush, and index build) can use to share allocation logic while tracking their own metrics. Users can now configure memory acquisition behavior via the \OnExhaustedPolicy\ (wait with a configurable timeout or fail immediately) and control allocation precision using \PermitGranularity\ (1 KB for fine-grained fairness or 1 MB for large operations). The manager supports both limited and unlimited modes, allows dynamic expansion of memory grants via \MemoryGuard\, and provides clear error handling for limits, timeouts, and semaphore issues.

src/common/memory-manager · high confidence

New common time module with Date, Duration, Interval, and Time types

The \src/common/time\ crate now provides a comprehensive set of time-related data types and utilities, including \Date\, \Duration\, \Interval\ (YearMonth, DayTime, MonthDayNano), \Time\, \Timestamp\, \TimestampMillis\, \GenericRange\, and \Timezone\ support. This introduces native support for date arithmetic, interval operations, and timezone-aware formatting and parsing, replacing previous ad-hoc implementations and enabling more robust time handling across the system.

src/common/time · high confidence

New configurable GC scheduler with soft-drop table support

The meta-srv now includes a new garbage collection scheduler that automatically identifies and cleans up regions based on configurable priority scores (SST count and file removal rate) and cooldown periods. It also introduces support for soft-dropped tables (behind the enterprise feature flag), allowing retention and automatic purging of dropped tables. The scheduler is fully configurable via \GcSchedulerOptions\, including concurrency limits, retry policies, and full file listing intervals.

src/meta-srv/src/gc · high confidence

New configurable Kafka WAL settings for Datanode and Metasrv

The Kafka Write-Ahead Log (WAL) configuration is now exposed as explicit, customizable settings for both Datanode and Metasrv components. Users can now configure SASL authentication (Plain, SCRAM-SHA-256, SCRAM-SHA-512) and TLS connections (server CA, client certificates) for secure Kafka communication. Additionally, Datanode-specific options like \auto\_create\_topics\ and \create\_index\ are now configurable, while Metasrv introduces controls for WAL pruning behavior (interval, logical delete mode, parallelism) and region flush triggers. These changes allow operators to fine-tune WAL performance, security, and maintenance intervals directly through configuration files.

src/common/wal/src/config/kafka · high confidence

New configuration structures for Object Store and Raft Engine WAL backends

The \src/common/wal/src/config\ module now exposes dedicated configuration structs for the two supported WAL backends. For the Object Store WAL, \ObjectStoreWalConfig\ allows users to tune storage provider selection, path prefixes, flush intervals, batch sizes, and introduces an \AckMode\ (defaulting to \Durable\, with an \Enqueued\ option for faster, background-persisted appends) along with handling for corrupted segments. For the Raft Engine WAL, \RaftEngineConfig\ exposes settings for file sizing, purge thresholds and intervals, sync behavior, and recovery parallelism. These changes enable users to fine-tune persistence guarantees and performance characteristics for each backend via configuration.

src/common/wal/src/config · high confidence

New datanode development and benchmarking subcommands

The datanode binary now includes several new subcommands for development, testing, and performance analysis: \objbench\ for benchmarking object storage operations, \parquetbench\ for scanning and benchmarking individual Parquet SST files, \scanbench\ for benchmarking region scans with structured JSON output, \parquet-rewrite\ for rewriting Parquet files with different compression/encoding settings, \parquet-meta\ for inspecting Parquet file metadata, and \sst-replace\ for replacing SST files and updating region manifests. These tools are gated behind the \dev-tools\ feature flag and are intended for internal development and diagnostics rather than production use.

src/cmd/src/datanode · high confidence

New datanode selection strategies with load-based and round-robin options

The meta-srv selector module now provides multiple strategies for choosing datanodes, including a new LoadBasedSelector that distributes workloads by computing weights from heartbeat statistics (such as region counts), a RoundRobinSelector for even distribution, and a LeaseBasedSelector for random selection among active nodes. These selectors support peer exclusion via \exclude\_peer\_ids\, workload filtering, and configurable duplication rules, giving users more control over how data nodes are selected for operations.

src/meta-srv/src/selector · high confidence

New event recording for WAL prune, batch GC, region migration, and repartition procedures

The meta-srv now emits structured lifecycle and report events for four internal maintenance procedures, allowing users to track their progress and outcomes via the event recorder. WAL prune events record the topic, prunable entry ID, and latest offset. Batch garbage collection events capture the configured regions, timeout, and detailed reports of deleted files and indexes per region. Region migration events expose the source and destination node IDs and addresses, the trigger reason, and the affected region IDs. Repartition events record the catalog, schema, and table identifiers, along with the submitted intent (source/target partition expressions and timeout) and group topology details.

src/meta-srv/src/event · high confidence

New file format implementations for CSV, JSON, ORC, and Parquet

The \src/common/datasource/src/file\_format\ module now includes dedicated implementations for CSV, JSON, ORC, and Parquet file formats. CSV support adds configurable options for delimiters, header handling, strict validation, and skipping bad records. JSON and CSV exports support customizable date, time, and timestamp formats. ORC support is implemented via the \datafusion-orc\ crate, and Parquet support includes a new \DefaultParquetFileReaderFactory\ and a packed reader for efficient concurrent reads. These changes provide the foundational file format handling required for COPY operations.

_src/common/datasource/src/file\format · high confidence

New gRPC protocol support for table alterations and schema management

The \src/common/grpc-expr\ module has been introduced to translate gRPC \AlterTableExpr\ and \CreateTableExpr\ messages into internal table requests. This enables users to modify table structures via the gRPC interface, including adding columns with specific placement (first or after existing columns), modifying column types, and managing index configurations such as full-text, inverted, and skipping indexes. The module also supports setting and unsetting table options and JSON settings, providing a comprehensive programmatic interface for schema evolution alongside the existing SQL capabilities.

src/common/grpc-expr · high confidence

New heartbeat instruction handlers for region lifecycle and maintenance operations

The datanode now handles a comprehensive set of new heartbeat instructions for managing region states and maintenance tasks. This includes batch operations to open and close regions, as well as specific handlers for downgrading regions (which gracefully flushes data before converting a leader to a follower), entering staging states for repartitioning, and applying staging manifests. Additionally, new handlers support garbage collection (GC) of regions and file references, allowing the system to clean up unused data, and provide mechanisms to flush regions synchronously or asynchronously with configurable error strategies.

src/datanode/src/heartbeat/handler · high confidence

New internal inspection tables and remote dynamic filter support in region server

The region server now exposes four new internal inspection tables (SST manifest, SST storage, SST index metadata, and region info) via a new catalog implementation, allowing users to query detailed storage and region state. Additionally, the server now supports applying remote dynamic filters during scans, including registration, update handling, and metrics tracking, to optimize query performance by pushing filter conditions to the data node.

_src/datanode/src/region\server · high confidence

New local procedure runner with lock conflict detection and panic recovery

The local procedure subsystem introduces a new \Runner\ component that manages procedure execution lifecycles, including automatic cleanup via a \ProcedureGuard\ that resets state to failed on panic and notifies parent procedures. It implements fine-grained lock management supporting both exclusive and shared keys, and adds logic to detect potential deadlocks by identifying conflicting lock keys between parent and child procedures. The runner also integrates with an event recorder for observability and handles retry logic with exponential backoff.

src/common/procedure/src/local · high confidence

New metadata key definitions and managers for catalog, schema, table, and flow scopes

The metadata key module now includes dedicated key structures and managers for catalog names, schema names (with TTL and extra options), table names and info, datanode-table mappings, flow metadata (info, route, state, name), node addresses, runtime switches (maintenance, pause, recovery), and table repartition tracking. These files define the key layouts (e.g., \_\_catalog\name/{catalog}, \\_schema\name/{catalog}/{schema}, \\_table\_info/{table\id}, \\_dn\_table/{datanode\_id}/{table\id}, \\_flow/...), serialization/deserialization, and manager APIs (create, get, batch\_get, exists, range streams) backed by the KV backend, providing the foundational metadata storage primitives for these entities.

src/common/meta/src/key · high confidence

New modular heartbeat handler pipeline in meta-srv

The meta-srv heartbeat processing has been restructured into a chain of dedicated handlers to improve clarity and extensibility. New handlers now manage specific responsibilities: \CheckLeaderHandler\ ensures only the leader processes heartbeats; \CollectClusterInfoHandler\ gathers node details for frontends, flownodes, and datanodes; \CollectLeaderRegionHandler\ tracks leader region manifests; \CollectStatsHandler\ and \PersistStatsHandler\ handle in-memory caching and periodic persistence of datanode statistics; \CollectTopicStatsHandler\ aggregates topic-level metrics; \ExtractStatHandler\ parses heartbeat payloads; \RegionFailureHandler\ feeds data to the region supervisor; \FilterInactiveRegionStatsHandler\ removes inactive regions from stats; \FlowStateHandler\ manages flownode state; \KeepLeaseHandler\ maintains node leases; \MailboxHandler\ processes mailbox messages; and \OnLeaderStartHandler\ resets state upon leadership changes. This modularization centralizes heartbeat logic within the handler module.

src/meta-srv/src/handler · high confidence

New modular key-value backend infrastructure with isolated storage backends

The metadata service now uses a unified, modular key-value backend architecture in \src/common/meta/src/kv\_backend\. This change introduces a standardized \KvBackend\ and \TxnService\ interface, along with a \Txn\ abstraction for atomic transactions. It provides concrete implementations for various storage engines: \EtcdStore\ for etcd, \MemoryKvBackend\ for in-memory testing, \RdsStore\ (with optional PostgreSQL and MySQL backends via features) for relational databases, and \ChrootKvBackend\ for namespacing. Additionally, it includes a \ReadOnlyKvBackend\ wrapper to enforce read-only access and a \MockKvBackend\ for testing. This refactoring centralizes backend logic, improves testability, and prepares the system for easier integration of new storage backends.

_src/common/meta/src/kv\backend · high confidence

New partition module with multi-dimensional partitioning and caching

The \src/partition\ crate has been introduced to centralize partitioning logic, featuring a new \MultiDimPartitionRule\ that allows tables to be partitioned across multiple columns using a set of boolean expressions. This module includes a \PartitionChecker\ to validate that partition rules are non-overlapping and fully cover the data space, and a \PartitionRuleManager\ that coordinates with a new \PartitionInfoCache\ to store and retrieve partition metadata efficiently. The implementation also provides utilities for splitting record batches into regions and handling physical versus logical table routes.

src/partition · high confidence

New plugins crate provides lifecycle hooks for service components

A new \plugins\ crate has been introduced to centralize plugin setup and lifecycle management across the system. It exposes standardized hooks for the frontend, datanode, flownode, metasrv, and standalone modes, allowing plugins to register during pre-build, post-build, and start phases. The crate also defines the \PluginOptions\ configuration structure, which supports filtering known options while safely ignoring and warning on unknown variants, and includes a CLI subcommand module to expose plugin-related tools.

src/plugins/src · high confidence

New procedural macros for SQL functions and error handling

The \src/common/macro\ crate now provides a suite of procedural macros to simplify the creation of SQL functions and improve error diagnostics. The \admin\_fn\ attribute macro converts Rust functions into SQL administration functions, supporting handlers for procedures, table mutations, and flows. The \range\_fn\ attribute macro transforms arithmetic functions into PromQL-compatible range functions, with optional support for specialized evaluators. Aggregate functions can now be created using the \as\_aggr\_func\_creator\ attribute and \AggrFuncTypeStore\ derive macro. Additionally, the \stack\_trace\_debug\ attribute generates stack-trace-style \Debug\ implementations for error types, and the \print\_caller\ macro injects call-site tracking into functions for debugging.

src/common/macro/src · high confidence

New procedure state and poison stores for large-value splitting and error isolation

The procedure store now includes a \PoisonStore\ trait and an \ObjectStateStore\ implementation backed by \ObjectStore\ to manage procedure state persistence. The \PoisonStore\ introduces a mechanism to mark resources as poisoned (inconsistent) when operations fail, preventing further operations until manual recovery. Additionally, the store supports automatic splitting of large procedure states into multiple segments via the \KeySet\ and \multiple\_value\_stream\ utilities, allowing efficient storage and retrieval of large values by merging split segments back together.

src/common/procedure/src/store · high confidence

A new \query\_perf\_fixture\ binary has been added to provide dedicated performance testing capabilities. It introduces a \PromRemoteWrite\ command to generate and send Prometheus remote-write traffic with configurable series, samples, and value patterns, and an \InspectFooter\ command to analyze Parquet file footers (column encodings, sizes, row counts) directly from object stores. The tool also supports \DirectSst\ scenarios for synthesizing SST files and \Plan\ for validating case configurations, enabling users to run structured performance benchmarks and inspect storage metadata.

_src/cmd/src/bin/query\_perf\fixture · high confidence

New record batch streaming and filtering utilities

The \src/common/recordbatch/src\ module now includes new components to enhance data streaming and query processing. A \RecordBatchStreamCursor\ allows applications to fetch data in controlled row-sized chunks from a stream, while a \SimpleFilterEvaluator\ provides an optimized, in-place evaluation for basic column comparisons and regex matches. Additionally, \ChainedRecordBatchStream\ and \LimitedRecordBatchStream\ enable the sequential chaining of multiple data streams and the enforcement of row limits, respectively, giving users finer control over how query results are consumed.

src/common/recordbatch/src · high confidence

New region lifecycle management components in meta-srv

The meta-srv now includes new internal modules for managing region health and operations: \failure\_detector.rs\ introduces a \RegionFailureDetector\ to track region heartbeats and detect failures; \flush\_trigger.rs\ adds a \RegionFlushTrigger\ that periodically triggers region flushes to manage WAL replay size and improve startup times; \lease\_keeper.rs\ implements a \RegionLeaseKeeper\ to handle region lease renewals and track operating regions; and \supervisor.rs\ provides a \RegionSupervisor\ that orchestrates these components, handling heartbeats, failure detection initialization, and region migration triggers.

src/meta-srv/src/region · high confidence

New region migration, repartition, and WAL pruning procedures

The meta-srv now includes dedicated procedure implementations for managing cluster topology and storage lifecycle. Region migration is handled by a multi-step procedure (start, flush, downgrade, upgrade, update metadata, abort, end) that moves regions between datanodes with configurable timeouts and trigger reasons (manual, auto-rebalance, failover). Repartitioning is supported via a new procedure that plans, allocates, and deallocates regions for tables, including handling for unpartitioned tables and group-based rollback. Additionally, a WAL pruning procedure has been added to manage Kafka-based WAL retention, supporting both physical deletion and logical pruning (metadata-only updates) to reduce storage overhead.

src/meta-srv/src/procedure · high confidence

New service configuration options for InfluxDB, Jaeger, MySQL, PostgreSQL, and OTLP

The frontend service now exposes dedicated configuration structs for several ingestion protocols, allowing users to control their behavior via TOML. InfluxDB ingestion can now be configured with a default merge mode (last\_non\_null or last\_row). Jaeger query APIs are enabled by default. MySQL and PostgreSQL servers now support configurable server-side keep-alive intervals and, for MySQL, a prepared statement cache size. OTLP trace ingestion allows configuring the chunk size and an experimental flag to synthesize resource descriptors for the semantic entity graph. These changes are reflected in the new service\_config modules for each protocol.

_src/frontend/src/service\config · high confidence

New system schema infrastructure and computed entity graph tables

The catalog now includes a new \system\_schema\ module that introduces a \MemoryTable\ implementation for static metadata tables and a \Predicate\ engine to filter \information\_schema\ query results. It also adds a \pg\_catalog\ provider to expose PostgreSQL-compatible system tables and introduces computed entity graph tables (\semantic\_entities\ and \semantic\_relationships\) in the \greptime\_private\ schema, which derive entity relationships from telemetry data at read time.

_src/catalog/src/system\schema · high confidence

New test utilities for Flight encoding, port allocation, and temporary files

The \src/common/test-util\ crate now provides shared helpers for internal testing: \flight.rs\ adds an \encode\_to\_flight\_data\ function to serialize \DfRecordBatch\ objects into Flight messages, \ports.rs\ offers a \get\_port\ function to allocate unique runtime ports for tests, \temp\_dir.rs\ re-exports \tempfile\ types and provides convenience functions to create temporary directories and named temp files, \recordbatch.rs\ includes a \check\_output\_stream\ helper to assert that query output matches expected pretty-printed text, and \lib.rs\ exposes a \find\_workspace\_path\ utility to resolve workspace-relative paths to absolute ones.

src/common/test-util · high confidence

New unified metadata caching and cluster information infrastructure

The metadata service now includes a new caching framework (\CacheContainer\, \CacheRegistry\) that supports configurable initialization strategies (version-checked vs. unchecked) and layered invalidation for table, schema, and flow metadata. Additionally, the \ClusterInfo\ trait and \NodeInfo\ struct have been expanded to report detailed node status, including CPU/memory usage, hostname, and environment variables, while \RegionStat\ now exposes granular storage metrics like SST file counts and written bytes.

src/common/meta/src · high confidence

New vector data types and equality logic introduced

The \src/datatypes/src/vectors\ module now includes implementations for Binary, Boolean, Date, Decimal128, Dictionary, Duration, Interval, JSON, List, Null, and Struct vectors, along with a comprehensive equality comparison system in \eq.rs\. This adds support for storing and comparing these specific data types as vectors, enabling users to work with binary data, dates, decimals, and complex nested structures in queries.

src/datatypes/src/vectors · high confidence

New vector operations: cast, filter, and take

The vector operations module now includes implementations for casting, filtering, and taking elements from vectors. The new \cast.rs\ file enables converting vector data between different types (e.g., float to int, date to timestamp) using Arrow's compute capabilities. The \filter.rs\ file allows selecting specific rows from a vector based on a boolean condition. The \take.rs\ file provides a way to extract specific elements from a vector by index. These operations are implemented as macros that delegate to the Arrow library's compute functions, ensuring consistent behavior with the underlying Arrow arrays.

src/datatypes/src/vectors/operations · high confidence

Physical table reconciliation procedure implementation

The meta service now includes a new reconciliation procedure for physical tables, implemented as a state machine in src/common/meta/src/reconciliation/reconcile\_table. The process starts by validating the table and listing region metadata, then resolves column metadata inconsistencies using a configurable strategy (UseLatest, UseMetasrv, or AbortOnConflict). If inconsistencies are found and resolved, the procedure syncs the column definitions to the datanodes via AlterRegion requests, updates the central table info, and invalidates the relevant table caches before completing.

_src/common/meta/src/reconciliation/reconcile\table · high confidence

PromQL engine scaffolding and extension plans introduced

The \src/promql\ area now contains the initial implementation of the PromQL query engine, including the core error handling (\error.rs\), a suite of DataFusion extension plans for query execution (\extension\_plan.rs\ and its submodules like \instant\_manipulate.rs\, \range\_manipulate.rs\, \histogram\_fold.rs\, and \absent.rs\), and the corresponding physical planner (\planner.rs\). This change also adds benchmarking infrastructure (\benches/bench\_main.rs\ and \benches/bench\_range\_fn.rs\) to measure the performance of PromQL range functions. This represents the foundational architecture for PromQL support rather than a specific user-facing feature update.

src/promql · high confidence

Repartition procedure implementation for table schema changes

The meta-srv now includes a new repartition procedure located in src/meta-srv/src/procedure/repartition that manages the lifecycle of table repartitioning. This implementation introduces state machines for allocating new regions, dispatching group-level sub-procedures, collecting their results, and deallocating old regions. It also handles metadata updates by applying and exiting staging states for region routes, ensuring data consistency during the transition. Additionally, a garbage collection requirement manager is included to enforce GC policies during and after the repartition process.

src/meta-srv/src/procedure/repartition · high confidence

Repository initialization with core documentation, build tooling, and licensing

The repository is initialized with essential project scaffolding, including the Apache 2.0 and Enterprise licenses, a security policy, and comprehensive contributor guidelines. Build infrastructure is established via a Makefile supporting native, Docker, and cross-compilation targets (Android, RISC-V), alongside Nix flake configurations for reproducible development environments. Developer workflows are standardized with pre-commit hooks for formatting and linting, and configuration examples are provided for various storage backends and testing scenarios.

(repo-wide) · high confidence

Support for full-text, inverted, and skipping index options in column definitions

The API layer now allows users to specify full-text, inverted, and skipping index configurations when defining columns. The new \column\_def.rs\ module maps these index options from gRPC column options into the internal column schema metadata and vice versa, enabling index settings to be passed through the API during table creation or alteration.

src/api/src/v1 · high confidence

Support for timestamp range metadata in Arrow Flight DoPut

The Arrow Flight DoPut implementation now accepts optional min and max timestamp metadata in requests, allowing clients to specify time-windowed batches. The server response includes the number of affected rows and the elapsed time for the bulk insert, enabling clients to track ingestion performance and coordinate requests via unique request IDs.

src/common/grpc/src/flight · high confidence

Telemetry module introduces uptime tracking and default-enabled reporting

The new \src/common/greptimedb-telemetry\ library establishes the core infrastructure for anonymous usage data collection, which is now enabled by default. It introduces an uptime metric that reports the system's running duration in privacy-friendly ranges (hours, days, weeks) alongside standard version and environment details. The module manages the telemetry lifecycle via \GreptimeDBTelemetryTask\, handling start/stop logic and printing a usage data disclaimer upon initialization, while configuring a 30-minute reporting interval to the specified telemetry endpoint.

src/common/greptimedb-telemetry · high confidence

Architecture

Centralized RPC definitions for DDL, routing, and store operations

The RPC layer in common-meta has been refactored to consolidate protocol definitions and request/response types into dedicated modules. This change introduces structured types for DDL tasks (including support for flows, views, and triggers), region routing logic (such as leader/follower distribution and state tracking), and key-value store operations (batch get/put, range queries). It also adds helper utilities for procedure state serialization and distributed locking, providing a unified interface for metadata operations across the system.

src/common/meta/src/rpc · high confidence

Extract mito codec into a dedicated crate

The Mito codec implementation has been extracted from the monolithic codebase into a new, standalone \mito-codec\ crate. This new module provides the core encoding and decoding logic for primary keys (both dense and sparse formats), manages key-value views for mutations, and includes utilities for index value encoding and primary key filtering. Users interacting with the storage layer will now rely on this isolated component for row conversion and serialization tasks.

src/mito-codec · high confidence

Extract standalone mode into a dedicated plugin-based crate

The standalone deployment mode has been refactored into a new \src/standalone\ crate, introducing a plugin-based router configuration. This change adds a \StandaloneDatanodeManager\ to handle local region and flow requests, a \StandaloneInformationExtension\ to expose node, procedure, and region statistics, and a \StandaloneRepartitionProcedureFactory\ that explicitly rejects repartition operations. Configuration is now managed via \StandaloneOptions\, which maps to frontend and datanode settings, and metadata storage is built using \RaftEngineBackend\.

src/standalone · high confidence

Metasrv initialization now uses a dedicated builder pattern

The \src/meta-srv/src/metasrv/builder.rs\ file introduces a \MetasrvBuilder\ struct that provides a fluent API for constructing the Metasrv instance. This builder allows users or internal callers to explicitly configure core components such as the \KvBackend\, \NodeManager\, \Selector\, and \HeartbeatHandlerGroupBuilder\ before instantiation, replacing previous direct construction or less structured initialization methods.

src/meta-srv/src/metasrv · high confidence

Refactored session and query context architecture

The session management layer has been restructured to improve clarity and performance. The \QueryContext\ is now a distinct, immutable-per-query structure that holds transient state (such as snapshot sequences, SQL dialect, and protocol-specific context), while the \Session\ struct manages persistent connection state (such as catalog/schema, timezone, and user info) using \RwLock\ for efficient concurrent access. This separation allows query contexts to be created cheaply from a session and enables protocol-specific data (like OTLP metric ingestion options) to be carried cleanly via a new \ProtocolCtx\ enum. Additionally, a new \QueryId\ type based on UUIDv7 is introduced to uniquely identify queries, and table name resolution now strictly enforces that the catalog in a query matches the current session catalog.

src/session · high confidence

Relocate CLI and server command implementations to src/cmd/src

The command-line interface and server entry points (CLI, Datanode, Flownode, Frontend, Metasrv, Standalone, and User tools) have been moved into the dedicated \src/cmd/src\ directory. This change reorganizes the codebase structure, extracting the \plugins\ crate and relocating the CLI subcommands and server \Instance\/\Command\ implementations to a new, centralized location for better modularity.

src/cmd/src · high confidence

Behavioural changes

Agent skills relocated to shared directory

The \.claude/skills\ path is now a symbolic link pointing to \../.agents/skills\, centralizing agent skill definitions in a shared location rather than keeping them local to the \.claude\ directory.

.claude · high confidence

Client transport lanes are now isolated

The gRPC client now maintains separate channel pools for query and control traffic, ensuring that query streams (Flight DoGet) and control operations (health checks, inserts, and other unary RPCs) no longer share connections. This isolation prevents a long-running query from blocking or delaying control-plane requests, improving reliability and latency for both read and write operations.

src/client · high confidence

Configurable default catalog and case-insensitive system schema handling

The default catalog name is now configurable via the \DEFAULT\_CATALOG\_NAME\ environment variable (defaulting to \greptime\), allowing users to customize the primary catalog identity at build time. Additionally, system schemas (\information\_schema\, \pg\_catalog\, and the private schema) are now matched case-insensitively during connection parsing, ensuring that queries using mixed-case system schema names behave consistently with standard SQL expectations.

src/common/catalog · high confidence

Customizable version information and optimized build-time compilation

The versioning system has been replaced with a custom build script that allows users to customize the product name via the GREPTIME\_PRODUCT\_NAME environment variable and derives the version from workspace metadata. To improve incremental compilation performance, Git-derived build info (branch, commit hash, etc.) is now only refreshed during release builds, while debug builds skip this overhead. The build script also dynamically watches workspace files and Git status changes to ensure version constants are updated only when necessary.

src/common/version · high confidence

Defensive limit handling in file stream queries

The file engine's query module now explicitly manages the \limit\ parameter when building record batch streams for CSV, JSON, and ORC formats. The \create\_stream\ function and its helpers (\new\_csv\_stream\, \new\_json\_stream\, \new\_orc\_stream\) only push the limit down to the DataFusion execution plan if no filters are present; otherwise, the limit is omitted to ensure correct result sets. This change prevents potential incorrect results or performance issues when filtering and limiting are combined, grounding the behavior in the specific file format handlers within \file\_stream.rs\.

src/file-engine/src/query · high confidence

Discovery logic refactored to use NodeInfo for active node detection

The meta-server's node discovery mechanism has been refactored to rely on the new \NodeInfo\ structure rather than legacy lease values for determining active nodes. This change introduces dedicated accessors (\NodeInfoAccessor\) and utility functions (\alive\_datanode\_infos\, \alive\_flownode\_infos\, etc.) that query node metadata, enabling the system to expose richer node details—such as environment variables, CPU/memory usage, and workload types—to placement selectors and other components. Users benefit from more accurate node liveness detection and the ability to make routing decisions based on detailed node attributes.

src/meta-srv/src/discovery · high confidence

Flow creation now supports deferred resolution of missing source tables

The CreateFlowProcedure implementation in the metadata module now includes logic to collect source table IDs while explicitly tracking unresolved (missing) source table names. This allows the flow creation process to proceed even when some source tables are not yet available, deferring their resolution rather than failing immediately.

_src/common/meta/src/ddl/create\flow · high confidence

Introduce common version reporting with build-time metadata

The \src/common/version\ library now provides a centralized way to access build information, including branch, commit hash, version, and other build-time details. This change replaces the previous shadow-rs dependency with a self-maintained version info system, allowing for more customization and control over how version data is generated and displayed. Users can now access this information through functions like \build\_info()\, \version()\, and \short\_version()\, which are used across the application to provide consistent version reporting.

src/common/version/src · high confidence

Introduce new API error handling and gRPC type conversion helpers

The \src/api\ module now includes a dedicated error handling module (\error.rs\) defining specific error variants for gRPC and data type operations, such as unknown column data types, JSON serialization failures, and time unit inconsistencies, along with their corresponding status codes. Additionally, \helper.rs\ provides comprehensive conversion logic between gRPC \ColumnDataType\ (including extensions for JSON, Decimal, and other types) and internal \ConcreteDataType\ representations, ensuring accurate type mapping for column definitions and data values in API requests.

src/api/src · high confidence

Introduce structured error handling for expression evaluation

The expression evaluation module now uses a dedicated \EvalError\ enum to handle errors during columnar evaluation, replacing previous ad-hoc error handling. This change introduces specific error variants for type mismatches, casting failures, invalid arguments, and internal issues, mapping them to appropriate status codes (e.g., \InvalidArguments\ for type mismatches) to provide clearer diagnostics for users encountering evaluation failures.

src/flow/src/expr · high confidence

Introduces structured error handling for jemalloc memory profiling

The jemalloc memory profiling module now uses a dedicated \Error\ enum to handle failures related to heap profiling activation, deactivation, status reading, and gdump flag updates. This change maps specific internal and storage errors to appropriate status codes and retry hints, ensuring that users receive clearer feedback when profiling operations fail due to issues like temporary file creation or I/O errors.

src/common/mem-prof/src/jemalloc · high confidence

Metasrv refactoring and new capabilities

The metasrv module has undergone a significant refactor, introducing a new plugin system and a cache invalidation mechanism that broadcasts cache updates to frontends, datanodes, and flownodes via the mailbox. A new event handler implementation now records lifecycle events (such as GC, region migration, and repartition) by inserting them into a dedicated table. Additionally, the service now supports distributed telemetry reporting, where the leader node collects and reports cluster information, and a new example demonstrates usage of the etcd key-value store.

src/meta-srv/src · high confidence

Metasrv service layer refactored into modular gRPC and HTTP handlers

The metasrv service implementation has been restructured into distinct, modular components for better maintainability and clarity. The \admin.rs\ file introduces a new HTTP admin interface (using Axum) that exposes endpoints for health checks, node leases, heartbeats, leader status, maintenance mode, and procedure management, while marking the legacy \make\_admin\_service\ as deprecated. The \cluster.rs\ file implements the gRPC cluster service, providing methods for batch operations, range queries, and retrieving metasrv peer information (including node stats like CPU and memory usage). The \heartbeat.rs\ file contains the core heartbeat handling logic, including session management, leader step-down detection, and pusher registration. The \mailbox.rs\ file defines the mailbox receiver and channel types for inter-node communication. The \procedure.rs\ file implements the procedure service, handling DDL tasks, region migrations, and procedure state queries. The \store.rs\ file implements the KV store service, exposing standard key-value operations (put, get, delete, range) with metrics instrumentation. Finally, \utils.rs\ provides a \check\_leader\ macro to enforce leader-only execution for specific requests. This refactoring separates concerns, improves observability, and modernizes the HTTP interface while maintaining backward compatibility for gRPC.

src/meta-srv/src/service · high confidence

Metric engine introduces sparse primary key encoding and replaces mur3 with fxhash for TSID generation

The metric engine now uses sparse primary key encoding by default, compacting table and time-series identity into the primary key rather than using separate internal columns. This change is accompanied by a performance optimization that replaces the mur3 hash algorithm with fxhash for generating Time Series IDs (TSIDs), significantly speeding up tag hashing. The engine also adds a benchmark suite to validate the TSID generator performance improvements.

src/metric-engine · high confidence

New DDL procedure implementations and module structure

The \src/common/meta/src/ddl\ module has been restructured to introduce dedicated procedure implementations for core database operations. A new \allocator.rs\ file now exposes submodules for region routes, resource IDs, and WAL options. New files have been added to handle \AlterDatabase\, \AlterLogicalTables\, \AlterTable\, \CommentOn\, \CreateDatabase\, \CreateFlow\, \CreateLogicalTables\, and \CreateTable\ operations, each implementing the standard procedure lifecycle (prepare, execute, and state management). This change centralizes the logic for these DDL tasks within the meta service.

src/common/meta/src/ddl · high confidence

New SQL value conversion and default constraint logic in common/sql

The \src/common/sql\ module now provides the core logic for converting SQL values to internal data types and parsing column default constraints. This includes support for parsing numeric strings into various integer and float types, handling boolean coercion (e.g., 0/1), and managing timestamp overflow checks. For default constraints, the system now explicitly rejects default values for JSON columns and correctly handles negative number defaults by parsing them as unary operations to prevent overflow errors. These changes centralize SQL-to-datatype conversion and default value validation within the \common/sql\ crate.

src/common/sql · high confidence

New configuration options for datanode client, memory profiling, and lenient plugin loading

This change introduces new configuration structures in the common options module. It adds \DatanodeClientOptions\ to allow tuning gRPC message size limits and timeouts for datanode clients. It also adds \MemoryOptions\ to control heap profiling activation. Additionally, it implements lenient plugin option deserialization, which allows the system to start even if the configuration contains plugin options not recognized by the current build, issuing a warning instead of failing.

src/common/options · high confidence

New metadata key infrastructure for Flow management

The metadata layer for Flows has been restructured into a dedicated key-value schema under the \\_\_flow\ namespace. This change introduces specific key managers and storage layouts for flow lifecycle and routing: \FlowInfoKey\ (storing flow metadata and schedule configuration like \FlowScheduleConfig\), \FlowNameKey\ (mapping catalog/flow names to IDs), \FlowRouteKey\ (routing flow partitions to specific flownodes), \FlowStateKey\ (tracking per-flownode flow state), \FlownodeFlowKey\ (indexing flows by flownode), and \TableFlowKey\ (mapping source tables to their consuming flows). This provides a structured, scalable foundation for flow metadata operations.

src/common/meta/src/key/flow · high confidence

New schema module with column constraints and index metadata support

The schema subsystem in \src/datatypes/src/schema\ has been restructured into dedicated modules. \column\_schema.rs\ now defines the \ColumnSchema\ struct, which manages column metadata including time index flags, comments, and specific index types (full-text, inverted, and skipping) via dedicated metadata keys. \constraint.rs\ introduces the \ColumnDefaultConstraint\ enum, enabling users to specify default values for columns, including support for \CURRENT\_TIMESTAMP\ and \NOW\ functions for timestamp columns. Additionally, \ext.rs\ provides utilities to detect JSON extension fields within Arrow schemas.

src/datatypes/src/schema · high confidence

OTLP trace ingestion now supports v2 pipeline and request-wide schema reconciliation

The frontend OTLP ingestion module has been restructured to support three trace models: the legacy v0 (fixed schema with JSON attributes), the dynamic v1 (flattened attributes with request-wide schema reconciliation), and the new v2 (fixed schema with JSON2 attributes). The v2 pipeline writes fixed-schema chunks without dynamic-column reconciliation but currently discards events and links. The v1 path now performs request-wide schema planning to reconcile compatible column observations across chunks, allowing safe widening of Int64 columns to Float64 while rejecting unsupported type mixes. Ingestion outcomes now include bounded, deduplicated failure details that name the specific cause of rejected spans, and the system distinguishes between full success, partial success (with rejected spans), and failure.

src/frontend/src/instance/otlp · high confidence

Ported query regression runner to Rust

The query regression runner tool has been rewritten in Rust, replacing the previous implementation. This new version introduces a structured CLI with subcommands for measuring query performance against base and candidate endpoints, preparing direct-SST fixtures, materializing fixtures into object stores, handling Prometheus remote-write scenarios, and running OTLP trace loads. The Rust implementation provides stricter validation for table schemas and region directories, integrates with the OpenDAL object store abstraction for flexible storage backends, and outputs structured JSON reports for regression analysis.

_src/cmd/src/bin/query\_regression\runner · high confidence

Preserve primary key index order in table creation templates

The table creation logic now explicitly preserves the order of primary key indices as defined in the original table metadata when generating region creation requests. This ensures that logical and physical table templates maintain the correct primary key column ordering, which is critical for schema reconstruction and storage layout consistency on the datanode side.

_src/common/meta/src/ddl/create\table · high confidence

Refactored ALTER TABLE execution into a modular executor and procedure

The ALTER TABLE logic has been restructured into a new \AlterTableExecutor\ and supporting modules (\metadata.rs\, \region\_request.rs\) within the \ddl/alter\_table\ directory. This change centralizes the coordination of table metadata updates and region alterations, introducing specific handling for semantic table options (SET/UNSET), JSON2 settings, inverted index modifications, and default value management (SET/DROP DEFAULT). The refactoring also ensures that column addition requests are idempotent by skipping existing columns and correctly updates partition key indices during alter operations.

_src/common/meta/src/ddl/alter\table · high confidence

Refactored alter logical tables procedure into separate executor, validator, and metadata update modules

The alter logical tables procedure has been restructured to separate validation, execution, and metadata update logic into distinct components. The new \AlterLogicalTableValidator\ ensures all alter expressions share the same schema and are limited to \AddColumns\ operations, while verifying that logical table routes correctly reference the same physical table. The \AlterLogicalTablesExecutor\ handles sending alter region requests to datanodes and updating physical table metadata, and the \update\_metadata\ module manages the persistence of new table info versions. This separation improves code clarity and maintainability of the logical table alteration workflow.

_src/common/meta/src/ddl/alter\_logical\tables · high confidence

Refactored drop table execution with soft-drop support and improved rollback logic

The table drop process has been restructured into a dedicated executor and metadata handling module. This change introduces soft-drop table support (gated behind the enterprise feature), allowing tables to be logically deleted with a retention lifecycle rather than immediately removed. It also refines rollback behavior, ensuring that metadata rollback is specifically triggered when dropping metric physical tables fails, while preventing conflicts with existing table tombstones.

_src/common/meta/src/ddl/drop\table · high confidence

Refactored heartbeat response handling into a modular handler pipeline

The heartbeat response processing logic in the meta service has been restructured to use a pluggable handler pattern. A new \HeartbeatResponseHandler\ trait and \HandlerGroupExecutor\ allow distinct response types (such as mailbox messages, cache invalidation, and suspend instructions) to be processed by dedicated, isolated modules. This change improves maintainability and ensures that errors in one handler do not prevent others from executing, while also introducing safer logging that excludes large mailbox payloads from error traces.

src/common/meta/src/heartbeat · high confidence

Refactored heartbeat response handling into modular, composable handlers

The heartbeat response processing logic has been restructured into a chain of dedicated, reusable handler components located in the common meta module. This change introduces specific handlers for distinct operations: \InvalidateCacheHandler\ now manages schema cache invalidation instructions, \ParseMailboxMessageHandler\ converts incoming mailbox payloads into internal instructions, and \SuspendHandler\ manages the suspension state of frontend and datanode components based on meta-server signals. By extracting these concerns into individual modules, the system ensures that cache updates, mailbox message parsing, and suspension controls are handled consistently and independently within the heartbeat loop.

src/common/meta/src/heartbeat/handler · high confidence

Refactored logical table creation into modular procedure steps

The \CreateLogicalTablesProcedure\ implementation has been restructured into distinct modules (\check\, \metadata\, \region\_request\, \update\_metadata\) to improve code clarity and maintainability. This refactoring introduces specific validation logic to ensure input tasks share the same schema and to handle cases where tables already exist, automatically merges physical partition columns into logical table schemas, and separates the generation of region creation requests from metadata updates. Users benefit from a more robust and organized internal process for creating logical tables, ensuring consistent schema handling and clearer separation of concerns during table provisioning.

_src/common/meta/src/ddl/create\_logical\tables · high confidence

Refactored meta-client with modular, leader-aware RPC clients

The meta-client implementation in \src/meta-client/src/client\ has been restructured into distinct, specialized modules (ask\_leader, cluster, config, heartbeat, procedure, store, util) that share a common leader discovery mechanism. This change introduces a \LeaderProvider\ abstraction that allows clients to dynamically discover and switch to the current Metasrv leader, improving resilience against leader changes and VIP/LB backend failures. The refactored client now supports pulling configuration from Metasrv during startup, handles heartbeat streams with configurable intervals, and exposes dedicated RPC clients for DDL tasks, cluster information, and key-value store operations, all while centralizing error handling and retry logic.

src/meta-client/src/client · high confidence

Refined region migration metadata updates with leader downgrading and rollback support

The region migration procedure now explicitly manages leader state transitions during metadata updates. It introduces a 'downgrading' state for the old leader region to prevent write conflicts while the candidate region is being prepared, and adds a rollback step to restore the previous leader if the migration fails. The upgrade step for the candidate region now strictly validates that the old leader matches expectations before switching to the new leader, ensuring consistency in the table route metadata.

_src/meta-srv/src/procedure/region\_migration/update\metadata · high confidence

Removal of default application entry point

The default \main.rs\ entry point, which previously printed a standard "Hello, world!" message, has been removed from the source tree. This change eliminates the default executable behavior for the application, likely as part of restructuring the project into a library or common crate.

src · high confidence

Unified CLI storage configuration and metadata backend support

The CLI now uses a unified configuration module for object storage and metadata backends. Object store settings (S3, GCS, Azblob, OSS, FS) are validated declaratively, ensuring required fields are present when a backend is enabled. The metadata store backend is configurable via the \--backend\ flag, supporting Etcd (default), Memory, RaftEngine, and optionally PostgreSQL or MySQL via feature flags. PostgreSQL support includes an \--auto-create-schema\ option (enabled by default) to automatically create the metadata schema if it does not exist. TLS settings for backend connections are also configurable via dedicated flags.

src/cli/src/common · high confidence

Fixes

Pin binstall installation to version 1.6.6

The Docker dev-builder now uses a dedicated script to install cargo-binstall, pinning the version to v1.6.6 instead of fetching the latest release. This ensures consistent behavior in the development environment by avoiding potential breaking changes from upstream updates.

docker/dev-builder/binstall · high confidence

Test coverage

Add TLS fixtures and configuration for integration tests; Add distributed mode SQLness regression tests; Add time-range filtering benchmark for BulkPart; Added CSV test fixtures for schema inference and type casting; Added benchmark dataset for pipeline performance testing; Added benchmarks for Flight decoder and channel manager; Added benchmarks for RecordBatch iteration and memory accounting; Added test data generation for ORC format; Added test utilities for creating local and Kafka log stores; Added tests for DDL event contracts across database, table, flow, and view operations; Added tests for DDL procedures in the meta service; Added tests for OTLP trace ingestion failure handling and schema logic; Added tests for configuration loading and standalone daemon mode; Added tests for derive macros ToRow, Schema, and IntoRow; Added tests for error extension utilities and retry hints; Added tests for gRPC channel manager TLS configuration and reloading; Added tests for heartbeat cache invalidation and task configuration; Added unit tests for the GC scheduler mock implementation; Added unit tests for the authentication permission checker; New DDL test utilities for table and region metadata construction; New sqlness test runner with multi-protocol and compatibility testing support; New test utilities for executing and mocking database procedures.

Dependencies

1665 commits updating dependencies (78 manifests)

A dependency / build maintenance change in (dependencies) — 1665 commits (204 fixs), 78 files.

(dependencies) · high confidence · unverified

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 56 → 60 (+4.5)
  • Rubric changed (rubric-2026.09.9 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 84 → 84 (+0.3)
  • Architecture 98 → 71 (-27.0)
  • Maturity 88 → 88 (+0.0)
  • Readiness 48 → 60 (+11.4)
  • Security 39 → 47 (+8.3)
  • Domain Modelling 93 → 91 (-1.3)
  • Event Sourcing 100 → 100 (+0.0)
  • Performance 94 (new)

Resolved (418)

  • Arrangement::get (cognitive 19) (src/flow/src/utils.rs)
  • BatchingTask::execute_logical_plan_unlocked (cognitive 21) (src/flow/src/batching_mode/task.rs)
  • Change coupling: adapter.rs ↔ map.rs (src/flow/src/adapter.rs)
  • Change coupling: data_region.rs ↔ alter.rs (src/metric-engine/src/data_region.rs)
  • Change coupling: data_region.rs ↔ create.rs (src/metric-engine/src/data_region.rs)
  • Change coupling: func.rs ↔ plan.rs (src/flow/src/expr/func.rs)
  • Change coupling: func.rs ↔ transform.rs (src/flow/src/expr/func.rs)
  • Change coupling: test_util.rs ↔ handle_open.rs (src/mito2/src/test_util.rs)
  • Change-coupling hub: metrics.rs → ddl.rs, ddl_manager.rs, error.rs (src/common/meta/src/metrics.rs)
  • Change-coupling hub: reader.rs → context.rs, part_reader.rs, range.rs (src/mito2/src/sst/parquet/reader.rs)
  • ClassTooLong: Batch (src/mito2/src/read.rs)
  • ClassTooLong: DatafusionQueryEngine (src/query/src/datafusion.rs)
  • ClassTooLong: ScalarExpr (src/flow/src/expr/scalar.rs)
  • ClassTooLong: StreamingEngine (src/flow/src/adapter.rs)
  • Documentation: no contributor guidance (README.md)
  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • Documentation: written for insiders (docs/style-guide.md)
  • Duplicated block (10 lines × 2) (src/cmd/src/standalone.rs)
  • Duplicated block (10 lines × 2) (src/datatypes/src/vectors/operations/replicate.rs)
  • …and 398 more

New (386)

  • Actor::start_creates (cognitive 18) (src/log-store/src/object_store_wal/store.rs)
  • BatchingEngine::create_flow_inner (cognitive 16) (src/flow/src/batching_mode/engine.rs)
  • BatchingTask::capture_recovery_windows_since (cognitive 16) (src/flow/src/batching_mode/task.rs)
  • BatchingTask::execute_plan_unlocked (cognitive 25) (src/flow/src/batching_mode/task.rs)
  • BatchingTask::execute_plan_unlocked (cyclomatic 16) (src/flow/src/batching_mode/task.rs)
  • Boundary-crossing change coupling: applier.rs ↔ bloom_filter_index.rs (src/index/src/bloom_filter/applier.rs)
  • Boundary-crossing change coupling: builder.rs ↔ options.rs (src/frontend/src/instance/builder.rs)
  • Boundary-crossing change coupling: lib.rs ↔ elasticsearch.rs (src/pipeline/src/lib.rs)
  • Boundary-crossing change coupling: prom_store.rs ↔ test_util.rs (src/frontend/src/service_config/prom_store.rs)
  • Change coupling: context.rs ↔ reader.rs (src/mito2/src/memtable/bulk/context.rs)
  • Change coupling: csv.rs ↔ json.rs (src/common/datasource/src/file_format/csv.rs)
  • Change coupling: ddl.rs ↔ metrics.rs (src/common/meta/src/ddl.rs)
  • Change coupling: ddl_manager.rs ↔ metrics.rs (src/common/meta/src/ddl_manager.rs)
  • Change coupling: range.rs ↔ reader.rs (src/mito2/src/read/range.rs)
  • ClassTooLong: Actor (src/log-store/src/object_store_wal/store.rs)
  • ClassTooLong: DatanodeBuilder (src/datanode/src/datanode.rs)
  • ClassTooLong: DfLogicalPlanner (src/query/src/planner.rs)
  • ClassTooLong: Error (src/datanode/src/error.rs)
  • ClassTooLong: Error (src/log-store/src/error.rs)
  • ClassTooLong: FrontendClient (src/flow/src/batching_mode/frontend_client.rs)
  • …and 366 more

Changes since last survey

  • 193 commits — 131 feature/other, 62 fixes

By area

  • src/mito2 — 31 commits
  • src/servers — 22 commits
  • src/common — 18 commits
  • tests/cases — 14 commits
  • .github/workflows — 12 commits
  • (root) — 11 commits
  • src/query — 11 commits
  • src/operator — 7 commits
  • src/flow — 6 commits
  • src/log-store — 6 commits
  • .github/runner-scale-sets — 4 commits
  • .github/scripts — 4 commits
  • src/cli — 4 commits
  • src/datanode — 4 commits
  • src/meta-srv — 4 commits
  • src/promql — 4 commits
  • .github/actions — 3 commits
  • docs/rfcs — 3 commits
  • src/frontend — 3 commits
  • src/datatypes — 2 commits

Notable commits

  • fix: ci(query-regression): bump RUNNER_IMAGE_EPOCH to 6 (#9381)
  • fix: ci(query-regression): bump RUNNER_IMAGE_EPOCH to 7 (#9395)
  • fix: fix(auth): follow symlink chains in watch_file_user_provider (#9365)
  • fix: fix(ci): build tests-integration lib with meta-srv/mock (#9299)
  • fix: fix(ci): grant PR write permission for CI command replies (#9331)
  • fix: fix(ci): increase query regression ECS disk to 80 GiB (#9123)
  • fix: fix(ci): repair agent observability dispatch and runner cleanup (#9204)
  • fix: fix(ci): repair draft PR command dispatch (#9271)
  • fix: fix(ci): rerun semantic PR checks after pushes (#9190)
  • fix: fix(ci): stabilize long-range benchmark execution and artifact collection (#9241)
  • fix: fix(ci): teach check-builder-rust-version.sh to handle stable channels (#9369)
  • fix: fix(ci): update the shared Actions runner to v2.337.0 (#9376)
  • fix: fix(client): complete transport lane isolation (#9030)
  • fix: fix(client): yield Flight batches and affected rows without waiting for next message (#8918)
  • fix: fix(json): fix JSONPath panic with jsonb 0.5.6 (follow-up to #9192) (#9228)
  • fix: fix(json2): restrict JSON2 type hints (#9316)
  • fix: fix(meta): populate physical metric table column ids (#9286)
  • fix: fix(meta-srv): use NoTls for disabled and Unix socket Postgres KV backends (#9059)
  • fix: fix(mito): preserve mixed JSON2 types during compaction (#9135)
  • fix: fix(mito2): cancel cache construction for incomplete scans (#9254)
  • …and 173 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

GreptimeTeam/greptimedb was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit a99cb56d7168c88727b34468db4eee1c0f3b4e2c — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-c4983f2d4e5c.