Skip to content
CAI
Software that uses CAICheck a score

apache/druid

45.5

Weak · 6 August 2026

892.6k

lines of production code

Java

with TypeScript

3

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Apache Druid is a distributed, columnar, real-time analytics database designed for high-performance data ingestion and interactive querying. The system supports a wide variety of data sources, including Kafka, Kinesis, HDFS, S3, and various cloud storage backends, while offering flexible storage options such as S3, GCS, Cassandra, and SQL databases. It provides extensive aggregation capabilities through DataSketches, Bloom filters, and histogram aggregations, enabling fast approximate and exact statistical analysis. Additionally, the platform features a comprehensive security model with LDAP and database-backed authentication, role-based authorization, and support for Kerberos and JWT-based access control.

How it got here

2012–2018 — Security, storage, and statistical extensions

93 changes.

This period focused on hardening the platform with comprehensive security features, including basic authentication, LDAP support, and Kerberos integration. It also expanded data handling capabilities by adding support for diverse storage backends like PostgreSQL, Cassandra, and S3, while introducing advanced statistical aggregations via DataSketches and histogram extensions.

2019–2020 — cloud storage and security extensions

50 changes.

This period focused on expanding cloud storage support by adding extensions for Google Cloud Storage, Aliyun OSS, and HDFS, alongside new aggregators like Bloom Filters and t-digest sketches. The release also introduced significant security enhancements, including LDAP and OIDC authentication, while providing comprehensive configuration templates and Docker setups for easier deployment.

2021–2026 — Extensibility and Observability

46 changes.

This period focused on expanding Druid's ecosystem by introducing new connectors for Iceberg, Delta Lake, and RabbitMQ, alongside advanced aggregation types like compressed BigDecimals and DDsketches. Concurrently, the project significantly enhanced observability through OpenTelemetry and Prometheus integrations, while establishing a robust testing infrastructure with Docker-based containers and SQL-level validation tools.

Features

Add AWS RDS token-based password provider

The Druid AWS RDS extension now includes a new \AWSRDSTokenPasswordProvider\ that generates temporary authentication tokens for Amazon RDS databases using the AWS SDK. This allows users to configure database credentials via JSON with the type \aws-rds-token\, specifying the database host, port, username, and AWS region, enabling secure, token-based authentication without storing static passwords.

extensions-core/druid-aws-rds-extensions · high confidence

Add Aliyun OSS input source support

Users can now ingest data from Alibaba Cloud Object Storage Service (OSS) using the new OssInputSource, which registers the 'oss' scheme in Jackson for ingestion specs. The change introduces configuration classes (OssClientConfig) for endpoint and credentials, an entity handler (OssEntity) for reading object streams, and a Druid module (OssInputSourceDruidModule) to wire up the input source, enabling batch ingestion from OSS buckets.

extensions-contrib/aliyun-oss-extensions/src/main/java/org/apache/druid/data · high confidence

Add Array-of-Doubles Sketch aggregation support

Users can now perform approximate set and statistical aggregations on arrays of double values using the Apache DataSketches library. This change introduces new aggregators (build, merge, and constant) and post-aggregators for the 'arrayOfDoublesSketch' complex type, enabling users to compute estimates, variances, and set operations (union, intersection, not) on tuple sketches directly within Druid queries.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/tuple · high confidence

Add Bloom Filter Aggregators to Druid

The druid-bloom-filter extension now includes a complete set of bloom filter aggregation implementations. This adds support for aggregating data into Bloom filters, enabling approximate set operations and cardinality estimation within Druid queries. The implementation includes specific aggregators for different data types (String, Long, Double, Float, and generic Object), as well as a merge aggregator for combining results. This allows users to perform efficient, memory-optimized aggregations on large datasets.

extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/query/aggregation/bloom · high confidence

Add Cassandra storage backend for data segments

Users can now store and retrieve data segments using a Cassandra backend. This change introduces the \CassandraDataSegmentConfig\, \CassandraDataSegmentPuller\, \CassandraDataSegmentPusher\, and \CassandraLoadSpec\ classes, which implement the segment loading and pushing interfaces for Cassandra. The \CassandraDruidModule\ registers the \c\*\ storage scheme, enabling Druid to read from and write to Cassandra clusters configured via the new \druid.storage\ configuration properties.

extensions-contrib/cassandra-storage/src/main/java · high confidence

Add Consul-based service discovery and leader election extension

A new contrib extension, \consul-extensions\, is added to provide Consul-based service discovery and leader election for Apache Druid, enabling a complete ZooKeeper replacement. This includes configuration classes (\ConsulDiscoveryConfig\, \ConsulSSLConfig\), a Guice module (\ConsulDiscoveryModule\) to wire the components, and implementations for node announcement (\ConsulDruidNodeAnnouncer\), node discovery (\ConsulDruidNodeDiscoveryProvider\), and leader election (\ConsulLeaderSelector\). The extension supports TLS/mTLS, ACL, and HTTP Basic Auth for secure Consul communication.

extensions-contrib/consul-extensions · high confidence

Add DDsketch aggregation and post-aggregation support

Users can now compute DDsketch-based quantile estimates on numeric columns. This change introduces a new \ddSketch\ aggregation type that builds sketches during ingestion or query time, along with \quantileFromDDSketch\ and \quantilesFromDDSketch\ post-aggregators that extract single or multiple quantile values from the aggregated sketches.

extensions-contrib/ddsketch · high confidence

Add Dropwizard metrics emitter extension

Users can now export Druid metrics to Dropwizard-compatible sinks. This change introduces the \dropwizard-emitter\ extension, which includes a \DropwizardEmitter\ that converts Druid \ServiceMetricEvent\s into Dropwizard metrics (counters, gauges, histograms, meters, and timers) and forwards \AlertEvent\s to configured alert emitters. The extension supports two reporter types—\console\ and \jmx\—and allows configuration of metric prefixes, host inclusion, dimension mapping, and registry size limits via \DropwizardEmitterConfig\. A \DropwizardConverter\ handles dimension filtering and metric type mapping, while \DropwizardMetricSpec\ defines the structure for metric specifications. The module registers the emitter via Guice, enabling seamless integration with existing Druid monitoring pipelines.

extensions-contrib/dropwizard-emitter/src/main/java · high confidence

Add Druid 26.0 release notebook and tutorial documentation

A new Jupyter notebook (Druid26.ipynb) and its accompanying README have been added to the quickstart releases directory. The notebook provides a hands-on tutorial for Druid 26.0, demonstrating new features such as type-aware schema auto-discovery, shuffle join, and UNNEST support. Users can run these examples against a quickstart cluster to learn about the latest capabilities.

examples/quickstart/releases · medium confidence

Add GCE autoscaling support for Druid workers

Users can now enable automatic scaling of worker nodes on Google Cloud Engine (GCE). This change introduces the GCE autoscaler implementation, which manages worker node lifecycle by provisioning and terminating instances based on workload. The update also improves the pending task-based provisioning strategy to handle scenarios where there are no currently running worker nodes, ensuring more robust scaling behavior.

extensions-contrib/gce-extensions, extensions-core/ec2-extensions · high confidence

Add Google Cloud Storage support for data ingestion, deep storage, and task logs

This change introduces a new Google Cloud Storage (GCS) extension for Apache Druid, enabling users to ingest data from GCS buckets, use GCS as a deep storage backend for segments, and store task logs in GCS. The diff adds new classes including \GoogleCloudStorageInputSource\ and its factory for reading data, \GoogleDataSegmentPusher\/\GoogleDataSegmentPuller\/\GoogleDataSegmentKiller\ for segment management, and \GoogleTaskLogs\ for log storage. It also introduces configuration classes like \GoogleAccountConfig\ and \GoogleInputDataConfig\ to manage bucket, prefix, and listing parameters. The \GoogleStorageDruidModule\ registers these components, allowing Druid to interact with Google Cloud Storage for various storage and ingestion needs.

extensions-core/google-extensions/src/main · high confidence

Add HDFS input source for indexing tasks

Users can now read data from HDFS files directly in indexing tasks. This change introduces the HdfsInputSource, HdfsInputEntity, and related configuration classes, enabling HDFS as a supported input source alongside local, S3, and other cloud storage options.

extensions-core/hdfs-storage/src/main/java/org/apache/druid/inputsource · high confidence

Add HDFS task log storage and streaming support

Users can now store and stream task logs, reports, and status files to HDFS. The new HdfsTaskLogs implementation handles pushing and streaming task payloads, with task IDs sanitized (colons replaced) to comply with HDFS path naming constraints.

extensions-core/hdfs-storage/src/main/java/org/apache/druid/storage/hdfs/tasklog · high confidence

Add JWT and OpenID Connect authentication via the pac4j extension

The druid-pac4j extension now supports two authentication methods: a new 'jwt' authenticator that validates ID tokens from a bearer header, and an updated 'pac4j' authenticator for full OpenID Connect flows. The OIDC configuration now includes configurable scope and client authentication method, and the extension supports custom SSL contexts for secure communication with the identity provider.

extensions-core/druid-pac4j/src/main · high confidence

Add KLL doubles sketch aggregation support

Users can now aggregate double values into KLL sketches using the new KllDoublesSketchAggregatorFactory and related classes. This includes support for building, merging, and serializing KLL doubles sketches, as well as post-aggregators to convert sketches to CDFs, histograms, and quantiles.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/kll · high confidence

Add Microsoft SQL Server metadata storage support

Users can now store metadata in a Microsoft SQL Server database. This change introduces the SQLServerConnector and SQLServerMetadataStorageModule, registering the 'sqlserver' storage type with the Druid metadata provider, allowing Druid to persist metadata using SQL Server-specific types like VARBINARY(MAX) and BIGID for efficient storage.

extensions-contrib/sqlserver-metadata-storage/src/main/java/org/apache/druid/metadata/storage · high confidence

Add MomentSketch aggregation extension

Introduces a new MomentSketch aggregation extension that provides approximate quantile estimation using a moments-based sketch. The change adds the \MomentSketchModule\ and associated classes to register the \momentSketch\ and \momentSketchMerge\ aggregators, along with post-aggregators for min, max, and quantile extraction. This enables users to perform memory-efficient statistical aggregations on numeric columns.

extensions-contrib/momentsketch/src/main · high confidence

Add ORC file format support for native batch ingestion

Users can now ingest ORC (Optimized Row Columnar) files using the native batch ingestion path. This change introduces the \orc-extensions\ module, which registers an \OrcInputFormat\ to read ORC files, handling type conversion and schema discovery. The implementation includes the necessary Guice bindings and Jackson modules to support ORC as a data source, enabling users to load data from ORC files directly into Druid.

extensions-core/orc-extensions · high confidence

Add OpenLineage request logging extension

Users can now enable OpenLineage event emission for Druid queries by setting druid.request.logging.type=openlineage. The new OpenLineageRequestLogger and its provider support both CONSOLE and HTTP transport modes, with configurable queue capacity, thread count, and TLS/SSL settings. Custom JSON schemas for query context and statistics facets are included, and the extension registers itself via the standard Druid module service file.

extensions-contrib/openlineage-emitter · high confidence

Add OpenTelemetry emitter extension for query tracing

Users can now enable the OpenTelemetry emitter to generate OpenTelemetry spans for Druid query/time metrics. This new extension extracts W3C trace context from query context to link Druid spans to parent traces, and adds query context entries as span attributes.

extensions-contrib/opentelemetry-emitter · high confidence

Add Parquet data source support

The Parquet extension is promoted from 'contrib' to 'core' and now includes the full implementation for reading Parquet files, including the ParquetExtensionsModule, ParquetInputFormat, ParquetReader, and related converters and flatteners. This enables users to ingest data from Parquet files directly into Druid.

extensions-core/parquet-extensions · high confidence

Add PostgreSQL metadata storage extension

Users can now use PostgreSQL as a metadata storage backend for Apache Druid. This change introduces a new extension that provides the necessary components to connect to a PostgreSQL database, including the \PostgreSQLConnector\ for database interactions, \PostgreSQLMetadataStorageActionHandler\ for handling metadata actions, and configuration classes for SSL and table schemas. The extension also registers the PostgreSQL connector as an option for metadata storage, allowing users to configure their Druid cluster to store metadata in a PostgreSQL database instead of the default or other supported databases.

(repo-wide) · high confidence

Add Protobuf input format and decoding support

The Protobuf extension now supports ingesting Protobuf data through a new \ProtobufInputFormat\ that decodes binary Protobuf messages into Druid rows. Users can configure decoding via three strategies: loading a \.desc\ file from the classpath or URL (\FileBasedProtobufBytesDecoder\), providing an inline Base64-encoded descriptor (\InlineDescriptorProtobufBytesDecoder\), or fetching schemas from a Confluent Schema Registry (\SchemaRegistryBasedProtobufBytesDecoder\). The extension includes a \ProtobufConverter\ to transform Protobuf messages into plain Java objects, a \ProtobufFlattenerMaker\ for nested field discovery, and a custom \ProtobufJsonProvider\ to support JSON path and JQ query flattening on Protobuf data.

extensions-core/protobuf-extensions · high confidence

Add Protobuf-based metrics publisher example

A new example in the quickstart directory demonstrates sending metrics via Protobuf serialization. The change introduces a \metrics.proto\ schema defining a \Metrics\ message with fields for unit, HTTP method, value, timestamp, HTTP code, page, metric type, and server. A corresponding Python module \metrics\_pb2.py\ is generated from this schema, and a \pb\_publisher.py\ script is added to read JSON lines from stdin, convert them to Protobuf format, and publish them to a Kafka topic named 'metrics\_pb'.

examples/quickstart/protobuf · medium confidence

Add Quidem UT module for SQL-level testing

The quidem-ut module has been added to enable writing SQL-level tests easily. It provides a test framework that can be used to write tests against existing test backends, allowing test cases to be moved closer to the exercised code. The module includes a launcher that exposes a QueryComponentSupplier as a Broker service, allowing for the capture of SQL queries and their results. This facilitates the creation of regression tests that can be run against different configurations and backends.

quidem-ut · high confidence

Add REST API for managing basic authentication users

The basic-security extension now exposes a new REST endpoint at /druid-ext/basic-security/authentication that allows administrators to manage users for basic HTTP authentication. The API supports creating, deleting, and listing users, as well as updating credentials and refreshing caches. Coordinator nodes handle the actual database operations, while non-coordinator nodes receive updates via a listener. This change enables programmatic management of the user database for basic authentication.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authentication/endpoint · high confidence

Add Rackspace CloudFiles storage support for data segments

The CloudFiles storage extension now supports pushing and pulling data segments to and from Rackspace CloudFiles. This includes new configuration classes for account and pusher settings, along with implementations for loading, pulling, and pushing segments via the CloudFiles API, enabling users to store their Druid data segments in Rackspace's object storage.

extensions-contrib/cloudfiles-extensions/src/main/java · high confidence

Add SQL parser configuration for custom SQL syntax

The SQL module now includes FMPP configuration files (config.fmpp and default\_config.fmpp) that extend the Calcite SQL parser. This setup defines new SQL keywords (CLUSTERED, OVERWRITE, PARTITIONED, EXtern), adds non-reserved keywords (OVERWRITE, EXTERN, VARIANT), and registers custom statement parsers (INSERT, EXPLaine, REPLACE) and data type parsers, enabling the SQL layer to recognize and parse these specific SQL constructs.

sql · high confidence

Add SQL support for BLOOM\_FILTER aggregation

Users can now use the BLOOM\_FILTER function in SQL queries to perform bloom filter aggregations. This new SQL aggregator maps the BLOom\_FILTER SQL function to the underlying Druid BloomFilterAggregatorFactory, allowing users to specify the column and maxNumEntries parameters directly in their SQL statements.

extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/query/aggregation/bloom/sql, extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/query/expressions · high confidence

Add SQL support for HLL sketch operations

Users can now use HLL sketch estimates, estimates with error bounds, set unions, and string conversions directly in SQL queries. New SQL aggregators (DS\_HLL, APPROX\_COUNT\_DISTINCT\_DS\_HLL, and a UTF-8 variant) allow building and merging HLL sketches, while operator conversions enable functions like HLL\_SKetch\_ESTIMATE, HLL\_SKETCH\_ESTIATE\_WITH\_ERROR\_BOUNDS, HLL\_SKETCH\_UNION, and HLL\_SKETCH\_TO\_STRING to be used in SQL expressions.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/hll/sql · high confidence

Add SQL support for non-expression sketch post-aggregations

Users can now use SQL functions to access various properties of double sketch aggregations, including approximate quantiles (DS\_GET\_QUANTILE, DS\_GET\_QUANTILES), rank (DS\_RANK), cumulative distribution function (DS\_CDF), histogram (DS\_HISTOGRAM), and summary (DS\_QUANTILE\_SUMMARY). This enables direct querying of sketch metadata and derived values within SQL queries.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/quantiles/sql · high confidence

Add SQL support for quantile aggregation on histograms

Users can now compute quantiles on histogram data directly via SQL. New SQL aggregators (APPROX\_QUANTILE and APPROX\_QUantile\_FIXED\_BUCKETS) map to Druid's ApproximateHistogram and FixedBucketsHistogram aggregators, enabling statistical analysis of histogram columns in SQL queries.

extensions-core/histogram/src/main/java/org/apache/druid/query/aggregation/histogram/sql · high confidence

Add SpectatorHistogram extension for histogram aggregation

The SpectatorHistogram extension is introduced, enabling users to aggregate and query histogram data. This includes new classes for serialization (NullableOffsetsHeader, SpectatorHistogramIndexed), aggregation (SpectatorHistogramAggregator, SpectatorHistogramAggregatorFactory), and post-aggregation (SpectatorHistogramCountPostAggregator). The module registers the 'spectatorHistogram' complex type, allowing Druid to ingest, store, and compute on Spectator-style histograms.

extensions-contrib/spectator-histogram/src/main · high confidence

Add Thrift input format extension

Users can now ingest data encoded in the Apache Thrift format. This change introduces a new Thrift input format that supports binary, compact, and JSON Thrift protocols (with optional Base64 encoding). The extension registers the ThriftInputFormat, allowing users to specify a thriftClass and an optional thriftJar to load custom Thrift schema classes for deserialization.

extensions-contrib/thrift-extensions/src/main/java · high confidence

Add approximate histogram aggregation support

Users can now perform approximate histogram aggregations on numeric columns using the new \approxHistogram\ aggregation type. This feature introduces a suite of components including \ApproximateHistogram\ and its associated aggregators (\ApproximateHistogramAggregator\, \ApproximateHistogramFoldingAggregator\, and vectorized/buffered variants) to efficiently compute and combine histogram data. The change also registers the necessary serialization and deserialization logic via \ComplexMetrics\ and \ApproximateHistogramDruidModule\, enabling the system to handle complex \approximateHistogram\ types in queries and post-aggregations.

extensions-core/histogram/src/main/java/org/apache/druid/query/aggregation/histogram · high confidence

Add compressed BigDecimal support for high-precision aggregations

The compressed-bigdecimal module introduces a new column type and associated aggregators (SUM, MIN, MAX) that store and process high-precision decimal values using a compressed, memory-efficient representation. This allows users to perform accurate financial or scientific calculations without the precision loss or memory overhead associated with standard floating-point types. The implementation includes specialized storage strategies and serialization logic to handle large decimal values efficiently within Druid segments.

extensions-contrib/compressed-bigdecimal/src/main · high confidence

Add configuration files for a single-machine ZooKeeper deployment example

The examples/conf/zk directory now includes three new configuration files—jvm.config, log4j2.xml, and zoo.cfg—that provide a ready-to-use setup for running ZooKeeper as a single-node instance. The JVM configuration sets memory and timezone, the logging config enables daily rolling log files with automatic cleanup, and the ZooKeeper config defines the data directory, client port, and autopurge settings.

examples/conf/zk · high confidence

Add configuration files for the historical node in the single-machine example

The single-machine deployment example now includes specific configuration files for the historical node, including JVM settings (jvm.config), the main entry point (main.config), and runtime properties (runtime.config) that define memory limits, thread counts, segment caching, and query caching settings for that component.

examples/conf/druid/cluster/data/historical · high confidence

Add configuration files for the query layer (broker and router)

New example configuration files have been added for the broker and router components in the query layer. The broker configuration sets up HTTP server and client settings, processing buffers, and disables the query cache. The router configuration defines HTTP proxy settings, service discovery for the broker and coordinator, and enables the management proxy for the unified web console.

examples/conf/druid/cluster/query · high confidence

Add configuration for the Druid indexer server

Added new configuration files (jvm.config, main.config, runtime.properties) for the Druid indexer, defining JVM memory settings, the main class, and runtime properties such as the service name, port, worker capacity, and processing thread configuration.

examples/conf/druid/cluster/data/indexer · high confidence

Add coordinator-overlord configuration for Druid cluster master node

New configuration files (jvm.config, main.config, runtime.properties) are added for the coordinator-overlord role on the master node. This sets up the JVM memory (15g heap), defines the server type as 'coordinator' with the overlord service enabled, and configures the segment metadata cache to poll every 5 seconds for faster segment loading.

examples/conf/druid/cluster/master · high confidence

Add default metric dimensions and module registration for the Dropwizard emitter

The Dropwizard emitter extension now includes a new \defaultMetricDimensions.json\ file that defines metric types, dimensions, and time units for query, segment, and ingest metrics. Additionally, the \META-INF/services\ file registers the \DropwizardEmitterModule\, enabling the extension's metrics to be automatically discovered and initialized.

extensions-contrib/dropwizard-emitter/src/main/resources · high confidence

Add dump-segment tool for inspecting segment metadata

A new 'dump-segment' tool has been added to the Druid examples, allowing users to dump v10 segment metadata. This change includes the necessary configuration files (common.runtime.properties and log4j2.xml) and JVM settings to support this new utility.

examples/conf/druid/tools · medium confidence

Add exact cardinality counting using 64-bit bitmaps

Users can now perform exact distinct counting on numeric columns using a 64-bit bitmap implementation. This new aggregation type, exposed as 'bitmap64ExactCount' in post-aggregators and 'BITMAP64\_EXACT\_COUNT' in SQL, provides precise cardinality results by tracking unique values in a RoaringBitmap64Counter. The extension includes build and merge aggregators, serialization, and SQL support, allowing queries to return exact counts of unique values without approximation.

extensions-contrib/druid-exact-count-bitmap · high confidence

Add example configurations for single-server and cluster deployments

Added new example configuration files for deploying Druid in both single-server and cluster topologies. The single-server examples (nano-quickstart, micro-quickstart, small, medium, large, and xlarge) provide pre-configured setups for running all Druid nodes on one machine, with varying resource allocations. Cluster examples include configurations for a master node (with and without ZooKeeper) and a query layer, enabling users to quickly spin up a multi-node cluster. These configs serve as templates for different deployment scales.

examples/conf/supervise · high confidence

Add gRPC query extension for SQL and Native queries

The gRPC query extension is introduced as a contrib module, enabling Druid Brokers to expose a gRPC API for executing SQL and Native queries. The extension supports multiple response formats including CSV, JSON, and Protobuf, and implements server-side interceptors for both anonymous and Basic authentication. The server lifecycle is managed via Guice, and a standard health check service is included for monitoring.

extensions-contrib/grpc-query · high confidence

Add git hooks for pre-commit and pre-push checks

Developers can now install a suite of git hooks that automatically run pre-commit and pre-push checks. The new \install-hooks.sh\ script copies hook scripts into the \.git/hooks\ directory, including a Python utility (\run-all-in-dir.py\) that executes all scripts in specific directories (e.g., \pre-commits\, \pre-pushes\). This includes a \checkstyle-check\ script for pre-push validation, ensuring code quality standards are met before commits and pushes.

hooks · high confidence

Add middleManager configuration for single-machine deployment

New configuration files (jvm.config, main.config, runtime.properties) are added for the middleManager in the single-machine deployment example. These files define JVM settings, the main class to run, and runtime properties such as worker capacity, task directories, and processing thread counts, enabling users to deploy and run the middleManager component in a single-node environment.

examples/conf/druid/cluster/data/middleManager · high confidence

Add native S3 input source support for SQL queries and catalog tables

The S3 extension now includes a native \S3InputSource\ and associated factory classes, enabling users to read data directly from S3 buckets via SQL queries and catalog table definitions. This change introduces new configuration classes (\S3InputSourceConfig\) to handle S3-specific properties such as access keys, session tokens, and role ARNs, allowing for secure and flexible S3 data ingestion without relying on Hadoop-based ingestion. The implementation supports defining S3 sources in the catalog, where users can specify buckets, prefixes, or glob patterns, and the system handles the underlying S3 client interactions and retries.

extensions-core/s3-extensions · high confidence

Add simple-client-sslcontext extension for internal HTTPS client SSL configuration

The simple-client-sslcontext extension is introduced to manage SSL/TLS configuration for Druid's internal HTTP client communication. It provides a configurable SSLContext via the new SSLClientConfig class, which exposes properties such as trust and key store paths, protocols, and hostname validation. The SSLContextModule registers the SSLContext for various Guice scopes (Global, Client, EscalatedGlobal, EscalatedClient, and Router), enabling secure internal node-to-node communication with customizable certificate and key management.

extensions-core/simple-client-sslcontext · high confidence

Add support for Aliyun OSS as a deep storage backend

Users can now use Alibaba Cloud Object Storage Service (OSS) as a deep storage backend for Druid. This change introduces a new extension that implements the core storage interfaces, including \OssDataSegmentPusher\ for uploading data segments, \OssDataSegmentPuller\ for retrieving them, \OssDataSegmentMover\ for moving segments, \OssDataSegmentArchiver\ for archiving, and \OssDataSegmentKiller\ for deletion. The module also registers \OssTaskLogs\ to store task logs in OSS, and provides configuration classes (\OssStorageConfig\, \OssInputDataConfig\, \OssDataSegmentArchiverConfig\) to manage bucket names, prefixes, and archive settings. A new \OssLoadSpec\ handles the serialization and deserialization of OSS segment metadata.

extensions-contrib/aliyun-oss-extensions/src/main/java/org/apache/druid/storage · high confidence

Add support for Theta sketch aggregations and post-aggregators

Users can now perform approximate distinct count and set operations using Theta sketches. This change introduces new aggregator and post-aggregator classes (such as SketchAggregator, SketchEstimatePostAggregator, and SketchMergeAggregatorFactory) that enable Theta sketch computations, including the ability to return estimates with error bounds. The implementation includes serialization, object strategies, and buffer aggregators specifically for the theta package, allowing these sketch types to be used in queries and expressions.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/theta · high confidence

Add support for ingesting from RabbitMQ super streams

Users can now ingest data from RabbitMQ super streams using the new \RabbitStreamSupervisor\ and \RabbitStreamIndexTask\ implementations. This feature introduces a complete indexing pipeline for RabbitMQ, including configuration classes (\RabbitStreamSupervisorIOConfig\, \RabbitStreamIndexTaskIOConfig\), task runners, and record suppliers that handle RabbitMQ-specific sequence numbers and offsets.

extensions-contrib/rabbit-stream-indexing-service/src/main · high confidence

Add timeMin and timeMin aggregations for timestamp columns

The time-min-max extension now provides new aggregators to compute the minimum and maximum timestamp values from a specified column. Users can now apply timeMin and timeMax aggregations to find the earliest and latest timestamps in their data, with support for custom time formats and field names.

extensions-contrib/time-min-max/src/main/java · high confidence

Add variance and standard deviation aggregators with SQL support

The stats extension now includes new aggregators for calculating variance and standard deviation, supporting both sample and population estimators. The implementation provides vectorized aggregators for float, double, and long columns, along with a folding aggregator for distributed computation. These aggregators are registered in the DruidStatsModule and exposed via SQL functions (var\_pop, var\_samp, stddev\_pop, stddev\_samp, etc.), enabling statistical analysis directly in SQL queries.

extensions-core/stats · high confidence

Added Ambari Metrics Emitter extension for Apache Druid

A new contrib extension, Ambari Metrics Emitter, has been added to Apache Druid. This includes a new DruidModule registration in \META-INF/services\ and a \defaultWhiteListMap.json\ that defines the set of metrics (such as \segment/scan/active\, \query/segments/count\, and various segment and ingest metrics) that will be emitted to Ambari.

extensions-contrib/ambari-metrics-emitter/src/main/resources · high confidence

Added Kerberos authentication and client escalation support

Introduced a new Kerberos security extension for Apache Druid, including the \DruidKerberosAuthenticationHandler\ for server-side authentication, a \KerberosAuthenticator\ for handling HTTP negotiation and signed cookies, and a \KerberosEscalator\ with a \KerberosHttpClient\ to allow internal services to make authenticated requests using Kerberos tickets. The module registers these components via \DruidKerberosModule\ and provides utility classes for cookie management and GSS-API context handling.

extensions-core/druid-kerberos/src/main/java · high confidence

Added build infrastructure and source files for the RADStack and Druid demo papers

The publications directory now includes the complete source files for two academic papers: the 'Druid' demo paper and the 'RADStack' paper. For each paper, the repository now contains the LaTeX source (.tex), bibliography (.bib), and auxiliary files (.aux, .bbl, .blg, .out), along with a Makefile to automate the PDF generation process using LuaLaTeX and BibTeX. Additionally, R scripts are provided to generate the data plots and tables used in the RADStack paper. This change enables users to reproduce the PDFs and understand the structure of the published work.

publications · high confidence

Added release automation and license compliance scripts

The distribution/bin directory now includes a suite of new Python and shell scripts to support release engineering and license compliance. These include tools to generate binary license and notice files (generate-binary-license.py, generate-binary-notice.py), list jar notices (jar-notice-lister.py), and generate dependency reports (generate-license-dependency-reports.py). Additionally, new scripts assist in managing release milestones and backports (find-missing-backports.py, tag-missing-milestones.py, get-milestone-prs.py, get-milestone-contributors.py), while others help format release notes (make-linkable-release-notes.py) and list web console dependencies (web-console-dep-lister.py).

distribution/bin · high confidence

Added support for raw input value extraction and query context configuration for DataSketches aggregations

The DataSketches extension now includes a new RawInputValueExtractor, enabling users to extract raw metric values from input rows for use in sketch aggregations. Additionally, a SketchQueryContext class has been introduced to support the 'sqlFinalizeOuterSketches' query context, allowing users to control whether outer sketches are finalized during SQL query execution.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches · high confidence

Added tools for scale and fault-tolerance testing

The testing-tools extension now includes a suite of components to simulate faults and delays across the Druid cluster for scalability and fault-tolerance testing. This includes a 'sleep' expression macro and SQL operator that allow queries to be intentionally delayed, as well as 'faulty' client implementations for the Coordinator, Overlord, and task actions that can introduce configurable delays or skip cleanup steps. Additionally, a custom 'eventCollector' service has been added to collect and forward metrics from other Druid services, facilitating the monitoring of these test scenarios.

extensions-core/testing-tools · high confidence

Basic authentication now supports database-backed user management

The basic security extension now includes a new \BasicAuthenticatorMetadataStorageUpdater\ interface and its implementations (\CoordinatorBasicAuthenticatorMetadataStorageUpdater\ and \NoopBasicAuthenticatorMetadataStorageUpdater\) in the \db/updater\ package. This change enables the system to store and manage basic authentication users and credentials directly in the metadata storage, allowing for persistent user management rather than relying solely on configuration files.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authentication/db/updater · high confidence

Configure default metrics whitelist for Graphite emitter

The Graphite emitter extension now includes a default whitelist of metrics to be exported, covering query timing, segment status, and JVM statistics. This configuration file defines which metrics are sent to Graphite by default, ensuring consistent monitoring of query performance and segment states.

extensions-contrib/graphite-emitter/src/main/resources · high confidence

Enable Aliyun OSS as a deep storage and input source

The Aliyun OSS extension is now registered as a Druid module, enabling users to use Alibaba Cloud OSS as a deep storage backend and as an input source for data ingestion.

extensions-contrib/aliyun-oss-extensions/src/main/resources · high confidence

Introduce Bloom filter support for dimension filtering

Added new classes (BloomDimFilter, BloomKFilter, BloomKFilterHolder) to enable filtering rows using a Bloom filter. This allows queries to leverage pre-computed Bloom filters for faster dimension matching, with serialization and caching support via SHA-512 hashing of the filter bytes.

extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/query/filter · high confidence

Introduce Derby-based task storage and new input source modules

The indexing-service now includes a new Derby-based task storage module (DerbyTaskStorageModule) that binds the MetadataStorageActionHandlerFactory to the Derby implementation, enabling the use of an embedded SQL database for task state management. Additionally, a new IndexingServiceInputSourceModule is introduced to register the GeneratorInputSource and DruidInputSource, providing a built-in data generator for testing and a standardized input source interface for indexing tasks.

indexing-service · high confidence

Introduce Iceberg extension with support for multiple catalogs and filter types

Users can now ingest data from Apache Iceberg tables using the new druid-iceberg-extensions module. This adds support for reading Iceberg data files via local, Hive, AWS Glue, and REST-based catalogs. The extension provides a new \IcebergInputSource\ and a suite of filter implementations (\IcebergAndFilter\, \IcebergOrFilter\, \IcebergNotFilter\, \IcebergEqualsFilter\, \IcebergIntervalFilter\, \IcebergRangeFilter\, \IcebergTimeWindowFilter\) that allow users to apply complex filtering conditions during ingestion. The module also includes a Guice module (\IcebergDruidModule\) to wire up these components, with specific handling for class loading issues in the Glue catalog.

extensions-contrib/druid-iceberg-extensions/src/main · high confidence

Introduce Kafka-based lookup namespace extraction

Added new classes to support Kafka as a source for lookup data. The module registers Jackson subtypes for the Kafka lookup extractor factory, and the factory itself consumes Kafka topics to populate the namespace cache. An introspection handler is also provided to check the status of the Kafka consumer.

extensions-core/kafka-extraction-namespace/src/main/java · high confidence

Introduce KinesisInputFormat and KinesisRecordEntity for streaming ingestion

The Kinesis indexing service now includes a new KinesisInputFormat and KinesisRecordEntity to handle reading data from AWS Kinesis streams. This change adds support for extracting Kinesis-specific record properties such as the approximate arrival timestamp and partition key, allowing these fields to be included in the ingested data. The implementation wraps the raw Kinesis record data and metadata into a structured entity that can be processed by the existing streaming ingestion pipeline.

extensions-core/kinesis-indexing-service · high confidence

Introduce LDAP and Metadata Store credential validators

Added new validator implementations for basic authentication: LDAPCredentialsValidator for connecting to LDAP directories and MetadataStoreCredentialsValidator for validating against the internal metadata store. The change introduces a common CredentialsValidator interface and a PasswordHashGenerator for secure password hashing, enabling users to configure their preferred authentication backend.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authentication/validator · high confidence

Introduce database-backed authorizer metadata storage updater

Added new classes to manage the persistence of basic security authorization data (users, roles, group mappings) in the metadata storage. The \BasicAuthorizerMetadataStorageUpdater\ interface and its implementations (\CoordinatorBasicAuthorizerMetadataStorageUpdater\ and \NoopBasicAuthorizerMetadataStorageUpdater\) enable the system to read and write authorizer state directly to the database, supporting features like superuser permissions and role-based access control stored in the metadata store.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authorization/db/updater · high confidence

Introduce filter support for the Druid Delta Lake connector

The Druid Delta Lake extension now supports filtering data during ingestion. This change adds a new \DeltaFilter\ interface and its implementations (e.g., \DeltaEqualsFilter\, \DeltaAndFilter\) that translate Druid query filters into Delta Kernel predicates. These filters are applied at the scan level to prune data, improving ingestion performance by reducing the amount of data read from Delta Lake tables.

extensions-contrib/druid-deltalake-extensions/src/main · high confidence

Introduce metadata catalog for table definitions

Adds a new metadata catalog system that stores and manages table definitions (schemas and table metadata) separate from the actual data sources. The Coordinator now maintains the authoritative catalog via a SQL-backed storage layer, while Brokers and Overlords cache and sync this metadata via a new internal REST API. This enables SQL queries to resolve table schemas from the catalog rather than relying on the legacy DruidLeaderClient.

extensions-core/druid-catalog/src/main · high confidence

Introduce mixed-mode task runner supporting both Kubernetes and traditional workers

The Kubernetes overlord extension now includes a new \KubernetesAndWorkerTaskRunner\ that allows tasks to be scheduled on either Kubernetes Jobs or traditional workers, depending on the configured runner strategy. This enables a hybrid execution model where operators can route specific tasks to the Kubernetes-based runner while keeping others on standard workers, providing flexibility in workload distribution and resource management within a single Druid cluster.

extensions-core/kubernetes-overlord-extensions · high confidence

Introduce namespace-based global cached lookups for JDBC and URI sources

Adds new classes to support global cached lookups based on extraction namespaces, including \NamespaceLookupExtractorFactory\ and \NamespaceLookupIntrospectHandler\ for managing cache lifecycle and introspection. The update introduces \ExtractionNamespace\ implementations for JDBC (\JdbcExtractionNamespace\) and URI (\UriExtractionNamespace\) sources, along with corresponding cache generators (\JdbcCacheGenerator\, \UriCacheGenerator\) that populate caches from these sources. A \MapPopulator\ utility is added to handle populating maps from byte sources, and a \StaticMapExtractionNamespace\ is provided for testing caching mechanisms. This change enables the system to cache and periodically update lookup data from external sources.

extensions-core/lookups-cached-global · high confidence

Introduce structured counter tracking for multi-stage queries

Adds a new, structured counter tracking system for multi-stage queries, replacing ad-hoc metrics with a formalized approach. The change introduces a \CounterTracker\ that manages named counters (such as CPU time, row counts, and file loads) and organizes snapshots into a \CounterSnapshotsTree\ keyed by stage and worker. This enables more granular and consistent reporting of query execution metrics, including CPU usage, shuffle progress, and storage statistics, which are then aggregated and exposed in task reports.

multi-stage-query · high confidence

Introduce t-digest sketch aggregators for quantile estimation

Adds support for t-digest based sketch aggregators, enabling users to compute quantile estimates on numeric data. This change introduces new aggregation and post-aggregation operators (such as \tDigestSketch\, \quantileFromTDigestSketch\, and \quantilesFromTDigestSketch\) that allow for efficient approximate quantile calculations. The implementation includes configuration for max compression, SQL function support for generating sketches and computing quantiles, and serialization/deserialization logic for the underlying \MergingDigest\ objects.

extensions-contrib/tdigestsketch/src/main · high confidence

Introduces configurable class composition and caching for basic authentication

The basic security extension now supports a pluggable architecture for authentication and authorization components. New configuration classes (BasicAuthClassCompositionConfig, BasicAuthCommonCacheConfig, BasicAuthDBConfig, BasicAuthLDAPConfig, BasicAuthSSLConfig) allow administrators to specify custom classes for metadata storage updaters, cache managers, resource handlers, and cache notifiers for both authenticators and authorizers. This change enables users to customize how credentials are stored, cached, and synchronized across Druid nodes, and introduces a new BasicSecurityDruidModule to wire these components via dependency injection.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic · high confidence

New AWS client configuration and credential providers for AWS SDK v2

The AWS common module now includes a complete set of configuration classes and utility providers for the AWS SDK v2, including \AWSClientConfig\, \AWSCredentialsConfig\, \AWSEndpointConfig\, and \AWSProxyConfig\. A new \AWSModule\ registers these configurations via Guice, wiring a default credentials provider chain that supports environment variables, system properties, web identity tokens, profile credentials, container credentials, and file-based session credentials. Additionally, \AWSClientUtil\ introduces a retry mechanism for transient AWS errors, improving resilience for S3 and Kinesis operations.

cloud/aws-common/src/main · high confidence

New Avro input formats and decoders for streaming and OCF data

Users can now ingest Avro data from streams and Object Container Files (OCF) using the new \AvroStreamInputFormat\ and \AvroOCFInputFormat\ classes. The \AvroBytesDecoder\ interface and its implementations (\InlineSchemaAvroBytesDecoder\, \SchemaRepoBasedAvroBytesDecoder\, \SchemaRegistryBasedAvroBytesDecoder\) provide flexible ways to decode Avro records, including support for schema registries and schema repositories. The \AvroFlattenerMaker\ enables JSONPath-based flattening of Avro records, allowing users to extract nested fields and handle unions. This change introduces a more structured approach to Avro parsing, supporting both inline and external schema sources.

extensions-core/avro-extensions/src/main/java · high confidence

New Dockerfile and compose setup for Druid

The distribution/docker directory now contains a complete, modernized Docker build setup. A new multi-stage Dockerfile builds the Druid distribution from source, uses a distroless Java 25 base image, and deduplicates JARs to reduce image size. A companion Dockerfile.mariadb and Dockerfile.mysql allow users to easily add the respective JDBC connectors to the base image. The setup includes a docker-compose.yml for running a full Druid cluster (coordinator, broker, historical, middlemanager, router) with PostgreSQL and Zookeeper, along with helper scripts (druid.sh, peon.sh, deduplicate\_jars.sh) and an environment file for quickstart configuration.

distribution/docker · high confidence

New DoublesSketch quantile aggregation support

Users can now perform quantile aggregations on double-precision floating-point columns using the DataSketches library. This change introduces a new \quantilesDoublesSketch\ aggregator that builds a sketch from input values, along with a merge aggregator for combining sketches. The implementation includes full support for buffered and vectorized aggregation, serialization, and SQL integration for quantile, histogram, CDF, and rank post-aggregations.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/quantiles · high confidence

New Kafka emitter implementation for emitting metrics, alerts, and other events

The Kafka emitter has been refactored into a new implementation in \KafkaEmitter.java\ and \KafkaEmitterConfig.java\, introducing support for emitting multiple event types (metrics, alerts, requests, and segment metadata) with configurable topics and producer settings. The module registers the emitter via \KafkaEmitterModule\, allowing users to configure event types, bootstrap servers, and producer properties through Druid configuration. This change enables more flexible and reliable event emission to Kafka, with improved thread management and error tracking for lost messages.

extensions-contrib/kafka-emitter/src/main/java · high confidence

New Kubernetes-based node discovery and leader election implementation

The kubernetes-extensions module now provides a complete implementation for Kubernetes-based node discovery and leader election. This includes a new \K8sDiscoveryConfig\ for configuration, a \K8sApiClient\ for interacting with the Kubernetes API, and a \K8sDruidNodeDiscoveryProvider\ that watches for pod changes. Additionally, leader election is handled via a \K8sDruidLeaderSelector\ that uses Kubernetes ConfigMaps for coordination. These components are wired into the application via the \K8sDiscoveryModule\ Guice module, enabling Druid to run on Kubernetes without requiring Zookeeper for node discovery or leader election.

extensions-core/kubernetes-extensions · high confidence

New MySQL metadata storage extension with SSL and driver configuration

The mysql-metadata-storage extension now provides a complete, standalone implementation for storing metadata in MySQL. This includes a new MySQL connector that supports configurable SSL settings (including trust stores, client certificates, and cipher suites) and allows specifying the JDBC driver class (supporting both standard MySQL and MariaDB drivers). The extension registers the necessary Guice bindings and Jackson modules to enable this storage backend.

extensions-core/mysql-metadata-storage · high confidence

New REST API for managing basic security authorizers

A new set of REST endpoints has been added to manage basic security authorizers. The \BasicAuthorizerResource\ exposes endpoints for creating, reading, updating, and deleting users, group mappings, and roles, as well as retrieving their permissions. The implementation is split into \CoordinatorBasicAuthorizerResourceHandler\ for coordinator nodes and \DefaultBasicAuthorizerResourceHandler\ for non-coordinator nodes, allowing administrators to manage authorization policies via the API.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authorization/endpoint · high confidence

New StatsD emitter extension with DogStatsD support

A new StatsD emitter extension has been added to Druid, enabling the export of metrics to a StatsD server. The implementation supports standard StatsD metrics and includes full DogStatsD compatibility, allowing users to send constant tags, include the service name as a tag, and emit DogStatsD-style events. Configuration is handled via the StatsDEmitterConfig class, which exposes parameters for hostname, port, prefix, separator, and various DogStatsD-specific options like constant tags and event emission.

extensions-contrib/statsd-emitter/src/main/java · high confidence

New automated quickstart configuration for Druid

A new set of example configuration files has been added to the Druid auto directory, providing a ready-to-use setup for a local Druid cluster. This includes common JVM and runtime properties, a log4j2 logging configuration, and specific runtime properties for each node type (broker, coordinator, historical, indexer, middleManager, and router). The configuration enables features like the ServiceStatusMonitor, SQL support, and routing task logs to dedicated files, offering a simplified starting point for users.

examples/conf/druid/auto · high confidence

New bar chart and multi-axis chart modules in the Explore view

The Explore view now includes two new visualization modules: a bar chart that displays a single measure against a split column with configurable time buckets and sorting, and a multi-axis chart that plots multiple measures over time with automatic or manual granularity selection. These modules are registered in the module repository and provide interactive charting capabilities within the web console.

web-console/src/views/explore-view/modules/bar-chart-module, web-console/src/views/explore-view/modules/multi-axis-chart-module · high confidence

New basic authorization components for role-based and read-only access control

Added new classes in the basic security extension to support role-based authorization and read-only access control. This includes the BasicRoleBasedAuthorizer for enforcing permissions based on user roles, a RoleProvider interface with implementations for metadata store and LDAP-based role resolution, and a ReadOnlyAuthorizer that permits all read operations while denying write and delete actions.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authorization · high confidence

New cache manager and notifier interfaces for basic authorization

The basic security extension now introduces a new caching layer for authorization data. A new \BasicAuthorizerCacheManager\ interface defines how cached user, role, and group mapping data is accessed and updated. Two implementations are provided: \CoordinatorPollingBasicAuthorizerCacheManager\, which periodically polls the coordinator for updates, and \MetadataStoragePollingBasicAuthorizerCacheManager\, which retrieves data from metadata storage. Additionally, a \BasicAuthorizerCacheNotifier\ interface and its implementations (\CoordinatorBasicAuthorizerCacheNotifier\ and \NoopBasicAuthorizerCacheNotifier\) handle broadcasting updates to non-coordinator services. This change restructures how the basic authorizer manages and distributes authorization state.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authorization/db/cache · high confidence

New continuous time-chart rendering for the Explore view

The Explore view now renders time-series data using a new continuous chart component (time-chart-module). This introduces support for line, area, and bar mark types with configurable curve styles (smooth, linear, step), faceting, and timezone-aware time axes. Users can interact with the chart to select and shift time ranges, with visual feedback for selection and shifting states.

web-console/src/views/explore-view/modules/time-chart-module · high confidence

New developer documentation and tooling in the dev/ directory

The repository now includes a comprehensive set of developer resources in the dev/ directory. This includes a guide for backporting changes, a script to manage heap dump permissions, and detailed instructions for committers on PR and issue management. Additionally, configuration files for code formatting in Eclipse and IntelliJ are provided, along with setup guides for IntelliJ, license management procedures, and XML application definitions for debugging Druid services.

dev · high confidence

New distinct count aggregation extension with pluggable bitmap backends

The distinct count aggregation feature is introduced in the extensions-contrib/distinctcount module, providing a new way to count distinct values in queries. The implementation includes a core DistinctCountAggregatorFactory that supports pluggable bitmap backends (Java, Concise, and Roaring) via a BitMapFactory interface, allowing users to choose the underlying data structure for distinct counting. The module registers the 'distinctCount' aggregation type with Jackson for serialization and provides Noop aggregators for edge cases.

extensions-contrib/distinctcount/src/main/java · high confidence

New entity classes for basic security authorization

The basic security extension introduces a new set of entity classes to represent authorization data, including BasicAuthorizerUser, BasicAuthorizerRole, BasicAuthorizerPermission, and their corresponding 'Full' and 'SimplifiedPermissions' variants. These classes define the structure for users, roles, and permissions, enabling the storage and retrieval of security metadata. The introduction of 'SimplifiedPermissions' variants suggests a shift in how permissions are represented in API responses, moving away from compiled regex patterns to simpler ResourceAction objects for better client compatibility.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authorization/entity · high confidence

New example configuration for cluster deployment

Added common.runtime.properties and log4j2.xml files for the cluster example, providing a ready-to-use configuration for a multi-node Druid cluster. The properties file configures local Derby for metadata storage, local filesystem for deep storage, and enables the ServiceStatusMonitor. The log4j2.xml configures rolling log files, routes task logs to dedicated files, and sets specific log levels for components like QueryResource and KafkaSupervisors.

_examples/conf/druid/cluster/\common · high confidence

New example scripts and utilities for managing Druid nodes and tasks

A new set of executable scripts has been added to the examples/bin directory to help users manage and interact with a Druid cluster. This includes node management scripts (broker, coordinator, historical, middleManager, overlord, router) that wrap a common node.sh launcher, alongside utility scripts for running the Druid server (run-druid, run-java, java-util), connecting to the cluster (jconsole.sh), and submitting indexing tasks (post-index-task with Python 2 and 3 versions). Additionally, a new generate-example-metrics script is provided to simulate web traffic for testing, and a dump-segment tool is added to inspect segment metadata.

examples/bin · high confidence

New single-server example configurations for micro, medium, and large deployments

Added new example configuration sets for single-server deployments, each tailored to a specific resource tier: micro, medium, and large. These configurations provide pre-tuned settings for all Druid node types (broker, coordinator, overlord, historical, middleManager, router) with appropriate JVM memory limits, thread counts, and buffer sizes. The common configuration enables the ServiceStatusMonitor, configures local Derby and storage for quickstarts, and sets up logging with rolling files and task-specific log routing. These examples serve as starting points for users to understand and adapt Druid's configuration for different scale levels.

examples/conf/druid/single-server · high confidence

Refactored lookup extension to support loading and polling cache strategies

The lookups-cached-single extension was refactored to introduce a new loading lookup strategy alongside the existing polling approach. This change adds a \DataFetcher\ interface and \LoadingLookup\/\PollingLookup\ classes that manage cache lifecycles and data retrieval. It also introduces \LoadingCache\ and \PollingCache\ abstractions with both on-heap (Guava-based) and off-heap (MapDB-based) implementations, allowing users to configure memory-efficient lookups. The \LookupExtractionModule\ was updated to register these new factory types, enabling dynamic selection of cache behavior based on configuration.

extensions-core/lookups-cached-single · high confidence

Register new DataSketches aggregation modules

The Druid application now automatically loads several new aggregation modules from the DataSketches library, including KllSketch, HllSketch, and Theta sketches. This is achieved by adding a new service provider file that registers the corresponding Java classes (e.g., KllSketchModule, HllSketchModule) with the Druid initialization system, enabling these statistical aggregation capabilities out-of-the-box.

extensions-core/datasketches/src/main/resources · high confidence

Repository configuration and compliance files added

The repository now includes essential configuration and compliance files to support the Apache Software Foundation (ASF) release process and community workflows. This includes \.asf.yaml\ for GitHub email subject formatting and notification settings, \.backportrc.json\ to define backport branches, \.codecov.yml\ to configure code coverage reporting, and standard repository files like \.gitignore\, \.dockerignore\, \.lgtm.yml\, \AGENTS.md\ (developer guidelines), \CONTRIBUTING.md\ (contribution guidelines), \LICENSE\, and \NOTICE\ files. These changes establish the necessary infrastructure for community contributions, automated testing, and legal compliance.

(repo-wide) · high confidence

SQL support for Array of Doubles Sketch aggregations and set operations

Users can now use SQL to aggregate and combine Array of Doubles Sketches. This change adds SQL operator conversions and aggregators for the DS\_TUPLE\_DOUBLES sketch, enabling the use of DS\_TUPLE\_DOUBLES in SELECT and aggregate clauses, as well as set operations (UNION, INTERSECT, NOT) on these sketches directly in SQL queries.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/tuple/sql · medium confidence

SQL support for Theta sketch aggregations and set operations

Users can now use Theta sketch aggregations and set operations directly in SQL queries. This change introduces SQL support for the \APprox\_COUNT\_DISTINCT\_DS\_THETA\ aggregator, as well as \theta\_sketch\_estimate\ and \theta\_sketch\_estimate\_with\_error\_bounds\ expressions. Additionally, SQL operators for Theta sketch set operations (UNION, INTERSECT, and NOT) are now available, allowing users to perform set-based analytics on sketch data directly within their SQL queries.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/theta/sql · high confidence

Support SQL expressions in Bloom filter queries

Users can now use SQL expressions and virtual columns within Bloom filter conditions. The new \BloomFilterOperatorConversion\ class enables the conversion of SQL expressions and virtual columns into \BloomDimFilter\ instances, allowing for more flexible filtering capabilities in historical query processing.

extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/query/filter/sql · medium confidence

Support for Kafka headers, key, and partition metadata in ingestion

The Kafka indexing service now supports ingesting Kafka headers, message keys, and partition metadata as columns in the resulting segments. This change introduces new classes (KafkaRecordEntity, KafkaInputFormat, etc.) that allow users to configure how headers, keys, and other Kafka record attributes are mapped to Druid columns, enabling richer data enrichment from Kafka messages.

extensions-core/kafka-indexing-service · high confidence

Support for Map-type virtual columns in SQL queries

Added support for Map-type virtual columns, enabling users to query map-typed columns in SQL. This includes the new MapVirtualColumn class and its associated dimension and value selectors (MapTypeMapVirtualColumnDimensionSelector, StringTypeMapVirtualColumnDimensionSelector), which allow retrieving map values by key or as a full map object. The change also registers the MapVirtualColumn in the DruidVirtualColumnsModule for Jackson serialization.

extensions-contrib/virtual-columns/src/main/java · high confidence

Website build and documentation infrastructure updates

The website build system and documentation tooling have been updated. A new \.spelling\ dictionary file was added to the \website\ directory to support spell-checking of documentation content. Additionally, a \link-lint.js\ script was introduced to validate internal links and anchors within the generated documentation, and a \notify-spellcheck-issues\ script was added to report spelling errors. The \docusaurus.config.js\ was updated to configure the Docusaurus 3 static site generator, including sidebar navigation, theme settings, and redirect rules, while custom CSS and JavaScript for code block copy buttons were added to improve the user experience on the documentation site.

website · high confidence

Removals

Removal of deprecated batch ingestion and legacy query modes

The server module has removed support for the Hadoop ingestion engine and the Appenderator-based batch ingestion tasks. Additionally, the deprecated 'native' and 'group-by v1' query modes have been removed, and the 'isDescending' flag has been removed from the Query interface in favor of specific implementations like TimeseriesQuery. These changes simplify the ingestion and query execution paths by eliminating legacy code paths.

server · high confidence

Remove support for Hadoop-based batch ingestion

The Hadoop ingestion task type has been removed from the processing engine. Users relying on Hadoop for batch data ingestion must migrate to alternative ingestion methods, such as the Multi-Stage Query (MSQ) engine or direct ingestion tasks, as the legacy Hadoop integration is no longer supported.

processing · high confidence

Security

Update Apache Kafka to version 3.9.1 to mitigate [CVE redacted]

The Apache Kafka client dependency has been upgraded to version 3.9.1. This update addresses the security vulnerability [CVE redacted], ensuring that the Kafka indexing service and related components operate with a patched version of the Kafka client library.

(dependencies) · high confidence

Architecture

Consolidate service entry points and configuration into the services module

The CLI entry points for all core Druid services (Broker, Coordinator, Historical, Indexer, MiddleManager) and their associated Guice modules have been moved into the \services\ module. This change centralizes the service startup logic, including the \run.sh\ script and Java classes like \CliBroker\ and \CliCoordinator\, ensuring that service-specific configurations and dependencies are managed within a single module rather than being scattered across the codebase.

services · high confidence

Behavioural changes

Add HadoopFsWrapper to handle Hadoop FileSystem.rename compatibility

A new wrapper class, HadoopFsWrapper, has been added to the hdfs-storage extension to invoke Hadoop's FileSystem.rename method via reflection. This change ensures compatibility across different Hadoop jar versions by avoiding direct calls to the protected rename method, which previously caused classloader issues. The wrapper specifically uses the Rename.NONE option to prevent temporary directories from being incorrectly moved into segment directories during replication tasks.

extensions-core/hdfs-storage/src/main/java/org/apache/hadoop · medium confidence

Added license files for third-party dependencies

The project now includes license text files for various third-party dependencies, including Apache 2.0, MIT, BSD-3-Clause, EPL-1.0, and OFL licenses. These files, located in the \licenses/\ directory, ensure compliance with the licensing terms of libraries such as Babel, Emotion, and others used in the codebase.

licenses · high confidence

Configure Kafka extension dependencies and module registration

The Kafka extraction namespace extension now explicitly declares its dependency on the druid-lookups-cached-global extension via a new druid-extension-dependencies.json file, and registers its module implementation in the service loader file to ensure proper initialization.

extensions-core/kafka-extraction-namespace/src/main/resources · high confidence

Expanded Prometheus metrics for queries, ingestion, and cluster state

The Prometheus emitter now exposes a comprehensive set of new metrics for monitoring query performance, ingestion pipelines, and cluster state. Key additions include detailed query metrics (e.g., \query/segment/time\, \query/segmentAndCache/time\), ingestion tracking (e.g., \ingest/rows/published\, \ingest/realtime/segmentUpgrade/\\), and segment lifecycle metrics (e.g., \segment/assigned/count\, \segment/loadQueue/\\). The configuration also adds Kafka consumer metrics (\kafka/consumer/\\), groupBy query metrics (\groupBy/\\), and task execution metrics (\task/\*\). This provides deeper observability into data flow, task health, and cluster balancing.

extensions-contrib/prometheus-emitter/src/main/resources · high confidence

GCP initialization is now lazy and modularized

The GCP module has been moved to the core extension and restructured to support lazy initialization. This change introduces a new \GcpModule\ that provides HTTP transport and JSON factory components as lazy singletons, allowing the system to defer GCP client setup until actually needed. A corresponding mock module and unit test have been added to verify the new initialization behavior.

cloud/gcp-common · medium confidence

HLL sketch build and merge aggregators support string encoding and vectorization

The HLL sketch aggregators in the datasketches extension now support configurable string encoding and vectorized processing. Users can specify the string encoding used when building sketches from string data, and the new implementations include optimized buffer and vector aggregators for improved performance.

extensions-core/datasketches/src/main/java/org/apache/druid/query/aggregation/datasketches/hll · high confidence

Migrate InfluxDB line protocol parsing to the new InputFormat interface

The InfluxDB extension has been refactored to use the new InputFormat interface, replacing the previous InfluxParser/InfluxParseSpec approach. This change introduces InfluxInputFormat and InfluxLineProtocolReader to handle InfluxDB line protocol data, allowing users to ingest InfluxDB metrics into Druid using the modern input format API. The extension also includes an ANTLR grammar for parsing the line protocol and registers the module via the standard service loader mechanism.

extensions-contrib/influx-extensions · high confidence

Moved Ranger security extension to contrib

The Ranger security extension has been moved from the core to the contrib directory. This change relocates the Ranger authorizer implementation, its Guice module, and associated test resources to the extensions-contrib/druid-ranger-security directory, making it an optional, community-maintained extension rather than a core component.

extensions-contrib/druid-ranger-security · high confidence

Prometheus emitter gains metric TTL and custom histogram bucket support

The Prometheus emitter now supports expiring stale metrics via a configurable TTL (flushPeriod) and allows defining custom histogram buckets for timer metrics. Additionally, it supports extra labels and optional removal of metrics from the PushGateway on shutdown.

extensions-contrib/prometheus-emitter/src/main/java · medium confidence

Redis cache extension refactored with TLS and cluster support

The Redis cache extension has been refactored to support both standalone and cluster topologies, with new configuration options to enable TLS encryption and skip hostname verification. The \RedisCacheConfig\ now includes \enableTls\ and \skipTlsHostnameVerification\ properties, and the \RedisCacheFactory\ builds \Jedis\ clients with \SslOptions\ when TLS is enabled. Additionally, \RedisClusterCache\ implements a slot-based \mget\ workaround for Jedis cluster limitations, and the provider registers subtypes via Jackson modules.

extensions-contrib/redis-cache/src/main/java · medium confidence

Refactor HDFS storage implementation to use new classes

The HDFS storage extension has been refactored to use new implementation classes for core storage operations. HdfsDataSegmentPusher, HdfsDataSegmentPuller, and HdfsDataSegmentKiller now handle segment push, pull, and deletion respectively, with updated configuration via HdfsDataSegmentPusherConfig. Additionally, HdfsStorageAuthentication and HdfsStorageAvailabilityChecker manage Kerberos-based security and cluster availability checks, while HdfsFileTimestampVersionFinder and HdfsLoadSpec support versioned data finding and loading. These changes streamline how Druid interacts with HDFS for deep storage and task logs.

extensions-core/hdfs-storage/src/main/java/org/apache/druid/storage/hdfs · high confidence

Refactor basic authentication cache into a pluggable interface

The basic authentication cache implementation has been refactored to use a new \BasicAuthenticatorCacheManager\ interface, separating the cache logic from the notifier. This introduces a \CoordinatorPollingBasicAuthenticatorCacheManager\ for non-coordinator services to poll the coordinator for updates, and a \MetadataStoragePollingBasicAuthenticatorCacheManager\ for the coordinator itself, allowing for more flexible configuration of how authentication state is cached and updated.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authentication/db/cache · medium confidence

Refactor basic security extension to use org.apache.druid package and introduce LDAP user principal

The basic security extension's authentication and escalation components, including BasicHTTPAuthenticator, BasicHTTPEscalator, and the new LdapUserPrincipal, have been migrated from the io.druid to the org.apache.druid package namespace. This change updates the internal structure of the basic security extension, introducing a new LdapUserPrincipal class to handle LDAP user state and verification, while the existing basic HTTP authentication and escalation logic is reorganized under the new package path.

extensions-core/druid-basic-security/src/main/java/org/apache/druid/security/basic/authentication · medium confidence

Refactored Azure storage and input source implementations

The Azure extension's storage and input source logic has been refactored to use new, dedicated classes such as AzureEntity, AzureInputSource, and AzureStorageAccountInputSource. This change introduces a cleaner separation between input sources and storage clients, adds support for system fields (URI, bucket, path) on input entities, and updates the client factory to support managed identity and credentials chain authentication. The refactoring also includes configuration classes for input sources and segment management, improving how Druid interacts with Azure Blob Storage for both data ingestion and deep storage operations.

extensions-core/azure-extensions/src/main · high confidence

Register Bloom Filter serialization and SQL integration in Guice module

The BloomFilterExtensionModule now registers Jackson serializers and deserializers for BloomKFilter and BloomKFilterHolder, and registers the BloomFilterSerde with ComplexMetrics. Additionally, the module binds the BloomFilterOperatorConversion for SQL filtering, the BloomFilterSqlAggregator for SQL aggregation, and registers the Create, Add, and Test expression macros for Bloom filters.

extensions-core/druid-bloom-filter/src/main/java/org/apache/druid/guice · high confidence

Register BloomFilterExtensionModule for Druid initialization

The extension now registers the BloomFilterExtensionModule via the Java SPI mechanism, enabling the Druid runtime to automatically discover and initialize the bloom filter functionality. This change adds the necessary service provider configuration file to ensure the extension is loaded correctly at startup.

extensions-core/druid-bloom-filter/src/main/resources · high confidence

SQL Server metadata storage handler overrides SQL dialect for limit clauses

The SQL Server-specific implementation of the metadata storage action handler now overrides the method that appends limit clauses to SQL queries. Instead of using the standard 'LIMIT' syntax, it generates 'SELECT TOP (:n)' syntax, which is required for Microsoft SQL Server. This ensures that metadata storage operations (such as task and lock table interactions) execute correctly on SQL Server databases.

extensions-contrib/sqlserver-metadata-storage/src/main/java/org/apache/druid/metadata · high confidence

Update service loader files to use new package names

The META-INF/services/org.apache.druid.initialization.DruidModule files across multiple extensions (including Cassandra, Distinct Count, Kafka Emitters, Redis Cache, SQL Server Metadata Storage, StatsD Emitters, Thrift Extensions, Time Min/Max, Virtual Columns, Avro Extensions, Basic Security, Kerberos, HDFS Storage, and Histograms) have been updated to reference the new org.apache.druid package structure instead of the previous io.druid namespace. This change ensures that the Java ServiceLoader can correctly locate and initialize the corresponding DruidModule implementations for each extension.

(repo-wide) · high confidence

Updated Druid module registration for Apache migration

The service registration file for the CloudFiles storage extension has been updated to reflect the package rename from io.druid to org.apache.druid. This ensures the CloudFilesStorageDruidModule is correctly discovered and initialized by the Apache Druid framework.

extensions-contrib/cloudfiles-extensions/src/main/resources · high confidence

Updated Hadoop 3.3.6 tutorial Docker environment

The Hadoop quickstart tutorial's Docker environment has been updated to use Hadoop 3.3.6 and Zulu OpenJDK 11, replacing the previous Java 8 setup. The Dockerfile now installs Zulu 11, configures Hadoop environment variables, and sets up SSH and Hadoop services. Configuration files (core-site.xml.template, hdfs-site.xml, mapred-site.xml, yarn-site.xml) and the bootstrap script have been added to support the Hadoop 3.x architecture, including YARN and MapReduce settings.

examples/quickstart/tutorial/hadoop · high confidence

Test coverage

Add Druid testcontainers for running Docker-based tests; Added Avro schema test fixture for complex data types; Added HllSketch test resources; Added KLL sketch test data for doubles and floats; Added SQL test coverage for HllSketch aggregators; Added SQL tests for histogram quantile aggregations; Added comprehensive tests for KafkaEmitter configuration and event handling; Added embedded tests for basic and LDAP authentication configurations; Added new JMH benchmark suite for the benchmarks module; Added sample data for histogram tests; Added test coverage for ArrayOfDoublesSketch SQL aggregation; Added test data files for the Compressed BigDecimal extension; Added test data for ArrayOfDoublesSketch and bucketed aggregations; Added test data for numeric quantiles sketch aggregator; Added tests for Aliyun OSS input source; Added tests for Aliyun OSS storage components; Added tests for ArrayOfDoublesSketch aggregations and post-aggregators; Added tests for CompressedBigDecimal aggregation and serialization; Added tests for DoublesSketch aggregators and post-aggregators; Added tests for HllSketch aggregators and utilities; Added tests for Iceberg filter implementations; Added tests for KllDoublesSketch aggregations and post-aggregators; Added tests for MapVirtualColumn in GroupBy and TopN queries; Added tests for RabbitMQ stream indexing configuration and record handling; Added tests for RabbitMQ stream supervisor configuration and behavior; Added tests for SQL-based quantile aggregation on double sketch data; Added tests for Theta Sketch SQL aggregators; Added tests for Thrift input format parsing; Added tests for compressed BigDecimal aggregators; Added tests for datasketches projection functionality; Added tests for t-digest sketch aggregators; Added tests for the Druid Catalog subsystem; Added tests for the Spectator Histogram extension; Added tests for the approximate histogram aggregation feature; Added tests for the druid-pac4j security extension; Added tests for theta sketch aggregations and post-aggregators; Added tests for timestamp min/max aggregations; Added unit and integration tests for the Kafka lookup extension; Added unit tests for AWS client configuration and utility classes; Added unit tests for AvroStreamInputFormat; Added unit tests for Azure input source components; Added unit tests for Azure storage components; Added unit tests for CloudFiles storage components; Added unit tests for DropwizardEmitterConfig serialization and deserialization; Added unit tests for Google Cloud Storage input source; Added unit tests for Google Cloud Storage output components; Added unit tests for HDFS task log operations; Added unit tests for Iceberg input extension components; Added unit tests for Kerberos authentication components; Added unit tests for Redis cache configuration and cache implementations; Added unit tests for SQL Server metadata storage components; Added unit tests for ToObjectVectorColumnProcessorFactory; Added unit tests for distinct count aggregation across query types; Added unit tests for the Druid Delta Lake extension; Added unit tests for the Moving Average query extension; Added unit tests for the OpenTSDB emitter; Added unit tests for the Prometheus emitter configuration and metrics handling; Web console test infrastructure and configuration scaffolding.

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 45.

Lenses

  • Code Health 76
  • Architecture 56
  • Maturity 75
  • Readiness 52
  • Security 41
  • Domain Modelling 100
  • Accessibility 38

Changes since last survey

  • 300 commits — 227 feature/other, 73 fixes

By area

  • processing/src — 52 commits
  • server/src — 50 commits
  • (root) — 47 commits
  • indexing-service/src — 31 commits
  • website/package-lock.json — 14 commits
  • web-console/package-lock.json — 12 commits
  • embedded-tests/src — 9 commits
  • extensions-core/s3-extensions — 8 commits
  • web-console/src — 8 commits
  • extensions-core/kafka-indexing-service — 5 commits
  • sql/src — 5 commits
  • .github/workflows — 4 commits
  • docs/tutorials — 4 commits
  • multi-stage-query/src — 4 commits
  • benchmarks/pom.xml — 2 commits
  • cloud/aws-common — 2 commits
  • distribution/bin — 2 commits
  • docs/ingestion — 2 commits
  • examples/bin — 2 commits
  • extensions-contrib/thrift-extensions — 2 commits

Notable commits

  • fix: Fix record supplier lock release (#19811)
  • fix: ci: fix unit tests, bump actions-timeline (#19509)
  • fix: fix(delta): drain all batches per scan file in DeltaInputSourceIterator (#19592)
  • fix: fix(test): wait for Kafka partitions before publishing (#19817)
  • fix: fix: Avoid sleep on stop in K8sDruidNodeDiscoveryProvider. (#19748)
  • fix: fix: Cast for comparison for LONG + IN. (#19834)
  • fix: fix: Close exposure to potential NPE in S3Utils.useHttps (#19504)
  • fix: fix: Emit publish metrics only when tasks actually publish. (#19395)
  • fix: fix: Fix clone historicals and inter-historical move handling of partial load profiles (#19843)
  • fix: fix: Fix for flaky embedded test QueryVirtualStorageTest
  • fix: fix: Fix invalid time comparison in SqlSegmentsMetadataManager (#19603)
  • fix: fix: Fix race in LocalIntermediaryDataManager.addSegment (#19446)
  • fix: fix: Fix undesireable load rule behavior for clustered segments with conforming shape but no matched groups (#19728)
  • fix: fix: Improve reserved keyword SQL parse errors (#19719)
  • fix: fix: Limit number of unused segment rows scanned (#19690)
  • fix: fix: Make sys.server_properties table filterable and resilient to unreachable servers (#19459)
  • fix: fix: OrcInputFormat concurrent FileSystem init race condition (#19491) (#19497)
  • fix: fix: Peon metrics emit same value (#19484)
  • fix: fix: Perform best effort cleanup of intermediary data in ParallelIndexSupervisorTask (#19647)
  • fix: fix: Raise start-druid middleManager memory to 256m (#19582)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/druid was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 6 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 594b1694c28c7a7f550acb58c93412be0c3c75b7 — the exact code this score is about.
  • Scored under rubric-2026.08.19 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer latest.