Skip to content
CAI
Software that uses CAICheck a score

apache/kyuubi

46.8

Weak · 27 September 2026

141k

lines of production code

Scala

with Java, TypeScript

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Apache Kyuubi is a multi-engine SQL gateway that provides a unified interface for executing SQL queries across diverse backends, including Spark, Flink, Hive, Trino, and various JDBC databases. It manages session lifecycles, enforces fine-grained column-level authorization via Apache Ranger, and supports advanced SQL optimizations and lineage tracking. The system offers robust operational tooling, including a REST API, CLI, and web UI, alongside comprehensive monitoring and high-availability features.

How it got here

2017–2021 — Apache Kyuubi project initialization and core architecture

61 changes.

This period established the foundational structure of Apache Kyuubi, including ASF compliance, build infrastructure, and core service abstractions for sessions, operations, and authentication. It introduced the initial implementations for Spark, Flink, and Trino SQL engines, alongside essential tooling for CLI management, metrics, and release automation.

2022 — New engine integrations and security extensions

59 changes.

This period focused on expanding Kyuubi's backend support by introducing dedicated SQL engines for Hive, Flink, and a multi-dialect JDBC engine, alongside comprehensive integration tests for each. It also delivered significant security and observability features, including column-level authorization via Apache Ranger, SQL lineage tracking with Atlas, and a pluggable event logging system. Additionally, the work established a new Web UI, a Docker-based playground for development, and a Java REST client SDK to improve usability and deployment flexibility.

2023–2026 — YARN deployment, Spark 4.x support, and authorization hardening

44 changes.

This period focused on expanding deployment capabilities by introducing YARN support for Kyuubi engines and adding comprehensive SQL extensions for Spark 4.0, 4.1, and 4.2. Significant work was also dedicated to strengthening security through a major refactor of the Spark authorization plugin, including stricter configuration checks and row-level filtering, alongside enhancements to LDAP authentication and the introduction of a new Data Agent Engine.

Features

Add Kyuubi Hive JDBC dialect for Spark

This change introduces a new JDBC dialect extension for Spark that enables connecting to Kyuubi via the JDBC source. The dialect recognizes both \jdbc:hive2://\ and \jdbc:kyuubi://\ connection URLs and adapts Spark SQL data types to their corresponding Hive definitions (e.g., mapping \IntegerType\ to \INT\ and \DoubleType\ to \DOUBLE\) to ensure compatibility with Hive 2.2.0+ type synonyms. It also provides custom identifier quoting logic. The extension is registered via \KyuubiSparkJdbcDialectExtension\ and includes unit tests verifying URL handling, identifier quoting, and type mapping.

extensions/spark/kyuubi-extension-spark-jdbc-dialect · high confidence

Add Spark 4.2 SQL extension with optimizer rules and configuration support

This change introduces the Kyuubi SQL extension module for Spark 4.2, providing the necessary grammar, parser, and optimizer rules to support this version. It includes a SQL grammar definition for OPTIMIZE ... ZORDER statements, a parser to build logical plans for these commands, and several optimizer rules: \DropIgnoreNonexistent\ to silently ignore drops of non-existent objects, \DynamicShufflePartitions\ to adjust shuffle partitions based on data size, \InferRebalanceAndSortOrders\ to automatically infer partitioning and sorting columns from join keys (including cast-wrapped keys), \InsertShuffleNodeBeforeJoin\ to ensure shuffle nodes exist for skewed join optimization, \KyuubiEnsureRequirements\ to manage distribution and ordering requirements, and \FinalStageConfigIsolation\ to allow different configuration for the final query stage. The module also defines the associated SQL configurations (e.g., \spark.sql.optimizer.\*\) and exception classes.

extensions/spark/kyuubi-extension-spark-4-2 · high confidence

Add TPC-DS benchmark tool and data generator

The \dev/kyuubi-tpcds\ module now includes a TPC-DS benchmark tool and data generator. Users can generate TPC-DS data using the \DataGenerator\ class with configurable scale factors and parallelism, and run benchmarks using the \RunBenchmark\ class. The benchmark tool supports options to include or exclude specific queries, set execution modes, and record detailed performance breakdowns.

dev/kyuubi-tpcds · high confidence

Add Zstd compression support for Arrow IPC query results

Users can now enable Zstandard (Zstd) compression for Arrow IPC query results, reducing network bandwidth usage and potentially improving query performance for large result sets. This change introduces a new \KyuubiArrowCompressionSupport\ module to isolate the optional Arrow compression dependency, ensuring the uncompressed code path remains free of it. The \KyuubiArrowConverters\ have been updated to support Zstd encoding and decoding, allowing clients to request compressed batches via the 'zstd' codec name.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/sql/execution/arrow · high confidence

The Kyuubi Flink SQL engine now supports running in Flink's YARN application mode. This is enabled by adding a service provider file for \org.apache.flink.core.execution.PipelineExecutorFactory\ that registers \EmbeddedExecutorFactory\, allowing the engine to be deployed as a YARN application. The change also includes the standard Apache 2.0 LICENSE and NOTICE files for the module.

externals/kyuubi-flink-sql-engine/src/main/resources · high confidence

Added dev scripts for regenerating golden files

New shell scripts have been added to the dev/gen directory to automate the regeneration of golden files and documentation. These scripts cover configuration documentation, Hive and Spark UDF documentation, Ranger policy and spec JSON files, TPC-DS output schemas, and TPC-DS/TPC-H query resources, ensuring these generated assets stay synchronized with the codebase.

dev/gen · high confidence

Added templates for release vote cancellation and result announcements

The release process now includes new email templates for handling vote outcomes. A template for canceling a release vote has been added, allowing release managers to notify the community of a cancellation with a link to the relevant issue. Additionally, a template for announcing the result of a passed vote has been introduced, providing a structured format to list binding and non-binding votes and link to the vote thread.

build/release/tmpl · high confidence

Data Agent Engine introduces datasource tooling, prompt composition, and pluggable LLM providers

The Data Agent Engine now includes the core infrastructure to connect to SQL datasources and interact with LLMs. This adds a datasource layer with a HikariCP-backed connection factory and dialect-specific SQL identifier quoting for Spark, Trino, MySQL, and SQLite, alongside a \TableRef\ model for normalizing table references across these systems. Prompt generation is handled by a \SystemPromptBuilder\ that composes base templates with datasource-specific guidelines and tool descriptions. The engine runtime features a \ReactAgent\ with a middleware pipeline (logging, compaction, approval) and a pluggable provider interface (\DataAgentProvider\) that includes a production-ready OpenAI-compatible \ChatCompletionProvider\ and a test \EchoProvider\ for simulating agent events.

externals/kyuubi-data-agent-engine · high confidence

The Flink SQL engine now includes core schema handling and result-set generation capabilities. New files in the schema package introduce FlinkTRowSetGenerator to convert Flink Row data into Thrift TRowSet structures, supporting types such as LocalZonedTimestamp, VarChar, Char, and various numeric types. RowSet provides utilities to map Flink LogicalTypes to Hive-compatible TTypeDesc and TTypeId definitions, including qualifiers for decimals and character lengths, and handles string representations for complex types like arrays, maps, and structs. SchemaHelper adds methods to query Flink catalogs for schemas and tables using pattern matching. This enables the engine to properly describe result sets and interact with catalog metadata.

externals/kyuubi-flink-sql-engine/src/main/scala/org/apache/kyuubi/engine/flink/schema · high confidence

The Flink SQL engine now registers several Kyuubi-defined functions (UDFs) accessible in SQL queries, including kyuubi\_version, kyuubi\_engine\_name, kyuubi\_engine\_id, kyuubi\_system\_user, and kyuubi\_session\_user, allowing users to retrieve engine and session metadata directly. Additionally, the engine has expanded its supported Flink runtime versions to include 1.20, 2.0, 2.1, 2.2, and 2.3, with Flink 2.0 specifically marked as deprecated. The engine also implements delegation token renewal capabilities and supports initializing SQL statements at startup.

externals/kyuubi-flink-sql-engine/src/main/scala/org/apache/kyuubi/engine/flink · high confidence

The Flink SQL engine now provides a structured result-handling layer and interactive command support. A new \IncrementalResultFetchIterator\ enables timeout-aware, incremental fetching of query results, while \ResultSet\ and \ResultSetUtil\ standardize how results are built and returned. Additionally, a \HELP\ command is now available to display a formatted list of supported Flink SQL commands and usage hints.

externals/kyuubi-flink-sql-engine/src/main/scala/org/apache/kyuubi/engine/flink/result · high confidence

The Kyuubi Flink extension now includes a new delegation token provider and receiver (\KyuubiDelegationTokenProvider\ and \KyuubiDelegationTokenReceiver\) that enable user impersonation. This implementation allows the Flink engine to obtain, renew, and manage Hadoop delegation tokens based on credentials passed from Kyuubi, ensuring that jobs run with the correct user identity for secure cluster access.

extensions/flink/kyuubi-flink-token-provider · high confidence

Hive SQL Engine event logging support

The Hive SQL engine now emits structured lifecycle events for sessions, operations, and the engine itself, logged as JSON files. This enables users to track session activity, query execution details, and engine state changes for auditing and debugging purposes.

externals/kyuubi-hive-sql-engine/src/main/scala/org/apache/kyuubi/engine/hive/events · high confidence

Hive SQL engine now supports deployment on YARN

Users can now run the Hive SQL engine on YARN clusters. This change introduces the HiveYarnModeSubmitter, which handles the submission process, including loading the correct Hadoop and Hive classpaths and configuration files. It also enforces Kerberos security requirements by mandating that the engine principal and keytab are configured when security is enabled on the YARN cluster.

externals/kyuubi-hive-sql-engine/src/main/scala/org/apache/kyuubi/engine/hive/deploy · high confidence

Hive engine operation handlers for metadata and statement execution

The Hive SQL engine now includes dedicated operation handlers for executing statements and retrieving metadata (catalogs, schemas, tables, columns, functions, primary keys, cross-references, and type info). A base HiveOperation class and HiveOperationManager coordinate these operations with the underlying Hive service, including support for operation logs and query IDs.

externals/kyuubi-hive-sql-engine/src/main/scala/org/apache/kyuubi/engine/hive/operation · high confidence

The Flink SQL engine now supports a comprehensive set of JDBC metadata and statement execution operations, including executing SQL statements (with plan-only mode support), retrieving catalogs, schemas, tables, columns, primary keys, and functions, as well as managing the current catalog and database. This change introduces the core operation classes and the \FlinkSQLOperationManager\ to handle these requests, enabling clients to interact with Flink SQL through standard metadata APIs.

externals/kyuubi-flink-sql-engine/src/main/scala/org/apache/kyuubi/engine/flink/operation · high confidence

Initial implementation of the Hive SQL engine session layer

This change introduces the core session management components for the new Hive SQL engine in Kyuubi. It adds \HiveSessionImpl\ to handle session lifecycle events (open/close), register UDFs, and expose server metadata (such as SQL keywords and version info) via the GetInfo API. It also adds \HiveSessionManager\ to create and manage underlying Hive sessions, supporting both Hive 2.3 and 3.1+ runtime versions through reflection, respecting the \hive.server2.enable.doAs\ configuration for impersonation, and stopping the engine when the session level share mode is set to CONNECTION.

externals/kyuubi-hive-sql-engine/src/main/scala/org/apache/kyuubi/engine/hive/session · high confidence

Initial release of the Kyuubi Python client (PyHive)

This change introduces the \python\ directory as a standalone package, effectively renaming the previous \pyhive\ module to \python\. It includes the core PyHive library code (DB-API and SQLAlchemy interfaces for Hive, Presto, and Trino), auto-generated Thrift service definitions, and a Docker Compose environment for local testing. The package is licensed under Apache 2.0 and includes development dependencies for testing and linting.

python · high confidence

Initial support for PySpark via a dedicated Python Gateway Server

Kyuubi now supports PySpark workloads by introducing a new KyuubiPythonGatewayServer component. This server initializes a Py4J gateway, binds to a specific port, and writes connection details (port and secret) to a temporary file so that Python processes can securely connect to the Spark engine. It also includes logic to gracefully shut down the gateway server during engine termination.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/api · high confidence

Initial support for executing Python code in Kyuubi

Added Python execution scripts (execute\_python.py and kyuubi\_util.py) that enable users to run Python code and interact with Spark sessions through Kyuubi. The new scripts handle connecting to the existing Py4J gateway, initializing Spark sessions, and parsing/executing Python code blocks, including support for notebook-style magic commands when enabled.

externals/kyuubi-spark-sql-engine/src/main/resources/python · high confidence

Introduce Dockerfile-based Kyuubi Playground images

This change adds a set of Dockerfiles (base, Hadoop, Spark, Hive Metastore, and Kyuubi) to the \docker/playground/image\ directory, establishing the foundational images for the Kyuubi Playground. These images provide a pre-configured environment for testing Kyuubi, including the installation of Java 8 and 17, Hadoop, Spark, Hive, and necessary connectors (Iceberg, TPC-DS, TPC-H, AWS SDK, PostgreSQL JDBC). The base image sets up the OS and Java runtime, while subsequent layers add specific components like Hadoop and Spark, culminating in the Kyuubi image which installs the Kyuubi server and its dependencies.

docker/playground/image · high confidence

Introduce JVM Quake plugin for Spark driver and executor monitoring

This change adds a new Spark plugin (available from version 1.10.0) that monitors JVM garbage collection behavior on both drivers and executors. When enabled via \spark.driver.jvmQuake.enabled\ or \spark.executor.jvmQuake.enabled\, the plugin periodically calculates a 'quake' value based on GC time and run time; if this value exceeds the configurable \spark.jvmQuake.killThreshold\, the process is terminated. Optionally, if heap dumping is enabled (\spark.driver.jvmQuake.heapDump.enabled\ or \spark.executor.jvmQuake.heapDump.enabled\), a heap dump is saved to the path specified by \spark.jvmQuake.heapDumpPath\ before the process exits.

extensions/spark/kyuubi-spark-jvm-quake · high confidence

Introduce Kyuubi Hive BeeLine module with SQL shell and JSON output support

Adds the kyuubi-hive-beeline module, providing a dedicated SQL shell for Kyuubi that includes command-line option parsing, interactive command completion, and multiple output formats including a new JSON format. The implementation introduces Kyuubi-specific hooks to track the current database in the prompt and integrates with Kyuubi's JDBC driver for connection management.

kyuubi-hive-beeline · high confidence

Introduce Kyuubi Spark AuthZ module with column-level fine-grained authorization

The Kyuubi Spark AuthZ module is introduced to provide column-level fine-grained authorization for Spark SQL queries. This new component defines core authorization primitives including \PrivilegeObject\, \PrivilegeObjectType\, \PrivilegeObjectActionType\, \OperationType\, and \ObjectType\ to model database, table, view, column, function, and URI resources. It includes an \AccessControlException\ for denials and a \PrivilegesBuilder\ that analyzes Spark logical plans to extract required privileges, supporting features like column pruning, row-level filtering, and data masking policies. The module integrates with Apache Ranger to enforce access control at the column level for various SQL operations.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz · high confidence

Introduce Kyuubi Spark TPC-DS Connector for in-memory benchmarking

This change adds the Kyuubi Spark TPC-DS Connector, a new extension that allows users to query the TPC-DS benchmark dataset directly in Spark without external storage. The connector implements a \TableCatalog\ and \BatchScan\ to generate data in-memory using the Trino TPC-DS generator, supporting multiple scale factors (including a new 'tiny' scale for quick testing) and parallelized batch reads. It provides configurable options to exclude specific databases, control ANSI string type usage (CHAR/VARCHAR vs STRING), and adjust partition sizes, while ensuring compatibility with Spark 4.2 through dynamic constructor handling for character types.

extensions/spark/kyuubi-spark-connector-tpcds/src/main/scala · high confidence

Introduce Kyuubi Spark TPC-H Connector for in-memory benchmarking

This change adds the initial implementation of the Kyuubi Spark TPC-H connector, enabling users to query TPC-H benchmark datasets directly within Spark without external storage. The connector introduces a new \TPCHCatalog\ that supports namespace listing and table resolution, backed by \TPCHBatchScan\ and \TPCHPartitionReader\ classes that generate data in-memory using the Trino TPC-H generator. It supports multiple data scales (from tiny 0.01 up to 100,000) and allows configuration via options such as \excludeDatabases\ and \useAnsiStringType\. The implementation includes schema utilities and statistics estimation to optimize query planning.

extensions/spark/kyuubi-spark-connector-tpch/src/main/scala · high confidence

Introduce Kyuubi Web UI with SQL Editor and Batch Management

The Kyuubi Web UI has been initialized with a new Vue 3 frontend, providing users with a SQL Editor for running queries and a Batch Management page to view and control batch jobs. The interface includes a login modal for authentication, a collapsible sidebar for navigation, and a dark/light theme toggle. This entry covers the frontend implementation of these pages and their corresponding API client code for interacting with the server.

kyuubi-server/web-ui · high confidence

Introduce Kyuubi-specific SQL parser for admin commands

The Kyuubi server now includes a dedicated SQL parser (KyuubiSqlBaseLexer and KyuubiSqlBaseParser) to handle Kyuubi-specific administrative commands such as DESCRIBE SESSION and DESCRIBE ENGINE. This parser intercepts statements prefixed with KYUUBI or KYUUBIADMIN, allowing the server to process these internal commands directly rather than passing them through to the underlying engine, while other SQL statements continue to be forwarded as before.

kyuubi-server · high confidence

Introduce Ranger-based authorization plugin for Spark SQL

This change adds a new Apache Ranger integration for the Kyuubi Spark authorization module, enabling table, column, and function-level access control via the \RangerSparkExtension\. It introduces core components including \AccessRequest\ and \AccessResource\ to map Spark operations to Ranger policies, and \RuleAuthorization\ to enforce privileges during query execution. The plugin supports batch verification of access requests for improved performance, allows overriding user groups via Ranger's UserStore, and integrates with Ranger's data masking and row-filtering capabilities.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/ranger · high confidence

Introduce Spark SQL Engine event tracking and logging infrastructure

The Spark SQL engine now emits structured events for engine lifecycle, user sessions, and individual SQL operations. New event classes (EngineEvent, SessionEvent, SparkOperationEvent) capture details such as application IDs, session names, client/server IPs, execution times, and CPU usage, while an event store and handler register these events for JSON logging and history server consumption. This enables administrators to monitor engine status, session statistics, and query performance directly from the Kyuubi UI and event logs.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/events · high confidence

Introduce Trino backend engine support

This change adds the initial implementation of the Trino backend engine for Kyuubi, enabling users to execute SQL queries against a Trino cluster. The new engine includes core components such as session management, statement execution with streaming data support, and a progress monitor that displays detailed stage-level query progress (splits, nodes, rows, and memory usage) in the operation logs. It also provides metadata introspection capabilities (GetCatalogs, GetColumns, GetSchemas, GetTables, etc.) and integrates with Kyuubi's event logging system for tracking engine, session, and operation lifecycles.

externals/kyuubi-trino-engine/src/main · high confidence

Introduce basic Hive SQL engine implementation

This change introduces the initial implementation of the Hive SQL engine within Kyuubi, adding core components such as HiveSQLEngine, HiveBackendService, and HiveTBinaryFrontendService to handle session management and Thrift binary protocol interactions. It includes support for delegating delegation token renewal from the Kyuubi server to the Hive engine, registers several Kyuubi-defined UDFs (like kyuubi\_version and engine\_name) for introspection, and ensures the engine properly recognizes hiveserver2-site.xml configurations for consistent security settings.

externals/kyuubi-hive-sql-engine/src/main/scala/org/apache/kyuubi/engine/hive · high confidence

Introduce dedicated event bus and extensible event logging handlers

The kyuubi-events module now provides a new EventBus that supports both synchronous and asynchronous event dispatching, allowing event handlers to be registered for non-blocking processing. It introduces a pluggable event handler system with built-in support for JSON and Kafka loggers, and enables the loading of custom event handlers via the Java ServiceLoader mechanism. The JSON logger also includes configurable file permission management for the logged event files.

kyuubi-events · high confidence

Introduce kyuubi-ctl command-line interface

The kyuubi-ctl tool is now available, providing a new command-line interface for managing Kyuubi resources. It supports creating, getting, listing, deleting, and submitting batch jobs, as well as creating, getting, listing, and deleting servers and engines. The tool also allows listing current sessions, fetching batch logs, and performing administrative actions like deleting engines, listing servers, and refreshing configurations. It uses a REST client for communication with the Kyuubi server, supporting basic and SPNEGO authentication, and allows configuration of connection and socket timeouts.

kyuubi-ctl · high confidence

Introduce kyuubi-util module with new utility classes

The new kyuubi-util module provides a collection of Java utility classes to support Kyuubi operations. IPStackUtils adds IPv4/IPv6 host:port parsing and formatting, while JavaUtils includes OS detection and local IP address resolution. SubjectUtil bridges JAAS Subject APIs across JDK versions (8 through 25+) to ensure compatibility with modern Java security deprecations. UuidUtils introduces RFC 9562 UUIDv7 generation, and the reflect package (DynClasses, DynConstructors, DynFields, DynMethods) offers dynamic class, constructor, field, and method handling adapted from upstream projects.

kyuubi-util · high confidence

Introduce kyuubi-util-scala with shared Scala utility modules

This change introduces the new kyuubi-util-scala module, consolidating shared Scala utilities into a dedicated library. It provides a SemanticVersion class for robust major/minor version parsing and comparison, EnumUtils for safe enumeration validation, and ReflectUtils for simplified reflection operations including safe class loadability checks and method/field access. Additionally, it adds CommandLineUtils for assembling and parsing command-line arguments with support for key-value pair generation and configuration redaction, along with a suite of test assertion helpers and golden-file utilities to support testing across the project.

kyuubi-util-scala · high confidence

Introduce new Java REST client SDK for Kyuubi

The kyuubi-rest-client module now provides a comprehensive Java SDK for interacting with the Kyuubi REST API. This new client includes dedicated API wrappers for administrative tasks (AdminRestApi), batch job management (BatchRestApi), session handling (SessionRestApi), and operation execution (OperationRestApi). It features a builder-based configuration for connection settings, authentication (Basic, SPNEGO, Custom), and timeouts, along with a retryable HTTP client that handles failover across multiple server URIs.

kyuubi-rest-client · high confidence

Introduce new Spark SQL engine implementation and utilities

This change introduces a new set of core components for the Spark SQL engine within Kyuubi. It adds the \SparkSQLEngine\ class which manages the engine lifecycle, including self-termination, lifetime checking, and graceful stopping. A new \SparkSQLBackendService\ and \SparkTBinaryFrontendService\ handle the backend session management and Thrift binary protocol interactions, including delegation token renewal. Utility functions in \KyuubiSparkUtil\ provide engine identification, URL extraction (supporting both Spark 3.x and 4.x YARN proxy configurations), and diagnostics. Additionally, event handling is updated with \SparkHistoryLoggingEventHandler\ and \SparkJsonLoggingEventHandler\ to integrate with Kyuubi's event bus, and a \DataFrameHolder\ is added to manage result passing for CLI operations.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark · high confidence

Introduce new metrics reporting framework with Prometheus as default

The kyuubi-metrics module now provides a new metrics system that exposes server-side metrics via multiple reporters, with Prometheus enabled by default. Users can configure reporters (Console, JMX, JSON, SLF4J, Prometheus) via kyuubi.metrics.reporters, and adjust intervals and paths (e.g., kyuubi.metrics.console.interval, kyuubi.metrics.json.location, kyuubi.metrics.prometheus.port and path). The Prometheus endpoint supports an optional instance label via kyuubi.metrics.prometheus.labels.instance.enabled. The system also registers JVM attributes (GC, memory, thread, class loading) and defines metric keys for connections, operations, engine startup, and backend service methods.

kyuubi-metrics · high confidence

Introduce new operation management infrastructure in kyuubi-common

This change introduces a new operation management layer in the kyuubi-common module, replacing the previous operation handling mechanisms. It adds core components including AbstractOperation, OperationManager, OperationHandle, and OperationState to manage the lifecycle of SQL operations. The update also introduces FetchIterator and FetchOrientation to support streaming data fetching, and adds PlanOnlyMode and PlanOnlyStyle to enable SQL lineage and query plan analysis features. Additionally, it includes Handle for operation identifier management and OperationAuditLogger for auditing operation state changes.

kyuubi-common/src/main/scala/org/apache/kyuubi/operation · high confidence

Introduce new session management architecture with dedicated handle and manager components

This change introduces a new session management layer in \kyuubi-common\, adding \SessionHandle\ for unique session identification, \SessionManager\ to orchestrate session lifecycle (open, close, idle timeout checking, and operation log directory cleanup), and \AbstractSession\ as the base implementation for session state and operation execution. The \Session\ trait defines the interface for session operations, while \package.scala\ establishes configuration prefixes (e.g., \hiveconf:\, \spark.\) for session configuration handling. This refactoring centralizes session state management, improves idle session cleanup, and ensures operation log directories are properly deleted upon session closure.

kyuubi-common/src/main/scala/org/apache/kyuubi/session · high confidence

Introduce pluggable service discovery client for high availability

The high availability module now supports pluggable service discovery backends, allowing users to choose between ZooKeeper and etcd via the \kyuubi.ha.client.class\ configuration. This change introduces a unified \DiscoveryClient\ interface and specific implementations (\ZookeeperDiscoveryClient\ and \EtcdDiscoveryClient\) to handle service registration, deregistration, and node discovery. It also adds dedicated configuration entries for etcd SSL settings and separates ZooKeeper authentication principals and keytabs for the server and engine components, enabling more granular control over Kerberos and digest authentication in HA deployments.

kyuubi-ha · high confidence

Introduce standalone JDBC engine with multi-dialect support

This change introduces a new standalone JDBC engine component that allows Kyuubi to connect directly to various SQL databases via JDBC. The implementation includes a new backend service, session manager, and Thrift frontend service to handle engine lifecycle and client communication. It features a pluggable dialect system with built-in support for ClickHouse, Doris, Impala, MySQL, Oracle, Phoenix, PostgreSQL, and StarRocks, each with specific connection providers, schema helpers, and row-set generators to handle database-specific quirks (such as Impala's metadata column naming or MySQL's catalog/schema handling). The engine also supports YARN deployment and configurable initialization SQL.

externals/kyuubi-jdbc-engine/src/main/scala · high confidence

Introduce unified authentication provider framework and internal engine security

The authentication subsystem in kyuubi-common has been refactored to support a pluggable provider model. Users can now configure authentication methods (NONE, LDAP, JDBC, CUSTOM) via the \AUTHENTICATION\_METHOD\ configuration, with the \AuthenticationProviderFactory\ instantiating the corresponding implementations (e.g., \LdapAuthenticationProviderImpl\, \JdbcAuthenticationProviderImpl\, \AnonymousAuthenticationProviderImpl\). The diff also introduces \InternalSecurityAccessor\ and \EngineSecureAuthenticationProviderImpl\ to handle secure internal authentication between the Kyuubi server and engines when \INTERNAL\_SECURITY\_ENABLED\ is true, using encrypted tokens. Additionally, \AuthTypes\ and \AuthMethods\ enumerations and \AuthUtils\ helper methods are added to manage SASL/PLAIN authentication types and proxy access verification.

kyuubi-common/src/main/scala/org/apache/kyuubi/service/authentication · high confidence

Introduces EngineType and ShareLevel enumerations for multi-engine and session sharing support

This change adds the \EngineType\ and \ShareLevel\ enumerations to the \kyuubi-common\ module, establishing the foundational types for Kyuubi's engine abstraction and session sharing mechanisms. \EngineType\ defines the supported engine variants, including SPARK\_SQL, FLINK\_SQL, TRINO, HIVE\_SQL, JDBC, and DATA\_AGENT. \ShareLevel\ defines the granularity for sharing engine applications across sessions, offering levels from CONNECTION (single session) to USER, GROUP, SERVER\_LOCAL, and SERVER (cluster-wide). These types enable the system to distinguish between different backend engines and manage how engine instances are shared among users and groups.

kyuubi-common/src/main/scala/org/apache/kyuubi/engine · high confidence

JDBC engine now supports ClickHouse, Doris, Impala, MySQL, Oracle, Phoenix, PostgreSQL, and StarRocks

The JDBC engine has been extended to support eight additional database dialects: ClickHouse, Doris, Impala, MySQL, Oracle, Phoenix, PostgreSQL, and StarRocks. This is achieved by registering the corresponding connection providers and dialect implementations in the Java SPI service files (JdbcConnectionProvider and JdbcDialect), enabling users to connect to these databases through the existing JDBC engine infrastructure.

externals/kyuubi-jdbc-engine/src/main/resources · high confidence

Kyuubi Spark 4.0 extension adds SQL optimization rules and configuration support

The Kyuubi extension for Spark 4.0 introduces several new SQL optimization capabilities and configuration options. It adds a SQL grammar for OPTIMIZE ZORDER statements and an AST builder to parse them. New query planning rules include DynamicShufflePartitions, which dynamically adjusts shuffle partition counts based on data size; FinalStageConfigIsolation, which allows different configurations for the final query stage; and InsertShuffleNodeBeforeJoin, which ensures shuffle nodes exist before joins to support skewed join optimization. The extension also adds InferRebalanceAndSortOrders to infer partitioning and sorting columns from join keys, and DropIgnoreNonexistent to suppress errors when dropping non-existent objects. These features are controlled by new configuration options such as spark.sql.optimizer.inferRebalanceAndSortOrders.enabled, spark.sql.optimizer.dynamicShufflePartitions.enabled, and spark.sql.optimizer.finalStageConfigIsolation.enabled.

extensions/spark/kyuubi-extension-spark-4-0 · high confidence

Kyuubi Spark 4.1 extension adds SQL optimization rules and configuration options

The Kyuubi extension for Spark 4.1 introduces several new SQL optimization capabilities and configuration options. Users can now enable dynamic shuffle partition adjustment based on data size via \spark.sql.optimizer.dynamicShufflePartitions.enabled\ and \spark.sql.optimizer.dynamicShufflePartitions.maxNum\. The extension also adds support for inferring rebalance and sort orders from join keys to improve compression, with options to skip non-cheap keys (\spark.sql.optimizer.inferRebalanceAndSortOrders.skipSort\) and limit inferred columns (\spark.sql.optimizer.inferRebalanceAndSortOrdersMaxColumns\). Additionally, a new \spark.sql.optimizer.forceShuffleBeforeJoin.enabled\ option ensures shuffle nodes exist before joins to support AQE's \OptimizeSkewedJoin\. The extension also introduces final stage config isolation (\spark.sql.optimizer.finalStageConfigIsolation.enabled\) to allow different configurations for the final query stage, and a \spark.sql.optimizer.dropIgnoreNonExistent\ option to prevent errors when dropping non-existent objects.

extensions/spark/kyuubi-extension-spark-4-1 · high confidence

Kyuubi Spark Hive Connector introduces external catalog pooling and delegation token support

The connector now manages external catalog instances through a configurable pooling strategy (ONE\_FOR\_ALL or ONE\_FOR\_ONE) to optimize resource usage across sessions, and includes a dedicated Hadoop delegation token provider to handle secure authentication with Kerberized Hive Metastores. These changes are supported by new configuration options for controlling catalog sharing and delegation token renewal, along with utility classes to ensure compatibility with various Spark versions.

extensions/spark/kyuubi-spark-connector-hive/src/main · high confidence

New Apache Atlas integration for SQL lineage tracking

The Kyuubi Spark lineage plugin now supports sending lineage data to Apache Atlas. This change introduces a new \ATLAS\ dispatcher option (configurable via \spark.kyuubi.plugin.lineage.dispatchers\) alongside the existing Spark and Kyuubi event dispatchers. When enabled, the plugin captures input and output table relationships, as well as column-level lineage, and pushes them to Atlas as \spark\_process\ and \spark\_column\_lineage\ entities. This allows users to visualize data flow and dependencies in Atlas for SQL queries executed through Kyuubi.

extensions/spark/kyuubi-spark-lineage/src/main · high confidence

New Docker Compose-based Kyuubi Playground

A new Docker Compose environment is now available in \docker/playground\ to quickly spin up a local Kyuubi development and testing environment. This setup includes Kyuubi (v1.11.0), Spark (v3.5.8), and a PostgreSQL metastore, along with RustFS for object storage and Prometheus/Grafana for monitoring. Users can start the stack with \docker compose up -d\ and immediately connect via Beeline or DBeaver, with built-in scripts to load TPC-DS and TPC-H sample datasets for immediate experimentation.

docker/playground · high confidence

New Kyuubi Query Engine tab in Spark UI with session details and graceful stop support

The Spark Web UI now includes a 'Kyuubi Query Engine' tab that displays engine version, compilation details, and background execution pool metrics. Users can view online and closed sessions, along with running, completed, and failed SQL statements. Clicking a session reveals its properties, user/IP info, creation/end times, and statement statistics. The tab also supports stopping the engine immediately or gracefully via new UI links. This UI integration works across Spark versions by adapting to both javax and jakarta servlet namespaces and is also available in the Spark History Server via a new plugin.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/ui · high confidence

New Kyuubi-specific SQL UDFs for engine and session metadata

The Spark SQL engine now registers a set of built-in user-defined functions that expose internal Kyuubi and Spark context details directly in SQL queries. Users can call \kyuubi\_version\ to get the server version, \engine\_name\ and \engine\_id\ to identify the Spark application, \system\_user\ for the OS user, \session\_user\ for the authenticated session user, and \engine\_url\ for the engine connection URL. These functions are automatically registered when the Spark session is initialized, with \session\_user\ intentionally skipped on Spark 4.0+ to avoid conflicts with Spark's native built-in function of the same name.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/udf · high confidence

New Serde layer for authorization command and plan specifications

The authorization plugin now uses a new serialization/deserialization (Serde) layer to define and extract privileges from Spark SQL commands and logical plans. This change introduces a set of Scala case classes (such as CommandSpec, Database, Table, Function, and Uri) and a collection of pluggable Extractors that map Spark plan nodes to authorization objects. The system loads command specifications from JSON files (e.g., database\_command\_spec.json, table\_command\_spec.json) and uses reflection to extract relevant details like table names, database names, and URIs, enabling more flexible and version-aware privilege checking.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/serde · high confidence

New Spark SQL engine listeners and progress monitoring components

The Spark SQL engine now includes a new set of listener and monitoring classes to improve operation tracking and user visibility. SQLOperationListener and SparkConsoleProgressBar provide statement-level logging and a real-time console progress bar showing stage and task completion. SparkProgressMonitor exposes detailed stage progress metrics (total, completed, running, pending, failed) for the UI. SparkSQLEngineEventListener manages the retention of session and statement events in the UI store, while SparkSQLEngineListener handles engine deregistration based on configurable exception classes, messages, and job failure thresholds. Supporting utilities in SparkContextHelper and SparkUtilsHelper enable delegation token updates, local property access, and sensitive information redaction.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/kyuubi · high confidence

New Spark SQL engine operation implementations for SQL, Python, Scala, and metadata

The Spark SQL engine now includes dedicated operation classes for executing SQL statements (ExecuteStatement), Python code (ExecutePython), and Scala scripts (ExecuteScala), alongside a suite of JDBC metadata operations (GetCatalogs, GetColumns, GetFunctions, GetSchemas, GetTables, GetTypeInfo, etc.). This change introduces support for asynchronous execution, incremental result collection, and saving large query results to files, while also providing the underlying infrastructure for catalog and schema introspection via the SparkCatalogUtils.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/operation · high confidence

New Spark connector for reading YARN aggregated logs

The \kyuubi-spark-connector-yarn\ extension now provides a Spark DataSource V2 implementation that allows users to query YARN aggregated logs as a table. This connector exposes a read-only \app\_logs\ table (via the \YarnCatalog\) containing columns for \app\_id\, \user\, \host\, \container\_id\, \log\_type\, \line\, \message\, and \mtime\. It supports filter pushdown for \app\_id\, \user\, and \host\ to optimize scans, reads TFile format logs from the NodeManager's remote log directory, and includes tests verifying the schema and filtering capabilities.

extensions/spark/kyuubi-spark-connector-yarn · high confidence

New authorization utility classes for session verification and Spark compatibility

The authorization module introduces several new utility components to enhance security and compatibility. AuthZUtils adds ECDSA-based session user verification to prevent impersonation, detects Spark 4.0+ to adapt to the new Derby JDBC driver package, and provides helpers for quoting identifiers. PathIdentifier enables access checks for path-based tables, while WithInternalChildren traits support internal logical plan manipulation. ReservedKeys defines the configuration keys for these new features.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/util · high confidence

New bin scripts for Kyuubi server, control clients, and Docker image management

The bin directory now includes a comprehensive set of shell scripts to manage Kyuubi operations. The main \bin/kyuubi\ script handles server lifecycle commands (start, stop, restart, run, kill, status) and displays a colored ASCII logo on startup. New control clients \bin/kyuubi-ctl\ and \bin/kyuubi-admin\ are provided for interacting with the Kyuubi Control Interface and performing administrative tasks like engine management. The legacy \bin/beeline\ is replaced by \bin/kyuubi-beeline\, which supports background execution and uses a dedicated classpath. Additionally, \bin/docker-image-tool.sh\ allows building and pushing Docker images, while \bin/kyuubi-zk-cli\ provides a ZooKeeper client using shaded dependencies. Environment setup is centralized in \bin/load-kyuubi-env.sh\, which configures paths for Spark, Flink, Trino, and Hive engines, and \bin/stop-application.sh\ assists in killing YARN applications.

bin · high confidence

New build infrastructure and testing Dockerfiles

This change introduces a new \build/\ directory containing the core scripts and configuration for building and distributing Kyuubi. It adds \build/dist\ to create binary distributions with support for optional Spark, Flink, and Hive provided profiles, and \build/mvn\ to manage the Maven version and execution environment. It also includes \build/dependency.sh\ to detect changes in the dependency list, \build/kyuubi-build-info\ to generate version information, and new Dockerfiles (\Dockerfile.CI\, \Dockerfile.HBase\, \Dockerfile.HMS\) to facilitate CI testing and local development of backend engines.

build · high confidence

New common utilities for configuration parsing and version handling in the Spark connector

The kyuubi-spark-connector-common module now provides shared utilities to simplify configuration and version management. JavaUtils adds static regex patterns for efficiently parsing human-readable time strings (e.g., '50s', '100ms') and byte strings (e.g., '100k', '250mb') into their respective units. SparkConfParser introduces a fluent API that allows connector code to read configuration values from runtime options, session configurations, or table properties with a defined priority order, supporting native parsing for bytes and time. Additionally, SparkUtils exposes the current Spark runtime version as a SemanticVersion object for easier version comparisons.

extensions/spark/kyuubi-spark-connector-common/src/main · high confidence

New developer utility scripts and dependency tracking

Added \dev/checkout\_pr.sh\ to simplify checking out pull requests locally, \dev/merge\_kyuubi\_pr.py\ to automate the PR merge workflow with GitHub API integration, and \dev/reformat\ to apply code formatting (including Python spotless checks via \black\). Also introduced \dev/dependencyList\ to explicitly track the versions of all runtime dependencies used in the build.

dev · high confidence

New release announcement and voting scripts added

The release process now includes two new executable scripts: \announce.sh\ and \dev\_kyuubi\_vote.sh\. The announcement script generates a standardized email template for release announcements, including download links and documentation URLs. The voting script generates a vote request template for release candidates, incorporating links to the staging repository, commit history, and KEYS file. These scripts streamline the creation of official Apache Kyuubi release communications.

build/release/script · high confidence

New release automation scripts and infrastructure

The build/release directory now includes a comprehensive set of new scripts to streamline the release process. This adds \release.sh\ for publishing and finalizing releases, \create-package.sh\ for generating source and binary tarballs, and \collect-licenses.sh\ for assembling license and NOTICE files. It also introduces Python utilities (\pre\_gen\_release\_notes.py\, \release\_utils.py\, \append\_notice.py\) to automate release note generation and contributor canonicalization using \known\_translations\, alongside an \asf-settings.xml\ for Maven deployment configuration.

build/release · high confidence

New utility classes for JDBC, Hadoop, threading, and signing

This change introduces a suite of new utility objects in the kyuubi-common module to centralize common functionality. JdbcUtils provides safe resource management and SQL execution helpers, including a redactPassword method for logging and isDuplicatedKeyDBErr for cross-database error detection. KyuubiHadoopUtils handles Hadoop Configuration creation, Hadoop Credentials serialization/deserialization, and delegation token management. ThreadUtils adds support for Java 21 virtual threads alongside standard thread pools, while NamedThreadFactory and KyuubiUncaughtExceptionHandler ensure all created threads are properly named and log uncaught exceptions. SignUtils implements ECDSA signing and verification using the secp521r1 curve, and TempFileCleanupUtils manages temporary file deletion via a shutdown hook. Additional utilities include ClassUtils for reflection-based instantiation, RowSetUtils for date/time formatting, SignalRegister for Unix signal handling, and ThriftUtils for verifying Thrift status codes.

kyuubi-common/src/main/scala/org/apache/kyuubi/util · high confidence

Pluggable session configuration and group resolution interfaces

The Kyuubi server plugin now exposes extensible interfaces for session management, allowing administrators to customize how session configurations are merged and how user groups are resolved. A new \SessionConfAdvisor\ interface enables plugins to provide configuration overlays that override default session settings based on the user and current configuration, while a \GroupProvider\ interface allows for custom logic to determine a user's primary group and associated groups. Default implementations (\DefaultSessionConfAdvisor\) are provided to maintain existing behavior, ensuring backward compatibility while enabling future extensibility.

extensions/server/kyuubi-server-plugin · high confidence

Project initialization and Apache Software Foundation (ASF) compliance setup

The repository has been initialized with the foundational structure required for an Apache project. This includes adding the standard Apache License 2.0, NOTICE files for source and binary distributions, and an .asf.yaml configuration that enables GitHub features (issues, discussions, projects) and sets up notification routing to Apache mailing lists. Additionally, developer tooling and compliance configurations have been added, such as .gitignore, .rat-excludes for release audits, .scalafmt.conf for code formatting, scalastyle-config.xml for Scala style checks, and .readthedocs.yaml for documentation hosting. The README has been updated with the official Apache Kyuubi branding, badges, and links to the new documentation and project sites.

(repo-wide) · high confidence

Register Kyuubi history server plugin for Spark UI integration

The Kyuubi history server plugin is now automatically discovered and loaded by the Spark UI. This is achieved by adding a new service provider configuration file that registers the \org.apache.spark.ui.KyuubiHistoryServerPlugin\ class, enabling seamless integration of Kyuubi-specific history server features into the standard Spark UI.

externals/kyuubi-spark-sql-engine/src/main/resources/META-INF/services · high confidence

Support for running Kyuubi engines on YARN

Kyuubi now supports deploying engines on YARN, allowing users to leverage YARN's resource management for engine execution. This change introduces a new \DeployMode.YARN\ option and implements the necessary infrastructure, including an Application Master to manage the engine lifecycle, a submitter to handle application submission and staging, and argument parsing for engine configuration. The implementation also includes support for Kerberos authentication, ensuring secure engine deployment in secured Hadoop clusters.

kyuubi-common/src/main/scala/org/apache/kyuubi/engine/deploy · high confidence

TPC-DS benchmark queries added for Spark 3.2

The Kyuubi Spark connector now includes a new set of TPC-DS benchmark queries (q1 through q18) and their corresponding expected output schemas and hashes, specifically organized under the \tpcds\_3.2\ resource folder. This addition provides users with standardized SQL workloads and validation artifacts tailored for Spark 3.2, enabling consistent performance testing and result verification for this version.

extensions/spark/kyuubi-spark-connector-tpcds/src/main/resources · high confidence

TPC-H benchmark queries added to Kyuubi Spark connector

The Kyuubi Spark connector now includes the full set of TPC-H benchmark queries (q1–q22) as bundled resources. Each query is provided with its SQL definition, expected output schema, and a content hash for validation, enabling users to run standard TPC-H workloads directly against the connector.

extensions/spark/kyuubi-spark-connector-tpch/src/main/resources · high confidence

Architecture

Refactored service lifecycle and backend abstraction in Kyuubi common

The service layer in kyuubi-common has been restructured to improve modularity and lifecycle management. A new base \AbstractService\ introduces strict state transitions (LATENT, INITIALIZED, STARTED, STOPPED) and configuration handling, which is extended by \CompositeService\ to manage the start/stop order of child services. The \Serverable\ abstraction now explicitly owns a \BackendService\ (handling session and operation logic) and a sequence of \FrontendService\s (handling client protocols), replacing the previous hierarchy. Additionally, a new \TempFileService\ has been added to manage server-side temporary files with size-based eviction and scheduled cleanup, addressing potential memory leaks from previous \deleteOnExit\ usage.

kyuubi-common/src/main/scala/org/apache/kyuubi/service · high confidence

Behavioural changes

Add internal configuration flag for test environment detection

A new internal configuration option, \kyuubi.testing\, has been added to the \kyuubi-common\ module. This boolean flag allows the system to detect when it is running within a test environment, supporting the infrastructure changes required for unit tests and MiniYARNCluster integration.

kyuubi-common/src/main/scala/org/apache/kyuubi/config/internal · high confidence

Arrow serialization now supports Zstd compression and optimized local execution

The Spark SQL engine's Arrow serialization path has been enhanced to support Zstd compression for query results, configurable via Spark's standard Arrow compression settings. Additionally, the new \SparkDatasetHelper\ introduces optimized handling for \LocalTableScanExec\ and \CommandResultExec\ plans to avoid unnecessary job triggers, and ensures that Arrow batch conversion uses Kyuubi-specific converters to properly apply compression codecs.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/sql/kyuubi · high confidence

AuthZ plugin removes JNA dependency for hostname resolution

The Kyuubi Spark Authorization plugin no longer relies on JNA libraries to determine the local machine's hostname. A new internal utility class, com.kstruct.gethostname4j.Hostname, has been added to handle hostname retrieval using standard Java APIs (InetAddress) and environment variables, effectively replacing the previous JNA-based implementation.

extensions/spark/kyuubi-spark-authz/src/main/java · high confidence

Centralized metadata schema constants for GetTables operations

A new \ResultSetSchemaConstant\ object has been introduced in the \kyuubi-common\ module to define standard column names for JDBC metadata result sets (such as \TABLE\_CAT\, \TABLE\_NAME\, \DATA\_TYPE\, etc.). This change consolidates common variable definitions previously duplicated across Flink and Spark engine implementations, ensuring consistent metadata schema handling for \GetTables\ operations across different execution engines.

kyuubi-common/src/main/scala/org/apache/kyuubi/operation/meta · high confidence

The Flink SQL engine now manages sessions through a dedicated session manager that wraps the Flink SQL Gateway, handling session lifecycle and configuration normalization. When a session is opened, the engine executes any configured initialization SQL statements before processing USE CATALOG and USE DATABASE commands, ensuring the correct context is established. Additionally, session closure behavior has changed: if the engine share level is set to CONNECTION, the Flink engine stops immediately after the last session closes, rather than remaining idle. The engine also now supports retrieving server and database metadata via GetInfo calls, returning 'Apache Flink' as the server name and the current Flink version.

externals/kyuubi-flink-sql-engine/src/main/scala/org/apache/kyuubi/engine/flink/session · high confidence

Introduce KyuubiHiveDriver as the primary JDBC driver

The kyuubi-hive-jdbc module now registers KyuubiHiveDriver as the main driver implementation, replacing the legacy KyuubiDriver. KyuubiDriver is retained only as a deprecated alias to ensure backward compatibility for existing applications. This change establishes the new driver class as the standard entry point for connecting to Kyuubi via the HiveServer2 thrift protocol.

kyuubi-hive-jdbc · high confidence

Introduce KyuubiSQLException with SQL state and vendor code support

The \kyuubi-common\ module now includes a new \KyuubiSQLException\ class that extends \java.sql.SQLException\, allowing Kyuubi to expose structured error information including SQL state codes (such as \0A000\ for feature not supported) and vendor-specific error codes to JDBC clients. This replaces the previous generic exception handling with a standard SQL exception format, improving client-side debugging and error handling compatibility. The change also introduces a new \KyuubiException\ base class and updates the \Logging\ and \Utils\ utilities to support the new exception hierarchy and build information retrieval.

kyuubi-common/src/main/scala/org/apache/kyuubi · high confidence

Introduce standalone embedded ZooKeeper server with configurable binding and paths

The kyuubi-zookeeper module now provides a dedicated EmbeddedZookeeper service that can be started and stopped independently, allowing users to run a local ZooKeeper instance for testing or lightweight deployments. This change introduces new configuration options (kyuubi.zookeeper.embedded.\*) to control the client port, data directories, and session timeouts, with relative paths automatically resolved against KYUUBI\_HOME. Users can now explicitly bind the embedded server to a specific IP address or hostname via kyuubi.zookeeper.embedded.client.port.address, overriding the previous default behavior of using the local canonical hostname. The module also includes a new log4j2 XML configuration for tests and unit tests verifying the connection, binding, and path resolution behaviors.

kyuubi-zookeeper · high confidence

Introduces audience-based configuration scoping to replace server-only flags

The configuration system in kyuubi-common has been refactored to support audience-based scoping, replacing the previous server-only restriction model. A new ConfigAudience enumeration defines specific targets (SERVER, SPARK, FLINK, HIVE, TRINO, JDBC, DATA\_AGENT, ALL\_ENGINES, ANY), and the ConfigBuilder now allows marking entries with these audiences. This change enables the system to distinguish between server-side, frontend, and engine-specific configurations, allowing for more granular control over which configuration keys are exposed or applied to specific components.

kyuubi-common/src/main/scala/org/apache/kyuubi/config · high confidence

Kyuubi Helm chart templates restructured and modernized

The Kyuubi Helm chart templates have been completely rewritten to support a more flexible and robust deployment model. The Kyuubi server is now deployed as a StatefulSet instead of a Deployment, ensuring stable network identities and supporting headless services for direct pod access. The chart now supports multiple enabled frontend protocols (e.g., REST, Thrift) with dedicated services and ports for each, rather than a single monolithic service. Configuration is now split into separate ConfigMaps for Kyuubi, Hadoop, and Spark settings, allowing users to mount distinct configuration files for each component. Additionally, the chart introduces improved metrics integration with support for PrometheusRule, ServiceMonitor, and PodMonitor resources, along with configurable labels, annotations, and priority classes for better observability and scheduling control.

charts/kyuubi/templates · high confidence

New SPI service providers for the authorization plan serialization layer

The authorization extension now registers a comprehensive set of Java SPI service providers under \META-INF/services/org.apache.kyuubi.plugin.spark.authz.serde\ to support the new plan serialization architecture. This change introduces extractors for actions, catalogs, columns, databases, functions, queries, tables, table types, and URIs, enabling the system to correctly serialize and deserialize authorization context for various Spark logical plan nodes and metadata structures.

extensions/spark/kyuubi-spark-authz/src/main/resources/META-INF · high confidence

New Spark catalog and JSON utility helpers for improved compatibility and performance

Added JsonUtils for standardized JSON serialization and SparkCatalogUtils to centralize interactions with Spark's catalog system. SparkCatalogUtils introduces reflective access to the CatalogManager to maintain binary compatibility across Spark versions (including Spark 4.2) and implements optimized logic for retrieving tables and views, addressing performance issues in GetTables operations and ensuring correct handling of catalog namespaces and identifiers.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/util · high confidence

New configuration and node-level security checks in Kyuubi authorization

The Kyuubi authorization extension now enforces stricter security controls by introducing two new rule checkers. AuthzConfigurationChecker prevents users from modifying or resetting restricted Spark configurations (such as spark.sql.extensions) and blocks attempts to exclude authorization rules via spark.sql.optimizer.excludedRules, ensuring security rules cannot be bypassed. NodeDenyListChecker rejects queries containing specific logical plan nodes defined in a denylist, adding a layer of protection against potentially unsafe query structures.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/rule/config · high confidence

New structured logging and configuration templates for Kyuubi

The conf directory now includes new template files (kyuubi-defaults.conf.template, kyuubi-env.sh.template, log4j2.xml.template, and log4j2-repl.xml.template) that provide a structured logging setup. This introduces dedicated audit log files for REST API requests, Kubernetes application state changes, and general operations, alongside environment variables for configuring JVM options and engine-specific paths (Spark, Flink, Hive).

conf · high confidence

Permanent view marker now inherits determinism and statistics from child plans

The PermanentViewMarker node now extends LeafNode and MultiInstanceRelation, ensuring that its determinism status and query statistics are correctly inherited from the underlying view definition. This change allows the optimizer to properly handle nondeterministic views (such as those using rand() or uuid()) and ensures that table caching works as expected, while also simplifying privilege object resolution by treating the marker as a leaf node.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/rule/permanentview · high confidence

Playground configuration updated to use RustFS and Java 17

The Kyuubi Playground's configuration files have been updated to switch the underlying object storage from MinIO to RustFS (configured via core-site.xml and hive-site.xml pointing to s3a:// endpoints) and to upgrade the runtime environment to Java 17 (updated in kyuubi-env.sh and spark-env.sh). Additionally, the setup now includes specific logging configurations for Kyuubi (kyuubi-log4j2.xml) and defines default Spark and Hive catalogs for TPC-DS, TPC-H, and PostgreSQL data sources.

docker/playground/conf · high confidence

Redesigned LDAP authentication with pluggable filters and custom queries

The LDAP authentication module in kyuubi-common has been refactored to support a chain of pluggable filters, allowing administrators to combine multiple validation steps (such as user existence checks, user filters, and group membership checks) into a single authentication flow. A new CustomQueryFilter enables users to define arbitrary LDAP search queries for additional validation, while the GroupFilterFactory now supports both group-based and user-membership-key-based group checks. The underlying search logic has been restructured around a new Query builder and SearchResultHandler to improve performance and prevent memory leaks from StringTemplate caches.

kyuubi-common/src/main/scala/org/apache/kyuubi/service/authentication/ldap · high confidence

Refactored TRowSet generation with new column and value generator traits

The result assembly logic in kyuubi-common has been restructured to improve performance and clarity. New traits TRowSetColumnGetter, TColumnGenerator, and TColumnValueGenerator have been introduced to separate concerns between accessing row data and converting it into Thrift types. TRowSetGenerator now leverages these traits to handle both row-based (pre-V6 protocol) and column-based (V6+ protocol) result set generation, allowing for more efficient column-based assembly and direct type mapping.

kyuubi-common/src/main/scala/org/apache/kyuubi/engine/result · high confidence

Restructured authorization rule implementation with improved tag management and TypeOf handling

The authorization rule logic in the \rule\ package has been reorganized into a new modular structure. A new \Authorization\ base class and companion object now manage privilege checking and introduce a \KYUUUBI\_AUTHZ\_TAG\ to mark nodes as checked, preventing redundant checks and fixing issues where cached catalog nodes bypassed authorization. Specific rules have been extracted or refined: \RuleEliminateMarker\ now handles the removal of data masking and row filter markers; \RuleEliminatePermanentViewMarker\ ensures subqueries within permanent views are correctly marked as checked to avoid redundant privilege checks; and \RuleEliminateTypeOf\ / \RuleApplyTypeOfMarker\ with a new \TypeOfPlaceHolder\ expression ensure \TypeOf\ expressions are correctly processed without leaking placeholders to execution. Helper utilities like \RuleHelper\ provide common functionality for Spark session and user group access.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/rule · high confidence

Row-level authorization for SHOW commands and table scans

The row filter authorization plugin now enforces access control on metadata discovery operations (SHOW TABLES, SHOW NAMESPACES, SHOW FUNCTIONS, SHOW COLUMNS) and table scans. For SHOW commands, results are filtered at execution time by checking the user's USE or SELECT permissions via the Ranger admin plugin, ensuring users only see objects they are authorized to access. For table scans, the logical plan is modified to inject a Filter node based on a filter expression retrieved from Ranger, restricting the rows returned by SELECT queries. This change moves authorization checks from the driver side to the executor side for SHOW operations to avoid unnecessary executor tasks, while maintaining row-level filtering for data access.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/rule/rowfilter · high confidence

Scala 2.13 interpreter initialization updated to use explicit Settings

The Kyuubi Spark SQL engine's Scala 2.13 REPL implementation now passes a new \Settings\ instance explicitly to \createInterpreter\ and initializes the compiler via \iMain.initializeCompiler()\, replacing the previous implicit initialization path. This change ensures the Scala interpreter correctly loads the Spark classpath and binds \spark\ and \sc\ objects for users executing code in the Kyuubi SQL engine.

externals/kyuubi-spark-sql-engine/src/main/scala-2.13 · high confidence

Spark 3.5 SQL extension introduces dynamic shuffle partitioning and optimized query planning

The Spark 3.5 extension adds a new SQL grammar for OPTIMIZE ... ZORDER BY statements and introduces several optimizer rules to improve query performance and stability. A new DynamicShufflePartitions rule automatically adjusts shuffle partition counts based on input data size when Adaptive Query Execution is enabled. The extension also includes rules to insert shuffle nodes before joins to support skewed join optimization, infer rebalance and sort orders from join keys and aggregates to improve compression, and isolate final-stage configurations for better control over write operations. Additionally, a DropIgnoreNonexistent rule allows DROP operations to proceed silently if the target does not exist, configurable via spark.sql.optimizer.dropIgnoreNonExistent.

extensions/spark/kyuubi-extension-spark-3-5/src/main · high confidence

Spark SQL engine adds compatibility support for Spark 4.0

The Spark SQL engine now includes helper utilities to support execution on Spark 4.0. New files in the execution package introduce dynamic method invocation to access the Spark session from a SparkPlan and to propagate SQL configuration via the new \withSQLConfPropagated\ API, ensuring compatibility with both the classic SparkSession (introduced in Spark 4.0) and the legacy version.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/spark/sql/execution · high confidence

Spark SQL engine session management and idle timeout support

The Spark SQL engine now introduces a dedicated session manager and implementation that support user-isolated Spark sessions with configurable idle timeouts, allowing unused sessions to be automatically cleaned up. It also enforces that catalog and database context changes are applied before other session configurations, and ensures that cached temporary views are properly uncached when a session closes to free resources.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/session · high confidence

Spark engine schema and row-set generation refactored for Spark 4.0 compatibility

The Spark SQL engine's schema handling and row-set generation have been restructured to support Spark 4.0 and newer data types. The new \RowSet\ object uses reflection to dynamically adapt to Spark 4.0's \HiveResult\ API changes, specifically handling the new \BinaryFormatter\ parameter required for \toHiveString\. \SchemaHelper\ now maps Spark's \VariantType\, \TimestampNTZType\, and interval types to their corresponding Thrift type IDs, ensuring accurate metadata reporting. Additionally, \SparkTRowSetGenerator\ and \SparkArrowTRowSetGenerator\ have been introduced to handle standard and Arrow-based result sets respectively, improving type safety and performance in data serialization.

externals/kyuubi-spark-sql-engine/src/main/scala/org/apache/kyuubi/engine/spark/schema · high confidence

The EmbeddedExecutorFactory in the Kyuubi Flink SQL engine now uses synchronized access and volatile flags to manage bootstrap job IDs, preventing race conditions during concurrent statement execution. This change ensures that multiple threads submitting jobs through the SQL gateway do not corrupt shared state when initializing the embedded Flink environment, addressing stability issues in multi-tenant scenarios.

externals/kyuubi-flink-sql-engine/src/main/java · high confidence

Unified operation logging with Log4j 2 support and seekable readers

The operation log system now supports both Log4j 1 and Log4j 2 through dedicated divert appenders that capture engine logs and write them to the operation log file. This change introduces a seekable buffered reader for random access to operation logs, allowing users to efficiently retrieve specific portions of large log files without reading from the beginning. The implementation also ensures proper resource management by closing existing seekable readers when adding extra logs and handling file-not-found scenarios gracefully.

kyuubi-common/src/main/scala/org/apache/kyuubi/operation/log · high confidence

Fixes

Fix data masking failures on UNION-ALL views over masked views

The data masking authorization rules have been restructured into a two-stage process (Stage 0 and Stage 1) to correctly handle complex query plans. This change specifically fixes an issue where data masking would fail or produce incorrect results when applied to a UNION-ALL view that references a masked view. The new implementation ensures that masker expressions are correctly resolved and applied even when expression IDs collide across different branches of a union, preventing out-of-scope maskers from shadowing the correct ones.

extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/rule/datamasking · high confidence

Test coverage

Added DDL command tests for the Kyuubi Spark Hive Connector; Added Flink SQL engine test infrastructure; Added LDAP test fixtures for Microsoft Active Directory and standard schemas; Added Oracle JDBC engine test suite; Added StarRocks integration tests for the JDBC engine; Added benchmark tests for TPC-DS table generation performance; Added integration tests for Doris and PostgreSQL JDBC engines; Added integration tests for Flink SQL engine operations; Added integration tests for Hive engine operation modes; Added integration tests for Kyuubi Server with Trino engine; Added integration tests for Kyuubi on Kubernetes with Spark engines; Added integration tests for Phoenix JDBC engine; Added integration tests for Spark on Kubernetes client and cluster modes; Added integration tests for Trino JDBC frontend; Added integration tests for Zookeeper service discovery; Added integration tests for the Apache Impala JDBC engine dialect; Added test base trait for Spark listener extension tests; Added test configuration and logging resources; Added test configuration resources for lineage plugin; Added test coverage for Kyuubi authentication providers and security components; Added test coverage for Spark SQL engine deregistration, UDTs, and operation listeners; Added test coverage for the Spark Hive Connector; Added test for Hive UDF documentation generation; Added test helper for SparkContext internals; Added test infrastructure and coverage for JDBC engine; Added test infrastructure and suites for Kyuubi service components; Added test infrastructure for generating and validating Ranger authorization policy files; Added test suite for Apache Phoenix JDBC engine integration; Added test suite for Kyuubi Spark SQL UDF documentation; Added test suites for Flink SQL engine operations and initialization; Added test suites for Kyuubi Spark 3.5 SQL extension features; Added test suites for Kyuubi configuration system; Added test suites for Spark SQL engine discovery, scheduling, and timeout handling; Added test suites for Spark SQL engine operations; Added test suites for Spark SQL engine session management; Added test suites for the Kyuubi Spark TPC-DS connector; Added test suites for the Kyuubi Spark TPC-H Connector; Added test utilities and suites for the Kyuubi Spark Connector; Added tests for Arrow compression and Spark plan limit handling; Added tests for EmbeddedExecutorFactory job ID reservation; Added tests for EngineEventsStore session and statement tracking; Added tests for FileWriterFactory task attempt ID handling; Added tests for Flink StringData type conversion in result sets; Added tests for Hive backend engine event logging; Added tests for Kyuubi JDBC driver functionality; Added tests for SQL lineage parsing helpers; Added tests for SeekableBufferedReader; Added tests for Spark SQL engine event logging; Added tests for Spark and Kyuubi lineage event capture; Added tests for Spark schema conversion and RowSet generation; Added tests for Trino engine, session, and operation event logging; Added tests for YARN engine submission logic; Added tests for operation metadata and fetch handling; Added tests for operation timeout scheduler lifecycle; Added tests for server operation log functionality; Added tests for session management and configuration validation; Added tests for the Apache Atlas lineage dispatcher; Added tests for the Kyuubi Engine Tab UI; Added unit tests for Kyuubi utility classes; Added unit tests for LDAP authentication components; Added unit tests for Trino engine schema handling and row set generation; Expanded test infrastructure for data lake formats, Kerberos security, and utility validation; Expanded unit test coverage for Kyuubi Spark authorization privilege checks.

Dependencies

New development tooling modules and documentation dependencies

This change introduces new Maven modules and configuration files to support development workflows. It adds the \kyuubi-codecov\ module, which aggregates project dependencies to facilitate code coverage analysis, and the \kyuubi-tpcds\ module, which packages a TPC-DS benchmark generator with shaded dependencies. Additionally, it adds a \docs/requirements.txt\ file to pin documentation build dependencies (such as Sphinx and MyST parser) and introduces the \kyuubi-flink-token-provider\ and \kyuubi-server-plugin\ modules to support Flink integration and server extension points.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 43 → 47 (+3.4)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 83 → 89 (+6.6)
  • Architecture 89 → 89 (-0.6)
  • Maturity 53 → 51 (-2.0)
  • Readiness 32 → 44 (+12.1)
  • Security 40 → 60 (+19.1)
  • Accessibility 39 (new)

Resolved (198)

  • BeeLine.connectUsingArgs (cognitive 25) (BeeLine)
  • BeeLine.connectUsingArgs (cyclomatic 18) (BeeLine)
  • BeeLine.getDefaultConnectionUrl (cognitive 25) (BeeLine)
  • BeeLine.getDefaultConnectionUrl (cyclomatic 16) (BeeLine)
  • BeeLine.handleSQLException (cyclomatic 16) (BeeLine)
  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (10 lines × 2) (externals/kyuubi-data-agent-engine/src/main/java/org/apache/kyuubi/engine/dataagent/runtime/MiddlewareDispatcher.java)
  • Duplicated block (10 lines × 2) (externals/kyuubi-data-agent-engine/src/main/java/org/apache/kyuubi/engine/dataagent/runtime/MiddlewareDispatcher.java)
  • Duplicated block (10 lines × 2) (externals/kyuubi-data-agent-engine/src/main/java/org/apache/kyuubi/engine/dataagent/runtime/MiddlewareDispatcher.java)
  • Duplicated block (10 lines × 2) (externals/kyuubi-data-agent-engine/src/main/java/org/apache/kyuubi/engine/dataagent/runtime/MiddlewareDispatcher.java)
  • Duplicated block (10 lines × 2) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/common/FastHiveDecimalImpl.java)
  • Duplicated block (102 lines × 3) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/adapter/SQLConnection.java)
  • Duplicated block (11 lines × 2) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/common/FastHiveDecimalImpl.java)
  • Duplicated block (11 lines × 4) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/KyuubiBaseResultSet.java)
  • Duplicated block (11 lines × 5) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/KyuubiArrowBasedResultSet.java)
  • Duplicated block (112 lines × 2) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/common/FastHiveDecimalImpl.java)
  • Duplicated block (113 lines × 2) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/JdbcColumn.java)
  • Duplicated block (12 lines × 2) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/common/FastHiveDecimalImpl.java)
  • Duplicated block (12 lines × 2) (kyuubi-rest-client/src/main/java/org/apache/kyuubi/client/AdminRestApi.java)
  • …and 178 more

New (700)

  • AccessType.apply (cognitive 22) (extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/ranger/AccessType.scala)
  • AccessType.apply (cyclomatic 22) (extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/ranger/AccessType.scala)
  • ArrowColumnVector.initAccessor (cognitive 17) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/arrow/ArrowColumnVector.java)
  • ArrowColumnVector.initAccessor (cyclomatic 17) (kyuubi-hive-jdbc/src/main/java/org/apache/kyuubi/jdbc/hive/arrow/ArrowColumnVector.java)
  • BatchJobSubmission.close (cognitive 17) (kyuubi-server/src/main/scala/org/apache/kyuubi/operation/BatchJobSubmission.scala)
  • BatchJobSubmission.submitAndMonitorBatchJob (cognitive 18) (kyuubi-server/src/main/scala/org/apache/kyuubi/operation/BatchJobSubmission.scala)
  • BatchJobSubmission.submitAndMonitorBatchJob (cyclomatic 16) (kyuubi-server/src/main/scala/org/apache/kyuubi/operation/BatchJobSubmission.scala)
  • BatchesResource.openBatchSessionInternal (cognitive 26) (kyuubi-server/src/main/scala/org/apache/kyuubi/server/api/v1/BatchesResource.scala)
  • BatchesResource.openBatchSessionInternal (cyclomatic 21) (kyuubi-server/src/main/scala/org/apache/kyuubi/server/api/v1/BatchesResource.scala)
  • BeeLine.connectUsingArgs (cognitive 25) (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • BeeLine.connectUsingArgs (cyclomatic 18) (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • BeeLine.getDefaultConnectionUrl (cognitive 25) (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • BeeLine.getDefaultConnectionUrl (cyclomatic 16) (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • BeeLine.handleSQLException (cyclomatic 16) (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • Boundary-crossing change coupling: AdminRestApi.java ↔ AdminResource.scala (kyuubi-rest-client/src/main/java/org/apache/kyuubi/client/AdminRestApi.java)
  • Boundary-crossing change coupling: KyuubiConf.scala ↔ MetadataStore.scala (kyuubi-common/src/main/scala/org/apache/kyuubi/config/KyuubiConf.scala)
  • Change coupling: BatchesResource.scala ↔ MetadataStore.scala (kyuubi-server/src/main/scala/org/apache/kyuubi/server/api/v1/BatchesResource.scala)
  • Change coupling: MetadataStore.scala ↔ KyuubiSessionManager.scala (kyuubi-server/src/main/scala/org/apache/kyuubi/server/metadata/MetadataStore.scala)
  • Change coupling: PrivilegesBuilder.scala ↔ PermanentViewMarker.scala (extensions/spark/kyuubi-spark-authz/src/main/scala/org/apache/kyuubi/plugin/spark/authz/PrivilegesBuilder.scala)
  • ClassTooLong: BeeLine (kyuubi-hive-beeline/src/main/java/org/apache/hive/beeline/BeeLine.java)
  • …and 680 more

Changes since last survey

  • 68 commits — 55 feature/other, 13 fixes

By area

  • kyuubi-util/src — 12 commits
  • .github/workflows — 7 commits
  • extensions/spark — 7 commits
  • kyuubi-server/src — 6 commits
  • externals/kyuubi-spark-sql-engine — 5 commits
  • (root) — 3 commits
  • externals/kyuubi-flink-sql-engine — 3 commits
  • kyuubi-server/web-ui — 3 commits
  • kyuubi-util-scala/src — 3 commits
  • build/dist — 2 commits
  • docs/connector — 2 commits
  • docs/deployment — 2 commits
  • docs/security — 2 commits
  • kyuubi-common/src — 2 commits
  • .github/PULL_REQUEST_TEMPLATE — 1 commit
  • dev/kyuubi-tpcds — 1 commit
  • dev/merge_kyuubi_pr.py — 1 commit
  • docs/client — 1 commit
  • docs/contributing — 1 commit
  • docs/quick_start — 1 commit

Notable commits

  • fix: [KYUUBI #6403][KSHC] Fix write into partitioned table with non-last partition column
  • fix: [KYUUBI #7627] [KSHC] Fix connector compatibility with Spark 4.2
  • fix: [KYUUBI #7640] Fix exception double-unwrap in DynMethods.UnboundMethod.invoke
  • fix: [KYUUBI #7643] [TESTS] Fix closed operation metric assertion
  • fix: [KYUUBI #7645][SERVER] Fix operation state metrics transition
  • fix: [KYUUBI #7649] Fix compaction summarizer tokens not accumulated
  • fix: [KYUUBI #7651] Fix NPE binding the AlwaysNull sentinel field in DynFields
  • fix: [KYUUBI #7674] Fix flaky PyHive SQLAlchemy insert tests
  • fix: [KYUUBI #7678] [INFRA] Fix greetings workflow for PRs from forks
  • fix: [KYUUBI #7690] Fix DynFields javadoc copied from the method variants
  • fix: [KYUUBI #7692][TESTS] Fix testFindLocalInetAddress on loopback-only hosts
  • fix: [KYUUBI #7715] Fix Scala 2.13 compilation of the TPC-DS generator
  • fix: [KYUUBI #7743] [DOC] Fix cross-references
  • change: [KYUUBI #6943][2/2] OrcScan and ParquetScan support DPP
  • change: [KYUUBI #6995] Support Flink 2.0, 2.1, 2.2 and 2.3
  • change: [KYUUBI #7293][KUBERNETES] Initialize in-cluster client automatically
  • change: [KYUUBI #7434][DOC] Reformat client docs from RST to Markdown
  • change: [KYUUBI #7434][DOC] Reformat connector docs from RST to Markdown
  • change: [KYUUBI #7434][DOC] Reformat quick start docs from RST to Markdown
  • change: [KYUUBI #7434][DOC] Reformat security docs from RST to Markdown
  • …and 48 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/kyuubi was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 0218ff4414b100c38b1040f5ffa561e47b5a59b5 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.