ClickHouse/spark-clickhouse-connector
70.2
Strong · 20 September 2026
19.6k
lines of production code
Scala
with Java
1
measurement over time
What this system is
This system is a Spark connector that enables Apache Spark to read from and write to ClickHouse databases. It supports Spark versions 3.3 through 4.0 and provides both catalog-based and format-based access patterns for interacting with single-node and distributed ClickHouse clusters. The project includes comprehensive test infrastructure and development tooling to ensure compatibility and stability across these environments.
How it got here
2020–2022 — Connector initialization and testing
11 changes.
This period established the foundational structure and build infrastructure for the Spark ClickHouse Connector, including support for Spark 3.3 through 4.0 and Java 21. It introduced core connector features such as catalog-based access and comprehensive unit and integration tests to validate data mapping and cluster operations. Additionally, a modular Docker-based development environment was created to streamline testing and deployment workflows.
2023–2024 — Spark 3.4/3.5 connector development
15 changes.
This period focused on building the foundational infrastructure for the ClickHouse Spark connector, including core client libraries, SQL parsing, and cluster management utilities. It introduced initial support for Spark 3.4 and 3.5 with catalog and table provider APIs, accompanied by extensive unit and integration tests to validate functionality across single-node, cluster, and cloud environments.
2025–2026 — Spark 4.0 support and test expansion
7 changes.
This period focused on enabling Spark 4.0 compatibility by implementing new catalog and table provider APIs, alongside adding corresponding integration tests and Scala examples. It also significantly expanded test coverage across Spark 3.4, 3.5, and core modules to validate write metrics, sharding logic, and scan stability.
Features
Add default configuration files for Kyuubi, Spark, Hive, and CloudBeaver
This change introduces a set of default configuration files for the playground environment. It configures Kyuubi to bind on 0.0.0.0:10009 with authentication disabled and pre-initializes namespaces for TPC-DS, PostgreSQL, and multiple ClickHouse clusters. Spark is configured to use Iceberg with S3 storage, define catalogs for TPC-DS, TPCH, PostgreSQL, and four ClickHouse instances (enabling async inserts), and set specific write parameters for ClickHouse. Hive is configured to use a PostgreSQL metastore and store its warehouse in S3. Finally, CloudBeaver is configured with a pre-defined connection to the Kyuubi service, specific security policies, and disabled drivers.
docker/conf · high confidence
Added Spark 4.0 Scala examples for batch and streaming debugging
New Scala example files have been added to the Spark 4.0 examples directory to assist with debugging and testing the ClickHouse connector. The SimpleBatchExample demonstrates writing sample employee data to ClickHouse using the catalog-aware API and performing aggregations, while the StreamingRateExample illustrates a micro-batch streaming pipeline that enriches rate-source data and writes it to ClickHouse using foreachBatch. Both examples configure the Spark session to connect to a local ClickHouse instance via environment variables and include logging to help trace connector behavior.
examples/scala/spark-4.0 · high confidence
Initial Spark 3.4 connector support with catalog and format APIs
The connector now supports Spark 3.4, introducing a new \ClickHouseCatalog\ for catalog-based table access and a \ClickHouseTableProvider\ that enables the standard \format("clickhouse")\ read/write pattern. This location provides the core wiring for these capabilities, including the \DataSourceRegister\ service file, helper traits for SQL and connection logic, a function registry for ClickHouse-specific hash functions, and custom metrics for monitoring read/write performance.
spark-3.4/clickhouse-spark/src/main · high confidence
Initial Spark 3.5 connector support with catalog and table provider APIs
The connector now supports Spark 3.5, introducing a new \ClickHouseCatalog\ for catalog-based access and a \ClickHouseTableProvider\ that enables the \format("clickhouse")\ read/write pattern. The catalog implementation discovers ClickHouse clusters and macros, registers built-in functions (such as \clickhouse\_shard\_num\ and hash functions), and manages table loading with proper timezone handling. The table provider allows ad-hoc queries and table creation via options like \host\, \database\, \table\, and \order\_by\, while also supporting external metadata inference. This location provides the core wiring and service registration (\DataSourceRegister\) that enables these new access modes.
spark-3.5/clickhouse-spark/src/main · high confidence
Initial core library and client infrastructure for the Spark connector
This change introduces the foundational components for the ClickHouse Spark connector, including a new ANTLR grammar for parsing ClickHouse SQL, a comprehensive Java enum mapping ClickHouse server error codes, and a bundled CityHash implementation for hashing. It also establishes the core Scala utilities, constants, and logging traits, alongside the initial client layer that manages connections to ClickHouse nodes and clusters, handling query execution, insertion, and runtime environment detection.
clickhouse-core/src/main · high confidence
Introduce Spark 3.3 connector with catalog-based and format-based access
The Spark 3.3 connector now supports both catalog-based and format-based access patterns. A new \ClickHouseCatalog\ implementation enables integration with Spark's catalog API, while a \ClickHouseTableProvider\ allows ad-hoc queries via the standard format API (e.g., \spark.read.format("clickhouse")\). The connector registers itself through \DataSourceRegister\ and includes helper utilities for SQL compilation, function registry management, and custom metrics tracking for read/write operations.
spark-3.3/clickhouse-spark/src/main · high confidence
Introduce modular Docker image build system for Spark and Kyuubi components
The build process now uses a layered set of Dockerfiles (scc-base, scc-hadoop, scc-spark, scc-kyuubi, scc-metastore) to construct the runtime environment. This structure installs base utilities in the base image, adds Hadoop and AWS SDK dependencies in the Hadoop image, and layers Spark with specific connectors (Iceberg, TPC-DS, TPC-H, ClickHouse) and JDBC drivers in the Spark image, ensuring a consistent and reproducible setup for the playground environment.
docker/image · high confidence
New Docker-based Kyuubi Playground with integrated CloudBeaver and PostgreSQL metastore
Users can now spin up a complete, one-click development and testing environment using Docker Compose. This playground bundles Kyuubi, Spark 3.5.2, ClickHouse (defaulting to 25.3), Iceberg 1.6.0, and Kyuubi 1.9.2, replacing the previous MySQL metastore with PostgreSQL 12. It includes a web-based SQL interface via CloudBeaver (accessible at localhost:8978) and supports both pre-built release images and local SNAPSHOT builds for developers. The setup is managed via \docker compose\ commands, with separate configurations for stable (\compose.yml\) and development (\compose-dev.yml\) environments.
docker · high confidence
New developer scripts for backporting and code formatting
Added two new executable scripts to the dev directory to streamline maintenance tasks. The \dev/backport\ script allows developers to automatically backport specific commits between supported Spark versions (3.3, 3.4, and 3.5) by generating and applying patches. The \dev/reformat\ script simplifies code style enforcement by running the Spotless formatter across all supported Spark binary versions (3.3, 3.4, and 3.5) in a single step.
dev · high confidence
New standalone Spark 3.5 example for ClickHouse integration
Added a new standalone Scala example (\Saprk-3.5.scala\) demonstrating how to run Apache Spark 3.5 as a standalone application to connect to ClickHouse. The example configures a SparkSession with specific catalog settings for ClickHouse (including host, port, SSL, and credentials) and executes a simple query to list tables, providing a reference implementation for users integrating these technologies.
examples/scala/spark-3.5 · high confidence
Spark 4.0 connector support via new catalog and table provider APIs
The ClickHouse Spark connector now supports Spark 4.0 by introducing a new \ClickHouseCatalog\ implementation that adheres to the Spark 4 \TableCatalog\ and \FunctionCatalog\ interfaces, enabling unified catalog-based table and function access. Alongside this, a new \ClickHouseTableProvider\ has been added to support the \format("clickhouse")\ data source API, allowing users to read and write ClickHouse tables without pre-configured catalogs by specifying connection options directly. The connector also registers itself via a new \DataSourceRegister\ service file, ensuring automatic discovery in Spark 4.0 environments.
spark-4.0/clickhouse-spark · high confidence
Test coverage
Add Cloud test tag for ClickHouse Cloud test filtering; Added integration tests for ClickHouse cluster operations in Spark 3.4; Added integration tests for Spark 3.5 single-node connector; Added integration tests for hash functions and utility parsing; Added log4j configuration for test logging; Added test fixtures for capturing log warnings; Added test infrastructure and utilities for Spark 3.3 integration tests; Added test suites for Spark 3.4 connector configuration, expression handling, and schema mapping; Added tests for ClickHouse SQL engine clause parsing; Added tests for cluster node sorting, shard calculation, and table engine resolution; Added tests for write metrics and read plan stability; Added unit tests for ClickHouse Spark connector configuration, expression, and schema utilities; Added unit tests for ClickHouse Spark connector internals; Added unit tests for NodeClient configuration warnings and WriteMetricsProjection logic; Added unit tests for column parsing and utility functions; Added unit tests for metrics, sharding, scan identity, and write sorting; Integration test suite for Spark 4.0 ClickHouse connector; Integration tests for ClickHouse cluster operations; Refactored test fixtures to support ClickHouse Cloud and cluster environments.
Dependencies
Build system upgraded to support Spark 4.0 and Java 21
The build configuration has been updated to introduce official support for Apache Spark 4.0 (version 4.0.1) alongside existing Spark 3.3, 3.4, and 3.5 versions. To accommodate Spark 4.0 and modern Java environments, the project now targets Java 17 for compilation and includes ASM 9.6 to support Java 21 class files. Additionally, the Spark 4.0 runtime module bundles and relocates Jackson dependencies to prevent conflicts with external libraries, and the Gradle build scripts have been restructured to manage these multi-version dependencies and publishing artifacts.
(dependencies) · high confidence
Upgrade Gradle wrapper to version 8.9
The Gradle wrapper configuration has been updated to use Gradle 8.9. This change ensures that builds will automatically download and use this specific version of the Gradle build tool, providing access to its features and improvements while maintaining consistent build environments across different machines.
gradle · high confidence
Housekeeping
Initial repository structure and documentation for Spark ClickHouse Connector
The repository is initialized with the core project documentation and build scaffolding. This includes a comprehensive CHANGELOG.md detailing the connector's history (from early gRPC-based versions through the current HTTP-only, ClickHouse Java client-based architecture), an AGENTS.md guide for AI-assisted development, and a CONTRIBUTING.md guide for human contributors. The README.md is expanded to provide a compatibility matrix for Spark (3.3–4.0) and ClickHouse versions, build instructions, and usage notes. Build infrastructure is established with a Gradle wrapper, a .scalafmt.conf for code formatting, and a .gitignore. The project license is updated to reflect ClickHouse, Inc. copyright (2016-2024), and a NOTICE file acknowledges the project's donation from The HousePower Organization.
(repo-wide) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 70.
Lenses
- Code Health 89
- Architecture 100
- Maturity 72
- Readiness 74
- Security 62
Changes since last survey
- 300 commits — 264 feature/other, 36 fixes
By area
- (root) — 75 commits
- spark-3.3/clickhouse-spark — 38 commits
- .github/workflows — 37 commits
- clickhouse-core/src — 35 commits
- spark-3.4/clickhouse-spark — 16 commits
- spark-3.3/clickhouse-spark-it — 14 commits
- spark-3.2/clickhouse-spark — 8 commits
- docker/.env — 7 commits
- docker/image — 7 commits
- docs/configurations — 7 commits
- docs/quick_start — 6 commits
- spark-4.0/clickhouse-spark — 5 commits
- docs/internals — 4 commits
- gradle/wrapper — 4 commits
- spark-3.5/clickhouse-spark — 4 commits
- docker/README.md — 3 commits
- docker/conf — 3 commits
- examples/scala — 3 commits
- spark-3.2/clickhouse-spark-it — 3 commits
- spark-3.3/build.gradle — 3 commits
Notable commits
- fix: Fix SonarQube (#422)
- fix: Fix UInt64 type mapping to prevent data loss and overflow (#477)
- fix: Fix sonarqube (#423)
- fix: Fix tests (#419)
- fix: Fix/cloud workflow no env (#496)
- fix: Fix/issue 521 arrow nested variant (#541)
- fix: Fixed ReplacingMergeTree EngineSpec parsing: is_deleted column presence caused error (#357)
- fix: Playground: Fix S3 magic committer confs
- fix: Playground: Fix dev setup
- fix: Remove website deploy & fix docs url (#327)
- fix: Spark 3.3: Fix Decimal precision in JSON mode on reading (#245)
- fix: Spark 3.3: Fix custom options (#234)
- fix: Spark 3.4: Fix Decimal precision in JSON mode on reading (#242)
- fix: Spark 3.4: Fix custom options (#231)
- fix: Spark 3.4: Fix custom rule ExprUtils$CustomResolveTimeZone
- fix: Spark: Fix ArrayIndexOutOfBoundsException when all columns are pruned and agg pushdown does not take effects (#256)
- fix: Spark: Fix Decimal reading in JSON format (#220)
- fix: Spark: Fix reading decimal values (#180)
- fix: Spark: Fix timestamp value transformation (#216)
- fix: Test/ci test fixes (#464)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ClickHouse/spark-clickhouse-connector was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f9befa6704f4c81755fd62846c135301bee58079 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.