apache/hbase-connectors
62.7
Adequate · 20 September 2026
11.2k
lines of production code
Scala
with Java
1
measurement over time
What this system is
This system is the Apache HBase Connectors project, providing integration bridges between HBase and external data processing ecosystems. It includes a Kafka connector that replicates HBase events to Kafka topics using configurable routing rules, and a Spark connector that enables bulk data operations and SQL predicate pushdown across Spark 3 and 4 versions. The codebase also establishes the necessary build infrastructure, testing suites, and operational tooling to manage these connectors.
How it got here
2018 — Initial HBase Connectors release
11 changes.
This period established the Apache HBase Connectors project as a multi-module Maven build, introducing the first releases of Kafka and Spark connectors. Key features included an Avro-based event schema, a configurable HBase-to-Kafka replication proxy, and core Spark integration classes for bulk operations. The release also provided essential infrastructure such as command-line scripts, logging configuration, and comprehensive unit and integration tests to validate the new functionality.
2019–2026 — Spark 4 support and CI infrastructure
6 changes.
This period focused on porting the HBase Spark connector to Spark 4 and refactoring the SQL push-down filter logic to support multiple connector versions. It also established comprehensive development infrastructure, including a dedicated Jenkins pre-commit pipeline, code coverage automation with SonarQube, and standardized code formatting and licensing configurations.
Features
Added Spark filter protocol buffer definitions
The hbase-spark-protocol module now includes a new SparkFilter.proto file defining the SQLPredicatePushDownFilter message and its associated cell-to-column mapping structure. This adds the underlying protocol buffer schema required for pushing down SQL predicates in Spark queries, enabling the serialization of filter logic between Spark and HBase.
spark/hbase-spark-protocol · high confidence
Added code formatting and license configuration files
The dev-support directory now includes configuration files to standardize code style and licensing: .scalafmt.conf defines formatting rules for Scala code (including import sorting and line length), eclipse.importorder sets import organization for Eclipse, hbase\_eclipse\_formatter.xml provides a Java code formatter profile, and license-header contains the standard Apache 2.0 license text template.
dev-support · high confidence
Introduce HBase-to-Kafka replication proxy with configurable routing and drop rules
The Kafka proxy component now provides a bridge that captures HBase replication events and forwards them to Kafka topics based on XML-defined routing and drop rules. Users can configure which tables, column families, or qualifiers are routed to specific topics or excluded entirely using wildcard support in rule definitions. The proxy runs as a standalone service, managing its own Zookeeper connection and Kafka producer, and exposes command-line options to specify broker addresses, rule files, peer names, and whether to automatically enable the replication peer on startup.
kafka/hbase-kafka-proxy/src/main · high confidence
New Avro schema for HBase Kafka events
The HBase Kafka connector now includes a new Avro schema file (HbaseKafkaEvent.avro) that defines the structure of events sent to Kafka. This schema specifies a record named 'HBaseKafkaEvent' with fields for the row key, timestamp, a delete flag, the value, column qualifier, column family, and table name, all encoded as bytes or primitives. This change establishes the data contract for the Kafka connector's event serialization.
kafka/hbase-kafka-model · high confidence
New HBase Connectors command-line scripts
Added the \hbase-connectors\ entry-point script, its configuration helper \hbase-connectors-config.sh\, and the daemon manager \hbase-connectors-daemon.sh\ to the \bin\ directory. These scripts provide the user-facing interface for launching and managing HBase Connector services, including the Kafka Proxy, with support for environment configuration, heap sizing, logging, and daemon lifecycle control (start, stop, restart, autostart).
bin · high confidence
New HBase Spark connector example applications and core support classes
This release adds a suite of Java example applications demonstrating HBase Spark integration patterns, including bulk delete, bulk get, bulk load, bulk put, distributed scanning, and streaming bulk put operations via JavaHBaseContext. It also introduces core Scala support classes required for these operations, such as BulkLoadPartitioner for region-aware data partitioning, ByteArrayWrapper and ColumnFamilyQualifierMapKeyWrapper for efficient map key handling, FamiliesQualifiersValues for organizing bulk load cells, FamilyHFileWriteOptions for HFile configuration, and HBaseConnectionCache for managing HBase connection lifecycle and reuse.
spark/hbase-spark/src/main · high confidence
New Jenkins pre-commit pipeline for HBase Connectors
The HBase Connectors project now has a dedicated CI pre-commit pipeline located in dev-support/jenkins. This change introduces a new Dockerfile (based on Maven 3.9 with JDK 8 and JDK 17) and a Jenkinsfile that orchestrates pre-commit checks using Apache Yetus. The pipeline includes a custom 'hbase-personality.sh' script to handle Maven build phases correctly for Spark modules (including the new Spark4 profile) and a driver script to execute Yetus checks in a Docker container, ensuring consistent testing environments for pull requests.
dev-support/jenkins · high confidence
New code coverage analysis script with SonarQube integration
A new \run-coverage.sh\ script and documentation have been added to the \dev-support/code-coverage\ directory. This script automates the generation of Java and Scala test coverage reports using JaCoCo and SCoverage via Maven. It also supports uploading these results to a SonarQube server when specific parameters (host URL, login credentials, and project key) are provided.
dev-support/code-coverage · high confidence
Port HBase Spark integration to Spark 4
The HBase Spark connector has been ported to support Spark 4. This update introduces the \HBaseContext\ facade for managing HBase connections and executing bulk operations (put, get, delete, scan) via RDD transformations, along with the \NewHBaseRDD\ for reading HBase tables. It also includes the \BulkLoadPartitioner\ and supporting classes to enable efficient bulk loading of HFiles by partitioning data according to HBase region splits, and adds the \HBaseConnectionCache\ to manage connection lifecycle and reuse.
spark4/hbase-spark4 · high confidence
Behavioural changes
Added default logging configuration for HBase Connectors
A new log4j.properties file has been added to the configuration directory to provide default logging settings for the HBase Connectors assembly. This configuration establishes INFO-level logging for the connector root logger and specific components like Kafka integration, while setting metrics-related logs to WARN, ensuring consistent and appropriate log output for users running the connectors.
conf · high confidence
Assembly includes legal notices and license metadata
The HBase connectors assembly now bundles a META-INF/LEGAL file and a supplemental-models.xml configuration. This ensures that third-party license information for dependencies such as stax-api, jettison, and bcprov-jdk18on is correctly included in the distribution, and provides the necessary metadata for Maven to resolve license details for these artifacts.
hbase-connectors-assembly/src/main/resources · high confidence
Refactored Spark SQL push-down filter to support multi-version connectors
The Spark SQL push-down filter logic has been moved into the common \hbase-spark-pushdown\ module to decouple it from specific Scala \Field\ types. This change introduces a new \PushdownMappedField\ Java interface, allowing both Spark 3 and Spark 4 connector modules to contribute column mappings without tight coupling. The \SparkSQLPushDownFilter\ now accepts this interface to build family/qualifier mappings, and supporting classes like \ByteArrayComparable\ and \DynamicLogicExpression\ have been added or updated to facilitate this shared, version-agnostic filter execution.
spark/hbase-spark-pushdown · high confidence
Test coverage
Added comprehensive test suite for HBase Spark connector; Added integration test for Spark-based bulk loading; Added test suite for Java HBaseContext; Added unit tests for HBase Kafka Proxy routing and drop rules.
Dependencies
Initial release of Apache HBase Connectors with Kafka and Spark modules
The project is now structured as a multi-module Maven build containing a Kafka connector (proxy and model) and a Spark connector (core, protocol, pushdown, and integration tests). This release establishes the build infrastructure, including dependency management for HBase 2.6.2, Hadoop 3.4.1, Spark, and Avro 1.11.4, and configures the assembly module to package the connectors into a distributable archive.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 63.
Lenses
- Code Health 91
- Architecture 100
- Maturity 50
- Readiness 56
- Security 93
Changes since last survey
- 151 commits — 135 feature/other, 16 fixes
By area
- (root) — 57 commits
- spark/hbase-spark — 37 commits
- kafka/hbase-kafka-proxy — 13 commits
- (repo) — 10 commits
- dev-support/jenkins — 6 commits
- dev-support/hbase-personality.sh — 5 commits
- hbase-connectors-assembly/src — 4 commits
- spark/pom.xml — 4 commits
- spark4/hbase-spark4 — 4 commits
- dev-support/Dockerfile — 2 commits
- kafka/README.md — 2 commits
- spark/README.md — 2 commits
- dev-support/code-coverage — 1 commit
- hbase-kafka-proxy/src — 1 commit
- kafka/pom.xml — 1 commit
- spark/hbase-spark-it — 1 commit
- spark/hbase-spark-pushdown — 1 commit
Notable commits
- fix: HBASE-18570 Fix NPE when HBaseContext was never initialized (#127)
- fix: HBASE-21431 Fix build and test issues
- fix: HBASE-21878 Fix hbase-checkstyle version reference
- fix: HBASE-22210 Fix hbase-connectors-assembly to include every jar (#20)
- fix: HBASE-22318 Fix for warning The POM for org.glassfish.javax.el is missing (#26)
- fix: HBASE-22319 Fix for warning The assembly descriptor contains a filesystem-root relative reference (#28)
- fix: HBASE-22329 Fix for warning The parameter forkMode is deprecated since version in hbase-spark-it (#30)
- fix: HBASE-23579 Fixed Checkstyle issues
- fix: HBASE-26211 Fix decoding of Long values in NaiveEncoder (#83)
- fix: HBASE-26863 fix incorrect rowkey pushdown (#95)
- fix: HBASE-27285: Fix sonar report paths (#103)
- fix: HBASE-28534 Fix Kerberos authentication failure in local mode (#128)
- fix: Revert "HBASE-21430 [hbase-connectors] Move hbase-spark* modules to hbase-connectors repo"
- fix: Revert "HBASE-21430 [hbase-connectors] Move hbase-spark* modules to hbase-connectors repo"
- fix: Revert "Preparing hbase-connectors release 1.0.1RC1; tagging and updates to CHANGESLOG.md and RELEASENOTES.md"
- fix: Revert "upgrade jackson version"
- change: Edit on the READMEs to accommodate new kafka proxy.
- change: First commit of LICENSE, README, and .gitignore
- change: Formatting fixup on Readmes
- change: HBASE-20934 Create an hbase-connectors repository; commit new kafka connect here
- …and 131 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
apache/hbase-connectors was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 9c73e6415328dac2288b8aaa860470b7c290a5e2 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.