Skip to content
CAI
Software that uses CAICheck a score

apache/hbase-connectors

62.7

Adequate · 20 September 2026

11.2k

lines of production code

Scala

with Java

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is the Apache HBase Connectors project, providing integration bridges between HBase and external data processing ecosystems. It includes a Kafka connector that replicates HBase events to Kafka topics using configurable routing rules, and a Spark connector that enables bulk data operations and SQL predicate pushdown across Spark 3 and 4 versions. The codebase also establishes the necessary build infrastructure, testing suites, and operational tooling to manage these connectors.

How it got here

2018 — Initial HBase Connectors release

11 changes.

This period established the Apache HBase Connectors project as a multi-module Maven build, introducing the first releases of Kafka and Spark connectors. Key features included an Avro-based event schema, a configurable HBase-to-Kafka replication proxy, and core Spark integration classes for bulk operations. The release also provided essential infrastructure such as command-line scripts, logging configuration, and comprehensive unit and integration tests to validate the new functionality.

2019–2026 — Spark 4 support and CI infrastructure

6 changes.

This period focused on porting the HBase Spark connector to Spark 4 and refactoring the SQL push-down filter logic to support multiple connector versions. It also established comprehensive development infrastructure, including a dedicated Jenkins pre-commit pipeline, code coverage automation with SonarQube, and standardized code formatting and licensing configurations.

Features

Added Spark filter protocol buffer definitions

The hbase-spark-protocol module now includes a new SparkFilter.proto file defining the SQLPredicatePushDownFilter message and its associated cell-to-column mapping structure. This adds the underlying protocol buffer schema required for pushing down SQL predicates in Spark queries, enabling the serialization of filter logic between Spark and HBase.

spark/hbase-spark-protocol · high confidence

Added code formatting and license configuration files

The dev-support directory now includes configuration files to standardize code style and licensing: .scalafmt.conf defines formatting rules for Scala code (including import sorting and line length), eclipse.importorder sets import organization for Eclipse, hbase\_eclipse\_formatter.xml provides a Java code formatter profile, and license-header contains the standard Apache 2.0 license text template.

dev-support · high confidence

Introduce HBase-to-Kafka replication proxy with configurable routing and drop rules

The Kafka proxy component now provides a bridge that captures HBase replication events and forwards them to Kafka topics based on XML-defined routing and drop rules. Users can configure which tables, column families, or qualifiers are routed to specific topics or excluded entirely using wildcard support in rule definitions. The proxy runs as a standalone service, managing its own Zookeeper connection and Kafka producer, and exposes command-line options to specify broker addresses, rule files, peer names, and whether to automatically enable the replication peer on startup.

kafka/hbase-kafka-proxy/src/main · high confidence

New Avro schema for HBase Kafka events

The HBase Kafka connector now includes a new Avro schema file (HbaseKafkaEvent.avro) that defines the structure of events sent to Kafka. This schema specifies a record named 'HBaseKafkaEvent' with fields for the row key, timestamp, a delete flag, the value, column qualifier, column family, and table name, all encoded as bytes or primitives. This change establishes the data contract for the Kafka connector's event serialization.

kafka/hbase-kafka-model · high confidence

New HBase Connectors command-line scripts

Added the \hbase-connectors\ entry-point script, its configuration helper \hbase-connectors-config.sh\, and the daemon manager \hbase-connectors-daemon.sh\ to the \bin\ directory. These scripts provide the user-facing interface for launching and managing HBase Connector services, including the Kafka Proxy, with support for environment configuration, heap sizing, logging, and daemon lifecycle control (start, stop, restart, autostart).

bin · high confidence

New HBase Spark connector example applications and core support classes

This release adds a suite of Java example applications demonstrating HBase Spark integration patterns, including bulk delete, bulk get, bulk load, bulk put, distributed scanning, and streaming bulk put operations via JavaHBaseContext. It also introduces core Scala support classes required for these operations, such as BulkLoadPartitioner for region-aware data partitioning, ByteArrayWrapper and ColumnFamilyQualifierMapKeyWrapper for efficient map key handling, FamiliesQualifiersValues for organizing bulk load cells, FamilyHFileWriteOptions for HFile configuration, and HBaseConnectionCache for managing HBase connection lifecycle and reuse.

spark/hbase-spark/src/main · high confidence

New Jenkins pre-commit pipeline for HBase Connectors

The HBase Connectors project now has a dedicated CI pre-commit pipeline located in dev-support/jenkins. This change introduces a new Dockerfile (based on Maven 3.9 with JDK 8 and JDK 17) and a Jenkinsfile that orchestrates pre-commit checks using Apache Yetus. The pipeline includes a custom 'hbase-personality.sh' script to handle Maven build phases correctly for Spark modules (including the new Spark4 profile) and a driver script to execute Yetus checks in a Docker container, ensuring consistent testing environments for pull requests.

dev-support/jenkins · high confidence

New code coverage analysis script with SonarQube integration

A new \run-coverage.sh\ script and documentation have been added to the \dev-support/code-coverage\ directory. This script automates the generation of Java and Scala test coverage reports using JaCoCo and SCoverage via Maven. It also supports uploading these results to a SonarQube server when specific parameters (host URL, login credentials, and project key) are provided.

dev-support/code-coverage · high confidence

Port HBase Spark integration to Spark 4

The HBase Spark connector has been ported to support Spark 4. This update introduces the \HBaseContext\ facade for managing HBase connections and executing bulk operations (put, get, delete, scan) via RDD transformations, along with the \NewHBaseRDD\ for reading HBase tables. It also includes the \BulkLoadPartitioner\ and supporting classes to enable efficient bulk loading of HFiles by partitioning data according to HBase region splits, and adds the \HBaseConnectionCache\ to manage connection lifecycle and reuse.

spark4/hbase-spark4 · high confidence

Behavioural changes

Added default logging configuration for HBase Connectors

A new log4j.properties file has been added to the configuration directory to provide default logging settings for the HBase Connectors assembly. This configuration establishes INFO-level logging for the connector root logger and specific components like Kafka integration, while setting metrics-related logs to WARN, ensuring consistent and appropriate log output for users running the connectors.

conf · high confidence

The HBase connectors assembly now bundles a META-INF/LEGAL file and a supplemental-models.xml configuration. This ensures that third-party license information for dependencies such as stax-api, jettison, and bcprov-jdk18on is correctly included in the distribution, and provides the necessary metadata for Maven to resolve license details for these artifacts.

hbase-connectors-assembly/src/main/resources · high confidence

Refactored Spark SQL push-down filter to support multi-version connectors

The Spark SQL push-down filter logic has been moved into the common \hbase-spark-pushdown\ module to decouple it from specific Scala \Field\ types. This change introduces a new \PushdownMappedField\ Java interface, allowing both Spark 3 and Spark 4 connector modules to contribute column mappings without tight coupling. The \SparkSQLPushDownFilter\ now accepts this interface to build family/qualifier mappings, and supporting classes like \ByteArrayComparable\ and \DynamicLogicExpression\ have been added or updated to facilitate this shared, version-agnostic filter execution.

spark/hbase-spark-pushdown · high confidence

Test coverage

Added comprehensive test suite for HBase Spark connector; Added integration test for Spark-based bulk loading; Added test suite for Java HBaseContext; Added unit tests for HBase Kafka Proxy routing and drop rules.

Dependencies

Initial release of Apache HBase Connectors with Kafka and Spark modules

The project is now structured as a multi-module Maven build containing a Kafka connector (proxy and model) and a Spark connector (core, protocol, pushdown, and integration tests). This release establishes the build infrastructure, including dependency management for HBase 2.6.2, Hadoop 3.4.1, Spark, and Avro 1.11.4, and configures the assembly module to package the connectors into a distributable archive.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 63.

Lenses

  • Code Health 91
  • Architecture 100
  • Maturity 50
  • Readiness 56
  • Security 93

Changes since last survey

  • 151 commits — 135 feature/other, 16 fixes

By area

  • (root) — 57 commits
  • spark/hbase-spark — 37 commits
  • kafka/hbase-kafka-proxy — 13 commits
  • (repo) — 10 commits
  • dev-support/jenkins — 6 commits
  • dev-support/hbase-personality.sh — 5 commits
  • hbase-connectors-assembly/src — 4 commits
  • spark/pom.xml — 4 commits
  • spark4/hbase-spark4 — 4 commits
  • dev-support/Dockerfile — 2 commits
  • kafka/README.md — 2 commits
  • spark/README.md — 2 commits
  • dev-support/code-coverage — 1 commit
  • hbase-kafka-proxy/src — 1 commit
  • kafka/pom.xml — 1 commit
  • spark/hbase-spark-it — 1 commit
  • spark/hbase-spark-pushdown — 1 commit

Notable commits

  • fix: HBASE-18570 Fix NPE when HBaseContext was never initialized (#127)
  • fix: HBASE-21431 Fix build and test issues
  • fix: HBASE-21878 Fix hbase-checkstyle version reference
  • fix: HBASE-22210 Fix hbase-connectors-assembly to include every jar (#20)
  • fix: HBASE-22318 Fix for warning The POM for org.glassfish.javax.el is missing (#26)
  • fix: HBASE-22319 Fix for warning The assembly descriptor contains a filesystem-root relative reference (#28)
  • fix: HBASE-22329 Fix for warning The parameter forkMode is deprecated since version in hbase-spark-it (#30)
  • fix: HBASE-23579 Fixed Checkstyle issues
  • fix: HBASE-26211 Fix decoding of Long values in NaiveEncoder (#83)
  • fix: HBASE-26863 fix incorrect rowkey pushdown (#95)
  • fix: HBASE-27285: Fix sonar report paths (#103)
  • fix: HBASE-28534 Fix Kerberos authentication failure in local mode (#128)
  • fix: Revert "HBASE-21430 [hbase-connectors] Move hbase-spark* modules to hbase-connectors repo"
  • fix: Revert "HBASE-21430 [hbase-connectors] Move hbase-spark* modules to hbase-connectors repo"
  • fix: Revert "Preparing hbase-connectors release 1.0.1RC1; tagging and updates to CHANGESLOG.md and RELEASENOTES.md"
  • fix: Revert "upgrade jackson version"
  • change: Edit on the READMEs to accommodate new kafka proxy.
  • change: First commit of LICENSE, README, and .gitignore
  • change: Formatting fixup on Readmes
  • change: HBASE-20934 Create an hbase-connectors repository; commit new kafka connect here
  • …and 131 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

apache/hbase-connectors was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 9c73e6415328dac2288b8aaa860470b7c290a5e2 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.