Skip to content
CAI
Software that uses CAICheck a score

AbsaOSS/ABRiS

58.4

Adequate · 20 September 2026

2.3k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

ABRiS is a library that enables the serialization and deserialization of Spark DataFrames to and from Avro format, with deep integration for Confluent Schema Registry. It provides configurable error handling for malformed records, supports schema conversion strategies, and ensures compatibility across Spark 3.x and 4.x versions. The system also includes utilities for managing schema retrieval, caching, and testing within Kafka streaming pipelines.

How it got here

2018 — Spark 4.2 migration and Avro refactoring

17 changes.

The project upgraded its baseline to Spark 4.2, Scala 2.13, and Java 17, replacing the legacy spark-avro-dataframes module with a modernized architecture. This period involved removing obsolete custom Avro parsing logic and IDE metadata while introducing new utilities for schema conversion, configurable schema registry integration, and comprehensive example applications.

2019–2022 — Spark 4.0 compatibility and API refactoring

7 changes.

This period focused on extending ABRiS support to Spark 4.0 through a reflection-based compatibility layer and refactoring the internal SQL expression implementations. The work also involved restructuring the configuration API, decoupling the Confluent Schema Registry client, and introducing configurable error handling for deserialization.

Features

Add Spark 4.0 compatibility via reflection-based column conversion layer

Users can now run ABRiS on Spark 4.0 in addition to Spark 3.x. This is achieved by introducing a new \SparkColumnCompat\ utility that uses reflection to handle the differences in how \Column\ and \Expression\ are converted between the two versions, and a new \functions\ object that exposes \to\_avro\ and \from\_avro\ methods using this compatibility layer. This allows the existing Avro conversion functionality to work seamlessly across both Spark major versions without code changes.

src/main/scala/za/co/absa/abris/avro · high confidence

Added Avro schema definitions for native type testing

New Avro schema files have been added to define record structures for testing native data types. The 'NativeComplete' schema includes fields for bytes, nullable strings, integers, longs, doubles, floats, booleans, arrays, maps, and fixed-size types. The 'NativeSimpleOuter' schema defines a record with a string field and a nested record containing integer and long fields.

src/main/avro · high confidence

Added Confluent Kafka Avro Reader and Writer examples

New example applications have been added to demonstrate how to read from and write to Kafka topics using Confluent Schema Registry with Avro serialization. The ConfluentKafkaAvroReader example shows how to configure a streaming Spark job to deserialize Avro data using the topic-name strategy, while the ConfluentKafkaAvroWriter example illustrates how to serialize Spark DataFrames into Avro format and publish them to a Kafka topic, including schema registration with the registry.

src/main/scala/za/co/absa/abris/examples · high confidence

Added example data generation utilities for Avro schemas

New helper classes have been added to the examples package to facilitate the creation of test and example data. ComplexRecordsGenerator provides methods to generate Spark Rows containing various data types (including bytes, strings, numbers, booleans, arrays, and maps) based on defined Avro schemas. FixedString offers a utility for handling Avro fixed-length fields, and TestSchemas exposes a collection of pre-defined Avro schema specifications (such as native complete, simple, array, map, decimal, date, and timestamp schemas) for use in testing and demonstration purposes.

src/main/scala/za/co/absa/abris/examples/data · high confidence

Introduces configurable Avro deserialization error handling

Users can now control how malformed Avro records are processed during deserialization through a new extensible error-handling framework. The \DeserializationExceptionHandler\ trait allows for custom strategies, with two built-in implementations provided: \FailFastExceptionHandler\, which throws a \SparkException\ to stop processing immediately, and \PermissiveRecordExceptionHandler\, which logs a warning and replaces the malformed record with a null-filled row to allow the pipeline to continue.

src/main/scala/za/co/absa/abris/avro/errors · high confidence

New Avro serialization and deserialization wrappers

Added AbrisAvroSerializer and AbrisAvroDeserializer classes in the org.apache.spark.sql.avro package to wrap Spark's internal AvroSerializer and AvroDeserializer. These wrappers provide a public API for serializing and deserializing Catalyst data types to and from Avro schemas, with the deserializer explicitly configured to use LEGACY mode and the serializer supporting nullable types.

src/main/scala/org · high confidence

New AvroSchemaUtils utility for schema conversion and loading

A new AvroSchemaUtils object has been added to the parsing utilities, providing methods to parse and load Avro schemas from file paths, as well as convert Spark DataFrame columns into Avro schemas. This includes support for converting single or multiple columns, specifying record names and namespaces, and wrapping schemas, facilitating easier schema management and generation within the library.

src/main/scala/za/co/absa/abris/avro/parsing · high confidence

New SparkAvroConversions utility for schema translation

A new \SparkAvroConversions\ object has been added to the \za.co.absa.abris.avro.format\ package to provide utilities for converting between Spark SQL types and Avro schemas. This component exposes methods to translate a Spark \StructType\ into an Avro \Schema\ (with support for custom names and namespaces) and to parse Avro schema strings or objects back into Spark \StructType\s, leveraging the underlying Spark-Avro library for the actual conversion logic.

src/main/scala/za/co/absa/abris/avro/format · high confidence

New example utilities for Spark session management and Avro encoding compatibility

Added two new utility classes to the examples package: CompatibleRowEncoder, which provides a unified way to create Row encoders that works across Spark versions prior to and including 3.5.0 by dynamically selecting the appropriate API, and ExamplesUtils, which offers helper methods for loading configuration properties, building Spark sessions, and applying options to DataStreamReaders and DataStreamWriters for both Row and byte-array payloads.

src/main/scala/za/co/absa/abris/examples/utils · high confidence

Removals

Removal of AvroDeserializer integration class

The AvroDeserializer object, which previously provided an implicit extension to Spark's DataStreamReader for decoding binary Avro records into Spark Rows, has been removed from the library. Users can no longer use the .avro(schema) method on DataStreamReader instances to automatically parse Avro data.

spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes · high confidence

Removal of AvroParser utility class

The \AvroParser\ class, which previously handled the conversion of Avro \GenericRecord\s to Spark \GenericRowWithSchema\ objects and translated Avro schemas to Spark \StructType\ using the Databricks Spark-Avro library, has been removed from the codebase. This change eliminates the custom parsing logic and implicit \Utf8Unwrapper\ extension, likely shifting responsibility for Avro-to-Spark data conversion to other components or a different implementation strategy.

spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/parsing · high confidence

Removal of SampleFilterApp example

The SampleFilterApp example, which demonstrated reading Avro data from a Kafka stream and applying a filter, has been removed from the examples directory. Users relying on this specific sample for reference or testing will no longer have access to this code snippet.

spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/examples · high confidence

Removal of custom Scala Avro parsing classes

The custom Scala-specific Avro parsing components—ScalaDatumReader, ScalaRecord, and ScalaSpecificData—have been removed from the spark-avro-dataframes module. These classes previously handled the runtime translation of Avro types (such as Strings, Arrays, and Maps) into Scala collections and converted nested Avro records into Spark Rows to ensure queryable nested structures. Their deletion indicates a shift away from this manual translation layer, likely relying on default Avro behavior or a different integration strategy for data conversion.

spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/avro · high confidence

Behavioural changes

Configurable schema converter and reduced logging noise

Users can now select a custom schema converter implementation by registering it in the META-INF/services/za.co.absa.abris.avro.sql.SchemaConverter file, replacing the previous default. Additionally, default logging levels have been adjusted to suppress verbose internal messages from Kafka, Spark, and Abris components, showing only errors by default to reduce console noise.

src/main/resources · high confidence

New SQL expression implementations and configurable schema conversion

The library introduces new internal Spark SQL expression classes, \AvroDataToCatalyst\ and \CatalystDataToAvro\, which handle the actual deserialization and serialization of Avro data within Spark queries. These expressions support Confluent Schema Registry integration (including magic byte handling and schema ID attachment) and allow for a configurable schema converter via a new \SchemaConverter\ trait and \DefaultSchemaConverter\ implementation, enabling users to customize how Avro schemas map to Spark SQL types.

src/main/scala/za/co/absa/abris/avro/sql · high confidence

Refactored Confluent Schema Registry client abstraction and subject naming

The registry package has been restructured to decouple the Confluent Schema Registry integration from internal naming logic. A new \AbrisRegistryClient\ trait and \AbstractConfluentRegistryClient\ base class now provide a cleaner interface for schema operations, while \SchemaSubject\ and \SchemaVersion\ types explicitly model registry subjects and versions. This change removes the dependency on Confluent's internal naming strategy code, allowing the library to support multiple Spark Avro versions and improving configuration flexibility without altering the external API.

src/main/scala/za/co/absa/abris/avro/registry · high confidence

Refactored Schema Registry integration with caching and custom client support

The Confluent Schema Registry integration in the read module has been refactored to improve reliability and extensibility. The new \SchemaManager\ and \SchemaManagerFactory\ introduce thread-safe caching of Schema Registry client instances to avoid redundant connections, support for custom registry client implementations via configuration, and a mock registry client for testing (triggered by URLs starting with \mock://\). Additionally, the schema retrieval logic now explicitly supports fetching schemas by both ID and subject/version, and includes better error handling and logging for registry interactions.

src/main/scala/za/co/absa/abris/avro/read · high confidence

Refactored configuration API with new builder classes and internal config parsers

The configuration system in \src/main/scala/za/co/absa/abris/config\ has been restructured to introduce a new, backward-compatible builder API. A new \Config.scala\ file defines fluent builder fragments (e.g., \ToSimpleAvroConfigFragment\, \FromSimpleAvroConfigFragment\) that allow users to configure schema sources (simple vs. Confluent), download strategies (by ID, version, or latest), and registration strategies. The core \ToAvroConfig\ class is now serializable and supports explicit schema and schema ID settings. Additionally, internal configuration parsing is handled by new \InternalFromAvroConfig\ and \InternalToAvroConfig\ classes, which extract reader/writer schemas, schema converters, and deserialization exception handlers from configuration maps, supporting more granular control over Avro conversion behavior.

src/main/scala/za/co/absa/abris/config · high confidence

Removal of Eclipse IDE configuration files

The project-specific Eclipse settings files (org.eclipse.jdt.core.prefs and org.eclipse.m2e.core.prefs) for the spark-avro-dataframes module have been removed. This change eliminates local IDE configuration for Java compiler compliance (set to 1.8) and M2E workspace resolution, likely to standardize build configurations or remove obsolete IDE-specific metadata.

spark-avro-dataframes/.settings · high confidence

Test coverage

Added unit tests for Avro deserialization error handling and schema management; Removal of legacy Avro parsing test infrastructure; Removal of legacy test resource files.

Dependencies

Upgrade to Spark 4.2 and Scala 2.13 with Java 17 baseline

The ABRiS library has been updated to target Spark 4.2.0 and Scala 2.13.17, raising the minimum Java version to 17. This change replaces the legacy \spark-avro-dataframes\ module with a modernized \pom.xml\ that uses the Confluent Platform 8.3.1 for Kafka and Schema Registry integration, updates Avro to 1.12.1, and refreshes test dependencies to ScalaTest 3.2.20 and Mockito 5.12.

(dependencies) · high confidence

Housekeeping

Removal of Eclipse IDE metadata files

The project has removed Eclipse-specific configuration files, including .classpath, .project, and .gitignore. This cleanup removes local IDE metadata that is typically managed by build tools like Maven, ensuring a cleaner repository state without affecting the core library functionality.

spark-avro-dataframes · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 58.

Lenses

  • Code Health 100
  • Architecture 79
  • Maturity 51
  • Readiness 70
  • Security 53

Changes since last survey

  • 300 commits — 247 feature/other, 53 fixes

By area

  • (root) — 118 commits
  • (repo) — 77 commits
  • src/main — 68 commits
  • src/test — 18 commits
  • .github/workflows — 13 commits
  • documentation/confluent-avro-documentation.md — 4 commits
  • documentation/python-documentation.md — 2 commits

Notable commits

  • fix: #170: Fix docs for schema generation
  • fix: #325 - bug fix - update jacoco version (#327)
  • fix: Abris #123 fix compiler warnings
  • fix: Abris #135 Fix unit tests that did not work with scala 2.11.
  • fix: Abris #246 fix profile for Spark 2.3
  • fix: Abris #256 fix deprecation for Scala 2.13 in non test code
  • fix: Abris #256 fix deprecation for Scala 2.13 in tests
  • fix: Abris #256 fix deprecation for apache.commons.io and scalatest
  • fix: Feature/350 fix tests (#351)
  • fix: Feature/fix scalastyle (#292)
  • fix: Fix Capitalization of names in READ.ME
  • fix: Fix README
  • fix: Fix bug when accessing schema registry from inside of executor
  • fix: Fix bugs with bare and multiply nested bugs #56 #58
  • fix: Fix confluent schema evolution functionality
  • fix: Fix of README.md
  • fix: Fix scm url
  • fix: Fix security settings example (#91)
  • fix: Fix subject string in debug messages (#177)
  • fix: Merge branch 'master' into fix/public
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

AbsaOSS/ABRiS was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 0d9ffb020eb4e76725a559dd1543971ed089dd8d — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.