AbsaOSS/ABRiS
58.4
Adequate · 20 September 2026
2.3k
lines of production code
Scala
primary language
1
measurement over time
What this system is
ABRiS is a library that enables the serialization and deserialization of Spark DataFrames to and from Avro format, with deep integration for Confluent Schema Registry. It provides configurable error handling for malformed records, supports schema conversion strategies, and ensures compatibility across Spark 3.x and 4.x versions. The system also includes utilities for managing schema retrieval, caching, and testing within Kafka streaming pipelines.
How it got here
2018 — Spark 4.2 migration and Avro refactoring
17 changes.
The project upgraded its baseline to Spark 4.2, Scala 2.13, and Java 17, replacing the legacy spark-avro-dataframes module with a modernized architecture. This period involved removing obsolete custom Avro parsing logic and IDE metadata while introducing new utilities for schema conversion, configurable schema registry integration, and comprehensive example applications.
2019–2022 — Spark 4.0 compatibility and API refactoring
7 changes.
This period focused on extending ABRiS support to Spark 4.0 through a reflection-based compatibility layer and refactoring the internal SQL expression implementations. The work also involved restructuring the configuration API, decoupling the Confluent Schema Registry client, and introducing configurable error handling for deserialization.
Features
Add Spark 4.0 compatibility via reflection-based column conversion layer
Users can now run ABRiS on Spark 4.0 in addition to Spark 3.x. This is achieved by introducing a new \SparkColumnCompat\ utility that uses reflection to handle the differences in how \Column\ and \Expression\ are converted between the two versions, and a new \functions\ object that exposes \to\_avro\ and \from\_avro\ methods using this compatibility layer. This allows the existing Avro conversion functionality to work seamlessly across both Spark major versions without code changes.
src/main/scala/za/co/absa/abris/avro · high confidence
Added Avro schema definitions for native type testing
New Avro schema files have been added to define record structures for testing native data types. The 'NativeComplete' schema includes fields for bytes, nullable strings, integers, longs, doubles, floats, booleans, arrays, maps, and fixed-size types. The 'NativeSimpleOuter' schema defines a record with a string field and a nested record containing integer and long fields.
src/main/avro · high confidence
Added Confluent Kafka Avro Reader and Writer examples
New example applications have been added to demonstrate how to read from and write to Kafka topics using Confluent Schema Registry with Avro serialization. The ConfluentKafkaAvroReader example shows how to configure a streaming Spark job to deserialize Avro data using the topic-name strategy, while the ConfluentKafkaAvroWriter example illustrates how to serialize Spark DataFrames into Avro format and publish them to a Kafka topic, including schema registration with the registry.
src/main/scala/za/co/absa/abris/examples · high confidence
Added example data generation utilities for Avro schemas
New helper classes have been added to the examples package to facilitate the creation of test and example data. ComplexRecordsGenerator provides methods to generate Spark Rows containing various data types (including bytes, strings, numbers, booleans, arrays, and maps) based on defined Avro schemas. FixedString offers a utility for handling Avro fixed-length fields, and TestSchemas exposes a collection of pre-defined Avro schema specifications (such as native complete, simple, array, map, decimal, date, and timestamp schemas) for use in testing and demonstration purposes.
src/main/scala/za/co/absa/abris/examples/data · high confidence
Introduces configurable Avro deserialization error handling
Users can now control how malformed Avro records are processed during deserialization through a new extensible error-handling framework. The \DeserializationExceptionHandler\ trait allows for custom strategies, with two built-in implementations provided: \FailFastExceptionHandler\, which throws a \SparkException\ to stop processing immediately, and \PermissiveRecordExceptionHandler\, which logs a warning and replaces the malformed record with a null-filled row to allow the pipeline to continue.
src/main/scala/za/co/absa/abris/avro/errors · high confidence
New Avro serialization and deserialization wrappers
Added AbrisAvroSerializer and AbrisAvroDeserializer classes in the org.apache.spark.sql.avro package to wrap Spark's internal AvroSerializer and AvroDeserializer. These wrappers provide a public API for serializing and deserializing Catalyst data types to and from Avro schemas, with the deserializer explicitly configured to use LEGACY mode and the serializer supporting nullable types.
src/main/scala/org · high confidence
New AvroSchemaUtils utility for schema conversion and loading
A new AvroSchemaUtils object has been added to the parsing utilities, providing methods to parse and load Avro schemas from file paths, as well as convert Spark DataFrame columns into Avro schemas. This includes support for converting single or multiple columns, specifying record names and namespaces, and wrapping schemas, facilitating easier schema management and generation within the library.
src/main/scala/za/co/absa/abris/avro/parsing · high confidence
New SparkAvroConversions utility for schema translation
A new \SparkAvroConversions\ object has been added to the \za.co.absa.abris.avro.format\ package to provide utilities for converting between Spark SQL types and Avro schemas. This component exposes methods to translate a Spark \StructType\ into an Avro \Schema\ (with support for custom names and namespaces) and to parse Avro schema strings or objects back into Spark \StructType\s, leveraging the underlying Spark-Avro library for the actual conversion logic.
src/main/scala/za/co/absa/abris/avro/format · high confidence
New example utilities for Spark session management and Avro encoding compatibility
Added two new utility classes to the examples package: CompatibleRowEncoder, which provides a unified way to create Row encoders that works across Spark versions prior to and including 3.5.0 by dynamically selecting the appropriate API, and ExamplesUtils, which offers helper methods for loading configuration properties, building Spark sessions, and applying options to DataStreamReaders and DataStreamWriters for both Row and byte-array payloads.
src/main/scala/za/co/absa/abris/examples/utils · high confidence
Removals
Removal of AvroDeserializer integration class
The AvroDeserializer object, which previously provided an implicit extension to Spark's DataStreamReader for decoding binary Avro records into Spark Rows, has been removed from the library. Users can no longer use the .avro(schema) method on DataStreamReader instances to automatically parse Avro data.
spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes · high confidence
Removal of AvroParser utility class
The \AvroParser\ class, which previously handled the conversion of Avro \GenericRecord\s to Spark \GenericRowWithSchema\ objects and translated Avro schemas to Spark \StructType\ using the Databricks Spark-Avro library, has been removed from the codebase. This change eliminates the custom parsing logic and implicit \Utf8Unwrapper\ extension, likely shifting responsibility for Avro-to-Spark data conversion to other components or a different implementation strategy.
spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/parsing · high confidence
Removal of SampleFilterApp example
The SampleFilterApp example, which demonstrated reading Avro data from a Kafka stream and applying a filter, has been removed from the examples directory. Users relying on this specific sample for reference or testing will no longer have access to this code snippet.
spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/examples · high confidence
Removal of custom Scala Avro parsing classes
The custom Scala-specific Avro parsing components—ScalaDatumReader, ScalaRecord, and ScalaSpecificData—have been removed from the spark-avro-dataframes module. These classes previously handled the runtime translation of Avro types (such as Strings, Arrays, and Maps) into Scala collections and converted nested Avro records into Spark Rows to ensure queryable nested structures. Their deletion indicates a shift away from this manual translation layer, likely relying on default Avro behavior or a different integration strategy for data conversion.
spark-avro-dataframes/src/main/scala/za/co/absa/avro/dataframes/avro · high confidence
Behavioural changes
Configurable schema converter and reduced logging noise
Users can now select a custom schema converter implementation by registering it in the META-INF/services/za.co.absa.abris.avro.sql.SchemaConverter file, replacing the previous default. Additionally, default logging levels have been adjusted to suppress verbose internal messages from Kafka, Spark, and Abris components, showing only errors by default to reduce console noise.
src/main/resources · high confidence
New SQL expression implementations and configurable schema conversion
The library introduces new internal Spark SQL expression classes, \AvroDataToCatalyst\ and \CatalystDataToAvro\, which handle the actual deserialization and serialization of Avro data within Spark queries. These expressions support Confluent Schema Registry integration (including magic byte handling and schema ID attachment) and allow for a configurable schema converter via a new \SchemaConverter\ trait and \DefaultSchemaConverter\ implementation, enabling users to customize how Avro schemas map to Spark SQL types.
src/main/scala/za/co/absa/abris/avro/sql · high confidence
Refactored Confluent Schema Registry client abstraction and subject naming
The registry package has been restructured to decouple the Confluent Schema Registry integration from internal naming logic. A new \AbrisRegistryClient\ trait and \AbstractConfluentRegistryClient\ base class now provide a cleaner interface for schema operations, while \SchemaSubject\ and \SchemaVersion\ types explicitly model registry subjects and versions. This change removes the dependency on Confluent's internal naming strategy code, allowing the library to support multiple Spark Avro versions and improving configuration flexibility without altering the external API.
src/main/scala/za/co/absa/abris/avro/registry · high confidence
Refactored Schema Registry integration with caching and custom client support
The Confluent Schema Registry integration in the read module has been refactored to improve reliability and extensibility. The new \SchemaManager\ and \SchemaManagerFactory\ introduce thread-safe caching of Schema Registry client instances to avoid redundant connections, support for custom registry client implementations via configuration, and a mock registry client for testing (triggered by URLs starting with \mock://\). Additionally, the schema retrieval logic now explicitly supports fetching schemas by both ID and subject/version, and includes better error handling and logging for registry interactions.
src/main/scala/za/co/absa/abris/avro/read · high confidence
Refactored configuration API with new builder classes and internal config parsers
The configuration system in \src/main/scala/za/co/absa/abris/config\ has been restructured to introduce a new, backward-compatible builder API. A new \Config.scala\ file defines fluent builder fragments (e.g., \ToSimpleAvroConfigFragment\, \FromSimpleAvroConfigFragment\) that allow users to configure schema sources (simple vs. Confluent), download strategies (by ID, version, or latest), and registration strategies. The core \ToAvroConfig\ class is now serializable and supports explicit schema and schema ID settings. Additionally, internal configuration parsing is handled by new \InternalFromAvroConfig\ and \InternalToAvroConfig\ classes, which extract reader/writer schemas, schema converters, and deserialization exception handlers from configuration maps, supporting more granular control over Avro conversion behavior.
src/main/scala/za/co/absa/abris/config · high confidence
Removal of Eclipse IDE configuration files
The project-specific Eclipse settings files (org.eclipse.jdt.core.prefs and org.eclipse.m2e.core.prefs) for the spark-avro-dataframes module have been removed. This change eliminates local IDE configuration for Java compiler compliance (set to 1.8) and M2E workspace resolution, likely to standardize build configurations or remove obsolete IDE-specific metadata.
spark-avro-dataframes/.settings · high confidence
Test coverage
Added unit tests for Avro deserialization error handling and schema management; Removal of legacy Avro parsing test infrastructure; Removal of legacy test resource files.
Dependencies
Upgrade to Spark 4.2 and Scala 2.13 with Java 17 baseline
The ABRiS library has been updated to target Spark 4.2.0 and Scala 2.13.17, raising the minimum Java version to 17. This change replaces the legacy \spark-avro-dataframes\ module with a modernized \pom.xml\ that uses the Confluent Platform 8.3.1 for Kafka and Schema Registry integration, updates Avro to 1.12.1, and refreshes test dependencies to ScalaTest 3.2.20 and Mockito 5.12.
(dependencies) · high confidence
Housekeeping
Removal of Eclipse IDE metadata files
The project has removed Eclipse-specific configuration files, including .classpath, .project, and .gitignore. This cleanup removes local IDE metadata that is typically managed by build tools like Maven, ensuring a cleaner repository state without affecting the core library functionality.
spark-avro-dataframes · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 58.
Lenses
- Code Health 100
- Architecture 79
- Maturity 51
- Readiness 70
- Security 53
Changes since last survey
- 300 commits — 247 feature/other, 53 fixes
By area
- (root) — 118 commits
- (repo) — 77 commits
- src/main — 68 commits
- src/test — 18 commits
- .github/workflows — 13 commits
- documentation/confluent-avro-documentation.md — 4 commits
- documentation/python-documentation.md — 2 commits
Notable commits
- fix: #170: Fix docs for schema generation
- fix: #325 - bug fix - update jacoco version (#327)
- fix: Abris #123 fix compiler warnings
- fix: Abris #135 Fix unit tests that did not work with scala 2.11.
- fix: Abris #246 fix profile for Spark 2.3
- fix: Abris #256 fix deprecation for Scala 2.13 in non test code
- fix: Abris #256 fix deprecation for Scala 2.13 in tests
- fix: Abris #256 fix deprecation for apache.commons.io and scalatest
- fix: Feature/350 fix tests (#351)
- fix: Feature/fix scalastyle (#292)
- fix: Fix Capitalization of names in READ.ME
- fix: Fix README
- fix: Fix bug when accessing schema registry from inside of executor
- fix: Fix bugs with bare and multiply nested bugs #56 #58
- fix: Fix confluent schema evolution functionality
- fix: Fix of README.md
- fix: Fix scm url
- fix: Fix security settings example (#91)
- fix: Fix subject string in debug messages (#177)
- fix: Merge branch 'master' into fix/public
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
AbsaOSS/ABRiS was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0d9ffb020eb4e76725a559dd1543971ed089dd8d — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.