Skip to content
CAI
Software that uses CAICheck a score

sksamuel/avro4s

55.0

Adequate · 28 September 2026

4.3k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Avro4s is a Scala library that automatically generates Avro schemas and handles the encoding and decoding of Scala types to and from Avro data formats. It supports a wide range of Scala constructs, including collections, temporal types, sealed traits, and custom enums, while providing utilities for schema merging and field naming strategies. The system also includes modules for integrating with Apache Kafka and the Cats library, enabling seamless serialization for distributed data pipelines.

How it got here

2015–2016 — Scala 3 migration and streaming API

7 changes.

The project modernized its build infrastructure by migrating to SBT 1.12 and Scala 3.3, replacing Shapeless with Magnolia for serialization. This period introduced a new streaming API for Avro I/O, added field naming strategies, and expanded governance documentation alongside comprehensive test coverage.

2018 — Comprehensive test coverage and Kafka integration

8 changes.

This period focused on significantly expanding test coverage for the core Avro4s library, adding extensive validation for record encoding, decoding, schema generation, and stream I/O operations. It also introduced a new GenericSerde feature for the Kafka module, enabling Avro record serialization and deserialization without requiring a Schema Registry.

2019–2021 — Scala 3 support and schema derivation rewrite

9 changes.

The project rewrote its core schema derivation logic to support Scala 3 using Magnolia V1, replacing the previous Magnolia V0 implementation. This period also involved expanding encoder and decoder coverage for complex Scala types, adding utility helpers for schema management, and integrating support for Cats NonEmpty collections.

Features

Added GenericSerde for Kafka without Schema Registry

Users can now serialize and deserialize Avro records to and from Kafka topics using the new GenericSerde class, which operates independently of the Confluent Schema Registry. This component supports Binary, JSON, and Data Avro formats and correctly handles null values (tombstones) during serialization and deserialization.

avro4s-kafka · high confidence

Added support for Cats NonEmpty collections in Avro encoding and decoding

The avro4s-cats module now provides built-in support for encoding and decoding Cats library types NonEmptyList, NonEmptyVector, and NonEmptyChain. These types are serialized as Avro arrays, allowing users to seamlessly integrate non-empty collection structures into their Avro schemas without manual conversion logic.

avro4s-cats · high confidence

New Avro I/O streams and field naming strategies

The library introduces a new streaming API for Avro serialization and deserialization, providing \AvroDataOutputStream\ (which embeds the schema, suitable for file storage), \AvroBinaryOutputStream\, and \AvroJsonInputStream\ to support different data formats and use cases like Kafka. It also adds a \FieldMapper\ trait with built-in strategies (\SnakeCase\, \PascalCase\, \LispCase\) to control how Scala field names are mapped to Avro schema field names, and introduces \ImmutableRecord\ as a lightweight implementation of Avro's \GenericRecord\ and \SpecificRecord\ interfaces.

repository · high confidence

New binary and data file input/output streams with Java enum annotations

The library now provides explicit \AvroBinaryInputStream\ and \AvroDataInputStream\ classes to handle reading Avro data, distinguishing between binary streams (which require a provided writer schema) and data files (which embed the schema). Corresponding \AvroBinaryOutputStream\ handles binary serialization. Additionally, new Java annotations (\@AvroJavaName\, \@AvroJavaNamespace\, \@AvroJavaProp\, \@AvroJavaEnumDefault\) are introduced to allow configuring Avro schema generation for Java enums and types directly from Java code.

avro4s-core/src/main/scala/com/sksamuel/avro4s · high confidence

New decoders for BigDecimal, collections, Either, and temporal types

The library now supports decoding Scala types that were previously missing or incomplete. Users can decode BigDecimal values stored as bytes, strings, or fixed-length fields; standard Scala collections (List, Seq, Set, Vector, Array) and Maps; Either types using type guards to distinguish left and right values; and temporal types including Instant, LocalDateTime, LocalDate, LocalTime, Date, Timestamp, and OffsetDateTime with support for Avro logical types like timestamp-millis and timestamp-micros. Additionally, decoders for primitive types (Byte, Short, Int, Long, Float, Double, Boolean), strings (String, Utf8, CharSequence, UUID), and byte arrays (Array\[Byte\], ByteBuffer) have been added or refined to handle various underlying Avro schema representations.

avro4s-core/src/main/scala/com/sksamuel/avro4s/decoders · high confidence

New encoder implementations for Scala types

The encoder module now includes dedicated encoders for BigDecimal (supporting bytes, fixed, and string schemas), UUIDs, temporal types (Instant, LocalDate, LocalTime, LocalDateTime, OffsetDateTime, and java.sql.Date/Timestamp), and various collection types (Array, List, Seq, Set, Vector, Map). It also adds encoders for primitive types, byte arrays/ByteBuffers, Option types, Either types, Tuples up to arity 5, and sealed traits (both as enums and unions).

avro4s-core/src/main/scala/com/sksamuel/avro4s/encoders · high confidence

New schema utility helpers for merging, union handling, and ByteBuffer support

The avroutils package introduces several new helper objects to improve schema generation and data handling. AvroSchemaMerge allows combining multiple Avro record schemas into a single schema, correctly handling field defaults and union type ordering. ByteBufferHelper provides a safe way to convert ByteBuffers to byte arrays, supporting cases where the buffer does not expose its underlying array directly. EnumHelpers adds support for setting default values on enum schemas. SchemaHelper expands union management capabilities by flattening nested unions, moving default or null types to the head of unions as required by the Avro spec, and extracting specific subschemas for Either types. It also includes utilities for mapping field names and marking schemas as error records.

avro4s-core/src/main/scala/com/sksamuel/avro4s/avroutils · high confidence

Behavioural changes

Migrate benchmarks from ScalaMeter to JMH

The benchmarking infrastructure in the benchmarks module has been migrated from ScalaMeter to JMH. This change introduces new Java-based Avro generated record classes (such as Empty, Invalid, and ValidInt) to serve as benchmark data models, alongside new Scala benchmark implementations (Decoding, Encoding) that utilize JMH annotations for warmup, measurement, and JVM configuration. The migration replaces the previous ScalaMeter-based approach with a standardized JMH setup for measuring encoding and decoding throughput.

benchmarks · high confidence

New schema derivation implementation for Scala 3

The schema derivation logic in the core library has been rewritten to support Scala 3, replacing the previous Magnolia V0 implementation with Magnolia V1. This change introduces new schema handlers for collections (Array, Seq, Set, Vector, List, Map), Either, tuples (up to Tuple6), and value types, while also adding support for Java enums with defaults, logical types like timestamp-nanos and datetime-with-offset, and the 'avro.java.string' property for strings.

avro4s-core/src/main/scala/com/sksamuel/avro4s/schemas · high confidence

Revised subtype ordering and annotation handling for sealed traits

The library now determines the order of sealed trait subtypes using a priority system driven by @AvroSortPriority and @AvroUnionPosition annotations, rather than relying on alphabetical symbol names. This change ensures that union positions and enum symbol orders match the explicit priorities defined in the schema or the original 4.x behavior, providing more predictable serialization results for complex type hierarchies.

avro4s-core/src/main/scala/com/sksamuel/avro4s/typeutils · high confidence

Test coverage

Added comprehensive encoder tests for Avro4s record encoding; Added regression tests for GitHub issues in avro4s-core; Added test coverage for Avro input streams; Added test for class in uppercase package; Added test resources for Avro schema generation; Added tests for Avro output stream codecs and data serialization; Added tests for README examples and sealed trait enumeration mapping; Added tests for record encoding, decoding, and type-guarded decoders; Added tests for recursive data structures; Expanded schema generation tests for collections, annotations, and types; Expanded test coverage for record decoders.

Dependencies

Build infrastructure modernized to SBT 1.12 and Scala 3.3

The project's build system has been updated to use SBT version 1.12.11, with the default Scala version set to 3.3.7. This change introduces a new centralized build configuration in \project/Build.scala\ that defines and manages dependency versions for core libraries including Avro (1.11.5), ScalaTest (3.2.17), and Magnolia (1.3.20). Additionally, the build plugins have been updated to sbt-pgp 2.3.1 and sbt-jmh 0.4.8, and publishing credentials are now configured for the new central.sonatype.com repository.

project · high confidence

Project structure modernized with Magnolia and JMH benchmarks

The build configuration has been restructured to define the project as a multi-module aggregate (core, cats, kafka) with explicit publishing controls. The core module now relies on Magnolia for serialization instead of Shapeless or json4s, and the kafka module explicitly depends on kafka-clients version 2.4.0. Additionally, a new benchmarks module has been added, migrating performance testing from ScalaMeter to JMH.

(dependencies) · high confidence

Housekeeping

Project governance and documentation overhaul

The repository has been updated with a comprehensive Contributor Covenant Code of Conduct to establish community standards, and the project license has been switched from MIT to Apache 2.0. Additionally, the README has been significantly expanded to include detailed documentation on schema generation, field naming overrides, and versioning strategies for Scala 2 and 3, while benchmark results and scripts have been added to track performance.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 54 → 55 (+0.7)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 94 → 94 (+0.0)
  • Architecture 100 → 100 (+0.0)
  • Maturity 51 → 51 (+0.0)
  • Readiness 50 → 53 (+3.6)
  • Security 47 → 47 (+0.0)

Resolved (3)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • Off-boarding risk: anonymized user #1

New (1)

  • Outdated: org.apache.kafka:kafka-clients

Architecture

  • Unchanged — 0 containers · 1 contexts · 0 edges

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

sksamuel/avro4s was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit a1a9785389905621aab2115f7595a6929310692d — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.