Skip to content
CAI
Software that uses CAICheck a score

spotify/magnolify

58.8

Adequate · 20 September 2026

7.1k

lines of production code

Scala

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

Magnolify is a Scala library that provides automatic typeclass derivation for serializing and deserializing data across various storage and processing formats, including Avro, BigQuery, Bigtable, Datastore, Parquet, Protobuf, TensorFlow, and Apache Beam. It leverages the Magnolia macro library to generate converters for case classes and sealed traits, supporting both standard and unsafe primitive conversions as well as integration with Cats, Guava, and ScalaCheck. The system also includes tools for schema parsing and migration rules to facilitate upgrades, ensuring compatibility with multiple Scala versions and Java runtimes.

How it got here

2019 — multi-module feature expansion and Scala 3 support

13 changes.

This period focused on expanding the magnolify library with automatic derivation features for various ecosystem integrations, including Cats, Guava, ScalaCheck, and TensorFlow. Significant work was also dedicated to enhancing BigQuery and Datastore modules with precise serialization and key handling, while simultaneously upgrading the build infrastructure to support Scala 3 and modern Java versions.

2020–2021 — multi-format serialization and unsafe type support

16 changes.

This period focused on expanding Magnolify's serialization capabilities across multiple storage backends, including initial support for Protobuf, Bigtable, and Parquet, alongside enhancements to Avro and BigQuery modules. Significant effort was dedicated to implementing 'unsafe' type conversions for primitive narrowing and enum handling across all formats to improve performance and compactness. The work also introduced support for Refined types, Joda-Time integration, and new schema parsing tools to facilitate code generation and schema interoperability.

2022 — comprehensive test coverage expansion

14 changes.

This period focused on significantly expanding test coverage across the project's various integration modules, including Neo4j, Avro, BigQuery, Bigtable, Datastore, Guava, Parquet, Scalacheck, and TensorFlow. The work involved adding comprehensive test suites to verify serialization, deserialization, and schema generation for diverse data types, while also introducing new features like Neo4j value type support and Avro schema documentation aliases.

2024 — Scalafix migration and Beam RowType support

6 changes.

This period focused on enhancing Scalafix with a new 0.7 migration rule for MapBasedConverters, including build updates and comprehensive test coverage. Simultaneously, the project expanded Beam integration by introducing RowType support with portable timestamp handling and Joda compatibility, while resolving implicit conflicts in Guava funnels.

Features

Add Parquet logical type support for Java 8 time and Joda-Time types

The Parquet module now includes explicit support for mapping Java 8 time types (such as \Instant\, \LocalDateTime\, \LocalTime\, and \OffsetTime\) and Joda-Time equivalents to Parquet logical types. This is implemented via new \TimeTypes.scala\ and \package.scala\ files in the \magnolify.parquet.logical\ package, which define implicit \ParquetField\ instances for milliseconds, microseconds, and nanos precision. The changes also introduce deprecated aliases (e.g., \pfTimestampMillis\) to maintain backward compatibility while encouraging the use of the new, more specific type mappings.

parquet/src/main/scala/magnolify/parquet/logical · high confidence

Add Refined type support for Magnolify serialization formats

Users can now serialize and deserialize \eu.timepit.refined\ types across multiple Magnolify-backed storage formats. This change introduces implicit conversions for Guava Funnel, Avro, BigQuery, Bigtable, Datastore, Protobuf, and TensorFlow, allowing refined values to be handled seamlessly in these ecosystems.

refined/src/main · high confidence

Add Scala 2.12 compatibility shims and Scala 3 macro support for EnumType derivation

The shared module now includes Scala 2.12-specific shims to provide serializable \CanBuildFrom\ instances for collections like \List\, \Vector\, and \Set\, which are not serializable by default in that version. Additionally, the library introduces Scala 3 macro implementations for \EnumType\ derivation, enabling automatic conversion for Scala enumerations and ADTs using the new \scala.deriving.Mirror\ API, while maintaining parallel Scala 2 macro implementations for backward compatibility.

shared · high confidence

Add automatic Guava Funnel derivation for case classes and sealed traits

Users can now automatically derive Guava \Funnel\ instances for their types using the new \FunnelDerivation\ object in the \magnolify.guava.semiauto\ package. This feature leverages Magnolia to recursively funnel case class fields (including support for value classes and distinguishing empty vs. populated collections via index/size injection) and sealed trait subtypes. The package also provides built-in implicit funnels for primitive types (Int, Long, Byte, Char, Short, Boolean), arrays, and CharSequences, enabling seamless integration with Guava's hashing utilities without manual implementation.

guava/src/main/scala/magnolify/guava/semiauto · high confidence

Add automatic derivation for ScalaCheck Arbitrary and Cogen typeclasses

Users can now automatically derive \Arbitrary\ and \Cogen\ instances for case classes and sealed traits using the new \magnolify.scalacheck.auto\ package and \semiauto\ derivation objects. This change introduces macro-based generation that leverages the Magnolia library to recursively derive generators and hashers, including a specific workaround to prevent invalid \Seed\ generation and support for \ReadOnlyCaseClass\ in Cogen derivation.

scalacheck/src/main · high confidence

Add support for Joda-Time types and BigQuery logical types in Avro serialization

The Avro module now supports serializing and deserializing Joda-Time types (such as joda.Instant, joda.LocalTime, joda.LocalDateTime, and joda.DateTime) alongside the standard Java 8 java.time types for timestamps and times in both milliseconds and microseconds. It also introduces BigQuery-specific logical type support, including a custom 'datetime' logical type for parsing and a numeric type mapping for BigDecimal, ensuring compatibility with BigQuery schemas.

avro/src/main/scala/magnolify/avro/logical · high confidence

Added scalafix 0.7 migration rule and updated build infrastructure

This change introduces a new migration rule for scalafix version 0.7, specifically targeting the \MapBasedConverters\ rule which updates how \ExampleType\ instances are invoked (changing from method calls like \etPerson.from(e)\ to direct application \etPerson(e)\). Additionally, the scalafix project's build infrastructure has been updated to use sbt version 1.10.7, and the project now includes specific JVM memory settings in \.sbtopts\ to support the build process.

scalafix · high confidence

Auto-derivation of Guava Funnels via macro

Users can now automatically derive Guava Funnels for case classes without manual implementation by importing the new \magnolify.guava.auto\ package. This change introduces a \LowPriorityImplicits\ trait containing a macro-based \genFunnel\ method, which resolves implicit conflicts for top-level collection funnels and reduces compiler warnings by providing a fallback derivation mechanism.

guava/src/main/scala/magnolify/guava/auto · high confidence

Derive Cats typeclasses for case classes and sealed traits

The \cats\ module now provides automatic derivation for Cats typeclasses (including Show, Eq, Hash, Semigroup, Monoid, Group, and their commutative variants) for case classes and sealed traits. Users can import \magnolify.cats.auto.\_\ to get implicit instances via macro-based derivation, or use \magnolify.cats.semiauto\ for explicit derivation. This enables seamless integration with Cats' typeclass system without manual implementation.

cats/src/main · high confidence

Expose Avro schema documentation type alias

The Avro module now exposes a \doc\ type alias (importable via \magnolify.avro.\_\) that delegates to the shared documentation type. This allows users to attach documentation metadata to Avro schema definitions using a consistent, reusable type reference.

avro/src/main/scala/magnolify/avro · high confidence

Initial project setup and configuration for magnolify

This change establishes the foundational configuration for the magnolify library, introducing a comprehensive scalafmt configuration (version 3.10.7) to standardize code formatting across Scala 2.13 and Scala 3 sources. It adds a Scala Steward configuration file to automate dependency updates while explicitly excluding specific groups like Apache Beam and pinning critical versions for Neo4j, Protobuf, and Scala 3 LTS. The project also includes JVM optimization options, a .gitignore update to exclude IDE and build artifacts, and documentation files (README, NOTICE, catalog-info.yaml) to define the library's modules (Avro, BigQuery, Bigtable, Cats, etc.) and release process.

(repo-wide) · high confidence

Initial support for automatic Protobuf serialization and deserialization

Introduces the ProtobufType and ProtobufField abstractions, enabling automatic derivation of converters between Scala case classes and Protocol Buffers messages. This change allows users to seamlessly serialize and deserialize data structures to/from Protobuf format without writing manual mapping code, leveraging the magnolia library for typeclass derivation.

protobuf/src/main/scala/magnolify/protobuf · high confidence

Introduce Beam RowType support with portable Timestamp logical types and Joda compatibility

This change adds a new \RowType\ implementation for Apache Beam \Row\ values, enabling automatic schema generation and conversion for case classes. It introduces a structured approach to temporal types via the \magnolify.beam.logical\ package, mapping \java.time.Instant\ to Beam's portable \Timestamp\ logical type at millisecond, microsecond, and nanosecond precisions (with truncation on write for finer precisions). It also provides \compat\ objects that preserve the legacy \DATETIME\/\INT64\/\NanosInstant\ encodings for backward compatibility and support for older Beam IOs, while adding explicit support for Joda Time types alongside the standard Java 8 time API.

beam · high confidence

Introduce BigNumeric support and configurable Parquet array encoding

Users can now store and process high-precision decimal values in BigQuery via the new BigNumeric type, which validates precision and scale constraints. Additionally, Parquet writers now support configurable array encoding (Ungrouped, ThreeLevelArray, or ThreeLevelList) via the MagnolifyParquetProperties configuration, allowing users to control schema structure for compatibility with downstream systems.

repository · high confidence

Introduce BigtableType and BigtableField for automatic serialization

Adds the BigtableType and BigtableField abstractions in the bigtable module, enabling automatic derivation of typeclasses for converting Scala case classes and primitives to and from Google Cloud Bigtable mutations and rows. This includes a new ByteStringComparator for Java-side byte ordering and support for value classes, records, and primitive types via Magnolia-based macro derivation.

bigtable/src/main · high confidence

Introduce automatic schema derivation for TensorFlow Example serialization

The library now automatically derives the TensorFlow Schema alongside data conversion for case classes via the new \ExampleType\ and \ExampleField\ abstractions. Users can serialize and deserialize TensorFlow \Example\ protos with built-in schema generation that respects field names, types, and \@doc\ annotations, eliminating the need for manual schema construction.

tensorflow/src/main/scala/magnolify/tensorflow · high confidence

Introduce typed Parquet field derivation and filter predicate support

The Parquet module now includes a new \ParquetField\ trait and associated implementation files (\ParquetField.scala\, \Predicate.scala\, \Schema.scala\, \SchemaUtil.scala\, \SerializationUtils.scala\, \TypeConverter.scala\) that enable automatic derivation of Parquet schemas from Scala types and support for generating \FilterPredicate\s on primitive columns. This change adds the infrastructure for typed Parquet annotations and filter predicate generation, allowing users to define custom filters on Parquet fields while ensuring schema compatibility and proper handling of repeated types and schema evolution.

parquet/src/main/scala/magnolify/parquet · high confidence

New schema parsing tools for Avro, Parquet, and BigQuery

Added new code-generation utilities in the \tools\ module that parse schema definitions from Avro, Parquet, and BigQuery formats into a unified internal representation. This enables users to convert existing schemas from these data formats into Magnolify-compatible types, supporting logical types such as UUIDs, decimals, dates, and timestamps across all three formats.

tools/src/main · high confidence

Support for Char and UnsafeEnum types in Parquet serialization

The \magnolify.parquet.unsafe\ package now provides implicit \ParquetField\ instances for \Char\ and \UnsafeEnum\[T\]\ types. \Char\ values are serialized as integers, while \UnsafeEnum\ values are serialized as strings using the provided \EnumType\ for conversion. This enables users to handle these specific data types when working with Parquet files in an unsafe context.

parquet/src/main/scala/magnolify/parquet/unsafe · high confidence

Support for Datastore entity keys and excludeFromIndexes annotation

The Datastore module now supports mapping entity keys and controlling index inclusion via annotations. Users can annotate case class fields with @key to define the entity's key structure (including project, namespace, and kind) and with @excludeFromIndexes to mark specific fields as non-indexed in Google Cloud Datastore. This is implemented through new KeyField and EntityField typeclasses in EntityType.scala, alongside a TimestampConverter for handling java.time.Instant conversions.

datastore/src/main/scala/magnolify/datastore · high confidence

Support for Neo4j Value types and value classes

Added a new \ValueType\ and \ValueField\ typeclass system in the Neo4j module to handle conversion between Scala types and Neo4j driver \Value\ objects. This includes support for primitive types, dates, durations, points, options, iterables, and maps. It also introduces specific handling for Scala value classes, ensuring they are treated as their underlying types during serialization and deserialization, and adds unsafe enum support via the \unsafe\ package.

neo4j/src/main · high confidence

Unsafe BigQuery type conversions and enum support

The \magnolify.bigquery.unsafe\ package now provides implicit \TableRowField\ instances for narrowing primitive conversions (Byte, Char, Short, Int from Long; Float from Double) and for handling Enums. Users can now serialize/deserialize these types via BigQuery's TableRow format, with enum support relying on the \EnumType\ typeclass for mapping between string representations and Scala types.

bigquery/src/main/scala/magnolify/bigquery/unsafe · high confidence

Unsafe numeric and enum conversions in Datastore

The \magnolify.datastore.unsafe\ package now provides implicit \EntityField\ instances for \Byte\, \Char\, \Short\, \Int\, and \Float\ that perform narrowing conversions from \Long\ and \Double\ respectively, as well as support for \Enum\ and \UnsafeEnum\ types via string serialization. These conversions are marked as unsafe because they may lose precision or throw exceptions if values fall outside the target type's range, allowing users to opt into more compact storage at the cost of safety.

datastore/src/main/scala/magnolify/datastore/unsafe · high confidence

Unsafe package introduces ProtobufField support for primitive types and UnsafeEnum

The new \magnolify.protobuf.unsafe\ package provides implicit \ProtobufField\ instances for primitive types (\Byte\, \Char\, \Short\) by mapping them to \Int\, and adds support for \UnsafeEnum\[T\]\ via a string-based conversion that handles empty strings gracefully. This allows users to serialize and deserialize these specific types within the unsafe protobuf context without manual conversion logic.

protobuf/src/main/scala/magnolify/protobuf/unsafe · high confidence

Unsafe primitive and enum type mappings for TensorFlow ExampleField

The new \magnolify.tensorflow.unsafe\ package provides implicit \ExampleField.Primitive\ instances for standard Scala types (Byte, Char, Short, Int, Double, Boolean, String) and enums. These mappings allow users to serialize and deserialize these types directly into TensorFlow's \ExampleField\ format using unsafe conversions (e.g., treating Longs as Bytes or Diffs as Floats), which is useful for performance-critical paths where strict type safety is traded for efficiency. It also includes support for \UnsafeEnum\ types, enabling enum serialization via string representation in the protobuf bytes.

tensorflow/src/main/scala/magnolify/tensorflow/unsafe · high confidence

Unsafe type conversions for primitive types and enums in Avro

The new \magnolify.avro.unsafe\ package provides implicit \AvroField\ instances for \Byte\, \Char\, and \Short\ by mapping them to and from \Int\, and introduces \UnsafeEnum\ support for Avro enums. This allows users to handle these specific primitive types and enum variants in Avro schemas, though the mappings are considered unsafe due to potential data loss or validation gaps.

avro/src/main/scala/magnolify/avro/unsafe · high confidence

Behavioural changes

Add migration rule for MapBasedConverters in Scalafix 0.7

A new Scalafix rule named MapBasedConverters has been added to assist with migrating code from version 0.7. This rule automatically refactors calls to magnolify's ExampleType converter functions (from/to), replacing the method-call syntax with a direct constructor-style syntax. The rule is registered via the standard Scala services mechanism and targets specific magnolify symbols to ensure safe transformation.

scalafix/rules · high confidence

Improved BigQuery timestamp and date serialization precision

The BigQuery integration now uses custom formatters in TimestampConverter to strictly adhere to BigQuery's SQL data-type specifications for Time, Date, and Timestamp fields. This ensures that fractional seconds (nanoseconds) are correctly truncated to microseconds and formatted with a consistent number of digits, preventing precision loss or parsing errors when reading and writing temporal data.

bigquery/src/main/scala/magnolify/bigquery · high confidence

Redirect for Scaladoc documentation

The site now includes a template that automatically redirects users from the old Scaladoc location to the new Magnolify API documentation page, ensuring visitors are directed to the correct reference material.

site · high confidence

Fixes

Fix implicit conflict for top-level collection funnels

Resolved an implicit resolution conflict that occurred when deriving funnels for top-level collections. The change introduces a macro in GuavaMacros.scala that delegates funnel generation to the semiauto.FunnelDerivation object, ensuring that collection types are handled correctly without ambiguity during compilation.

guava/src/main/scala/magnolify/guava · high confidence

Test coverage

Added Avro schema comparison utility for tests; Added Avro type serialization tests and BigQuery data generation utilities; Added BigQuery data type testing script; Added JMH benchmarks for Magnolify serialization and type conversions; Added comprehensive test suites for Parquet module; Added test coverage for scalacheck Arbitrary and Cogen derivation; Added test suite for EnumType and UnsafeEnum support; Added test suite for scalafix 0.7 migration rule; Added test suites for Avro, BigQuery, Parquet, and SchemaPrinter tools; Added test suites for cats typeclass derivation; Added tests for BigQuery TableRowType and TimestampConverter; Added tests for Bigtable type serialization and field mapping; Added tests for Datastore EntityType serialization and key handling; Added tests for ExampleType serialization and schema generation; Added tests for Guava Funnel derivation and scope resolution; Added tests for Neo4j value type serialization; Added tests for ProtobufType serialization and Bug1001 builder reuse fix; Added tests for Refined type integration across storage backends.

Dependencies

Upgrade to Scala 2.13.17, 2.12.20, and 3.3.7 with Java 17/25 support

The build system has been updated to support Scala versions 2.12.20, 2.13.17, and 3.3.7, with the default Scala version set to 2.13.17. The project now targets Java 17 and Java 25 (Corretto), dropping support for Java 8. This change also introduces a new \scalafix\ sub-project to manage migration rules for upgrading from magnolify versions 0.6.0 and 0.7.0, including test infrastructure to verify these migrations.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 59.

Lenses

  • Code Health 97
  • Architecture 100
  • Maturity 36
  • Readiness 93
  • Security 64

Changes since last survey

  • 300 commits — 291 feature/other, 9 fixes

By area

  • (root) — 192 commits
  • project/plugins.sbt — 34 commits
  • project/build.properties — 15 commits
  • .github/workflows — 14 commits
  • scalafix/build.sbt — 9 commits
  • parquet/src — 7 commits
  • (repo) — 5 commits
  • avro/src — 3 commits
  • beam/src — 3 commits
  • bigquery/src — 3 commits
  • protobuf/src — 3 commits
  • cats/src — 1 commit
  • docs/cats.md — 1 commit
  • docs/mapping.md — 1 commit
  • docs/parquet.md — 1 commit
  • guava/src — 1 commit
  • jmh/src — 1 commit
  • scalafix/project — 1 commit
  • scalafix/rules — 1 commit
  • shared/src — 1 commit

Notable commits

  • fix: (fix #1153) Support xmap in magnolify-bigquery (#1217)
  • fix: (fix #766) Deprecate AvroCompat, replace automatic schema detection on read + Configurable write (#996)
  • fix: Fix actions-gh-pages usage (#882)
  • fix: Fix mdoc wiring in site (#886)
  • fix: Fix proto builder re-use bug (#1236)
  • fix: Fix workflow name (#1249)
  • fix: Fix workflow name v2 (#1250)
  • fix: [avro] fix: copy bytes when reading from avro (#937)
  • fix: [guava] fix: resolve implicit conflict for top level collection funnels (#936)
  • change: (chore) Update to codecov-action (#960)
  • change: (magnolify-beam) Map Instant to Beam's portable Timestamp logical type (#1392)
  • change: (magnolify-beam) RowType#from should look up field values by name, not index (#1350)
  • change: 2.13.17 etc (#1237)
  • change: Add BigNumeric (#1235)
  • change: Add benchmarks for magnolify-parquet vs parquet-avro R/W (#1040)
  • change: Add joda support (#1234)
  • change: Add manual GHA site publish (#880)
  • change: Add removed avro/parquet logical fields for compatibility (#1285)
  • change: Add scalafix 0.7 migration rule (#911)
  • change: Add support for Beam schemas (#1027)
  • …and 280 more

Architecture

  • 0 containers · 2 bounded contexts · 1 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

spotify/magnolify was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit a90c8ea0b313b7f6b635b6dfc9682d9e7494c255 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.