Skip to content
CAI
Software that uses CAICheck a score

airbnb/chronon

54.3

Adequate · 28 September 2026

33.7k

lines of production code

Scala

with Python

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Chronon is a feature engineering and serving platform that defines, computes, and serves machine learning features through GroupBy and Join abstractions. It supports batch and streaming execution via Spark and Flink, providing online serving capabilities with configurable backends like MongoDB and REST APIs. The system includes comprehensive tooling for orchestration, lineage tracking, and statistical monitoring to manage feature pipelines end-to-end.

How it got here

2022 — Build system modernization and API expansion

26 changes.

This period focused on modernizing the build infrastructure by introducing Bazel alongside SBT and upgrading Scala and Spark compatibility. It significantly expanded the Python and Java APIs with new lineage tracking, model integration, and online serving capabilities, supported by comprehensive test coverage and dependency updates.

2023 — Orchestration and streaming expansion

12 changes.

This period focused on expanding Chronon's execution capabilities by introducing initial Airflow integration for job orchestration and a new Flink-based streaming engine for GroupBy jobs. The team also enhanced online serving options with a MongoDB implementation and improved multi-version Scala compatibility, while significantly broadening test coverage across API, online, and sample components.

2024–2025 — Feature serving and build system expansion

7 changes.

This period focused on extending the platform's capabilities by introducing a new standalone HTTP-based feature fetching service for external applications. Concurrently, the project expanded its build infrastructure by adding Bazel support for tools and IntelliJ integration, while significantly increasing test coverage for lineage parsing and group-by logic.

Features

Added Bazel project configuration for IntelliJ

A new \.bazelproject\ file has been added to the \.ijwb\ directory to enable Bazel-based builds within IntelliJ IDEA. This configuration automatically derives targets from the workspace directories and explicitly enables support for Python, Scala, and Java languages, allowing users to leverage Bazel alongside sbt for their development workflow.

.ijwb · high confidence

Initial Airflow integration for Chronon orchestration

This change introduces the initial set of Airflow components required to orchestrate Chronon jobs, including DAG constructors for GroupBy backfills/uploads, Join backfills, Staging Queries, and Online/Offline consistency metrics. It provides the core Python operators (ChrononOperator) to download and execute Chronon packages, helper utilities for task dependency management and scheduling logic, and configuration constants for concurrency and upload cadence.

airflow · high confidence

Initial Bazel build system support for tools

This change introduces Bazel as a build system for the \tools\ directory, adding \BUILD\ files and Starlark rules in \tools/build\_rules\ to handle JVM/Scala compilation, Maven artifact resolution, Python testing via pytest, and Thrift code generation. It also adds configuration settings in \tools/flags\ to enable conditional compilation for specific versions of Flink (1.16, 1.19, 2.0) and Spark (2.4, 3.1, 3.2, 3.5).

tools · high confidence

Adds a new Flink module that provides a streaming implementation for Chronon GroupBy jobs, supporting both untiled (raw event) and tiled (pre-aggregated) processing modes. The implementation includes operators for Spark expression evaluation, Avro serialization, and async writes to the KV store, along with custom windowing triggers and watermark strategies to handle event-time processing and late data.

flink/src · high confidence

Introduces Two-Stack Lite algorithm and enhances online aggregation with tiled processing and mutation fixes

The windowing module now supports the Two-Stack Lite algorithm via new \TwoStackLiteAggregationBuffer\ and \TwoStackLiteAggregator\ classes, providing an alternative method for maintaining sliding window aggregations. The \SawtoothOnlineAggregator\ gains a tiled processing path (\lambdaAggregateIrTiled\) that accepts pre-aggregated \TiledIr\ objects, enabling more efficient batch merging. Additionally, the online aggregation logic now correctly filters streaming events based on \mutationTs\ (for mutations) or \row.ts\ (for events) relative to a \batchResponseMaxTs\, preventing over-counting from stale data. The \SawtoothMutationAggregator\ also adopts a bitset to prevent duplicate updates to tail hops and uses bulk merging to avoid intermediate mutations. Finally, \TsUtils\ switches from \SimpleDateFormat\ to \FastDateFormat\ for thread-safe date formatting, and \HopsAggregator\ replaces \println\ with SLF4J logging.

aggregator/src/main/scala/ai/chronon/aggregator/windowing · high confidence

Introduction of Bazel build system alongside SBT

Chronon now supports building with Bazel in addition to the existing SBT build system. This change introduces Bazel configuration files (including WORKSPACE, BUILD.bazel, .bazelrc, and .bazeliskrc) that define build rules for Scala, Java, and Python dependencies, allowing users to compile the project using the Bazel toolchain.

(repo-wide) · high confidence

Java API overhaul: structured responses, feature flags, and builder pattern

The Java online client has been significantly refactored to improve type safety and usability. Fetch responses now use a new \JTry\ wrapper instead of raw maps, allowing Java consumers to handle success and failure states explicitly without manual exception parsing. A new \JavaFeatureRecord\ class provides a strongly-typed, name-based view of feature data (supporting strings, numbers, booleans, structs, and lists) and surfaces errors separately. The \JavaFetcher\ now uses a builder pattern, enabling configuration of optional components like a \FlagStore\ for runtime feature toggles, custom execution contexts, and error-throwing behavior. Additionally, new classes (\JavaStatsRequest\, \JavaStatsResponse\, \JavaSeriesStatsResponse\, \JavaStructuredResponse\) expose stats and structured data fetching capabilities, while \TBaseDecoderFactory\ and \ThriftDecoder\ add support for decoding complex Thrift-based structured data types.

online/src/main/java · high confidence

MongoDB-backed online serving implementation for the quickstart

The quickstart now includes a complete MongoDB-based online serving implementation, allowing users to serve feature lookups and log responses using MongoDB instead of the default in-memory or other stores. This adds a new \mongo-online-impl\ module containing a \ChrononMongoOnlineImpl\ API adapter, a \MongoKvStore\ for key-value operations, and utilities like \Spark2MongoLoader\ and \MongoLoggingDumper\ to sync batch data and logs into MongoDB. The setup is supported by a new Dockerfile that installs Spark, Hadoop, Thrift, and Scala, along with configuration files for the quickstart environment.

quickstart · high confidence

New Chronon Feature Fetching Service for external feature lookups

A new standalone feature service module has been added to allow non-JVM applications to retrieve features via HTTP. Built on the Vert.x framework, this service exposes REST endpoints at \/v1/features/groupby/:name\ and \/v1/features/join/:name\ to support bulk feature lookups. It dynamically loads a concrete \Api\ implementation (such as a MongoDB connector) from a specified JAR file at startup, enabling users to quickly deploy a feature serving layer without managing the full JVM stack.

service · high confidence

New Python APIs for lineage tracking, model transforms, and PySpark notebook execution

This change introduces several new capabilities in the Python API. First, it adds a Lineage Parser that extracts column-level lineage metadata from Chronon configurations, enabling impact analysis and data quality tracing. Second, it provides a new Model API (model.py) with helper functions to define InferenceSpec, Model, and ModelTransform objects for integrating ML models into feature pipelines. Third, it introduces a PySpark Interface that allows users to execute GroupBy, Join, and StagingQuery operations directly within notebook environments like Databricks and Jupyter, facilitating rapid prototyping and iteration.

api/py/ai · high confidence

New approximate frequency and histogram aggregators; base classes refactored to traits

The aggregator base library now includes new \FrequentItems\ and \ApproxHistogram\ aggregators, enabling users to track the most frequent values in a stream and maintain approximate histograms with configurable error tolerance. Additionally, \BaseAggregator\ and \SimpleAggregator\ have been converted from abstract classes to traits, and a \bulkMerge\ method has been added to \BaseAggregator\ to allow more efficient merging of intermediate results.

aggregator/src/main/scala/ai/chronon/aggregator/base · high confidence

New online serving components and serialization safety checks

This change introduces several new components to the online serving layer: an \ApiSerializationCheck\ utility that validates API objects survive Spark driver-to-executor serialization (preventing silent null-pointer failures), a \CatalystUtil\ with pooled Spark sessions for efficient SQL derivation execution, and a \FetcherCache\ trait using Caffeine to cache batch IRs and reduce KV store latency. It also adds an \ExternalSourceRegistry\ for dynamic registration of external feature handlers, a \DataStreamBuilder\ for constructing Spark DataFrames from topics, and \DerivationUtils\ to apply SQL-based derivations during fetching. These additions enhance reliability, performance, and extensibility for online feature serving.

online/src/main/scala · high confidence

Support for element-wise, map, and bucketed aggregations with improved stats and error handling

The row aggregator now supports aggregating over complex input types: element-wise operations on vectors (arrays/lists), map-based aggregations, and bucketed aggregations, allowing users to compute statistics on nested data structures. A new StatsGenerator module introduces drift detection metrics (PSI and L-infinity distance) and configurable percentile calculations for feature monitoring. Additionally, the system now includes a bulk merge capability for performance, sanitizes NaN/Infinity values in aggregations to prevent downstream errors, and provides clearer error messages when constructing aggregators for unsupported column types.

aggregator/src/main/scala/ai/chronon/aggregator/row · high confidence

Behavioural changes

Added Scala 2.11-specific compatibility helpers

New Scala 2.11-specific source files have been added to the online module to support version-specific APIs. This includes a helper for converting between Scala Futures and Java CompletionStages, and a Catalyst helper to evaluate filter predicates using the InterpretedPredicate API, ensuring compatibility with Spark 2.4.0 on Scala 2.11.

online/src/main/scala-2.11 · high confidence

Added Scala 2.12 and 2.13 specific online helper utilities

New version-specific source files have been added for the online module to handle Scala version differences. For Scala 2.12, \FutureConverters\ uses \scala.compat.java8\ for \Future\/\CompletionStage\ conversion, while Scala 2.13 uses \scala.jdk\. Additionally, \ScalaVersionSpecificCatalystHelper\ provides a consistent API for evaluating Spark Catalyst predicates, with the Scala 2.13 version explicitly handling \Seq\ conversion for attributes. These files are now exported via new \BUILD\ targets in both version directories.

online/src/main/scala-2.12, online/src/main/scala-2.13 · high confidence

Configurable logging via log4j properties

The application now uses a log4j configuration file to manage logging behavior. By default, log messages are set to the INFO level and are output to the console (System.out) with a specific timestamp and pattern, allowing users to see structured log entries directly in the application output.

online/src/main/resources · high confidence

Python API build, release, and documentation overhaul

The Python API build process has been migrated from a legacy shell script to Bazel, introducing new BUILD.bazel rules for packaging (py\_wheel) and publishing (twine\_upload) alongside a new python-api-build.sh script that supports build and release actions. The setup.py versioning logic has been updated to be PEP440 compliant, handling branch names and SNAPSHOT suffixes via environment variables. Additionally, the API documentation has been significantly expanded to include user-facing examples for Sources, Group Bys, and Joins, and pre-commit hooks (black, isort, autoflake) have been added to enforce code style.

api/py · high confidence

Refactor Scala Java/Scala collection conversions and expand API model builders

The API layer replaces the deprecated \scala.collection.JavaConverters\ with version-specific \ScalaJavaConversions\ (using \scala.jdk.CollectionConverters\ for Scala 2.12/2.13) to improve compatibility and remove deprecation warnings. This change is accompanied by a significant expansion of the \Builders\ object, which now supports constructing \Join\ configurations with \externalParts\, \labelPart\, \bootstrapParts\, \derivations\, \modelTransforms\, and \skewKeys\, as well as \GroupBy\ with \derivations\. New API types like \ExternalJoinPart\, \ExternalSource\, and \FetchException\ are introduced to support external feature joins and more granular error handling, while \Constants\ and \Extensions\ are updated to reflect new table naming conventions (e.g., \\_bootstrap\, \\_labels\) and metadata fields.

api/src/main · high confidence

Spark 3.5+ encoder compatibility and structured logging support

The Spark engine now includes a version-specific encoder utility for Spark 3.5+ that uses the native Encoders API, ensuring compatibility with newer Spark releases, while the default path continues to use the Catalyst RowEncoder. Additionally, a standard log4j configuration is provided to control log verbosity across the Spark application, and Kryo serialization is extended to support Iceberg and Delta Lake internal classes for improved shuffle performance.

spark/src/main · high confidence

Thrift API schema expansion and build system integration

The Thrift API definition has been significantly expanded to support new data modeling capabilities, including the addition of a \JoinSource\ struct to allow joins to feed into other computations, an \ExternalSource\ struct with factory configuration for dynamic handler creation, and new aggregation operations such as \SKEW\, \KURTOSIS\, \APPROX\_HISTOGRAM\_K\, and \BOUNDED\_UNIQUE\_COUNT\. The \StagingQuery\ struct now supports \createView\ flags and additional date macros like \{{ latest\_date }}\ and \{{ max\_date }}\, while \EventSource\ and \Query\ structs have gained new fields for subpartitioning and partition columns. Alongside these schema changes, a Bazel build file has been introduced to generate Python and Java libraries from the Thrift definition, enabling consistent code generation across languages.

api/thrift · high confidence

Test coverage

Added Python unit tests for Chronon configuration and runtime components; Added Scala unit tests for API data types, extensions, and macros; Added sample GroupBy definitions for the quickstart test suite; Added sample join configuration files for testing; Added sample join configurations for unit testing; Added sample scripts for testing Chronon workflows; Added sample staging query definitions for testing; Added sample test data for checkout, purchase, return, and user entities; Added test coverage for Kaggle Outbrain GroupBys and Joins; Added test coverage for Kaggle Outbrain source definition; Added test coverage for new aggregation operations and utilities; Added test coverage for online fetcher, serialization, and caching components; Added test coverage for online serving components; Added test fixtures for staging query and real-time event sources; Added test fixtures for standalone staging queries; Added test sample for Chronon join training set configuration; Added tests for lineage metadata parsing; Added unit test fixtures for group-by configurations; Expanded sample GroupBy definitions for test coverage; Expanded sample join definitions for comprehensive test coverage; Expanded test coverage for Spark feature components; Updated sample group-by test fixtures for validation.

Dependencies

51 commits updating dependencies (5 manifests)

A dependency / build maintenance change in (dependencies) — 51 commits (11 fixs), 5 files.

(dependencies) · medium confidence · unverified

Added PySpark and SQLGlot dependencies for feature development

The Python API environment now includes PySpark (version 3.3.1) and SQLGlot (version 26.16.1) as new dependencies, enabling PySpark and notebook integration for Chronon feature development. This change also updates several existing packages, including Click (8.1.3 to 8.1.8), Six (1.16.0 to 1.17.0), and various development tools like Black, Pytest, and Tox, while adding pre-commit hooks and related linting tools (isort, autoflake) to the development requirements.

api/py/requirements · high confidence

Build infrastructure modernization and logging improvements

The SBT build system has been upgraded from version 1.2.8 to 1.8.2, and the sbt-assembly plugin has been updated from 0.14.10 to 2.1.1. Several new SBT plugins were added to enhance the build process, including sbt-scalafmt for code formatting, sbt-release for release management, sbt-git for versioning, sbt-buildinfo for build information generation, and sbt-sonatype for publishing. Additionally, internal build scripts (FolderCleaner and ThriftGen) now use SLF4J logging instead of println statements, and a new VersionDependency utility was introduced to manage Scala version-specific dependencies.

project · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 48 → 54 (+6.1)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 86 → 85 (-0.2)
  • Architecture 96 → 76 (-19.9)
  • Maturity 70 → 70 (+0.3)
  • Readiness 32 → 39 (+7.0)
  • Security 47 → 62 (+15.2)

Resolved (7)

  • Documentation: no installation or build instructions (README.md)
  • Documentation: written for insiders (docs/source/setup/Overview.md)
  • Hotspot: online/src/main/scala/ai/chronon/online/FetcherBase.scala (online/src/main/scala/ai/chronon/online/FetcherBase.scala)
  • Hotspot: spark/src/main/scala/ai/chronon/spark/Driver.scala (spark/src/main/scala/ai/chronon/spark/Driver.scala)
  • Hotspot: spark/src/main/scala/ai/chronon/spark/GroupBy.scala (spark/src/main/scala/ai/chronon/spark/GroupBy.scala)
  • Hotspot: spark/src/main/scala/ai/chronon/spark/streaming/JoinSourceRunner.scala (spark/src/main/scala/ai/chronon/spark/streaming/JoinSourceRunner.scala)
  • Off-boarding risk: anonymized user #1

New (27)

  • Duplicated block (10 lines × 2) (airflow/group_by_dag_constructor.py)
  • Duplicated block (10–11 lines × 2) (airflow/helpers.py)
  • Duplicated block (12–13 lines × 3) (api/py/ai/chronon/pyspark/executables.py)
  • Duplicated block (13 lines × 2) (api/py/ai/chronon/group_by.py)
  • Duplicated block (16–17 lines × 2) (api/py/ai/chronon/pyspark/executables.py)
  • Duplicated block (17–19 lines × 2) (airflow/decorators.py)
  • Duplicated block (6 lines × 2) (online/src/main/java/ai/chronon/online/TBaseDecoderFactory.java)
  • Duplicated block (8 lines × 2) (api/py/ai/chronon/repo/compile.py)
  • Duplicated block (9 lines × 2) (api/py/ai/chronon/lineage/lineage_parser.py)
  • No ADRs found
  • Off-boarding risk: anonymized user #1
  • Outdated: ch.qos.logback:logback-classic
  • Outdated: com.datadoghq:java-dogstatsd-client
  • Outdated: com.github.ben-manes.caffeine:caffeine
  • Outdated: com.google.code.gson:gson
  • Outdated: com.typesafe:config
  • Outdated: io.micrometer:micrometer-registry-statsd
  • Outdated: io.netty:netty-all
  • Outdated: io.vertx:vertx-config
  • Outdated: io.vertx:vertx-core
  • …and 7 more

Changes since last survey

  • 8 commits — 7 feature/other, 1 fixes

By area

  • spark/src — 4 commits
  • api/py — 2 commits
  • (root) — 1 commit
  • online/src — 1 commit

Notable commits

  • fix: [BugFix][Python] Skip invalid nested GroupBys in explore (#1154)
  • change: Add option to grant public read access on Chronon-created tables (#1152)
  • change: Adding a new parameter: group by upload cadence (#1153)
  • change: Chronon release 0.0.115 (#1162)
  • change: Parallelize MetadataExporter's per-entity analysis (#1157)
  • change: Stop incrementing push_notification.count from the KV-write path (#1158)
  • change: [Online] Bind CatalystUtil evaluators to the Chronon session conf on every thread (#1156)
  • change: [Streaming] Add Api serialization check and an end-to-end push-mode streaming test (#1159)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

airbnb/chronon was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 5ecd78c2406dd3a44eeb140f91ecd6f9486d0170 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.