Skip to content
CAI
Software that uses CAICheck a score

typelevel/frameless

60.7

Adequate · 28 September 2026

8k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a type-safe Scala library for Apache Spark that provides a strongly-typed Dataset API, replacing untyped operations with compile-time checked expressions. It extends Spark's capabilities by integrating with the Cats ecosystem for effect handling and offering typed wrappers for Spark ML algorithms, including classification, regression, and clustering. The library ensures data integrity through automatic encoders for case classes and refined types, while maintaining compatibility across multiple Spark versions.

How it got here

2015–2016 — Initial release and Cats integration

8 changes.

This period established the Frameless library with a type-safe API for Apache Spark Datasets, introducing core components like TypedDataset and comprehensive encoder support. It expanded the ecosystem by integrating the Cats library for effect-type operations and refining the underlying Catalyst type classes to ensure precise type mapping and aggregation behavior.

2017 — Typed ML wrappers and testing

5 changes.

This period focused on introducing type-safe wrappers for Apache Spark ML algorithms within the Frameless library, enabling seamless integration with TypedDatasets. The work included implementing core traits and concrete models for classification, regression, and feature transformation. Comprehensive test coverage was added to verify encoder correctness, type constraints, and end-to-end pipeline functionality.

2018–2023 — Spark 3 compatibility and test expansion

4 changes.

This period focused on establishing Spark 3 compatibility by implementing internal API wrappers and logical plan support for newer Spark versions. Concurrently, the project expanded test coverage for clustering algorithms and query optimization features, while adding support for refined types in encoders.

Features

Add Cats integration with effect-type parameterized Spark operations

This change introduces the \cats\ module, providing integration with the Cats ecosystem. It adds syntax for running Spark operations within effect types (via \cats.effect.Sync\ and \cats.mtl.Ask\), allowing users to set local Spark properties like job group IDs and descriptions within their effect chains. It also provides commutative aggregation operations (sum, min, max) for RDDs using Cats typeclasses, along with semigroup instances for combining RDDs via union, inner join, and outer join.

cats · high confidence

Initial project setup with Spark 4.2 support and standardized formatting

The repository is initialized with core configuration files including a \.scalafmt.conf\ (version 3.11.5) to enforce code style, a \.gitignore\ for build artifacts, and a \README.md\ documenting the library. The documentation establishes compatibility with Spark 4.2 (requiring Scala 2.13 and JDK 17+) alongside Spark 3.5, and introduces environment variables \FRAMELESS\_GEN\_MIN\_SIZE\ and \FRAMELESS\_GEN\_SIZE\_RANGE\ to allow users to customize the size of generated collections in property tests to prevent memory issues. A \github.sbt\ file configures the CI workflow to run tests with coverage reporting and upload results to Codecov.

(repo-wide) · high confidence

Initial release of Frameless typed Dataset API

Introduces the core Frameless library, providing a type-safe interface for Apache Spark Datasets. This includes the \TypedDataset\ class for strongly-typed operations, \TypedEncoder\ and \RecordEncoder\ for automatic serialization of case classes and primitive types, and \TypedColumn\ for type-safe column expressions. The release also adds syntax extensions via \FramelessSyntax\ for converting untyped Spark columns and datasets to their typed equivalents, along with comprehensive aggregate and non-aggregate function implementations.

dataset/src/main/scala · high confidence

Introduce typed ML wrappers for Spark ML algorithms

This change adds a new \frameless.ml\ package that provides type-safe wrappers around Apache Spark ML components. It introduces \TypedEstimator\ and \TypedTransformer\ traits to bridge Frameless \TypedDataset\s with Spark ML pipelines, handling column mapping via internal input checkers (e.g., \TreesInputsChecker\, \VectorInputsChecker\). Concrete implementations are provided for several algorithms, including \TypedLinearRegression\, \TypedRandomForestClassifier\, \TypedRandomForestRegressor\, \TypedKMeans\, and \TypedBisectingKMeans\, as well as feature transformers like \TypedStringIndexer\, \TypedIndexToString\, and \TypedVectorAssembler\. The package also includes necessary parameter types (e.g., \KMeansInitMode\, \LossStrategy\) and internal utilities to manage Spark ML's private UDTs for Vector and Matrix types.

ml/src/main · high confidence

Support for refined types in Frameless encoders

Frameless now supports encoding and decoding of types from the refined library. This change adds implicit TypedEncoder and RecordFieldEncoder instances for refined types (including optional refined fields), allowing users to use refined types like NonEmptyString directly in TypedDatasets and case classes without manual conversion.

refined · high confidence

Behavioural changes

Added Spark 3 compatibility layer for Frameless internals

This change introduces a new compatibility layer in the Spark 3 module, adding \MapGroups.scala\ to support the \MapGroups\ logical plan and \FramelessInternals.scala\ to provide internal utilities for expression resolution, dataset creation, and join handling. These additions enable the library to function correctly on Spark 3.4.0 and DBR 12.2 by implementing the necessary internal APIs and logical plan wrappers specific to that version.

dataset/src/main/spark-3 · high confidence

New core type classes for typed Catalyst operations

The core module now includes a comprehensive set of new type classes (such as CatalystAverageable, CatalystBitwise, CatalystCast, CatalystNumeric, and CatalystOrdered) that explicitly define how Scala types map to Spark Catalyst operations. This change introduces precise return-type handling for aggregations (e.g., Int/Long summing to Long, averaging to Double), adds support for bitwise and shift operations, and enables automatic derivation of ordering for case classes via Shapeless. It also adds support for Java time types (Duration, Instant, Period) in ordering and defines specific behaviors for rounding, variance, and pivoting.

core · high confidence

Restructured documentation and CI publishing scripts

The project introduces new shell scripts to streamline documentation generation and continuous integration workflows. The \docs-build.sh\ script automates the process of copying the README, running \mdoc\, and building the GitBook site, while \docs-publish.sh\ handles merging the generated documentation into the \gh-pages\ branch. Additionally, \travis-publish.sh\ reorganizes the CI pipeline into three distinct phases: generating documentation and running tutorials, collecting code coverage reports via Codecov, and publishing artifacts locally, effectively replacing the previous monolithic build logic.

scripts · high confidence

Test coverage

Added tests for Frameless ML Vector and Matrix encoders; Added tests for K-Means and Bisecting K-Means clustering; Added tests for TypedIndexToString, TypedStringIndexer, and TypedVectorAssembler; Added tests for TypedLinearRegression and TypedRandomForestRegressor; Added tests for literal pushdown in Spark 3.2; Added tests for the new Frameless ML classification components; Expanded test coverage for TypedDataset operations; Migrate test logging configuration from Log4j 1 to Log4j 2.

Dependencies

Update dependencies to Spark 3.5.9, Scala 2.13.18, and Cats 2.13.0

The build configuration has been updated to use Spark 3.5.9 as the primary version, with support for Spark 4.0, 4.1, and 4.2 also defined. The Scala version is set to 2.13.18 (with 2.12.21 as a cross-compilation target). Key library updates include Cats Core to 2.13.0, Cats Effect to 3.7.1, Cats MTL to 1.7.0, Shapeless to 2.3.13, ScalaCheck to 1.20.0, and Refined to 0.11.4. This ensures compatibility with the latest stable releases of the underlying frameworks and improves security and stability for users depending on these versions.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 61 → 61 (+0.1)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 96 → 96 (+0.0)
  • Architecture 71 → 72 (+0.8)
  • Maturity 54 → 54 (+0.0)
  • Readiness 63 → 63 (+0.0)
  • Security 62 → 62 (+0.0)

Resolved (3)

  • Documentation: no architecture or design documentation (docs/Cats.md)
  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)

New (3)

  • Dependency hygiene PARTLY measured — sbt declarations read, no exact version to grade for currency
  • Documentation: no installation or build instructions (docs/TypedML.md)
  • Projects may be oversized for their cohesion

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

typelevel/frameless was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 592bffe7a2535db40ff1de33ac6a3a043a6c0353 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.