typelevel/frameless
60.7
Adequate · 28 September 2026
8k
lines of production code
Scala
primary language
2
measurements over time
What this system is
This system is a type-safe Scala library for Apache Spark that provides a strongly-typed Dataset API, replacing untyped operations with compile-time checked expressions. It extends Spark's capabilities by integrating with the Cats ecosystem for effect handling and offering typed wrappers for Spark ML algorithms, including classification, regression, and clustering. The library ensures data integrity through automatic encoders for case classes and refined types, while maintaining compatibility across multiple Spark versions.
How it got here
2015–2016 — Initial release and Cats integration
8 changes.
This period established the Frameless library with a type-safe API for Apache Spark Datasets, introducing core components like TypedDataset and comprehensive encoder support. It expanded the ecosystem by integrating the Cats library for effect-type operations and refining the underlying Catalyst type classes to ensure precise type mapping and aggregation behavior.
2017 — Typed ML wrappers and testing
5 changes.
This period focused on introducing type-safe wrappers for Apache Spark ML algorithms within the Frameless library, enabling seamless integration with TypedDatasets. The work included implementing core traits and concrete models for classification, regression, and feature transformation. Comprehensive test coverage was added to verify encoder correctness, type constraints, and end-to-end pipeline functionality.
2018–2023 — Spark 3 compatibility and test expansion
4 changes.
This period focused on establishing Spark 3 compatibility by implementing internal API wrappers and logical plan support for newer Spark versions. Concurrently, the project expanded test coverage for clustering algorithms and query optimization features, while adding support for refined types in encoders.
Features
Add Cats integration with effect-type parameterized Spark operations
This change introduces the \cats\ module, providing integration with the Cats ecosystem. It adds syntax for running Spark operations within effect types (via \cats.effect.Sync\ and \cats.mtl.Ask\), allowing users to set local Spark properties like job group IDs and descriptions within their effect chains. It also provides commutative aggregation operations (sum, min, max) for RDDs using Cats typeclasses, along with semigroup instances for combining RDDs via union, inner join, and outer join.
cats · high confidence
Initial project setup with Spark 4.2 support and standardized formatting
The repository is initialized with core configuration files including a \.scalafmt.conf\ (version 3.11.5) to enforce code style, a \.gitignore\ for build artifacts, and a \README.md\ documenting the library. The documentation establishes compatibility with Spark 4.2 (requiring Scala 2.13 and JDK 17+) alongside Spark 3.5, and introduces environment variables \FRAMELESS\_GEN\_MIN\_SIZE\ and \FRAMELESS\_GEN\_SIZE\_RANGE\ to allow users to customize the size of generated collections in property tests to prevent memory issues. A \github.sbt\ file configures the CI workflow to run tests with coverage reporting and upload results to Codecov.
(repo-wide) · high confidence
Initial release of Frameless typed Dataset API
Introduces the core Frameless library, providing a type-safe interface for Apache Spark Datasets. This includes the \TypedDataset\ class for strongly-typed operations, \TypedEncoder\ and \RecordEncoder\ for automatic serialization of case classes and primitive types, and \TypedColumn\ for type-safe column expressions. The release also adds syntax extensions via \FramelessSyntax\ for converting untyped Spark columns and datasets to their typed equivalents, along with comprehensive aggregate and non-aggregate function implementations.
dataset/src/main/scala · high confidence
Introduce typed ML wrappers for Spark ML algorithms
This change adds a new \frameless.ml\ package that provides type-safe wrappers around Apache Spark ML components. It introduces \TypedEstimator\ and \TypedTransformer\ traits to bridge Frameless \TypedDataset\s with Spark ML pipelines, handling column mapping via internal input checkers (e.g., \TreesInputsChecker\, \VectorInputsChecker\). Concrete implementations are provided for several algorithms, including \TypedLinearRegression\, \TypedRandomForestClassifier\, \TypedRandomForestRegressor\, \TypedKMeans\, and \TypedBisectingKMeans\, as well as feature transformers like \TypedStringIndexer\, \TypedIndexToString\, and \TypedVectorAssembler\. The package also includes necessary parameter types (e.g., \KMeansInitMode\, \LossStrategy\) and internal utilities to manage Spark ML's private UDTs for Vector and Matrix types.
ml/src/main · high confidence
Support for refined types in Frameless encoders
Frameless now supports encoding and decoding of types from the refined library. This change adds implicit TypedEncoder and RecordFieldEncoder instances for refined types (including optional refined fields), allowing users to use refined types like NonEmptyString directly in TypedDatasets and case classes without manual conversion.
refined · high confidence
Behavioural changes
Added Spark 3 compatibility layer for Frameless internals
This change introduces a new compatibility layer in the Spark 3 module, adding \MapGroups.scala\ to support the \MapGroups\ logical plan and \FramelessInternals.scala\ to provide internal utilities for expression resolution, dataset creation, and join handling. These additions enable the library to function correctly on Spark 3.4.0 and DBR 12.2 by implementing the necessary internal APIs and logical plan wrappers specific to that version.
dataset/src/main/spark-3 · high confidence
New core type classes for typed Catalyst operations
The core module now includes a comprehensive set of new type classes (such as CatalystAverageable, CatalystBitwise, CatalystCast, CatalystNumeric, and CatalystOrdered) that explicitly define how Scala types map to Spark Catalyst operations. This change introduces precise return-type handling for aggregations (e.g., Int/Long summing to Long, averaging to Double), adds support for bitwise and shift operations, and enables automatic derivation of ordering for case classes via Shapeless. It also adds support for Java time types (Duration, Instant, Period) in ordering and defines specific behaviors for rounding, variance, and pivoting.
core · high confidence
Restructured documentation and CI publishing scripts
The project introduces new shell scripts to streamline documentation generation and continuous integration workflows. The \docs-build.sh\ script automates the process of copying the README, running \mdoc\, and building the GitBook site, while \docs-publish.sh\ handles merging the generated documentation into the \gh-pages\ branch. Additionally, \travis-publish.sh\ reorganizes the CI pipeline into three distinct phases: generating documentation and running tutorials, collecting code coverage reports via Codecov, and publishing artifacts locally, effectively replacing the previous monolithic build logic.
scripts · high confidence
Test coverage
Added tests for Frameless ML Vector and Matrix encoders; Added tests for K-Means and Bisecting K-Means clustering; Added tests for TypedIndexToString, TypedStringIndexer, and TypedVectorAssembler; Added tests for TypedLinearRegression and TypedRandomForestRegressor; Added tests for literal pushdown in Spark 3.2; Added tests for the new Frameless ML classification components; Expanded test coverage for TypedDataset operations; Migrate test logging configuration from Log4j 1 to Log4j 2.
Dependencies
Update dependencies to Spark 3.5.9, Scala 2.13.18, and Cats 2.13.0
The build configuration has been updated to use Spark 3.5.9 as the primary version, with support for Spark 4.0, 4.1, and 4.2 also defined. The Scala version is set to 2.13.18 (with 2.12.21 as a cross-compilation target). Key library updates include Cats Core to 2.13.0, Cats Effect to 3.7.1, Cats MTL to 1.7.0, Shapeless to 2.3.13, ScalaCheck to 1.20.0, and Refined to 0.11.4. This ensures compatibility with the latest stable releases of the underlying frameworks and improves security and stability for users depending on these versions.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 61 → 61 (+0.1)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 96 → 96 (+0.0)
- Architecture 71 → 72 (+0.8)
- Maturity 54 → 54 (+0.0)
- Readiness 63 → 63 (+0.0)
- Security 62 → 62 (+0.0)
Resolved (3)
- Documentation: no architecture or design documentation (docs/Cats.md)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
New (3)
- Dependency hygiene PARTLY measured — sbt declarations read, no exact version to grade for currency
- Documentation: no installation or build instructions (docs/TypedML.md)
- Projects may be oversized for their cohesion
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
typelevel/frameless was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 592bffe7a2535db40ff1de33ac6a3a043a6c0353 — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-2d9048c36d26.