VirtusLab/iskra
45.5
Weak · 20 September 2026
1.4k
lines of production code
Scala
primary language
1
measurement over time
What this system is
Iskra is a Scala library that provides a type-safe DataFrame API for Apache Spark, enforcing schema correctness at compile time through Scala 3 macros. It enables developers to perform structured data operations—such as filtering, joining, and aggregating—while ensuring column names, types, and nullability are validated before runtime. The system replaces legacy untyped approaches with a unified interface for building and manipulating typed and untyped DataFrames.
Features
Add typed aggregate and literal column functions
New files in src/main/functions introduce typed aggregation functions (sum, max, min, avg) and a literal column constructor (lit). The aggregates are implemented as classes wrapping Spark SQL functions, accepting typed Col arguments and returning typed results, while lit creates a Col from a primitive value using a PrimitiveEncoder. A when.scala file re-exports the existing When.when functionality into the functions package.
src/main/functions · high confidence
Introduce typed DataFrame API with compile-time schema validation
The library now provides a new typed DataFrame abstraction in \src/main\ that enforces schema correctness at compile time. This introduces \StructDataFrame\ for schema-aware operations and \ClassDataFrame\ for encoding case classes, replacing the previous untyped approach. Users can now leverage type-safe column operations (via \ColumnOps\), structured selection (\Select\), filtering (\Where\), joining (\Join\), and grouping (\GroupBy\) where the compiler validates column names, data types, and nullability. The API uses Scala 3 macros to generate schema views and validate operations, ensuring that errors like mismatched types or missing columns are caught before runtime.
src/main · high confidence
Introduce unified API surface for Iskra DataFrame operations
The \src/main/api\ module now exposes a consolidated entry point (\org.virtuslab.iskra.api\) that aggregates core types, operations, and functions for users. This includes exports for primitive and struct data types (with nullable variants), column and DataFrame classes (including \ClassDataFrame\ and \StructDataFrame\), and key functions like \lit\ and \when\. It also re-exports \SparkSession\ and utility objects, providing a single import location for building and manipulating typed and untyped DataFrames.
src/main/api · high confidence
Introduces a new Scala 3 type system for data columns and encoders
The \src/main/types\ module now provides a comprehensive type-level framework for defining column data types, nullability, and serialization. It introduces a hierarchy of \DataType\ traits (e.g., \BooleanLike\, \StringLike\) with distinct nullable and non-nullable variants (e.g., \boolean\_?\, \boolean\), enabling the compiler to enforce type safety for column operations. The module also defines \Encoder\ traits and implementations that map Scala types to Spark SQL catalyst types, supporting both primitive values and case classes via Scala 3 mirrors. Additionally, it includes utilities for type coercion, equality, and ordering within specific type families (numeric, boolean, string).
src/main/types · high confidence
Removals
Removal of legacy TypedSpark implementation and example
The \TypedSpark\ library code, including its core \TypedSpark.scala\ implementation, debugging utilities (\ShowType.scala\), and the \HelloSpark\ example file, has been removed from the source tree. This change eliminates the previous \TableSchema\-based API and its associated macro-driven schema derivation, clearing the way for the new \FrameSchema\-based architecture introduced in parallel commits.
src/main/scala · high confidence
Behavioural changes
Added publish configuration metadata for the Iskra project
A new \.publish/publish.scala\ file has been introduced to define publication metadata for the Iskra artifact. This configuration specifies the organization as \org.virtuslab\, the project name as \iskra\, and sets the version to be derived from git tags. It also establishes the repository URL, version control location, Apache-2.0 license, central repository target, and identifies Michał Pałka as the developer.
.publish · high confidence
Project rebranded to Iskra with updated documentation and build configuration
The project has been renamed from 'Typed Spark' to 'Iskra', reflected in the updated README, the new USAGE-DEV.md guide, and the package structure. The build system has been migrated to use scala-cli (via the \project.scala\ and \scala-cli\ scripts), establishing Scala 3.3.7 and Spark 3.2.0 as the core dependencies. An Apache 2.0 LICENSE file has been added to the repository, and the .gitignore has been updated to exclude scala-cli-specific build artifacts.
(repo-wide) · high confidence
Test coverage
Added example tests for Books, Countries, and Workers data processing; Added test coverage for core DataFrame operations.
Dependencies
Migration from SBT to Scala CLI
The project has removed the \build.sbt\ file, indicating a migration away from the SBT build tool (likely to Scala CLI, as suggested by commit messages). This change removes the explicit definition of the Scala 3.1.1-RC1 version and the Spark 3.2.0 dependencies (spark-core, spark-sql, spark-scala3) from the build configuration.
(dependencies) · medium confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 46.
Lenses
- Code Health 93
- Architecture 69
- Maturity 31
- Readiness 46
- Security 65
Changes since last survey
- 87 commits — 85 feature/other, 2 fixes
By area
- (repo) — 32 commits
- src/main — 31 commits
- (root) — 18 commits
- .github/workflows — 4 commits
- src/test — 2 commits
Notable commits
- fix: Publishing config quick fix
- fix: Publishing config quick fix 2
- change: * Allow nonstructural schema views for selections * Introduce Row class for grouping multiple named columns * Split the selection operation into .select and .selectRow
- change: * Bump scala version * Update using directives syntax for publish
- change: * Make .withNamedColumns accept NamedColumns * Drop .withNamedColumn * Bump scala to 3.3.0-RC5
- change: * Use Columns instead of Row for grouping (named) columns * Merge .select and .selectRow back into .select supporting both functionalities
- change: Add publishing config
- change: Basic join on condition without ambiguous columns
- change: Basic support for nested types
- change: Bump scala to 3.3.3 and scala-cli to 1.2.2
- change: Bump scala to 3.3.4, scala-cli to 1.5.1 and scalatest to 3.2.19
- change: Bump scala to LTS 3.3.6; Bump scala-cli to 1.9.1; Run CI for scala Next 3.7.3
- change: Bump scala to LTS 3.3.7; Bump scala-cli to 1.11.0; Run CI for scala Next 3.7.4; Run with java 11
- change: Bump scala-cli to 1.4.0
- change: Bump versions to 0.0.2 in README
- change: Centralize using directives
- change: Clean up named columns abstraction
- change: Cleanups + runtime representation of FrameSchema
- change: Code cleanup
- change: Configure basic CI
- …and 67 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
VirtusLab/iskra was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 6ef821e2dd7c89445247d8aa0a8703d40021cbe4 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.