delta-io/delta
71.9
Strong · 27 September 2026
272.8k
lines of production code
Scala
with Java
3
measurements over time
What this system is
This system is the Delta Lake open-source storage layer, providing a transactional data lakehouse format built on top of cloud object storage. It enables reliable, high-performance data operations including ACID-compliant reads and writes, schema evolution, and complex transformations like MERGE and streaming ingestion. The codebase implements the core protocol, a new Java-based Kernel for connector development, and integrations with major engines like Spark, Flink, and Unity Catalog.
How it got here
2019–2023 — Delta Kernel and V2 Connector Development
106 changes.
This period focused on establishing the Delta Kernel Java API and implementing the Delta V2 DataSource API to modernize the Spark connector architecture. It involved significant refactoring of core components like MERGE, CDC, and storage logic into modular traits, alongside the introduction of advanced features such as Deletion Vectors, Identity Columns, and Hilbert clustering.
2024–2025 — Kernel API expansion and Unity Catalog integration
91 changes.
This period focused on significantly expanding the Delta Kernel API with new interfaces for engines, metrics, and catalog-managed table commits, while introducing Unity Catalog integration for coordinated transactions. It also delivered major feature additions including Delta Connect support, Hudi Universal Format conversion, and advanced table management capabilities like row tracking and Z-ORDER clustering.
2026 — Flink kernel integration and Spark 4.2 compatibility
28 changes.
This period focused on introducing a Delta Kernel-backed implementation for the Flink connector, enabling advanced sink features like upserts, incremental checkpointing, and Unity Catalog integration. Concurrently, the project addressed source compatibility with Spark 4.2 API changes and expanded the Spark V2 connector with changelog and metadata-only delete support.
Features
Add Delta Connect Python Protobuf definitions
The \python/delta/connect/proto\ package now includes the generated Python Protocol Buffer code (\.py\ and \.pyi\ stubs) for the Delta Connect interface. This adds the foundational message types for accessing Delta tables (by path or name) and defines the command and relation structures required for operations such as cloning, vacuuming, upgrading protocol, generating manifests, creating tables, managing feature support, scanning, describing history/detail, converting to Delta, restoring, checking table status, deleting, updating, merging, and optimizing tables.
python/delta/connect/proto · high confidence
Add Delta Connect protocol definitions and Python code generation configuration
This change introduces the core Protobuf schema definitions for Delta Connect in the \spark-connect/common\ module, establishing the contract for remote execution of Delta Lake operations. The new files define the \delta.connect\ package with base messages for table access (\DeltaTable\) and a comprehensive set of commands and relations, including \CloneTable\, \VacuumTable\, \UpgradeTableProtocol\, \Generate\, \CreateDeltaTable\, \AddFeatureSupport\, \DropFeatureSupport\, \Scan\, \DescribeHistory\, \DescribeDetail\, \ConvertToDelta\, \RestoreTable\, \IsDeltaTable\, \DeleteFromTable\, \UpdateTable\, and \MergeIntoTable\. Additionally, it adds the \buf.gen.yaml\ configuration to enable Python Protobuf code generation for these definitions, ensuring that Python clients can generate the necessary stubs to interact with the Delta Connect server.
spark-connect/common · high confidence
Add DynamoDB-based Coordinated Commits client for Delta Lake
Delta Lake now supports using Amazon DynamoDB as a commit coordinator for Coordinated Commits. This change introduces the \DynamoDBCommitCoordinatorClient\ and its builder, which manage commit state (such as table version, timestamps, and unbackfilled commit metadata) in a DynamoDB table. Users can configure this backend via Spark SQL configuration keys (e.g., \spark.delta.coordinatedCommits.dynamoDBTableName\) to enable coordinated commit semantics backed by DynamoDB, replacing or supplementing file-system-based coordination.
spark/src/main/java/io/delta/dynamodbcommitcoordinator · high confidence
Add GoldenTableUtils for discovering and accessing golden test tables
A new utility object, GoldenTableUtils, has been added to the golden-tables connector to support the consolidation of testing infrastructure. This utility provides methods to locate the root 'golden' resource directory, retrieve specific table paths or files by name, and automatically discover all available golden tables by scanning for directories containing a '\_delta\_log' folder. This simplifies how tests access and iterate over the set of golden table fixtures.
connectors/golden-tables/src/main/scala · high confidence
Add Hilbert curve clustering and sort-within-files support for multi-dimensional clustering
Delta Lake now supports Hilbert curve-based multi-dimensional clustering alongside the existing Z-order method, allowing users to choose the clustering algorithm via the \curve\ parameter. The implementation introduces new internal functions (\hilbert\_index\, \range\_partition\_id\, \interleave\_bits\) to compute clustering keys. Additionally, a new configuration option \MDC\_SORT\_WITHIN\_FILES\ enables sorting data within partitions after clustering, which can improve query performance by keeping related rows physically closer together on disk.
spark/src/main/scala/org/apache/spark/sql/delta/skipping · high confidence
Add Iceberg-to-Delta conversion utilities and integration test
This change introduces the core implementation for converting Iceberg tables to Delta format, including utilities for schema translation, partition value conversion, and statistics mapping, along with a new integration test demonstrating the end-to-end conversion workflow.
iceberg · high confidence
Add JSON serialization utilities for Delta Kernel rows
The kernel-defaults module now includes a new \JsonUtils\ class that provides methods to serialize and deserialize Delta \Row\ objects to and from JSON strings. This utility supports a specific set of data types (including primitives, strings, structs, arrays, and maps with string keys) and is currently utilized in tests, with plans to support broader refactoring efforts for JSON file handling.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/json · high confidence
Add Mumbling bitmap reader and PFOR codec for Delta Lake deletion vectors
Delta Lake now includes a new read-only Mumbling bitmap implementation and a Patched Frame of Reference (PFOR) codec within the deletion vectors module. This adds the \BitPacking\, \ByteBuffers\, \MumblingBitmap\, and \PFOREncoding\ classes, enabling the system to decode Mumbling-compressed bitmap data used for tracking deleted rows. Users benefit from improved efficiency in reading and processing deletion vectors, which can lead to faster query performance on tables with frequent row deletions.
spark/src/main/java/org/apache/spark/sql/delta/deletionvectors · high confidence
Add Terraform infrastructure for AWS and GCP benchmarks
New Terraform configurations have been added to provision the cloud infrastructure required for running benchmarks. The AWS module (in \benchmarks/infrastructure/aws/terraform\) creates a VPC, S3 buckets, a MySQL metastore, and an EMR cluster using the \hashicorp/aws\ provider (v4.15.1). The GCP module (in \benchmarks/infrastructure/gcp/terraform\) provisions a Google Storage bucket, a Dataproc Metastore, and a Dataproc cluster using the \hashicorp/google\ and \hashicorp/google-beta\ providers (v4.22.0). Both modules include README documentation and variable definitions to configure regions, credentials, and cluster sizes.
(repo-wide) · high confidence
Add script to run Delta Kernel examples locally or against staged artifacts
Users can now execute Delta Kernel example programs (single-threaded and multi-threaded table readers) and integration tests via a new Python runner script. The script supports running examples against locally built JARs or against specific versions from a Maven repository, automatically managing artifact caches and invoking the necessary Maven commands to demonstrate API usage.
kernel/examples · high confidence
Add support for IBM Cloud Object Storage and Oracle Cloud Infrastructure
Delta Lake now includes new LogStore implementations for IBM Cloud Object Storage (IBMCOSLogStore) and Oracle Cloud Infrastructure (OracleCloudLogStore), enabling users to store transaction logs in these cloud environments. The IBM Cloud Object Storage implementation requires the Hadoop configuration \fs.cos.atomic.write\ to be set to true to ensure atomic writes, while the Oracle Cloud Infrastructure implementation utilizes atomic rename operations for non-overwrite writes. Both components are marked as unstable and include corresponding test suites to verify their behavior.
contribs · high confidence
Add utility to validate Parquet files for Iceberg Compatibility V2
A new utility class, ParquetIcebergCompatV2Utils, has been added to the Delta Lake UniForm module to help users verify whether existing Parquet data files meet the requirements for Iceberg Compatibility V2. This tool inspects Parquet file footers to ensure that timestamps are not stored as INT96 and that all fields (including nested types like LIST and MAP) possess the required field IDs, which are mandatory for Iceberg V2 compatibility. This enables users to check data readiness without performing a full table rewrite.
spark/src/main/scala/org/apache/spark/sql/delta/uniform · high confidence
Added Flink 2.0 Docker Compose configuration with S3 support
A new Docker Compose setup for Flink version 2.0.1 is now available in the flink/docker/2.0 directory. This configuration defines a JobManager and two TaskManagers, enabling users to run a local Flink cluster. The setup includes an initialization script that copies user-provided JARs and an S3 configuration file (core-site.xml) into the appropriate Flink directories, allowing for S3 integration out of the box.
flink/docker · high confidence
Added Scala build launcher and configuration for the Scala example
The Scala example now includes a dedicated build environment with an sbt launcher script, a library for managing sbt execution, and a repository configuration file. This setup allows users to build and run the Scala example using sbt, with repository sources configured to use Maven Central, Typesafe, and the Spark Packages repository (repos.spark-packages.org) instead of the deprecated Bintray service.
examples/scala/build · high confidence
Added internal utility classes for lazy evaluation and list operations
The kernel API now includes two new internal utility classes: \Lazy\, which provides a simple lazy-initialization wrapper using a supplier and optional, and \ListUtils\, which offers helper methods for partitioning lists and retrieving the first or last element (with a note to remove the latter once JDK 21+ is the minimum supported version). These additions support internal implementation details but do not expose new public APIs to users.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/lang · high confidence
Async telemetry of commit metrics to Unity Catalog
A new post-commit hook now asynchronously sends commit metrics for Unity Catalog-managed Delta tables to the Unity Catalog REST API. This feature is controlled by the \spark.conf\ setting \delta.uc.commit.metrics.enabled\ and is disabled by default. When enabled, the hook runs in a background thread pool to avoid impacting commit latency, constructing a fresh catalog client per commit to report telemetry data (including file size histograms from the snapshot checksum). Errors during transmission are logged but do not fail the commit.
spark/src/main/scala/org/apache/spark/sql/delta/hooks/metrics · high confidence
Delta Lake Spark Connect client API for table operations
The Spark Connect thin client now includes a full set of Scala APIs for managing Delta tables remotely. This adds a \DeltaTable\ class for standard operations like vacuum, history, and delete, alongside builder classes (\DeltaTableBuilder\, \DeltaColumnBuilder\) for creating and replacing tables with support for generated and identity columns. It also introduces \DeltaMergeBuilder\ for upserts (including \whenNotMatchedBySource\ updates) and \DeltaOptimizeBuilder\ for compaction and Z-Ordering. To ensure compatibility across Spark 4.0, 4.1, and 4.2, the client uses version-specific shims to handle the relocation of the \SparkStringUtils\ class.
spark-connect/client · high confidence
Enable Flink SQL and Catalog integration via service provider and configuration files
Users can now utilize Delta Lake with Flink SQL and external catalogs. This change introduces the \org.apache.flink.table.factories.Factory\ service provider file, registering \DeltaDynamicTableSinkFactory\ and \FlinkUnityCatalogFactory\ to enable automatic discovery of Delta sinks and Unity Catalog support in Flink SQL sessions. Additionally, a new \delta-flink.properties\ configuration file is added, providing default settings for sink retry attempts, table caching behavior, and credential refresh parameters, allowing users to tune performance and reliability without code changes.
flink/src/main/resources · high confidence
Flink kernel adds incremental v2 checkpointing, upsert expression helpers, and deletion-vector I/O
The Flink kernel module now includes new classes that enable incremental v2 checkpointing for Flink sink tables (CheckpointActionRow and CheckpointWriter), provide expression utilities for building row-IN predicates used in upsert merge paths (ExpressionUtils), and implement deletion-vector read/write/merge support via a new DVAccess abstraction and BinDVAccess file format (along with RoaringBitmapArray serialization). These changes add the underlying kernel capabilities that Flink sink tables rely on for efficient checkpointing, upserts, and deletion-vector operations.
flink/src/main/java/io/delta/flink/kernel · high confidence
Initial project scaffolding and repository setup
The repository is initialized with essential project infrastructure files, including the Apache 2.0 license and notice documents, a code of conduct, and a contributing guide. A comprehensive .gitignore is added to filter build artifacts, IDE settings, and test outputs, while .gitattributes standardizes line endings for Windows scripts. The build environment is configured with .sbtopts for JVM memory settings and .scalafmt.conf for code formatting. Additionally, a Dockerfile is introduced to create a reproducible Python 3.8 testing environment using the uv package manager with hash-verified dependencies, and the core Delta Transaction Protocol specification is documented in PROTOCOL.md.
(repo-wide) · high confidence
Intent-based builders for table creation, replacement, and updates
The Delta Kernel API now provides intent-based builders to manage table lifecycle operations with greater control. Users can create new tables using CreateTableTransactionBuilder, which supports configuring table properties, data layout specifications (partitioning and clustering), and custom commit logic via a committer interface. Existing tables can be replaced atomically using ReplaceTableTransactionBuilder, allowing updates to schema, properties, and layout. Additionally, UpdateTableTransactionBuilder enables schema evolution (adding, renaming, or dropping columns), property management, clustering updates, and idempotent operations via transaction identifiers, along with configurable retry and log compaction settings.
kernel/kernel-api/src/main/java/io/delta/kernel/transaction · high confidence
Introduce Adaptive Metadata Tree (AMT) checkpoint infrastructure
This change adds the core Scala components for the Adaptive Metadata Tree (AMT) feature, enabling Delta Lake to maintain a manifest-tree-based checkpoint structure. The new files in \spark/src/main/scala/org/apache/spark/sql/delta/amt\ define the \AMTCheckpointProvider\ to reconstruct table state from the manifest tree, \AMTWriteHelper\ and \AMTWriterManager\ to handle full and incremental tree writes, and \AMTUtils\ for path resolution and conflict-resolution metrics. It also includes \AMTPartitionValues\ and \AMTContentStats\ to convert Delta's string-based partition maps and JSON statistics into Iceberg V4's typed partition structs and content stats, ensuring compatibility with Iceberg readers.
spark/src/main/scala/org/apache/spark/sql/delta/amt · high confidence
Introduce Catalog-Owned table feature with Unity Catalog commit coordination
This change introduces the Catalog-Owned table feature, enabling Delta tables to be managed by Unity Catalog as their commit coordinator. It adds a new \CatalogOwnedCommitCoordinatorBuilder\ interface and a dedicated provider to register and retrieve catalog-specific commit coordinators, allowing tables to leverage Unity Catalog for transaction coordination. The implementation includes an in-memory mock (\InMemoryUCCommitCoordinator\ and \InMemoryUCClient\) to support testing of catalog-owned table operations without requiring a live Unity Catalog service. Additionally, it adds usage logging for catalog-owned tables and coordinated commits to track operational metrics and errors, such as invalid path-based access attempts or missing coordinator implementations.
spark/src/main/scala/org/apache/spark/sql/delta/coordinatedcommits · high confidence
Introduce CommitRange API for incremental table changes
The Delta Kernel now exposes a new CommitRange API that allows users to retrieve actions and commit metadata for a specific range of table versions. This feature enables incremental processing by letting callers define start and end boundaries (by version or timestamp) and retrieve the corresponding delta log files and actions. The implementation includes a builder pattern for configuring the range, validation logic to ensure version consistency, and factory methods to resolve boundaries and fetch the necessary data from the log.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/commitrange · high confidence
Introduce DefaultEngine and Hadoop-based handler implementations
The kernel-defaults module now provides a concrete DefaultEngine implementation that serves as the entry point for Delta Lake operations, replacing the previous direct Hadoop coupling with a generic FileIO abstraction. This engine wires together default implementations for expression evaluation, JSON parsing and writing, Parquet reading and writing, and file system operations, all backed by Hadoop APIs. Additionally, a LoggingMetricsReporter is included to output metrics reports as JSON via SLF4J, giving users visibility into engine performance and operation details.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/engine · high confidence
Introduce Delta Connect Python client for remote Spark sessions
This change adds the initial implementation of the Delta Connect Python client in the \python/delta/connect\ module, enabling Delta Lake operations over Spark Connect. It introduces a \DeltaTable\ class compatible with the remote execution model, supporting core table operations such as scan, delete, update, merge, vacuum, history, and detail. The module also includes the necessary logical plan definitions to serialize these operations to the server and registers custom exception classes to ensure Delta-specific errors are correctly converted and surfaced to Python users.
python/delta/connect · high confidence
Introduce Delta Format Sharing for Spark queries
Adds a new Delta Format Sharing implementation for Spark, enabling batch, streaming, and Change Data Feed (CDF) queries on Delta Sharing tables using the Delta log format. This includes a new \DeltaSharingDataSource\ and \DeltaFormatSharingSource\ to handle these workloads, a synthetic \DeltaSharingLogFileSystem\ to serve the constructed delta log from the block manager, and a \DeltaSharingFileIndex\ for batch access. The change also introduces \DeltaSharingDeltaLogBuilder\ to construct the synthetic log and manage presigned URL caching, \DeltaSharingCDFUtils\ for CDF preparation, \DeltaSharingJsonPredicates\ for server-side filter pushdown, and a \DeltaFormatSharingLimitPushDown\ optimizer rule to push limits to the sharing server.
sharing/src/main · high confidence
Introduce Delta Kernel Checksum (CRC) infrastructure
The Delta Kernel now includes a new internal checksum (CRC) system to store and verify table state. This change adds the \CRCInfo\ data model, a \ChecksumWriter\ to persist checksum files, and a \ChecksumReader\ to load them, along with \ChecksumUtils\ to compute the checksums. The checksum files capture critical table metadata and statistics, including protocol, metadata, file size histograms, domain metadata, deletion-vector metrics, and the full list of active files (when the table is small enough). This infrastructure enables the kernel to efficiently validate table state and supports incremental checksum updates.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/checksum · high confidence
Introduce Delta Kernel Unity Catalog integration layer
This change adds the \kernel/unitycatalog\ module, providing the core Java components for Delta Kernel to interact with Unity Catalog-managed tables. It introduces \UCCatalogManagedClient\ for loading snapshots and \UCCatalogManagedCommitter\ for handling catalog create and write operations, including support for Uniform metadata and domain metadata (e.g., clustering). The implementation includes utility classes like \UnityCatalogUtils\ for parsing delta files and extracting table properties, adapters for protocol and metadata actions, and a telemetry framework (\UcCommitTelemetry\, \UcLoadSnapshotTelemetry\, \UcPublishTelemetry\) to track performance metrics for these operations.
kernel/unitycatalog/src/main · high confidence
Introduce Delta Kernel data representation interfaces
The \io.delta.kernel.data\ package now provides the core Java interfaces for representing data within the Delta Kernel, including \ColumnVector\ for columnar access, \Row\ for record-level access, and \ColumnarBatch\ for batch operations. It also introduces \FilteredColumnarBatch\ to handle row selection via optional selection vectors, and \ArrayValue\/\MapValue\ to support complex nested types. These interfaces, marked as \@Evolving\ since version 3.0.0, form the foundation for reading and manipulating Delta table data in a columnar format.
kernel/kernel-api/src/main/java/io/delta/kernel/data · high confidence
Introduce Delta Spark V2 Connector with Changelog and Metadata-Only Delete support
The Delta Spark V2 Connector is a new implementation that bridges Apache Spark and Delta Kernel using the DataSource V2 (DSV2) APIs, allowing Spark to read and write Delta tables via the Kernel. This release adds support for reading table changelogs (CDC) through the Spark V2 Changelog API, enabling users to query change data between specific table versions. Additionally, for Spark 4.2 and later, the connector now supports metadata-only deletes, which allows deleting files based on partition predicates without scanning or rewriting data. The connector also includes shims to handle version-specific capabilities, such as schema evolution and partition-aware deletes, ensuring compatibility across different Spark versions.
spark/v2 · high confidence
Introduce Delta Storage module with cloud-specific LogStore implementations
The new \delta-storage\ module provides a pluggable \LogStore\ interface and concrete implementations for Azure, GCS, HDFS, and local file systems, enabling Delta Lake to write transaction logs with storage-system-specific atomicity and consistency guarantees. Azure and HDFS implementations use atomic rename operations to ensure files are visible only when fully written, while the GCS implementation handles GCS's strong consistency and precondition failures by writing in a separate thread, and the S3 implementation uses a per-JVM path lock to prevent concurrent writes to the same commit file. This modularization allows users to configure the appropriate storage backend via Hadoop configuration properties.
storage · high confidence
Introduce Delta Suite Generator for automated test suite creation
A new suite generator tool has been added to the project, allowing developers to automatically generate modularized and parameterized test suites. The generator, located in spark/delta-suite-generator, uses a configuration file (SuiteGeneratorConfig.scala) to define base test suites and various dimensions (such as table access methods, merge/update/delete operations, and feature flags) to create combinations. This automation helps manage the complexity of Delta Lake's test matrix by splitting large suites into smaller, more manageable files and ensuring consistent naming and formatting via scalafmt.
spark/delta-suite-generator · high confidence
Introduce Delta V2 DataSource and Source APIs
The Delta Lake Spark connector now exposes a new set of core classes for the V2 DataSource API, including \DeltaDataSource\, \DeltaSource\, \DeltaSink\, and \DeltaSourceCDCSupport\, alongside a consolidated \DeltaSQLConf\ configuration registry. This change provides the implementation foundation for the V2 streaming source and sink interfaces, enabling features such as schema evolution tracking, Change Data Feed (CDC) support, and admission control within the streaming engine, while maintaining the existing V1 integration layer.
spark/src/main/scala/org/apache/spark/sql/delta/sources · high confidence
Introduce FileSizeHistogram for tracking file size distributions
The kernel now includes a new \FileSizeHistogram\ class that tracks the distribution of file sizes within a Delta table. This component defines a specific schema with bin boundaries, file counts, and total bytes, and provides utilities to create default histograms, serialize them to rows, and deserialize them from column vectors. This enables the system to collect and propagate file size statistics through transaction metrics.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/stats · high confidence
Introduce Flink sink merge strategies for append and upsert workloads
The Flink Delta sink now supports configurable merge strategies via the new \mergestrategy\ package. For append-only sinks, the \AppendOnly\ strategy writes rows directly without modifying existing files. For upsert sinks, the system now implements \MoRUpsert\ (Merge-on-Read) as the default, which updates data by writing deletion vectors to existing Parquet files, and \CoWUpsert\ (Copy-on-Write), which rewrites affected files entirely. These strategies are coordinated by a \RowLocator\ interface (with a default \ScanLocator\ implementation) to identify candidate files based on primary keys, enabling efficient row-level updates in Flink Delta sinks.
flink/src/main/java/io/delta/flink/sink/mergestrategy · high confidence
Introduce Hudi Universal Format support for Delta tables
Delta Lake now supports Hudi as a Universal Format, allowing Delta tables to be automatically converted into Hudi tables. This change introduces the core conversion engine, including schema translation utilities that map Delta data types to Hudi Avro schemas, and transaction handling that commits Delta actions as Hudi replace commits. The implementation includes an asynchronous conversion thread to manage the transformation of Delta snapshots into Hudi format in the background, ensuring that the Hudi table at the same path stays synchronized with Delta commits.
hudi/src/main · high confidence
Introduce JsonMetadataDomain base class for structured metadata storage
The Delta Kernel now includes a new abstract base class, JsonMetadataDomain, which provides a standardized mechanism for serializing and deserializing metadata domain configurations to and from JSON strings using Jackson. This class simplifies the implementation of specific metadata domains (such as RowTrackingMetadataDomain) by handling the JSON conversion logic, allowing features to store structured configuration data within table metadata more reliably.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/metadatadomain · high confidence
Introduce Kernel-backed Delta table implementation for Flink
The Flink connector now supports a new table implementation backed by the Delta Kernel, providing a consistent foundation for interacting with Delta tables. This change introduces a new \AbstractKernelTable\ base class and specific implementations like \HadoopTable\ for filesystem-based tables and \CatalogManagedTable\ for Unity Catalog-managed tables, enabling unified handling of metadata, schema, and commit operations. It also adds a new \DeltaCatalog\ interface to abstract table discovery and credential vending, a \CredentialManager\ for proactive credential refresh, and a global \Conf\ class for operational tuning of sinks (e.g., retry policies, thread pools, caching).
flink/src/main/java/io/delta/flink/table · high confidence
Introduce Row Tracking Backfill and Unbackfill Commands
Delta Lake now provides commands to backfill and unbackfill row tracking metadata for existing tables. The new RowTrackingBackfillCommand re-commits files that lack base row IDs, enabling row tracking on tables that were created before this feature was available, while RowTrackingUnBackfillCommand removes this metadata to disable row tracking. These operations are implemented via a new backfill framework (BackfillCommand, BackfillExecutor, BackfillBatch) that processes files in configurable batches to avoid large commits, and includes protocol upgrades to support the RowTrackingFeature when necessary.
spark/src/main/scala/org/apache/spark/sql/delta/commands/backfill · high confidence
Introduce S3DynamoDBLogStore for concurrent Delta Lake commits on S3
This change adds the S3DynamoDBLogStore implementation, enabling Delta Lake to use an external DynamoDB table as a coordination layer for commit operations on S3. This allows multiple concurrent writers to safely commit to the same Delta table by using DynamoDB for mutual exclusion, preventing data loss in eventually consistent or non-mutually-exclusive storage environments. The implementation includes a base class (BaseExternalLogStore) for external log stores, a retryable iterator to handle transient S3 errors, and configuration options for DynamoDB table settings, credentials, and retry behavior.
storage-s3-dynamodb · high confidence
Introduce Table Redirect feature for routing queries to destination tables
Adds the core implementation for the Table Redirect feature, allowing Delta Lake tables to route read and write operations to a specified destination location. The new \TableRedirect.scala\ file defines the state machine (NoRedirect, EnableRedirectInProgress, RedirectReady, DropRedirectInProgress) and the \PathBasedRedirectSpec\ to manage source and destination paths, ensuring query consistency and preventing data loss during the redirection lifecycle.
spark/src/main/scala/org/apache/spark/sql/delta/redirect · high confidence
Introduce UnityCatalog as the committing catalog for Delta UniForm tables
Added \UnityCatalog\ and \UnityCatalogTableOperations\ to enable Iceberg metadata commits for Delta UniForm tables via Unity Catalog instead of HiveCatalog. This change allows Delta Lake tables to maintain their metadata in Unity Catalog while using Iceberg for storage, ensuring that schema and partition spec field IDs assigned by Delta are preserved during conversion to Iceberg V3 format.
icebergShaded/src/main/java/org/apache/iceberg/unityCatalog · high confidence
Introduce V2 checkpoint support and checkpoint protection in Delta Kernel
The Delta Kernel checkpoint subsystem has been refactored to support the new V2 checkpoint format alongside existing classic and multi-part checkpoints. This change introduces a \CheckpointInstance\ class that distinguishes between CLASSIC, MULTI\_PART, and V2 formats, enabling the engine to correctly parse and compare checkpoint files. A new \Checkpointer\ class centralizes checkpoint writing and reading logic, including the ability to write V2 checkpoints and handle sidecar files. Additionally, the system now supports \checkpointProtection\, which prevents the cleanup of expired log files when this feature is enabled, ensuring data integrity for protected tables. The \\_last\_checkpoint\ file handling has also been updated to capture and relay opaque JSON data for snapshot hints.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/checkpoints · high confidence
Introduce ZCube-based incremental clustering for Delta tables
Delta Lake now supports incremental clustering using a ZCube approach, allowing tables to be automatically organized by specified clustering columns to improve query performance. This change introduces new internal components, including ZCube file grouping, clustering statistics collection, and utility functions for managing clustering metadata via DomainMetadata. Users can now define clustering columns using the CLUSTER BY clause, and the system will automatically maintain data organization during writes and optimize existing data using the OPTIMIZE command with full support for clustered tables.
spark/src/main/scala/org/apache/spark/sql/delta/skipping/clustering · high confidence
Introduce collation-aware data skipping in the Kernel
The Kernel's data skipping engine now supports collation identifiers, allowing it to correctly prune files for string columns that use specific collations. This change introduces a new \DataSkippingPredicate\ class to track referenced columns and collations, updates \StatsSchemaHelper\ to include a \statsWithCollation\ field in the statistics schema, and modifies \DataSkippingUtils\ to construct and apply skipping filters that respect these collation settings.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/skipping · high confidence
Introduce default ColumnVector implementations for the Delta Kernel
The \kernel-defaults\ module now provides a complete set of default \ColumnVector\ implementations for reading data, including scalar types (boolean, byte, short, int, long, float, double, decimal, string, binary), complex types (array, map, struct), and generic views. This change introduces \AbstractColumnVector\ as a base class, specific implementations like \DefaultIntVector\ and \DefaultStructVector\, and utility vectors such as \DefaultViewVector\ for sub-setting data and \DefaultSubFieldVector\ for accessing nested struct fields. These components form the internal data access layer for the kernel's default reader, enabling efficient columnar data retrieval.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/data/vector · high confidence
Introduce default columnar batch and JSON row implementations
The kernel-defaults module now includes concrete implementations for core data structures: \DefaultColumnarBatch\ provides a standard columnar batch backed by \ColumnVector\ arrays, \DefaultRowBasedColumnarBatch\ wraps a list of \Row\ objects with lazy column vector generation, and \DefaultJsonRow\ decodes JSON data into typed row values. These classes form the internal data handling foundation for the kernel's default engine.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/data · high confidence
Introduce internal Deletion Vectors implementation for Delta Kernel
Added a new internal implementation for Deletion Vectors in the Delta Kernel, including classes for Base85 encoding/decoding, RoaringBitmapArray handling, and loading deletion vector bitmaps from inline or on-disk storage. This enables the Kernel to read and process deletion vectors, which are used to track deleted rows in Delta tables.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/deletionvectors · high confidence
Introduce internal Path utility for file system operations
The Delta Kernel now includes an internal \Path\ class in the \io.delta.kernel.internal.fs\ package to handle file and directory naming. This utility, adapted from Apache Hadoop, provides standardized path construction, resolution against parent paths, and normalization of URI components, including specific handling for Windows drive letters and directory separators.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/fs · high confidence
Introduce new internal bitmap and filtering infrastructure for Deletion Vectors
This change adds new internal Scala classes to the Deletion Vectors module to support manifest-based deletion vectors and optimized row filtering. It introduces the \ManifestBitmap\ trait and \RoaringManifestBitmap\ implementation to handle immutable bitmaps for AMT manifest and CDF tracking data, along with \RoaringBitmapArray\ for efficient 64-bit bitmap operations. Additionally, it adds \RowIndexMarkingFilters\ (including \DropMarkedRowsFilter\ and \KeepMarkedRowsFilter\) to enable predicate pushdown by filtering rows based on deletion vector presence, and \StoredBitmap\ to manage the loading and validation of deletion vector data from inline or on-disk sources.
spark/src/main/scala/org/apache/spark/sql/delta/deletionvectors · high confidence
Introduce post-commit hooks for automatic maintenance and format conversion
This change introduces a new \PostCommitHook\ framework in the Delta Lake hooks package, allowing automatic background tasks to run after successful table commits. Key hooks added include \AutoCompact\ for automatically compacting small files, \CheckpointHook\ and \ChecksumHook\ for maintaining table metadata integrity, \GenerateSymlinkManifest\ for creating Hive-style manifests for external tools like Presto, and \IcebergConverterHook\ and \HudiConverterHook\ for synchronizing Delta tables to Iceberg and Hudi formats (UniForm). The \UpdateCatalog\ hook is also present to keep the metastore schema and properties in sync. This enables users to automate maintenance and interoperability tasks without manual intervention.
spark/src/main/scala/org/apache/spark/sql/delta/hooks · high confidence
Introduce server-side scan planning for Delta tables without credentials
This change adds a new server-side planning infrastructure that allows Spark to offload file-list generation to a remote catalog service (currently Iceberg REST via Unity Catalog) when a table lacks local storage credentials. The implementation introduces a \ServerSidePlannedTable\ that acts as a fallback for credential-less tables, utilizing a factory pattern (\ServerSidePlanningClientFactory\) to instantiate planning clients. It includes specific metadata handling for Unity Catalog (extracting OAuth tokens and REST endpoints) and default metadata for testing. The feature is gated by the \spark.sql.delta.enableServerSidePlanning\ configuration flag and is designed to improve performance and security by avoiding local credential exposure for tables managed by external catalogs.
spark/src/main/scala/org/apache/spark/sql/delta/serverSidePlanning · high confidence
Introduces Deletion Vectors and Table Features support in Delta Lake
This change adds core support for Deletion Vectors (DV) and the Table Features protocol to Delta Lake. It introduces the \DeletionVectorDescriptor\ action to track deleted rows via inline or external storage, and updates \InMemoryLogReplay\ to reconcile file actions using DV object identity for accurate state reconstruction. Additionally, it implements the \TableFeatureSupport\ trait on the \Protocol\ class, enabling tables to explicitly declare supported reader and writer features (starting at protocol versions 3/7) and allowing features to be added or removed without downgrading the base protocol version.
spark/src/main/scala/org/apache/spark/sql/delta/actions · high confidence
Introduces Unity Catalog API routing for Delta table operations
The Delta catalog now supports routing managed table operations (CREATE, REPLACE, and loading) through the Unity Catalog Delta REST API instead of the standard Spark delegate catalog. This is enabled by default via the \deltaRestApi.enabled\ configuration and is implemented by introducing \AbstractDeltaCatalogClient\ and its \UCDeltaCatalogClientImpl\ to handle staging, credential vending, and property augmentation. The catalog also advertises support for generated columns, identity columns, and column default values, and includes a configuration to block Spark 4.0 clients from writing to Variant tables when using Unity Catalog.
spark/src/main/scala/org/apache/spark/sql/delta/catalog · high confidence
Introduces internal metrics collection and reporting for Delta Kernel operations
The Delta Kernel now collects and reports detailed performance metrics for snapshot construction, scans, and transactions. This change adds internal classes (Counter, Timer) to track operation durations and event counts, and specific metric collectors (SnapshotMetrics, ScanMetrics, TransactionMetrics) that record data such as snapshot load times, file add/remove counts and sizes, and commit durations. These metrics are packaged into immutable result objects and exposed via report implementations (SnapshotReportImpl, ScanReportImpl, TransactionReportImpl) that can be serialized to JSON using a pre-configured Jackson serializer, enabling observability into kernel operation performance.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/metrics · high confidence
Introduces post-commit hooks for checkpoints, checksums, and log compaction
The kernel now supports executing specific maintenance operations automatically after a transaction commits via a new \PostCommitHook\ interface and its implementations. Users benefit from streamlined table maintenance as the system can now trigger checkpoint writing, simple or full checksum file generation, and inline log compaction immediately following a commit, reducing the need for separate manual maintenance steps.
kernel/kernel-api/src/main/java/io/delta/kernel/hook, kernel/kernel-api/src/main/java/io/delta/kernel/internal/hook · high confidence
Introduction of AdmittableFile interface for Delta streaming admission control
A new AdmittableFile interface has been added to the Delta Lake streaming sources to support admission control logic. This interface defines methods to check if a file has an associated action and to retrieve its size, providing a unified abstraction that allows both DSv1 and DSv2 IndexedFile implementations to be used with the new admission control mechanism.
spark/src/main/java/org/apache/spark/sql/delta/sources · high confidence
Introduction of Deletion Vector storage abstraction and validation
The Delta Lake module now includes a new \DeletionVectorStore\ trait and its \HadoopFileSystemDVStore\ implementation, providing the core infrastructure for reading, writing, and managing Deletion Vectors. This change introduces strict validation during read operations, throwing exceptions if the file size or CRC32 checksum of a Deletion Vector does not match expectations, thereby ensuring data integrity. It also adds utility functions for calculating total file sizes, handling path escaping, and generating unique file names for Deletion Vectors within a table.
spark/src/main/scala/org/apache/spark/sql/delta/storage/dv · high confidence
Introduction of Delta Kernel Java type system and metadata APIs
The Delta Kernel API now exposes a complete Java type system for schema representation, including primitive types (Boolean, Byte, Short, Integer, Long, Float, Double, Date, Timestamp, TimestampNTZ, Binary, String, Variant, Decimal), complex types (Struct, Array, Map), and geospatial types (Geometry, Geography). StringType now supports collation identifiers, and MapType enforces that keys use the default UTF8\_BINARY collation. The API also introduces FieldMetadata for storing field-level metadata, MetadataColumnSpec for built-in metadata columns (row\_index, row\_id, row\_commit\_version), and TypeChange tracking for schema evolution.
kernel/kernel-api/src/main/java/io/delta/kernel/types · high confidence
Introduction of DeltaSqlParser for custom SQL command parsing
The Delta SQL parser component has been introduced to handle Delta-specific SQL commands. This parser intercepts SQL statements, attempting to parse them into Delta logical plans (such as those involving clustering or table features) before delegating standard queries to the underlying Spark parser. This change enables Delta to support its own syntax extensions and command structures while maintaining compatibility with standard Spark SQL.
spark/src/main/scala/io/delta/sql/parser · high confidence
Introduction of LogStoreProvider for scheme-based storage selection
The Delta Kernel now includes a new LogStoreProvider utility that automatically selects the appropriate LogStore implementation based on the file scheme (e.g., S3, Azure, GCS, HDFS) of the path being accessed. This allows the kernel to dynamically instantiate the correct storage handler (such as S3SingleDriverLogStore or AzureLogStore) either via configuration properties or by default, enabling seamless support for multiple cloud storage backends without manual configuration for standard schemes.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/logstore · high confidence
Kernel implementation of row tracking metadata and materialized columns
The Delta Kernel now includes internal support for row tracking, enabling tables to maintain unique row identifiers and commit versions. This change introduces the \RowTrackingMetadataDomain\ to persist the row ID high watermark as domain metadata, and the \MaterializedRowTrackingColumn\ class to manage the configuration and physical mapping of materialized \\_row-id\ and \\_row-commit-version\ columns. Additionally, the \RowTracking\ utility handles the assignment of base row IDs and default commit versions to new files, including logic for conflict resolution during concurrent transactions. For users, this provides the foundational kernel capability to read and write row tracking metadata, ensuring consistent row identity across table operations.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/rowtracking · high confidence
Kernel support for clustering column metadata storage
The Delta Kernel now includes internal infrastructure to persist and retrieve clustering configuration. A new \ClusteringMetadataDomain\ class handles serialization of clustering column names (supporting both logical and physical names depending on column mapping settings) into the \delta.clustering\ metadata domain, while \ClusteringUtils\ provides a helper to generate the corresponding domain metadata. This enables the kernel to store and read clustering definitions as part of the transaction log.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/clustering · high confidence
New Delta Kernel Java interfaces for table operations
The Delta Kernel module introduces a set of new Java interfaces to standardize table interactions. This includes \CommitActions\ and \CommitRange\ for iterating over specific commit versions and their actions, \Snapshot\ and \SnapshotBuilder\ for resolving table states by version or timestamp, and \Scan\ and \ScanBuilder\ for reading data with filtering and pagination support. Additionally, \Table\ serves as the entry point for resolving tables, while \Operation\ defines standard transaction types and \DataWriteContext\ provides context for data writes.
kernel/kernel-api/src/main/java/io/delta/kernel · high confidence
New Delta Kernel example programs for table reads and writes
Added a set of Java example programs in the \kernel-examples\ module that demonstrate how to use the Delta Kernel APIs for both reading and writing Delta tables. The new code includes \BaseTableReader\ and \BaseTableWriter\ as base utilities, along with concrete examples: \SingleThreadedTableReader\ and \MultiThreadedTableReader\ for querying tables with column pruning and filtering, and \CreateTable\, \CreateTableAndInsertData\ for creating unpartitioned and partitioned tables and inserting data (including CTAS and idempotent writes). A \RowSerDe\ utility is also provided to support serialization of scan state and files for distributed reading scenarios.
kernel/examples/kernel-examples · high confidence
New Delta Kernel expressions API for predicates and scalars
The \io.delta.kernel.expressions\ package now provides a structured API for defining and evaluating filter predicates and scalar expressions within Delta Kernel. This includes a base \Expression\ interface, \ScalarExpression\ for value transformations (such as \ELEMENT\_AT\, \COALESCE\, \ADD\, \TIMEADD\, and \SUBSTRING\), and \Predicate\ for boolean filters (including \AND\, \OR\, \NOT\, \IS\_NULL\, \IS\_NOT\_NULL\, \LIKE\, \IN\, and \StGeometryBoxesIntersect\). The API supports advanced features like nested column references via \Column\, literal values via \Literal\, and string collation for consistent comparison semantics. Connectors can implement \ExpressionEvaluator\ and \PredicateEvaluator\ to optimize these expressions against columnar data batches.
kernel/kernel-api/src/main/java/io/delta/kernel/expressions · high confidence
New Delta V2 connector mode and type-widening support
This change introduces a centralized configuration mechanism (DeltaV2Mode) that allows users to enable the new Delta V2 connector via the spark.databricks.delta.v2.enableMode setting, supporting NONE, AUTO (for Unity Catalog managed tables), and STRICT modes. It also adds support for automatic data type widening during schema evolution in the V2 write path, controlled by the delta.allow.automatic.widening configuration, while ensuring type changes are blocked if they would violate generated column or check constraints.
spark/src/main/java/org/apache/spark/sql/delta · high confidence
New Delta benchmark framework with TPC-DS and Merge workload support
The \benchmarks/src\ directory now contains a new open-source benchmarking framework for Delta Lake workloads. This includes a base \Benchmark\ class for measuring SQL query durations and generating JSON reports, alongside specific implementations for TPC-DS performance benchmarks (\TPCDSBenchmark\, \TPCDSDataLoad\) and Delta \MERGE\ operation benchmarks (\MergeBenchmark\, \MergeDataLoad\, \MergeTestCases\). The framework supports configurable scale factors (1GB and 3TB for TPC-DS), multiple iterations, and cloud storage paths for data and reports. A utility object \SparkUtils\ is also added to capture Spark environment information and provide a median calculation helper.
benchmarks/src · high confidence
New Delta performance benchmark framework with TPC-DS and Merge workloads
A new benchmarking framework has been introduced to measure Delta Lake performance on Spark clusters running on AWS EMR or Google Cloud Dataproc. The framework includes predefined benchmarks for TPC-DS (at 1GB and 3TB scales) and Merge operations, allowing users to load data and run 99 standard queries against Delta tables. It supports configuration for external Hive Metastores (RDS MySQL or Dataproc Metastore) and provides Terraform templates for infrastructure setup. The tooling includes a Python runner (\run-benchmark.py\) and SBT build scripts, targeting Delta version 2.3.0.
benchmarks · high confidence
New Flink SQL API for Delta Lake sinks and Unity Catalog integration
This change introduces a new Flink SQL interface for writing to Delta tables, allowing users to define sinks directly via Flink SQL DDL using the 'delta' connector identifier. The implementation supports both Hadoop-based file paths and Unity Catalog-managed tables, enabling upsert write modes when a primary key is specified in the table definition. It also includes a new 'unitycatalog' catalog factory for registering Unity Catalog as a Flink catalog, which maps Unity Catalog schemas and tables to Flink's internal representation, including support for complex data types and partitioning metadata.
flink/src/main/java/io/delta/flink/sink/sql · high confidence
New Iceberg compatibility metadata validators and updaters
The Delta Kernel now includes a new set of classes in the \icebergcompat\ package to validate and update table metadata for Iceberg compatibility. This includes \IcebergCompatMetadataValidatorAndUpdater\ as a base class, with specific implementations for \IcebergCompatV2\, \IcebergCompatV3\, \IcebergWriterCompatV1\, and \IcebergWriterCompatV3\. These validators enforce required table properties (like column mapping modes), check for supported data types and table features, and ensure metadata consistency when Iceberg compatibility features are enabled. Additionally, \IcebergUniversalFormatMetadataValidatorAndUpdater\ ensures that Universal Format settings are consistent with Iceberg compatibility settings.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/icebergcompat · high confidence
New JSON serialization and type widening validation in Kernel API
The Kernel API now includes a dedicated \DataTypeJsonSerDe\ class for serializing and deserializing Delta data types to and from JSON, replacing previous ad-hoc handling. Additionally, a new \TypeWideningChecker\ utility has been added to validate schema evolution operations, enforcing protocol rules that allow widening integer, floating-point, date, and decimal types (e.g., Byte to Long, Float to Double, Date to TimestampNTZ, and specific precision/scale increases for Decimals).
kernel/kernel-api/src/main/java/io/delta/kernel/internal/types · high confidence
New Scala APIs for Identity Columns, Merge Execution, and Table Optimization
The Delta Lake Scala API in \io.delta.tables\ has been expanded with new capabilities. \DeltaColumnBuilder\ now supports defining identity columns via \generatedAlwaysAsIdentity\ and \generatedByDefaultAsIdentity\, allowing users to create auto-incrementing columns. The \DeltaMergeBuilder\ has been refactored to execute MERGE operations using the DataFrame API, returning a \DataFrame\ instead of \Unit\ to provide execution metrics and enable chaining. Additionally, \DeltaOptimizeBuilder\ introduces \executeCompaction\ and \executeZOrderBy\ methods, giving users programmatic control over file compaction and Z-Order optimization directly from Scala code.
spark/src/main/scala/io/delta/tables · high confidence
New Scala examples for Delta Lake features
The examples/scala directory now includes a comprehensive set of runnable Scala programs demonstrating core and advanced Delta Lake capabilities. Users can explore Change Data Feed (CDC) with streaming and batch reads, table clustering, schema evolution for map types, and Iceberg compatibility (UniForm) with REORG operations. The collection also covers basic CRUD operations via both Scala APIs and SQL (Quickstart, QuickstartSQL, QuickstartSQLOnPaths), structured streaming with upserts, Unity Catalog integration for managed tables, and the new VARIANT data type. A Utilities example provides reference implementations for common administrative tasks like vacuum, history, and detail inspection.
examples/python, examples/scala/src/main/scala · high confidence
New Scala extension utilities for conditional option handling
A new ScalaExtensions utility object has been added to the Spark SQL utilities package, providing implicit classes that enhance standard Scala types. For Option values, it introduces ifDefined for clearer conditional execution, and for the Option companion object, it adds when and whenNot for conditional value creation, as well as sum for aggregating optional numeric values. Additionally, it provides condDo on Any to apply a partial function conditionally, improving code readability for conditional logic.
spark/src/main/scala/org/apache/spark/sql/util · high confidence
New benchmarking framework for Delta Lake workloads
This change introduces a new Python-based benchmarking framework located in \benchmarks/scripts\ to evaluate Delta Lake performance. It provides a structured way to define and run benchmarks, specifically adding support for TPC-DS and Merge workloads. The framework includes utility scripts for command execution and SSH operations, along with specification classes that configure Spark sessions with the necessary Delta extensions and Maven artifacts for running these specific benchmark types.
benchmarks/scripts · high confidence
New build and CI helper scripts for Spark, Unity Catalog, and test analytics
Added several new scripts to the project to streamline local development and CI workflows. The \build\_spark.sh\ script automates building a specific Apache Spark source reference and publishing it to the local Maven repository for Delta tests. The \setup\_unitycatalog\_main.sh\ script clones and publishes Unity Catalog artifacts to local Ivy/Maven repositories, supporting pinned SHAs, floating main canaries, and release pipelines. The \get\_spark\_version\_info.py\ script provides utilities for CI/CD to resolve Spark version metadata and compute artifact versions. Additionally, \install-buf.sh\ and \install-uv.sh\ provide SHA-verified installation of the Buf and uv tools, while \collect\_test\_durations.py\ aggregates test duration data from CI runs to update a project-level CSV file.
project/scripts · high confidence
New concurrency testing framework for Delta Lake transactions
Added a new concurrency testing framework in the Delta Lake fuzzer package, introducing \AtomicBarrier\, \ExecutionPhaseLock\, and \PhaseLockingExecutionObserver\ to control the execution order of transaction phases. This allows tests to enforce lock-step execution of operations like INSERT REPLACE and standard transactions, ensuring source materialization and commit behaviors are tested under controlled concurrent conditions.
spark/src/main/scala/org/apache/spark/sql/delta/fuzzer · high confidence
New default Parquet reader and writer implementations in Delta Kernel
The Delta Kernel default module now includes a complete, new implementation for reading and writing Parquet files, replacing previous approaches. The new reader (ParquetFileReader, ParquetBatchRowReader) supports column projection, row-group level predicate pushdown, and materializes data into columnar batches using dedicated type-specific readers (ArrayColumnReader, DecimalColumnReader, MapColumnReader, etc.). The new writer (ParquetFileWriter, ParquetColumnWriters) handles writing columnar batches to Parquet files, supporting multiple output files, atomic writes, and statistics collection. This change enables more efficient and flexible Parquet I/O operations within the Delta Kernel.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/parquet · high confidence
New default expression evaluation engine for the Delta Kernel
This change introduces the default implementation of the expression evaluation framework in the Delta Kernel, enabling the execution of a wide range of SQL expressions directly on columnar data. The new \DefaultExpressionEvaluator\ and \DefaultPredicateEvaluator\ handle validation, implicit type casting, and evaluation for predicates (including \AND\, \OR\, \=\, \\<\, \\>\, \IS NOT DISTINCT FROM\, \LIKE\, \STARTS\_WITH\, \IN\, and \IS NULL\/\IS NOT NULL\) as well as scalar functions (\element\_at\, \partition\_value\, \COALESCE\, \ADD\, \SUBSTRING\, \TIMEADD\). It also supports data skipping capabilities for specific predicates and geometry intersections, providing the foundational logic for query execution and partition pruning within the kernel.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/expressions · high confidence
New development tooling for code quality and Delta Connect code generation
The \dev\ directory now includes a suite of new scripts and configuration files to enforce code standards and manage generated assets. A new Python script (\check-delta-connect-codegen-python.py\) and shell script (\delta-connect-gen-protos.sh\) have been added to generate and verify the synchronization of Delta Connect Python protobuf code. Additionally, Checkstyle configurations (\connectors-checkstyle.xml\, \kernel-checkstyle.xml\) and a suppression file (\checkstyle-suppressions.xml\) are introduced to standardize Java code formatting, while a new Python linting script (\lint-python\) and configuration (\tox.ini\) enforce style checks using tools like flake8 and pycodestyle. A structured logging style checker (\spark\_structured\_logging\_style.py\) and a copyright header template (\copyrightHeader\) are also added to the development workflow.
dev · high confidence
New experimental commit API for catalog-managed tables
The Delta Kernel now exposes a new set of experimental APIs in the \io.delta.kernel.commit\ package to support committing changes to catalog-managed tables. This includes the \CatalogCommitter\ interface, which extends the standard \Committer\ to handle catalog-specific operations such as publishing ratified catalog commits to the Delta log and enforcing required table properties. The update introduces \CommitMetadata\ to carry detailed commit context (including protocol, metadata, and domain metadata), \CommitResponse\ to return commit results, and \CommitFailedException\ to distinguish between retryable and conflict-based failures. Utility classes like \CatalogCommitterUtils\ provide helpers for extracting protocol properties, while \PublishMetadata\ and \PublishFailedException\ manage the publishing workflow. These changes enable connectors to implement custom commit logic for catalog-managed tables while maintaining compatibility with existing filesystem-managed workflows.
kernel/kernel-api/src/main/java/io/delta/kernel/commit · high confidence
New expression allowlist and catalog maintenance operation definitions
This change introduces the AllowedUserProvidedExpressions allowlist, which restricts the custom expressions permitted in generated columns and check constraints to a specific set of deterministic, non-aggregate functions (such as math, string, and logical operations), and defines the CatalogManagedTableMaintenanceOperation constants (DATA\_CLEANUP, DATA\_REORGANIZATION, METADATA\_CLEANUP) that specify which maintenance operations are permitted on catalog-managed tables.
spark/src/main/scala/org/apache/spark/sql/delta · high confidence
New expression functions for Hilbert clustering, Z-ordering, and Variant stats serialization
This location introduces several new Catalyst expression classes to support advanced data organization and Variant type handling. Hilbert clustering is enabled via HilbertLongIndex and HilbertByteArrayIndex, which compute Hilbert curve distance keys from integer columns, supported by HilbertIndex, HilbertStates, and HilbertUtils. Z-ordering is supported by InterleaveBits, which interleaves bits from integer columns into a binary key, with an optional fast path for up to 9 columns. Variant stats serialization is handled by EncodeNestedVariantAsZ85String (encoding VariantVal fields to Z85 strings for JSON) and DecodeNestedZ85EncodedVariant (decoding Z85 strings back to VariantVal). Additionally, JoinedProjection provides a helper for binding attributes in joined projections used during statistics collection, and RangePartitionId/PartitionerExpr support range partitioning by rewriting into partitioner expressions.
spark/src/main/scala/org/apache/spark/sql/delta/expressions · high confidence
New implicit conversion methods for Spark data structures
The Delta Lake Spark module now provides new implicit classes in the \org.apache.spark.sql.delta.implicits\ package to simplify working with common data structures. Users can now call \toDF()\ directly on \Seq\[AddFile\]\, \Seq\[String\]\, and \Seq\[Int\]\ to convert them into DataFrames. Additionally, \StructType\ gains a \findNestedFieldIgnoreCase\ method for case-insensitive field lookup within nested maps and arrays, and \LogicalPlan\ gains \transformAllExpressionsUp\ and \transformAllExpressionsUpWithPruning\ methods for more flexible expression transformation during query planning.
spark/src/main/scala/org/apache/spark/sql/delta/implicits · high confidence
New intent-based transaction builders and commit action handling in Delta Kernel
Delta Kernel introduces intent-based builders for table operations, specifically the \CreateTableTransactionBuilder\ (implemented by \CreateTableTransactionBuilderImpl\) and the \CommitActions\ interface (implemented by \CommitActionsImpl\). The new builder allows users to construct and execute \CREATE TABLE\ transactions with support for table properties, data layout specs (partitioning and clustering), custom committers, and retry logic, while ensuring the table does not already exist. The \CommitActions\ implementation provides a memory-efficient way to read and iterate over Delta commit log files, handling protocol validation and column stripping. These changes are accompanied by new internal utilities for data write contexts (\DataWriteContextImpl\), error handling (\DeltaErrors\, \DeltaErrorsInternal\), and log action utilities (\DeltaLogActionUtils\), establishing the foundation for intent-driven write operations in the kernel.
kernel/kernel-api/src/main/java/io/delta/kernel/internal · high confidence
New internal action models for Delta Log entries
The Delta Kernel API now includes internal Java classes to represent Delta Log actions, including AddFile, AddCDCFile, CommitInfo, DeletionVectorDescriptor, DomainMetadata, Format, Metadata, Protocol, and GenerateIcebergCompatActionUtils. These classes provide schema definitions, serialization/deserialization logic, and utility methods for handling Delta protocol actions such as file additions, commit metadata, deletion vectors, domain metadata, and Iceberg compatibility conversions.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/actions · high confidence
New internal utility classes for Delta Kernel
The kernel's internal utility package now includes several new classes to support core Delta features: CaseInsensitiveMap for case-insensitive key lookups, Clock for testable time operations, ColumnMapping for handling logical-to-physical schema conversions and column ID/name mapping modes, DateTimeConstants for time unit conversions, DirectoryCreationUtils for managing Delta log directories (including staged commits and sidecars), DomainMetadataUtils for validating and populating domain metadata, ExpressionUtils for predicate and collation handling, FileNames for parsing Delta log file patterns and versions, GeometryUtils for parsing WKT geospatial stats, InCommitTimestampUtils for ICT enablement and extraction, InternalUtils for common data and path operations, and IntervalParserUtils for parsing Delta table interval configurations.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/util · high confidence
New kernel statistics APIs for data file stats, snapshot checksums, and table metadata
The kernel now exposes new classes and interfaces in the \io.delta.kernel.statistics\ package to support advanced statistics handling. \DataFileStatistics\ allows connectors to construct, serialize, and deserialize per-file statistics (including min/max values, null counts, and a new \tightBounds\ flag) to/from JSON. \SnapshotStatistics\ provides an interface for snapshots to report checksum write modes, incremental checksum load costs, and aggregate \TableStats\. \TableStats\ exposes aggregate table-level metadata, specifically \tableSizeBytes\ and \numFiles\, which can be computed incrementally from checksum files.
kernel/kernel-api/src/main/java/io/delta/kernel/statistics · high confidence
New logical plan nodes for Delta Lake SQL operations
Added new Catalyst logical plan nodes to support Delta Lake-specific SQL commands and internal operations. This includes \TimeTravel\ and \RestoreTableStatement\ for time-travel queries, \CloneTableStatement\ for table cloning, and \DeltaDelete\/\DeltaUpdateTable\ for row-level modifications. The update also introduces \DeltaMergeAction\ and related traits to handle MERGE INTO logic, including schema evolution behaviors for struct fields. Additionally, new nodes \AlterTableAddConstraint\, \AlterTableDropConstraint\, and \AlterTableDropFeature\ enable constraint and table feature management, while \SyncIdentity\ supports the ALTER TABLE SYNC IDENTITY command. A new \BitmapAggregator\ expression is also added to support bitmap-based aggregation.
spark/src/main/scala/org/apache/spark/sql/catalyst · high confidence
New metrics reporting interfaces for Delta operations
The Delta Kernel API now exposes structured metrics reporting interfaces for Snapshot, Scan, and Transaction operations. Users can access detailed performance data through new report types (SnapshotReport, ScanReport, TransactionReport) and their corresponding metrics result interfaces (SnapshotMetricsResult, ScanMetricsResult, TransactionMetricsResult). These interfaces provide granular timing information (e.g., snapshot construction duration, scan planning time, commit duration), file counts and sizes (add/remove files, total bytes), and operational metadata (table version, schema, filters, clustering columns). All reports implement a common MetricsReport interface with JSON serialization support, enabling monitoring and debugging of Delta table operations.
kernel/kernel-api/src/main/java/io/delta/kernel/metrics · high confidence
New native IncrementMetric and ConditionalIncrementMetric expressions for Delta Lake
Delta Lake introduces two new Catalyst expressions, IncrementMetric and ConditionalIncrementMetric, to replace previous UDF-based approaches for counting rows in DML commands. IncrementMetric unconditionally increments a specified SQLMetric for every row passing through, while ConditionalIncrementMetric only increments the metric when a provided boolean condition evaluates to true. An optimization rule is also included to simplify ConditionalIncrementMetric instances with constant conditions (e.g., converting always-true conditions to IncrementMetric or removing the metric logic entirely for always-false conditions). These changes are currently accessible via the Scala DSL and are designed to provide more accurate and efficient metric tracking during query execution.
spark/src/main/scala/org/apache/spark/sql/delta/metric · high confidence
New schema tracking log for Delta streaming sources
A new \SchemaTrackingLog\ class has been introduced in the Delta streaming package to manage and persist schema evolution for Delta sources. This component serializes partition and data schemas using a versioned JSON format, allowing the system to track the sequence of schema changes and detect concurrent modifications to prevent data inconsistencies during streaming operations.
spark/src/main/scala/org/apache/spark/sql/delta/streaming · high confidence
New telemetry structures and logging abstraction for Delta Lake
This change introduces a new \DeltaLogging\ trait that decouples logging call sites from \DeltaLog\ by using a \DeltaLoggingProvider\ interface, allowing telemetry to be recorded without direct dependencies on the log instance. It also adds a \ScanReport\ case class to capture detailed metrics about data scans, including partition filters, data filters, and data sizes, which enhances observability into scan performance and filter usage.
spark/src/main/scala/org/apache/spark/sql/delta/metering · high confidence
New utility classes for Delta Lake internal operations
This change introduces a suite of new utility classes in the \spark/src/main/scala/org/apache/spark/sql/delta/util\ package to support internal Delta Lake operations. Key additions include \DeltaCommitFileProvider\ to resolve commit file paths for coordinated commits, \DatasetRefCache\ to manage Spark Dataset references across sessions, \BinPackingIterator\ and \BinPackingUtils\ for grouping files by size, and \DeltaEncoders\ for reusable Spark encoders. It also adds \AnalysisHelper\ for expression resolution, \Codec\ for UUID and Base85 encoding, \DeltaFileOperations\ for file system interactions, \DeltaFileSystemOptions\ for Hadoop configuration, \DeltaLogGroupingIterator\ for grouping log files, \DeltaProgressReporter\ for job status reporting, \DeltaSparkPlanUtils\ for logical plan analysis, \DeltaSqlParserUtils\ for SQL parsing, and \DeltaStatsJsonUtils\ for Variant statistics encoding.
spark/src/main/scala/org/apache/spark/sql/delta/util · high confidence
New utility classes for Unity Catalog detection and Parquet format versioning
Added two new utility classes to the Delta Spark connector: CatalogTableUtils, which provides methods to detect whether a table is managed by Unity Catalog by inspecting specific storage properties (delta.feature.catalogManaged, delta.feature.catalogOwned-preview, and the UC table ID), and ParquetFormatVersion, an enum that maps Parquet format versions (V1\_0\_0 and V2\_12\_0) to their corresponding Parquet WriterVersion and supports semantic version resolution. These utilities support the connector's ability to handle catalog-managed tables and configure Parquet output formats correctly.
spark/src/main/java/org/apache/spark/sql/delta/util · high confidence
New utility classes for resource-safe iteration and file metadata in Delta Kernel
The \io.delta.kernel.utils\ package now includes \CloseableIterator\ and \CloseableIterable\ interfaces, which extend standard Java iteration with explicit \close()\ methods to ensure resources are released safely during Delta log and data file processing. Additionally, \FileStatus\ and \DataFileStatus\ classes have been added to encapsulate file metadata, with the latter supporting optional column-level statistics, and \PartitionUtils\ provides a helper to check if a partition contains data.
kernel/kernel-api/src/main/java/io/delta/kernel/utils · high confidence
Support for removing column mapping from Delta tables
Users can now remove column mapping from Delta tables using the new RemoveColumnMappingCommand. This operation rewrites table data to eliminate column mapping metadata, updates the table schema and configuration to disable column mapping, and preserves row tracking information. The command also validates that schema field names are valid before proceeding and prevents execution on catalog-managed tables.
spark/src/main/scala/org/apache/spark/sql/delta/commands/columnmapping · high confidence
Support for writing log compaction files
The Delta Kernel now includes the capability to write log compaction files. A new \LogCompactionWriter\ utility has been added to the \kernel-api\ module, which compacts commit logs between specified versions into a single file. This feature allows the system to reduce the number of individual log files by generating a consolidated checkpoint-like structure for a range of versions, improving efficiency in log management.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/compaction · high confidence
Unified Delta connector introduces Kernel-backed V2 streaming and Change Data Capture support
The new spark-unified module merges the legacy V1 connector and the new Kernel-backed V2 connector into a single published artifact, providing a unified entry point via \DeltaCatalog\ and \DeltaSparkSessionExtension\. This change enables V2 streaming reads for catalog-owned tables by rewriting V1 \StreamingRelation\ plans to \StreamingRelationV2\ backed by \DeltaV2Table\, and adds support for read-time Change Data Capture (CDF) on V2 tables through the \ResolveTableChangesV2\ rule and \ChangelogSupport\ trait (gated by \DELTA\_CHANGELOG\_V2\_ENABLED\). The implementation also includes a Kernel-backed \DeltaV2OptimisticTransaction\ for writes and version-specific shims to ensure compatibility across Spark 4.0, 4.1, and 4.2.
spark-unified · high confidence
Removals
Removal of legacy Delta Lake core implementation files
The legacy core implementation files for the Delta Lake transaction log system have been removed from the source code. This includes the deletion of Checkpoints.scala, Checksum.scala, DeltaConfig.scala, DeltaErrors.scala, DeltaFileFormat.scala, DeltaHistoryManager.scala, DeltaLog.scala, DeltaOperations.scala, and DeltaOptions.scala. These files contained the previous logic for managing checkpoints, checksums, configuration, error handling, file formats, history, logging, operations, and options.
src/main · high confidence
API
New Engine API for connector-provided file system and data handlers
The Delta Kernel now exposes a new \Engine\ interface in the \io.delta.kernel.engine\ package, allowing connectors to provide their own implementations for core operations. This interface aggregates handlers for expression evaluation (\ExpressionHandler\), JSON parsing and writing (\JsonHandler\), Parquet file reading and writing (\ParquetHandler\), and file system access (\FileSystemClient\). Connectors must implement these interfaces to support reading Delta tables, enabling custom logic for file listing, metadata retrieval, atomic file operations, and data format processing.
kernel/kernel-api/src/main/java/io/delta/kernel/engine · high confidence
Architecture
CDC reader refactored into modular base and implementation traits
The Change Data Capture (CDC) reading logic has been restructured to improve maintainability and extensibility. The previous monolithic \CDCReader\ implementation has been split into a new \CDCReaderBase\ trait containing shared functionality (such as version resolution, schema handling, and filter construction) and a \CDCReaderImpl\ trait that extends it. The \CDCReader\ object now extends \CDCReaderImpl\, and the \DeltaCDFRelation\ class has been moved into the implementation layer while \DeltaCDFRelationBase\ handles the common Spark \BaseRelation\ logic. This change isolates the core CDC reading mechanics, making it easier to support different CDC reader implementations or configurations in the future without duplicating code.
spark/src/main/scala/org/apache/spark/sql/delta/commands/cdc · high confidence
Introduce generic FileIO abstraction for Hadoop-based storage operations
The Delta Kernel defaults engine now uses a new, generic \FileIO\ interface to handle all file system interactions, replacing direct Hadoop API usage within the engine. This change introduces a set of new interfaces (\FileIO\, \InputFile\, \OutputFile\, \SeekableInputStream\, \PositionOutputStream\) and a concrete \HadoopFileIO\ implementation that wraps Hadoop's \FileSystem\ and \LogStore\. For users, this provides a more modular and testable I/O layer, enabling the engine to support different storage backends in the future. The implementation also includes optimizations such as passing known file sizes to skip HEAD probes during reads and ensuring atomic file copy and write operations.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/engine/hadoopio · high confidence
Introduces abstract traits for Delta metadata, protocol, and commit info to enable V2 connector interoperability
The Delta Lake Spark connector now provides a set of abstract traits (AbstractMetadata, AbstractProtocol, AbstractCommitInfo) in the v2.interop package. These abstractions allow the V2 connector to reuse existing V1 utilities by implementing adapters that wrap Kernel's action types, ensuring that V1 code depends on these interfaces rather than concrete implementations. This change facilitates code reuse and smoother integration between the V1 and V2 connector codebases.
spark/src/main/scala/org/apache/spark/sql/delta/v2 · high confidence
Introduces comprehensive SBT build infrastructure for multi-version support and CI automation
The build system is restructured with new SBT plugins to enable cross-compilation and publishing for multiple Spark and Flink versions, enforce Java and Scala style checks during compilation and testing, and parallelize test execution across shards and JVMs. This change also centralizes binary compatibility checks via MiMa and manages shaded Iceberg dependencies, providing a unified source of truth for version matrices and release workflows.
project · high confidence
Refactored Delta table command execution into modular command classes
The code in the \spark/src/main/scala/org/apache/spark/sql/delta/commands\ directory has been restructured to replace monolithic command implementations with a modular, trait-based architecture. This change introduces dedicated base classes and traits—such as \CloneTableBase\ for cloning logic, \CreateDeltaTableLike\ for shared table creation utilities, and \DMLWithDeletionVectorsHelper\ for deletion vector handling—which are now used by specific commands like \CloneTableCommand\, \CreateDeltaTableCommand\, and \ConvertToDeltaCommand\. This refactoring centralizes common behaviors (like catalog updates, metric reporting, and transaction handling) into reusable components, improving code maintainability and separation of concerns across Delta Lake's DDL and DML operations.
spark/src/main/scala/org/apache/spark/sql/delta/commands · high confidence
Refactored MERGE execution into modular executor traits
The MERGE command implementation has been restructured into distinct Scala traits to improve code organization and support specialized execution paths. The new \ClassicMergeExecutor\ trait handles the standard two-phase MERGE logic (finding touched files and writing changes), while \InsertOnlyMergeExecutor\ provides an optimized path for MERGE operations that only insert new data, avoiding unnecessary joins. \MergeIntoMaterializeSource\ manages the materialization of the source DataFrame to ensure deterministic reads, including retry logic for RDD block loss. \MergeOutputGeneration\ centralizes the logic for transforming MERGE clauses into output expressions, and \MergeStats\ defines the comprehensive metrics and statistics collected during the operation. This refactoring isolates specific concerns like source materialization, insert-only optimization, and output generation within the \spark/src/main/scala/org/apache/spark/sql/delta/commands/merge\ directory.
spark/src/main/scala/org/apache/spark/sql/delta/commands/merge · high confidence
Refactored data skipping and auto-compaction statistics tracking
The data skipping and auto-compaction logic in the stats module has been restructured for better maintainability and extensibility. The large DataSkippingReaderBase class was extracted into focused components: DataFiltersBuilder handles predicate construction, DataSkippingPredicateBuilder defines the interface for skipping logic, and AutoCompactPartitionStats manages partition-level statistics for auto-compaction. New histogram classes (DeletedRecordCountsHistogram, FileSizeHistogram) were added to track deletion and file size distributions, while DeltaScan and DeltaScanGenerator were introduced to standardize scan result generation. These changes improve code readability and enable more granular control over data skipping and auto-compaction behaviors.
spark/src/main/scala/org/apache/spark/sql/delta/stats · high confidence
Behavioural changes
Add Spark 4.0–4.2 compatibility shims for catalog, expressions, and streaming APIs
This change introduces version-specific shim implementations in spark/src/main/scala-shims to keep Delta Lake compatible with Spark 4.0, 4.1, and 4.2. The shims adapt to API changes across these versions, including: DataSourceV2Relation constructor differences (5 params in 4.0 vs 6 in 4.1+), ParseException handling (separate start/stop origins in 4.0 vs single origin in 4.1+), LogKey trait vs interface differences, QualifiedColType default value representation (String in 4.0 vs DefaultValueExpression in 4.1+), View constructor parameter changes, streaming class relocation (e.g., CheckpointFileManager, StreamExecution moved to new packages in 4.1), geospatial type support (no-op in 4.0, full implementation in 4.1+), variant shredding/stats collection (no-op in 4.0, active in 4.1+), schema evolution propagation rules (no-op in 4.0/4.1, active in 4.2), and table creation overwrite behavior differentiation between V1 and V2 writers.
spark/src/main/scala-shims · high confidence
Added default logging configuration for Scala examples
A new log4j2.properties file has been added to the Scala examples resources, configuring the root logger to output only error-level messages to the console. This ensures that example runs are less verbose by default, suppressing informational and debug logs that might clutter the output for users running the Delta Lake examples.
examples/scala/src/main/resources · high confidence
Added source-compatibility shims for Spark 4.2 API changes
New shim files have been introduced to maintain source compatibility with Spark 4.2, which introduced breaking API changes. The \StreamingRelationV2Shim\ provides an 8-field pattern extractor for \StreamingRelationV2\, bypassing the new 9th parameter (\sourceIdentifyingName\) added in Spark 4.2. Similarly, \CharVarcharShims\ offers extractors for \CharType\ and \VarcharType\ that expose only the length, ignoring the new \collation\ parameter introduced in Spark 4.2, ensuring Delta Lake's internal callsites continue to function without modification.
spark/src/main/scala/org/apache/spark/sql/delta/shims, spark/src/main/scala/org/apache/spark/sql/types · high confidence
Centralized Spark session extension logic for Delta connector
The Delta Spark connector now uses a new \AbstractDeltaSparkSessionExtension\ class to centralize the registration of SQL parsers, optimizer rules, and analysis rules. This change consolidates the extension logic previously scattered across V1 and V2 implementations, ensuring consistent behavior for features like implicit casting, schema evolution, and time travel preprocessing. Users benefit from a unified extension point that supports both legacy (DSv1) and modern (DSv2) connector modes without requiring changes to their existing configurations.
spark/src/main/scala/io/delta/sql · high confidence
Enhanced SBT launcher with proxy support, mirror fallbacks, and reduced memory usage
The SBT build launcher now supports dependency resolution through a configurable Maven proxy via the MAVEN\_PROXY\_URL environment variable, allowing users to route requests through an internal mirror. It also introduces a fallback mechanism for downloading the SBT launch JAR, prioritizing specific mirror URLs (for versions 0.13.18 and 1.5.5) or a Google-hosted Maven Central mirror before falling back to the official Maven repository. Additionally, the default memory allocation for SBT has been reduced from 4000MB to 1000MB to lower resource consumption during builds.
build · high confidence
Enhanced usage tracking and new Spark implicits for Delta Lake
This change introduces two main updates for users. First, it adds implicit classes to the \io.delta.implicits\ package, allowing users to read, write, and stream Delta tables using concise syntax like \spark.read.delta(path)\ and \df.write.delta(path)\. Second, it significantly expands the \DatabricksLogging\ utility to support detailed usage metrics; the \recordUsage\ and \recordEvent\ methods now capture data into a trackable \UsageRecord\ structure, and new methods \recordProductUsage\ and \recordProductEvent\ are added to support product-specific telemetry, with \recordEvent\ now defaulting to tracking operation duration.
spark/src/main/resources/META-INF, spark/src/main/scala/com, spark/src/main/scala/io/delta/implicits · high confidence
Expanded SQL grammar for Delta Lake table management commands
The Delta SQL parser grammar has been updated to support a broader range of table management operations. Users can now use SQL syntax to drop table features (\ALTER TABLE ... DROP FEATURE\), manage clustering (\ALTER TABLE ... CLUSTER BY\ and \CLUSTER BY NONE\), and synchronize identity columns (\ALTER TABLE ... SYNC IDENTITY\). Additionally, the parser now handles \OPTIMIZE ... FULL\ for clustered tables, \REORG TABLE\ commands for purging or upgrading to Iceberg compatibility, and string coalescing in identifiers.
spark/src/main/antlr4 · high confidence
Iceberg library upgraded to 1.11.0 with Delta UniForm compatibility patches
The shaded Iceberg library has been upgraded to version 1.11.0. To support Delta UniForm, the integration includes specific behavioral changes: the \Delegates\ class now preserves the explicit \first\_row\_id\ on data files instead of suppressing it, ensuring Delta row IDs are maintained during conversion. Additionally, \PartitionSpec\ validation conflicts are suppressed by default when converting from Delta to honor Delta-assigned field IDs, and deprecated APIs required for schema management have been restored in \MetadataUpdate\ and \TableMetadata\. The \HiveCatalog\ and \HiveTableOperations\ have been modified to accept and apply \MetadataUpdate\ lists, allowing the use of schema and partition specs with field IDs assigned by Delta Lake, and \HiveTableOperations\ now correctly handles \NoSuchIcebergTableException\ to facilitate table creation transactions.
icebergShaded/src/main/java/org/apache/iceberg · high confidence
Improved error reporting for empty schema writes and constraint violations
Delta Lake now provides more specific and actionable error messages when writing data that results in an empty schema (e.g., all columns are NullType/VOID or structs have no fields) or when data violates table constraints. The new \EmptySchemaWriteException\ utility distinguishes between tables with no columns, tables with all VOID columns, and structs with no fields, throwing distinct errors for each case. Additionally, \InvariantViolationException\ handling has been refined to include specific error classes for NOT NULL and CHECK constraint violations, ensuring users receive clear feedback on which constraint was violated and the values involved.
spark/src/main/scala/org/apache/spark/sql/delta/schema · high confidence
Introduce DefaultFileSystemManagedTableOnlyCommitter with catalog-managed table protection
The kernel now includes a new \DefaultFileSystemManagedTableOnlyCommitter\ implementation that handles commit logic for file-system managed tables. This component validates protocol versions before writing the JSON commit file atomically and explicitly rejects attempts to use this committer on catalog-managed tables by throwing an error if such a protocol is detected. It also exposes a protected constructor, allowing subclasses to extend its behavior while ensuring the default instance remains restricted to file-system managed scenarios.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/commit · high confidence
Introduce Flink Sink V2 with upsert and merge-on-read support
The Delta Flink connector now implements the Flink Sink V2 API, replacing the previous sink implementation. This new architecture introduces a two-phase commit protocol with dedicated writer and committer components, enabling exactly-once semantics and robust checkpointing. Users can now configure the sink to operate in 'upsert' mode, which merges incoming changes into the table based on a declared primary key using a merge-on-read strategy, or stick to the default 'append' mode. The update also includes new configuration options for file rolling strategies (by size or count), schema evolution policies, and primary key definitions, along with comprehensive data type conversions between Flink and Delta structures.
flink/src/main/java/io/delta/flink/sink · high confidence
Introduce ZCube metadata tracking for Z-ORDER files
Added ZCubeInfo to track and store metadata (ZCubeID and Z-ORDER columns) directly in file tags for files written by OPTIMIZE Z-ORDER BY commands. This enables the system to identify which Z-ORDER optimization run produced each file, supporting incremental clustering by allowing the system to distinguish between files from different Z-ORDER operations.
spark/src/main/scala/org/apache/spark/sql/delta/zorder · high confidence
Migrate LogStore implementations to Java-based Delta Storage classes
The Delta Spark module now delegates file system operations for S3, Azure, and HDFS to the corresponding Java-based \io.delta.storage\ classes (e.g., \S3SingleDriverLogStore\, \AzureLogStore\, \HDFSLogStore\). The \DelegatingLogStore\ in this package has been updated to resolve and instantiate these Java implementations by file system scheme, replacing the previous Scala-only defaults. This change shifts the core storage logic to the external Delta Storage library while maintaining the \LogStore\ interface for Delta Lake operations.
spark/src/main/scala/org/apache/spark/sql/delta/storage · high confidence
New constraint validation engine and char/varchar length enforcement
Delta now validates NOT NULL and CHECK constraints at write time using a new physical operator (DeltaInvariantCheckerExec) and expression (CheckDeltaInvariant), providing detailed error messages that include the violating column names and values. The system also enforces char/varchar string length limits via a new CharVarcharConstraint module, and supports adding or dropping constraints through the TableChange API (AddConstraint/DropConstraint).
spark/src/main/scala/org/apache/spark/sql/delta/constraints · high confidence
New internal data structures for row and column vector handling in the Kernel API
The \kernel/kernel-api\ module introduces a suite of new internal classes to manage row and columnar data representations. \ChildVectorBasedRow\ and its subclasses (\ColumnarBatchRow\, \StructRow\) provide abstractions for accessing data from column vectors by row ID, while \GenericRow\ and \GenericColumnVector\ offer flexible implementations wrapping maps and lists of values. Additionally, \DelegateRow\ allows creating modified views of existing rows without mutation, and \SelectionColumnVector\ handles boolean selection masks. Specific state containers \ScanStateRow\ and \TransactionStateRow\ encapsulate table metadata, schemas, and protocol information for scan and transaction operations, providing utility methods to extract logical/physical schemas, partition columns, and configuration.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/data · high confidence
New internal utility classes for error handling and timestamp conversion
The default kernel implementation introduces two new internal classes: \DefaultEngineErrors\ and \DefaultKernelUtils\. \DefaultEngineErrors\ provides factory methods for specific exceptions, including support for invalid escape sequences in LIKE expressions and errors for unsupported expression evaluation. \DefaultKernelUtils\ adds utility methods for converting between Julian days and microseconds, parsing TimestampNTZ strings, and resolving column data types within a schema, supporting more robust timestamp handling and expression evaluation in the default engine.
kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal · high confidence
Parallelized Delta operations with thread-local context propagation
Delta Lake now executes listFrom and commitStore getCommits calls in parallel using a new thread pool infrastructure. This change introduces a custom thread pool that captures and forwards Spark thread-local properties (such as TaskContext and local properties) to worker threads, ensuring that parallel tasks run with the correct caller's SparkSession context. It also includes a NonFateSharingFuture mechanism to prevent error propagation across multiple waiting threads, improving reliability and performance for concurrent Delta operations.
spark/src/main/scala/org/apache/spark/sql/delta/util/threads · high confidence
Preserve Variant stats during checkpointing and state reconstruction
Added version-specific shim implementations for Variant stats in Spark 4.0 and 4.1+ to ensure correct handling of variant metadata during checkpointing and state reconstruction. The Spark 4.0 shim provides a custom implementation for reading unsigned integers because \VariantUtil.readUnsigned\ is private, while the Spark 4.1+ shim leverages the now-public \VariantUtil.readUnsigned\ method directly.
spark/src/main/java-shims · high confidence
Python client API restructure and exception handling improvements
The Python client package has been restructured to expose the Delta Lake version via \delta.\_\version\\_\ and provide a \configure\_spark\_with\_delta\_pip\ utility for easier local installation. Concurrent modification errors are now raised as distinct, typed Python exceptions (e.g., \DeltaConcurrentModificationException\) rather than generic errors, and the \DeltaTable.merge\ API now returns a \DataFrame\ to allow for chaining and inspection of the merge results.
python/delta · high confidence
Refactor DeltaTable execution layer to use new command-based operations
The \spark/src/main/scala/io/delta/tables/execution\ package has been restructured to decouple the Scala \DeltaTable\ API from its underlying Spark SQL command implementations. New files introduce \DeltaTableOperations\ as the central trait executing operations like \delete\, \update\, \vacuum\, \restore\, and \clone\ via specific logical plan commands (e.g., \VacuumTableCommand\, \CloneTableStatement\). This includes adding \DeltaConvert\ for converting external tables to Delta and \DeltaTableBuilderOptions\ for create/replace semantics. The \VacuumTableCommand\ specifically supports external inventory sources for garbage collection. This change standardizes how DeltaTable API calls are translated into Spark execution plans, improving modularity and enabling features like shallow cloning and inventory-based vacuuming.
spark/src/main/scala/io/delta/tables/execution · high confidence
Refactor Kernel exceptions into a dedicated hierarchy
The Delta Kernel API now provides a structured set of specific exception classes under \io.delta.kernel.exceptions\ to replace generic error handling. This change introduces a base \KernelException\ and granular types for common scenarios, including \TableNotFoundException\, \TableAlreadyExistsException\, \InvalidTableException\, and \ConcurrentWriteException\ (with subclasses like \ConcurrentTransactionException\ and \MetadataChangedException\). It also adds new exceptions for versioning and protocol issues, such as \StartVersionNotFoundException\, \CommitRangeNotFoundException\, \UnsupportedProtocolVersionException\, and \UnsupportedTableFeatureException\, along with \CommitStateUnknownException\ to handle ambiguous commit states during retries. This allows users to catch and handle specific failure modes more precisely.
kernel/kernel-api/src/main/java/io/delta/kernel/exceptions · high confidence
Refactor streaming snapshot initialization to use incremental loading
The DeltaSourceSnapshot implementation has been refactored to support incremental loading for streaming queries. Instead of caching the entire set of initial files, the code now filters files based on partition and data predicates immediately and caches the resulting filtered dataset. This change optimizes the startup process for new streaming queries by reducing the amount of data processed and cached during the initial snapshot phase.
spark/src/main/scala/org/apache/spark/sql/delta/files · high confidence
Refactored CONVERT TO DELTA file listing and schema inference
The internal implementation of the CONVERT TO DELTA command has been restructured to improve performance and reliability. A new \ConvertUtils\ module centralizes utility logic, while file manifest handling is split into specialized classes (\ManualListingFileManifest\, \CatalogFileManifest\) that support rebalancing file listings across tasks to prevent skewed processing during schema inference. Additionally, the \ParquetTable\ implementation now exposes \numFiles\ and \sizeInBytes\ properties, providing users with accurate metadata about the source table during the conversion process.
spark/src/main/scala/org/apache/spark/sql/delta/commands/convert · high confidence
Refactored Delta log file parsing into a typed class hierarchy
The internal log-file parsing logic in the Kernel API has been restructured from a monolithic approach into a typed class hierarchy under \io.delta.kernel.internal.files\. New classes now represent specific Delta log components—\ParsedPublishedDeltaData\ and \ParsedCatalogCommitData\ for commits, \ParsedClassicCheckpointData\, \ParsedV2CheckpointData\, and \ParsedMultiPartCheckpointData\ for checkpoints, plus \ParsedLogCompactionData\ and \ParsedChecksumData\—all extending a common \ParsedLogData\ base. This change introduces explicit type prioritization for checkpoints (V2 \> MultiPart \> Classic) and adds utility methods in \LogDataUtils\ to validate log data ordering and merge published versus ratified commits, ensuring the kernel correctly handles and prioritizes different log file types.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/files · high confidence
Refactored Delta log replay to use columnar iterators and pagination support
The log replay mechanism in the Delta Kernel has been rewritten to use a columnar, iterator-based design for improved performance and memory efficiency. This change introduces new internal classes such as \ActionsIterator\, \ActiveAddFilesIterator\, and \CreateCheckpointIterator\ to process Delta log actions (adds, removes, metadata) in batches. It also adds support for paginated scans via the \PageToken\ class, allowing large log segments to be read in chunks. Additionally, the \ConflictChecker\ has been updated to handle concurrent delete-vs-delete conflicts and row tracking metadata during conflict resolution.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/replay · high confidence
Refactored snapshot construction with intent-based builders and catalog version support
The internal snapshot building process has been restructured to use a new intent-based builder pattern (SnapshotBuilderImpl) and a dedicated factory (SnapshotFactory). This change introduces support for catalog-managed tables by allowing users to specify a maximum catalog version, which is enforced during snapshot resolution to prevent time-travel beyond allowed limits. The refactoring also improves timestamp-based time travel by correctly handling catalog commits and reading the \_last\_checkpoint file for latest queries, ensuring consistent snapshot state across different table management modes.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/table · high confidence
Refactored snapshot construction with new LogSegment and MetadataCleanup components
The snapshot loading logic in the kernel has been restructured to improve clarity and maintainability. A new \LogSegment\ class now encapsulates the validation and representation of the transaction log files (deltas, checkpoints, and compactions) required to build a snapshot, replacing previous inline logic. \SnapshotManager\ has been updated to utilize this new structure, delegating log segment generation and validation to \LogSegment\. Additionally, a new \MetadataCleanup\ class has been introduced to handle the deletion of expired delta log files based on retention policies, separating this concern from the snapshot building process. These changes support features like time-travel by timestamp and version-specific snapshot retrieval while ensuring robust validation of the Delta log state.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/snapshot · high confidence
Standardized structured logging keys for Delta Lake
Delta Lake now defines a centralized set of keys for Mapped Diagnostic Context (MDC) in its structured logging system via the new \DeltaLogKeys\ trait. This change standardizes how log data is tagged across Delta operations, ensuring consistent and parseable log output for users relying on structured logging frameworks.
spark/src/main/scala/org/apache/spark/sql/delta/logging · high confidence
Structured error classes for Delta concurrent modification exceptions
The Delta Spark library now exposes a dedicated set of exception classes in the \io.delta.exceptions\ package for handling concurrent modification conflicts. These include \ConcurrentWriteException\, \MetadataChangedException\, \ProtocolChangedException\, \ConcurrentAppendException\, and \ConcurrentDeleteReadException\. Each class implements the \DeltaThrowable\ interface, providing structured error codes (such as \DELTA\_CONCURRENT\_WRITE\ and \DELTA\_METADATA\_CHANGED\) and message parameters, allowing applications to programmatically detect and handle specific types of commit conflicts rather than relying on generic exception messages.
spark/src/main/scala/io/delta/exceptions · high confidence
Support for ALTER TABLE CLUSTER BY and clustering column management
Users can now modify the clustering configuration of existing Delta tables using the ALTER TABLE ... CLUSTER BY command, allowing them to add, change, or remove clustering columns without recreating the table. This change introduces the necessary logical plan and table change structures to handle clustering specifications, enabling dynamic updates to how data is clustered for query optimization.
spark/src/main/scala/org/apache/spark/sql/delta/skipping/clustering/temp · high confidence
Unified table feature framework in Delta Kernel
The Delta Kernel now uses a structured, unified framework for managing table features, replacing the previous ad-hoc logic. This change introduces a base \TableFeature\ class and a \FeatureAutoEnabledByMetadata\ interface, allowing features like Change Data Feed, Column Mapping, and Invariants to be explicitly defined with their protocol requirements (reader/writer versions) and metadata triggers. This provides a consistent mechanism for the Kernel to determine feature support, handle legacy features, and automatically enable features based on table properties.
kernel/kernel-api/src/main/java/io/delta/kernel/internal/tablefeatures · high confidence
Test coverage
Add integration test for Delta Uniform Hudi support; Added DefaultKernelTestUtils for test data extraction; Added Delta Kernel benchmark suite with JMH infrastructure and workload models; Added Delta log fixtures for identity columns and deletion vectors; Added GoldenTables test infrastructure for generating reference table states; Added Hive Metastore schema resources for embedded HMS tests; Added Java unit tests for Delta Lake SQL commands and deletion vector encodings; Added comprehensive test suites for Delta Format Sharing; Added golden table test fixtures for decimal decoding, iterator handling, and transactional operations; Added integration tests for DynamoDB Commit Coordinator; Added integration tests for Flink Delta Lake features; Added integration tests for Unity Catalog commit coordinator and managed tables; Added log4j 2.x test configuration files; Added parser tests for Delta SQL commands; Added shared test coverage for Delta table refresh, caching, and self-join behaviors; Added test coverage for Delta Kernel Parquet I/O; Added test coverage for Delta Kernel checksum and checkpoint features; Added test coverage for Delta Kernel type system features; Added test coverage for Flink SQL Delta sink and Unity Catalog integration; Added test coverage for Kernel expressions and predicates; Added test coverage for commit metadata, default committer, and publish metadata validation; Added test coverage for the Flink Delta Sink implementation; Added test coverage for the new Flink Kernel-backed table implementation; Added test fixtures for deletion vectors with and without checkpoints; Added test infrastructure for Delta Connect; Added test infrastructure for Unity Catalog and Delta-Tables API integration; Added test logging configuration files for Flink; Added test logging configuration for structured logging support; Added test shims for Spark 4.0 and 4.1+ streaming classes; Added test suite for Delta Kernel table feature validation; Added test suites for Delta Kernel internal action classes; Added test suites for DeltaTable builder, forName resolution, and core table operations; Added test suites for Iceberg compatibility and column defaults validation; Added test suites for kernel expression evaluation; Added test suites for the Default Engine's core components; Added test to audit delta-iceberg JAR contents; Added test utilities for building columnar batches and rows; Added test utilities for configuring Delta's Spark extensions; Added tests and benchmarks for the Path utility class; Added tests for ActionsIterator resource cleanup and interrupt handling; Added tests for CRCInfo serialization, backward compatibility, and checksum state tracking; Added tests for CatalogCommitterUtils.extractProtocolProperties; Added tests for CloseableIterator utilities and Transaction partition handling; Added tests for ClusteringMetadataDomain serialization and column handling; Added tests for DataLayoutSpec in transaction module; Added tests for DataType JSON serialization and Type Widening logic; Added tests for Deletion Vectors Base85 codec and RoaringBitmapArray; Added tests for Delta Connect Python client exception handling and API coverage; Added tests for Delta Kernel checkpoint instance logic and checkpointer behavior; Added tests for Delta Kernel log data parsing and validation utilities; Added tests for Delta SQL extension and catalog configuration; Added tests for Delta concurrent exception handling; Added tests for Delta-to-Hudi conversion; Added tests for FileSizeHistogram utility; Added tests for Flink kernel checkpointing and deletion vector access; Added tests for Hadoop input and output file handling; Added tests for JSON serialization and deserialization utilities; Added tests for JsonMetadataDomain serialization and deserialization; Added tests for Kernel metrics reporting and collection; Added tests for Kernel metrics serialization and timing utilities; Added tests for LogStore provider and listFrom behavior; Added tests for catalog-managed log segment construction and snapshot builder validation; Added tests for catalog-managed table E2E reads and property validation; Added tests for clustering column info resolution; Added tests for data skipping utilities and stats schema handling; Added tests for published artifact naming conventions across Spark and Flink versions; Added unit and end-to-end tests for the Unity Catalog kernel integration; Added unit tests for Counter and StringType in kernel-api; Added unit tests for Delta Kernel internal components; Added unit tests for Flink upsert merge strategies and scan locator; Added unit tests for Kernel exception handling and error classes; Added unit tests for LogSegment validation and metadata cleanup logic; Added unit tests for kernel internal utility classes; Expanded Unity Catalog integration test coverage for Spark; Expanded test coverage for Delta Lake Spark integration; Initial Python test suite for Delta Lake; New Python test infrastructure and configuration for Delta Lake; New test utilities and mock infrastructure for Delta Kernel; New test utilities for Delta Kernel defaults; Removed Delta test suites; Test shim infrastructure for Spark 4.0, 4.1, and 4.2 compatibility.
Dependencies
Configures additional Maven and SBT repository mirrors for dependency resolution
The build system now explicitly defines a prioritized list of artifact repositories in the new \build/sbt-config/repositories\ file. This configuration adds a Google-hosted mirror of Maven Central as the primary remote source to reduce load on the central repository, includes MuleSoft as a fallback, and integrates specific repositories for Typesafe, Spark packages, and Apache snapshots. This change ensures more reliable and faster dependency resolution by distributing requests across multiple mirrors and handling specific plugin repositories.
build/sbt-config · high confidence
New build manifests for benchmarks, dev tools, docs, and examples
Added build configuration files for several project areas: benchmarks (Scala 2.12.18, Spark 3.5.3, scopt 4.0.1, jackson-module-scala 2.13.1), dev tooling (mypy 1.8.0, flake8 3.9.0, black 23.12.1, grpcio \>=1.67.0, mypy-protobuf 3.3.0), documentation site (Node 22.18.0, Astro 5.15.9, Starlight 0.36.2, TypeScript 5.9.3, ESLint 9.39.1, Prettier 3.6.2), and Scala examples (Scala 2.13.17, Iceberg 1.4.1, Jackson 2.15.4, with dynamic Spark/Delta version resolution). Also added a Maven POM for kernel examples (delta-kernel 3.2.0-SNAPSHOT, Hadoop 3.3.1, Jackson 2.13.5).
(dependencies) · high confidence
Housekeeping
Added placeholder for Flink Java source directory
A placeholder file was added to the Flink Java source directory to ensure the directory structure is recognized by the build system.
flink/src/main/java · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 48 → 72 (+23.8)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 83 → 84 (+0.9)
- Architecture 94 → 99 (+4.9)
- Maturity 54 → 71 (+16.7)
- Readiness 27 → 67 (+39.7)
- Security 73 → 75 (+2.2)
Resolved (110)
- Coverage not measured — test suite did not build
- Critical IaC: AWS-0104 (benchmarks/infrastructure/aws/terraform/modules/processing/main.tf)
- Critical IaC: AWS-0104 (benchmarks/infrastructure/aws/terraform/modules/processing/main.tf)
- Critical IaC: AWS-0104 (benchmarks/infrastructure/aws/terraform/modules/processing/main.tf)
- Dimension evaluation failed
- Duplicated block (103 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/actions/GenerateIcebergCompatActionUtils.java)
- Duplicated block (12 lines × 7) (kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/parquet/ParquetColumnReaders.java)
- Duplicated block (121 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/stats/FileSizeHistogram.java)
- Duplicated block (13 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/DeltaErrors.java)
- Duplicated block (14 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/types/DataTypeJsonSerDe.java)
- Duplicated block (15 lines × 2) (kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/expressions/DefaultExpressionUtils.java)
- Duplicated block (15 lines × 3) (kernel/examples/kernel-examples/src/main/java/io/delta/kernel/examples/BaseTableWriter.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (16 lines × 4) (kernel/kernel-api/src/main/java/io/delta/kernel/types/FieldMetadata.java)
- Duplicated block (162 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/util/ColumnMapping.java)
- Duplicated block (164 lines × 2) (kernel/kernel-api/src/main/java/io/delta/kernel/statistics/DataFileStatistics.java)
- …and 90 more
New (829)
- 1.writeNextFile (cognitive 16) (kernel/kernel-defaults/src/main/java/io/delta/kernel/defaults/internal/parquet/ParquetFileWriter.java)
- AbstractDeltaCatalog.alterTable (cognitive 46) (spark/src/main/scala/org/apache/spark/sql/delta/catalog/AbstractDeltaCatalog.scala)
- AbstractDeltaCatalog.alterTable (cyclomatic 40) (spark/src/main/scala/org/apache/spark/sql/delta/catalog/AbstractDeltaCatalog.scala)
- AbstractDeltaCatalog.createDeltaTable (cognitive 31) (spark/src/main/scala/org/apache/spark/sql/delta/catalog/AbstractDeltaCatalog.scala)
- AbstractDeltaCatalog.createDeltaTable (cyclomatic 31) (spark/src/main/scala/org/apache/spark/sql/delta/catalog/AbstractDeltaCatalog.scala)
- AbstractDeltaCatalog.loadTable (cognitive 16) (spark/src/main/scala/org/apache/spark/sql/delta/catalog/AbstractDeltaCatalog.scala)
- ActiveAddFilesIterator.prepareNext (cognitive 27) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/replay/ActiveAddFilesIterator.java)
- ActiveAddFilesIterator.prepareNext (cyclomatic 16) (kernel/kernel-api/src/main/java/io/delta/kernel/internal/replay/ActiveAddFilesIterator.java)
- AlterTableChangeColumnDeltaCommand.run (cognitive 21) (spark/src/main/scala/org/apache/spark/sql/delta/commands/alterDeltaTableCommands.scala)
- AutoCompactUtils.reserveTablePartitions (cognitive 24) (spark/src/main/scala/org/apache/spark/sql/delta/hooks/AutoCompactUtils.scala)
- AutoCompactUtils.reserveTablePartitions (cyclomatic 22) (spark/src/main/scala/org/apache/spark/sql/delta/hooks/AutoCompactUtils.scala)
- BaseExternalLogStore.write (cognitive 21) (storage-s3-dynamodb/src/main/java/io/delta/storage/BaseExternalLogStore.java)
- BitPacking.packBits2 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- BitPacking.packBits3 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- BitPacking.packBits4 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- BitPacking.packBits5 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- BitPacking.packBits6 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- BitPacking.packBits7 (cognitive 16) (spark/src/main/java/org/apache/spark/sql/delta/deletionvectors/mumbling/BitPacking.java)
- CDCReaderImpl.changesToDF (cognitive 27) (spark/src/main/scala/org/apache/spark/sql/delta/commands/cdc/CDCReader.scala)
- CDCReaderImpl.changesToDF (cyclomatic 18) (spark/src/main/scala/org/apache/spark/sql/delta/commands/cdc/CDCReader.scala)
- …and 809 more
Changes since last survey
- 189 commits — 179 feature/other, 10 fixes
By area
- spark/src — 122 commits
- spark/v2 — 20 commits
- spark-unified/src — 17 commits
- flink/src — 7 commits
- kernel/kernel-api — 4 commits
- protocol_rfcs/iceberg-v4-metadata.md — 4 commits
- (root) — 3 commits
- kernel/kernel-defaults — 3 commits
- protocol_rfcs/README.md — 3 commits
- .devcontainer/devcontainer.json — 1 commit
- .github/workflows — 1 commit
- build/sbt-config — 1 commit
- iceberg/src — 1 commit
- sharing/src — 1 commit
- storage/src — 1 commit
Notable commits
- fix: Fix DataEntry tracking and CDF on the incremental AMT path (#7481)
- fix: Fix DataManifestEntry tracking.status across AMT write paths (#7443)
- fix: Fix clusteringColumns example in PROTOCOL.md to use path arrays (#7531)
- fix: Revert "[Kernel] Promote Snapshot/Scan/CommitRange accessors used by Delta Spark (#7269)" (#7472)
- fix: Revert "[Spark] DeltaV2OptimisticTransaction: Kernel-backed blind-append commit path (#7519)" (#7541)
- fix: [Kernel] Fix kernel corrupt checkpoint bug (#7221)
- fix: [Spark] Fix flaky AMT partition values test (#7487)
- fix: [Spark] Fix flaky identity column streaming test (#7488)
- fix: [Spark] Revert staged commit filesystem filtering (#7616)
- fix: [Storage] Fix S3 fast listing to exclude nested keys (#7451)
- change: AMT conflict resolution: winning AMT commit vs losing log or inline-tree commit (#7609)
- change: Abstract IncrementalAMTWriter old-AMT inputs behind an actions provider (#7607)
- change: Add AMT conflict-resolution round metrics (#7694)
- change: Add DeltaTableProvider test trait to parameterize table format in test suites (#7468)
- change: Add Kernel-backed concurrent-writer conflict checking for DeltaV2OptimisticTransaction (#7628)
- change: Add dataChange as a top-level CommitInfo field (#7500)
- change: Add kernel factory for constructing Default Engine (#7384)
- change: Add query context for Delta V2 snapshot operations (#7702)
- change: Add query-context Delta V2 snapshot manager APIs (#7703)
- change: Add read-only Mumbling bitmap reader and PFOR codec (#7715)
- …and 169 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
delta-io/delta was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit ca2e1952187f05936045dd7c5d6f2c8c97926ca7 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.