apache/datafusion
72.7
Strong · 29 September 2026
744.6k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a modular, extensible query engine built in Rust that processes SQL and DataFrame APIs against diverse data sources. It provides a comprehensive execution runtime featuring a logical and physical optimizer, support for standard file formats like Parquet and CSV, and a rich library of scalar, aggregate, and window functions. The architecture emphasizes separation of concerns through dedicated crates for cataloging, expression evaluation, and protocol serialization, while offering extensive benchmarking and testing infrastructure to validate performance and correctness.
How it got here
2016–2023 — modularization and benchmarking infrastructure
62 changes.
This period focused on restructuring the DataFusion codebase into a modular architecture, extracting core components like the SQL planner, physical expressions, and execution runtime into dedicated crates. Concurrently, the project established a comprehensive benchmarking infrastructure with standardized suites for TPC-H and TPC-DS, alongside extensive integration and fuzz testing to ensure stability across the new modular design.
2024 — modularization and FFI integration
48 changes.
This period focused on restructuring the codebase by extracting shared logic, physical optimizers, and catalog implementations into dedicated crates to improve modularity. It also introduced a new Foreign Function Interface (FFI) crate to enable cross-library interoperability and dynamic module loading. Concurrently, the project expanded its function library with new array, window, and table functions, while significantly increasing test coverage and benchmarking infrastructure.
2025 — modularization and spark compatibility
49 changes.
This period focused on decoupling the codebase by extracting core components like catalogs, datasources, and pruning logic into dedicated crates to improve modularity and reduce compile times. Simultaneously, significant effort was directed toward expanding Spark compatibility through a new dedicated crate and extensive function implementations. The work was supported by comprehensive testing, benchmarking, and the consolidation of example code to ensure stability across these architectural and functional changes.
2026 — SQL benchmark suite expansion and infrastructure
30 changes.
This period focused on establishing a comprehensive SQL benchmarking infrastructure by adding standardized suites for TPC-H, TPC-DS, ClickBench, and H2O, alongside specialized benchmarks for join algorithms and Parquet pruning. Concurrently, the project refactored core file scanning components to support dynamic work stealing and improved sort pushdown, while introducing a new crate for shared protobuf schemas to streamline serialization.
Features
Add ClickBench benchmark queries and documentation
The benchmarks/queries/clickbench directory now includes a README documenting the standard ClickBench queries alongside a new set of "Extended" queries designed to test DataFusion-specific features such as high-cardinality aggregations, MEDIAN, APPROX\_PERCENTILE\_CONT, STDDEV, VAR, and complex string filtering. A shell script (update\_queries.sh) is provided to automatically download and split the official ClickBench SQL file into individual query files, ensuring the benchmark suite stays synchronized with the upstream ClickHouse repository.
benchmarks/queries/clickbench · high confidence
Add ClickBench sorted SQL benchmark suite
A new benchmark suite named 'clickbench\_sorted' has been added to the SQL benchmarks, allowing users to run ClickBench queries against a pre-sorted dataset. The suite includes a configuration file that exposes command-line arguments to select the sort column (defaulting to 'EventTime') and sort order (defaulting to 'ASC'), enabling users to test performance with different sorting configurations. The benchmark setup configures DataFusion to leverage existing sorts via the 'prefer\_existing\_sort' optimizer setting and creates an external table pointing to a sorted Parquet file, with an initial query example provided to verify data loading.
(repo-wide) · high confidence
Add DataFusion WebAssembly (WASM) test harness and demo application
Introduces a new \wasmtest\ crate and a companion \datafusion-wasm-app\ demo to verify that DataFusion compiles and runs correctly on the \wasm32-unknown-unknown\ target. The test crate provides a suite of \wasm-pack\ tests that validate core functionality—including expression evaluation, SQL parsing, and query execution—within a browser environment. The included demo application demonstrates these capabilities by invoking DataFusion and logging results to the browser console, while the updated README documents the setup process for building and testing on various platforms, including Apple Silicon.
datafusion/wasmtest · high confidence
Add H2O SQL benchmark suite for group-by, join, and window operations
The \benchmarks/sql\_benchmarks/h2o\ directory now contains a complete benchmark suite for the H2O dataset, covering group-by, join, and window query subgroups. The suite includes a \h2o.suite\ configuration file that allows users to run benchmarks via the \benchmark\_runner\ binary, with support for varying dataset sizes (small, medium, big) and file formats (CSV, Parquet). It provides specific SQL queries for group-by aggregations, various join types (inner, left, etc.), and window functions (including Top-N, moving averages, and rank-based operations), along with the necessary initialization scripts to load the data.
_benchmarks/sql\benchmarks/h2o · high confidence
Add H2O.ai Database-like Ops benchmark suite
Added a new set of SQL benchmark queries in the \benchmarks/queries/h2o\ directory, covering GROUP BY aggregations, JOINs, and complex window functions (including Top-N with ROW\_NUMBER and RANK). These queries are designed to evaluate performance on database-like operations such as partitioned aggregations, moving averages, and handling of ties in ranking functions.
benchmarks/queries/h2o · high confidence
Add IMDB (JOB) benchmark with data conversion and query execution utilities
Users can now run the Join Order Benchmark (JOB) using the IMDB dataset. This change introduces a new benchmark module in \benchmarks/src/imdb\ that includes a \convert\ utility to transform CSV source data into Parquet or CSV formats, and a \run\ utility to execute 113 benchmark queries against the dataset. The implementation provides command-line options to specify input/output paths, file formats, batch sizes, and execution modes such as loading data into a MemTable or preferring hash joins.
benchmarks/src/imdb · high confidence
Add SQL-to-logical-plan example
A new example file \datafusion/sql/examples/sql.rs\ demonstrates how to parse SQL statements and convert them into logical query plans using DataFusion. It provides a concrete implementation of the \ContextProvider\ trait (using \MyContextProvider\) to define schema tables (\customer\, \state\, \orders\) and register aggregate UDAFs (\sum\, \count\), showing users how to wire up the SQL parser, \SqlToRel\ planner, and custom context for query execution.
datafusion/sql/examples · high confidence
Add TPC-DS SQL benchmark suite
Added 99 benchmark definition files (q01–q99) for the TPC-DS workload, along with SQL scripts to load and clean up the required external Parquet tables. These files enable running standardized TPC-DS queries against the benchmark engine.
_benchmarks/sql\benchmarks/tpcds · high confidence
Add TPC-H SQL benchmark suite with configurable data formats
The TPC-H benchmark suite is now available in the SQL benchmarks directory, providing 22 standard queries (q01–q22) that can be executed via the benchmark runner. Users can configure the data source format using the \--format\ flag (supporting parquet, csv, and mem modes) and adjust the dataset scale factor with \--scale-factor\. The suite includes initialization scripts to load data from the specified \DATA\_DIR\, run the queries, and clean up tables, with results saved to CSV files.
_benchmarks/sql\benchmarks/tpch · high confidence
Add TPC-H benchmark queries 1–22
The benchmarks/queries directory now includes the complete set of TPC-H benchmark queries (q1 through q22). These SQL files provide standardized workloads for evaluating query performance against the TPC-H schema, covering complex operations such as aggregations, joins, subqueries, and windowing functions.
benchmarks/queries · high confidence
Add TPC-H sorting benchmark suite
A new benchmark suite named \sort\_tpch\ has been added to evaluate SQL sorting performance using the TPC-H \lineitem\ table. The suite includes 11 distinct queries (Q01–Q11) that test various sorting characteristics, including different key types (INTEGER, BIGINT, VARCHAR), cardinalities, the number of sort keys (1 to 4), and payload column counts (thin vs. wide). Users can run these benchmarks via the \benchmark\_runner\ binary, with options to select the TPC-H scale factor and control whether the data is pre-sorted.
_benchmarks/sql\_benchmarks/sort\tpch · high confidence
Add benchmarks for StringView and BinaryView column spilling
Added a new benchmark suite in \benchmarks/sql\_benchmarks/spill\_views\ that measures the performance of spilling StringView and BinaryView columns during sort and aggregation operations. The suite includes load scripts for distinct and low-cardinality datasets, SQL queries for sorting and grouping, and configuration files that allow memory limits to be overridden via environment variables (\SPILL\_VIEWS\_LIMIT\_REPEATED\ and \SPILL\_VIEWS\_LIMIT\_DISTINCT\).
_benchmarks/sql\_benchmarks/spill\views · high confidence
Add build environment script and historical changelogs
The \dev/build-set-env.sh\ script was added to automatically export the \DATAFUSION\_VERSION\ environment variable by parsing the version from \datafusion/core/Cargo.toml\. Additionally, historical changelog files for versions 10.0.0, 11.0.0, and 12.0.0 were added to the \dev/changelog/\ directory, documenting breaking changes, enhancements, and bug fixes for those releases.
dev · high confidence
Add extended ClickBench query suite
The benchmarks/queries/clickbench/extended directory now contains individual SQL files (q0–q16) for an extended set of ClickBench queries. These additions introduce new benchmark patterns including grouped COUNT(DISTINCT) on string columns, FIRST\_VALUE window functions with GROUP BY, and various statistical aggregations like covar\_samp and APPROX\_PERCENTILE\_CONT, providing a broader set of performance test cases for the engine.
benchmarks/queries/clickbench/extended · high confidence
Add hash join SQL benchmarks for performance validation
Added a new suite of SQL benchmarks under \benchmarks/sql\_benchmarks/hj\ to validate HashJoinExec performance. The suite includes 25 query definitions (Q01–Q25) covering various join scenarios, including inner, right semi, and right anti joins, with varying build/probe side sizes, densities, and hit rates. It also introduces benchmarks for high-fanout and skewed string-key joins (Q23–Q25) to test candidate equality filtering. The suite is configured to run against TPC-H data and allows users to execute specific queries or the full set via the benchmark runner with configurable scale factors.
_benchmarks/sql\benchmarks/hj · high confidence
Add nested-loop join SQL benchmarks
Added a new suite of SQL benchmarks in the \nlj\ directory to measure NestedLoopJoinExec performance. The suite includes 17 query definitions (q01–q17) covering various join types (INNER, LEFT/RIGHT/FULL OUTER, SEMI, ANTI, MARK) with different table sizes and selectivity levels, along with a \nlj.suite\ configuration file that allows running all queries or specific ones via the benchmark runner.
_benchmarks/sql\benchmarks/nlj · high confidence
Added IMDB benchmark queries 10–26
The IMDB benchmark suite in benchmarks/queries/imdb has been expanded with 47 new SQL query files (10a through 26a). These additions cover a wide range of complex movie database queries, including filtering by production companies, country codes, release dates, genres, ratings, and keywords, as well as queries involving cast information, character names, and movie links.
benchmarks/queries/imdb · high confidence
Added Sort-Merge Join SQL benchmarks
The \benchmarks/sql\_benchmarks/smj\ directory now includes a new suite of SQL benchmarks focused on Sort-Merge Join performance. This addition introduces 26 individual benchmark queries (q01–q26) covering various join types (INNER, LEFT, FULL, LEFT SEMI, LEFT ANTI, LEFT MARK) and data scales, along with a suite configuration file that enables running the full set or specific queries via the benchmark runner.
_benchmarks/sql\benchmarks/smj · high confidence
Added focused Parquet pruning benchmark and SQL benchmark harness
Users can now run a new focused benchmark (\parquet\_pruning\_setup\_cache\) that measures end-to-end Parquet scan costs for a cache-favourable workload using 128 files with a shared schema, providing a baseline for comparing cache-disabled and cache-enabled branches. Additionally, a new SQL benchmark harness (\sql.rs\) has been added to run SQL benchmarks defined in \.benchmark\ files under \sql\_benchmarks\, executable via \benchmarks/bench.sh\ or directly with Cargo (e.g., \BENCH\_NAME=tpch cargo bench --bench sql\), supporting both \snmalloc\ and \mimalloc\ allocators.
benchmarks/benches · high confidence
Added protobuf code generation tool
A new Rust binary has been added to the project to automate the generation of DataFusion protobuf serialization code. This tool compiles the \datafusion.proto\ definition using \prost\ and \pbjson\, then copies the resulting generated files into the \datafusion/proto-models/src/generated\ directory, ensuring the serialization logic stays in sync with the protocol buffer schema.
datafusion/proto-models/gen · high confidence
Automatic PostgreSQL container management for SQL logic tests
The sqllogictest runner now automatically starts and manages a PostgreSQL test container when the \postgres\ feature is enabled. A new \postgres\_container.rs\ module handles the lifecycle of the container (using \testcontainers-modules\), exposing the host and port via channels so the runner can set the \PG\_URI\ environment variable. This allows users to run PostgreSQL-compatible SQL logic tests without manually configuring or starting a database instance.
datafusion/sqllogictest/bin · high confidence
Avro write support added to DataFusion
The \datasource-avro\ crate now supports writing Avro files. The \AvroFormat\ implementation includes a \create\_writer\_physical\_plan\ method that constructs a \DataSinkExec\ for appending data, utilizing \arrow\_avro\ for serialization. This enables users to insert data into Avro files via DataFusion's DML capabilities.
datafusion/datasource-avro · high confidence
Benchmark utilities now support simulated latency and per-query memory peak reporting
The benchmarking harness in benchmarks/src/util now includes a LatencyObjectStore wrapper that injects realistic S3-like latency (GET: 25–200ms, LIST: 40–400ms) into object store operations when the --simulate-latency flag is enabled, allowing users to evaluate performance under remote storage conditions. Additionally, the benchmark run utilities now track and report the peak memory pool reservation for each query via the PeakRecordingPool, exposing this metric in the JSON output alongside elapsed time and row counts to help users identify memory-bound queries.
benchmarks/src/util · high confidence
CLI now supports custom session contexts via the CliSessionContext trait
The datafusion-cli now allows users to provide custom session contexts by implementing the new \CliSessionContext\ trait, enabling advanced use cases such as plan modification or custom object-store registration. This is demonstrated by the new \cli-session-context.rs\ example, which shows how to wrap the standard context to automatically union query results with themselves. The core CLI execution logic in \exec.rs\ and \command.rs\ has been refactored to operate against this trait rather than a concrete \SessionContext\, making the CLI extensible for specialized workloads.
datafusion-cli/src · high confidence
Consolidated Arrow Flight examples into a single runnable entry point
The Flight examples have been reorganized into a unified \flight\ example with a \main.rs\ dispatcher that accepts subcommands (\client\, \server\, \sql\_server\, or \all\) to run specific demonstrations. This change introduces a new client example that connects to a remote Flight server to retrieve schemas and execute SQL queries, alongside existing server implementations for basic Flight and FlightSQL (JDBC-compatible) protocols, all driven by the \strum\ library for argument parsing.
datafusion-examples/examples/flight · high confidence
Consolidated Data I/O examples with new catalog, in-memory store, and JSON shredding demos
The \data\_io\ example module has been restructured into a single entry point (\main.rs\) that runs a suite of distinct data I/O demonstrations via subcommands. New examples include \catalog\ (registering tables into a custom catalog implementation), \in\_memory\_object\_store\ (reading CSV from an in-memory object store), and \json\_shredding\ (implementing custom filter rewriting for semi-structured data). Existing examples for Parquet advanced/embedded indexing, encryption (with and without KMS), object-store-backed spill files, HTTP CSV queries, and partitioned file schemas are also consolidated here, providing a unified location for learning these specific DataFusion capabilities.
_datafusion-examples/examples/data\io · high confidence
Consolidated benchmarking infrastructure and suite
The benchmarks directory has been reorganized into a unified structure featuring a new \bench.sh\ orchestration script, a \compare.py\ tool for performance diffing, and a SQL-based benchmark harness (\dfbench\). This change introduces a wide array of new benchmark suites including TPC-H, TPC-DS, ClickBench, H2O.ai, IMDB, and various micro-benchmarks (e.g., sort pushdown, predicate evaluation, spill views), alongside new utility scripts for compile profiling and LineProtocol export.
benchmarks · high confidence
Consolidated built-in function examples into a unified module
The \datafusion-examples/examples/builtin\_functions\ directory has been reorganized into a single, unified example application. This change consolidates previously scattered examples for date/time functions, regular expressions, and the \FunctionFactory\ (SQL macros) into a new modular structure. Users can now run all built-in function demonstrations via a single entry point (\cargo run --example builtin\_functions\) with subcommands (\date\_time\, \regexp\, \function\_factory\, or \all\), simplifying how to explore and test DataFusion's built-in capabilities.
_datafusion-examples/examples/builtin\functions · high confidence
Consolidated query planning examples
The query planning examples have been consolidated into a single executable with a modular structure. A new \main.rs\ entry point allows users to run all examples or select specific ones (such as \analyzer\_rule\, \expr\_api\, \optimizer\_rule\, \parse\_sql\_expr\, \plan\_to\_sql\, \planner\_api\, \pruning\, and \thread\_pools\) via command-line arguments. The individual example files have been refactored into separate modules to demonstrate distinct capabilities, including custom analyzer and optimizer rules, expression API usage, SQL parsing and unparsing, physical planning, file pruning, and custom thread pool configuration.
_datafusion-examples/examples/custom\_data\_source, datafusion-examples/examples/execution\_monitoring, datafusion-examples/examples/query\planning · high confidence
Embedded example datasets and automated documentation generation
Example datasets (cars.csv and regex.csv) are now stored directly in the repository under datafusion-examples/data, eliminating the need for external test files or submodules. To support these embedded files, new utility modules were added: a dataset schema registry (datasets/cars.rs, datasets/regex.rs) and a helper to convert CSVs to Parquet in temporary directories (csv\_to\_parquet.rs). Additionally, an automated documentation tool (examples-docs binary and example\_metadata utilities) now scans example groups and their main.rs doc comments to generate a README, keeping example documentation in sync with the code.
datafusion-examples/src · high confidence
Initial release of the DataFusion FFI crate for cross-library interoperability
The new \datafusion-ffi\ crate provides a stable Foreign Function Interface (FFI) boundary, enabling DataFusion libraries to share execution plans, table providers, and user-defined functions across different Rust compiler versions and runtime loads. This location introduces the core FFI structs and wrappers—including \FFI\_ExecutionPlan\, \FFI\_CatalogProvider\, \FFI\_TaskContext\, and \FFI\_ConfigOptions\—along with the \stabby\-based ABI stability mechanisms and memory management patterns required to safely pass complex DataFusion objects between separate compiled units.
datafusion/ffi · high confidence
Initial release of the datafusion-sql crate
The SQL query planner has been extracted into a standalone \datafusion-sql\ crate, providing a general-purpose SQL parser and logical plan generator that is decoupled from the physical execution engine. This new crate includes the core SQL planning logic (such as CTE handling and expression parsing), documentation, and licensing files, allowing projects to use DataFusion's SQL planning capabilities without depending on the full DataFusion execution stack.
datafusion/sql · high confidence
Initial repository structure and Apache Software Foundation governance setup
This change establishes the foundational repository structure for Apache DataFusion, introducing essential ASF governance and configuration files. It adds \.asf.yaml\ to configure GitHub repository settings, including branch protection rules for \main\ and release candidate branches, required status checks, and merge button configurations. The commit also introduces standard project files such as \LICENSE.txt\, \NOTICE.txt\, \CODE\_OF\_CONUT.md\, and \CONTRIBUTING.md\, alongside developer tooling configurations like \rust-toolchain.toml\ (pinning Rust 1.98.1), \clippy.toml\, \rustfmt.toml\, and \taplo.toml\. Additionally, it sets up AI agent guidelines via \AGENTS.md\ and \CLAUDE.md\, configures pre-commit hooks, and initializes Python dependency management with \uv.lock\ and \pyproject.toml\-adjacent structures, while adding submodules for testing data.
(repo-wide) · high confidence
Introduce DynamicFileCatalog for file-based table discovery
Added a new \DynamicFileCatalog\ component in \datafusion-catalog\ that wraps an existing catalog provider list to enable dynamic table creation from file paths. This feature introduces a chain of providers (\DynamicFileCatalog\, \DynamicFileCatalogProvider\, and \DynamicFileSchemaProvider\) that delegate standard catalog operations to an inner provider while intercepting table lookups; if a table is not found in the inner schema, the system attempts to create a \TableProvider\ using a pluggable \UrlTableFactory\ based on the table name (interpreted as a URL/path).
_datafusion/catalog/src/dynamic\file · high confidence
Introduce PhysicalExprAdapter for schema adaptation in file scans
A new \datafusion-physical-expr-adapter\ crate has been added to provide utilities for adapting physical expressions to different schemas during file scans. This includes the \PhysicalExprAdapter\ trait and \DefaultPhysicalExprAdapter\ implementation, which rewrite expressions to handle schema differences such as type casting, missing columns, and partition values. The crate also introduces helpers in \rewrite.rs\ to transform scalar UDFs like \file\_row\_index\ and \input\_file\_name\ into concrete physical expressions bound to the current file, and provides \BatchAdapter\ to simplify mapping \RecordBatch\ between schemas.
datafusion/physical-expr-adapter · high confidence
Introduce datafusion-functions crate with new Arrow metadata UDFs
The \datafusion/functions\ crate is introduced as a dedicated library for DataFusion functions, containing core utilities like \arrow\_cast\, \arrow\_field\, and \arrow\_metadata\. These new scalar UDFs allow users to cast expressions to specific Arrow data types, inspect the field metadata (name, type, nullability) of expressions, and retrieve custom metadata attached to Arrow fields. The crate also includes foundational binary concatenation builders and re-exports the project license and notice files.
datafusion/functions · high confidence
Introduce datafusion-proto-common crate for shared protobuf definitions
A new \datafusion-proto-common\ crate has been added to the DataFusion workspace to house shared Protocol Buffers definitions and serialization logic. This location contributes the \datafusion\_common.proto\ schema, the generated Rust code (prost and pbjson), and the \from\_proto\/\to\_proto\ conversion modules for common types such as schemas, data types, and constraints. This change factors out common datafusion types into a separate proto file to support serialization/deserialization of DataFusion primitive types, serving as a foundation for the \datafusion-proto\ crate.
datafusion/proto-common · high confidence
Introduce datafusion-proto-models crate for shared protobuf schemas
A new \datafusion-proto-models\ crate has been added to host the generated Rust types for DataFusion's logical and physical plan protobuf schemas, along with the \From\/\TryFrom\ conversions to \datafusion-common\ types. This separation allows other crates to reference the proto schema types without pulling in the full \datafusion-proto\ surface, while \datafusion-proto\ continues to re-export these types for existing users.
datafusion/proto-models · high confidence
Introduce datafusion-substrait crate for Substrait plan serialization
A new \datafusion-substrait\ crate has been added to provide a Substrait producer and consumer for DataFusion plans. This enables serializing and deserializing both logical and physical execution plans to and from the Substrait protocol buffer format, allowing DataFusion to run plans created by other systems (like Apache Calcite) and to pass query plans across FFI boundaries or between nodes.
datafusion/substrait · high confidence
Introduce generate\_series and range table functions
The \datafusion/functions-table\ crate now provides two new table functions: \generate\_series\ and \range\. These allow users to generate sequences of integer or timestamp values directly within SQL queries, supporting various step sizes and timezone-aware timestamp ranges. The implementation uses a lazy memory execution plan to efficiently produce these sequences without materializing the entire result set in memory upfront.
datafusion/functions-table · high confidence
Introduce logical type system and canonical extension types
DataFusion now provides a new logical type system in \datafusion/common/src/types\ to distinguish logical semantics from physical storage. This includes a \NativeType\ enum and \LogicalType\ trait for representing types like \Null\, \Boolean\, \Int32\, and \Decimal\ with their own signatures. Additionally, support for Arrow canonical extension types has been added, allowing DataFusion to handle \Bool8\, \UUID\, \JSON\, \Opaque\, \TimestampWithOffset\, and both fixed and variable shape tensors. These extension types are implemented via the new \DFExtensionType\ trait, enabling custom behaviors such as pretty-printing and storage validation.
datafusion/common/src/types · high confidence
Introduce xtask for local CI command execution
Developers can now run DataFusion's CI test and check commands locally using the new \cargo xtask\ tool, ensuring local runs match GitHub Actions. This adds a Rust-based automation layer (the \xtask\ binary) that exposes CI steps like \test\ and \check\, allowing developers to reproduce CI behavior and inspect underlying commands with the \--explain\ flag.
xtask · high confidence
Introduces structured file writer options for CSV, JSON, Parquet, Arrow, and Avro
DataFusion now exposes dedicated configuration structs (CsvWriterOptions, JsonWriterOptions, ParquetWriterOptions, ArrowWriterOptions, AvroWriterOptions) in the file\_options module to manage how data is written to disk. This change centralizes writer settings, allowing users to explicitly configure compression levels for CSV and JSON, and granularly control Parquet writer properties such as bloom filters, encoding, and column-specific statistics via SQL statement options. It also standardizes the handling of file type extensions and ensures that format options passed through CLI or SQL statements are correctly propagated to the underlying Arrow and Parquet writers.
_datafusion/common/src/file\options · high confidence
Introduction of the datafusion-physical-expr crate
The \datafusion-physical-expr\ crate has been introduced as a dedicated submodule for physical expression types and utilities, separating them from the core DataFusion crate. This new module provides the foundational data structures and logic required for evaluating expressions during the physical execution phase of query plans, including aggregate function builders, expression analysis contexts, and asynchronous scalar function support.
datafusion/physical-expr · high confidence
New ANY\_VALUE aggregate function
The \any\_value\ aggregate function is now available in DataFusion. It returns an arbitrary non-null value from a group, or NULL if the group contains only NULL values. This function is implemented as a User Defined Aggregate Function (UDAF) and reuses the existing \TrivialFirstValueAccumulator\ logic.
datafusion/functions-aggregate · high confidence
New CI scripts for local linting and breaking-change detection
The CI area now includes a comprehensive suite of new helper scripts under ci/scripts/ to standardize local development checks and improve the merge-queue safety. These scripts allow developers to run CI validations locally, including breaking-change detection (changed\_crates.sh, check\_asf\_yaml\_status\_checks.py), documentation generation and verification (check\_docs\_html.sh, check\_examples\_docs.sh, check\_generated\_docs.sh), code quality and formatting (rust\_clippy.sh, rust\_fmt.sh, rust\_toml\_fmt.sh, doc\_prettier\_check.sh, license\_header.sh, typos\_check.sh), and dependency/security auditing (check\_unused\_dependencies.sh, security\_audit.sh). Additionally, new scripts enforce repository hygiene by checking for large files (check\_large\_files.sh), ensuring examples only use the main datafusion crate (check\_examples\_datafusion\_crates.py), preventing direct cargo installs in workflows (check\_no\_cargo\_install\_in\_workflows.sh), and validating markdown links (markdown\_link\_check.sh). A release version labeler script (release\_version\_labeler.js) and a retry utility (retry) are also added to support CI workflows.
ci · high confidence
New CLI tools for generating DataFusion configuration and function documentation
Three new command-line binaries have been added to the DataFusion core crate to automate the generation of documentation from embedded code. The \print\_config\_docs\ binary outputs Markdown for \ConfigOptions\, \print\_runtime\_config\_docs\ outputs Markdown for \RuntimeEnvBuilder\ settings, and \print\_functions\_docs\ generates structured documentation for aggregate, scalar, and window functions (including aliases and SQL examples) by querying the session state defaults. These tools support the project's shift toward attribute-based, code-embedded documentation.
datafusion/core/src/bin · high confidence
New DataFrame write and parquet API
The DataFrame API now includes methods to write data to Parquet files. Users can use \write\_parquet\ to export DataFrame results, configuring output behavior via \DataFrameWriteOptions\ which supports setting the insert operation, forcing single-file output, partitioning by column, and sorting the output. The implementation leverages the \copy\_to\ logical plan builder and respects optional \TableParquetOptions\ for writer-level settings like compression.
datafusion/core/src/dataframe · high confidence
New FFI Table Provider example for dynamic module loading
Added a new example library (\ffi\_example\_table\_provider\) that demonstrates how to implement a DataFusion TableProvider as a dynamically loadable FFI module. The example exports a \ffi\_example\_get\_module\ entry point that constructs a simple in-memory table with integer and float columns, allowing users to see how to expose DataFusion tables across language boundaries using the FFI interface.
_datafusion-examples/examples/ffi/ffi\_example\_table\provider · high confidence
New FFI module loader example for dynamic table providers
Added a new example demonstrating how to dynamically load a DataFusion TableProvider from a shared library at runtime. The \ffi\_module\_loader\ executable uses \libloading\ to load a compiled library (e.g., \ffi\_example\_table\_provider.so\), retrieves a \TableProviderModule\ struct containing a factory function, and invokes it to create an \FFI\_TableProvider\. This provider is then converted into a standard DataFusion \TableProvider\, registered in a \SessionContext\, and queried, illustrating the full cycle of cross-boundary table provider integration.
_datafusion-examples/examples/ffi/ffi\_module\loader · high confidence
New SQL predicate evaluation micro-benchmark suite
Added a new \predicate\_eval\ benchmark suite under \benchmarks/sql\_benchmarks/predicate\_eval\ to measure conjunctive filter evaluation performance. The suite includes a shared template, a suite definition with configurable row and string-width knobs, and a comprehensive set of test cases covering cardinality, correlation, cost, selectivity, drift, width, scale, and nullable predicates. It also provides SQL scripts to generate synthetic datasets (e.g., \ints\, \markers\, \drift\) and defines the expected result CSVs for validation.
_benchmarks/sql\_benchmarks/predicate\eval · high confidence
New ScalarStructBuilder and scalar caching infrastructure
The \datafusion/common/src/scalar\ module now includes a \ScalarStructBuilder\ to simplify the creation of \ScalarValue::Struct\ instances, including dedicated support for constructing null structs. Additionally, a new caching layer (\cache.rs\) has been introduced to reuse dictionary key and null arrays, reducing allocation overhead during scalar-to-array conversions. Constant definitions for decimal powers and floating-point bounds have also been consolidated into a new \consts.rs\ module.
datafusion/common/src/scalar · high confidence
New analyzer rule for expression-to-function rewrites
The logical optimizer now includes an \ApplyFunctionRewrites\ analyzer rule that transforms expressions into function calls (such as converting \\|\|\ to \array\_concat\) before other analysis passes like type coercion. This allows for more flexible expression handling and optimization during the planning phase.
datafusion/optimizer · high confidence
New array functions: array\_add, array\_avg, array\_any\_match, array\_compact, array\_filter, array\_first, array\_normalize
The \datafusion/functions-nested\ crate now includes several new array functions. \array\_add\ computes the element-wise sum of two numeric arrays. \array\_avg\ returns the arithmetic mean of elements in a numeric array, skipping NULLs. \array\_any\_match\ is a higher-order function that returns true if any element in an array satisfies a given predicate. \array\_compact\ removes null values from an array. \array\_filter\ filters array values using a boolean lambda. \array\_first\ returns the first element of an array that satisfies a predicate. \array\_normalize\ returns the L2-normalized vector for a numeric array.
datafusion/functions-nested · high confidence
New benchmark runner binaries and utilities for DataFusion
The benchmarks directory now includes several new command-line tools to run and profile DataFusion workloads. The \benchmark\_runner\ binary provides a unified interface for SQL benchmarks with support for listing suites, running simple or Criterion-based benchmarks, and outputting results as JSON. The \dfbench\ binary serves as the main entry point for a wide range of existing benchmark suites including TPC-H, TPC-DS, ClickBench, H2O, IMDB, and various join/sort benchmarks. New utilities include \external\_aggr\ for testing memory-limited aggregation performance, \mem\_profile\ for profiling memory usage across benchmarks, and \gen\_wide\_data\ for synthesizing wide-schema Parquet datasets to measure metadata overhead.
benchmarks/src/bin · high confidence
New benchmark suite for join performance and query cancellation
The benchmark runner now includes dedicated micro-benchmarks for Hash Join (hj.rs), Nested Loop Join (nlj.rs), and Sort-Merge Join (smj.rs) that allow users to measure performance across various join types, input sizes, and selectivities. Additionally, a new cancellation benchmark (cancellation.rs) has been added to verify that queries stop executing quickly after being cancelled, and the existing ClickBench runner (clickbench.rs) has been updated to support the \--pushdown\ flag for enabling Parquet filter pushdown.
benchmarks/src · high confidence
New benchmarking data utilities module
Added a new \data\_utils\ module to the benchmarking suite, providing helper functions to generate realistic in-memory tables with various data distributions (including wide, mid, and narrow integer ranges, string views, and dictionary-encoded columns) for more accurate performance testing.
_datafusion/core/benches/data\utils · high confidence
New boundary-aligned streaming and generic decoder infrastructure for CSV/JSON reads
The datasource module introduces \AlignedBoundaryStream\ to handle newline-delimited JSON and CSV range reads by lazily aligning byte boundaries to record (newline) characters, automatically issuing additional bounded GET requests when the terminating newline is not found in the initial fetch window. It also adds a generic \DecoderDeserializer\ and \BatchDeserializer\ trait that wrap Arrow's JSON and CSV decoders, allowing file formats to process byte streams into \RecordBatch\ objects through a unified, buffered interface.
datafusion/datasource/src · high confidence
New common library module for DataFusion core types
The \datafusion/common\ crate now includes a new \alias\ module providing an \AliasGenerator\ for creating unique query aliases, a \cast\ module with safe downcasting functions for Arrow array types, a \column\ module defining the \Column\ struct for qualified field references, a \config\ module implementing the \ConfigOptions\ tree and \config\_namespace!\ macro for runtime configuration, and a \cse\ module containing the Common Subexpression Elimination logic with \HashNode\ and \NormalizeEq\ traits.
datafusion/common/src · high confidence
New common-runtime crate for task spawning and tracing
The \datafusion/common-runtime\ crate is introduced to provide shared utilities for managing asynchronous tasks. It exposes a \SpawnedTask\ wrapper that ensures tasks are aborted on drop for cancel-safety and supports unwinding panics via \join\_unwind\. Additionally, it provides a \JoinSet\ wrapper around Tokio's \JoinSet\ that allows injecting a global \JoinSetTracer\ to instrument spawned futures and blocking closures for context propagation, with a no-op fallback when no tracer is configured.
datafusion/common-runtime · high confidence
New consolidated DataFrame examples with custom caching and struct deserialization
The \datafusion-examples/examples/dataframe\ directory now provides a unified entry point (\main.rs\) that runs three distinct examples: a core \dataframe\ example demonstrating reading from Parquet, CSV, and in-memory sources and writing output; a \deserialize\_to\_struct\ example showing how to convert Arrow \RecordBatch\ results into native Rust structs; and a new \cache\_factory\ example that demonstrates implementing custom lazy caching strategies for DataFrames using the \CacheFactory\ API. The examples use \strum\ for argument parsing and consolidate previously scattered examples into a single runnable module.
datafusion-examples/examples/dataframe · high confidence
New datafusion-datasource-arrow crate for Arrow IPC file support
A new \datafusion-datasource-arrow\ crate has been introduced to handle Apache Arrow IPC file formats. This module provides the \ArrowFormat\ implementation, which supports both the standard Arrow IPC file format (with footer, enabling parallel range-based reading) and the Arrow IPC stream format (sequential reading). It includes the \ArrowFormatFactory\ for registration, schema inference logic that merges schemas from multiple files, and execution plans via \ArrowSource\ that utilize \ArrowFileOpener\ and \ArrowStreamFileOpener\ for efficient data retrieval from object stores.
datafusion/datasource-arrow · high confidence
New datafusion-execution crate introduces execution runtime and caching infrastructure
The \datafusion-execution\ crate is introduced as a new submodule providing core execution runtime components, including the \SessionConfig\ for query configuration, \DiskManager\ for managing temporary spill files, and a new caching subsystem (\cache\) with an LRU-based \DefaultCache\ implementation for file statistics, file listings, and metadata. It also exposes \async\_stream\ utilities for creating streams from async generators and an \ArrowMemoryPool\ adapter to integrate DataFusion's memory management with Arrow's allocation APIs.
datafusion/execution · high confidence
New datafusion-functions-aggregate-common crate for shared aggregate logic
The \datafusion/functions-aggregate-common\ crate has been introduced to centralize common functionality for aggregate and window functions. This new location provides shared implementations for \AVG DISTINCT\ (supporting both numeric and decimal types) and \COUNT DISTINCT\ (optimized for primitive, byte, dictionary, and grouped scenarios), along with the \GroupsAccumulatorAdapter\ that bridges standard accumulators to the grouped execution path. It also includes benchmarks for the accumulation and adapter update paths, and establishes the \AccumulatorArgs\ and \StateFieldsArgs\ structures that define the context passed to aggregate function implementations.
datafusion/functions-aggregate-common · high confidence
New datafusion-physical-expr-common crate for shared physical expression logic
A new \datafusion-physical-expr-common\ crate has been introduced to host shared APIs and implementations for physical expressions, including the \PhysicalExpr\ and \PhysicalSortExpr\ traits. This location now provides optimized data structures for handling string and binary data, specifically \ArrowBytesMap\ and \ArrowBytesViewMap\, which are designed to minimize copying when storing distinct values for operations like \COUNT DISTINCT\ and \GROUP BY\. The crate also includes a \datum\ module for applying Arrow \Datum\ kernels to DataFusion's \ColumnarValue\ abstraction, a \metrics\ module for tracking execution statistics like elapsed compute time and output skew, and benchmarks for comparing nested types and byte maps.
datafusion/physical-expr-common · high confidence
New datafusion-session crate for shared session and catalog interfaces
A new \datafusion-session\ crate has been introduced to centralize shared interfaces for session-related APIs and extension points. This crate defines the core traits and structures that constitute the runtime context for query execution, including \Session\ (for configuration, function registry, and planning), \CatalogProvider\/\SchemaProvider\/\TableProvider\ (for metadata hierarchies), and \QueryPlanner\/\PhysicalPlanner\/\PhysicalOptimizerRule\ (for plan generation and optimization). By extracting these shared contracts into a dedicated crate, DataFusion enables concrete query-engine implementations to depend on stable extension points without tight coupling to the core engine, while most projects should continue using the main \datafusion\ crate which re-exports these modules.
datafusion/session · high confidence
New datafusion-spark crate with Spark-compatible functions
A new \datafusion-spark\ crate has been introduced to provide Apache Spark-compatible expressions for DataFusion. This location contributes the implementation of specific Spark-compatible scalar, aggregate, and window functions (such as \avg\, \collect\_list\, \collect\_set\, \try\_sum\, \array\_contains\, \shuffle\, \slice\, and \array\), along with the necessary module structure and registration helpers. Users can enable this library via the \--spark\ CLI flag or by registering the functions in their session, allowing DataFusion to execute queries using Spark's specific semantics for these operations.
datafusion/spark · high confidence
New display utilities for human-readable metrics and GraphViz execution plans
The \datafusion/common\ crate now includes a \display\ module that provides utilities for formatting execution plan output. Users can now view execution plans as GraphViz DOT graphs for visualization, and see size, count, and duration metrics in human-readable formats (e.g., '1.53 MB', '10.10 K', '1.23s'). This also introduces the \PlanType\ enum and \StringifiedPlan\ struct, which allow EXPLAIN commands to display various stages of the query plan (logical and physical) with support for verbose output to show intermediate steps.
datafusion/common/src/display · high confidence
New documentation submodule for user-defined functions
A new \datafusion/doc\ crate has been introduced to provide the structures and macros used for documenting user-defined functions (UDFs). This module defines the \Documentation\ struct and \DocSection\ types, along with specific section definitions for scalar, aggregate, and window functions, enabling the automatic generation of SQL function documentation. It also includes standard license and notice file references to ensure proper attribution in the generated docs.
datafusion/doc · high confidence
New extension\_types example demonstrating custom DataFusion extension types
Added a new example in datafusion-examples/examples/extension\_types that demonstrates how to create and use custom extension types in DataFusion. The example implements a TemperatureExtensionType that wraps Float32/Float64 data with unit metadata (Celsius, Fahrenheit, Kelvin), showing how to register the type via MemoryExtensionTypeRegistry, define field schemas with Arrow extension metadata, and query the data using a SessionContext.
_datafusion-examples/examples/extension\types · high confidence
New external dependency examples for Amazon S3 integration
Added new example files in the \external\_dependency\ directory (\dataframe\_to\_s3.rs\, \query\_aws\_s3.rs\, and \main.rs\) that demonstrate how to use DataFusion with Amazon S3. These examples show how to register the S3 object store, query public datasets (like the NYC TLC open dataset) and local S3 buckets using CSV and Parquet formats, and write query results back to S3 in Parquet, JSON, and CSV formats. The examples use the \object\_store\ crate for S3 connectivity and \strum\ for command-line argument parsing to run specific examples or all of them.
_datafusion-examples/examples/external\dependency · high confidence
New factory-based table creation and dynamic file support
The \datasource\ module now uses dedicated factory implementations to create tables, introducing \ListingTableFactory\ for standard external tables and \StreamTableFactory\ for unbounded streams, coordinated by a \DefaultTableFactory\ that routes \CREATE EXTERNAL TABLE\ commands based on the \unbounded\ option. A new \DynamicListTableFactory\ enables creating \ListingTable\ instances directly from URLs via the \UrlTableFactory\ interface, automatically inferring schema and partitioning. Additionally, \MemTable\ now preserves constraint metadata (such as primary keys) when loaded, ensuring that constraint information is maintained across table operations.
datafusion/core/src/datasource · high confidence
New proto examples for composed codecs and expression deduplication
The \datafusion-examples/examples/proto\ directory now includes two new examples demonstrating advanced serialization patterns. The \composed\_extension\_codec\ example shows how to combine multiple \PhysicalExtensionCodec\ implementations to serialize execution plans containing nodes from different crates (e.g., Ballista and Delta Lake). The \expression\_deduplication\ example demonstrates using the \PhysicalProtoConverterExtension\ trait to cache and deduplicate physical expressions during deserialization, reducing memory usage and enabling optimizations based on pointer equality.
datafusion-examples/examples/proto · high confidence
New protobuf code generation tool for datafusion-common
A new standalone binary has been added to the datafusion/proto-common/gen module to automate the generation of Rust bindings for the datafusion\_common.proto file. This tool compiles the protobuf definition into both prost and pbjson formats, producing the resulting source files in the src/generated directory. This change refactors the build process by extracting the common code generation logic into a dedicated, reusable tool rather than embedding it within other build scripts.
datafusion/proto-common/gen · high confidence
New public API for serializing DataFusion plans and expressions to bytes
The \datafusion/proto\ crate now exposes a public \bytes\ module that allows users to serialize \Expr\, \LogicalPlan\, and \ExecutionPlan\ objects to opaque byte streams (and JSON) for transport or storage. This includes the \Serializeable\ trait for expressions, convenience functions like \logical\_plan\_to\_bytes\ and \physical\_plan\_from\_bytes\, and support for custom extension codecs to handle user-defined functions during round-trips.
datafusion/proto · high confidence
New relation planner examples for custom SQL operators
Added a new \relation\_planner\ example demonstrating how to extend DataFusion's SQL syntax with custom table operators using the \RelationPlanner\ extension API. The example includes three sub-examples: \match\_recognize\ (logical planning for pattern matching), \pivot\_unpivot\ (rewriting PIVOT/UNPIVOT to standard SQL), and \table\_sample\ (full logical and physical planning for TABLESAMPLE).
_datafusion-examples/examples/relation\planner · high confidence
New sqllogictest runner with config matrix and crash recovery
The sqllogictest runner now supports a \configMatrix\ directive that sweeps DataFusion configuration settings across repeated runs of a single test file, and includes a \CurrentlyExecutingSqlTracker\ that records the last executed SQL statements to aid debugging in case of a crash.
datafusion/sqllogictest · high confidence
New standalone dev tool for detecting circular dependencies
A new standalone binary, \dev/depcheck\, has been added to the repository to verify that there are no circular dependencies among DataFusion crates. This tool parses the Cargo.toml files and checks the dependency graph, ensuring that publishing to crates.io remains possible by preventing circular references between packages.
dev/depcheck · high confidence
New statistics example demonstrating join reorder via StatisticsRegistry
Added a new \statistics\ example that demonstrates how to use the \StatisticsRegistry\ to plug in refined cardinality estimation. The \join\_reorder\ subcommand shows how supplying custom column statistics (specifically refining distinct counts using a survival formula) can change the optimizer's decision on which side of a hash join to build, flipping the build side from the larger table to the smaller one for better performance.
datafusion-examples/examples/statistics · high confidence
New stdin object store and CLI profiling instrumentation
The CLI now supports reading data from standard input via a new \stdin://\ object store scheme, allowing piped data (e.g., \cat data.csv \| datafusion-cli\) to be queried as external tables. Additionally, an instrumented object store wrapper has been added to the CLI, enabling users to profile object store operations (such as PUT, GET, LIST, and COPY) with configurable modes (Disabled, Summary, Trace) to monitor performance metrics like request duration and throughput.
_datafusion-cli/src/object\storage · high confidence
New test utilities for generating and scanning CSV and Parquet files
The \datafusion/core/src/test\_util\ module now includes dedicated helpers for creating test data files. A new \TestCsvFile\ struct allows tests to write RecordBatches to CSV files and retrieve their schema, while \TestParquetFile\ provides similar functionality for Parquet, including a \create\_scan\ method that builds a \DataSourceExec\ with optional filter pushdown support. The module also exposes \populate\_csv\_partitions\ for generating partitioned CSV datasets and retains existing utilities like \scan\_empty\ and \register\_aggregate\_csv\.
_datafusion/core/src/test\util · high confidence
New test utilities for random data generation and TPC benchmark schemas
The \test-utils\ crate now includes a comprehensive \array\_gen\ module with generators for random Binary, Boolean, Decimal, Primitive, and String arrays, enabling more robust fuzz testing across diverse data types. It also introduces a \data\_gen\ module for creating synthetic access-log-style RecordBatches and provides \tpch\ and \tpcds\ modules that define the schemas and primary-key constraints for the TPC-H and TPC-DS benchmark suites, simplifying the setup of standardized test data.
test-utils · high confidence
New user\_doc macro and documentation infrastructure for DataFusion macros crate
The datafusion/macros crate now includes a new \user\_doc\ procedural macro that automatically generates user-facing documentation for DataFusion functions (such as ScalarUDF, AggregateUDF, and WindowUDF) by parsing custom attributes like \doc\_section\, \description\, and \sql\_example\. This change also adds standard LICENSE and NOTICE symlink files and a README to the crate, establishing the foundational documentation infrastructure for this subcrate.
datafusion/macros · high confidence
New utility modules in datafusion-common for aggregation, hex encoding, and memory accounting
The \datafusion/common/src/utils\ module now includes several new sub-modules that provide reusable utilities for internal and downstream use. The \aggregate\ module adds \scalar\_add\ and \precision\_add\ helpers for in-place accumulation of \ScalarValue\ and \Precision\<ScalarValue\>\, supporting null propagation and overflow handling. The \hex\ module introduces \HexCase\ and functions like \encode\_bytes\, \encode\_bytes\_into\, and \encode\_bytes\_to\_slice\ for efficient hex encoding of bytes and integers with configurable case. The \memory\ module provides \estimate\_memory\_size\ for pre-allocation sizing of hash tables and \RecordBatchMemoryCounter\ for accurately tracking shared buffer memory across multiple \RecordBatch\es. Additionally, \proxy.rs\ adds \VecAllocExt\ and \HashTableAllocExt\ traits for tracking memory allocations on \Vec\ and \HashTable\ respectively, while \string\_utils.rs\ offers \string\_array\_to\_vec\ for converting Arrow string arrays to vectors of string slices.
datafusion/common/src/utils · high confidence
SessionContext now supports reading and writing Avro, CSV, JSON, and Parquet files
The \SessionContext\ API in \datafusion/core/src/execution/context\ has been expanded with dedicated methods for four data formats. Users can now read and write Avro, CSV, JSON, and Parquet files directly via \read\\<format\>\, \register\\<format\>\, and \write\_\<format\>\ methods (e.g., \read\_csv\, \write\_parquet\). These methods are implemented in new format-specific modules (\avro.rs\, \csv.rs\, \json.rs\, \parquet.rs\) and integrate with the existing \ListingTable\ infrastructure, allowing users to easily ingest and export data in these common formats without manual table registration or complex configuration.
datafusion/core/src/execution/context · high confidence
TPC-H and TPC-DS benchmarks now support primary key constraints and configurable join strategies
The TPC-H and TPC-DS benchmark runners have been updated to register primary key constraints on their respective tables, enabling the query optimizer to leverage this metadata for more efficient execution plans. Additionally, users can now control join behavior via new command-line options: \--prefer-hash-join\ and \--enable-piecewise-merge-join\ allow selecting between hash and sort-merge joins, while \--hash-join-buffering-capacity\ lets users tune the memory buffer size for the probe side of hash joins. The TPC-H runner also introduces a \--scale-factor\ option to explicitly set the scale factor for query substitutions.
benchmarks/src/tpch · high confidence
Architecture
CSV datasource extracted into standalone crate
The CSV file source implementation has been split out from the main DataFusion crate into its own dedicated \datafusion-datasource-csv\ crate. This modularization provides a standalone \CsvSource\ and \CsvFormat\ implementation, allowing users to depend on or extend CSV handling independently of the broader DataFusion execution engine.
datafusion/datasource-csv · high confidence
Catalog implementation code moved to the datafusion-catalog crate
The concrete implementations for catalog components—such as the Information Schema, CTE work tables, streaming tables, views, and memory-based providers—have been moved from the core crate into the new \datafusion-catalog\ crate. This change decouples the core logical and physical planning logic from specific catalog implementations, allowing the core to remain independent of these storage and metadata details while exposing the same catalog interfaces through re-exports.
datafusion/catalog/src · high confidence
DataFusion core crate restructured with new module layout and documentation
The \datafusion/core\ crate has been reorganized into a new modular structure. The main entry point (\lib.rs\) now includes updated architecture documentation and examples, while error handling is centralized in a new \error.rs\ module that re-exports types from \datafusion\_common\. A new \prelude.rs\ module simplifies importing common types like \SessionContext\, \DataFrame\, and expression functions. The physical planning logic is now contained in \physical\_planner.rs\, and schema equivalence checks are handled by \schema\_equivalence.rs\. Additionally, optimizer rule documentation and corresponding tests have been added in \optimizer\_rule\_reference.rs\ and \optimizer\_rule\_reference.md\ to ensure the documented rule order matches the actual implementation.
datafusion/core/src · high confidence
Introduce datafusion-expr-common crate for shared expression logic
A new \datafusion-expr-common\ crate has been added to the DataFusion workspace to host shared types and traits used by both logical and physical expressions, such as \Accumulator\, \GroupsAccumulator\, \ColumnarValue\, \Operator\, \Signature\, and type coercion utilities. This refactoring isolates common expression logic to prevent physical expressions from depending on logical expressions, providing a cleaner architectural boundary while re-exporting the module through the main \datafusion\ crate for existing users.
datafusion/expr-common · high confidence
Introduce dedicated datafusion-catalog-listing crate for ListingTable
The \ListingTable\ implementation and its associated configuration (\ListingTableConfig\, \ListingOptions\) have been moved from the core \datafusion\ crate into a new, separate \datafusion-catalog-listing\ crate. This architectural change decouples the file-listing table provider from the core execution engine, allowing for independent evolution and reduced compile times for the core crate. Users relying on the \ListingTable\ should now depend on the \datafusion-catalog-listing\ crate directly, although the main \datafusion\ crate continues to re-export these components for backward compatibility.
datafusion/catalog-listing · high confidence
ListingTable implementation moved to datafusion-catalog-listing
The \ListingTable\, \ListingOptions\, and \ListingTableConfig\ types have been moved from \datafusion/core\ into the new \datafusion-catalog-listing\ crate. This location now re-exports these types from \datafusion\_catalog\_listing\ and provides a \ListingTableConfigExt\ trait for session-aware schema inference, while maintaining backward compatibility by re-exporting \PartitionedFileStream\ from \datafusion\_datasource\.
datafusion/core/src/datasource/listing · high confidence
MemTable and in-memory catalog providers moved to datafusion-catalog
The in-memory table implementation (MemTable) and its associated catalog components (MemoryCatalogProvider, MemorySchemaProvider, and MemoryCatalogProviderList) have been relocated from the core datafusion crate to the new datafusion-catalog crate. This architectural change reduces the compile time and binary size of the core datafusion crate by decoupling these specific in-memory implementations, while preserving their functionality for users who rely on them for testing or lightweight in-memory data sources.
datafusion/catalog/src/memory · high confidence
New datafusion-pruning crate for modular file and row-group pruning
The pruning logic previously embedded in the main DataFusion crate has been extracted into a dedicated \datafusion-pruning\ crate to improve modularity. This new crate exposes the \FilePruner\ for skipping files based on partition values and file-level statistics, and the \PruningPredicate\ for analyzing filter expressions against statistics (such as min/max values and null counts) to skip row groups or other containers. It includes optimized handling for IN-list predicates on both primitive and string types, and benchmarks are provided to measure performance. Most projects should continue using the \datafusion\ crate directly, which re-exports this module.
datafusion/pruning · high confidence
Parquet datasource restructured into modular components
The \datafusion-datasource-parquet\ crate has been refactored to split its internal logic into distinct, dedicated modules. The new structure includes \access\_plan.rs\ for managing row-group and row-level read selections, \bloom\_filter.rs\ for handling Split Block Bloom Filter pruning statistics, \decoder\_projection.rs\ for constructing projection masks and batch mapping, \metadata.rs\ for consolidating file metadata and statistics fetching into \DFParquetMetadata\, and \metrics.rs\ for detailed scan performance tracking. This modularization improves code organization and maintainability for the Parquet file source implementation.
datafusion/datasource-parquet · high confidence
Physical optimizer rules moved to dedicated crate
The physical optimizer rules (including \AggregateStatistics\, \CombinePartialFinalAggregate\, \EnsureCooperative\, and the \EnsureRequirements\ distribution/sorting enforcement logic) have been moved from the core \datafusion\ crate into a new \datafusion-physical-optimizer\ crate. This change improves modularity by separating the optimizer implementation from the core execution plan definitions, while the \datafusion\ crate continues to re-export these components for existing users.
datafusion/physical-optimizer · high confidence
Session state management moved to a dedicated execution module
The core session state logic, including \SessionState\, \SessionStateBuilder\, and \SessionStateDefaults\, has been relocated from the legacy \core\ module into a new \datafusion/core/src/execution\ module. This change reorganizes the codebase to better separate execution concerns, exposing the session state builder and defaults via the new \execution\ module while maintaining backward compatibility through re-exports.
datafusion/core/src/execution · high confidence
Behavioural changes
ClickBench query suite restructured with per-file queries and byte-length fixes
The ClickBench benchmark queries in benchmarks/queries/clickbench/queries have been split into individual files (q0.sql through q42.sql) for easier maintenance, and a new sorted\_data subdirectory has been added with its own query set. Additionally, queries q27 and q28 now use octet\_length instead of length to correctly measure byte-length semantics for URLs and referers, aligning the benchmark results with ClickBench expectations.
benchmarks/queries/clickbench/queries · high confidence
Consolidated SQL operations examples into a single runnable module
The \sql\_ops\ example directory has been reorganized into a unified module (\main.rs\) that groups four distinct SQL operation demonstrations—querying data, analyzing logical plans, building plans via a custom frontend, and extending the SQL parser—into a single executable. Users can now run all demonstrations at once or select specific ones (e.g., \analysis\, \custom\_sql\_parser\, \frontend\, \query\) via command-line arguments, replacing the previous scattered example structure.
_datafusion-examples/examples/sql\ops · high confidence
Consolidated and updated User-Defined Function examples
The \datafusion-examples/examples/udf\ directory has been restructured into a single unified example application. A new \main.rs\ entry point now manages and runs all UDF demonstrations via command-line arguments, including advanced implementations for scalar functions (UDF), aggregate functions (UDAF), window functions (UDWF), and table functions (UDTF), as well as new examples for asynchronous UDFs and struct-returning aggregates. The example code has been updated to reflect the latest DataFusion APIs, such as the removal of \as\_any\ from trait definitions and the adoption of the \ScalarFunctionArgs\ interface for scalar UDFs.
datafusion-examples/examples/udf · high confidence
File format modules re-exported from datafusion-datasource
The file format modules in \datafusion/core/src/datasource/file\format\ (including CSV, Parquet, JSON, Avro, and Arrow) now re-export their implementations from the corresponding \datafusion-datasource-\\ crates. This change centralizes the file format logic in the dedicated datasource crates while keeping the public API in the core crate, ensuring that existing code using these formats continues to work without modification.
_datafusion/core/src/datasource/file\format · high confidence
FileScanConfig refactored into modular components with improved sort pushdown and serialization
The file scanning configuration logic has been reorganized into a dedicated module (\file\_scan\_config\) split into three focused files: the core configuration (\mod.rs\), a new sort pushdown optimization module (\sort\_pushdown.rs\), and shared protocol buffer serialization (\proto.rs\). This refactoring extracts statistics-based file sorting and NULL handling logic into \sort\_pushdown.rs\ to enable more precise SortExec elimination when file groups can be reordered to match query requirements. The \proto.rs\ module centralizes serialization, ensuring that \FileScanConfig\ fields are correctly round-tripped across process boundaries using the new \ExecutionPlanEncodeCtx\/\DecodeCtx\ context, while preventing silent data loss for non-serializable components like expression adapter factories. Users benefit from more efficient query execution plans that avoid unnecessary sorting operations and reliable serialization of file scan configurations.
_datafusion/datasource/src/file\_scan\config · high confidence
FileStream refactored with builder pattern and dynamic work stealing
The file scanning implementation in \datafusion/datasource/src/file\_stream\ has been rewritten to support dynamic work stealing and a new construction API. A \FileStreamBuilder\ now handles the creation of file streams, replacing the previous direct constructor. The scan logic has been restructured around a \ScanState\ and \Morsel\ API, enabling \SharedWorkSource\ to distribute file processing across sibling streams for improved parallelism. Additionally, \FileStreamMetrics\ has been extracted into its own module to provide detailed wall-clock timing for file opening, scanning, and processing, along with error and file-count counters.
_datafusion/datasource/src/file\stream · high confidence
JSON datasource refactored into new modular crate with streaming array support
The JSON file source implementation has been restructured into a dedicated \datafusion/datasource-json\ crate, replacing the previous monolithic \JsonExec\ with a new \FileSource\-based architecture. This change introduces a streaming \JsonArrayToNdjsonReader\ that converts large JSON arrays to newline-delimited JSON on-the-fly, significantly reducing memory usage for array-format files. The new module includes a \JsonFormatFactory\ for format registration, a \JsonSource\ for execution planning, and updated licensing/notice symlinks, providing a more modular and efficient foundation for JSON data ingestion.
datafusion/datasource-json · high confidence
New common library for user-defined window function arguments
The \datafusion/functions-window-common\ crate has been introduced to provide shared argument structures for user-defined window functions. It adds \ExpressionArgs\ for passing input expressions and fields, \WindowUDFFieldArgs\ for defining result field metadata, and \PartitionEvaluatorArgs\ for passing execution context (including \is\_reversed\ and \ignore\_nulls\ flags) to partition evaluators. This refactors the API to use \FieldRef\ instead of \Field\ for better performance and consistency.
datafusion/functions-window-common · high confidence
Re-export unified DataSourceExec and FileSource implementations for all file formats
The physical plan module now re-exports the unified \DataSourceExec\ and \FileSource\ implementations for Arrow, Avro, CSV, JSON, and Parquet from their respective \datafusion\datasource\\*\ crates. This change replaces the previous format-specific execution plan types (such as \CsvExec\, \ParquetExec\, \JsonExec\, and \ArrowExec\) with a single, consistent execution plan structure, simplifying the API and unifying how different file formats are read and processed.
_datafusion/core/src/datasource/physical\plan · high confidence
Refactored file writing into modular components with configurable buffer and compression settings
The file write logic in the datasource module has been reorganized into three distinct modules: \demux.rs\ handles splitting the input stream into multiple output files based on row counts and partition columns; \mod.rs\ defines the core traits and builders, including \ObjectWriterBuilder\ which now supports configurable buffer sizes and compression levels; and \orchestration.rs\ manages the parallel serialization and writing of record batches to the object store. This refactoring introduces new configuration options for writer buffer size and compression, allowing users to tune performance and output file characteristics.
datafusion/datasource/src/write · high confidence
Window functions are now user-defined (UDWFs) with a new crate
The \datafusion/functions-window\ crate now implements all standard window functions—including \row\_number\, \rank\, \dense\_rank\, \percent\_rank\, \ntile\, \cume\_dist\, \lead\, \lag\, \first\_value\, \last\_value\, and \nth\_value\—as User-Defined Window Functions (UDWFs). This architectural shift replaces the previous built-in implementation, providing a unified extension API, improved documentation generation, and a fluent expression API for building window queries.
datafusion/functions-window · high confidence
Fixes
Fix duplicate groups in legacy hash aggregation with fallback group keys
Resolves a correctness issue in legacy hash aggregation where duplicate groups were produced when using a fallback group key that lacked a dedicated implementation. This fix ensures accurate aggregation results for queries relying on this specific fallback path.
datafusion/physical-plan · high confidence
Improved nullability tracking for CASE expressions
The \datafusion-expr\ crate now correctly computes the nullability of \CASE\ expressions, ensuring that the resulting schema accurately reflects whether the expression can return null values. This change fixes previous inaccuracies where nullability was not properly inferred, which can lead to more precise query planning and execution.
datafusion/expr · high confidence
Test coverage
1 commit adding/updating tests in datafusion/core/tests/data/partitioned\_table\_arrow; Add ASOF join SQL benchmarks; Add benchmarks for memory accounting, scalar conversion, statistics merging, and hashing; Add optimizer integration tests using insta snapshots; Added FFI test utility for session context and codec setup; Added SQL benchmark suite for null-aware NOT IN joins; Added SQL integration tests for decimal parsing; Added SQL logic tests for Spark LIKE and ILIKE predicates; Added SQL logic tests for Spark concat, reverse, and size functions; Added SQL logic tests for Spark conditional functions; Added SQL logic tests for Spark-compatible float and integer to timestamp casting; Added SQL logic tests for Spark-compatible hash functions; Added TPC-DS benchmark SQL queries to the test suite; Added TPC-H benchmark SQL logic tests; Added TPC-H query plan regression tests; Added aggregation fuzzer test framework; Added benchmark suite for Parquet RowFilter skip optimization; Added benchmarks for Parquet metadata statistics and filter pushdown; Added benchmarks for Spark-compatible string, math, and hash functions; Added benchmarks for TopK pushdown through joins; Added benchmarks for inexact sort pushdown scenarios; Added comprehensive test suite for Substrait integration; Added execution tests for cooperative scheduling, batch splitting, and Arrow registration; Added fuzz tests for EquivalenceProperties ordering and projection; Added integration tests for FFI components; Added integration tests for Parquet content-defined chunking, custom readers, dynamic pruning, encryption, expression adapters, external access plans, file statistics, and filter pushdown; Added integration tests for configuration, set comparison, and TPC-DS planning; Added integration tests for datafusion-cli; Added macro hygiene tests; Added memory limit validation tests for sort, nested loop join, and sort-merge join; Added memory-limit regression tests for spilling and resource exhaustion; Added micro-benchmarks for optimizer rules; Added optimizer integration tests for SQL query planning; Added partitioned CSV test data for sqllogictests; Added regression tests for async task tracing; Added snapshot tests for datafusion-cli features; Added test data and documentation for the Substrait integration; Added test data fixtures for CSV and JSON parsing features; Added test data for CSV parsing with null values; Added test data for recursive CTE scenarios; Added test data for regex functions; Added test fixtures for CSV files with empty files and null columns; Added test infrastructure for SQL planning and context providers; Added test utilities for partitioned CSV scanning, object store concurrency, and system variables; Added tests for CSV schema inference and object store access patterns; Added tests for DML planning, filter pushdown, and statistics propagation; Added tests for FIFO-based unbounded stream processing; Added tests for catalog partition listing and pruning; Added tests for extension type pretty printing; Added tests for memory catalog schema and table deregistration; Added tests for partition-aware metrics snapshots; Added tests for the DataFusion expression API; Added tests for user-defined SQL planning and execution extensions; Added tests to validate SQL example table formatting; Automated string function testing across Utf8, LargeUtf8, and StringView types; Comprehensive SQL logic tests for DataFusion regular expression functions; Consolidated Substrait integration test infrastructure; Consolidated physical optimizer test suite; Expanded PostgreSQL compatibility test coverage; Expanded SQL logic tests for date, time, and interval arithmetic; Expanded Spark SQL string function test coverage; Expanded Spark-compatible datetime function tests; Expanded Spark-compatible math function test coverage; Expanded and refined array function test coverage in SQL logic tests; Expanded fuzz testing coverage for core DataFusion operations; Expanded physical plan round-trip tests for DataFusion Proto; Expanded test coverage for aggregate functions with dictionary columns, nulls, and stricter nested nullability; Massive expansion of SQL logic test coverage; Migrated DataFrame tests to insta snapshot testing; New SQL and Data-Source Benchmarks for Aggregate, Parquet, and CSV Workloads; New SQL integration tests for DataFusion core; New benchmarks for aggregate functions; New benchmarks for crypto, datetime, encoding, and math functions; New benchmarks for nested array functions; New benchmarks for physical expression performance; New benchmarks for physical-plan performance analysis; New test suite for SQL planning, diagnostics, and unparsing; Updated TPC-H benchmark answer files for SQL logic tests.
Dependencies
Update sqlparser requirement from 0.51.0 to 0.52.0
The sqlparser dependency has been upgraded from version 0.51.0 to 0.52.0, bringing in the latest parser updates and potential SQL syntax or semantic changes provided by the upstream library.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 70 → 73 (+2.7)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.
Lenses
- Code Health 83 → 83 (+0.0)
- Architecture 97 → 95 (-2.7)
- Maturity 67 → 67 (+0.0)
- Readiness 75 → 77 (+1.6)
- Security 64 → 71 (+7.0)
- Performance 100 (new)
Resolved (191)
- Change coupling: listing_schema.rs ↔ mod.rs (datafusion/catalog/src/listing_schema.rs)
- Dependency advisory scan runs only on code events
- Documentation: no installation or build instructions (README.md)
- Documentation: no project overview (README.md)
- Documentation: no project overview (datafusion/physical-expr-adapter/README.md)
- Documentation: no project overview (datafusion/sqllogictest/README.md)
- Duplicated block (10 lines × 2) (datafusion/functions/src/regex/regexpcount.rs)
- Duplicated block (10 lines × 2) (datafusion/functions/src/string/levenshtein.rs)
- Duplicated block (10 lines × 3) (datafusion/physical-plan/src/joins/hash_join/exec.rs)
- Duplicated block (10–12 lines × 3) (datafusion/functions/src/regex/regexpcount.rs)
- Duplicated block (11 lines × 2) (datafusion/functions-nested/src/array_add.rs)
- Duplicated block (11 lines × 2) (datafusion/functions-nested/src/range.rs)
- Duplicated block (11 lines × 2) (datafusion/functions/src/core/arrow_field.rs)
- Duplicated block (11 lines × 3) (datafusion/functions/src/regex/regexpcount.rs)
- Duplicated block (11 lines × 3) (datafusion/physical-plan/src/aggregates/hash_stream.rs)
- Duplicated block (11 lines × 3) (datafusion/physical-plan/src/aggregates/ordered_final_stream.rs)
- Duplicated block (11 lines × 4) (datafusion/physical-plan/src/aggregates/ordered_final_stream.rs)
- Duplicated block (11–12 lines × 2) (datafusion/functions-nested/src/array_add.rs)
- Duplicated block (12 lines × 2) (datafusion/datasource-csv/src/file_format.rs)
- Duplicated block (12 lines × 2) (datafusion/functions/src/regex/regexpcount.rs)
- …and 171 more
New (163)
- AsOfJoinExecNode::deserialize (cognitive 25) (datafusion/proto-models/src/generated/pbjson.rs)
- AsOfJoinExecNode::deserialize (cyclomatic 22) (datafusion/proto-models/src/generated/pbjson.rs)
- AsOfJoinNode::deserialize (cognitive 28) (datafusion/proto-models/src/generated/pbjson.rs)
- AsOfJoinNode::deserialize (cyclomatic 25) (datafusion/proto-models/src/generated/pbjson.rs)
- AsOfJoinNode::serialize (cognitive 16) (datafusion/proto-models/src/generated/pbjson.rs)
- AsOfJoinNode::serialize (cyclomatic 17) (datafusion/proto-models/src/generated/pbjson.rs)
- BinaryTypeCoercer::signature_inner (cognitive 16) (datafusion/expr-common/src/type_coercion/binary.rs)
- Change coupling: listing_schema.rs ↔ listing_table_factory.rs (datafusion/catalog/src/listing_schema.rs)
- Change coupling: listing_schema.rs ↔ statement.rs (datafusion/catalog/src/listing_schema.rs)
- ClassTooLong: AsOfJoinExec (datafusion/physical-plan/src/joins/asof_join.rs)
- ClassTooLong: MultiLevelMergeBuilder (datafusion/physical-plan/src/sorts/multi_level_merge.rs)
- CsvWriterOptions::try_from (cognitive 31) (datafusion/proto-common/src/from_proto/mod.rs)
- CsvWriterOptions::try_from (cyclomatic 22) (datafusion/proto-common/src/from_proto/mod.rs)
- DecorrelatePredicateSubquery::rewrite (cognitive 21) (datafusion/optimizer/src/decorrelate_predicate_subquery.rs)
- DerivedRelationBuilder::build (cognitive 17) (datafusion/sql/src/unparser/ast.rs)
- Documentation: no installation or build instructions (datafusion/ffi/README.md)
- Documentation: no usage examples (datafusion/ffi/README.md)
- Duplicated block (10 lines × 2) (datafusion/common/src/hash_utils.rs)
- Duplicated block (10 lines × 2) (datafusion/datasource-csv/src/file_format.rs)
- Duplicated block (10 lines × 2) (datafusion/physical-plan/src/aggregates/ordered_single_stream.rs)
- …and 143 more
Changes since last survey
- 248 commits — 147 feature/other, 101 fixes
By area
- datafusion/physical-plan — 44 commits
- .github/workflows — 20 commits
- datafusion/functions — 17 commits
- datafusion/optimizer — 17 commits
- datafusion/core — 15 commits
- datafusion/sqllogictest — 14 commits
- datafusion/physical-expr — 11 commits
- benchmarks/sql_benchmarks — 10 commits
- datafusion/functions-nested — 10 commits
- datafusion/datasource-parquet — 8 commits
- (root) — 7 commits
- docs/source — 7 commits
- datafusion/functions-aggregate — 6 commits
- datafusion/sql — 6 commits
- datafusion/substrait — 6 commits
- ci/scripts — 5 commits
- datafusion/common — 5 commits
- datafusion/expr — 5 commits
- datafusion/session — 4 commits
- datafusion/datasource — 3 commits
Notable commits
- fix: Fix EliminateLimit not eliminating limits under LogicalPlan::Extension nodes. (#25577)
- fix: Fix Numeric signature coercion to properly handle null types (#24988)
- fix: Fix outdated Python DataFrame API documentation link (#25225)
- fix: Fix panic from avg(x ORDER BY y) (and more) (#25711)
- fix: Fix panic in PercentileContGroupsAccumulator::convert_to_state() (#25755)
- fix: chore(docs): Fix the documentation for TableProvider::scan()'s limit argument (#25397)
- fix: chore(docs): Fix the rendering of an expression in docstring (#25379)
- fix: fix datafusion-cli CI: switch from minio to rustfs (#25706)
- fix: fix(core): reject a DELETE or an UPDATE whose WHERE clause cannot reach the provider (#24657)
- fix: fix(datetime): floor negative scalar timestamps in date_trunc (#25430)
- fix: fix(docs): Fix code snippet in LogicalPlan docs (#25377)
- fix: fix(ffi): don't discard downcast-delegating wrappers of foreign plans (#25749)
- fix: fix(physical-plan): CoalescePartitionsExec panic on wasm32-unknown-unknown (#24890)
- fix: fix(physical-plan): honor distinct soft limits in SingleHashAggregateStream (#25158)
- fix: fix(proto): preserve CSV sink writer options (#25058)
- fix: fix(proto): preserve Parquet source and sink state (#25057)
- fix: fix(substrait): consume chained window functions whose default names collide (#25181)
- fix: fix: Correct array_repeat NULL handling (#25230)
- fix: fix: Derive Substrait intersection nullability from every input (#25091)
- fix: fix: Emit equality conditions for Substrait CASE base expressions (#25191)
- …and 228 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
apache/datafusion was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit bdad988f23f4f8a45a71409e7d7f24a326ca909c — the exact code this score is about.
- Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fbec9b1e08c2.