Eventual-Inc/Daft
53.8
Adequate · 29 September 2026
295.9k
lines of production code
Rust
with Python
2
measurements over time
What this system is
Daft is a high-performance, distributed data processing engine built in Rust with a Python API, designed for scalable data engineering and AI workloads. It provides a unified interface for reading, transforming, and writing data across diverse formats and storage systems, including Parquet, Delta Lake, Iceberg, and various cloud object stores. The system supports complex analytical queries via SQL and DataFrame operations, integrates natively with AI model providers for embedding and classification, and offers robust distributed execution through Ray with detailed observability and dashboarding.
How it got here
2022–2024 — Native engine rewrite and arrow-rs migration
96 changes.
This period focused on a comprehensive architectural overhaul, migrating the core data structures from Arrow2 to arrow-rs and replacing the legacy execution runners with a new native and Flotilla-based engine. The work established a robust Rust foundation featuring a new logical plan builder, unified I/O architecture, and native readers for formats like Parquet, CSV, and JSON. It also introduced extensive new capabilities including SQL support, catalog integrations, and a wide array of expression functions, all underpinned by significant improvements in performance and type safety.
2025 — Flotilla distributed engine and observability
87 changes.
This period focused on replacing the distributed execution engine with the new Rust-based Flotilla runner, introducing a comprehensive subscriber framework for real-time query observability and a dedicated web dashboard. It also involved a major architectural overhaul of core data structures, including the introduction of RecordBatch, MicroPartition, and ScalarColumn, alongside extensive native implementations for joins, aggregations, and shuffle operations.
2026 — Native extensions and checkpoint infrastructure
25 changes.
This period focused on establishing a robust native extension SDK for custom Rust and C++ functions, alongside implementing a new S3-backed checkpoint store to enable reliable two-phase commit sinks. Significant engineering efforts also included rewriting the Parquet reader using the arrow-rs public API and expanding distributed execution capabilities with new join strategies and partition reference types.
Features
Add AI benchmarking suite for Daft, Ray Data, and Spark
Added a new benchmarking suite in the \benchmarking/ai\ directory that compares the performance of Daft against Ray Data and Spark across four multimodal workloads: audio transcription (Whisper), document embedding (MiniLM), image classification (ResNet18), and video object detection (YOLO11). The entry includes the benchmark definitions, cluster configuration files for AWS (g6.xlarge), and the specific implementation scripts for each engine to facilitate reproducible performance comparisons.
benchmarking/ai · high confidence
Add Apache Paimon read and write support
Users can now read from and write to Apache Paimon tables using the new \read\_paimon\ function and \PaimonDataSink\. This integration supports S3-compatible and OSS storage backends, automatically converting catalog options to the appropriate IO configuration. Reads leverage Daft's native Parquet reader for performance, falling back to pypaimon's native reader for primary-key tables requiring LSM-tree merges or non-Parquet formats. Writes support both append and overwrite modes, with automatic type casting and a patch to handle complex types (lists, maps, structs) during statistics computation.
daft/io/paimon · high confidence
Add Arrow FFI utilities for PyArrow interoperability
Introduces a new \src/common/arrow-ffi\ crate providing Rust utilities to convert between PyArrow objects and \arrow-rs\ types. This enables direct data exchange via the Arrow PyCapsule Interface (using \\_\_arrow\_c\array\\_\) and fallbacks to legacy C-FFI methods, allowing Python Arrow data to be consumed as Rust \RecordBatch\ and \ArrayData\ structures without requiring external dependencies like \arrow-pyarrow\.
src/common/arrow-ffi · high confidence
Add Avro file format read and write support
Users can now read from and write to Avro files. This change introduces the \daft-avro\ crate, providing \read\_avro\ and \read\_avro\_schema\ functions for Python and Rust, along with \write\_record\_batch\_to\_avro\ for output. The implementation supports Null and Deflate compression codecs, handles both local and remote file access via the existing IO client, and includes schema conversion between Daft and Arrow types.
src/daft-scan · high confidence
Add BPE-based text tokenization and detokenization functions
The \daft-functions-tokenize\ module now provides \tokenize\_encode\ and \tokenize\_decode\ functions, enabling users to convert text strings into integer token lists and back using Byte Pair Encoding (BPE). These functions support built-in tokenizers (cl100k\_base, o200k\_base, p50k\_base, p50k\_edit, r50k\_base) and allow loading custom token files from remote paths. Additionally, the Llama 3 special token set is now supported for encoding and decoding.
src/daft-functions-tokenize · high confidence
Add Bigtable data sink for writing to Google Cloud Bigtable
Users can now write data to Google Cloud Bigtable using the new \BigtableDataSink\. This feature allows specifying a project, instance, and table, along with row key and column family mappings. The sink automatically handles type compatibility by serializing non-integer, non-binary, and non-string columns to JSON (with an option to disable this), filters out rows with invalid or null row keys, and supports emulator configuration for local testing.
daft/io/bigtable · high confidence
Add ClickBench SQL queries for benchmarking
Added the standard ClickBench SQL query suite (\queries.sql\) to the \benchmarking/clickbench\ directory, enabling performance benchmarking against this specific workload.
benchmarking/clickbench · high confidence
Add ClickHouse data sink for writing DataFrames
Users can now write DataFrames to ClickHouse databases using the new \ClickHouseDataSink\. This change introduces the \daft.io.clickhouse\ module, exposing \ClickHouseDataSink\ which accepts connection parameters (host, port, user, password, database) and table name. The sink writes data via the \clickhouse\_connect\ library and returns a summary of total written rows and bytes.
daft/io/clickhouse · high confidence
Add DDSketch serialization to Arrow format
The daft-sketch crate now includes functionality to serialize and deserialize DDSketch data structures into and from Arrow arrays. This allows users to efficiently exchange sketch data using the Arrow columnar format, supporting round-trip conversion for single sketches, null values, and empty vectors.
src/daft-sketch · high confidence
Add Float16 support to NdArray and unify Python type conversions
The NdArray module now supports half-precision floating-point data (F16) alongside existing integer and single/double-precision types, enabling users to work with lower-precision numerical data. Additionally, the module introduces a unified conversion layer between Rust NdArray and Python NumpyArray types, ensuring consistent handling of all supported data types during interop.
src/common/ndarray · high confidence
Add GPU device detection utilities
The \daft/internal/gpu.py\ module has been added to provide low-level utilities for detecting available CUDA devices. It includes a function to count devices via NVML and another to retrieve the list of visible CUDA devices, respecting the \CUDA\_VISIBLE\_DEVICES\ environment variable when set.
daft/internal · high confidence
Add Gravitino catalog integration for Iceberg and PostgreSQL tables
Users can now connect Daft to Apache Gravitino catalogs to read Iceberg and PostgreSQL tables. This change introduces the \GravitinoCatalog\ and \GravitinoTable\ classes, along with a \GravitinoClient\ that supports simple and OAuth2 authentication, enabling access to tables and filesets managed by Gravitino.
daft/catalog/\\gravitino · high confidence
Add HLL cardinality and sketch percentile aggregation functions
Users can now use the \hll\_cardinality\ function to estimate the number of distinct values in a column using HyperLogLog sketches, and the \sketch\_percentile\ function to compute approximate percentiles from sketch data. The \hll\_cardinality\ function accepts a single HLL sketch input and returns a UInt64 count, while \sketch\_percentile\ accepts a struct input containing sketch data and returns either a single Float64 value or a list of Float64 values depending on the number of requested percentiles and the \force\_list\_output\ flag.
src/daft-dsl/src/functions/sketch · high confidence
Add JSON serialization and deserialization support
Users can now serialize and deserialize data to and from JSON format. This change introduces \serialize\ and \deserialize\ functions for JSON, allowing conversion of Series to JSON strings and parsing JSON strings back into typed Series. It also includes \try\_serialize\ and \try\_deserialize\ variants that handle parsing errors by inserting nulls instead of failing, with specific type coercion logic for integers, floats, booleans, and strings during deserialization.
src/daft-functions-serde/src/format · high confidence
Add JSON serialization for logical plans
Users can now serialize Daft logical plans into a JSON representation via the new \repr\_json\ capability. The implementation in \src/daft-logical-plan/src/display/json.rs\ introduces a \to\_json\_value\ function and a \JsonVisitor\ that traverse the plan tree, converting each node (such as Source, Project, Filter, Join, AsofJoin, Sample, etc.) into a structured JSON object containing relevant parameters like predicates, join types, and batch sizes. This enables programmatic inspection, debugging, and external tooling integration by providing a machine-readable format of the execution plan.
src/daft-logical-plan/src/display · high confidence
Add LM Studio provider with configurable embedding dimensions
Users can now connect Daft to a local LM Studio instance via a new LMStudioProvider, which defaults to the local server at localhost:1234 and automatically appends the /v1 path segment to the base URL. This provider extends the existing OpenAI provider logic and adds support for specifying the dimensions parameter when creating text embeddings, allowing users to control the output size of their embedding models.
_daft/ai/lm\studio · high confidence
Add LM Studio text embedder protocol with dynamic dimension detection
Users can now embed text using LM Studio models via a new \LMStudioTextEmbedder\ protocol. This implementation automatically detects the embedding dimensions of the loaded model by probing the local server, allowing it to support models with varying dimension sizes. It reuses the existing OpenAI text embedder logic for the actual embedding requests while handling LM Studio-specific configuration such as batch token limits and input text token limits.
_daft/ai/lm\studio/protocols · high confidence
Add TPC-DS benchmarking infrastructure for local and Ray-based execution
Introduces a new benchmarking suite for TPC-DS queries in the \benchmarking/tpcds\ directory. This includes a data generation script (\datagen.py\) to create Parquet datasets using DuckDB, helper utilities for parsing query ranges and loading bindings, and a main runner (\\_\main\\_.py\) that executes queries locally with validation against DuckDB. Additionally, a Ray entrypoint (\ray\_entrypoint.py\) is provided to run individual queries on a Ray cluster, capturing execution and planning statistics. This enables users to benchmark Daft's SQL performance against the TPC-DS standard.
benchmarking/tpcds · high confidence
Add Transformers provider for AI capabilities
Introduces the \TransformersProvider\ in \daft/ai/transformers\, enabling users to leverage local Hugging Face models for text and image embedding, text and image classification, and prompting. The provider includes default model configurations (e.g., \sentence-transformers/all-MiniLM-L6-v2\ for text embedding) and validates that required dependencies like \torch\ and \numpy\ are available before initialization.
daft/ai/transformers · high confidence
Add Transformers provider for AI protocols
The \daft/ai/transformers/protocols\ module now provides a complete implementation of the AI protocol interfaces using the Hugging Face \transformers\ and \sentence-transformers\ libraries. This adds support for zero-shot image classification (\TransformersImageClassifier\), image embedding (\TransformersImageEmbedder\), text classification (\TransformersTextClassifier\), text embedding (\TransformersTextEmbedder\), and text generation prompting (\TransformersPrompter\). The implementation includes specific behaviors such as using a global lock to prevent concurrent model loading meta-tensor errors, supporting text file inputs for the prompter, and allowing explicit specification of embedding dimensions for text models.
daft/ai/transformers/protocols · high confidence
Add WARC file reader with structured column extraction
Introduces a new WARC reader in the \daft-warc\ crate that parses WARC files and exposes their contents as a structured RecordBatch. The reader extracts key metadata fields—including WARC type, target URI, date, content length, and identified payload type—into dedicated columns, alongside the raw content and header data, enabling direct analysis of web archive data within Daft.
src/daft-warc · high confidence
Add ability to extract fields from Struct columns
Users can now extract specific fields from struct-typed columns using the new \get\ function. This change introduces a \StructExpr::Get\ expression and its corresponding \GetEvaluator\, allowing users to retrieve the value of a named field from a struct, with proper schema validation and error handling for missing fields or non-struct inputs.
src/daft-dsl/src/functions/struct\ · high confidence_
Add cosine, dot product, and Euclidean distance functions
Users can now compute vector similarity and distance metrics directly in expressions. This change introduces three new scalar functions: \cosine\_distance\, \dot\_product\, and \euclidean\_distance\, which are registered in the function registry and accept two vector expressions as inputs.
src/daft-functions/src/distance · high confidence
Add great\_circle\_distance function for geographic calculations
Users can now calculate the great-circle distance in meters between two geographic points using the new \great\_circle\_distance\ function. This feature, implemented in \src/daft-geo\, accepts four numeric arguments (latitude and longitude for both the start and end points) and returns the distance using the haversine formula. The function handles null inputs and invalid coordinates by returning null values, and is registered as part of the spatial functions module.
src/daft-geo · high confidence
Add map\_get and map\_keys functions for Map type columns
Users can now extract specific values or keys from Map type columns using the new \map\_get\ and \map\_keys\ functions. The \map\_get\ function retrieves the value associated with a specified key from a map, while \map\_keys\ returns a list of all keys present in the map. These functions are implemented in the DSL layer and integrate with the existing function evaluation framework.
src/daft-dsl/src/functions/map · high confidence
Add parquet column and table statistics conversion
The \src/daft-parquet/src/statistics\ module now provides logic to convert arrow-rs parquet statistics into Daft's internal \ColumnRangeStatistics\ and \TableStatistics\. This enables Daft to read min/max values for columns (including Boolean, Int32, Int64, Int96, Float, Double, and ByteArray types) from Parquet files, supporting logical type conversions for dates, timestamps, and decimals to improve query optimization and pruning.
src/daft-parquet/src/statistics · high confidence
Add support for reading Apache Hudi tables
Users can now read Apache Hudi tables using the new \daft.read\_hudi\ function, which creates a DataFrame from a Hudi table URI (supporting local paths and remote object stores like S3 or GCS). This change introduces a new \daft/io/hudi\ module containing a \HudiDataSource\ implementation that handles Hudi-specific metadata parsing, timeline management, partition pruning, and column statistics for optimized scanning. The reader also supports optional checkpointing for progress tracking across runs, requiring the Ray runner.
daft/io/hudi · high confidence
Add utility functions for date, duration, and engine name formatting
A new utility module has been added to the frontend library to support dashboard UI enhancements. This includes functions for converting epoch timestamps to human-readable dates and durations, as well as a helper to map internal runner identifiers (such as 'Native' and 'Ray') to user-facing engine names ('Swordfish' and 'Flotilla').
src/daft-dashboard/frontend/src/lib · high confidence
Added ASOF join benchmarking suite
A new benchmarking suite for the \join\_asof\ operation has been added to \benchmarking/asof\_join\. It includes scripts to generate reproducible Parquet datasets at small, medium, and large scales with clustered timestamps and Zipf-skewed entity distributions, a benchmark runner that measures wall time and memory usage (via memray) for both native and Ray runners, and a Ray cluster configuration for distributed AWS execution.
_benchmarking/asof\join · high confidence
Added Common Crawl text extraction benchmark
A new benchmarking script has been added to extract plain text from Common Crawl WARC files. The script uses a Daft UDF to process HTML content, handling encoding fixes and filtering out excessively large documents, then writes the extracted text to Parquet.
_benchmarking/common\crawl · high confidence
Added TPC-DS benchmark queries 01–28
Added 28 SQL benchmark queries (01.sql through 28.sql) to the TPC-DS benchmarking suite, providing standard analytical workloads for performance testing.
benchmarking/tpcds/queries · high confidence
Added TPC-H benchmark queries
Added 22 SQL files (01.sql through 22.sql) to the \benchmarking/tpch/queries\ directory, providing the standard TPC-H benchmark suite for performance testing.
benchmarking/tpch/queries · high confidence
Added vLLM benchmarking scripts and configuration
Added a new set of benchmarking scripts in the \benchmarking/vllm\ directory to evaluate batch inference performance using vLLM. The change includes a configuration file (\config.py\) defining model parameters (defaulting to Qwen/Qwen3-8B) and a data generation notebook (\generate\_data.ipnb\) for creating test datasets. Several execution scripts are provided to test different batching strategies: naive batch inference (\naive-batch.py\, \naive-batch-sorted.py\), continuous batching with prefix caching (\continuous-batch.py\, \continuous-batch-sorted.py\), prefix bucketing (\prefix-bucketing.py\), and Ray Data integration (\ray-data.py\).
benchmarking/vllm · high confidence
Configurable OpenTelemetry tracing and metrics via environment variables
The tracing subsystem now supports OpenTelemetry (OTLP) configuration through standard environment variables (e.g., \OTEL\_EXPORTER\_OTLP\_ENDPOINT\, \OTEL\_EXPORTER\_OTLP\_PROTOCOL\) and a new \DAFT\_TRACE\ variable for console output formatting. Users can now enable distributed tracing, metrics, and logs export by setting these environment variables, with the system automatically detecting endpoints and protocols. The \DAFT\_TRACE\ variable allows users to control console log verbosity and format (compact, pretty, or JSON) for local debugging, while the underlying OTLP exporters handle the actual data transmission to configured backends.
src/common/tracing · high confidence
Dashboard application root layout and styling established
The dashboard frontend now includes its core application shell, featuring a dark theme with Daft-specific magenta accents defined in the new global styles. The root layout wires up essential providers for notifications, server state, and tooltips, while the home page automatically redirects users to the queries section.
src/daft-dashboard/frontend/src/app · high confidence
Distributed execution dashboard and progress bar integration
The Python bindings for distributed execution now include a dashboard subscriber and a progress bar component. The dashboard subscriber listens to task lifecycle events (submitted, scheduled, completed, failed, cancelled) and emits per-operator statistics and task counters to a configured dashboard URL, allowing users to monitor distributed query execution in real-time. Additionally, a progress bar is integrated into the Python API, providing visual feedback on task completion status for distributed pipelines.
src/daft-distributed/src/python · high confidence
Distributed join execution engine for Flotilla
The distributed pipeline node implementation for joins has been rewritten to support the Flotilla execution engine. This change introduces dedicated pipeline nodes for Hash, Broadcast, Sort-Merge, Cross, and Asof joins, along with a KeyFiltering join strategy. The new implementation handles distributed execution logic including range repartitioning, shuffle reads, and task building via SwordfishTaskBuilder, replacing the previous approach with a modular, engine-specific node structure that also includes comprehensive runtime statistics tracking for join operations.
_src/daft-distributed/src/pipeline\node/join · high confidence
Enable HTML visualization for NumPy arrays and PIL images
The visualization module now supports rendering NumPy arrays and PIL images directly in HTML contexts. When these libraries are available, NumPy arrays are displayed with their shape and dtype, while PIL images are rendered as base64-encoded thumbnails (max 128px height) using a lazy registration pattern to avoid unnecessary imports.
daft/viz · high confidence
Enable Rust-based catalogs to interact with Python table and function implementations
Users can now use Rust-based catalog implementations (such as the in-memory catalog) to manage tables and functions defined in Python. This change introduces \PyCatalogWrapper\ and \PyTableWrapper\ in the \src/daft-catalog/src/python\ module, allowing the Rust catalog layer to call into Python objects for operations like creating, listing, and retrieving tables and functions, as well as appending or overwriting data via logical plans.
src/daft-catalog/src/python · high confidence
Experimental Turbopuffer write sink with multi-namespace support
Users can now write Daft DataFrames to Turbopuffer using the new experimental \TurbopufferDataSink\. This sink supports writing to a single namespace or dynamically partitioning data across multiple namespaces using an expression. It handles ID and vector column mapping, allows custom client and write parameters, filters out null IDs, and gracefully handles transient API errors to ensure writes continue where possible.
daft/io/turbopuffer · high confidence
Flight shuffle backend and pre-shuffle merge optimization for distributed shuffles
The distributed shuffle pipeline now supports a Flight-based shuffle backend alongside the existing Ray backend, allowing users to switch to \flight\_shuffle\ via configuration to avoid head-node memory pressure on large shuffles. The system now includes a configurable pre-shuffle merge strategy that merges data on workers before the shuffle phase when the partition product exceeds a threshold, reducing shuffle overhead. Additionally, the Flight backend implementation optimizes coordinator memory usage during the gather phase by folding output streams into per-server input lists, keeping memory consumption at O(map\_tasks + partitions) rather than O(map\_tasks × partitions).
_src/daft-distributed/src/pipeline\node/shuffles · high confidence
Initial frontend configuration for the Daft dashboard
The frontend environment for the Daft dashboard is now established with a standard Next.js setup. This includes configuration for TypeScript, ESLint, Prettier, and PostCSS with Tailwind CSS. The styling system is configured to use the 'new-york' style from shadcn/ui with a 'Geist Mono' font family and specific design tokens for colors and borders. Security hardening is applied via .npmrc settings that ignore scripts and enforce exact dependency versions.
src/daft-dashboard/frontend · high confidence
Introduce AI functions module with embedding, classification, and prompt capabilities
The new \daft.functions.ai\ module exposes AI-powered expressions including \embed\_text\, \embed\_image\, \classify\_text\, \classify\_image\, and \prompt\. Users can now generate text and image embeddings, classify content, and interact with LLMs directly within Daft expressions, specifying providers (such as OpenAI or Transformers) and models. The module includes a compatibility layer for Google Colab to resolve Pydantic serialization issues with cloudpickle, ensuring these functions work reliably in that environment.
daft/functions/ai · high confidence
Introduce Daft CLI with dashboard start/stop commands
Users can now control the Daft dashboard directly from the Python environment via a new CLI interface. The \daft.cli\ module exposes a \dashboard\ subcommand that supports \start\ and \stop\ actions. The \start\ command accepts options to bind to a specific address and port, enable verbose logging, and run the dashboard in daemon mode (background process) on Unix systems. It also supports importing query data from an event log file. The \stop\ command allows users to terminate a running background dashboard instance using a PID file.
src/daft-cli · high confidence
Introduce Daft Quickstart Helm chart for Kubernetes deployment
Adds a new Helm chart (\k8s/charts/quickstart\) that enables running Daft data processing workloads on Kubernetes. The chart supports two modes: a simple single-job execution for quick prototyping and a distributed mode that provisions a Ray cluster (head and worker deployments) for large-scale processing. It includes templates for the Ray head service (exposing the dashboard on port 8265 and metrics on 8080), optional Prometheus and Grafana monitoring sidecars, and a Job resource that supports either script injection or custom container images.
k8s/charts/quickstart · high confidence
Introduce Flight and Ray partition reference types for Python integration
The \daft-partition-refs\ crate now exposes concrete \FlightPartitionRef\ and \RayPartitionRef\ types alongside their Python bindings (\PyFlightPartitionRef\, \PyFlightPartitions\, \RayPartitionRef\). These types implement the \Partition\ trait and are registered in the Python module, enabling Python code to interact with partition references originating from Flight (shuffle) and Ray execution backends.
src/daft-partition-refs · high confidence
Introduce Flight-based shuffle transport with Rust client and server
The shuffle subsystem now supports a new Flight-based transport layer implemented in Rust. This adds a \ShuffleFlightClient\ for reading partition data from remote shuffle servers via gRPC, a \FlightClientManager\ to handle connection pooling, and a \ShuffleFlightServer\ to serve shuffle data. The server supports both local in-process reads and remote gRPC reads, utilizing a \FlightDataStreamReader\ to stream Arrow IPC data directly from disk without decoding into RecordBatches first. Additionally, a \oneshot\_writer\ is introduced to write all output partitions of a single map task into one combined IPC stream file, enabling efficient ranged reads by the Flight server.
src/daft-shuffles · high confidence
Introduce Flotilla distributed task scheduler and dispatcher
The distributed execution engine now uses a new Flotilla scheduler and dispatcher to manage task execution. The dispatcher handles submitting tasks to workers, awaiting results, and managing task lifecycle events (success, failure, cancellation, worker death) while updating statistics. A generic Worker and WorkerManager trait system allows for flexible backend integration, including a new LocalSwordfishWorker for in-process testing that mimics Ray worker behavior. Task metadata, including source information and resource requests, is now explicitly tracked and passed through the scheduling pipeline.
src/daft-distributed/src/scheduling · high confidence
Introduce LocalPhysicalPlan for native task execution
The \src/daft-local-plan\ crate now defines the \LocalPhysicalPlan\ type, which represents the executable plan for a single distributed task. This plan includes operators for scanning data (in-memory, physical files, glob patterns), transforming data (project, filter, sort, sample, explode, limit), and aggregating (hash aggregate, window functions, distinct). It also supports distributed coordination nodes like shuffle reads, repartitioning, and joins (hash, sort-merge, asof, cross). The crate provides a translation layer from the logical plan to this local physical representation and exposes Python bindings (\PyLocalPhysicalPlan\) to construct and serialize these plans for the native runner.
src/daft-local-plan · high confidence
Introduce MicroPartition as the core in-memory data structure
This change introduces the \MicroPartition\ struct and its associated \MicroPartitionSet\ as the fundamental in-memory representation for tabular data in Daft. Users can now interact with data via this new internal primitive, which supports loading from Arrow arrays, reading from CSV, JSON, and Parquet sources, and managing partitioned data sets. The implementation provides methods for schema inspection, column access, and basic operations like concatenation and length calculation, serving as the underlying engine for higher-level DataFrame operations.
src/daft-micropartition/src · high confidence
Introduce PySeries Python bindings with Arrow PyCapsule Interface support
The \src/daft-core/src/python\ module now exposes the core \Series\ data structure to Python via \PySeries\. This change adds explicit Python bindings for creating Series from Python lists and PyArrow arrays, converting Series back to Python lists (with configurable map handling), and exporting data to Arrow. Crucially, it implements the Arrow PyCapsule Interface (\\_\_arrow\_c\schema\\\ and \\\_arrow\_c\array\\_\), enabling zero-copy data exchange with other Arrow-compatible libraries.
src/daft-core/src/python · high confidence
Introduce ResourceRequest struct for query fragment resource management
The \src/common/resource-request\ crate now provides a dedicated \ResourceRequest\ struct to define and manage resource requirements (CPUs, GPUs, and memory) for query fragment tasks. This change introduces validation logic to ensure GPU requests are non-negative and integer-valued when greater than one, and adds methods to combine, compare, and serialize these requests. Users can now explicitly specify resource constraints for operations, enabling the scheduler to better allocate hardware resources and prevent pipeline incompatibilities between CPU and GPU-bound tasks.
src/common/resource-request · high confidence
Introduce Rust-based in-memory catalog implementation
Users can now utilize a new in-memory catalog implementation written in Rust, located in \src/daft-catalog/src/impls/memory.rs\. This \MemoryCatalog\ supports creating, listing, and dropping tables and namespaces (with single-level namespace support only), as well as registering and retrieving functions. It is exposed via the \impls\ module in \mod.rs\ and provides a thread-safe, in-memory storage backend for catalog operations.
src/daft-catalog/src/impls · high confidence
Introduce S3-backed checkpoint store with Parquet key encoding
The \src/daft-checkpoint/src/impls\ module now provides an \S3CheckpointStore\ implementation for persisting checkpoint data to S3-compatible storage. This implementation introduces a new serialization format for checkpoint keys, encoding them as single-column Parquet files via \keys\_codec.rs\ to support efficient round-trip conversion of key series. The store manages checkpoint lifecycle states (Checkpointed, Committed) through manifest files and handles file metadata staging. Additionally, the implementation includes a fix for Windows compatibility by using \strip\_file\_uri\_to\_path\ when writing local files, ensuring correct path handling for \file://\ URIs.
src/daft-checkpoint/src/impls · high confidence
Introduce SQL image processing functions and structured error reporting
The SQL module now supports image manipulation via new \image\_encode\, \image\_decode\, \image\_crop\, \image\_resize\, and \image\_to\_mode\ functions, allowing users to process image data directly within SQL queries. Additionally, the SQL planner now provides detailed error reporting with line and column context (caret errors) for SQL parsing and execution failures, improving the debugging experience for users.
src/daft-sql · high confidence
Introduce ScalarColumn for O(1) memory storage of repeated values
The \RecordBatch\ column implementation now supports a \ScalarColumn\ variant that stores a single literal value repeated for a given logical length. This allows row-preserving operations like slicing, filtering, and broadcasting to adjust the logical length without allocating memory for the expanded array, significantly reducing memory usage for broadcasted literals. The full series is only materialized on demand via \as\_materialized\_series\ and cached for subsequent access.
src/daft-recordbatch/src · high confidence
Introduce TPC-H benchmarking suite with distributed Ray support
Adds a new TPC-H benchmarking framework under \benchmarking/tpch\ that executes all 22 standard questions using Daft DataFrames. The suite supports both local execution and distributed runs via Ray jobs, including a warmup step to mitigate cold-start latency. It provides comprehensive metrics collection (wall-time per query, environment details, commit hashes) exported to CSV, and includes profiling capabilities via VizTracer and Ray timelines. The package also features automated data generation (CSV to Parquet conversion) and a utility to optimize S3 read performance by sub-partitioning object prefixes.
benchmarking/tpch · high confidence
Introduce UDF v2 with async, batch, and aggregate capabilities
The \daft/udf\ module has been replaced with a new v2 implementation that supports asynchronous execution, batch processing, and user-defined aggregate functions (UDAFs). Users can now decorate functions with \@daft.func\ to enable async concurrency via \max\_concurrency\, process data in batches using \@daft.func.batch\, or define stateful aggregations with \@daft.udaf\. The new system also adds resource controls (CPU/GPU limits), retry logic (\max\_retries\, \on\_error\), and custom metric tracking for UDFs.
daft/udf · high confidence
Introduce \`daft.sql\` for executing SQL queries on DataFrames and external databases
This change adds the \daft.sql\ module, exposing \daft.sql()\ and \daft.sql\_expr()\ to users. \daft.sql()\ allows executing SQL queries against Python DataFrame variables (automatically registered as tables) or explicit CTE bindings, returning a new DataFrame. \daft.sql\_expr()\ enables parsing SQL snippets into Daft Expressions for use in DataFrame operations. The module also includes \SQLConnection\ and \SQLScanOperator\ to support reading from external SQL databases via SQLAlchemy or ConnectorX, handling dialect translation (e.g., mssql to tsql), schema inference, and partitioned scanning with configurable bound strategies.
daft/sql · high confidence
Introduce checkpoint store for tracking staged and committed data files
A new \daft-checkpoint\ crate provides a \CheckpointStore\ trait and an S3-backed implementation to track source keys and output file metadata during distributed execution. This enables two-phase commit (2PC) sinks like Iceberg and Delta Lake to atomically commit only the files that were successfully checkpointed, while also allowing the system to skip already-processed rows on re-runs. The store exposes a Python API (\daft.daft.CheckpointStore\) for listing checkpoints, retrieving file metadata, and marking checkpoints as committed, with support for staged, checkpointed, and committed lifecycle states.
src/daft-checkpoint/src · high confidence
Introduce checkpoint-based idempotent commits for Delta Lake and Iceberg writes
The \daft/dataframe\ module now includes a checkpointing infrastructure to ensure idempotent catalog commits for Delta Lake and Iceberg writers. New utilities in \\_checkpoint\_commit.py\ handle the storage and decoding of file metadata via Arrow IPC streams, allowing the system to safely recover from failures or re-run writes without duplicating catalog entries. This change is supported by updated write logic in \dataframe.py\ and new display and preview formatting options in \display.py\ and \preview.py\.
daft/dataframe · high confidence
Introduce column range statistics for predicate pushdown optimization
The \daft-stats\ crate now tracks minimum and maximum values for columns via \ColumnRangeStatistics\, enabling the query engine to evaluate filter expressions against metadata without scanning data. This change adds arithmetic, comparison, and logical operators to range statistics, allowing the planner to short-circuit scans when predicates can be proven false based on column bounds. It also introduces a strict allowlist of supported data types (e.g., numeric, string, temporal) while explicitly excluding complex types like lists, structs, and tensors from range statistics to ensure safe predicate evaluation.
src/daft-stats · high confidence
Introduce common image types and Python bindings
The \src/common/image\ module now provides core image data structures (\Image\, \CowImage\, \BBox\) and serialization support, enabling images to be hashed, resized, cropped, and converted to/from NumPy arrays. It also adds Python bindings that allow PIL images to be extracted and converted into the internal \Image\ format for use within the system.
src/common/image · high confidence
Introduce configurable hash function variants
The \src/daft-hash\ library now exposes a \HashFunctionKind\ enum that allows users to select from multiple hashing algorithms, including MurmurHash3, XxHash32, XxHash64, XxHash3\_64, and Sha1. This change provides the underlying infrastructure for hash-based expressions to support different hashing strategies, enabling users to choose specific hash functions for their data processing needs.
src/daft-hash · high confidence
Introduce core enums for count modes, join types, and operators
The \daft-core\ library now exposes explicit enums for \CountMode\ (All, Valid, Null), \JoinType\ (Inner, Left, Right, Outer, Anti, Semi), \JoinStrategy\ (Hash, SortMerge, Broadcast), \AsofJoinStrategy\ (Backward, Forward, Nearest), and \Operator\ (comparison, arithmetic, logical, etc.). These types are registered with Python via PyO3, allowing users to specify these configurations directly in Python code, and are re-exported in the core prelude for easy access.
src/daft-core/src · high confidence
Introduce daft-decoding for schema inference and data deserialization
The new \daft-decoding\ crate provides core utilities for decoding raw data into Arrow arrays and inferring schema metadata. It includes a \deserialize\ module that handles conversion of byte sequences into typed Arrow arrays (including primitives, decimals, booleans, and UTF-8 strings) and an \inference\ module that determines data types from byte content, supporting detection of dates, times, timestamps (with timezone awareness), integers, floats, and booleans. This module serves as the foundational decoding layer for other components like CSV and JSON readers.
src/daft-decoding · high confidence
Introduce daft-distributed library structure
The new \src/daft-distributed/src/lib.rs\ file establishes the module structure for the distributed computing layer, exposing \pipeline\_node\, \plan\, \scheduling\, \statistics\, and \utils\ modules, while conditionally exposing Python bindings via the \python\ feature gate.
src/daft-distributed/src · high confidence
Introduce daft.File datatype with media-specific subtypes and metadata functions
This change introduces the new \daft.File\ datatype, allowing users to work with files as first-class objects in DataFrames. It includes specific subtypes for video, audio, image, and HDF5 files, enabling media-aware operations. Users can now use scalar expressions like \file\_exists\ and \file\_size\ to check file availability and size without loading content, and the \file()\ function converts strings to file references. The implementation supports byte-range reads for efficient partial file access and includes MIME type detection capabilities.
src/daft-file · high confidence
Introduce experimental subscriber framework for query lifecycle events
This change adds a new \daft.subscribers\ module that provides an extensible framework for observing Daft query execution. It introduces a \Subscriber\ abstract base class with a unified \on\_event\ dispatch mechanism, allowing implementations to handle specific lifecycle events such as query start/finish, optimization phases, execution steps, operator progress, and process-level statistics. The module includes a built-in \EventLogSubscriber\ that writes these events to JSONL files for diagnostics, a Ray-based \EventLogSink\ for distributed logging, and a \launch\ function to start the Daft dashboard server for real-time visualization.
daft/subscribers · high confidence
Introduce internal cloudpickle-based serialization for UDFs
The \daft/pickle\ module has been added to provide a dedicated serialization layer for user-defined functions (UDFs). By embedding a vendored copy of cloudpickle (v3.1.1), Daft can now serialize interactively defined functions, classes, and lambdas for distributed execution, addressing previous limitations in UDF serializability.
daft/pickle · high confidence
Introduce local execution runtime stats and process monitoring
The local execution engine now includes a comprehensive runtime statistics system. A new \RuntimeStatsManager\ collects per-operator metrics (rows in/out, bytes in/out, duration, and task counts) and exposes them via a subscriber framework. Additionally, a process-level monitor tracks CPU, memory (RSS), and jemalloc allocation metrics, which are exported as OpenTelemetry gauges. A terminal progress bar has been added to visualize execution, with support for persisting the view after job completion.
_src/daft-local-execution/src/runtime\stats · high confidence
Introduce local execution sink operators
The local execution engine now includes a comprehensive set of sink operators to handle data aggregation, partitioning, and output. This adds implementations for \AggregateSink\ and \GroupedAggregateSink\ to perform single and grouped aggregations, \DedupSink\ for removing duplicates, and \PivotSink\ for pivoting data. It also introduces \RepartitionSink\ and \IntoPartitionsSink\ to handle data redistribution using either Ray or Flight backends, \GatherSink\ to collect results, and \CommitWriteSink\ to finalize writes to storage (including handling empty files and success markers). These components are built on a new \BlockingSink\ trait that manages state accumulation and finalization.
src/daft-local-execution/src/sinks · high confidence
Introduce native MicroPartition operations
This change introduces a new set of native operations for the MicroPartition type, including aggregation, deduplication, concatenation, expression evaluation (with async and parallel variants), filtering (with statistics-based short-circuiting), multiple join strategies (hash, sort-merge, cross), partitioning (by hash, random, range, and value), pivot/unpivot, slicing, sorting, and sampling. These operations form the core execution layer for data transformations within the native runner, enabling efficient in-memory processing and optimization via statistics.
src/daft-micropartition/src/ops · high confidence
Introduce native Rust writers for Avro, CSV, JSON, and IPC formats
The \src/daft-writers\ crate now includes native Rust implementations for writing Avro, CSV, JSON, and Arrow IPC files, alongside existing Parquet support. These writers handle Hive-style partitioned paths, support custom date and timestamp formatting for CSV and JSON, allow ignoring null fields in JSON, and enable single-file writes. The implementation also introduces a batch-based writing strategy for Parquet to target specific row group sizes and a file-size rotation mechanism for Parquet to target specific file sizes.
src/daft-writers · high confidence
Introduce new local execution source implementations
The local execution engine now includes new source implementations for glob scans, in-memory data, scan tasks, and shuffle reads. These sources handle data ingestion from various formats (Parquet, CSV, JSON, etc.) and internal data flows, providing a foundation for the new execution model.
src/daft-local-execution/src/sources · high confidence
Introduce runtime-swappable global logger
Added a new \SwappableLogger\ component in \src/common/logging\ that wraps the standard \log\ interface to allow the underlying logger implementation to be swapped out atomically at runtime. This enables dynamic log configuration changes without requiring application restarts, supporting features like temporarily redirecting output or changing log levels on the fly.
src/common/logging · high confidence
Introduce structured AI provider and protocol layer
The \daft/ai\ package now exposes a structured API for integrating with external AI models, centered on a \Provider\ abstraction that manages connections to backends like OpenAI, Google, LM Studio, and vLLM. Users can now instantiate providers (e.g., \load\_openai()\) to access specific capabilities such as text/image embedding, text/image classification, and prompt-based chat completions via dedicated protocols. This change introduces a unified interface for model access, including support for configurable options like token limits, retry policies, and dimension specifications, while also adding metrics tracking for token usage and request counts.
daft/ai · high confidence
Introduce structured File types and media-specific operations
The \daft.file\ module now provides a structured file interface with a base \File\ class and specialized subclasses for \AudioFile\, \ImageFile\, \VideoFile\, and \Hdf5File\. Users can now access file metadata (e.g., audio sample rates, image dimensions, video frame rates, HDF5 group structures) and perform media-specific operations like decoding images to PIL objects, resampling audio, or iterating through video frames with metadata. The module also includes a \DaftFileIO\ wrapper and an \open\_file\ convenience function for standard Python file-like interactions with remote or local files.
daft/file · high confidence
Introduce typed FileReference with media type support and byte-range reading
The file module now uses a new FileReference struct that explicitly tracks media types (Unknown, Video, Audio, Image, HDF5) and supports optional byte-range reads via position and size fields. This allows users to reference files with specific content types and read subsets of data efficiently. The change also improves the display format of FileReference objects to include media type, path, configuration status, and range information when applicable.
src/daft-core/src/file · high confidence
Introduce unified DataSource/Task I/O architecture and new readers
The \daft/io\ module has been refactored to use a new \DataSource\ and \DataSourceTask\ interface, replacing the older \ScanOperator\ pattern. This change introduces native Python implementations for several readers, including Avro (\read\_avro\), Blob (\read\_blob\), Kafka (\read\_kafka\), MCAP (\read\_mcap\), and WebDataset (\read\_webdataset\), alongside updated implementations for CSV, JSON, Parquet, and SQL. The module now exports a comprehensive set of configuration classes (e.g., \AzureConfig\, \GCSConfig\, \GravitinoConfig\) and supports checkpointing for progress tracking across runs. Additionally, new APIs like \from\_files\ for lazy file references and \read\_generator\ for custom data sources have been added.
daft/io · high confidence
Introduces FileFormat and WriteMode enums for file handling configuration
The \src/common/file-formats\ module now exposes the \FileFormat\ and \WriteMode\ enums, centralizing the definition of supported file types and write behaviors. \FileFormat\ explicitly lists Parquet, CSV, JSON, WARC, Text, MCAP, and Avro, providing string parsing and Python bindings for these formats. \WriteMode\ defines Overwrite, OverwritePartitions, and Append strategies, also with string parsing and Python integration. This change establishes the core type definitions used to configure file I/O operations across the system.
src/common/file-formats · high confidence
Introduces Rust-based in-memory catalog and table abstractions
This change adds the core Rust implementation for the Daft catalog system, defining the \Catalog\, \Table\, and \Function\ traits along with supporting types like \Identifier\ and \Bindings\. It provides the foundational API for managing namespaces, tables, and functions, including support for case-sensitive and case-normalized lookups, pattern-based filtering for listing tables, and Python bindings for interoperability.
src/daft-catalog/src · high confidence
Introduces Rust-based partitioning abstractions and cache entry types
The \src/common/partitioning\ crate now provides the core Rust traits and types for dataset partitioning, including \Partition\, \PartitionSet\, and \PartitionSetCache\. This introduces a shared interface to avoid circular dependencies and defines \PartitionCacheEntry\ to support both Rust and Python-backed cache storage via PyO3 serialization. This change establishes the foundational data structures for partition management in the Rust codebase, replacing or supplementing the previous Python-based partitioning logic.
src/common/partitioning · high confidence
Introduces batch UDF support, retry logic, and structured metrics for Python UDFs
This change adds a new \BatchPyFn\ execution path alongside the existing row-wise UDFs, allowing Python functions to process data in batches rather than row-by-row. It introduces a retry mechanism with exponential backoff and jitter for both synchronous and asynchronous UDF calls, respecting \RetryAfterError\ headers from upstream services. Additionally, the implementation captures and exposes structured operator metrics from Python UDFs and supports configuration options such as concurrency limits, resource overrides (CPU/GPU), and Ray-specific options.
_src/daft-dsl/src/python\udf · high confidence
Introduces new local execution engine with dynamic batching and memory management
Adds a new native local execution model that replaces the previous runner, introducing a \BatchManager\ for dynamic, strategy-controlled batch extraction and a \MemoryManager\ to handle memory limits (including via the \DAFT\_MEMORY\_LIMIT\ environment variable). The engine now executes pipelines using a channel-based architecture with explicit flush signals, supporting operators like Concat, CheckpointTerminus, and various joins and aggregations within a unified pipeline node structure.
src/daft-local-execution/src · high confidence
Introduces per-operator dynamic batching with latency-constrained strategy
The local execution engine now supports dynamic batching that adapts batch sizes per operator based on execution metrics. A new \BatchingStrategy\ trait and \BatchManager\ abstraction allow pluggable strategies, including a \LatencyConstrainedBatchingStrategy\ that uses a binary search algorithm to maximize throughput while keeping batch latency within a specified SLA target, and a \StaticBatchingStrategy\ for fixed-size requirements. This change enables the runtime to automatically tune batch sizes for better performance without manual configuration.
_src/daft-local-execution/src/dynamic\batching · high confidence
Introduces the Session API for managing connection state and attached resources
Users can now use the new \Session\ class to manage connection state, including attaching and detaching catalogs, providers, tables, and functions (UDFs/UDAFs) by alias. The session exposes methods to query current context (current catalog, namespace, provider, model) and list available resources, providing a unified entry point for organizing and resolving data assets during query planning.
src/daft-session · high confidence
Introduction of TreeNode API for tree data structure traversal and rewriting
The \src/common/treenode\ module introduces a new \TreeNode\ trait and associated APIs that allow users to inspect and rewrite tree data structures (such as execution plans and expression trees) efficiently. This API provides distinct methods for top-down, bottom-up, and combined traversals (e.g., \apply\, \visit\, \transform\, \rewrite\), enabling algorithms to be expressed separately from the tree structure to avoid code duplication. The implementation is optimized to avoid cloning during transformations and supports both borrowed inspection and owned rewriting.
src/common/treenode · high confidence
Introduction of core Series operations in daft-core
The \src/daft-core/src/series/ops\ directory has been populated with a comprehensive set of new modules implementing fundamental data operations on Series. This includes arithmetic and logical operators (add, sub, mul, div, and, or, xor), comparison and membership checks (equal, lt, gt, between, is\_in), and aggregation functions (sum, product, count). The update also introduces utility and transformation operations such as casting (including a new \try\_cast\ for safe conversions), filtering, broadcasting, hashing (Murmur3 and MinHash), and mathematical functions (abs, floor, log, ln).
src/daft-core/src/series/ops · high confidence
Logical plan serialization and display capabilities
The logical plan now supports JSON serialization and multiple display formats. Users can export the plan structure as JSON via the new \repr\_json\ method, and visualize the plan as a Mermaid flowchart or ASCII tree for debugging and inspection. The plan also includes accumulated selectivity statistics to aid in query optimization and performance analysis.
src/daft-logical-plan/src · high confidence
Native Azure and GCS backends with expanded authentication and Gravitino support
The I/O module now includes native Rust implementations for Azure Blob Storage and Google Cloud Storage, replacing previous dependencies. The Azure backend (\azure\_blob.rs\, \azure\_auth.rs\) supports multiple authentication methods including SharedKey (account keys), SAS tokens, and static bearer tokens (for Microsoft Fabric/OneLake), while the GCS backend (\google\_cloud.rs\) provides native object listing, reading, and deletion. Additionally, a new Gravitino virtual file system source (\gravitino.rs\) enables reading and writing via the \gvfs://\ protocol, and a \CountingReader\ utility has been added to track bytes read for I/O statistics.
src/daft-io/src · high confidence
Native MCAP reader with indexed chunk traversal
The \src/daft-mcap\ crate now provides a native MCAP reader that decodes MCAP files into Daft record batches. It supports indexed traversal for log-time ordered message retrieval using the file's summary chunk indexes, with a linear fallback for files lacking a usable index. The reader applies topic and time-range filters during decoding, supports residual predicates and projections, and handles remote file access via range requests with appropriate buffering and validation.
src/daft-mcap · high confidence
Native RecordBatch operations for aggregation, grouping, and data transformation
This change introduces a new set of native Rust operations for the RecordBatch type, replacing or supplementing previous implementations. Key additions include a new aggregation engine (\agg.rs\) with fast-path inline aggregation for common functions (count, sum, min, max) and support for Python UDFs via \map\_groups\. It also implements core grouping logic (\groups.rs\, \hash.rs\) using hash-based and sort-based groupers, enabling efficient groupby operations. New data transformation capabilities include \explode\ with optional index column generation, \pivot\ and \unpivot\ for reshaping data, \partition\ by hash, random, range, or value, and \search\_sorted\ for multi-column sorted lookups. Sorting and top-N operations (\sort.rs\) are also implemented natively. A benchmark file (\bench\_agg.rs\) is added to measure aggregation performance.
src/daft-recordbatch/src/ops · high confidence
Native extension support for scalar and aggregate functions
The \daft-ext-internal\ crate now provides the internal implementation for loading native extension modules (shared libraries) via a stable C ABI. This change introduces a global module registry that handles dynamic library loading (\dlopen\) and manages the lifecycle of extension functions. Specifically, it adds \ScalarFunctionHandle\ and \AggregateFunctionHandle\ wrappers that bridge the FFI boundary, allowing users to register and execute custom scalar and aggregate functions defined in external native code directly within Daft expressions.
src/daft-ext-internal · high confidence
Native streaming JSON reader with parallel processing and array support
The JSON reader in \src/daft-json\ has been rewritten to support native streaming and parallel reading of JSONL/NDJSON files, replacing the previous synchronous approach. This change introduces a new \simd\_json\-based deserializer and a streaming architecture that splits files into byte-range chunks for concurrent processing, significantly improving performance on large datasets. The reader now supports reading JSON arrays (e.g., \\[{...}, {...}\]\) in addition to newline-delimited JSON, and includes robust schema inference that handles mixed types, nulls, and nested structures. Users benefit from faster read times, lower memory usage due to streaming, and the ability to process large JSON files that previously might have failed or been slow.
src/daft-json · high confidence
Native streaming and parallel CSV reader with advanced parsing options
The \daft-csv\ crate now provides a native Rust implementation for reading CSV files, replacing previous approaches. This new reader supports streaming and parallel processing for improved performance, including handling of compressed files (deflate, zlib) and local file paths. Users gain access to granular control via \CsvParseOptions\ (custom delimiters, quote/escape characters, comment lines, variable columns) and \CsvConvertOptions\ (row limits, column projections, schema overrides, and predicate pushdown). The implementation also introduces robust error handling for corrupt files and detailed I/O statistics.
src/daft-csv · high confidence
New Axum-based dashboard server with event log import and interactive data display
The dashboard has been rewritten to use the Axum web framework, introducing a new server architecture that serves static assets and provides REST and Server-Sent Events (SSE) APIs for real-time query and task monitoring. Users can now import historical execution data via the new event log import feature, which replays past query states into the dashboard. Additionally, the dashboard now supports interactive Jupyter-style display of DataFrames, allowing users to view cell details in a side pane directly within the browser interface.
src/daft-dashboard/src · high confidence
New Flotilla distributed execution engine
The distributed execution engine has been replaced with a new Flotilla runner, introducing a \PipelineNode\-based architecture that translates logical plans into distributed tasks. This change adds new node implementations for core operations—including aggregation, distinct, explode, filter, glob scanning, in-memory sources, and partitioning (into\_batches, into\_partitions)—and introduces a new Actor UDF model that manages Python UDF actors via Ray. The new engine also includes execution-time clustering logic to optimize data movement and provides detailed per-operator observability metrics.
_src/daft-distributed/src/pipeline\node · high confidence
New FunctionArgs derive macro for declarative function argument parsing
A new \FunctionArgs\ procedural macro has been added to the \src/common/macros\ crate, allowing developers to derive argument parsing for function implementations. By applying \\#\[derive(FunctionArgs)\]\ to a struct, users can declaratively define required, optional, and variadic arguments, as well as override argument names via \\#\[arg(...)\]\ attributes. The macro automatically generates \TryFrom\<FunctionArgs\<T\>\>\ implementations that handle both expression (\ExprRef\) and series (\Series\) contexts, including automatic literal conversion for concrete types, simplifying the creation of new data functions.
src/common/macros · high confidence
New Hugging Face IO implementation with Xet support and improved error handling
The Hugging Face IO module has been rewritten to support reading from Xet-backed repositories, providing a new high-performance data access path for compatible datasets. The implementation introduces a dedicated error type system that distinguishes between private datasets, unauthorized bucket access, and generic connection failures, offering clearer diagnostics for users. Path parsing has been refined to correctly handle Hugging Face URL structures, including normalization of bucket tree paths and support for specific file revisions, ensuring that glob patterns and direct file references resolve accurately.
src/daft-io/src/huggingface · high confidence
New JSON functions: jq, json\_array\_length, json\_object\_keys, and json\_tuple
This change introduces four new JSON processing functions to the \daft-functions-json\ module. Users can now apply \jq\ filters to JSON strings, retrieve the length of the outermost JSON array via \json\_array\_length\, and extract top-level keys as a sorted list using \json\_object\_keys\. Additionally, \json\_tuple\ extracts values for specified top-level keys from a JSON object, returning them as a Struct (with field names matching the keys) rather than multiple columns, aligning with Spark's behavior but adapted for Daft's single-output UDF model. All functions are registered in the function registry and handle NULL inputs, malformed JSON, and type errors gracefully.
src/daft-functions-json · high confidence
New Lance I/O module with namespace support and lazy loading
The \daft/io/lance\ package has been restructured to provide a dedicated entry point for LanceDB integration, featuring a new \read\_lance\ function that supports reading from Lance Namespaces (via \table\_id\, \namespace\_impl\, and \namespace\_properties\) in addition to standard URIs. This module also introduces lazy loading of the underlying \daft\_lance\ library to speed up \import daft\, and exposes \DatasetOpenContext\ for namespace operations.
daft/io/lance · high confidence
New MinHash library with SIMD performance and alternative hashers
The \src/daft-minhash\ crate introduces a new MinHash implementation for estimating string similarity via Jaccard coefficients. This library replaces the previous \arrow2\ dependency with \arrow\ and leverages Rust's \portable\_simd\ feature to accelerate hash computations. It supports alternative hashers (such as Murmur3 and xxHash) and includes a \windowed\ module for generating word n-grams, along with comprehensive benchmarks and tests to validate the SIMD-based remainder and permutation logic.
src/daft-minhash · high confidence
New OpenAI protocol implementations for prompter and text embedder
This change introduces the \OpenAIPrompter\ and \OpenAITextEmbedder\ implementations within the \daft/ai/openai/protocols\ module. The prompter now supports both the Chat Completions API and the Responses API, allowing users to configure \use\_chat\_completions\ and handle multi-modal inputs (images, videos, files) via HTTP URLs or bytes. The text embedder adds support for configurable embedding dimensions for models like \text-embedding-3-small\ and \text-embedding-3-large\, respects model-specific token limits, and forwards \extra\_body\ options to embedding requests. These protocols serve as the concrete OpenAI-specific implementations for Daft's AI provider interface.
daft/ai/openai/protocols · high confidence
New Python API for dynamic scalar function resolution
A new Python module (\src/daft-functions/src/python\) exposes a \get\_function\_from\_registry\ function and a \PyScalarFunction\ wrapper, allowing Python code to dynamically resolve and invoke scalar functions by name. This change introduces the infrastructure for dynamic function resolution in the Python layer, enabling features like function overloads and custom UDFs to be accessed and executed through a unified registry interface.
src/daft-functions/src/python · high confidence
New Python module initialization with jemalloc and file type functions
The \src/lib.rs\ entry point now initializes the Python \daft\ module by registering a comprehensive set of submodules (including \daft\_file\, \daft\_session\, \daft\_sql\, and others) and populating the function registry with new file-related expressions such as \File\, \FileExists\, \FileSize\, \VideoFile\, \AudioFile\, \ImageFile\, \Hdf5File\, and \GuessMimeType\. Additionally, the library enables the jemalloc allocator with background threads on non-MSVC platforms to improve memory performance, and exposes Python-accessible functions for version, build type, logging control, and compute thread configuration.
src · high confidence
New Python type stubs for core, dashboard, and testing modules
The package now ships with comprehensive Python type stubs (\.pyi\ files) for the core \daft\ module, the \dashboard\ subsystem, and the \testing\ utilities. These stubs expose the public API surface—including data types like \ImageMode\, \WindowSpec\, and \JoinType\, as well as dashboard launch and registration functions—to improve static type checking and IDE autocomplete for users.
daft/daft · high confidence
New Queries page with detailed status tracking and Ray UI links
The dashboard now includes a dedicated Queries page that lists all queries with sortable columns for Name, Status, Duration, Entrypoint, Start Time, and Engine. Users can view real-time status updates (including Pending, Optimizing, Executing, Finalizing, Finished, Canceled, Failed, and Dead states) with human-readable durations. For queries running on Ray, a direct link to the Ray UI is provided. The page also handles empty states with a copyable command snippet for starting new queries.
src/daft-dashboard/frontend/src/app/queries · high confidence
New Ray-based distributed execution engine (Flotilla/Swordfish)
Users can now run Daft workloads on Ray using a new distributed execution engine called Flotilla (also referred to as Swordfish). This change introduces a new set of Rust modules in \src/daft-distributed/src/python/ray\ that implement the core scheduling and worker management logic, including \RaySwordfishWorker\ for handling task execution and state, \RayWorkerManager\ for managing the Ray cluster lifecycle and autoscaling, and \RayTaskResultHandle\ for retrieving task outcomes. The implementation supports both Ray and Flight partition references, allows configuration of autoscaling strategies (gradual or bisect), and integrates with Python via PyO3 to interact with the Ray runtime.
src/daft-distributed/src/python/ray · high confidence
New Rust SDK for building Daft native extensions
This change introduces the \daft-ext\ crate and its companion \daft-ext-macros\ crate, providing a stable C ABI and Rust tooling for developers to create custom Daft extensions. The SDK exposes \\#\[daft\_extension\]\ and \\#\[daft\_func\]\ proc macros that allow users to define scalar and aggregate functions in Rust, which are then compiled into shared libraries (cdylibs) that Daft can load at runtime. The implementation includes a zero-dependency Arrow C Data Interface layer for safe FFI data exchange, feature-gated compatibility with arrow-rs versions 56 through 59, and a session context API for registering these functions within a Daft execution session.
src/daft-ext · high confidence
New Spark-compatible string functions and case-conversion utilities
This release adds a suite of new string functions to the \daft-functions-utf8\ module, including \ascii\, \chr\, \capitalize\, \contains\, \count\_matches\, \find\, \ilike\, \like\, \startswith\, and \endswith\. It also introduces case-conversion functions (\to\_camel\_case\, \to\_snake\_case\, \to\_kebab\_case\, etc.) and string distance/similarity metrics (\levenshtein\_distance\, \damerau\_levenshtein\_distance\, \hamming\_distance\, \jaro\_similarity\, \jaro\_winkler\_similarity\). These functions are implemented as Scalar UDFs and registered in the \Utf8Functions\ module, with \ascii\ and \chr\ specifically matching Spark's semantics (e.g., signed byte values for non-ASCII characters in \ascii\).
src/daft-functions-utf8 · high confidence
New UI component library added to the dashboard frontend
The dashboard frontend now includes a comprehensive set of reusable UI components in the \src/daft-dashboard/frontend/src/components/ui\ directory. This addition provides foundational building blocks for the interface, including layout and navigation elements (Sidebar, Breadcrumb, Sheet, Collapsible), form controls (Input, Label, Button, Tabs), data display widgets (Table, Card, Badge, Empty state, Note), and interactive primitives (DropdownMenu, Tooltip, Resizable panels). These components are styled with Tailwind CSS and leverage Radix UI primitives to ensure accessibility and consistent behavior across the application.
src/daft-dashboard/frontend/src/components/ui · high confidence
New URI functions for downloading, uploading, and parsing URLs
This change introduces a new \daft-functions-uri\ crate that provides three new functions for working with URIs. The \url\_download\ function fetches content from URLs and returns binary data, supporting configuration for connection limits and error handling (raising or returning null on failure). The \url\_upload\ function uploads binary or string data to specified URL locations, also with configurable connection limits and error handling, and supports uploading to single or multiple folder paths. The \url\_parse\ function parses URL strings into structured components including scheme, username, password, host, port, path, query, and fragment.
src/daft-functions-uri · high confidence
New binary encoding and decoding functions
Users can now encode and decode binary data using Base64, Hex, Gzip, Deflate, Zlib, and UTF-8 codecs via the new \encode\, \decode\, \try\_encode\, and \try\_decode\ functions. These functions support both variable-length and fixed-size binary inputs, allowing seamless conversion between binary and text representations with configurable compression and encoding schemes.
src/daft-functions-binary · high confidence
New binary, comparison, hashing, and search kernels in daft-core
The \src/daft-core/src/kernels\ module now provides dedicated implementations for binary array concatenation (\binary.rs\), NaN-aware float comparisons and the \is\_nearer\ utility (\cmp.rs\), a full suite of hash functions (Murmur3, SHA1, xxHash32/64/3\_64) with support for Float16 and other types (\hashing.rs\), optimized binary-search lookups for sorted arrays (\search\_sorted.rs\), and string concatenation (\utf8.rs\). These kernels enable users to perform binary concatenation, precise float comparisons that respect NaN ordering, flexible hashing across more data types, and efficient sorted lookups.
src/daft-core/src/kernels · high confidence
New catalog implementations for AWS Glue, Iceberg, Paimon, PostgreSQL, S3 Tables, and Unity Catalog
Users can now connect Daft to a variety of external catalog services using new \Catalog\ implementations. This change introduces \GlueCatalog\ (with support for creating instances from a \boto3\ or \botocore\ session), \IcebergCatalog\ (wrapping PyIceberg catalogs), \PaimonCatalog\ (for Apache Paimon Lake Format), \PostgresCatalog\ (for PostgreSQL with automatic extension setup like \pgvector\), \S3Catalog\ (for AWS S3 Tables via Iceberg REST), and \UnityCatalog\ (for Databricks Unity Catalog). These implementations allow users to discover, access, and query data stored in these systems using Daft's unified catalog interface.
daft/catalog · high confidence
New checkpoint configuration types and Python bindings
Introduces the \checkpoint-config\ crate, which defines serializable configuration structures for checkpoint stores (\CheckpointStoreConfig\, \CheckpointConfig\) and tunable filtering settings (\KeyFilteringSettings\). This decouples the configuration schema from the live store implementations, allowing logical plan nodes to reference checkpoint settings without pulling in heavy IO dependencies. It also exposes these types to Python via PyO3 wrappers, enabling users to configure checkpoint backends (currently ObjectStore) and tune key-filtering parameters (such as \num\_workers\ and \filter\_batch\_size\) directly from Python code.
src/common/checkpoint-config · high confidence
New common display module for tree, table, and diagram rendering
A new \src/common/display\ crate has been introduced to centralize visualization logic, providing traits and implementations for rendering execution plans and data structures. It adds ASCII tree formatting (including a \git log\-style graph and indented styles) via \ascii.rs\, Mermaid diagram generation with configurable options (simple mode, bottom-up layout, subgraphs) via \mermaid.rs\, and enhanced table rendering with dynamic column truncation and configurable styling via \table\_display.rs\. The module also defines a \TreeDisplay\ trait for hierarchical data and utility functions like byte-to-human-readable conversion, establishing a shared foundation for explain output and DataFrame previews.
src/common/display · high confidence
New daft.functions module centralizes expression APIs
The \daft/functions\ package has been introduced to provide a unified, organized entry point for Daft's expression functions. This change consolidates previously scattered implementations into a structured module, exposing APIs for aggregations (\daft.functions.agg\), date/time operations (\daft.functions.datetime\), file handling (\daft.functions.file\_\), image processing (\daft.functions.image\), and more. Users can now import these functions directly from \daft.functions\ (e.g., \from daft.functions import date, file, resize\), replacing or supplementing previous access patterns and providing a consistent, discoverable interface for scalar, aggregate, and file-specific expressions.
daft/functions · high confidence
New dashboard components for query tracking and notifications
The dashboard frontend now includes new components to support query management and user notifications. A new \ServerProvider\ component establishes a Server-Sent Events (SSE) connection to the \/client/queries/subscribe\ endpoint, allowing the UI to receive real-time updates on query status (Pending, Finished) and display browser notifications when enabled. A \NotificationsProvider\ manages user preferences for these alerts and handles the logic for requesting browser notification permissions. Additionally, new \icons.tsx\ and \loading.tsx\ components provide visual assets and a placeholder loading state for the dashboard interface.
src/daft-dashboard/frontend/src/components · high confidence
New dashboard navbar with connection status and notification controls
The dashboard now features a new navigation bar that includes a logo with a terminal-style animation, a link to the 'All Queries' page displaying a live count of active queries, and a link to the documentation. A new connection status indicator is visible on the right side, showing whether the client is connected or disconnected by periodically pinging the server. Additionally, users can manage their notification preferences via a dropdown menu that allows toggling alerts for query start and end events.
src/daft-dashboard/frontend/src/components/navbar · high confidence
New dataset integrations for Common Crawl, DROID, and LeRobot
The \daft.datasets\ module now exposes dedicated loaders for three new data sources. Users can load Common Crawl web archives (WARC, WET, WAT formats) via \daft.datasets.common\_crawl\, with support for fetching data from AWS S3, HuggingFace, or HTTP. Robotics data is supported via \daft.datasets.droid\, which loads raw DROID episodes with metadata and lazy references to HDF5 trajectories and MP4 camera feeds. Additionally, \daft.datasets.lerobot\ provides access to LeRobot datasets, supporting both v2.0/v2.1 and v3 layouts, including optimized batched video frame decoding for efficient access to episode video shards.
daft/datasets · high confidence
New developer tooling for debugging, observability, and CI workflows
The \tools\ directory now includes several new utilities to improve the development experience. Developers can use \attach\_debugger.py\ to automatically attach the LLDB debugger to Python processes for Rust debugging, and \aggregate\_test\_durations.py\ (along with \capture-durations.sh\) to parse and summarize pytest execution times. For observability, a new OpenTelemetry setup in \tools/observability/opentelemetry/\ provides Docker Compose configurations to run a local collector, Jaeger, and Prometheus for tracing and metrics. Additionally, \gha\_run\_cluster\_job.py\ allows triggering cluster workflows via GitHub Actions, \convert\_md\_to\_notebook.py\ converts documentation to Jupyter notebooks, and \trace\_to\_speedscope.py\ converts JSON traces into SpeedScope profiles for performance visualization.
tools · high confidence
New distributed runtime utilities for channels, streaming, and output transposition
The \src/daft-distributed/src/utils\ module now provides core infrastructure for the distributed execution engine. It introduces a channel abstraction wrapping Tokio's synchronous and unbounded channels, a runtime initializer that configures a dedicated single-threaded Tokio runtime for the scheduler, and a \JoinableForwardingStream\ that allows background tasks to surface errors while an input stream is pending. Additionally, it includes utilities for transposing materialized outputs from streams and vectors to reorganize data by partition.
src/daft-distributed/src/utils · high confidence
New distributed statistics and task lifecycle event system
This change introduces a new statistics management subsystem in \src/daft-distributed/src/statistics\ to track distributed execution metrics and task lifecycle events. It adds a \StatisticsManager\ and \RuntimeNodeManager\ to aggregate per-node operator statistics (such as active, completed, failed, and cancelled task counts) and emit \OperatorStart\/\OperatorEnd\ events. Additionally, it implements a \TaskLifecycleEventSubscriber\ that translates internal task states (Submitted, Scheduled, Completed, Failed, Cancelled) into external \TaskSubmit\, \TaskScheduled\, and \TaskEnd\ events, which are dispatched via the context's event bus. This functionality is gated by the \DAFT\_TASK\_EVENTS\_ENABLED\ environment variable, allowing users to enable detailed task-level observability for distributed queries.
src/daft-distributed/src/statistics · high confidence
New example demonstrating native extension functions
The \examples/hello\ directory now includes a Rust library (\src/lib.rs\) that serves as a concrete example of how to register native extension functions with Daft. It demonstrates defining a scalar function (\greet\) and an aggregate function (\string\_count\) using the \\#\[daft\_extension\]\ and \\#\[daft\_func\]\ macros, showing users how to implement custom logic via the \DaftExtension\ trait.
examples/hello · high confidence
New example demonstrating native vector distance functions
The \examples/dvector\ directory now includes a working example that registers native Daft extension functions for computing vector distances. Users can see how to implement and register six specific distance metrics—L2, inner product, cosine, L1, Hamming, and Jaccard—using the \\#\[daft\_extension\]\ and \\#\[daft\_func\_batch\]\ macros, along with the supporting logic in \vectors.rs\ for handling fixed-size and variable-length float and boolean vector arrays.
examples/dvector · high confidence
New execution module with native executor, UDF workers, and distributed operators
The \daft/execution\ package is introduced, providing the core runtime components for query execution. This includes a \NativeExecutor\ that streams results from the native engine into Python, a \UDFActor\ implementation for Ray-based distributed UDFs with resource-aware scheduling, and a \LimitCounterActor\ for distributed limit operations. It also adds a \SharedMemoryTransport\ and \UdfHandle\ for efficient, deadlock-free communication between parent processes and subprocess-based UDF workers, and an experimental \VLLMExecutor\ interface for integrating large language model inference into the execution pipeline.
daft/execution · high confidence
New expression functions for data processing and similarity analysis
This release adds a suite of new expression functions to the \daft-functions\ crate, expanding capabilities for data manipulation and analysis. Users can now use \coalesce\ to return the first non-null value from a sequence of columns, \concat\_ws\ to join strings with a separator, and \length\ to determine the size of strings, binary data, or lists. For data integrity and deduplication, the \hash\ function now supports multiple inputs and configurable algorithms, while \minhash\, \simhash\, and \hamming\_distance\ provide tools for near-duplicate detection. Vector analysis is enhanced with \cosine\_similarity\, \jaccard\_similarity\, and \pearson\_correlation\ for comparing vector columns. Additional utilities include \slice\ for extracting sub-sequences from lists or binary data, \to\_struct\ for merging columns into a struct type, and \random\_int\ for generating random integer columns.
src/daft-functions/src · high confidence
New expression visitor pattern and PyArrow integration
The \daft/expressions\ module now exposes an \ExpressionVisitor\ base class and a \PredicateVisitor\ helper, enabling users and internal systems to traverse and transform expression trees (e.g., for column extraction or optimization). Additionally, a new \\_PyArrowExpressionVisitor\ allows Daft expressions to be translated into PyArrow compute expressions, facilitating interoperability with other PyArrow-based tools.
daft/expressions · high confidence
New float utility functions: is\_nan, is\_inf, not\_nan, and fill\_nan
Users can now detect and handle special floating-point values directly in expressions. The new \is\_nan\ and \is\_inf\ functions return boolean indicators for NaN and infinite values respectively, while \not\_nan\ filters out NaN entries. Additionally, \fill\_nan\ allows replacing NaN values with a specified scalar or vector fill value. These functions are implemented as scalar UDFs in the \daft-functions\ crate and are registered in the function registry, supporting Float16, Float32, and Float64 data types.
src/daft-functions/src/float · high confidence
New hooks for mobile detection and query status management
Added \use-mobile.tsx\ to provide a \useIsMobile\ hook for responsive UI logic based on a 768px breakpoint. Introduced \use-queries.ts\ which defines types for various query states (Pending, Optimizing, Setup, Executing, Finalizing, Finished, Failed, Canceled, Dead) and provides \useQueries\ and \useActiveQueries\ hooks to fetch and filter query summaries, enabling the dashboard to display detailed query lifecycle information and distinguish active queries from completed or failed ones.
src/daft-dashboard/frontend/src/hooks · high confidence
New image processing functions and expression API
This change introduces a new set of image processing functions available as expressions in the \daft-image\ module. Users can now decode images from binary data or files (\image\_decode\, \decode\_image\_file\), encode images to binary formats (\image\_encode\), and extract metadata such as dimensions and format (\image\_file\_metadata\, \image\_attribute\). The update also adds image manipulation capabilities including cropping (\image\_crop\), resizing (\image\_resize\), and mode conversion (\to\_mode\). Additionally, it supports converting images to tensors (\to\_tensor\) for machine learning workflows and computing perceptual hashes (\image\_hash\) for near-duplicate detection using various algorithms like pHash, dHash, and crop-resistant hashing.
src/daft-image/src/functions · high confidence
New interactive query plan visualization and task tracking
The dashboard now features a dedicated query detail page with an interactive physical plan tree that visualizes operator execution, including live wall-clock and CPU duration, row/byte throughput, and heatmap-based bottleneck detection (CPU vs. queue depth). Users can view edge labels showing data amplification or reduction, toggle between tree and JSON views, and drill down into Flotilla task groups via a new tasks sidebar that supports node filtering and hover-based plan highlighting. Operator status is now more accurate, accounting for Flotilla task retries and cancellations, and query errors are surfaced in a dedicated modal.
src/daft-dashboard/frontend/src/app/query · high confidence
New join reordering optimizer rule with DP-ccp and brute-force algorithms
A new \ReorderJoins\ optimizer rule has been added to automatically determine the most efficient order for joining tables in a query. The rule constructs a join graph from the logical plan and uses two algorithms to find the optimal order: a Dynamic Programming via Connected subgraph Complement Pairs (DP-ccp) algorithm for queries with up to 12 relations, and a brute-force algorithm for up to 7 relations. Both algorithms consider table cardinalities and join conditions to minimize intermediate result sizes, potentially improving query performance by avoiding expensive join orders.
_src/daft-logical-plan/src/optimization/rules/reorder\joins · high confidence
New key-filtering anti-join execution path
This change introduces a new execution strategy for anti-joins that rely on key existence checks. It adds a Python-based infrastructure in \daft/execution/key\_filtering\_join\ that manages Ray actors to store and filter keys, along with a Rust-Python bridge (\bridge.py\) to integrate this logic into the \KeyFilteringJoinNode\. This provides a distributed, actor-based mechanism for filtering rows during anti-joins, distinct from previous implementations.
_daft/execution/key\_filtering\join · high confidence
New list manipulation functions: append, contains, chunk, and explode options
This change introduces several new list expression functions to the product. Users can now append values to lists using \list\_append\, check for element presence with \list\_contains\, and split lists into fixed-size chunks via \list\_chunk\. Additionally, the \explode\ function now supports an \ignore\_empty\_and\_null\ parameter, allowing users to control whether empty lists and nulls are preserved or dropped during expansion.
src/daft-functions-list · high confidence
New local execution engine for joins
The local execution engine now includes a complete set of join operators, introducing support for ASOF joins, anti/semi joins, cross joins, and standard inner/left/right/outer joins. This adds the capability to execute these join types locally within the pipeline, moving beyond the previous implementation which likely relied on other execution paths or lacked these specific join variants in the local runtime.
src/daft-local-execution/src/join · high confidence
New local execution engine intermediate operators
The local execution engine now includes a new set of intermediate operators for the native runner, including distributed actor pool projects, explode, filter, project, unpivot, and UDF execution. These operators support dynamic batching, runtime statistics tracking, and checkpoint key staging, enabling more efficient and observable data processing within the local execution model.
_src/daft-local-execution/src/intermediate\ops · high confidence
New logical plan optimization rules for joins, offsets, and projections
The logical plan optimizer now includes several new rules that improve query performance and correctness. The \EliminateCrossJoin\ rule rewrites inner joins with equality predicates into standard joins, while \FilterNullJoinKey\ inserts pre-join filters to remove null keys, reducing data volume. \EliminateOffsets\ simplifies offset chains and merges them with limits. \ExtractWindowFunction\ converts window expressions in projections into explicit Window operations for better execution. \SplitGranularProjection\ isolates async expressions like URL downloads into separate projection steps to allow granular batch sizing. Additionally, \DetectMonotonicId\ optimizes \monotonically\_increasing\_id\ calls, and \DropIntoBatches\/\DropRepartition\ remove redundant partitioning operators.
src/daft-logical-plan/src/optimization/rules · high confidence
New metrics library with OpenTelemetry integration and operator-level tracking
The \src/common/metrics\ module has been introduced to centralize and standardize observability. It provides a new metrics abstraction backed by OpenTelemetry, exposing counters, gauges, and up-down counters with automatic name normalization. The library defines a comprehensive set of metric keys for execution statistics (rows, bytes, duration), checkpointing, joins, and process-level resource usage (CPU, memory). It introduces operator-specific metrics via \OperatorMetrics\ and \MetricsCollector\, allowing per-operator tracking with descriptions and attributes. A new \StatSnapshot\ system provides typed, serializable snapshots for different node types (sources, filters, joins, etc.), and \TaskExternalIo\ aggregates these into task-level I/O totals. Python bindings are included to expose \OperatorMetrics\ and \StatType\ to the Python runtime.
src/common/metrics · high confidence
New native extension examples: dvector, hello, and hello\_cpp
Added three new example projects in the examples directory to demonstrate Daft native extension development. The dvector example provides vector distance functions (L2, inner product, cosine, L1, Hamming, Jaccard) implemented in Rust. The hello example offers a minimal Rust-based extension with scalar and aggregate functions. The hello\_cpp example demonstrates a C++ extension using the Apache Arrow C++ library, providing a scalar greet function. Each example includes Python bindings, build configuration, and comprehensive tests.
examples · high confidence
New numeric functions added to the expression engine
The \src/daft-functions/src/numeric\ module now exposes a comprehensive set of new numeric expression functions, including \abs\, \bin\, \cbrt\, \ceil\, \clip\, \conv\, \e\, \exp\, \expm1\, \factorial\, \floor\, \hypot\, \log\ (with \log2\, \log10\, \ln\, \log1p\), \pi\, \pmod\, \pow\, \power\, \round\, \sign\, \negate\, \sqrt\, and full trigonometric support (\sin\, \cos\, \tan\, \csc\, \sec\, \cot\, \sinh\, \cosh\, \tanh\, and their inverses). These functions are implemented as \ScalarUDF\ instances and registered via the \NumericFunctions\ module, enabling users to perform advanced mathematical, base-conversion, and trigonometric operations directly within their data pipelines.
src/daft-functions/src/numeric · high confidence
New partitioning expression functions for Iceberg and time-based partitioning
This change introduces new DSL functions for partitioning expressions, specifically \years\, \months\, \days\, and \hours\ for temporal data, as well as \iceberg\_bucket\ and \iceberg\_truncate\ for Iceberg-specific partitioning strategies. These functions are implemented via new evaluators that handle type checking and execution, allowing users to define partition schemes using these expressions in their data pipelines.
src/daft-dsl/src/functions/partitioning · high confidence
New pattern matching utility for SQL LIKE expressions
Added a new \pattern\ library that converts SQL \LIKE\ patterns (supporting \%\, \\_\, and escape characters) into anchored regular expressions. This utility provides the core logic for pattern filtering, enabling features such as filtering results in \SHOW TABLES\ commands.
src/common/pattern · high confidence
New read\_video\_frames function for streaming video frames into DataFrames
Users can now read video files directly into a Daft DataFrame using the new \daft.read\_video\_frames\ function. This feature streams video frames as image data, allowing for resizing to specific dimensions and filtering by key frames. It also supports time-interval sampling to select frames at regular intervals and handles various video formats, including YouTube URLs. The function requires the PyAV library and returns a DataFrame containing frame metadata (such as timestamps and indices) alongside the image data.
daft/io/av · high confidence
New runtime abstraction with JoinSet and Python coroutine execution
The \src/common/runtime\ module now provides a structured runtime environment featuring \JoinSet\ and \OrderedJoinSet\ wrappers around Tokio tasks, enabling controlled spawning and joining of asynchronous work with support for specific compute or I/O runtime pools. It introduces a \Runtime\ type that manages worker thread counts and task execution, including panic handling for blocking tasks. Additionally, it adds Python interoperability via \execute\_python\_coroutine\, which allows Rust code to safely execute Python async coroutines within the proper asyncio event loop context.
src/common/runtime · high confidence
New serialize and deserialize functions for JSON support
This change introduces new \serialize\, \try\_serialize\, \deserialize\, and \try\deserialize\ functions to the Daft expression API, enabling users to convert data to and from string representations using specified formats (including JSON). The \serialize\ functions convert arbitrary expressions into string columns, while \deserialize\ functions parse string columns back into specified data types. The \try\\ variants handle parsing or serialization failures gracefully by inserting null values instead of raising errors. These functions are registered in the \SerdeFunctions\ module and are accessible via the DSL.
src/daft-functions-serde/src · high confidence
New streaming sink operators for local execution
The local execution engine now supports streaming sinks for several new operations, allowing them to process data in a streaming fashion rather than requiring full materialization. This change introduces implementations for \Limit\ (with offset support), \DistributedLimit\ (coordinating limits across distributed actors), \MonotonicallyIncreasingId\ (generating unique sequential IDs), \Sample\ (with/without replacement and configurable seeds), and \VLLM\ (integrating with local or remote LLM executors). Additionally, a new \AsyncUdfSink\ enables asynchronous execution of Python UDFs with detailed runtime metrics (rows, bytes, duration, custom counters) and configurable concurrency via the \DAFT\_MAX\_ASYNC\_UDF\_INFLIGHT\_TASKS\ environment variable. These sinks share a common \StreamingSink\ trait and base execution context, enabling consistent batching, concurrency control, and observability across all sink types.
_src/daft-local-execution/src/streaming\sink · high confidence
New subscriber framework for runtime event streaming
A new subscriber system has been introduced to allow external consumers to receive real-time execution events. This includes a \DashboardSubscriber\ that automatically connects to a dashboard service (via the \DAFT\_DASHBOARD\_URL\ environment variable) to visualize query progress, and a \PySubscriberWrapper\ that exposes these events to Python scripts. The framework defines a comprehensive set of typed events covering the full query lifecycle, including query start/end, optimization phases, operator start/end, and task scheduling/execution, enabling detailed observability and debugging.
src/daft-context/src/subscribers · high confidence
New temporal functions for date arithmetic, construction, and timezone handling
The \src/daft-functions-temporal\ module now provides a comprehensive set of new scalar UDFs for working with temporal data. Users can construct dates and timestamps from components using \make\_date\, \make\_timestamp\, and \make\_timestamp\_ltz\, and retrieve the current date, timestamp, or timezone via \current\_date\, \current\_timestamp\, and \current\_timezone\. Date arithmetic is supported through \date\_add\, \date\_sub\, \date\_diff\, \add\_months\, and \months\_between\, while navigation helpers like \last\_day\ and \next\_day\ allow finding specific calendar boundaries. Timezone management includes \convert\_time\_zone\, \replace\_time\_zone\, \from\_utc\_timestamp\, and \to\_utc\_timestamp\. Additional utilities include \truncate\ for interval-based rounding, \strftime\ for formatting, \from\_unixtime\ for conversion from epoch integers, and various extraction functions like \year\, \month\, \day\, \hour\, \minute\, \second\, \quarter\, \day\_of\_week\, \day\_of\_year\, \week\_of\_year\, \unix\_date\, \unix\_timestamp\, and duration totals (\total\_seconds\, \total\_minutes\, etc.).
src/daft-functions-temporal · high confidence
New tutorials for Common Crawl, Delta Lake, and document processing
Added new tutorial notebooks demonstrating how to work with Common Crawl data, perform local and distributed batch inference and training on Delta Lake tables, and build a PDF document processing pipeline. Also included a Flyte integration example and an embeddings tutorial for StackExchange data.
tutorials · high confidence
New utility crate for hashable float wrappers
A new utility crate, \hashable-float-wrapper\, has been introduced to provide a \FloatWrapper\ type that enables hashing and ordering for floating-point values. This wrapper supports \f16\ (half-precision), \f32\, and \f64\ types, implementing \Hash\, \Eq\, and \Ord\ traits to ensure consistent behavior in collections, including specific handling of NaN values where they are treated as equal to each other and less than any other value.
src/common/hashable-float-wrapper · high confidence
New utility modules for display, hashing, and type resolution
The \src/daft-core/src/utils\ directory now includes new modules that provide core formatting, hashing, and type-resolution capabilities. The \display\ module adds functions to format dates, times, timestamps (including timezone handling), durations, and decimal128 values for user-facing output, as well as series literal display. The \identity\_hash\_set\ module introduces an \IdentityHasher\ to bypass redundant hashing when working with pre-hashed values, improving performance for set operations. The \supertype\ module implements logic to determine the common supertype between two data types (e.g., promoting integers to floats or aligning time units), which is essential for type coercion in expressions. Additionally, \stats.rs\ provides utilities for calculating mean, variance, standard deviation, and skewness, while \ord.rs\ defines a dynamic comparator type for array comparisons.
src/daft-core/src/utils · high confidence
Read and write Hugging Face datasets with fallback and WebDataset support
Users can now read Hugging Face datasets via \read\_huggingface\, which automatically falls back to the \datasets\ library if Parquet files are unavailable and supports reading WebDataset TAR shards. Additionally, \DataFrame.write\_huggingface\ allows writing DataFrames directly to Hugging Face repositories as Parquet files, with configurable chunking and overwrite behavior.
daft/io/huggingface · high confidence
Repository initialization and developer workflow setup
The repository has been initialized with a comprehensive development environment setup. This includes a Makefile for managing the Python virtual environment (using uv), building the Rust/Python project (via maturin), running tests, and generating documentation. Pre-commit hooks are configured to enforce code quality using Ruff, mypy, and Rust tooling (clippy, fmt). Documentation is built using MkDocs with Material theme, and ReadTheDocs is configured for automated builds. The project also includes standard governance files like CONTRIBUTING.md, CODE\_OF\_CONDUCT.md, and SECURITY.md, along with configuration for IDEs (JetBrains, VS Code, Zed, Helix) and debugging (LLDB).
(repo-wide) · high confidence
Rust-based DaftContext with unified configuration and event subscriber framework
The core execution context has been ported to Rust, introducing a centralized \DaftContext\ singleton that manages shared execution and planning configurations alongside a new event subscriber system. Users can now attach and detach named subscribers to receive lifecycle events (such as query start, heartbeat, optimization, and end) for observability and dashboard integration. The context also exposes a partition cache utility to register in-memory data partitions for scanning, and provides Python bindings to access configuration and query metadata from the Rust side.
src/daft-context/src · high confidence
Unified IOConfig with support for new storage backends and OpenDAL integration
The IO configuration has been refactored into a single, unified \IOConfig\ struct that consolidates settings for S3, Azure, GCS, HTTP, Unity, Gravitino, Hugging Face, Tencent Cloud COS, GooseFS, and HDFS. This change introduces native configuration structs for previously unsupported or loosely integrated backends (Tencent Cloud COS, GooseFS, HDFS, Gravitino) and adds support for arbitrary OpenDAL-compatible backends via a generic \opendal\_backends\ map. Users can now configure protocol aliases to map custom URI schemes to existing backends, and the system automatically infers which backend configurations are relevant for a given set of URIs to optimize query plan displays.
src/common/io-config · high confidence
Unified literal handling and Python type conversions in daft-core
The literal system in \daft-core\ has been restructured to support a comprehensive set of data types and robust conversions between Rust and Python. The \Literal\ enum now explicitly includes variants for UUID, Float16, File (with media type subtypes like Video, Audio, Image, and HDF5), Tensor, SparseTensor, Embedding, and Map, enabling users to work with these types directly in expressions. New conversion modules (\conversions.rs\, \deserializer.rs\, \python.rs\) implement \FromLiteral\ and \IntoPyObject\ traits, allowing seamless creation of literals from Python objects (including dicts, lists, and PyObjects) and accurate round-tripping of complex types like timestamps with timezones, decimals, and file references. This change ensures that literal values are consistently represented and converted across the Python-Rust boundary, supporting the new \daft.File\ datatype and other advanced types.
src/daft-core/src/lit · high confidence
Architecture
Extracted schema and data types into a standalone daft-schema crate
The \daft-schema\ crate has been introduced to centralize the definition of core data types and schema structures. This change moves the \DataType\ enum (including support for Float16, UUID, Interval, and various media types like Image, Video, Audio, and HDF5), the \Field\ struct, and the \Schema\ class out of the core library into this dedicated module. It also includes the associated Python bindings (\PyDataType\, \PyField\, \PySchema\) and supporting enums such as \ImageMode\, \ImageFormat\, \MediaType\, and \TimeUnit\, providing a unified location for schema-related logic and serialization.
src/daft-schema · high confidence
Introduction of the new LogicalPlanBuilder and expression resolution system
The logical plan construction logic has been replaced with a new \LogicalPlanBuilder\ implementation located in \src/daft-logical-plan/src/builder\. This change introduces a fluent, immutable builder interface for constructing logical plans, replacing the previous mechanism. It also includes a new \resolve\_expr\ module that handles the resolution of unresolved columns, wildcard expansion, and list evaluation contexts within expressions. This refactoring centralizes the plan building logic and improves how expressions are validated and resolved against the plan schema.
src/daft-logical-plan/src/builder · high confidence
Logical plan operators are extracted into dedicated modules
The logical plan operators in \src/daft-logical-plan/src/ops\ have been reorganized into individual, dedicated source files (e.g., \agg.rs\, \asof\_join.rs\, \concat.rs\, \distinct.rs\, \explode.rs\, \filter.rs\, \join.rs\, \limit.rs\, \project.rs\, \window.rs\). This change improves code maintainability and clarity by separating the definitions and logic for each distinct plan node, while the \mod.rs\ file now serves as the central registry re-exporting these components.
src/daft-logical-plan/src/ops · high confidence
Restructured daft-dsl with dedicated modules for arithmetic, joins, and optimization
The \daft-dsl\ crate has been reorganized into distinct modules to improve code clarity and maintainability. Arithmetic operations (add, sub, mul, div, rem) are now implemented via a macro in a dedicated \arithmetic\ module. Join logic has been extracted into a \join\ module, which handles schema inference for different join types (inner, left, right, outer, semi, anti) and normalizes join keys by casting them to a common supertype. An \optimization\ module provides utilities for expression analysis, including identifying required columns, determining if an expression requires computation, and replacing columns with other expressions. The crate's public API in \lib.rs\ now explicitly exports these new modules alongside existing expression, function, and Python bindings.
src/daft-dsl/src · high confidence
Series construction and serialization refactored into dedicated modules
The \src/daft-core/src/series\ module has been reorganized to improve code clarity and maintainability. Construction logic is now split into \from.rs\ (handling conversions from external formats like Arrow and NdArray) and \from\_lit.rs\ (handling literal value creation and type compatibility). Serialization and deserialization logic has been moved to \serdes.rs\, implementing \serde\ traits for \Series\. The core \Series\ struct and its public API are defined in \mod.rs\, while the internal \SeriesLike\ trait in \series\_like.rs\ standardizes the interface for underlying array implementations.
src/daft-core/src/series · high confidence
Behavioural changes
Dashboard build process now supports cross-filesystem OUT\_DIR locations
The dashboard build script (build.rs) has been updated to handle cases where the Cargo output directory (OUT\_DIR) resides on a different filesystem than the source tree. Previously, moving frontend assets to the output directory would fail with a cross-device link error in such environments (e.g., CI caches, bind mounts, or separate volumes). The new implementation detects this condition and falls back to a recursive copy-and-remove strategy, ensuring the dashboard builds successfully regardless of where the target directory is located. Additionally, the build process now explicitly requires Node.js ≥22 and npm ≥11.10.0 to leverage supply-chain security features like min-release-age.
src/daft-dashboard · high confidence
Delta Lake I/O migrated to new DataSource architecture with performance and compatibility improvements
The Delta Lake read and write implementation has been refactored to use the new DataSource interface, replacing the previous scan mechanism. This change introduces several user-facing improvements: read performance is enhanced via predicate pushdowns that filter files using Delta Lake statistics, and column mapping is now supported for reading tables. The \read\_deltalake\ API now allows users to ignore deletion vectors during reads and correctly expands the tilde (\\~\) in local paths to the home directory. Additionally, file path construction for object storage (S3, GCS, Azure) now consistently uses forward slashes, and the codebase has been updated to support Delta Lake library versions 1.0.0 and later.
_daft/io/delta\lake · high confidence
Enforce lazy imports and add runtime type checking to the public API
Daft now enforces lazy imports for heavy dependencies (such as PyArrow, NumPy, Pandas, Ray, and fsspec) to significantly reduce startup time, and introduces a \PublicAPI\ decorator that performs runtime type checking on function arguments. This ensures that users receive immediate, clear errors when passing incorrect types to DataFrame and expression methods, while also preventing undeclared optional dependencies from being required at import time.
daft · high confidence
Iceberg integration refactored to new DataSource architecture with expanded read capabilities
The Iceberg I/O module has been restructured into a new \DataSource\-based implementation (\IcebergDataSource\), replacing the previous operator model. This change introduces support for reading Iceberg tables via specific branches and tags (in addition to snapshot IDs), automatically configures cloud storage credentials (including Alibaba Cloud OSS) from table properties, and enables more efficient filtering by pushing down \starts\_with\, \is\_nan\, and \not\_nan\ predicates. It also adds an \overwrite\_filter\ for static partition overwrites, a \count()\ pushdown optimization, and the \ignore\_corrupt\_files\ option to silently skip unreadable data files.
daft/io/iceberg · high confidence
Image operations now run in parallel and support hashing
Image processing operations such as resizing, cropping, and mode conversion in \src/daft-image\ are now executed in parallel using the Rayon library, improving performance for large batches. Additionally, the module introduces image hashing capabilities (supporting aHash, dHash, pHash, and wHash) to enable image deduplication, and includes new internal utilities like \CountingWriter\ and \ImageBufferIter\ to support these operations.
src/daft-image/src · high confidence
Introduce Flotilla scheduler with default and linear scheduling strategies
The distributed scheduling system now uses the new Flotilla scheduler, implemented in \src/daft-distributed/src/scheduling/scheduler\. This change introduces a \Scheduler\ trait with two concrete implementations: \DefaultScheduler\, which supports spread and soft worker-affinity scheduling strategies and triggers autoscaling based on a configurable threshold (default 1.25, overridable via the \DAFT\_AUTOSCALING\_THRESHOLD\ environment variable), and \LinearScheduler\, which schedules tasks one at a time and only when no workers have active tasks. The scheduler actor (\scheduler\_actor.rs\) manages the event loop, handling task submission, worker state updates, autoscaling requests, task dispatching, and idle worker retirement, while emitting task lifecycle events (Submitted, Scheduled, Cancelled) to the statistics manager.
src/daft-distributed/src/scheduling/scheduler · high confidence
Introduce Rust-based runner abstraction with Ray autoscaling and timeout controls
The \src/daft-runners\ module now provides a Rust implementation for managing execution runners, exposing Python bindings to configure and retrieve the active runner. Users can now explicitly set the Ray runner via \set\_runner\_ray\, which supports new parameters for \worker\_startup\_timeout\ and autoscaling configuration (\autoscale\_strategy\ and \autoscale\_bisect\_timeout\_secs\). The implementation validates that the autoscaling strategy is either 'gradual' or 'bisect' and ensures the bisect timeout is greater than zero, failing fast if these constraints are violated. This change centralizes runner initialization logic in Rust while maintaining compatibility with the existing Python runner classes.
src/daft-runners · high confidence
Introduce dedicated Python serialization module using Daft pickle and bincode
This change adds a new \py-serde\ module that provides serialization and deserialization capabilities for Python objects. It replaces previous serialization mechanisms by using Daft's internal pickle implementation (\daft.pickle\) for converting Python objects to byte arrays, and bincode for serializing Rust state. The module introduces a \PyObjectWrapper\ struct that makes Python UDFs and other Python objects serializable and hashable within the Rust ecosystem, ensuring consistent behavior for Python-side logic during distributed execution.
src/common/py-serde · high confidence
Introduce native Rust-based system information module
A new \SystemInfo\ module has been added to \src/common/system-info\, providing a Rust implementation for retrieving CPU count and total memory. This module leverages the \sysinfo\ crate to gather system metrics, including support for cgroup memory limits, and exposes these capabilities via a Python binding (\SystemInfo\ class) when the \python\ feature is enabled, effectively replacing the previous reliance on the \psutil\ library for these specific system queries.
src/common/system-info · high confidence
Introduce new Python UDF execution path in the DSL layer
The \src/daft-dsl/src/functions/python\ module has been restructured to define a \LegacyPythonUDF\ function expression and its associated \FunctionEvaluator\. This change introduces a dedicated code path for evaluating Python UDFs, handling initialization via \MaybeInitializedUDF\, managing resource requests, and invoking the Python runtime through \pyo3\ to execute the underlying Python logic. This serves as the DSL-side foundation for the new UDF execution model, separating the expression definition from the legacy implementation details.
src/daft-dsl/src/functions/python · high confidence
Introduce new logical plan builder and schema abstractions
The \daft/logical\ module has been restructured to introduce a new \LogicalPlanBuilder\ class that wraps the underlying Rust planner, providing Python-facing methods for operations such as \select\, \with\_columns\, \filter\, \limit\, \offset\, \shard\, \explode\, and \describe\. This builder automatically applies the current \DaftPlanningConfig\ to instantiated plans. Additionally, a new \Schema\ class is exposed from \daft.logical.schema\, re-exporting \DataType\, \Field\, and \Schema\ from the core \daft.schema\ module, and a \MapPartitionOp\ base class with an \ExplodeOp\ implementation is added to handle partition-level transformations.
daft/logical · high confidence
Introduce new runner architecture with Flotilla and Native runners
The execution engine has been replaced with a new runner architecture. The \daft/runners\ module now exposes a \daft.context\ API to configure execution via \set\_runner\_native()\ (for local multi-threaded processing) or \set\_runner\_ray()\ (for distributed Ray execution). The previous RayRunner and PyRunner have been removed and replaced by the Flotilla runner (a Ray-based distributed runner) and the Native runner. The new Ray runner includes configurable autoscaling strategies (gradual and bisect), worker downscaling, and configurable worker startup timeouts. Additionally, a new \Heartbeat\ mechanism and \query\_id\ emission have been added to support observability and query tracking.
daft/runners · high confidence
Introduces centralized planning and execution configuration structs
This change introduces \DaftPlanningConfig\ and \DaftExecutionConfig\ as the central configuration structures for the Daft engine, replacing scattered environment variable usage with a unified, serializable config system. \DaftPlanningConfig\ exposes settings for join reordering, strict filter pushdown, and the experimental DP-ccp join ordering algorithm, while \DaftExecutionConfig\ provides granular control over scan task sizing, shuffle algorithms, parquet/CSV/JSON write parameters, actor UDF timeouts, and dynamic batching. These configs are exposed to Python via \PyDaftPlanningConfig\ and \PyDaftExecutionConfig\, allowing users to programmatically tune performance characteristics such as \native\_parquet\_writer\, \shuffle\_algorithm\, and \pre\_shuffle\_merge\_threshold\ without relying solely on environment variables.
src/common/daft-config · high confidence
Introduction of Flotilla distributed execution plan runner
The \src/daft-distributed/src/plan\ module has been replaced with a new implementation that introduces the Flotilla execution engine. This change adds a \PlanRunner\ and \PlanExecutionContext\ to manage the lifecycle of distributed queries, including task scheduling via a \SchedulerHandle\, worker management, and statistics collection. It defines \DistributedPhysicalPlan\ to hold logical plans and execution configurations, and \PlanResult\ to stream materialized outputs back to the user. This new structure decouples the plan execution logic from the previous implementation, enabling features like fault tolerance, progress tracking, and OTEL metrics export within the distributed scheduler.
src/daft-distributed/src/plan · high confidence
Migrate array implementations from Arrow2 to arrow-rs
The array implementations in daft-core have been rewritten to use the arrow-rs library instead of the previous Arrow2 dependency. This migration updates core array types—including DataArray, ListArray, FixedSizeListArray, StructArray, UnionArray, and specialized arrays like FileArray and ImageArray—to use arrow-rs types for storage, iteration, and serialization. Users benefit from improved compatibility with the broader Rust Arrow ecosystem and more efficient, native Rust array operations, while the public API for creating and manipulating arrays remains consistent.
src/daft-core/src/array · high confidence
Migrate array operations to arrow-rs and implement native arithmetic kernels
The array operations in \daft-core\ have been migrated from arrow2 to arrow-rs, introducing native implementations for arithmetic, sorting, and aggregation. Arithmetic operations (add, sub, mul) now use custom, vectorized kernels that handle broadcasting and nulls directly, avoiding the conversion overhead of arrow-rs kernels. Sorting has been rewritten with specialized primitive sort paths that support null ordering and limits. Aggregation functions like \bool\_and\, \bool\_or\, and \approx\_count\_distinct\ are now implemented natively using arrow-rs types and the \daft\_sketch\ crate.
src/daft-core/src/array/ops · high confidence
Native Rust implementation of hash and merge joins
The join operations in \daft-recordbatch\ have been rewritten in native Rust, replacing the previous implementation. This change introduces dedicated \hash\_join.rs\ and \merge\_join.rs\ modules, providing native support for inner, left, right, outer, semi, and anti joins via hash tables, as well as a state-machine-based merge join for sorted data. Users benefit from improved performance and stability as these core relational operations are now executed directly in Rust without external dependencies like arrow2 for the join logic itself.
src/daft-recordbatch/src/ops/joins · high confidence
Native parquet reader now uses arrow-rs with field ID mapping and Iceberg positional delete support
The parquet reader in \src/daft-parquet\ has been rewritten to use the arrow-rs library instead of the previous arrow2/parquet2 stack. This change introduces a new \DaftParquetMetadata\ adapter to decouple the crate from the underlying parquet implementation, enabling support for reading Parquet files with Field IDs (renaming and filtering columns based on a mapping) and Iceberg Merge-on-Read positional deletes (skipping deleted rows via \RowSelection\). The reader also now supports predicate pushdown, row-group level statistics for scan task estimation, and configurable string encoding (UTF-8 vs Raw). Error handling has been restructured to provide more specific error variants (e.g., \MissingParquetFieldIds\, \InvalidParquetFile\) and the Python bindings expose these new capabilities via \read\_parquet\ and \read\_parquet\_bulk\ functions.
src/daft-parquet/src · high confidence
New aggregation and scalar function infrastructure
The \daft-dsl\ functions module has been restructured to support a new two-stage User-Defined Aggregate Function (UDAF) pipeline via the \AggFn\ trait, which manages aggregation, combination, and finalization stages with typed intermediate state. Scalar functions now utilize a \FunctionArgs\ system that standardizes argument parsing across SQL and Python frontends, supporting named, unnamed, optional, and variadic parameters. Additionally, a \merge\_mean\ function is introduced to correctly compute means from partial sums and counts, with specific handling for Decimal128 precision and scale.
src/daft-dsl/src/functions · high confidence
New arrow-based comparison and sort modules
The \src/daft-core/src/array/ops/arrow\ directory now exposes \comparison\ and \sort\ modules. The comparison module introduces \build\_is\_equal\ and \build\_multi\_array\_is\_equal\ functions that handle equality checks for arrays, including specific logic for floating-point NaN values and null handling, leveraging the \arrow\ crate's \make\_comparator\ and primitive array utilities.
src/daft-core/src/array/ops/arrow · high confidence
New boolean algebra crate for expression simplification and CNF/DNF conversion
A new \daft-algebra\ crate has been introduced to house symbolic and boolean algebra logic, specifically providing functions to convert boolean expressions into Conjunctive Normal Form (CNF) and Disjunctive Normal Form (DNF), split conjunctions/disjunctions, and determine if filter predicates remove nulls. This refactoring moves these capabilities out of the core DSL, enabling more efficient query optimization by allowing the optimizer to push filter predicates through joins and simplify join types using null-eliminating predicates.
src/daft-algebra/src · high confidence
New boolean, numeric, and null expression simplification rules
The expression simplifier in \src/daft-algebra/src/simplify\ has been expanded with new optimization rules. Boolean expressions are now simplified by combining common sub-expressions in AND/OR chains (e.g., \(A OR B) AND (A OR C)\ becomes \A OR (B AND C)\), and \IS IN\ checks on small literal lists are converted into OR chains of equality checks. Numeric expressions are simplified by removing identity operations (e.g., \A \* 1\, \A + 0\). Null handling is also optimized, such as reducing \is\_null(NULL)\ to \true\ and propagating nulls through most binary operators. These changes improve query performance by reducing expression complexity before execution.
src/daft-algebra/src/simplify · high confidence
New datatype module with aggregation supertype inference and logical array support
The \src/daft-core/src/datatypes\ module has been reorganized into dedicated files, introducing \agg\_ops.rs\ which defines the output data types for aggregations like sum, product, mean, standard deviation, variance, skew, and percentiles (e.g., numeric inputs cast to Float64, Decimal128 inputs to Float64 for mean/stddev). It also adds \infer\_datatype.rs\ to handle type inference for logical operations (AND/OR/XOR), comparisons, and clipping, and \logical.rs\ to implement \LogicalArrayImpl\ for wrapping physical arrays with semantic types like Date, Time, Timestamp, and Map.
src/daft-core/src/datatypes · high confidence
New incremental window aggregation state management
The window aggregation engine in \daft-recordbatch\ has been refactored to use a new \WindowAggStateOps\ trait that supports incremental updates. This allows window functions to efficiently add and remove rows from the aggregation state as the window slides, rather than recomputing from scratch. The implementation includes new state managers for \sum\, \count\, \count\_distinct\, \mean\, \first\_value\, \last\_value\, and \min/max\ (though min/max is currently commented out in the factory). This change enables more efficient processing of dynamic window frames and partition-by windows.
_src/daft-recordbatch/src/ops/window\states · high confidence
New internal grouping primitives and Float16 support in group-by operations
This change introduces new internal helper traits (\IntoGroups\, \IntoUniqueIdxs\) and implementations in the \daft-groupby\ module to optimize how group indices and unique value indices are computed for DataFrames. It adds explicit support for Float16 (half-precision) arrays, ensuring NaN values are canonicalized so they hash consistently into a single group. These low-level optimizations improve the performance of group-by and list-aggregation operations by providing more efficient underlying array grouping logic.
src/daft-groupby · high confidence
New optimizer infrastructure with cycle detection and join key tracking
The optimization module now includes a \JoinKeySet\ for efficiently tracking equality join keys and a \LogicalPlanTracker\ that uses plan digests to detect and prevent infinite optimization cycles. The \Optimizer\ has been restructured to use \RuleBatch\es with configurable execution strategies (Once or FixedPoint), allowing for more robust and ordered application of optimization rules such as filter pushdown, join reordering, and expression simplification.
src/daft-logical-plan/src/optimization · high confidence
Refactor binary operations to delegate to Python utilities
The binary operation logic in the Rust core has been refactored to delegate execution to Python utility functions. New files in \src/daft-core/src/series/utils\ introduce \python\_fn.rs\, which handles binary operators, membership checks, and between-checks by casting series to Python types, invoking specific functions in \daft.utils\ (such as \map\_operator\_arrow\_semantics\ and \python\_list\_membership\_check\), and reconstructing the result series. A new \cast\ module provides a macro for downcasting operations. This change shifts the implementation of these specific series operations from native Rust to the Python layer.
src/daft-core/src/series/utils · high confidence
Refactor expression column binding to use bound indices
The expression system now distinguishes between unresolved, resolved, and bound columns, introducing a \BoundExpr\ type that enforces column binding at the type level. This change shifts column resolution from name-based lookups to index-based access, ensuring that expressions used in logical plan operations contain only bound columns. The refactoring includes new modules for aggregation expression handling, window specifications, and display logic, providing a stricter boundary between unbound and bound states in the DSL.
src/daft-dsl/src/expr · high confidence
Refactored Series implementation to use a unified ArrayWrapper pattern
The internal implementation of Series in daft-core has been restructured to use a new \ArrayWrapper\ type and dedicated modules (\data\_array\, \logical\_array\, \nested\_array\, \python\_array\) for different array categories. This change standardizes how various array types (including primitive, logical, nested, and Python arrays) implement the \SeriesLike\ trait, ensuring consistent behavior for operations like filtering, casting, and aggregation across all data types.
_src/daft-core/src/series/array\impl · high confidence
Reliable benchmarking utilities with retry logic and metadata collection
The benchmarking workflow now includes a new shared utility module that enhances reliability and traceability. Google Sheets uploads are protected by automatic retries (up to 5 attempts with exponential backoff) for transient network errors (HTTP 429, 500, 502, 503, 504), ensuring results are not lost due to temporary API failures. Additionally, benchmark runs now automatically capture and report precise metadata, including the exact Daft version, the specific Git reference name, and the commit SHA, providing clearer context for each result set.
benchmarking · high confidence
Renamed Table to RecordBatch and introduced MicroPartition
The internal data structures have been renamed to better reflect their semantics: the \Table\ class is now \RecordBatch\, and a new \MicroPartition\ class has been introduced to manage collections of record batches. This change updates the public API in the \daft.recordbatch\ module, exposing \RecordBatch\ and \MicroPartition\ as the primary types for data manipulation, along with helper functions for reading Parquet files directly into PyArrow. The \MicroPartition\ class provides methods for creating, slicing, concatenating, and exporting data (to Arrow, Pandas, Python dicts/lists), while \RecordBatch\ handles single-batch operations and I/O for CSV, JSON, and Parquet formats.
daft/recordbatch · high confidence
Replaced Arrow2-based growable arrays with a new high-performance Arrow-rs implementation
The internal array growth mechanism in \daft-core\ has been rewritten to use the \arrow-rs\ crate instead of the previous \arrow2\ dependency. This change introduces a new set of specialized growable types (e.g., \ArrowGrowable\, \ListGrowable\, \StructGrowable\) that handle buffer manipulation directly, resulting in approximately 2x faster performance for operations like concatenation and filtering compared to the previous implementation. Users benefit from improved performance in data processing pipelines without any changes to the public API.
src/daft-core/src/array/growable · high confidence
Rewritten Parquet reader using arrow-rs public decoder API
The Parquet reader implementation in \src/daft-parquet/src/reader\ has been completely rewritten to use the public decoder API from \arrow-rs\. This change introduces new internal modules (\chunk\_source\, \field\_reader\, \rg\_processor\, \util\) that handle file opening, column chunk sourcing, and row group processing. The new reader uses \ArrowReaderMetadata\ for metadata handling and \SerializedPageReader\ for page iteration, replacing the previous internal reader logic. It also integrates with the \daft\_io\ client for remote file access and applies Iceberg field-id mappings before column filtering to ensure correct prefetch matching. The batch size for array readers is now derived from the chunk size, and the reader supports both local and remote file sources with optimized column projection and predicate pushdown.
src/daft-parquet/src/reader · high confidence
Specialized hash-join probe tables for integer types
The join probe mechanism now uses a dedicated, optimized hash table for integer columns (Int8 through UInt64) to accelerate hash joins. This change introduces a specialized \IntProbeTable\ that stores integer keys directly, bypassing the generic multi-column comparison logic used for other data types, which improves performance for integer-based join keys.
src/daft-recordbatch/src/probeable · high confidence
Unified error handling with cleaner Python exception messages
The \src/common/error\ module now provides a centralized \DaftError\ enum and a dedicated formatting pipeline for Python exceptions. This change introduces a new \format\_error\_for\_user\ function that uses \snafu::CleanedErrorText\ to deduplicate overlapping text in nested error chains, resulting in cleaner, more readable error messages for users. It also maps specific Rust error variants (such as timeouts, throttling, and file-not-found) to distinct Python exception types in \daft.exceptions\, improving error classification and debugging for Python developers.
src/common/error · high confidence
Test coverage
1503 commits adding/updating tests in tests; Add Parquet benchmarking suite for performance comparison; Added Parquet read benchmarks for nested types, codecs, and filter pushdown; Added benchmarks for image operations; Added integration tests for S3-backed CheckpointStore; Added test helpers for logical plan optimization; Added test utilities for dummy scan operators in logical plan tests; Benchmarking suite for LeRobot video decode performance.
Dependencies
Dependency updates for AI benchmarks and example extensions
Updated the \daft\ and \ray\[default\]\ dependencies in the AI benchmarking projects (audio transcription, document embedding, image classification, and video object detection) to versions 0.6.2 and 2.49.2 respectively. Additionally, updated the \dvector\ and \hello\ example extension crates to use \arrow\ versions 57.1.0 and 58 respectively, and added a new \nav-hide-children\ MkDocs plugin dependency.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 56 → 54 (-2.6)
- Rubric changed (rubric-2026.09.10 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 66 → 66 (-0.1)
- Architecture 98 → 98 (+0.6)
- Maturity 72 → 72 (-0.0)
- Readiness 56 → 47 (-8.7)
- Security 58 → 60 (+2.7)
- Event Sourcing 100 → 100 (+0.0)
- Accessibility 49 → 48 (-1.7)
- Performance 100 (new)
Resolved (66)
- Boundary-crossing change coupling: expressions.py ↔ mod.rs (daft/expressions/expressions.py)
- Boundary-crossing change coupling: expressions.py ↔ mod.rs (daft/expressions/expressions.py)
- Change coupling: repartition.rs ↔ translate.rs (src/daft-distributed/src/pipeline_node/shuffles/repartition.rs)
- DeltaLakeDataSource.get_tasks (cognitive 55) (daft/io/delta_lake/delta_lake_scan.py)
- DeltaLakeDataSource.get_tasks (cyclomatic 30) (daft/io/delta_lake/delta_lake_scan.py)
- Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (no pnpm-resolved versions to grade)
- Documentation: no architecture or design documentation (README.rst)
- Duplicated block (10 lines × 2) (src/daft-core/src/array/struct_array.rs)
- Duplicated block (10 lines × 2) (src/daft-functions/src/slice.rs)
- Duplicated block (13 lines × 2) (src/daft-csv/src/metadata.rs)
- Duplicated block (16 lines × 2) (src/daft-writers/src/batch_file_writer.rs)
- Edited copy of a member (13 corresponding lines) (src/daft-functions-utf8/src/lpad.rs)
- Edited copy of a member (14 corresponding lines) (src/daft-distributed/src/pipeline_node/sort.rs)
- Edited copy of a member (18 corresponding lines) (src/daft-functions-utf8/src/regexp_extract.rs)
- Edited copy of a member (20 corresponding lines) (src/daft-dsl/src/python_udf/batch.rs)
- Edited copy of a member (40 corresponding lines) (src/daft-functions-temporal/src/date_construction.rs)
- Edited copy of a member (77 corresponding lines) (src/daft-functions-temporal/src/date_construction.rs)
- High CVE: [CVE redacted] (src/daft-dashboard/frontend/package-lock.json)
- High CVE: [CVE redacted] (src/daft-dashboard/frontend/package-lock.json)
- High CVE: [CVE redacted] (src/daft-dashboard/frontend/package-lock.json)
- …and 46 more
New (387)
- AzureBlobSource::list_containers_stream (cognitive 18) (src/daft-io/src/azure_blob.rs)
- AzureBlobSource::list_directory_delimiter_stream (cognitive 52) (src/daft-io/src/azure_blob.rs)
- AzureBlobSource::list_directory_delimiter_stream (cyclomatic 17) (src/daft-io/src/azure_blob.rs)
- Boundary-crossing change coupling: expressions.py ↔ mod.rs (daft/expressions/expressions.py)
- Boundary-crossing change coupling: expressions.py ↔ mod.rs (daft/expressions/expressions.py)
- Boundary-crossing change coupling: flotilla.py ↔ worker_manager.rs (daft/runners/flotilla.py)
- ClassTooLong: AzureBlobSource (src/daft-io/src/azure_blob.rs)
- ClassTooLong: RayWorkerManager (src/daft-distributed/src/python/ray/worker_manager.rs)
- DeltaLakeDataSource._get_scan_tasks (cognitive 55) (daft/io/delta_lake/delta_lake_scan.py)
- DeltaLakeDataSource._get_scan_tasks (cyclomatic 30) (daft/io/delta_lake/delta_lake_scan.py)
- Duplicated block (10 lines × 2) (daft/catalog/__postgres.py)
- Duplicated block (10 lines × 2) (src/daft-core/src/array/struct_array.rs)
- Duplicated block (10–11 lines × 2) (src/daft-io/src/azure_blob.rs)
- Duplicated block (11 lines × 2) (daft/io/_sql.py)
- Duplicated block (12 lines × 2) (daft/catalog/__postgres.py)
- Duplicated block (12 lines × 2) (daft/recordbatch/micropartition.py)
- Duplicated block (12 lines × 2) (daft/recordbatch/recordbatch.py)
- Duplicated block (12 lines × 2) (daft/udf/execution.py)
- Duplicated block (12–13 lines × 3) (daft/recordbatch/recordbatch_io.py)
- Duplicated block (12–13 lines × 3) (src/daft-writers/src/avro_writer.rs)
- …and 367 more
Changes since last survey
- 28 commits — 20 feature/other, 8 fixes
By area
- daft/io — 6 commits
- src/daft-io — 5 commits
- .github/workflows — 2 commits
- tests/integration — 2 commits
- (root) — 1 commit
- daft/ai — 1 commit
- daft/catalog — 1 commit
- daft/functions — 1 commit
- daft/logging.py — 1 commit
- daft/runners — 1 commit
- docs/extensions — 1 commit
- src/common — 1 commit
- src/daft-avro — 1 commit
- src/daft-core — 1 commit
- src/daft-mcap — 1 commit
- src/daft-sql — 1 commit
- tests/io — 1 commit
Notable commits
- fix: fix!: Don't use HTTPConfig.bearer_token for HF auth (#7525)
- fix: fix(ai): forward extra_body to embedding requests (#7516)
- fix: fix(iceberg): apply partition predicates to count pushdown (#7421)
- fix: fix(io): enforce timeouts on OpenDAL clients (#7523)
- fix: fix(io): parse container@host Azure URIs on Microsoft Fabric / OneLake hosts (#7533)
- fix: fix(io): preserve Hugging Face glob revisions (#7322)
- fix: fix(sql): report read_iceberg argument errors instead of panicking (#7354)
- fix: fix: eliminate TOCTOU race in CREATE TABLE IF NOT EXISTS (#7387)
- change: chore!: remove deprecated llm_generate (#7450)
- change: chore!: remove deprecated min_cpu_per_task from ExecutionContext (#7465)
- change: chore(deps): bump the all group across 1 directory with 7 updates (#7501)
- change: chore(deps): upgrade daft-lance to 0.5.0 (#7512)
- change: chore(lance)!: remove deprecated daft.io.lance module (#7514)
- change: chore(lance): deprecate write_lance merge mode (#7526)
- change: chore(logging)!: remove deprecated setup_debug_logger() (#7427)
- change: chore: Add check for mutable values in default arguments (#7494)
- change: chore: Upgrade Rust nightly (#7393)
- change: ci: Hide Dependency Diff when no changes (#7552)
- change: ci: Use SeaweedFS instead of MinIO (#7550)
- change: docs: add daft-doris to Community Extensions (#7509)
- …and 8 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Eventual-Inc/Daft was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit dadd8a0b290be148d92acb6f9e6fc4b6e36f221f — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5ff527f25b99.