Skip to content
CAI
Software that uses CAICheck a score

lance-format/lance

70.0

Adequate · 29 September 2026

776.5k

lines of production code

Rust

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a high-performance columnar data storage and vector search engine, implemented primarily in Rust with Python and Java SDKs. It provides a modern file format (Lance) optimized for efficient I/O, supporting complex nested data, blob storage, and various scalar and vector index types like IVF-PQ and HNSW. The platform enables scalable data ingestion, concurrent updates via MemWAL, and seamless integration with Apache DataFusion for SQL querying and PyTorch for machine learning workflows.

How it got here

2022–2023 — modularization and vector index v2

53 changes.

This period focused on restructuring the Rust codebase into specialized crates and implementing a new v2 architecture for IVF vector indexes. It also introduced comprehensive benchmarking suites, PyTorch integration, and DataFusion support to enhance performance profiling and ecosystem compatibility.

2024–2025 — Lance 2.0 format and Java SDK launch

82 changes.

This period focused on the foundational rewrite of the Lance file format to version 2.x, introducing a modular encoding architecture, stable row IDs, and data overlay capabilities. It also marked the initial release of the Java SDK, providing comprehensive bindings for dataset operations, vector indexing, and full-text search. Concurrently, the project expanded its vector index support with new HNSW and Scalar Quantization implementations while establishing robust CI and benchmarking infrastructure.

2026 — MemWAL and file format evolution

51 changes.

This period focused on implementing MemWAL, an in-memory write-ahead log with LSM-style tiering, to enable high-throughput, low-latency writes and unified in-memory search across scalar, vector, and full-text indexes. Concurrently, the project advanced its storage layer by introducing the Lance v2.0 file format and its subsequent versions, alongside a comprehensive refactor of the inverted index and distributed index merging infrastructure.

Features

Add BigANN benchmark dataset preparation scripts

Added a new benchmark suite for the BigANN dataset, including a Python script (\dataset.py\) and documentation (\README.md\) to prepare text-to-image datasets in Lance format. The script supports creating \text2image-10m\ and \yfcc-10m\ datasets, converting raw binary data into Lance files with ground truth queries, and includes a \.gitignore\ to exclude generated artifacts.

benchmarks/bigann · high confidence

Add Cohere Wikipedia embedding benchmark for index build performance

A new benchmark located in benchmarks/wiki allows users to generate a 35M vector dataset from the Cohere Wikipedia embeddings and build an IVF\_PQ index using a disk-based shuffler. Users can run datagen.py to create the Lance dataset and index.py to construct the index with configurable metrics (L2, cosine, dot), partition counts, and sub-vector counts, specifically designed to measure indexing speed rather than recall or latency.

benchmarks/wiki · high confidence

Add Dbpedia-entities benchmark for vector search performance evaluation

Users can now run a benchmark using the Dbpedia-entities-openai dataset (1M embeddings) to evaluate vector search performance. This includes a data generation script (\datagen.py\) to convert the Hugging Face dataset into Lance format, and a benchmark runner (\benchmarks.py\) that tests top-k recall across various IVF\_PQ index configurations (IVF 256/512/1024, PQ 32/96/192) and refine factors.

benchmarks/dbpedia-openai · high confidence

Add FSST compression library and benchmark example

Introduces the FSST (Fast Static Symbol Table) compression implementation in \rust/compression/fsst\, including the core logic in \src/fsst.rs\ and library entry point in \src/lib.rs\. Adds a benchmark example (\examples/benchmark.rs\) that demonstrates compressing and decompressing string data using the FSST algorithm, highlighting the output buffer size contract (8x expansion) and performance metrics.

rust/compression/fsst · high confidence

Add OpenTelemetry metrics bridge for Lance

The Java library now includes a new \org.lance.otel\ package that bridges Lance's internal Rust metrics to OpenTelemetry. Users can call \LanceMetrics.instrument()\ to register observable instruments (counters, gauges, and histograms) on the global OpenTelemetry meter provider, allowing Lance's native metrics to be exported via standard OTel pipelines. The bridge supports querying the metric catalog and taking point-in-time snapshots of recorded data.

java/src/main/java/org/lance/otel · high confidence

Add Python benchmark dataset generation utilities

The \ci\_benchmarks/datagen\ module now provides Python scripts to generate synthetic datasets for CI benchmarks, including a basic dataset, a TPC-H lineitems dataset (using DuckDB), a Wikipedia dataset for full-text search (streaming from HuggingFace), a count\_rows dataset with various scalar indexes, and merge-insert datasets with specific fragmentation and deletion patterns.

_python/python/ci\benchmarks/datagen · high confidence

Add owned, SIMD-accelerated bitpacking codecs

The \lance-bitpacking\ crate now includes its own internal bitpacking implementation (\bitpacker\_internal\) featuring SIMD-optimized kernels for x86\_64 (SSE3, AVX2) and aarch64 (NEON). This change introduces new owned codec types (\BitPacker4x\, \BitPacker8x\) that provide high-performance compression and decompression for \u32\ blocks, replacing or augmenting previous external dependencies with a Lance-owned, byte-compatible implementation.

rust/compression/bitpacking · high confidence

Add standalone CLI benchmark for PK-based point lookups across LSM levels

A new standalone CLI benchmark (\mem\_wal\_point\_lookup\_bench\) has been added to measure point lookup latency against three tiers of the LSM tree: the base table (on-disk, compacted data), SSTables (on-disk L0), and the active MemTable (in-memory write buffer). The tool supports two phases—\prepare\ to create the base dataset and initialize MemWAL, and \lookup\ to ingest rows via ShardWriter and time point lookups—allowing users to configure parameters such as base rows, maximum MemTable rows, number of SSTables, and query count, with results optionally output as JSON.

_rust/lance/benches/mem\_wal/point\lookup · high confidence

Added ExpLinkedList for memory-efficient storage

A new \ExpLinkedList\ container has been added to \lance-core\. This data structure stores elements in a linked list of vectors, where each vector's capacity doubles when full, providing a memory-efficient way to handle large numbers of elements. It supports standard operations like push, pop, and iteration, and implements \DeepSizeOf\ for accurate memory accounting.

rust/lance-core/src/container · high confidence

Added IOTracker for monitoring and metrics of local storage operations

A new \IOTracker\ utility has been added to \rust/lance-io/src/utils/tracking\_store.rs\ to wrap ObjectStore implementations and track I/O statistics. This component enables the recording of read and write operations, including byte counts and IOPS, specifically for local reads and writes that bypass the standard ObjectStore layer. It also integrates with the metrics system to publish object store metrics for these local operations, ensuring that performance data is accurately captured even when optimized local paths are used.

rust/lance-io/src/utils · high confidence

Added Rust crate documentation and build configuration for Lance

This change introduces the initial structure for the Rust implementation of the Lance format. It adds a README file detailing installation via Cargo and providing code examples for creating datasets, reading data, performing vector index operations, and other core features. Additionally, it includes a build script (build.rs) that configures the Protobuf compilation process for ANN (Approximate Nearest Neighbor) protocols and sets up external path mappings for generated code.

rust/lance · high confidence

Added TPCH benchmark script to compare Lance and Parquet performance

Users can now run a Python-based benchmark in the \benchmarks/tpch\ directory to compare query performance between Lance and Parquet formats. The new \benchmark.py\ script supports TPCH queries Q1 and Q6, allowing users to generate sample datasets locally using DuckDB and measure latency differences between the two storage engines.

benchmarks/tpch · high confidence

Added code agent skills for the Lance user guide

This change introduces a new 'skills' directory containing structured documentation and scripts designed to guide code agents in assisting Lance users. It includes a primary skill definition (\lance-user-guide/SKILL.md\) that instructs agents on how to help with dataset operations (write, read, scan), vector and scalar index creation, and troubleshooting. The entry also adds reference materials for index selection and I/O operations, along with a Python end-to-end script (\python\_end\_to\_end.py\) that demonstrates a complete workflow of writing data, building indices, and performing queries.

skills · high confidence

Added flat vector index benchmark script

A new benchmark script has been added to the \benchmarks/flat\ directory to measure the latency of flat vector search. Running the script generates synthetic datasets with varying dimensions and lengths, executes nearest-neighbor searches using L2, cosine, and dot-product metrics, and outputs the results to \benchmark.csv\ along with a latency plot in \benchmark.html\.

benchmarks/flat · high confidence

Added fragment reuse index management and cleanup

The dataset now exposes a public accessor to retrieve the fragment reuse index (FRI), which tracks how physical row addresses have moved during compactions that defer index remapping. Additionally, a new cleanup function automatically trims the FRI by removing older reuse versions once all existing indices have caught up to them, ensuring that stale mapping data does not accumulate unnecessarily.

rust/lance/src/dataset/index · high confidence

Builder-style API for scalar index configuration and zone statistics

Java users can now configure scalar indices (B-Tree, Bitmap, Inverted, LabelList, NGram, ZoneMap) using dedicated builder classes (e.g., BTreeIndexParams.builder()) instead of raw JSON or maps. The InvertedIndexParams builder exposes full-text tuning options including analyzer presets (text/code), base tokenizers (simple, whitespace, raw, ngram, code, ICU, Lindera, Jieba), stemming, stop words, n-gram ranges, posting block size, and format version gating. B-Tree and ZoneMap builders allow setting zone/row sizing, while Bitmap supports explicit shard IDs for distributed builds. Additionally, a new ZoneStats class exposes per-zone min/max/null counts from zonemap indices via JNI, enabling users to inspect index statistics directly in Java.

java/src/main/java/org/lance/index/scalar · high confidence

Calibrated auto-probing for dot-product IVF\_FLAT searches

This change introduces a new benchmarking suite and protocol for calibrating the initial probe budget in IVF\_FLAT indices using dot-product metrics. The new \benchmarks/auto-ivf-dot\ directory contains Python scripts (\prepare.py\, \calibrate.py\, \measure.py\, \audit.py\) and documentation (\PROTOCOL.md\, \README.md\, \RESULTS.md\) that define a frozen evaluation contract on the Wiki-Cohere 35M and DPR Wikipedia datasets. The calibration process selects specific floor, cap, and margin parameters for the auto-probing policy to minimize scanned partitions while maintaining at least 95% recall, addressing issues with the legacy signed-distance multiplier that previously performed poorly on dot products with varying query norms.

benchmarks/auto-ivf-dot · high confidence

Establish automated release and dependency governance tooling

The repository now includes configuration for \bumpversion\ to manage synchronized version bumps across Rust, Python, and Java artifacts, alongside \cargo-deny\ to enforce dependency license compliance, security advisory checks, and workspace dependency consistency. Pre-commit hooks are added to enforce code formatting, spell-checking, and lockfile synchronization, while a formal release process document and CI infrastructure (including LocalStack for S3 testing) standardize how beta, RC, and stable releases are built and published.

(repo-wide) · high confidence

Expose Lance datasets as DataFusion TableProviders

Users can now query Lance datasets directly using DataFusion's SQL and DataFrame APIs. This change introduces a \TableProvider\ implementation for \Dataset\ (in \logical\_plan.rs\) and a configurable \LanceTableProvider\ (in \dataframe.rs\) that supports filter, projection, and limit pushdown. The provider allows optional inclusion of system columns (\\_row\_id\, \\_row\_addr\), custom blob handling policies, and batch size configuration, enabling seamless integration of Lance data into DataFusion-based query engines.

rust/lance/src/datafusion · high confidence

Initial Java SDK release with Maven wrapper and documentation

This change introduces the initial Java bindings and SDK for Lance, providing a new Maven-based project structure under the \java/\ directory. It includes a Maven wrapper (version 3.3.2) for consistent builds, a \README.md\ with quick-start examples for dataset creation, random access, and schema evolution, and configuration files for code formatting (Spotless, Scalafmt). The release also bundles generated third-party license documents for both Java and Rust dependencies to ensure compliance.

java · high confidence

Introduce Binary Quantization (BQ) and RaBitQ support for vector indexes

Added a new Binary Quantization (BQ) module and RaBitQ (Rabit Quantization) support to the vector index pipeline. This includes a new \BinaryQuantization\ transformer that converts float vectors to binary codes based on sign bits, and a \RabitQuantizer\ that supports configurable bit-depths (1-9 bits) and rotation types (Fast/Matrix) for IVF\_RQ indexes. The change also integrates these quantizers into the IVF transformer pipeline, allowing users to leverage binary and reduced-bit quantization for more compact storage and faster search performance.

rust/lance-index/src/vector · high confidence

Introduce Bitmap and Bloom Filter scalar indices

Added Bitmap and Bloom Filter scalar index implementations to the Lance index library. The Bitmap index is designed for low-cardinality columns, storing a bitmap for each unique value to quickly identify matching rows. The Bloom Filter index provides a probabilistic data structure for efficient existence checks, helping to prune non-matching data during scans. Both indices integrate with the existing scalar index plugin framework, supporting training, querying, and caching.

rust/lance-index/src/scalar · high confidence

Introduce CacheCodec for FTS index entries and cross-column scoring

The inverted index module now supports persistent caching of posting lists and positions via a new \CacheCodec\ implementation, enabling efficient serialization and zero-copy deserialization of FTS index data for backends. Additionally, a new \combined\_fields\ scoring mode allows cross-column BM25F queries by treating target columns as a single virtual field, and a \cross\_column\ module enables compound FTS queries across multiple indexed columns by mapping document identities to a shared row-address domain.

rust/lance-index/src/scalar/inverted · high confidence

Introduce DataFilePart and external blob base resolution for staged writes

The dataset write path now supports caller-managed data file parts, allowing large writes to be staged and assembled later. A new DataFilePart struct serializes the identity and metadata of completed staging parts, which can be checkpointed and later assembled into a complete data file via DataFilePart::open\_all. This also introduces ExternalBaseResolver to map external blob URIs to specific storage bases, enabling multi-base blob storage and proper resolution of external blob paths during writes.

rust/lance/src/dataset · high confidence

Introduce FFI bindings for RecordBatchStream and new SpillStore for scratch storage

The lance-io crate now exposes a new \ffi\ module that wraps a \RecordBatchStream\ into an \FFI\_ArrowArrayStream\, enabling cross-language data exchange via the Arrow FFI interface. Additionally, a new \SpillStore\ trait and its \LocalSpillStore\ implementation have been added to provide reclaimable scratch storage on the local disk, allowing temporary data (such as index shuffle runs) to be written to disk and read back within the same process, with automatic cleanup and configurable disk capacity limits.

rust/lance-io/src · high confidence

Introduce IVF\_FLAT vector index implementation

Added the \IVF\_FLAT\ index type to the \lance-index\ crate, enabling exact vector search within IVF partitions. This change introduces the \FlatIndex\ as a sub-index implementation, along with \FlatFloatStorage\ for in-memory vector storage, \FlatPairScorer\ for exact distance calculations (supporting Float16, Float32, Float64, and binary vectors with L2, Cosine, Dot, and Hamming metrics), and \FlatTransformer\ for column normalization. The implementation includes optimized search paths using a binary heap for top-k results and supports distance range queries.

rust/lance-index/src/vector/flat · high confidence

Introduce JSONB scalar UDFs for DataFusion

Adds a new set of user-defined functions (UDFs) in \rust/lance-datafusion/src/udf/json.rs\ that enable querying JSONB data stored as LargeBinary arrays within DataFusion. These functions provide type-aware extraction capabilities, allowing users to retrieve JSON fields and array elements by key or index, and convert JSONB values into standard Arrow types such as strings, integers (Int64), and floats (Float64).

rust/lance-datafusion/src/udf · high confidence

Introduce LSM scanner for unified reads across base table, SSTables, and in-memory memtables

The \mem\_wal/scanner\ module now provides a new \LsmScanner\ that unifies queries across the base Lance table, persisted SSTable generations, and active/frozen in-memory memtables. It automatically handles primary-key deduplication to ensure the newest version of each row is returned, supports vector (KNN) and full-text search (BM25) across all tiers, and applies cross-generation block-lists to suppress stale reads from older sources. The scanner also supports point lookups, nested projections, and schema reconciliation for generations written under previous table schemas.

_rust/lance/src/dataset/mem\wal/scanner · high confidence

Introduce Lance file format versions 2.1 and 2.2

Added new file format versions 2.1 and 2.2, each with dedicated writers, readers, and compression strategies. Version 2.1 uses a dense u16 miniblock layout and supports standard compression codecs, while version 2.2 upgrades to a dense u32 layout, enables large miniblock chunks, and adds support for variable packed struct per-value compression, general block compression, and structural blob encoding. Both versions provide explicit and lazy writer creation with configurable compression parameters.

_rust/lance-file/src/versions/v2\_1, rust/lance-file/src/versions/v2\2 · high confidence

Introduce Lance v2.0 file format support

Added the v2.0 file format implementation, including the reader, writer, and composition logic in the \rust/lance-file/src/versions/v2\_0\ module. This enables users to read and write data files using the v2.0 grammar, featuring specific handling for field encoding strategies, column metadata decoding, and page metadata spilling to bound memory usage during writes.

_rust/lance-file/src/versions/v2\0 · high confidence

Introduce Lance-native in-memory HNSW vector index for MemWAL

Added a new HNSW graph implementation within the MemWAL subsystem that enables fast approximate nearest neighbor search on in-memory vector data. This change introduces a zero-copy vector store backed by Arrow batches, allowing the index to borrow vector data without duplicating memory, and supports configurable build parameters (such as graph levels, edge count, and construction beam width) to tune performance. The implementation includes specific search and build parameter structures, distance computation kernels for L2, Dot, and Cosine metrics, and metadata serialization to persist the graph structure within Lance's batch format.

_rust/lance/src/dataset/mem\wal/hnsw · high confidence

Introduce MemWAL MemTable scanner with vector and full-text search support

This change introduces the new MemTable scanner implementation for the MemWAL (Write-Ahead Log) system, enabling in-memory search capabilities directly on the active memtable. Users can now perform vector searches (including HNSW-based nearest neighbor queries with configurable distance metrics and bounds) and full-text searches (supporting term, phrase, fuzzy, and boolean queries with cross-column support) on data that has not yet been flushed to disk. The scanner also handles nested column projections and ensures read-your-writes consistency by including the mutable tail in search results by default.

_rust/lance/src/dataset/mem\wal/memtable/scanner · high confidence

Introduce MemWAL-based MemTable with lock-free batch storage and flush pipeline

The MemTable implementation now uses a new MemWAL architecture featuring a lock-free \BatchStore\ for concurrent reads and serialized appends, a dedicated flusher that writes data to persistent SSTables with optional cache warming, and a DataFusion-integrated scanner supporting full scans and index queries (BTree, HNSW, FTS) with MVCC visibility. This change replaces the previous in-memory storage mechanism with a persistent, WAL-backed design that improves concurrency and durability for dataset writes.

_rust/lance/src/dataset/mem\wal/memtable · high confidence

Introduce MemWAL: an in-memory write-ahead log with LSM-style tiering for Lance datasets

This change adds the MemWAL subsystem to the Lance dataset API, enabling high-throughput, low-latency writes by buffering data in an in-memory MemTable before flushing to persistent storage. Users can initialize MemWAL on a dataset using a builder API that supports sharding strategies (manual, unsharded, bucket, or identity) and configurable maintained indexes (BTree, HNSW, and FTS). The system introduces an LSM-like architecture with a fresh tier (active MemTable) and a frozen tier (SSTables), providing features such as read-your-writes consistency, schema evolution reconciliation, and configurable backpressure. This allows applications to perform rapid batch inserts and updates with immediate visibility, while maintaining durability and supporting vector and full-text search indexes on the buffered data.

_rust/lance/src/dataset/mem\wal · high confidence

Introduce PyTorch integration for Lance datasets

This change adds a new \lance.torch\ module that enables using Lance datasets directly within the PyTorch ecosystem. It introduces \LanceDataset\ and \SafeLanceDataset\ to wrap Lance data as PyTorch datasets, supporting automatic conversion of PyArrow types (including fixed-size lists and bfloat16) to tensors. The module also provides GPU-accelerated distance computations (L2, cosine, dot) via \torch.compile\, a PyTorch-native K-Means implementation, utilities for distributed and multiprocessing training, and benchmarking helpers for ground-truth calculation.

python/python/lance/torch · high confidence

This change adds a new RaBitQ (Randomized Binary Quantization) index implementation in the \lance-index\ crate, enabling multi-bit IVF\_RQ storage and search. The new module includes a builder for creating quantizers with configurable bit-widths (defaulting to 5 bits), a storage layer for persisting binary and extended codes, and a transformer for computing query factors. Search performance is significantly improved through dedicated SIMD kernels (AVX2/AVX-512) for distance table quantization, ex-code dot products, and lower-bound pruning, alongside a fast random rotation pipeline. The implementation supports both approximate and accurate search modes and includes validation to ensure metadata consistency.

rust/lance-index/src/vector/bq · high confidence

Introduce bfloat16 extension type and commit conflict error handling

The Python SDK now supports the bfloat16 data type via a new PyArrow extension type (BFloat16Array) with native NumPy conversion and Pandas integration, enabling efficient handling of 16-bit floating-point vectors. Additionally, a new CommitConflictError is exposed to allow users to programmatically detect and handle concurrent write conflicts with retry logic.

python/python/lance · high confidence

Introduce bfloat16 support and data generation utilities in Python SDK

The Python SDK now includes native support for the bfloat16 data type, exposing a \BFloat16\ class and a \bfloat16\_array\ function to create and manipulate bfloat16 Arrow arrays. Additionally, a new \datagen\ module has been added to the Python bindings, providing the \rand\_batches\ function to generate random data batches for a given schema, which is useful for testing and benchmarking.

python/src · high confidence

Introduce data overlay files and stable row ID format support

The \lance-table\ crate now supports data overlay files, which allow updating specific cell values within a fragment without rewriting the underlying data files, and introduces the v2 data storage format (stable row IDs). This includes new \DataOverlayFile\ and \OverlayCoverage\ structures for managing these overlays, along with logic to handle index staleness when overlays modify indexed fields. The manifest format is updated to track the data storage format version, and \IndexMetadata\ now includes file sizes and creation timestamps.

rust/lance-table · high confidence

Introduce distributed execution support for vector search and filtered reads

Added protobuf serialization for the \ANNIvfPartitionExec\ and \FilteredReadExec\ execution nodes, enabling these plans to be serialized and sent to remote workers for distributed execution. Additionally, introduced a \LanceFilterExec\ wrapper around DataFusion's \FilterExec\ to preserve the original logical expression, allowing filter predicates to be serialized to Substrait for remote planning. These changes provide the necessary infrastructure for distributing vector search and filtered read operations across a cluster.

rust/lance/src/io/exec · high confidence

Introduce exact v2.x file format readers and writers

The \lance-file\ crate now supports reading and writing the exact v2.x file formats (v2.0 through v2.3) with version-specific encodings and metadata handling. This change replaces the legacy v1 reader with a modular, version-dispatched architecture that validates file structure, handles blob v2 descriptors, and supports structural projections. Users can now explicitly target stable (v2.2) or next (v2.3) file versions for new data, while existing v1 files remain readable via a separate legacy path. The new format includes improved encoding mechanisms, exact version identity, and better compatibility testing across versions.

rust/lance-file/src · high confidence

Introduce experimental Lance v2.3 file format

Added a new v2.3 file format version that implements a distinct encoding and compression strategy. This version introduces a new compression strategy supporting miniblock, per-value, and block-level compressors (including bitpacking, RLE, and byte-stream split) and a field encoding strategy that composes primitive page encodings (sparse, constant, and dense) with structural encodings for blobs, maps, lists, and structs. The module provides writer functions (with and without explicit compression tuning) and reader validation logic, but emits a warning that this is an unstable format intended for experimentation only, as future compatibility is not guaranteed.

_rust/lance-file/src/versions/v2\3 · high confidence

Introduce flat B-tree index page implementation

Added a new \FlatIndex\ structure in the scalar B-tree module to represent a single index page as a sorted batch of value/row-id pairs. This implementation enables efficient on-demand filtering of row IDs using Arrow's vectorized operations, supporting optimized handling of \IsIn\ predicates and null tracking without materializing intermediate structures per page.

rust/lance-index/src/scalar/btree · high confidence

Introduce hierarchical, type-safe session caches for metadata and indices

The session layer now uses a new hierarchical caching system to organize dataset metadata and index data, preventing collisions between different datasets and indices. This change introduces \GlobalMetadataCache\ and \GlobalIndexCache\ as top-level namespaces, with sub-caches for specific datasets and indices. It also implements type-safe cache keys for manifest data, transactions, deletion files, row address masks, and index metadata, ensuring that cache entries are correctly scoped and versioned. This improves cache reliability and performance by avoiding stale or incorrect data reuse across different datasets and index types.

rust/lance/src/session · high confidence

The MemTable now uses dedicated in-memory index structures to accelerate point lookups and scans before data is flushed to disk. This includes a custom single-writer, lock-free-read skiplist (arena\_skiplist) optimized for the MemTable's append-only access pattern, a B-tree index (btree) for scalar fields with compact key encodings for integers and strings, a partition-structured full-text search index (fts) supporting BM25 scoring and complex queries, and an HNSW vector index (hnsw) with build parameters tuned for fast flush cycles. These components replace previous generic or less efficient indexing mechanisms within the MemWAL layer, providing faster read performance and better memory locality for active memtable operations.

_rust/lance/src/dataset/mem\wal/index · high confidence

Introduce io\_uring-based file reader for high-performance local I/O

Added a new io\_uring-based file reader implementation in \lance-io\ that supports both a dedicated background thread pool (\thread.rs\) and a thread-local mode (\current\_thread.rs\) for current-thread runtimes. This new reader provides a \file+uring://\ URI scheme, handles file caching, manages concurrent reads, and includes comprehensive tests for various read scenarios.

rust/lance-io/src/uring · high confidence

Introduce lance-arrow-scalar crate for efficient scalar comparisons

The rust/arrow-scalar crate (now named lance-arrow-scalar) provides an ArrowScalar type that wraps single-element Arrow arrays to enable O(1) comparison and hashing. By delegating to arrow\_row::OwnedRow, it ensures correct total ordering, proper NaN handling, and consistent null ordering, while also supporting serde serialization and conversion from primitive types.

rust/arrow-scalar · high confidence

Introduce lance-linalg crate with SIMD-accelerated distance metrics

The lance-linalg crate is introduced as an internal sub-crate containing native linear algebra algorithms for Lance. It provides distance metrics (L2, cosine, dot, hamming) for float types (f16, bf16, f32, f64) and integer types (u8, i8), featuring runtime-dispatched SIMD backends for x86\_64 (AVX2, AVX-512, AMX-FP16), aarch64 (NEON), and loongarch64 (LSX/LASX). The crate includes a build system that compiles C-based SIMD kernels and a \DeepSizeOf\ derive macro for Arrow-aware memory accounting.

rust/lance-linalg · high confidence

Introduce lance-tools CLI for inspecting Lance file metadata

A new \lance-tools\ command-line utility is added to the Rust workspace, providing a \meta\ subcommand that displays key metadata (version, row count, byte sizes, and schema) for Lance files. The tool supports both local filesystem paths and URI-based sources by leveraging the existing \ObjectStore\ registry and \ScanScheduler\ infrastructure.

rust/lance-tools · high confidence

Introduce memory allocation tracking for Python tests

Added a new \memtest\ utility that allows Python test suites to track memory allocations made by the Python interpreter and its libraries. The tool works by preloading a shared library (built in Rust) that intercepts memory calls, exposing statistics like peak usage and allocation counts via a Python API and CLI. It supports both Linux and macOS, enabling developers to detect memory leaks or regressions in their Python code.

memtest · high confidence

Introduce new physical encoding implementations for basic, binary, bitmap, and other data types

This change adds a suite of new physical encoding and decoding components in the \lance-encoding\ module, including \BasicPageScheduler\/\BasicEncoder\ for primitive fields, \BinaryPageScheduler\ for variable-length strings/binary data, \DenseBitmapScheduler\ for bitmaps, \BitpackedForNonNegScheduler\ for bit-packing, \DictionaryPageScheduler\ for dictionary-encoded arrays, \FixedSizeBinaryPageScheduler\ for fixed-size binaries, \FixedListScheduler\ for fixed-size lists, \FsstPageScheduler\ for FSST compression, and \PackedStructPageScheduler\ for packed structs. These new schedulers and decoders replace or supplement the previous encoding mechanisms, providing the underlying infrastructure for reading and writing these specific data layouts efficiently.

_rust/lance-encoding/src/array\encoding/physical · high confidence

Introduce non-blocking async scanner and blob file APIs in Java SDK

The Java SDK now exposes an \AsyncScanner\ API that executes scans on a dedicated Tokio runtime and returns results via \CompletableFuture\, allowing non-blocking I/O for large queries. Additionally, new JNI bindings for \BlobFile\ provide direct read capabilities (\read\, \readUpTo\, \readRange\), and the \Dataset\ API now includes methods to retrieve blob data by row IDs or indices. These changes are supported by a new dispatcher mechanism that safely bridges asynchronous Rust tasks back to the Java thread.

java/lance-jni · high confidence

Introduce optional geo module with spatial UDFs and bounding box utilities

The \rust/lance-geo\ crate now provides a new optional \geo\ feature that exposes spatial capabilities. When enabled, it registers DataFusion UDFs for geometric measurements (Area, Distance, Length) and spatial relationships (Contains, Intersects, etc.) via the \geodatafusion\ integration. It also introduces a new \bbox\ module containing a \BoundingBox\ struct and helper functions for computing and managing spatial bounds, which are conditionally compiled and exported only when the \geo\ feature is active.

rust/lance-geo · high confidence

Introduce pluggable cache backend architecture with URI configuration

The cache system now supports pluggable backends via a new \CacheBackend\ trait, allowing users to swap the underlying storage implementation. The module exposes a URI-based configuration system (\build\_from\_uri\) that parses strings like \moka://?capacity=1073741824\ into backend configurations, enabling consistent setup across Python, Java, and Rust bindings. Two built-in implementations are provided: \MokaCacheBackend\ for general use and \QuickCacheBackend\ for high-contention scenarios like session metadata caching. The change also introduces a versioned serialization envelope (\LCE1\) for cache entries, allowing persistent backends to store and retrieve data across restarts with backward-compatible decoding that treats invalid formats as cache misses.

rust/lance-core/src/cache · high confidence

Introduce protobuf schema guidelines and new wire-format definitions

The repository now includes a Protobuf Guidelines document (AGENTS.md) that establishes the change process, compatibility requirements, and schema design rules for protobuf schemas, including the requirement for PMC votes on persisted format changes. Additionally, new protobuf files have been added to define wire contracts and encoding specifications: ann.proto defines vector query parameters and execution plan serialization (including IVF sub-index and partition execution nodes); filtered\_read.proto defines serialization for filtered read options and plans; index.proto and index\_old.proto define vector index metadata and details; file.proto and file2.proto define file descriptors, schemas, and the v2 file format layout; encodings\_v2\_0.proto and encodings\_v2\_1.proto specify array and structural encodings for Lance file formats 2.0 and 2.1; and fragment\_metadata.proto defines data fragment, file, and overlay file metadata structures.

protos · high confidence

Introduce public datagen API and benchmarks for array generation

The \lance-datagen\ crate now exposes a public API for generating test data arrays, including support for step, fill, random, and null patterns across various primitive types (Int8/16/32/64, Float32/64) and binary/string variants. This change adds a new \generator\ module with traits and implementations for array generation, along with a comprehensive benchmark suite (\benches/array\_gen.rs\) to measure performance of these generators. Users can now programmatically generate synthetic datasets with controlled characteristics for testing and development purposes.

rust/lance-datagen · high confidence

Introduce structural encodings for Lance 2.1+ (blob, constant, dictionary, miniblock, sparse)

This change adds the core structural encoding implementations for the Lance 2.1+ format, introducing new files for blob, constant, dictionary, miniblock, and sparse encodings. Users can now read and write data using these new structural layouts, which include support for out-of-line blob storage, constant-value optimization, dictionary encoding with normalized null handling, configurable miniblock chunk sizes (via LANCE\_MINIBLOCK\_MAX\_VALUES), and automatic sparse structural page planning for nested data.

rust/lance-encoding/src/encodings/logical/primitive · high confidence

Introduce system indices for fragment reuse and MemWAL

The index library now supports two new system index types: FragmentReuse and MemWAL. FragmentReuse tracks stable partition mappings to enable deferred compaction and efficient row-remapping during data rewrites, while MemWAL provides a write-ahead log mechanism for index state. These are exposed as first-class index types with dedicated adapters, plugin registry entries, and progress monitoring, allowing the system to manage internal index metadata and write-ahead state transparently.

rust/lance-index/src · high confidence

Introduces Python type stubs for the lance package

This change adds comprehensive .pyi type stub files for the Python lance module, covering the main package, debug utilities, fragment handling (including DeletionFile, RowIdMeta, and RowIdSequence), optimization operations (Compaction, RewriteResult), schema definitions (LanceSchema, LanceField), tracing, and the Bitmap binding. These stubs provide static type information for the underlying Rust-backed Python API, enabling better IDE autocomplete, type checking, and documentation for users of the library.

python/python/lance/lance · high confidence

Introduces versioned protobuf schema for index cache serialization

Adds a new \cache.proto\ file that defines the serialization format for index cache entries, including headers for FTS posting lists (compressed and plain), positions, posting groups, B-tree indices, and IVF vector partitions. This schema establishes the structure for cache keys and data, supporting features like configurable posting block sizes, impact data, and various quantizer types (PQ, Flat, SQ) with specific distance metrics and storage codecs.

rust/lance-index/protos-cache · high confidence

Introduction of the internal lance-index-core crate

The \lance-index-core\ crate has been introduced as an internal library containing the core traits and types used to implement index plugins for Lance. This new module defines the foundational \Index\ trait, the \IndexType\ enumeration (covering scalar types like BTree, Bitmap, and MinHashLsh, as well as vector types like IVF-PQ), and the \BuiltinIndexType\ enum for scalar index configuration. It also provides the \MetricsCollector\ trait for coarse-grained ANN stage timing and I/O statistics, and introduces asynchronous batch row-ID remapping utilities to handle row translation during index loading. This crate is explicitly marked as internal and not intended for external usage.

(repo-wide) · high confidence

Introduction of the lance-arrow internal sub-crate

The \lance-arrow\ crate has been introduced as an internal sub-crate containing Apache Arrow extensions for Lance. It provides support for bfloat16 vectors (including zero-copy array construction and round-tripping), JSON/JSONB read/write capabilities, and blob v2 schema handling. The crate also includes utilities for deep-copying Arrow arrays and RecordBatches, zero-copy IPC stream serialization, and memory accounting for Arrow buffers.

rust/lance-arrow · high confidence

Introduction of the lance-namespace Rust crate with core namespace APIs and compatibility shims

The \rust/lance-namespace\ crate has been added to provide the core Rust client and trait definitions for Lance namespaces, including the \LanceNamespace\ interface for managing namespaces and tables, schema conversion utilities between Arrow and JSON representations, and a comprehensive set of fine-grained error codes. To ensure backward compatibility with older SDK builds (Java and Python) that send scalar values for merge-insert keys, a deserialization shim in \compat.rs\ automatically promotes single-column string keys to the expected list format, while the error module standardizes error handling across all Lance implementations.

rust/lance-namespace · high confidence

Java SDK API overhaul with new core classes and session support

The Java SDK has been significantly refactored to align with the Rust engine, introducing a comprehensive set of new core classes including Dataset, CommitBuilder, Fragment, Branch, and BlobFile. This update adds support for session management, cache backends, and branch/tag metadata, while deprecating legacy fragment operations in favor of the new Transaction-based CommitBuilder. Users will now interact with a more modern, consistent API that exposes detailed dataset metadata, index build progress, and native blob handling capabilities.

java/src/main/java/org/lance · high confidence

Java bindings for MemWAL write, scan, and search operations

This change introduces a new set of Java classes in the \org.lance.memwal\ package that expose the MemWAL (Memory Write-Ahead Log) functionality to Java applications. Users can now initialize MemWAL on a dataset with configurable sharding (bucket, identity, or unsharded) and index maintenance via \InitializeMemWalParams\. The API provides \ShardWriter\ for inserting and deleting rows with configurable backpressure and flush intervals, and \LsmScanner\ for scanning data across the base table, SSTables, and active MemTable with SQL filtering and projection. Additionally, \LsmPointLookupPlanner\ and \LsmVectorSearchPlanner\ enable efficient primary-key lookups and IVF-PQ vector KNN searches across all MemWAL levels, with support for pre-filtering and stale-row suppression. The bindings also expose detailed statistics for MemTables and writers, and handle native resource management through JNI.

java/src/main/java/org/lance/memwal · high confidence

Java bindings for Merge Insert operations

Added Java classes (MergeInsertParams, MergeInsertResult, MergeInsertStats) to support merge insert operations, including configuration for matched/unmatched row handling, write modes, and statistics reporting.

java/src/main/java/org/lance/merge · high confidence

Java cleanup API: new policy, explanation, and stats models

The Java cleanup module now exposes a structured API for configuring and inspecting dataset cleanup operations. Users can build cleanup policies via a builder that supports filtering by timestamp or version, targeting specific dataset versions, cleaning referenced branches, deleting unverified files, enforcing a delete rate limit, and erroring on tagged old versions. The API also returns a CleanupExplanation containing RemovalStats (bytes removed, counts of data/transaction/index/deletion files removed, and failed deletes) and a list of candidate files with their kind, size, and verification status, along with warnings and branch references.

java/src/main/java/org/lance/cleanup · high confidence

Java compaction API with source limits, exclusion, and binary copy modes

The Java SDK now exposes a distributed compaction entry point in the \org.lance.compaction\ package, providing \planCompaction\ and \commitCompaction\ methods that wrap the native implementation. Users can now configure compaction with new limits (\maxSourceRows\, \maxSourceBytes\, \maxSourceFragments\) to bound the scope of planning, exclude specific fragments via \excludedFragmentIds\, and choose a \CompactionMode\ (reencode, try\_binary\_copy, or force\_binary\_copy) to control how data is rewritten. The API also supports configuring binary copy batch sizes, materializing deletions, and specifying a data storage version, while ensuring thread safety by acquiring a dataset read lock during native calls.

java/src/main/java/org/lance/compaction · high confidence

Java index metadata and configuration model

The \org.lance.index\ package now provides the core Java data models and configuration builders for index operations. This includes \Index\ and \IndexDescription\ for representing index metadata (such as segment sizes, creation times, and covering fields), \IndexType\ and \DistanceType\ enums to define supported index algorithms and distance metrics, and \IndexOptions\/\IndexParams\ builders for configuring index creation. Additionally, \IndexBuildProgress\ exposes callbacks to track build stages, and \OptimizeOptions\ allows users to configure index optimization strategies like merging and retraining.

java/src/main/java/org/lance/index · high confidence

Java namespace client now supports dynamic per-request context and mTLS certificate reloading

The Java namespace client (DirectoryNamespace and RestNamespace) now accepts a DynamicContextProvider, allowing applications to inject per-request context such as rotating authentication headers before each operation. Additionally, the RestNamespace client automatically reloads rotated mTLS certificates based on the new tls.reload\_interval\_seconds configuration property, ensuring secure connections persist through certificate rotations without requiring a restart.

java/src/main/java/org/lance/namespace · high confidence

Java schema model supports unenforced keys and column alterations

The Java schema package now includes \ColumnAlteration\ for modifying dataset columns (renaming, nullability, type casting) and extends \LanceField\ and \LanceSchema\ to support unenforced primary and clustering keys with explicit ordering positions. This allows users to define composite key ordering and clustering strategies in Java, while \LanceField\ also handles \FixedSizeList\ logical types during Arrow schema conversion.

java/src/main/java/org/lance/schema · high confidence

Java transaction operation models added

Added Java classes for dataset transaction operations (Append, Clone, CreateIndex, DataOverlay, DataReplacement, Delete, KeyExistenceFilter, Merge, Overwrite, Project, ReserveFragments, Restore, Rewrite, SchemaOperation) to align with the Rust transaction model, enabling Java clients to construct and manage dataset mutations.

java/src/main/java/org/lance/operation · high confidence

New CI benchmark infrastructure for tracking IO and memory performance

A new benchmark suite has been added to the Python package to monitor performance regressions in CI. This includes a custom pytest fixture that measures IO statistics (IOPS, bytes read/written) and memory usage (peak bytes, total allocations) during dataset operations, with optional support for memory tracking via the \lance-memtest\ library. The suite provides utilities for generating test datasets (including TPC-H and basic 10M row sets), resolving dataset URIs for local or Google Cloud environments, and managing data overlay benchmarks. It also includes helpers for clearing OS file caches to ensure consistent benchmark conditions.

_python/python/ci\benchmarks · high confidence

New CI tooling for release automation, security, and governance

The CI directory now includes a comprehensive suite of scripts and checks to support a structured release process and stricter quality gates. Release management is handled by new shell scripts (approve\_rc.sh, create\_rc.sh, create\_release\_branch.sh, publish\_beta.sh) that automate the creation of release branches, release candidates (RCs), and stable promotions, including automatic version bumping and release note generation via generate\_release\_notes.py. Security is enhanced by check\_dependency\_age.py, which enforces a 48-hour minimum age for crates.io dependencies to mitigate supply-chain risks, with an allowlist in dependency-age-allowlist.toml. Governance is supported by format\_vote\_gate.py, which enforces PMC voting requirements for format-specification changes, and check\_breaking\_changes.py, which validates that minor versions are bumped when breaking changes are detected. Additional utilities include check\_proto\_comments.py for enforcing multi-line comment styles in .proto files, coverage.py for Rust code coverage analysis, and new\_contributors.py for tracking contributor statistics.

ci · high confidence

New CLI tools for FM index management and benchmarking

Added three new command-line utilities in the Rust binary directory: \fm\_index\_tool\ for listing, creating, and dropping FM (Fuzzy Matching) scalar indexes on datasets; \fm\_contains\_bench\ for benchmarking FM index query performance with configurable thread counts, caching, and output to JSON/CSV; and \lq\ for basic dataset inspection, querying, and vector index creation. These tools provide dedicated interfaces for managing and evaluating FM indexes, which are distinct from the existing vector index workflows.

rust/lance/src/bin · high confidence

New DataFusion integration for querying Lance namespaces via SQL

The \lance-namespace-datafusion\ crate introduces a bridge that allows Lance namespaces to be queried as native DataFusion catalogs, schemas, and tables. It exposes a \SessionBuilder\ to construct a DataFusion \SessionContext\ with dynamic \CatalogProvider\ and \SchemaProvider\ implementations backed by a \LanceNamespace\. This enables read-only SQL access to Lance datasets, mapping top-level namespaces to catalogs, child namespaces to schemas, and loading tables on-demand with caching.

rust/lance-namespace-datafusion · high confidence

New DatasetDelta API to inspect row-level changes between dataset versions

Added DatasetDelta and DatasetDeltaBuilder classes in the Java API, enabling users to compute and stream differences between two versions of a dataset. Users can specify a version range or compare against a specific version to retrieve inserted and updated rows as Arrow streams, list transactions, and obtain deleted row IDs. The implementation holds a read lock on the underlying dataset during native calls to ensure thread safety.

java/src/main/java/org/lance/delta · high confidence

New FragmentSession API for efficient repeated fragment reads

Users can now use the new FragmentSession API to perform repeated reads on a fragment without the overhead of opening a new reader each time. This session-based approach maintains internal state, such as sorted deletion vectors, allowing for more efficient row lookups and automatic handling of deleted rows across multiple operations. The implementation also ensures that JSON columns are correctly converted from the internal Lance format to standard Arrow JSON for user-facing outputs.

rust/lance/src/dataset/fragment · high confidence

New HD-Vila dataset benchmark with data generation script

Added a new benchmark entry for the HD-Vila-100M dataset, including a README with setup instructions and a \datagen.py\ script. The script enables users to generate the dataset in Lance format by downloading videos from YouTube, cutting specific clips, and processing them using Ray and PyArrow.

benchmarks/hd-vila · high confidence

New HNSW index implementation with online builder support

The HNSW vector index implementation has been replaced with a new version in the \lance-index\ crate. This update introduces a new \HNSWBuilder\ for offline graph construction and an \OnlineHnswBuilder\ that supports concurrent search during index building, which is particularly useful for in-memory MemTable indexes. The new implementation includes deterministic graph construction via a fixed random seed, configurable build parameters (such as \m\, \ef\_construction\, and \max\_level\), and support for both float and binary (Hamming distance) vectors. It also integrates with the existing IVF framework, allowing HNSW to be used as a sub-index within IVF partitions.

rust/lance-index/src/vector/hnsw · high confidence

New IVF index builder, shuffler, and transformer components

The IVF index module now includes dedicated builder, shuffler, storage, and transformer components. The builder introduces \IvfBuildParams\ with a new \target\_partition\_size\ option (replacing the deprecated \num\_partitions\), support for streaming k-means training via \streaming\_sample\_rate\ and \streaming\_coreset\_rate\, and the ability to load precomputed partition mappings or shuffle buffers. The shuffler handles disk-based partitioning of data streams, while the transformer computes partition IDs and distances for vectors. Storage is managed via the \IvfModel\ struct, which persists centroids, partition offsets, and lengths to Lance files.

rust/lance-index/src/vector/ivf · high confidence

New IVF index building and search implementation (v2)

The IVF vector index implementation in \rust/lance/src/index/vector/ivf\ has been replaced with a new v2 architecture. This change introduces a new builder (\builder.rs\) and I/O layer (\io.rs\) for constructing index partitions, a new v2 index definition (\v2.rs\) that supports streaming search with configurable batch sizes, and a new serialization format (\partition\_serde.rs\) for caching IVF partitions. Users will see this as the underlying engine for IVF index creation and search, enabling features like distributed segment builds, parallel partition search, and improved memory management through chunked global top-k scoring.

rust/lance/src/index/vector/ivf · high confidence

New JSON serialization support for Arrow Schema types

Added a new \json.rs\ module that enables serializing and deserializing Apache Arrow \DataType\ and \Field\ structures to and from JSON. This feature supports a wide range of data types including primitives, lists, structs, and fixed-size binary/list types, allowing users to easily convert schema definitions into a portable JSON format for storage or interchange.

rust/lance/src/arrow · high confidence

New Java API for vector index configuration and training

The Java API now exposes a comprehensive set of builder classes for configuring vector index parameters, including HnswBuildParams, IvfBuildParams, PQBuildParams, RQBuildParams, and SQBuildParams. These are unified in VectorIndexParams, which provides factory methods to create IVF-based index configurations (IVF Flat, IVF PQ, IVF RQ, and IVF HNSW with PQ/SQ) and enforces validation rules such as mutual exclusivity of quantizers. Additionally, the new VectorTrainer utility allows users to pre-train IVF centroids and PQ codebooks with explicit distance type support, ensuring the training geometry matches the index build geometry to prevent silent recall degradation.

java/src/main/java/org/lance/index/vector · high confidence

New Java IPC layer with async scanning, full-text search, and vector search controls

The \org.lance.ipc\ package introduces a new Java interface to the native (Rust) engine, providing both synchronous (\LanceScanner\) and non-blocking asynchronous (\AsyncScanner\) data scanning via \CompletableFuture\. This release adds support for full-text search through the \FullTextQuery\ API, vector search tuning with the \ApproxMode\ enum (FAST, NORMAL, ACCURATE) for RQ-quantized indexes, and query optimization via \MaterializationStyle\ (early/late column fetching) and \ColumnOrdering\. Additionally, \ScanStats\ exposes detailed performance metrics, including per-query index cache hit/miss counts, to help users monitor and tune scan efficiency.

java/src/main/java/org/lance/ipc · high confidence

New Java file I/O API with configurable blob reading and write options

The \org.lance.file\ package now provides dedicated \LanceFileReader\ and \LanceFileWriter\ classes for direct file-level operations. Users can control how blob-encoded columns are returned during reads via the new \BlobReadMode\ enum (materialized content or position/size descriptors) through \FileReadOptions\. The writer supports configuring buffering and page sizes via \FileWriteOptions\, allows specifying a data storage version, and enables attaching custom schema metadata that is flushed to the file footer on close.

java/src/main/java/org/lance/file · high confidence

New Java fragment metadata and result classes

Added new Java classes in the \org.lance.fragment\ package to support fragment-level metadata and update operations. \DataFile\ and \DeletionFile\ (with \DeletionFileType\) expose file paths, versions, and deletion details. \FragmentUpdateResult\ and \FragmentMergeResult\ provide structured outputs for column updates and merges, including row offset handling via JNI. \RowIdMeta\ and \VersionMeta\ wrap serialized Rust metadata for row IDs and version sequences, enabling stable row identification and version tracking in Java.

java/src/main/java/org/lance/fragment · high confidence

New LSM scanner execution nodes for MemTable scans, deduplication, and indexed searches

The MemTable scanner now includes dedicated DataFusion execution nodes to handle complex read patterns within the active MemTable. A new deduplication scan ensures that primary-key updates are resolved correctly by emitting only the newest version of each key, preventing stale rows from leaking through filters. Indexed searches are now supported directly on the MemTable via BTree, Full-Text Search (FTS), and Vector (HNSW) execution nodes, allowing scalar, text, and vector queries to leverage in-memory indexes with MVCC visibility. For vector searches where no HNSW index is available, a brute-force KNN executor provides exact distance calculations. All these nodes enforce visibility constraints, ensuring that only readable batches are scanned.

_rust/lance/src/dataset/mem\wal/memtable/scanner/exec · high confidence

New LSM scanner execution nodes for point lookups and cross-generation deduplication

The mem-wal scanner now includes a suite of new DataFusion execution nodes to optimize point lookups and handle multi-generation data. BloomFilterGuardExec skips generations that definitely do not contain a requested primary key, while CoalesceFirstExec short-circuits evaluation to return the newest matching row immediately. PkBlockFilterExec removes superseded rows by checking primary-key membership against newer generations, and FirstByPkExec deduplicates results (such as those from cross-column full-text search) by keeping only the first occurrence per primary key. MemtableGenTagExec annotates rows with their generation number for ordering, SchemaRelabelExec aligns storage and logical schemas, and ReconcileExec applies schema plans to batches.

_rust/lance/src/dataset/mem\wal/scanner/exec · high confidence

New Product Quantization implementation with 4-bit support and SIMD optimizations

The Product Quantization (PQ) vector index module has been rewritten to support 4-bit quantization alongside the existing 8-bit mode, enabling significantly higher compression ratios for vector data. The new implementation includes a dedicated builder for training PQ codebooks, a storage backend that persists PQ codes and row IDs, and a transformer for converting vector columns into PQ codes. Performance is improved through SIMD-optimized distance calculations, including AVX512-VBMI byte-plane lookups for 4-bit codes and pre-transposed distance tables to avoid runtime transposition overhead. The module also introduces a pairwise scorer for symmetric PQ code-to-code scoring, supporting L2, cosine, and dot distance metrics.

rust/lance-index/src/vector/pq · high confidence

New Python API for distributed index building and model checkpointing

The \lance.indices\ package introduces a new \IndicesBuilder\ class that allows advanced users to construct vector indices (IVF, PQ, IVF\_PQ, IVF\_SQ) in discrete, checkpointable steps rather than as a single monolithic operation. This enables distributed and segmented index builds on large datasets. The package also exposes \IvfModel\ and \PqModel\ classes with \save\ and \load\ methods that support \storage\_options\, allowing users to persist and resume training progress to cloud storage backends.

python/python/lance/indices · high confidence

New Python bindings for dataset internals and operations

This change introduces several new Python-level components in the \python/src/dataset\ module. It exposes a \BlobFile\ API (\LanceBlobFile\) allowing users to read blob data via methods like \readall\, \read\_range\, and \read\_ranges\. It adds Python bindings for dataset cleanup operations, including \CleanupStats\, \CleanupCandidateFile\, and \CleanupExplanation\ classes to inspect and report on cleanup results. A \PyCommitLock\ implementation is added to support custom Python-based commit handlers. Additionally, it exposes \IOStats\ for tracking read/write I/O metrics, \DataStatistics\ and \FieldStatistics\ for reporting data stats, and comprehensive Python bindings for the compaction optimization system, including \CompactionPlan\, \CompactionTask\, and \CompactionMetrics\.

python/src/dataset · high confidence

New Python type stubs for index training, transformation, and metadata APIs

The \lance.indices\ module now exposes Python type stubs (\\_\init\\_.pyi\) that define the public API for vector index operations. Users can now rely on static type checking for standalone training functions (\train\_ivf\_model\, \train\_pq\_model\), vector transformation (\transform\_vectors\), and residual quantization model building (\build\_rq\_model\). The stubs also formalize the structure of index metadata classes (\IndexConfig\, \IndexSegment\, \IndexDescription\, \IndexSegmentDescription\), including fields for segment UUIDs, fragment IDs, sizes, and covering fields, ensuring consistent type hints for index inspection and management.

python/python/lance/lance/indices · high confidence

New Rust examples for vector search, full-text search, and dataset I/O

Added several new Rust example programs in the \rust/examples\ directory to demonstrate core Lance capabilities. The \full\_text\_search\ example shows how to create and query an inverted index for text data. The \hnsw\ and \ivf\_hnsw\ examples provide benchmarks and usage patterns for HNSW and IVF-HNSW vector indices, including building indexes and performing nearest-neighbor searches. The \llm\_dataset\_creation\ example demonstrates downloading and tokenizing text data from Hugging Face for LLM training. Finally, \write\_read\_ds\ illustrates basic dataset creation, reading, and cleanup operations.

rust/examples · high confidence

New SIFT/GIST-1M benchmark suite with LanceDB integration

The benchmarks/sift directory now provides a complete, reproducible benchmarking workflow for the SIFT-1M and GIST-1M datasets using the LanceDB API. This includes Python scripts for data generation (datagen.py), ground truth creation (gt.py), index building (index.py), and performance metric calculation (metrics.py), alongside a Jupyter notebook for result analysis. The suite supports multiple vector data types (f32, f16, bf16) and index configurations (IVF-PQ, DiskANN), with pre-computed statistics files (lance\_sift1m\_stats.csv, lance\_gist1m\_stats.csv) demonstrating recall and latency trade-offs across various partition and probe settings.

benchmarks/sift · high confidence

New Scalar Quantization (SQ) index implementation

This change introduces the core components for a new Scalar Quantization vector index within the \lance-index\ crate. It adds an SQ builder (\builder.rs\) to configure quantization parameters, a transformer (\transform.rs\) to convert input vectors into 8-bit codes, a storage layer (\storage.rs\) to manage the quantized data chunks, and a pairwise scorer (\pairwise.rs\) that performs exact integer-based distance calculations for L2, Cosine, and Dot metrics. This provides a new, efficient indexing option for users requiring lower memory footprint and faster search speeds through scalar quantization.

rust/lance-index/src/vector/sq · high confidence

New \`lance-test-macros\` crate for test-scoped tracing integration

A new \lance-test-macros\ crate has been added, providing a \\#\[test\]\ attribute macro that wraps existing test functions to automatically configure the \tracing\ library. When the \LANCE\_TRACING\ environment variable is set (e.g., to \debug\), tests wrapped with this macro initialize a tracing subscriber with a Chrome layer, generating \.json\ trace files compatible with Chrome DevTools or Perfetto. This allows developers to easily capture detailed performance and execution traces for individual tests without manual setup.

rust/lance-test-macros · high confidence

New arrow-stats crate for computing column statistics

Added the \lance-arrow-stats\ crate, which provides a \StatisticsAccumulator\ to compute min, max, null count, NaN count, and buffer memory usage for Apache Arrow arrays. This new capability supports numeric, temporal, boolean, string, binary, and list types, enabling more efficient predicate pushdown and query planning in Lance's columnar storage layer by tracking page-level statistics.

rust/arrow-stats · high confidence

New benchmark examples for ACORN-1 HNSW traversal and hierarchical k-means quality

Added three new example programs in the lance-index crate to evaluate specific index features. \acorn\_bench.rs\ and \acorn\_bench\_sift.rs\ benchmark the new ACORN-1 mask-aware HNSW traversal against the standard basic traversal and flat scans, measuring latency and recall under various filter selectivities and mask shapes (using synthetic data and the SIFT1M dataset respectively). \kmeans\_quality.rs\ measures the training time and clustering quality (WCSS, partition balance) of the new proportional hierarchical k-means algorithm for IVF training on large \.fbin\ datasets.

rust/lance-index/examples · high confidence

New benchmarking tool and credential vending infrastructure for Directory Namespace

This release introduces a copy-on-write \\_\_manifest\ commit benchmark (\manifest\_bench\) to measure Directory Namespace throughput under continuous and concurrent workloads, alongside a new \ConnectBuilder\ API for configuring namespace connections. It also adds a comprehensive credential vending system that automatically generates scoped, temporary credentials for AWS (STS AssumeRole), Azure (SAS tokens), and GCP, with built-in caching to reduce vendor call frequency and dynamic context providers for per-request header injection in the REST namespace.

rust/lance-namespace-impls · high confidence

New binary copy compaction and index remapping infrastructure

This change introduces a new binary copy mechanism for compaction in \rust/lance/src/dataset/optimize/binary\_copy.rs\, which merges small Lance files into larger ones by performing page-level binary copies to preserve stable row IDs and reduce I/O overhead. It also adds a new \remapping.rs\ module that provides the \IndexRemapper\ trait and utilities for remapping row addresses and index metadata during compaction, ensuring that indices remain valid when fragment row IDs change.

rust/lance/src/dataset/optimize · high confidence

New core utility modules for row addressing, rate limiting, and retry logic

The \rust/lance-core/src/utils\ directory now includes several new foundational modules that support internal operations. \address.rs\ introduces the \RowAddress\ type, which encodes fragment IDs and row offsets into a single 64-bit value to replace the previous \row\_id\/\row\_addr\ distinction. \aimd.rs\ provides an Additive Increase / Multiplicative Decrease (AIMD) rate controller for dynamically adjusting request rates, including validation to reject non-finite configuration values. \backoff.rs\ adds configurable exponential backoff and \SlotBackoff\ strategies to help spread out concurrent retries. Additional utilities include \assume.rs\ for release-checked invariants, \bit.rs\ for bit manipulation and alignment, \blob.rs\ for generating obfuscated blob sidecar paths, and \futures.rs\ for sharing streams between consumers.

rust/lance-core/src/utils · high confidence

New distributed index merging infrastructure

The distributed vector indexing subsystem now includes a new \index\_merger\ module and shared \partition\_merger\ helpers. This introduces the core logic for merging distributed index segments, including support for detecting various IVF index types (Flat, PQ, SQ, RQ, and HNSW variants) and writing unified metadata. It also adds strict and tolerant equality checks for vector data to ensure consistency during the merge process.

rust/lance-index/src/vector/distributed · high confidence

New encoding utilities for data accumulation and byte-packed integer encoding

The \lance-encoding\ crate now includes new utility modules for handling data buffering and efficient integer encoding. An \AccumulationQueue\ has been added to buffer Arrow arrays until a specified byte threshold is reached, allowing for optimized flushing of column data during encoding. Additionally, a \BytepackedIntegerEncoder\ provides a simple, fast way to encode integers by automatically selecting the smallest suitable byte width (u8, u16, u32, or u64) based on the maximum value, rejecting values that exceed the selected width to prevent overflow.

rust/lance-encoding/src/utils · high confidence

New lance-datafusion crate with DataFusion integration

A new \lance-datafusion\ crate has been introduced to provide a bridge between Lance and Apache DataFusion. This includes a build script for generating protobuf bindings, an \Aggregate\ struct for handling group-by and aggregate expressions, a \BatchReaderChunker\ for streaming record batches, and a \DataFrameExt\ trait that adds a \group\_by\_stream\ method to DataFusion DataFrames. The crate also implements core execution utilities like \OneShotExec\ for wrapping streams into execution plans, expression coercion logic in \expr.rs\ and \logical\_expr.rs\ to handle type conversions and filter simplification, and a \ProjectionBuilder\ to manage column selection and system column handling during query planning.

rust/lance-datafusion/src · high confidence

New lance-testing crate for internal test utilities

A new \lance-testing\ crate has been added to the repository, providing internal-only utilities for unit tests and benchmarks. It includes a \datagen\ module with generators for creating test data (such as incrementing integers and random vectors), a \progress\ module offering a macro to define in-memory progress recorders for testing distributed operations, and a \pprof\ module (Linux-only) that integrates with Criterion to generate flamegraph profiles during benchmarks.

rust/lance-testing · high confidence

New logical array encoders for binary, blob, list, primitive, and struct types

The logical encoding layer in \rust/lance-encoding/src/array\_encoding/logical\ now includes dedicated schedulers and decoders for Binary, Blob, List, Primitive, and Struct data types. These new components handle the scheduling and decoding of their respective data layouts, enabling more efficient handling of variable-width data, large binary blobs, nested lists, and structured records within the Lance encoding format.

_rust/lance-encoding/src/array\encoding/logical · high confidence

New logical encoding layer for complex data types

The logical encoding module now includes dedicated structural encoders, schedulers, and decoders for Blob, FixedSizeList, List, Map, and Struct types. This adds support for storing large binary data in external buffers (Blob), correctly handling nested structures like lists of structs (FixedSizeList), and managing repetition/definition levels for variable-length lists and maps. These components form the foundation for encoding complex nested data in the Lance format.

rust/lance-encoding/src/encodings/logical · high confidence

New modular write API with staged transactions and retry support

The write operations (insert, delete, update, merge insert) now use a builder pattern that separates data preparation from the final commit. Users can stage changes using methods like \execute\_uncommitted\ to obtain a \Transaction\, which is then committed via the new \CommitBuilder\. This architecture introduces configurable retry logic for concurrent write conflicts, allowing users to set \conflict\_retries\ and \retry\_timeout\ on delete and update operations, and exposes a \write\_progress\ callback on inserts to track throughput. The \CommitBuilder\ also centralizes configuration for storage formats, object stores, and session caching.

rust/lance/src/dataset/write · high confidence

New physical encoding implementations for Lance 2.1

The \rust/lance-encoding/src/encodings/physical\ directory now contains the core physical encoding implementations for the Lance 2.1 format. This includes \binary.rs\ for variable-width data, \bitpacking.rs\ for fixed-width bitpacking, \block.rs\ for general block compression (LZ4, Zstd), \byte\_stream\_split.rs\ for floating-point optimization, \constant.rs\ for constant/null handling, \fsst.rs\ for string compression, \general.rs\ for wrapping inner compressors with general compression, \packed.rs\ for struct packing, and \rle.rs\ for run-length encoding. These files provide the leaf-level encoders and decompressors that the logical layer uses to store data.

rust/lance-encoding/src/encodings/physical · high confidence

New quickstart and YouTube Q&A bot notebooks added

Added a new quickstart notebook demonstrating how to create and write Lance datasets using PyArrow and Pandas, including conversion from Parquet. Also added a new notebook for building a question-and-answer bot that searches YouTube transcripts using natural language, leveraging OpenAI embeddings and the Hugging Face datasets library.

notebooks · high confidence

New utility classes for JSON handling and range definitions

Added three new utility classes to the org.lance.util package: JsonFields provides helpers to create Arrow fields annotated with the 'arrow.json' extension for seamless JSON text handling; JsonUtils offers static methods to serialize and deserialize JSON using Jackson; and Range defines a simple immutable value object for representing integer ranges with inclusive start and exclusive end boundaries.

java/src/main/java/org/lance/util · high confidence

New utility modules for async primitives, time mocking, and test infrastructure

The \rust/lance/src/utils\ directory now includes three new modules: \future.rs\ introduces a \SharedPrerequisite\ struct that allows spawning an asynchronous background task and sharing its result (or error) across multiple threads, with synchronous access once the task completes; \temporal.rs\ provides a \utc\_now()\ function that abstracts the current system time, enabling unit tests to mock time via \mock\_instant\ while production code uses the real system clock; and \test.rs\ adds a \TestDatasetGenerator\ capable of creating datasets with randomized, 'hostile' column layouts (splitting fields across files, shuffling field IDs, and varying field orders) to improve test coverage for dataset operations under diverse storage configurations.

rust/lance/src/utils · high confidence

New v3 vector index architecture with shuffler and sub-index traits

The v3 vector index implementation introduces a new internal architecture for building and searching IVF-based indexes. This location adds the core shuffler component (shuffle\_bench.rs, shuffler.rs) which handles partitioning data streams into IVF partitions, and the sub-index trait definitions (subindex.rs) that standardize how sub-indexes like Flat and HNSW are built, searched, and remapped within the v3 format. Users benefit from the underlying performance and structural improvements of the v3 index format, including configurable CPU thread usage and optimized shuffle operations, as this code forms the foundational I/O and partitioning layer for the new index version.

rust/lance-index/src/vector/v3 · high confidence

New vector index details serialization and bounded partition stream for IVF builds

The vector index module now includes a new \details\ module that serializes and deserializes \VectorIndexDetails\ (including compression types like PQ, SQ, and RQ, and runtime hints) for index metadata, and a new \bounded\_partition\_stream\ module that manages memory-bounded, ordered execution of IVF partition builds. These changes support richer index introspection and more controlled resource usage during index construction.

rust/lance/src/index/vector · high confidence

New vector index performance benchmark report

Added a new benchmark suite in the \benchmarks/full\_report\ directory to evaluate vector index performance. This includes a Python library (\\_lib.py\) for generating and processing vector data (specifically using the NYT dataset with TF-IDF and random projection), and a Jupyter notebook (\report.ipnb\) that runs recall tests against Lance IVF\_PQ indexes. The benchmark measures recall rates across different \nprobes\ and \refine\_factor\ settings for both in-sample and out-of-sample queries, visualizing the results as heatmaps.

_benchmarks/full\report · high confidence

Optimized execution paths for delete-only and in-place merge insert operations

The merge insert engine now uses specialized execution nodes to improve performance for specific operation types. A new delete-only path (\DeleteOnlyMergeInsertExec\) handles \WhenMatched::Delete\ scenarios by reading only row addresses and action columns, skipping the write step entirely to efficiently mark existing rows as deleted. Additionally, an in-place update path (\InPlaceMergeInsertExec\) allows narrow updates to existing fragments by patching only the source columns into existing data files and tombstoning old versions, avoiding the overhead of rewriting entire rows. These changes optimize bulk deletes and partial updates without creating new fragments.

_rust/lance/src/dataset/write/merge\insert/exec · high confidence

Restructure lance crate into modular sub-modules

The \rust/lance/src\ directory has been refactored from a flat structure into distinct sub-modules (\arrow\, \blob\, \datafusion\, \dataset\, \index\, \io\, \session\, \table\, \utils\). This change introduces new public APIs for blob v2 storage (including \BlobFieldOptions\ and dedicated writers), exposes DataFusion integration via \LanceTableProvider\, and adds object store metrics support (documented in \metrics.md\). The \lib.rs\ entry point now re-exports these modules and their specific capabilities, such as the new blob field helpers and the DataFusion table provider, while maintaining backward compatibility for core dataset operations.

rust/lance/src · high confidence

Segmented scalar index merge support

Scalar index implementations (Bitmap, BloomFilter, BTree, FM-Index, Inverted, LabelList, MinHash LSH, NGram, RTree, and ZoneMap) now support merging multiple index segments into a single consolidated segment. This enables distributed index builds and compaction workflows to combine partial results without rebuilding from raw data, improving performance and scalability for large datasets.

rust/lance/src/index/scalar · high confidence

Support for dynamic, auto-refreshing storage credentials

Lance now supports dynamic storage options and credentials that are fetched at runtime and automatically refreshed before expiration. This is enabled by the new \StorageOptionsProvider\ trait and \StorageOptionsAccessor\, which allow integration with external sources like LanceNamespace to retrieve temporary credentials (e.g., AWS, Azure, GCP) or other configuration. The \NamespaceCredentialsProvider\ and \DynamicOpenDalStore\ use these options to build and cache object stores, ensuring that short-lived credentials are renewed seamlessly without requiring application restarts or manual intervention.

_rust/lance-io/src/object\store · high confidence

Unified object store providers with OpenDAL and dynamic credentials

The object store providers in \rust/lance-io/src/object\_store/providers\ have been refactored to implement a common \ObjectStoreProvider\ trait, standardizing how storage backends are initialized. This change introduces native support for Hugging Face (\hf://\), GooseFS (\goosefs://\), and Alibaba Cloud OSS (\oss://\), while extending Azure support to include the \abfss://\ scheme for ADLS Gen2 via OpenDAL. AWS, GCP, and Azure providers now utilize OpenDAL as an alternative backend (controlled by configuration) and support dynamic credential vending, allowing credentials to be refreshed without recreating the store. Additionally, a new \shared-memory://\ scheme enables process-wide in-memory state sharing across components.

_rust/lance-io/src/object\store/providers · high confidence

Vendored tokenizer stack with Lindera and Jieba support

The tokenizer implementation for full-text search has been moved into the lance-index crate, introducing dedicated tokenizers for Japanese and Korean text via Lindera and Chinese text via Jieba. This change adds new configuration handling for these language models and implements specific tokenization logic for JSON and text document types within the inverted index.

rust/lance-index/src/scalar/inverted/tokenizer · high confidence

Architecture

Extracted row-selection primitives into a new lance-select crate

The row-selection logic previously embedded in lance-core and lance-index has been extracted into a new, standalone lance-select crate. This crate provides the core mask types (RowAddrMask, NullableRowAddrMask) and index expression result wrappers (IndexExprResult, NullableIndexExprResult) that define which rows survive a filter, along with benchmarks to track their performance. This change decouples these primitives from larger dependencies, allowing downstream filtering code and benchmarks to depend on the mask substrate without pulling in the full lance-core or lance-index crates.

rust/lance-select · high confidence

Introduce lance-core crate with core data types, error handling, and blob v2 schema definitions

This change introduces the \lance-core\ crate, consolidating foundational components previously scattered across the codebase. It defines the core data types and schema structures, including the new Blob v2 logical and descriptor field layouts, and provides a comprehensive error handling system with Levenshtein-based suggestions for missing fields. Additionally, it includes a custom \DeepSizeOf\ trait for accurate Arrow-aware memory accounting, replacing the previous \deepsize\ crate, and exposes system column metadata and dataset traits to decouple row-taking logic from the main dataset implementation.

rust/lance-core/src · high confidence

Migrate schema and data types to lance-core

The core schema and data type definitions (Field and Schema structs) have been moved into the lance-core crate. This refactoring centralizes the schema logic, introducing support for unenforced primary and clustering keys with explicit ordering, and adds validation for FixedSizeList dimensions to prevent runtime panics on zero-dimension lists.

rust/lance-core/src/datatypes · high confidence

Refactor encoding mechanisms to be version-free

The encoding implementation has been refactored to decouple the core encoding logic from specific file format versions. The \array\_encoding\ module now exposes a reusable \ArrayFieldEncodingStrategy\ that implements the persisted \ArrayEncoding\ grammar, while the \encoder/structural\ module provides version-free structural field encoder builders. This change allows the same encoding mechanisms to be composed and accepted across different file versions, with version-specific decisions handled separately.

rust/lance-encoding/src · high confidence

Behavioural changes

Centralized dataset version policies and write validation

The dataset write logic now enforces strict version policies through a new \versions\ module, preventing the mixing of V1 and V2 data files within the same dataset. This change introduces version-specific behaviors for schema comparison, scan stream creation, and blob handling (including legacy vs. V2 blob validation), ensuring that write operations target the correct file format version and that schema evolution rules are applied consistently across different Lance file versions.

rust/lance/src/dataset/versions · high confidence

Input validation and structural writer refactoring in file writing

The file writer now validates that input arrays match the file schema's Arrow types and extensions before encoding, rejecting mismatches with clear error messages to prevent data corruption. Additionally, the underlying writing logic has been refactored to use a new \StructuralFileSink\ for managing page metadata and spilling, which supports the exact current-format writers and improves memory management during large writes.

rust/lance-file/src/writer · high confidence

Introduce Fragment Reuse Index for deferred index remapping

The index subsystem now uses a Fragment Reuse Index (FRI) to track row-address translations during compaction and deletions, allowing index remapping to be deferred until search time. This change introduces a new \IndexSegment\ structure to represent physical index segments with their own provenance (UUID, dataset version, and fragment coverage) and adds a \QueryRowIdRemapper\ that applies FRI history to translate row IDs for both scalar and vector index queries. Legacy index formats that cannot support this batch remapping are automatically excluded from coverage to ensure correct search results.

rust/lance/src/index · high confidence

Introduce internal sharded batch iterator and caching utilities

The internal dataset module now includes a ShardedBatchIterator class that enables reading Lance datasets in parallel across distributed PyTorch workers by sharding data at either the fragment or batch level, along with a CachedDataset helper for streaming data to temporary Arrow files. The ShardedBatchIterator is marked as deprecated in favor of the new Sampler API, and the module exposes these internal APIs for use in PyTorch data loading workflows.

_python/python/lance/\dataset · high confidence

Introduces memory-bounded manifest scanning and typed file schemas

The dataset file handling now uses a shared, memory-bounded manifest walker that lists manifest locations and reads manifests with a 1 GB in-flight memory budget, preventing unbounded memory usage during scans. This infrastructure supports new, strongly-typed Arrow schemas for file metadata: tracked files are now represented with a dictionary-encoded type field (mapping discriminants for manifest, data, deletion, transaction, and index files) to ensure consistent labeling, while all files include size and last-modified timestamp fields. These changes provide a more robust and performant foundation for dataset introspection and file tracking.

rust/lance/src/dataset/files · high confidence

Inverted index rewritten with new format, caching, and search logic

The scalar inverted index implementation has been completely refactored into a new modular structure (cache, doc\_set, flat\_search, format, inverted\_index, partition, posting builders). This introduces a new default FTS format version (V2) with varint-delta posting tails and shared position streams, replacing the legacy Arrow-based storage. The change adds a new slice-aware posting list cache to reduce memory overhead, a new Unicode-aware Levenshtein automaton for fuzzy matching, and a flat search path for unindexed data. Users will see changes in index file formats and potentially different memory usage patterns, with existing V1 indexes being automatically migrated to the new format on update.

rust/lance-index/src/scalar/inverted/index · high confidence

Merge insert now supports conditional updates, row failures, and column-level rewrite modes

The merge insert operation has been significantly enhanced with new behavioral options and execution paths. Users can now specify \WhenMatched::Fail\ to abort the operation if a matching row is found, and \WhenMatched::UpdateIf\ to conditionally update rows based on a provided expression. The system also supports \WhenNotMatchedBySource::DeleteIf\ to conditionally delete target rows that have no corresponding source row. Additionally, the write path now distinguishes between \RewriteRows\ (deleting and re-inserting full rows) and \RewriteColumns\ (attaching new data files for source columns while tombstoning old versions), allowing for more efficient updates when only specific columns change. These changes are implemented via a new action-based logical plan node and sentinel column logic to ensure NULL-safe row detection.

_rust/lance/src/dataset/write/merge\insert · high confidence

New commit retry timeout configuration and refactored I/O module structure

Users can now configure the commit conflict-retry timeout via the LANCE\_COMMIT\_RETRY\_TIMEOUT\_SECS environment variable, allowing long-running maintenance operations to succeed under high contention without hitting the default 30-second limit. Additionally, the I/O module has been restructured: commit handling logic is now centralized in a new commit.rs file with configurable backoff, deletion file reading is extracted into its own module with caching, and execution nodes are reorganized under io/exec.rs to expose updated scan, filter, and KNN execution interfaces.

rust/lance/src/io · high confidence

New conflict resolution logic for tagged fragment reuse indices

The commit handler now includes a comprehensive conflict resolver that correctly handles concurrent operations on tables with tagged fragment reuse indices. This ensures that mutations like deletes, updates, and merge inserts do not break the index's provenance tracking, and that the index coverage is correctly derived from live fragments rather than stale stored provenance.

rust/lance/src/io/commit · high confidence

New internal graph builder and I/O components for HNSW index construction

The \rust/lance-index/src/vector/graph\ module now includes a new \builder.rs\ file defining the \GraphBuilderNode\ struct, which manages neighbor relationships and ranking during the construction of HNSW graphs, and an \io.rs\ file stub for graph input/output operations. These changes provide the internal infrastructure required to build and persist HNSW graph structures, supporting the broader HNSW implementation updates in the vector index.

rust/lance-index/src/vector/graph · medium confidence

Python SDK development environment and documentation overhaul

The Python SDK now mandates \uv\ for all local environment setup, dependency management, and command execution (replacing \pip\, \make test\, etc.), with a new \AGENTS.md\ guide detailing this workflow. A comprehensive \DEVELOPMENT.md\ document has been added to cover building the Rust extension, running tests, and profiling benchmarks, while \CONTRIBUTING.md\ and \README.md\ have been updated to reflect these changes. Additionally, the \Makefile\ has been rewritten to support \uv\-based linting, formatting, and testing, and third-party license files for both Python and Rust dependencies have been added.

python · high confidence

Refactored Lance v1 file I/O into a modular, self-contained implementation

The v1 file format implementation has been reorganized into a dedicated module (\rust/lance-file/src/versions/v1\) with a clear separation of concerns. The new structure includes specific encoding modules for binary, dictionary, and plain data types, a format layer for metadata and page tables, and distinct reader and writer components. This refactoring establishes v1 as the canonical legacy file owner and composes exact current-format readers, ensuring that the v1 wire layout and schema dictionary payloads are handled by these dedicated codecs without leaking into the version-free I/O layer.

rust/lance-file/src/versions/v1 · high confidence

Spill row lineage to data files and validate stable row IDs

Users can now enable the \lance.row\_lineage.spill\ table configuration to move large row lineage sequences (row IDs and version metadata) out of the manifest and into hidden columns of the fragment's data files, preventing manifest bloat during compaction and updates. This change also adds integrity checks in \Dataset::validate\ to ensure stable row IDs remain unique and consistent with physical row counts across fragments.

rust/lance/src/dataset/rowids · high confidence

Stricter validation and error handling for Lance v2.0 encoding

The array encoding module now enforces stricter invariants for the v2.0 format. It rejects lossy null structures in structs and returns explicit errors for unsupported v2.0 field types, ensuring data integrity during encoding. Additionally, the module introduces new logical and physical encoding strategies (including packed struct and basic page schedulers) to support these validation rules and improve the robustness of the encoding pipeline.

_rust/lance-encoding/src/array\encoding · high confidence

Vendored internal tokenizer stack into lance-tokenizer

The \rust/lance-tokenizer\ crate now contains a self-contained implementation of the tokenizer stack, including the core \TextAnalyzer\ API, multiple tokenizer implementations (ICU, Jieba, Lindera, Code, N-gram, Simple, Raw), and a suite of token filters (ASCII folding, lowercasing, stemming, stop-word removal, and more). This change consolidates the tokenization logic previously scattered across the codebase into a single, internal library used by Lance, ensuring consistent text processing for full-text search without relying on external Tantivy components.

rust/lance-tokenizer · high confidence

Fixes

1 commit (1 fix) fixing test\_image\_dataset

A fix in test\_image\_dataset — 1 commit (1 fix), 1 file.

_test\_image\dataset · low confidence · unverified

Test coverage

Add benchmarks for v2 file reader performance and schema reconstruction; Added JNI test helper for FFI validation; Added Java SDK test coverage for async scanning, cleanup, compaction, and commit operations; Added Java SDK tests for dataset operations; Added Java tests for scalar/vector index creation, distributed indexing, and zonemap statistics; Added MemTable vs RocksDB KV point-lookup benchmark; Added benchmarks and tests for cache key performance; Added benchmarks for I/O scheduler performance; Added benchmarks for manifest interning, row ID indexing, and system column streaming; Added historical test fixtures for backward compatibility; Added integration and regression tests for memory management and durability; Added integration tests for GCS, GooseFS, and TOS object stores; Added integration tests for Java MemWAL bindings; Added integration tests for SQL aggregate pushdown optimization; Added integration tests for query execution across data types and index strategies; Added memory usage tests for vector indexing and data writing; Added test coverage for PyTorch integration utilities; Added test fixtures for jieba and lindera tokenization models; Added test resource for Jieba language model; Added test to verify panic location in distance functions; Added test utilities for covering indexes, storage failure injection, and cache serialization; Added tests for InvertedIndexParams serialization and validation; Added tests for Java Full-Text Search API and Scanner integration; Added tests for Java namespace implementations and dynamic context support; Added tests for MemWAL future depth and cross-generation fencing; Added tests for OpenTelemetry metrics bridge; Added tests for binary copy compaction behavior; Added tests for fragment update results and metadata classes; Expanded Python test coverage for query coercion, empty indices, and object store registry; Expanded test coverage for dataset operations and features; New CI benchmark suite for Python API performance; New FineWeb FTS benchmark suite for MemWAL; New LSM vector search benchmarks for MemWAL; New MemTable read performance benchmark; New MemWAL HNSW parity and recall benchmarks; New MemWAL write and replay benchmarks; New Python benchmark suite for data I/O, indexing, and vector operations; New benchmark suite for lance-linalg distance and utility kernels; New benchmarking suite for Lance encoding performance; New benchmarks for concurrent appends, COUNT pushdown, and distributed vector builds; New benchmarks for scalar and vector index performance; New encoding module structure and fuzz testing infrastructure; New library compatibility test suite for cross-version validation.

Dependencies

Add new benchmarking and testing utility packages

Introduces several new packages to support development and testing workflows: a \lance-memtest\ Rust/Python crate for memory allocation testing utilities, and multiple new benchmark projects (including \dbpedia-openai\, \bigann\, \hd-vila\, \sift\, \wiki\, and \full\_report\) with their respective Python dependencies and configurations. Additionally, adds new Rust crates for Arrow scalar support (\lance-arrow-scalar\), Arrow statistics (\lance-arrow-stats\), and various compression algorithms (\lance-bitpacking\, \fsst\), alongside updated example and core library manifests.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 69 → 70 (+1.0)
  • Rubric changed (rubric-2026.09.9 → rubric-2026.09.18) — scores are not directly comparable.

Lenses

  • Code Health 82 → 82 (-0.0)
  • Architecture 97 → 96 (-1.4)
  • Maturity 76 → 78 (+2.2)
  • Readiness 93 → 75 (-18.5)
  • Security 59 → 62 (+3.5)
  • Accessibility 72 → 75 (+2.8)
  • Performance 82 (new)

Resolved (152)

  • BitmapIndexPlugin::streaming_build_and_write (cognitive 16) (rust/lance-index/src/scalar/bitmap.rs)
  • Change coupling: feature_flags.rs ↔ transaction.rs (rust/lance-table/src/feature_flags.rs)
  • ClassTooLong: BitmapIndexPlugin (rust/lance-index/src/scalar/bitmap.rs)
  • ComplexAllNullScheduler::initialize (cognitive 16) (rust/lance-encoding/src/encodings/logical/primitive.rs)
  • Dataset::open_vector_index (cognitive 17) (rust/lance/src/index.rs)
  • Dataset::open_vector_index (cyclomatic 20) (rust/lance/src/index.rs)
  • Duplicated block (10 lines × 2) (rust/lance-index/src/vector/v3/shuffler.rs)
  • Duplicated block (10 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
  • Duplicated block (10 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/planner.rs)
  • Duplicated block (10 lines × 3) (rust/lance-encoding/src/decoder.rs)
  • Duplicated block (11 lines × 2) (rust/lance-index/src/scalar/bloomfilter.rs)
  • Duplicated block (11 lines × 2) (rust/lance/src/dataset/mem_wal/memtable/scanner/builder.rs)
  • Duplicated block (11 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/point_lookup.rs)
  • Duplicated block (11 lines × 2) (rust/lance/src/dataset/scanner.rs)
  • Duplicated block (11 lines × 3) (rust/lance-linalg/src/distance/cosine.rs)
  • Duplicated block (12 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
  • Duplicated block (12 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/planner.rs)
  • Duplicated block (12 lines × 4) (rust/lance/src/dataset/fragment.rs)
  • Duplicated block (12 lines × 4) (rust/lance/src/io/commit/conflict_resolver.rs)
  • Duplicated block (12–13 lines × 2) (rust/lance/src/dataset/write/merge_insert.rs)
  • …and 132 more

New (330)

  • Ambiguous distinction between update_columns and update_columns_with_offsets. Without documentation, it is unclear if the latter is a performance optimization, a different join strategy, or a bug-prone variant. They appear to do the same high-level operation (updating columns via join) but with a signature difference that isn't self-explanatory.
  • ClassTooLong: CompoundQueryExec (rust/lance/src/io/exec/fts.rs)
  • ClassTooLong: InvertedIndexParams (rust/lance-index/src/scalar/inverted/tokenizer.rs)
  • ClassTooLong: LsmFtsSearchPlanner (rust/lance/src/dataset/mem_wal/scanner/fts_search.rs)
  • ClassTooLong: Manifest (rust/lance-table/src/format/manifest.rs)
  • ClassTooLong: MinHashLshIndex (rust/lance-index/src/scalar/minhash_lsh/index.rs)
  • ClassTooLong: U64Segment (rust/lance-table/src/rowids/segment.rs)
  • Critical CVE: [GHSA redacted] (java/pom.xml)
  • Dataset::open_vector_index_from_metadata (cognitive 20) (rust/lance/src/index.rs)
  • Dataset::open_vector_index_from_metadata (cyclomatic 22) (rust/lance/src/index.rs)
  • Documentation: no project overview (rust/arrow-stats/README.md)
  • Duplicated block (10 lines × 2) (java/src/main/java/org/lance/fragment/RowIdMeta.java)
  • Duplicated block (10 lines × 2) (rust/lance-index/src/scalar/inverted.rs)
  • Duplicated block (10 lines × 2) (rust/lance-index/src/scalar/label_list.rs)
  • Duplicated block (10 lines × 2) (rust/lance-index/src/vector/bq/pairwise.rs)
  • Duplicated block (10 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
  • Duplicated block (10 lines × 2) (rust/lance-table/src/transaction/manifest_build.rs)
  • Duplicated block (10 lines × 2) (rust/lance/src/index/frag_reuse.rs)
  • Duplicated block (10 lines × 3) (rust/lance-encoding/src/decoder.rs)
  • Duplicated block (10–11 lines × 2) (java/src/main/java/org/lance/OpenDatasetBuilder.java)
  • …and 310 more

Changes since last survey

  • 204 commits — 131 feature/other, 73 fixes

By area

  • rust/lance — 70 commits
  • rust/lance-index — 29 commits
  • (root) — 22 commits
  • python/python — 17 commits
  • rust/lance-table — 12 commits
  • rust/lance-encoding — 11 commits
  • docs/src — 10 commits
  • rust/lance-linalg — 9 commits
  • rust/lance-io — 5 commits
  • .github/workflows — 4 commits
  • java/src — 4 commits
  • rust/lance-namespace-impls — 3 commits
  • rust/lance-core — 2 commits
  • rust/lance-datafusion — 2 commits
  • benchmarks/auto-ivf-dot — 1 commit
  • java/lance-jni — 1 commit
  • java/pom.xml — 1 commit
  • test_data/v8.0.0 — 1 commit

Notable commits

  • fix: fix!: derive merge_columns field ids from the dataset manifest (#9547)
  • fix: fix(core): parse fixed-offset timezones in timestamp logical types (#8950)
  • fix: fix(datafusion): coerce numeric literals to and from Float16 (#8847)
  • fix: fix(datafusion): report the scan range in scan statistics (#9411)
  • fix: fix(datafusion): size the memory pool by the effective partition count (#9183)
  • fix: fix(dataset): correct BlobFile seek semantics (#9358)
  • fix: fix(dataset): fix open branch URI with tag to non-latest version (#9227)
  • fix: fix(dataset): handle lance dataset blobs on deep clone (#9185)
  • fix: fix(dataset): keep carried index bases through chained shallow clones (#9176)
  • fix: fix(dataset): reject a zero add_columns batch size instead of panicking (#9233)
  • fix: fix(dataset): reuse the prefetch window in BlobFile::read (#9330)
  • fix: fix(dataset): stop the fragment-reuse index blocking row id migration (#9099)
  • fix: fix(deps): update rustls for RUSTSEC-2026-0285 (#9212)
  • fix: fix(encoding): derive full-zip max_visible_def like the writer (#9254)
  • fix: fix(encoding): emit NIL control words when rep/def levels are non-empty but zero-width (#9018)
  • fix: fix(encoding): preserve special slots when recording validity (#9268)
  • fix: fix(encoding): skip full-width block bitpacking (#9598)
  • fix: fix(encoding): split legacy decode batches before i32 offset overflow (#9225)
  • fix: fix(exec): preserve exact KNN ordering through late materialization (#9469)
  • fix: fix(format): move FLAG_UNKNOWN off the tagged FRI bit (#9420)
  • …and 184 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

lance-format/lance was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 5483093251925b8e90703f216e150b263d04f587 — the exact code this score is about.
  • Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-c4983f2d4e5c.