lance-format/lance
70.0
Adequate · 29 September 2026
776.5k
lines of production code
Rust
primary language
2
measurements over time
What this system is
This system is a high-performance columnar data storage and vector search engine, implemented primarily in Rust with Python and Java SDKs. It provides a modern file format (Lance) optimized for efficient I/O, supporting complex nested data, blob storage, and various scalar and vector index types like IVF-PQ and HNSW. The platform enables scalable data ingestion, concurrent updates via MemWAL, and seamless integration with Apache DataFusion for SQL querying and PyTorch for machine learning workflows.
How it got here
2022–2023 — modularization and vector index v2
53 changes.
This period focused on restructuring the Rust codebase into specialized crates and implementing a new v2 architecture for IVF vector indexes. It also introduced comprehensive benchmarking suites, PyTorch integration, and DataFusion support to enhance performance profiling and ecosystem compatibility.
2024–2025 — Lance 2.0 format and Java SDK launch
82 changes.
This period focused on the foundational rewrite of the Lance file format to version 2.x, introducing a modular encoding architecture, stable row IDs, and data overlay capabilities. It also marked the initial release of the Java SDK, providing comprehensive bindings for dataset operations, vector indexing, and full-text search. Concurrently, the project expanded its vector index support with new HNSW and Scalar Quantization implementations while establishing robust CI and benchmarking infrastructure.
2026 — MemWAL and file format evolution
51 changes.
This period focused on implementing MemWAL, an in-memory write-ahead log with LSM-style tiering, to enable high-throughput, low-latency writes and unified in-memory search across scalar, vector, and full-text indexes. Concurrently, the project advanced its storage layer by introducing the Lance v2.0 file format and its subsequent versions, alongside a comprehensive refactor of the inverted index and distributed index merging infrastructure.
Features
Add BigANN benchmark dataset preparation scripts
Added a new benchmark suite for the BigANN dataset, including a Python script (\dataset.py\) and documentation (\README.md\) to prepare text-to-image datasets in Lance format. The script supports creating \text2image-10m\ and \yfcc-10m\ datasets, converting raw binary data into Lance files with ground truth queries, and includes a \.gitignore\ to exclude generated artifacts.
benchmarks/bigann · high confidence
Add Cohere Wikipedia embedding benchmark for index build performance
A new benchmark located in benchmarks/wiki allows users to generate a 35M vector dataset from the Cohere Wikipedia embeddings and build an IVF\_PQ index using a disk-based shuffler. Users can run datagen.py to create the Lance dataset and index.py to construct the index with configurable metrics (L2, cosine, dot), partition counts, and sub-vector counts, specifically designed to measure indexing speed rather than recall or latency.
benchmarks/wiki · high confidence
Add Dbpedia-entities benchmark for vector search performance evaluation
Users can now run a benchmark using the Dbpedia-entities-openai dataset (1M embeddings) to evaluate vector search performance. This includes a data generation script (\datagen.py\) to convert the Hugging Face dataset into Lance format, and a benchmark runner (\benchmarks.py\) that tests top-k recall across various IVF\_PQ index configurations (IVF 256/512/1024, PQ 32/96/192) and refine factors.
benchmarks/dbpedia-openai · high confidence
Add FSST compression library and benchmark example
Introduces the FSST (Fast Static Symbol Table) compression implementation in \rust/compression/fsst\, including the core logic in \src/fsst.rs\ and library entry point in \src/lib.rs\. Adds a benchmark example (\examples/benchmark.rs\) that demonstrates compressing and decompressing string data using the FSST algorithm, highlighting the output buffer size contract (8x expansion) and performance metrics.
rust/compression/fsst · high confidence
Add OpenTelemetry metrics bridge for Lance
The Java library now includes a new \org.lance.otel\ package that bridges Lance's internal Rust metrics to OpenTelemetry. Users can call \LanceMetrics.instrument()\ to register observable instruments (counters, gauges, and histograms) on the global OpenTelemetry meter provider, allowing Lance's native metrics to be exported via standard OTel pipelines. The bridge supports querying the metric catalog and taking point-in-time snapshots of recorded data.
java/src/main/java/org/lance/otel · high confidence
Add Python benchmark dataset generation utilities
The \ci\_benchmarks/datagen\ module now provides Python scripts to generate synthetic datasets for CI benchmarks, including a basic dataset, a TPC-H lineitems dataset (using DuckDB), a Wikipedia dataset for full-text search (streaming from HuggingFace), a count\_rows dataset with various scalar indexes, and merge-insert datasets with specific fragmentation and deletion patterns.
_python/python/ci\benchmarks/datagen · high confidence
Add owned, SIMD-accelerated bitpacking codecs
The \lance-bitpacking\ crate now includes its own internal bitpacking implementation (\bitpacker\_internal\) featuring SIMD-optimized kernels for x86\_64 (SSE3, AVX2) and aarch64 (NEON). This change introduces new owned codec types (\BitPacker4x\, \BitPacker8x\) that provide high-performance compression and decompression for \u32\ blocks, replacing or augmenting previous external dependencies with a Lance-owned, byte-compatible implementation.
rust/compression/bitpacking · high confidence
Add standalone CLI benchmark for PK-based point lookups across LSM levels
A new standalone CLI benchmark (\mem\_wal\_point\_lookup\_bench\) has been added to measure point lookup latency against three tiers of the LSM tree: the base table (on-disk, compacted data), SSTables (on-disk L0), and the active MemTable (in-memory write buffer). The tool supports two phases—\prepare\ to create the base dataset and initialize MemWAL, and \lookup\ to ingest rows via ShardWriter and time point lookups—allowing users to configure parameters such as base rows, maximum MemTable rows, number of SSTables, and query count, with results optionally output as JSON.
_rust/lance/benches/mem\_wal/point\lookup · high confidence
Added ExpLinkedList for memory-efficient storage
A new \ExpLinkedList\ container has been added to \lance-core\. This data structure stores elements in a linked list of vectors, where each vector's capacity doubles when full, providing a memory-efficient way to handle large numbers of elements. It supports standard operations like push, pop, and iteration, and implements \DeepSizeOf\ for accurate memory accounting.
rust/lance-core/src/container · high confidence
Added IOTracker for monitoring and metrics of local storage operations
A new \IOTracker\ utility has been added to \rust/lance-io/src/utils/tracking\_store.rs\ to wrap ObjectStore implementations and track I/O statistics. This component enables the recording of read and write operations, including byte counts and IOPS, specifically for local reads and writes that bypass the standard ObjectStore layer. It also integrates with the metrics system to publish object store metrics for these local operations, ensuring that performance data is accurately captured even when optimized local paths are used.
rust/lance-io/src/utils · high confidence
Added Rust crate documentation and build configuration for Lance
This change introduces the initial structure for the Rust implementation of the Lance format. It adds a README file detailing installation via Cargo and providing code examples for creating datasets, reading data, performing vector index operations, and other core features. Additionally, it includes a build script (build.rs) that configures the Protobuf compilation process for ANN (Approximate Nearest Neighbor) protocols and sets up external path mappings for generated code.
rust/lance · high confidence
Added TPCH benchmark script to compare Lance and Parquet performance
Users can now run a Python-based benchmark in the \benchmarks/tpch\ directory to compare query performance between Lance and Parquet formats. The new \benchmark.py\ script supports TPCH queries Q1 and Q6, allowing users to generate sample datasets locally using DuckDB and measure latency differences between the two storage engines.
benchmarks/tpch · high confidence
Added code agent skills for the Lance user guide
This change introduces a new 'skills' directory containing structured documentation and scripts designed to guide code agents in assisting Lance users. It includes a primary skill definition (\lance-user-guide/SKILL.md\) that instructs agents on how to help with dataset operations (write, read, scan), vector and scalar index creation, and troubleshooting. The entry also adds reference materials for index selection and I/O operations, along with a Python end-to-end script (\python\_end\_to\_end.py\) that demonstrates a complete workflow of writing data, building indices, and performing queries.
skills · high confidence
Added flat vector index benchmark script
A new benchmark script has been added to the \benchmarks/flat\ directory to measure the latency of flat vector search. Running the script generates synthetic datasets with varying dimensions and lengths, executes nearest-neighbor searches using L2, cosine, and dot-product metrics, and outputs the results to \benchmark.csv\ along with a latency plot in \benchmark.html\.
benchmarks/flat · high confidence
Added fragment reuse index management and cleanup
The dataset now exposes a public accessor to retrieve the fragment reuse index (FRI), which tracks how physical row addresses have moved during compactions that defer index remapping. Additionally, a new cleanup function automatically trims the FRI by removing older reuse versions once all existing indices have caught up to them, ensuring that stale mapping data does not accumulate unnecessarily.
rust/lance/src/dataset/index · high confidence
Builder-style API for scalar index configuration and zone statistics
Java users can now configure scalar indices (B-Tree, Bitmap, Inverted, LabelList, NGram, ZoneMap) using dedicated builder classes (e.g., BTreeIndexParams.builder()) instead of raw JSON or maps. The InvertedIndexParams builder exposes full-text tuning options including analyzer presets (text/code), base tokenizers (simple, whitespace, raw, ngram, code, ICU, Lindera, Jieba), stemming, stop words, n-gram ranges, posting block size, and format version gating. B-Tree and ZoneMap builders allow setting zone/row sizing, while Bitmap supports explicit shard IDs for distributed builds. Additionally, a new ZoneStats class exposes per-zone min/max/null counts from zonemap indices via JNI, enabling users to inspect index statistics directly in Java.
java/src/main/java/org/lance/index/scalar · high confidence
Calibrated auto-probing for dot-product IVF\_FLAT searches
This change introduces a new benchmarking suite and protocol for calibrating the initial probe budget in IVF\_FLAT indices using dot-product metrics. The new \benchmarks/auto-ivf-dot\ directory contains Python scripts (\prepare.py\, \calibrate.py\, \measure.py\, \audit.py\) and documentation (\PROTOCOL.md\, \README.md\, \RESULTS.md\) that define a frozen evaluation contract on the Wiki-Cohere 35M and DPR Wikipedia datasets. The calibration process selects specific floor, cap, and margin parameters for the auto-probing policy to minimize scanned partitions while maintaining at least 95% recall, addressing issues with the legacy signed-distance multiplier that previously performed poorly on dot products with varying query norms.
benchmarks/auto-ivf-dot · high confidence
Establish automated release and dependency governance tooling
The repository now includes configuration for \bumpversion\ to manage synchronized version bumps across Rust, Python, and Java artifacts, alongside \cargo-deny\ to enforce dependency license compliance, security advisory checks, and workspace dependency consistency. Pre-commit hooks are added to enforce code formatting, spell-checking, and lockfile synchronization, while a formal release process document and CI infrastructure (including LocalStack for S3 testing) standardize how beta, RC, and stable releases are built and published.
(repo-wide) · high confidence
Expose Lance datasets as DataFusion TableProviders
Users can now query Lance datasets directly using DataFusion's SQL and DataFrame APIs. This change introduces a \TableProvider\ implementation for \Dataset\ (in \logical\_plan.rs\) and a configurable \LanceTableProvider\ (in \dataframe.rs\) that supports filter, projection, and limit pushdown. The provider allows optional inclusion of system columns (\\_row\_id\, \\_row\_addr\), custom blob handling policies, and batch size configuration, enabling seamless integration of Lance data into DataFusion-based query engines.
rust/lance/src/datafusion · high confidence
Initial Java SDK release with Maven wrapper and documentation
This change introduces the initial Java bindings and SDK for Lance, providing a new Maven-based project structure under the \java/\ directory. It includes a Maven wrapper (version 3.3.2) for consistent builds, a \README.md\ with quick-start examples for dataset creation, random access, and schema evolution, and configuration files for code formatting (Spotless, Scalafmt). The release also bundles generated third-party license documents for both Java and Rust dependencies to ensure compliance.
java · high confidence
Introduce Binary Quantization (BQ) and RaBitQ support for vector indexes
Added a new Binary Quantization (BQ) module and RaBitQ (Rabit Quantization) support to the vector index pipeline. This includes a new \BinaryQuantization\ transformer that converts float vectors to binary codes based on sign bits, and a \RabitQuantizer\ that supports configurable bit-depths (1-9 bits) and rotation types (Fast/Matrix) for IVF\_RQ indexes. The change also integrates these quantizers into the IVF transformer pipeline, allowing users to leverage binary and reduced-bit quantization for more compact storage and faster search performance.
rust/lance-index/src/vector · high confidence
Introduce Bitmap and Bloom Filter scalar indices
Added Bitmap and Bloom Filter scalar index implementations to the Lance index library. The Bitmap index is designed for low-cardinality columns, storing a bitmap for each unique value to quickly identify matching rows. The Bloom Filter index provides a probabilistic data structure for efficient existence checks, helping to prune non-matching data during scans. Both indices integrate with the existing scalar index plugin framework, supporting training, querying, and caching.
rust/lance-index/src/scalar · high confidence
Introduce CacheCodec for FTS index entries and cross-column scoring
The inverted index module now supports persistent caching of posting lists and positions via a new \CacheCodec\ implementation, enabling efficient serialization and zero-copy deserialization of FTS index data for backends. Additionally, a new \combined\_fields\ scoring mode allows cross-column BM25F queries by treating target columns as a single virtual field, and a \cross\_column\ module enables compound FTS queries across multiple indexed columns by mapping document identities to a shared row-address domain.
rust/lance-index/src/scalar/inverted · high confidence
Introduce DataFilePart and external blob base resolution for staged writes
The dataset write path now supports caller-managed data file parts, allowing large writes to be staged and assembled later. A new DataFilePart struct serializes the identity and metadata of completed staging parts, which can be checkpointed and later assembled into a complete data file via DataFilePart::open\_all. This also introduces ExternalBaseResolver to map external blob URIs to specific storage bases, enabling multi-base blob storage and proper resolution of external blob paths during writes.
rust/lance/src/dataset · high confidence
Introduce FFI bindings for RecordBatchStream and new SpillStore for scratch storage
The lance-io crate now exposes a new \ffi\ module that wraps a \RecordBatchStream\ into an \FFI\_ArrowArrayStream\, enabling cross-language data exchange via the Arrow FFI interface. Additionally, a new \SpillStore\ trait and its \LocalSpillStore\ implementation have been added to provide reclaimable scratch storage on the local disk, allowing temporary data (such as index shuffle runs) to be written to disk and read back within the same process, with automatic cleanup and configurable disk capacity limits.
rust/lance-io/src · high confidence
Introduce IVF\_FLAT vector index implementation
Added the \IVF\_FLAT\ index type to the \lance-index\ crate, enabling exact vector search within IVF partitions. This change introduces the \FlatIndex\ as a sub-index implementation, along with \FlatFloatStorage\ for in-memory vector storage, \FlatPairScorer\ for exact distance calculations (supporting Float16, Float32, Float64, and binary vectors with L2, Cosine, Dot, and Hamming metrics), and \FlatTransformer\ for column normalization. The implementation includes optimized search paths using a binary heap for top-k results and supports distance range queries.
rust/lance-index/src/vector/flat · high confidence
Introduce JSONB scalar UDFs for DataFusion
Adds a new set of user-defined functions (UDFs) in \rust/lance-datafusion/src/udf/json.rs\ that enable querying JSONB data stored as LargeBinary arrays within DataFusion. These functions provide type-aware extraction capabilities, allowing users to retrieve JSON fields and array elements by key or index, and convert JSONB values into standard Arrow types such as strings, integers (Int64), and floats (Float64).
rust/lance-datafusion/src/udf · high confidence
Introduce LSM scanner for unified reads across base table, SSTables, and in-memory memtables
The \mem\_wal/scanner\ module now provides a new \LsmScanner\ that unifies queries across the base Lance table, persisted SSTable generations, and active/frozen in-memory memtables. It automatically handles primary-key deduplication to ensure the newest version of each row is returned, supports vector (KNN) and full-text search (BM25) across all tiers, and applies cross-generation block-lists to suppress stale reads from older sources. The scanner also supports point lookups, nested projections, and schema reconciliation for generations written under previous table schemas.
_rust/lance/src/dataset/mem\wal/scanner · high confidence
Introduce Lance file format versions 2.1 and 2.2
Added new file format versions 2.1 and 2.2, each with dedicated writers, readers, and compression strategies. Version 2.1 uses a dense u16 miniblock layout and supports standard compression codecs, while version 2.2 upgrades to a dense u32 layout, enables large miniblock chunks, and adds support for variable packed struct per-value compression, general block compression, and structural blob encoding. Both versions provide explicit and lazy writer creation with configurable compression parameters.
_rust/lance-file/src/versions/v2\_1, rust/lance-file/src/versions/v2\2 · high confidence
Introduce Lance v2.0 file format support
Added the v2.0 file format implementation, including the reader, writer, and composition logic in the \rust/lance-file/src/versions/v2\_0\ module. This enables users to read and write data files using the v2.0 grammar, featuring specific handling for field encoding strategies, column metadata decoding, and page metadata spilling to bound memory usage during writes.
_rust/lance-file/src/versions/v2\0 · high confidence
Introduce Lance-native in-memory HNSW vector index for MemWAL
Added a new HNSW graph implementation within the MemWAL subsystem that enables fast approximate nearest neighbor search on in-memory vector data. This change introduces a zero-copy vector store backed by Arrow batches, allowing the index to borrow vector data without duplicating memory, and supports configurable build parameters (such as graph levels, edge count, and construction beam width) to tune performance. The implementation includes specific search and build parameter structures, distance computation kernels for L2, Dot, and Cosine metrics, and metadata serialization to persist the graph structure within Lance's batch format.
_rust/lance/src/dataset/mem\wal/hnsw · high confidence
Introduce MemWAL MemTable scanner with vector and full-text search support
This change introduces the new MemTable scanner implementation for the MemWAL (Write-Ahead Log) system, enabling in-memory search capabilities directly on the active memtable. Users can now perform vector searches (including HNSW-based nearest neighbor queries with configurable distance metrics and bounds) and full-text searches (supporting term, phrase, fuzzy, and boolean queries with cross-column support) on data that has not yet been flushed to disk. The scanner also handles nested column projections and ensures read-your-writes consistency by including the mutable tail in search results by default.
_rust/lance/src/dataset/mem\wal/memtable/scanner · high confidence
Introduce MemWAL-based MemTable with lock-free batch storage and flush pipeline
The MemTable implementation now uses a new MemWAL architecture featuring a lock-free \BatchStore\ for concurrent reads and serialized appends, a dedicated flusher that writes data to persistent SSTables with optional cache warming, and a DataFusion-integrated scanner supporting full scans and index queries (BTree, HNSW, FTS) with MVCC visibility. This change replaces the previous in-memory storage mechanism with a persistent, WAL-backed design that improves concurrency and durability for dataset writes.
_rust/lance/src/dataset/mem\wal/memtable · high confidence
Introduce MemWAL: an in-memory write-ahead log with LSM-style tiering for Lance datasets
This change adds the MemWAL subsystem to the Lance dataset API, enabling high-throughput, low-latency writes by buffering data in an in-memory MemTable before flushing to persistent storage. Users can initialize MemWAL on a dataset using a builder API that supports sharding strategies (manual, unsharded, bucket, or identity) and configurable maintained indexes (BTree, HNSW, and FTS). The system introduces an LSM-like architecture with a fresh tier (active MemTable) and a frozen tier (SSTables), providing features such as read-your-writes consistency, schema evolution reconciliation, and configurable backpressure. This allows applications to perform rapid batch inserts and updates with immediate visibility, while maintaining durability and supporting vector and full-text search indexes on the buffered data.
_rust/lance/src/dataset/mem\wal · high confidence
Introduce PyTorch integration for Lance datasets
This change adds a new \lance.torch\ module that enables using Lance datasets directly within the PyTorch ecosystem. It introduces \LanceDataset\ and \SafeLanceDataset\ to wrap Lance data as PyTorch datasets, supporting automatic conversion of PyArrow types (including fixed-size lists and bfloat16) to tensors. The module also provides GPU-accelerated distance computations (L2, cosine, dot) via \torch.compile\, a PyTorch-native K-Means implementation, utilities for distributed and multiprocessing training, and benchmarking helpers for ground-truth calculation.
python/python/lance/torch · high confidence
Introduce RaBitQ multi-bit quantization with SIMD-optimized search
This change adds a new RaBitQ (Randomized Binary Quantization) index implementation in the \lance-index\ crate, enabling multi-bit IVF\_RQ storage and search. The new module includes a builder for creating quantizers with configurable bit-widths (defaulting to 5 bits), a storage layer for persisting binary and extended codes, and a transformer for computing query factors. Search performance is significantly improved through dedicated SIMD kernels (AVX2/AVX-512) for distance table quantization, ex-code dot products, and lower-bound pruning, alongside a fast random rotation pipeline. The implementation supports both approximate and accurate search modes and includes validation to ensure metadata consistency.
rust/lance-index/src/vector/bq · high confidence
Introduce bfloat16 extension type and commit conflict error handling
The Python SDK now supports the bfloat16 data type via a new PyArrow extension type (BFloat16Array) with native NumPy conversion and Pandas integration, enabling efficient handling of 16-bit floating-point vectors. Additionally, a new CommitConflictError is exposed to allow users to programmatically detect and handle concurrent write conflicts with retry logic.
python/python/lance · high confidence
Introduce bfloat16 support and data generation utilities in Python SDK
The Python SDK now includes native support for the bfloat16 data type, exposing a \BFloat16\ class and a \bfloat16\_array\ function to create and manipulate bfloat16 Arrow arrays. Additionally, a new \datagen\ module has been added to the Python bindings, providing the \rand\_batches\ function to generate random data batches for a given schema, which is useful for testing and benchmarking.
python/src · high confidence
Introduce data overlay files and stable row ID format support
The \lance-table\ crate now supports data overlay files, which allow updating specific cell values within a fragment without rewriting the underlying data files, and introduces the v2 data storage format (stable row IDs). This includes new \DataOverlayFile\ and \OverlayCoverage\ structures for managing these overlays, along with logic to handle index staleness when overlays modify indexed fields. The manifest format is updated to track the data storage format version, and \IndexMetadata\ now includes file sizes and creation timestamps.
rust/lance-table · high confidence
Introduce distributed execution support for vector search and filtered reads
Added protobuf serialization for the \ANNIvfPartitionExec\ and \FilteredReadExec\ execution nodes, enabling these plans to be serialized and sent to remote workers for distributed execution. Additionally, introduced a \LanceFilterExec\ wrapper around DataFusion's \FilterExec\ to preserve the original logical expression, allowing filter predicates to be serialized to Substrait for remote planning. These changes provide the necessary infrastructure for distributing vector search and filtered read operations across a cluster.
rust/lance/src/io/exec · high confidence
Introduce exact v2.x file format readers and writers
The \lance-file\ crate now supports reading and writing the exact v2.x file formats (v2.0 through v2.3) with version-specific encodings and metadata handling. This change replaces the legacy v1 reader with a modular, version-dispatched architecture that validates file structure, handles blob v2 descriptors, and supports structural projections. Users can now explicitly target stable (v2.2) or next (v2.3) file versions for new data, while existing v1 files remain readable via a separate legacy path. The new format includes improved encoding mechanisms, exact version identity, and better compatibility testing across versions.
rust/lance-file/src · high confidence
Introduce experimental Lance v2.3 file format
Added a new v2.3 file format version that implements a distinct encoding and compression strategy. This version introduces a new compression strategy supporting miniblock, per-value, and block-level compressors (including bitpacking, RLE, and byte-stream split) and a field encoding strategy that composes primitive page encodings (sparse, constant, and dense) with structural encodings for blobs, maps, lists, and structs. The module provides writer functions (with and without explicit compression tuning) and reader validation logic, but emits a warning that this is an unstable format intended for experimentation only, as future compatibility is not guaranteed.
_rust/lance-file/src/versions/v2\3 · high confidence
Introduce flat B-tree index page implementation
Added a new \FlatIndex\ structure in the scalar B-tree module to represent a single index page as a sorted batch of value/row-id pairs. This implementation enables efficient on-demand filtering of row IDs using Arrow's vectorized operations, supporting optimized handling of \IsIn\ predicates and null tracking without materializing intermediate structures per page.
rust/lance-index/src/scalar/btree · high confidence
Introduce hierarchical, type-safe session caches for metadata and indices
The session layer now uses a new hierarchical caching system to organize dataset metadata and index data, preventing collisions between different datasets and indices. This change introduces \GlobalMetadataCache\ and \GlobalIndexCache\ as top-level namespaces, with sub-caches for specific datasets and indices. It also implements type-safe cache keys for manifest data, transactions, deletion files, row address masks, and index metadata, ensuring that cache entries are correctly scoped and versioned. This improves cache reliability and performance by avoiding stale or incorrect data reuse across different datasets and index types.
rust/lance/src/session · high confidence
Introduce in-memory MemWAL index backends for scalar, vector, and full-text search
The MemTable now uses dedicated in-memory index structures to accelerate point lookups and scans before data is flushed to disk. This includes a custom single-writer, lock-free-read skiplist (arena\_skiplist) optimized for the MemTable's append-only access pattern, a B-tree index (btree) for scalar fields with compact key encodings for integers and strings, a partition-structured full-text search index (fts) supporting BM25 scoring and complex queries, and an HNSW vector index (hnsw) with build parameters tuned for fast flush cycles. These components replace previous generic or less efficient indexing mechanisms within the MemWAL layer, providing faster read performance and better memory locality for active memtable operations.
_rust/lance/src/dataset/mem\wal/index · high confidence
Introduce io\_uring-based file reader for high-performance local I/O
Added a new io\_uring-based file reader implementation in \lance-io\ that supports both a dedicated background thread pool (\thread.rs\) and a thread-local mode (\current\_thread.rs\) for current-thread runtimes. This new reader provides a \file+uring://\ URI scheme, handles file caching, manages concurrent reads, and includes comprehensive tests for various read scenarios.
rust/lance-io/src/uring · high confidence
Introduce lance-arrow-scalar crate for efficient scalar comparisons
The rust/arrow-scalar crate (now named lance-arrow-scalar) provides an ArrowScalar type that wraps single-element Arrow arrays to enable O(1) comparison and hashing. By delegating to arrow\_row::OwnedRow, it ensures correct total ordering, proper NaN handling, and consistent null ordering, while also supporting serde serialization and conversion from primitive types.
rust/arrow-scalar · high confidence
Introduce lance-linalg crate with SIMD-accelerated distance metrics
The lance-linalg crate is introduced as an internal sub-crate containing native linear algebra algorithms for Lance. It provides distance metrics (L2, cosine, dot, hamming) for float types (f16, bf16, f32, f64) and integer types (u8, i8), featuring runtime-dispatched SIMD backends for x86\_64 (AVX2, AVX-512, AMX-FP16), aarch64 (NEON), and loongarch64 (LSX/LASX). The crate includes a build system that compiles C-based SIMD kernels and a \DeepSizeOf\ derive macro for Arrow-aware memory accounting.
rust/lance-linalg · high confidence
Introduce lance-tools CLI for inspecting Lance file metadata
A new \lance-tools\ command-line utility is added to the Rust workspace, providing a \meta\ subcommand that displays key metadata (version, row count, byte sizes, and schema) for Lance files. The tool supports both local filesystem paths and URI-based sources by leveraging the existing \ObjectStore\ registry and \ScanScheduler\ infrastructure.
rust/lance-tools · high confidence
Introduce memory allocation tracking for Python tests
Added a new \memtest\ utility that allows Python test suites to track memory allocations made by the Python interpreter and its libraries. The tool works by preloading a shared library (built in Rust) that intercepts memory calls, exposing statistics like peak usage and allocation counts via a Python API and CLI. It supports both Linux and macOS, enabling developers to detect memory leaks or regressions in their Python code.
memtest · high confidence
Introduce new physical encoding implementations for basic, binary, bitmap, and other data types
This change adds a suite of new physical encoding and decoding components in the \lance-encoding\ module, including \BasicPageScheduler\/\BasicEncoder\ for primitive fields, \BinaryPageScheduler\ for variable-length strings/binary data, \DenseBitmapScheduler\ for bitmaps, \BitpackedForNonNegScheduler\ for bit-packing, \DictionaryPageScheduler\ for dictionary-encoded arrays, \FixedSizeBinaryPageScheduler\ for fixed-size binaries, \FixedListScheduler\ for fixed-size lists, \FsstPageScheduler\ for FSST compression, and \PackedStructPageScheduler\ for packed structs. These new schedulers and decoders replace or supplement the previous encoding mechanisms, providing the underlying infrastructure for reading and writing these specific data layouts efficiently.
_rust/lance-encoding/src/array\encoding/physical · high confidence
Introduce non-blocking async scanner and blob file APIs in Java SDK
The Java SDK now exposes an \AsyncScanner\ API that executes scans on a dedicated Tokio runtime and returns results via \CompletableFuture\, allowing non-blocking I/O for large queries. Additionally, new JNI bindings for \BlobFile\ provide direct read capabilities (\read\, \readUpTo\, \readRange\), and the \Dataset\ API now includes methods to retrieve blob data by row IDs or indices. These changes are supported by a new dispatcher mechanism that safely bridges asynchronous Rust tasks back to the Java thread.
java/lance-jni · high confidence
Introduce optional geo module with spatial UDFs and bounding box utilities
The \rust/lance-geo\ crate now provides a new optional \geo\ feature that exposes spatial capabilities. When enabled, it registers DataFusion UDFs for geometric measurements (Area, Distance, Length) and spatial relationships (Contains, Intersects, etc.) via the \geodatafusion\ integration. It also introduces a new \bbox\ module containing a \BoundingBox\ struct and helper functions for computing and managing spatial bounds, which are conditionally compiled and exported only when the \geo\ feature is active.
rust/lance-geo · high confidence
Introduce pluggable cache backend architecture with URI configuration
The cache system now supports pluggable backends via a new \CacheBackend\ trait, allowing users to swap the underlying storage implementation. The module exposes a URI-based configuration system (\build\_from\_uri\) that parses strings like \moka://?capacity=1073741824\ into backend configurations, enabling consistent setup across Python, Java, and Rust bindings. Two built-in implementations are provided: \MokaCacheBackend\ for general use and \QuickCacheBackend\ for high-contention scenarios like session metadata caching. The change also introduces a versioned serialization envelope (\LCE1\) for cache entries, allowing persistent backends to store and retrieve data across restarts with backward-compatible decoding that treats invalid formats as cache misses.
rust/lance-core/src/cache · high confidence
Introduce protobuf schema guidelines and new wire-format definitions
The repository now includes a Protobuf Guidelines document (AGENTS.md) that establishes the change process, compatibility requirements, and schema design rules for protobuf schemas, including the requirement for PMC votes on persisted format changes. Additionally, new protobuf files have been added to define wire contracts and encoding specifications: ann.proto defines vector query parameters and execution plan serialization (including IVF sub-index and partition execution nodes); filtered\_read.proto defines serialization for filtered read options and plans; index.proto and index\_old.proto define vector index metadata and details; file.proto and file2.proto define file descriptors, schemas, and the v2 file format layout; encodings\_v2\_0.proto and encodings\_v2\_1.proto specify array and structural encodings for Lance file formats 2.0 and 2.1; and fragment\_metadata.proto defines data fragment, file, and overlay file metadata structures.
protos · high confidence
Introduce public datagen API and benchmarks for array generation
The \lance-datagen\ crate now exposes a public API for generating test data arrays, including support for step, fill, random, and null patterns across various primitive types (Int8/16/32/64, Float32/64) and binary/string variants. This change adds a new \generator\ module with traits and implementations for array generation, along with a comprehensive benchmark suite (\benches/array\_gen.rs\) to measure performance of these generators. Users can now programmatically generate synthetic datasets with controlled characteristics for testing and development purposes.
rust/lance-datagen · high confidence
Introduce structural encodings for Lance 2.1+ (blob, constant, dictionary, miniblock, sparse)
This change adds the core structural encoding implementations for the Lance 2.1+ format, introducing new files for blob, constant, dictionary, miniblock, and sparse encodings. Users can now read and write data using these new structural layouts, which include support for out-of-line blob storage, constant-value optimization, dictionary encoding with normalized null handling, configurable miniblock chunk sizes (via LANCE\_MINIBLOCK\_MAX\_VALUES), and automatic sparse structural page planning for nested data.
rust/lance-encoding/src/encodings/logical/primitive · high confidence
Introduce system indices for fragment reuse and MemWAL
The index library now supports two new system index types: FragmentReuse and MemWAL. FragmentReuse tracks stable partition mappings to enable deferred compaction and efficient row-remapping during data rewrites, while MemWAL provides a write-ahead log mechanism for index state. These are exposed as first-class index types with dedicated adapters, plugin registry entries, and progress monitoring, allowing the system to manage internal index metadata and write-ahead state transparently.
rust/lance-index/src · high confidence
Introduces Python type stubs for the lance package
This change adds comprehensive .pyi type stub files for the Python lance module, covering the main package, debug utilities, fragment handling (including DeletionFile, RowIdMeta, and RowIdSequence), optimization operations (Compaction, RewriteResult), schema definitions (LanceSchema, LanceField), tracing, and the Bitmap binding. These stubs provide static type information for the underlying Rust-backed Python API, enabling better IDE autocomplete, type checking, and documentation for users of the library.
python/python/lance/lance · high confidence
Introduces versioned protobuf schema for index cache serialization
Adds a new \cache.proto\ file that defines the serialization format for index cache entries, including headers for FTS posting lists (compressed and plain), positions, posting groups, B-tree indices, and IVF vector partitions. This schema establishes the structure for cache keys and data, supporting features like configurable posting block sizes, impact data, and various quantizer types (PQ, Flat, SQ) with specific distance metrics and storage codecs.
rust/lance-index/protos-cache · high confidence
Introduction of the internal lance-index-core crate
The \lance-index-core\ crate has been introduced as an internal library containing the core traits and types used to implement index plugins for Lance. This new module defines the foundational \Index\ trait, the \IndexType\ enumeration (covering scalar types like BTree, Bitmap, and MinHashLsh, as well as vector types like IVF-PQ), and the \BuiltinIndexType\ enum for scalar index configuration. It also provides the \MetricsCollector\ trait for coarse-grained ANN stage timing and I/O statistics, and introduces asynchronous batch row-ID remapping utilities to handle row translation during index loading. This crate is explicitly marked as internal and not intended for external usage.
(repo-wide) · high confidence
Introduction of the lance-arrow internal sub-crate
The \lance-arrow\ crate has been introduced as an internal sub-crate containing Apache Arrow extensions for Lance. It provides support for bfloat16 vectors (including zero-copy array construction and round-tripping), JSON/JSONB read/write capabilities, and blob v2 schema handling. The crate also includes utilities for deep-copying Arrow arrays and RecordBatches, zero-copy IPC stream serialization, and memory accounting for Arrow buffers.
rust/lance-arrow · high confidence
Introduction of the lance-namespace Rust crate with core namespace APIs and compatibility shims
The \rust/lance-namespace\ crate has been added to provide the core Rust client and trait definitions for Lance namespaces, including the \LanceNamespace\ interface for managing namespaces and tables, schema conversion utilities between Arrow and JSON representations, and a comprehensive set of fine-grained error codes. To ensure backward compatibility with older SDK builds (Java and Python) that send scalar values for merge-insert keys, a deserialization shim in \compat.rs\ automatically promotes single-column string keys to the expected list format, while the error module standardizes error handling across all Lance implementations.
rust/lance-namespace · high confidence
Java SDK API overhaul with new core classes and session support
The Java SDK has been significantly refactored to align with the Rust engine, introducing a comprehensive set of new core classes including Dataset, CommitBuilder, Fragment, Branch, and BlobFile. This update adds support for session management, cache backends, and branch/tag metadata, while deprecating legacy fragment operations in favor of the new Transaction-based CommitBuilder. Users will now interact with a more modern, consistent API that exposes detailed dataset metadata, index build progress, and native blob handling capabilities.
java/src/main/java/org/lance · high confidence
Java bindings for MemWAL write, scan, and search operations
This change introduces a new set of Java classes in the \org.lance.memwal\ package that expose the MemWAL (Memory Write-Ahead Log) functionality to Java applications. Users can now initialize MemWAL on a dataset with configurable sharding (bucket, identity, or unsharded) and index maintenance via \InitializeMemWalParams\. The API provides \ShardWriter\ for inserting and deleting rows with configurable backpressure and flush intervals, and \LsmScanner\ for scanning data across the base table, SSTables, and active MemTable with SQL filtering and projection. Additionally, \LsmPointLookupPlanner\ and \LsmVectorSearchPlanner\ enable efficient primary-key lookups and IVF-PQ vector KNN searches across all MemWAL levels, with support for pre-filtering and stale-row suppression. The bindings also expose detailed statistics for MemTables and writers, and handle native resource management through JNI.
java/src/main/java/org/lance/memwal · high confidence
Java bindings for Merge Insert operations
Added Java classes (MergeInsertParams, MergeInsertResult, MergeInsertStats) to support merge insert operations, including configuration for matched/unmatched row handling, write modes, and statistics reporting.
java/src/main/java/org/lance/merge · high confidence
Java cleanup API: new policy, explanation, and stats models
The Java cleanup module now exposes a structured API for configuring and inspecting dataset cleanup operations. Users can build cleanup policies via a builder that supports filtering by timestamp or version, targeting specific dataset versions, cleaning referenced branches, deleting unverified files, enforcing a delete rate limit, and erroring on tagged old versions. The API also returns a CleanupExplanation containing RemovalStats (bytes removed, counts of data/transaction/index/deletion files removed, and failed deletes) and a list of candidate files with their kind, size, and verification status, along with warnings and branch references.
java/src/main/java/org/lance/cleanup · high confidence
Java compaction API with source limits, exclusion, and binary copy modes
The Java SDK now exposes a distributed compaction entry point in the \org.lance.compaction\ package, providing \planCompaction\ and \commitCompaction\ methods that wrap the native implementation. Users can now configure compaction with new limits (\maxSourceRows\, \maxSourceBytes\, \maxSourceFragments\) to bound the scope of planning, exclude specific fragments via \excludedFragmentIds\, and choose a \CompactionMode\ (reencode, try\_binary\_copy, or force\_binary\_copy) to control how data is rewritten. The API also supports configuring binary copy batch sizes, materializing deletions, and specifying a data storage version, while ensuring thread safety by acquiring a dataset read lock during native calls.
java/src/main/java/org/lance/compaction · high confidence
Java index metadata and configuration model
The \org.lance.index\ package now provides the core Java data models and configuration builders for index operations. This includes \Index\ and \IndexDescription\ for representing index metadata (such as segment sizes, creation times, and covering fields), \IndexType\ and \DistanceType\ enums to define supported index algorithms and distance metrics, and \IndexOptions\/\IndexParams\ builders for configuring index creation. Additionally, \IndexBuildProgress\ exposes callbacks to track build stages, and \OptimizeOptions\ allows users to configure index optimization strategies like merging and retraining.
java/src/main/java/org/lance/index · high confidence
Java namespace client now supports dynamic per-request context and mTLS certificate reloading
The Java namespace client (DirectoryNamespace and RestNamespace) now accepts a DynamicContextProvider, allowing applications to inject per-request context such as rotating authentication headers before each operation. Additionally, the RestNamespace client automatically reloads rotated mTLS certificates based on the new tls.reload\_interval\_seconds configuration property, ensuring secure connections persist through certificate rotations without requiring a restart.
java/src/main/java/org/lance/namespace · high confidence
Java schema model supports unenforced keys and column alterations
The Java schema package now includes \ColumnAlteration\ for modifying dataset columns (renaming, nullability, type casting) and extends \LanceField\ and \LanceSchema\ to support unenforced primary and clustering keys with explicit ordering positions. This allows users to define composite key ordering and clustering strategies in Java, while \LanceField\ also handles \FixedSizeList\ logical types during Arrow schema conversion.
java/src/main/java/org/lance/schema · high confidence
Java transaction operation models added
Added Java classes for dataset transaction operations (Append, Clone, CreateIndex, DataOverlay, DataReplacement, Delete, KeyExistenceFilter, Merge, Overwrite, Project, ReserveFragments, Restore, Rewrite, SchemaOperation) to align with the Rust transaction model, enabling Java clients to construct and manage dataset mutations.
java/src/main/java/org/lance/operation · high confidence
New CI benchmark infrastructure for tracking IO and memory performance
A new benchmark suite has been added to the Python package to monitor performance regressions in CI. This includes a custom pytest fixture that measures IO statistics (IOPS, bytes read/written) and memory usage (peak bytes, total allocations) during dataset operations, with optional support for memory tracking via the \lance-memtest\ library. The suite provides utilities for generating test datasets (including TPC-H and basic 10M row sets), resolving dataset URIs for local or Google Cloud environments, and managing data overlay benchmarks. It also includes helpers for clearing OS file caches to ensure consistent benchmark conditions.
_python/python/ci\benchmarks · high confidence
New CI tooling for release automation, security, and governance
The CI directory now includes a comprehensive suite of scripts and checks to support a structured release process and stricter quality gates. Release management is handled by new shell scripts (approve\_rc.sh, create\_rc.sh, create\_release\_branch.sh, publish\_beta.sh) that automate the creation of release branches, release candidates (RCs), and stable promotions, including automatic version bumping and release note generation via generate\_release\_notes.py. Security is enhanced by check\_dependency\_age.py, which enforces a 48-hour minimum age for crates.io dependencies to mitigate supply-chain risks, with an allowlist in dependency-age-allowlist.toml. Governance is supported by format\_vote\_gate.py, which enforces PMC voting requirements for format-specification changes, and check\_breaking\_changes.py, which validates that minor versions are bumped when breaking changes are detected. Additional utilities include check\_proto\_comments.py for enforcing multi-line comment styles in .proto files, coverage.py for Rust code coverage analysis, and new\_contributors.py for tracking contributor statistics.
ci · high confidence
New CLI tools for FM index management and benchmarking
Added three new command-line utilities in the Rust binary directory: \fm\_index\_tool\ for listing, creating, and dropping FM (Fuzzy Matching) scalar indexes on datasets; \fm\_contains\_bench\ for benchmarking FM index query performance with configurable thread counts, caching, and output to JSON/CSV; and \lq\ for basic dataset inspection, querying, and vector index creation. These tools provide dedicated interfaces for managing and evaluating FM indexes, which are distinct from the existing vector index workflows.
rust/lance/src/bin · high confidence
New DataFusion integration for querying Lance namespaces via SQL
The \lance-namespace-datafusion\ crate introduces a bridge that allows Lance namespaces to be queried as native DataFusion catalogs, schemas, and tables. It exposes a \SessionBuilder\ to construct a DataFusion \SessionContext\ with dynamic \CatalogProvider\ and \SchemaProvider\ implementations backed by a \LanceNamespace\. This enables read-only SQL access to Lance datasets, mapping top-level namespaces to catalogs, child namespaces to schemas, and loading tables on-demand with caching.
rust/lance-namespace-datafusion · high confidence
New DatasetDelta API to inspect row-level changes between dataset versions
Added DatasetDelta and DatasetDeltaBuilder classes in the Java API, enabling users to compute and stream differences between two versions of a dataset. Users can specify a version range or compare against a specific version to retrieve inserted and updated rows as Arrow streams, list transactions, and obtain deleted row IDs. The implementation holds a read lock on the underlying dataset during native calls to ensure thread safety.
java/src/main/java/org/lance/delta · high confidence
New FragmentSession API for efficient repeated fragment reads
Users can now use the new FragmentSession API to perform repeated reads on a fragment without the overhead of opening a new reader each time. This session-based approach maintains internal state, such as sorted deletion vectors, allowing for more efficient row lookups and automatic handling of deleted rows across multiple operations. The implementation also ensures that JSON columns are correctly converted from the internal Lance format to standard Arrow JSON for user-facing outputs.
rust/lance/src/dataset/fragment · high confidence
New HD-Vila dataset benchmark with data generation script
Added a new benchmark entry for the HD-Vila-100M dataset, including a README with setup instructions and a \datagen.py\ script. The script enables users to generate the dataset in Lance format by downloading videos from YouTube, cutting specific clips, and processing them using Ray and PyArrow.
benchmarks/hd-vila · high confidence
New HNSW index implementation with online builder support
The HNSW vector index implementation has been replaced with a new version in the \lance-index\ crate. This update introduces a new \HNSWBuilder\ for offline graph construction and an \OnlineHnswBuilder\ that supports concurrent search during index building, which is particularly useful for in-memory MemTable indexes. The new implementation includes deterministic graph construction via a fixed random seed, configurable build parameters (such as \m\, \ef\_construction\, and \max\_level\), and support for both float and binary (Hamming distance) vectors. It also integrates with the existing IVF framework, allowing HNSW to be used as a sub-index within IVF partitions.
rust/lance-index/src/vector/hnsw · high confidence
New IVF index builder, shuffler, and transformer components
The IVF index module now includes dedicated builder, shuffler, storage, and transformer components. The builder introduces \IvfBuildParams\ with a new \target\_partition\_size\ option (replacing the deprecated \num\_partitions\), support for streaming k-means training via \streaming\_sample\_rate\ and \streaming\_coreset\_rate\, and the ability to load precomputed partition mappings or shuffle buffers. The shuffler handles disk-based partitioning of data streams, while the transformer computes partition IDs and distances for vectors. Storage is managed via the \IvfModel\ struct, which persists centroids, partition offsets, and lengths to Lance files.
rust/lance-index/src/vector/ivf · high confidence
New IVF index building and search implementation (v2)
The IVF vector index implementation in \rust/lance/src/index/vector/ivf\ has been replaced with a new v2 architecture. This change introduces a new builder (\builder.rs\) and I/O layer (\io.rs\) for constructing index partitions, a new v2 index definition (\v2.rs\) that supports streaming search with configurable batch sizes, and a new serialization format (\partition\_serde.rs\) for caching IVF partitions. Users will see this as the underlying engine for IVF index creation and search, enabling features like distributed segment builds, parallel partition search, and improved memory management through chunked global top-k scoring.
rust/lance/src/index/vector/ivf · high confidence
New JSON serialization support for Arrow Schema types
Added a new \json.rs\ module that enables serializing and deserializing Apache Arrow \DataType\ and \Field\ structures to and from JSON. This feature supports a wide range of data types including primitives, lists, structs, and fixed-size binary/list types, allowing users to easily convert schema definitions into a portable JSON format for storage or interchange.
rust/lance/src/arrow · high confidence
New Java API for vector index configuration and training
The Java API now exposes a comprehensive set of builder classes for configuring vector index parameters, including HnswBuildParams, IvfBuildParams, PQBuildParams, RQBuildParams, and SQBuildParams. These are unified in VectorIndexParams, which provides factory methods to create IVF-based index configurations (IVF Flat, IVF PQ, IVF RQ, and IVF HNSW with PQ/SQ) and enforces validation rules such as mutual exclusivity of quantizers. Additionally, the new VectorTrainer utility allows users to pre-train IVF centroids and PQ codebooks with explicit distance type support, ensuring the training geometry matches the index build geometry to prevent silent recall degradation.
java/src/main/java/org/lance/index/vector · high confidence
New Java IPC layer with async scanning, full-text search, and vector search controls
The \org.lance.ipc\ package introduces a new Java interface to the native (Rust) engine, providing both synchronous (\LanceScanner\) and non-blocking asynchronous (\AsyncScanner\) data scanning via \CompletableFuture\. This release adds support for full-text search through the \FullTextQuery\ API, vector search tuning with the \ApproxMode\ enum (FAST, NORMAL, ACCURATE) for RQ-quantized indexes, and query optimization via \MaterializationStyle\ (early/late column fetching) and \ColumnOrdering\. Additionally, \ScanStats\ exposes detailed performance metrics, including per-query index cache hit/miss counts, to help users monitor and tune scan efficiency.
java/src/main/java/org/lance/ipc · high confidence
New Java file I/O API with configurable blob reading and write options
The \org.lance.file\ package now provides dedicated \LanceFileReader\ and \LanceFileWriter\ classes for direct file-level operations. Users can control how blob-encoded columns are returned during reads via the new \BlobReadMode\ enum (materialized content or position/size descriptors) through \FileReadOptions\. The writer supports configuring buffering and page sizes via \FileWriteOptions\, allows specifying a data storage version, and enables attaching custom schema metadata that is flushed to the file footer on close.
java/src/main/java/org/lance/file · high confidence
New Java fragment metadata and result classes
Added new Java classes in the \org.lance.fragment\ package to support fragment-level metadata and update operations. \DataFile\ and \DeletionFile\ (with \DeletionFileType\) expose file paths, versions, and deletion details. \FragmentUpdateResult\ and \FragmentMergeResult\ provide structured outputs for column updates and merges, including row offset handling via JNI. \RowIdMeta\ and \VersionMeta\ wrap serialized Rust metadata for row IDs and version sequences, enabling stable row identification and version tracking in Java.
java/src/main/java/org/lance/fragment · high confidence
New LSM scanner execution nodes for MemTable scans, deduplication, and indexed searches
The MemTable scanner now includes dedicated DataFusion execution nodes to handle complex read patterns within the active MemTable. A new deduplication scan ensures that primary-key updates are resolved correctly by emitting only the newest version of each key, preventing stale rows from leaking through filters. Indexed searches are now supported directly on the MemTable via BTree, Full-Text Search (FTS), and Vector (HNSW) execution nodes, allowing scalar, text, and vector queries to leverage in-memory indexes with MVCC visibility. For vector searches where no HNSW index is available, a brute-force KNN executor provides exact distance calculations. All these nodes enforce visibility constraints, ensuring that only readable batches are scanned.
_rust/lance/src/dataset/mem\wal/memtable/scanner/exec · high confidence
New LSM scanner execution nodes for point lookups and cross-generation deduplication
The mem-wal scanner now includes a suite of new DataFusion execution nodes to optimize point lookups and handle multi-generation data. BloomFilterGuardExec skips generations that definitely do not contain a requested primary key, while CoalesceFirstExec short-circuits evaluation to return the newest matching row immediately. PkBlockFilterExec removes superseded rows by checking primary-key membership against newer generations, and FirstByPkExec deduplicates results (such as those from cross-column full-text search) by keeping only the first occurrence per primary key. MemtableGenTagExec annotates rows with their generation number for ordering, SchemaRelabelExec aligns storage and logical schemas, and ReconcileExec applies schema plans to batches.
_rust/lance/src/dataset/mem\wal/scanner/exec · high confidence
New Product Quantization implementation with 4-bit support and SIMD optimizations
The Product Quantization (PQ) vector index module has been rewritten to support 4-bit quantization alongside the existing 8-bit mode, enabling significantly higher compression ratios for vector data. The new implementation includes a dedicated builder for training PQ codebooks, a storage backend that persists PQ codes and row IDs, and a transformer for converting vector columns into PQ codes. Performance is improved through SIMD-optimized distance calculations, including AVX512-VBMI byte-plane lookups for 4-bit codes and pre-transposed distance tables to avoid runtime transposition overhead. The module also introduces a pairwise scorer for symmetric PQ code-to-code scoring, supporting L2, cosine, and dot distance metrics.
rust/lance-index/src/vector/pq · high confidence
New Python API for distributed index building and model checkpointing
The \lance.indices\ package introduces a new \IndicesBuilder\ class that allows advanced users to construct vector indices (IVF, PQ, IVF\_PQ, IVF\_SQ) in discrete, checkpointable steps rather than as a single monolithic operation. This enables distributed and segmented index builds on large datasets. The package also exposes \IvfModel\ and \PqModel\ classes with \save\ and \load\ methods that support \storage\_options\, allowing users to persist and resume training progress to cloud storage backends.
python/python/lance/indices · high confidence
New Python bindings for dataset internals and operations
This change introduces several new Python-level components in the \python/src/dataset\ module. It exposes a \BlobFile\ API (\LanceBlobFile\) allowing users to read blob data via methods like \readall\, \read\_range\, and \read\_ranges\. It adds Python bindings for dataset cleanup operations, including \CleanupStats\, \CleanupCandidateFile\, and \CleanupExplanation\ classes to inspect and report on cleanup results. A \PyCommitLock\ implementation is added to support custom Python-based commit handlers. Additionally, it exposes \IOStats\ for tracking read/write I/O metrics, \DataStatistics\ and \FieldStatistics\ for reporting data stats, and comprehensive Python bindings for the compaction optimization system, including \CompactionPlan\, \CompactionTask\, and \CompactionMetrics\.
python/src/dataset · high confidence
New Python type stubs for index training, transformation, and metadata APIs
The \lance.indices\ module now exposes Python type stubs (\\_\init\\_.pyi\) that define the public API for vector index operations. Users can now rely on static type checking for standalone training functions (\train\_ivf\_model\, \train\_pq\_model\), vector transformation (\transform\_vectors\), and residual quantization model building (\build\_rq\_model\). The stubs also formalize the structure of index metadata classes (\IndexConfig\, \IndexSegment\, \IndexDescription\, \IndexSegmentDescription\), including fields for segment UUIDs, fragment IDs, sizes, and covering fields, ensuring consistent type hints for index inspection and management.
python/python/lance/lance/indices · high confidence
New Rust examples for vector search, full-text search, and dataset I/O
Added several new Rust example programs in the \rust/examples\ directory to demonstrate core Lance capabilities. The \full\_text\_search\ example shows how to create and query an inverted index for text data. The \hnsw\ and \ivf\_hnsw\ examples provide benchmarks and usage patterns for HNSW and IVF-HNSW vector indices, including building indexes and performing nearest-neighbor searches. The \llm\_dataset\_creation\ example demonstrates downloading and tokenizing text data from Hugging Face for LLM training. Finally, \write\_read\_ds\ illustrates basic dataset creation, reading, and cleanup operations.
rust/examples · high confidence
New SIFT/GIST-1M benchmark suite with LanceDB integration
The benchmarks/sift directory now provides a complete, reproducible benchmarking workflow for the SIFT-1M and GIST-1M datasets using the LanceDB API. This includes Python scripts for data generation (datagen.py), ground truth creation (gt.py), index building (index.py), and performance metric calculation (metrics.py), alongside a Jupyter notebook for result analysis. The suite supports multiple vector data types (f32, f16, bf16) and index configurations (IVF-PQ, DiskANN), with pre-computed statistics files (lance\_sift1m\_stats.csv, lance\_gist1m\_stats.csv) demonstrating recall and latency trade-offs across various partition and probe settings.
benchmarks/sift · high confidence
New Scalar Quantization (SQ) index implementation
This change introduces the core components for a new Scalar Quantization vector index within the \lance-index\ crate. It adds an SQ builder (\builder.rs\) to configure quantization parameters, a transformer (\transform.rs\) to convert input vectors into 8-bit codes, a storage layer (\storage.rs\) to manage the quantized data chunks, and a pairwise scorer (\pairwise.rs\) that performs exact integer-based distance calculations for L2, Cosine, and Dot metrics. This provides a new, efficient indexing option for users requiring lower memory footprint and faster search speeds through scalar quantization.
rust/lance-index/src/vector/sq · high confidence
New \`lance-test-macros\` crate for test-scoped tracing integration
A new \lance-test-macros\ crate has been added, providing a \\#\[test\]\ attribute macro that wraps existing test functions to automatically configure the \tracing\ library. When the \LANCE\_TRACING\ environment variable is set (e.g., to \debug\), tests wrapped with this macro initialize a tracing subscriber with a Chrome layer, generating \.json\ trace files compatible with Chrome DevTools or Perfetto. This allows developers to easily capture detailed performance and execution traces for individual tests without manual setup.
rust/lance-test-macros · high confidence
New arrow-stats crate for computing column statistics
Added the \lance-arrow-stats\ crate, which provides a \StatisticsAccumulator\ to compute min, max, null count, NaN count, and buffer memory usage for Apache Arrow arrays. This new capability supports numeric, temporal, boolean, string, binary, and list types, enabling more efficient predicate pushdown and query planning in Lance's columnar storage layer by tracking page-level statistics.
rust/arrow-stats · high confidence
New benchmark examples for ACORN-1 HNSW traversal and hierarchical k-means quality
Added three new example programs in the lance-index crate to evaluate specific index features. \acorn\_bench.rs\ and \acorn\_bench\_sift.rs\ benchmark the new ACORN-1 mask-aware HNSW traversal against the standard basic traversal and flat scans, measuring latency and recall under various filter selectivities and mask shapes (using synthetic data and the SIFT1M dataset respectively). \kmeans\_quality.rs\ measures the training time and clustering quality (WCSS, partition balance) of the new proportional hierarchical k-means algorithm for IVF training on large \.fbin\ datasets.
rust/lance-index/examples · high confidence
New benchmarking tool and credential vending infrastructure for Directory Namespace
This release introduces a copy-on-write \\_\_manifest\ commit benchmark (\manifest\_bench\) to measure Directory Namespace throughput under continuous and concurrent workloads, alongside a new \ConnectBuilder\ API for configuring namespace connections. It also adds a comprehensive credential vending system that automatically generates scoped, temporary credentials for AWS (STS AssumeRole), Azure (SAS tokens), and GCP, with built-in caching to reduce vendor call frequency and dynamic context providers for per-request header injection in the REST namespace.
rust/lance-namespace-impls · high confidence
New binary copy compaction and index remapping infrastructure
This change introduces a new binary copy mechanism for compaction in \rust/lance/src/dataset/optimize/binary\_copy.rs\, which merges small Lance files into larger ones by performing page-level binary copies to preserve stable row IDs and reduce I/O overhead. It also adds a new \remapping.rs\ module that provides the \IndexRemapper\ trait and utilities for remapping row addresses and index metadata during compaction, ensuring that indices remain valid when fragment row IDs change.
rust/lance/src/dataset/optimize · high confidence
New core utility modules for row addressing, rate limiting, and retry logic
The \rust/lance-core/src/utils\ directory now includes several new foundational modules that support internal operations. \address.rs\ introduces the \RowAddress\ type, which encodes fragment IDs and row offsets into a single 64-bit value to replace the previous \row\_id\/\row\_addr\ distinction. \aimd.rs\ provides an Additive Increase / Multiplicative Decrease (AIMD) rate controller for dynamically adjusting request rates, including validation to reject non-finite configuration values. \backoff.rs\ adds configurable exponential backoff and \SlotBackoff\ strategies to help spread out concurrent retries. Additional utilities include \assume.rs\ for release-checked invariants, \bit.rs\ for bit manipulation and alignment, \blob.rs\ for generating obfuscated blob sidecar paths, and \futures.rs\ for sharing streams between consumers.
rust/lance-core/src/utils · high confidence
New distributed index merging infrastructure
The distributed vector indexing subsystem now includes a new \index\_merger\ module and shared \partition\_merger\ helpers. This introduces the core logic for merging distributed index segments, including support for detecting various IVF index types (Flat, PQ, SQ, RQ, and HNSW variants) and writing unified metadata. It also adds strict and tolerant equality checks for vector data to ensure consistency during the merge process.
rust/lance-index/src/vector/distributed · high confidence
New encoding utilities for data accumulation and byte-packed integer encoding
The \lance-encoding\ crate now includes new utility modules for handling data buffering and efficient integer encoding. An \AccumulationQueue\ has been added to buffer Arrow arrays until a specified byte threshold is reached, allowing for optimized flushing of column data during encoding. Additionally, a \BytepackedIntegerEncoder\ provides a simple, fast way to encode integers by automatically selecting the smallest suitable byte width (u8, u16, u32, or u64) based on the maximum value, rejecting values that exceed the selected width to prevent overflow.
rust/lance-encoding/src/utils · high confidence
New lance-datafusion crate with DataFusion integration
A new \lance-datafusion\ crate has been introduced to provide a bridge between Lance and Apache DataFusion. This includes a build script for generating protobuf bindings, an \Aggregate\ struct for handling group-by and aggregate expressions, a \BatchReaderChunker\ for streaming record batches, and a \DataFrameExt\ trait that adds a \group\_by\_stream\ method to DataFusion DataFrames. The crate also implements core execution utilities like \OneShotExec\ for wrapping streams into execution plans, expression coercion logic in \expr.rs\ and \logical\_expr.rs\ to handle type conversions and filter simplification, and a \ProjectionBuilder\ to manage column selection and system column handling during query planning.
rust/lance-datafusion/src · high confidence
New lance-testing crate for internal test utilities
A new \lance-testing\ crate has been added to the repository, providing internal-only utilities for unit tests and benchmarks. It includes a \datagen\ module with generators for creating test data (such as incrementing integers and random vectors), a \progress\ module offering a macro to define in-memory progress recorders for testing distributed operations, and a \pprof\ module (Linux-only) that integrates with Criterion to generate flamegraph profiles during benchmarks.
rust/lance-testing · high confidence
New logical array encoders for binary, blob, list, primitive, and struct types
The logical encoding layer in \rust/lance-encoding/src/array\_encoding/logical\ now includes dedicated schedulers and decoders for Binary, Blob, List, Primitive, and Struct data types. These new components handle the scheduling and decoding of their respective data layouts, enabling more efficient handling of variable-width data, large binary blobs, nested lists, and structured records within the Lance encoding format.
_rust/lance-encoding/src/array\encoding/logical · high confidence
New logical encoding layer for complex data types
The logical encoding module now includes dedicated structural encoders, schedulers, and decoders for Blob, FixedSizeList, List, Map, and Struct types. This adds support for storing large binary data in external buffers (Blob), correctly handling nested structures like lists of structs (FixedSizeList), and managing repetition/definition levels for variable-length lists and maps. These components form the foundation for encoding complex nested data in the Lance format.
rust/lance-encoding/src/encodings/logical · high confidence
New modular write API with staged transactions and retry support
The write operations (insert, delete, update, merge insert) now use a builder pattern that separates data preparation from the final commit. Users can stage changes using methods like \execute\_uncommitted\ to obtain a \Transaction\, which is then committed via the new \CommitBuilder\. This architecture introduces configurable retry logic for concurrent write conflicts, allowing users to set \conflict\_retries\ and \retry\_timeout\ on delete and update operations, and exposes a \write\_progress\ callback on inserts to track throughput. The \CommitBuilder\ also centralizes configuration for storage formats, object stores, and session caching.
rust/lance/src/dataset/write · high confidence
New physical encoding implementations for Lance 2.1
The \rust/lance-encoding/src/encodings/physical\ directory now contains the core physical encoding implementations for the Lance 2.1 format. This includes \binary.rs\ for variable-width data, \bitpacking.rs\ for fixed-width bitpacking, \block.rs\ for general block compression (LZ4, Zstd), \byte\_stream\_split.rs\ for floating-point optimization, \constant.rs\ for constant/null handling, \fsst.rs\ for string compression, \general.rs\ for wrapping inner compressors with general compression, \packed.rs\ for struct packing, and \rle.rs\ for run-length encoding. These files provide the leaf-level encoders and decompressors that the logical layer uses to store data.
rust/lance-encoding/src/encodings/physical · high confidence
New quickstart and YouTube Q&A bot notebooks added
Added a new quickstart notebook demonstrating how to create and write Lance datasets using PyArrow and Pandas, including conversion from Parquet. Also added a new notebook for building a question-and-answer bot that searches YouTube transcripts using natural language, leveraging OpenAI embeddings and the Hugging Face datasets library.
notebooks · high confidence
New utility classes for JSON handling and range definitions
Added three new utility classes to the org.lance.util package: JsonFields provides helpers to create Arrow fields annotated with the 'arrow.json' extension for seamless JSON text handling; JsonUtils offers static methods to serialize and deserialize JSON using Jackson; and Range defines a simple immutable value object for representing integer ranges with inclusive start and exclusive end boundaries.
java/src/main/java/org/lance/util · high confidence
New utility modules for async primitives, time mocking, and test infrastructure
The \rust/lance/src/utils\ directory now includes three new modules: \future.rs\ introduces a \SharedPrerequisite\ struct that allows spawning an asynchronous background task and sharing its result (or error) across multiple threads, with synchronous access once the task completes; \temporal.rs\ provides a \utc\_now()\ function that abstracts the current system time, enabling unit tests to mock time via \mock\_instant\ while production code uses the real system clock; and \test.rs\ adds a \TestDatasetGenerator\ capable of creating datasets with randomized, 'hostile' column layouts (splitting fields across files, shuffling field IDs, and varying field orders) to improve test coverage for dataset operations under diverse storage configurations.
rust/lance/src/utils · high confidence
New v3 vector index architecture with shuffler and sub-index traits
The v3 vector index implementation introduces a new internal architecture for building and searching IVF-based indexes. This location adds the core shuffler component (shuffle\_bench.rs, shuffler.rs) which handles partitioning data streams into IVF partitions, and the sub-index trait definitions (subindex.rs) that standardize how sub-indexes like Flat and HNSW are built, searched, and remapped within the v3 format. Users benefit from the underlying performance and structural improvements of the v3 index format, including configurable CPU thread usage and optimized shuffle operations, as this code forms the foundational I/O and partitioning layer for the new index version.
rust/lance-index/src/vector/v3 · high confidence
New vector index details serialization and bounded partition stream for IVF builds
The vector index module now includes a new \details\ module that serializes and deserializes \VectorIndexDetails\ (including compression types like PQ, SQ, and RQ, and runtime hints) for index metadata, and a new \bounded\_partition\_stream\ module that manages memory-bounded, ordered execution of IVF partition builds. These changes support richer index introspection and more controlled resource usage during index construction.
rust/lance/src/index/vector · high confidence
New vector index performance benchmark report
Added a new benchmark suite in the \benchmarks/full\_report\ directory to evaluate vector index performance. This includes a Python library (\\_lib.py\) for generating and processing vector data (specifically using the NYT dataset with TF-IDF and random projection), and a Jupyter notebook (\report.ipnb\) that runs recall tests against Lance IVF\_PQ indexes. The benchmark measures recall rates across different \nprobes\ and \refine\_factor\ settings for both in-sample and out-of-sample queries, visualizing the results as heatmaps.
_benchmarks/full\report · high confidence
Optimized execution paths for delete-only and in-place merge insert operations
The merge insert engine now uses specialized execution nodes to improve performance for specific operation types. A new delete-only path (\DeleteOnlyMergeInsertExec\) handles \WhenMatched::Delete\ scenarios by reading only row addresses and action columns, skipping the write step entirely to efficiently mark existing rows as deleted. Additionally, an in-place update path (\InPlaceMergeInsertExec\) allows narrow updates to existing fragments by patching only the source columns into existing data files and tombstoning old versions, avoiding the overhead of rewriting entire rows. These changes optimize bulk deletes and partial updates without creating new fragments.
_rust/lance/src/dataset/write/merge\insert/exec · high confidence
Restructure lance crate into modular sub-modules
The \rust/lance/src\ directory has been refactored from a flat structure into distinct sub-modules (\arrow\, \blob\, \datafusion\, \dataset\, \index\, \io\, \session\, \table\, \utils\). This change introduces new public APIs for blob v2 storage (including \BlobFieldOptions\ and dedicated writers), exposes DataFusion integration via \LanceTableProvider\, and adds object store metrics support (documented in \metrics.md\). The \lib.rs\ entry point now re-exports these modules and their specific capabilities, such as the new blob field helpers and the DataFusion table provider, while maintaining backward compatibility for core dataset operations.
rust/lance/src · high confidence
Segmented scalar index merge support
Scalar index implementations (Bitmap, BloomFilter, BTree, FM-Index, Inverted, LabelList, MinHash LSH, NGram, RTree, and ZoneMap) now support merging multiple index segments into a single consolidated segment. This enables distributed index builds and compaction workflows to combine partial results without rebuilding from raw data, improving performance and scalability for large datasets.
rust/lance/src/index/scalar · high confidence
Support for dynamic, auto-refreshing storage credentials
Lance now supports dynamic storage options and credentials that are fetched at runtime and automatically refreshed before expiration. This is enabled by the new \StorageOptionsProvider\ trait and \StorageOptionsAccessor\, which allow integration with external sources like LanceNamespace to retrieve temporary credentials (e.g., AWS, Azure, GCP) or other configuration. The \NamespaceCredentialsProvider\ and \DynamicOpenDalStore\ use these options to build and cache object stores, ensuring that short-lived credentials are renewed seamlessly without requiring application restarts or manual intervention.
_rust/lance-io/src/object\store · high confidence
Unified object store providers with OpenDAL and dynamic credentials
The object store providers in \rust/lance-io/src/object\_store/providers\ have been refactored to implement a common \ObjectStoreProvider\ trait, standardizing how storage backends are initialized. This change introduces native support for Hugging Face (\hf://\), GooseFS (\goosefs://\), and Alibaba Cloud OSS (\oss://\), while extending Azure support to include the \abfss://\ scheme for ADLS Gen2 via OpenDAL. AWS, GCP, and Azure providers now utilize OpenDAL as an alternative backend (controlled by configuration) and support dynamic credential vending, allowing credentials to be refreshed without recreating the store. Additionally, a new \shared-memory://\ scheme enables process-wide in-memory state sharing across components.
_rust/lance-io/src/object\store/providers · high confidence
Vendored tokenizer stack with Lindera and Jieba support
The tokenizer implementation for full-text search has been moved into the lance-index crate, introducing dedicated tokenizers for Japanese and Korean text via Lindera and Chinese text via Jieba. This change adds new configuration handling for these language models and implements specific tokenization logic for JSON and text document types within the inverted index.
rust/lance-index/src/scalar/inverted/tokenizer · high confidence
Architecture
Extracted row-selection primitives into a new lance-select crate
The row-selection logic previously embedded in lance-core and lance-index has been extracted into a new, standalone lance-select crate. This crate provides the core mask types (RowAddrMask, NullableRowAddrMask) and index expression result wrappers (IndexExprResult, NullableIndexExprResult) that define which rows survive a filter, along with benchmarks to track their performance. This change decouples these primitives from larger dependencies, allowing downstream filtering code and benchmarks to depend on the mask substrate without pulling in the full lance-core or lance-index crates.
rust/lance-select · high confidence
Introduce lance-core crate with core data types, error handling, and blob v2 schema definitions
This change introduces the \lance-core\ crate, consolidating foundational components previously scattered across the codebase. It defines the core data types and schema structures, including the new Blob v2 logical and descriptor field layouts, and provides a comprehensive error handling system with Levenshtein-based suggestions for missing fields. Additionally, it includes a custom \DeepSizeOf\ trait for accurate Arrow-aware memory accounting, replacing the previous \deepsize\ crate, and exposes system column metadata and dataset traits to decouple row-taking logic from the main dataset implementation.
rust/lance-core/src · high confidence
Migrate schema and data types to lance-core
The core schema and data type definitions (Field and Schema structs) have been moved into the lance-core crate. This refactoring centralizes the schema logic, introducing support for unenforced primary and clustering keys with explicit ordering, and adds validation for FixedSizeList dimensions to prevent runtime panics on zero-dimension lists.
rust/lance-core/src/datatypes · high confidence
Refactor encoding mechanisms to be version-free
The encoding implementation has been refactored to decouple the core encoding logic from specific file format versions. The \array\_encoding\ module now exposes a reusable \ArrayFieldEncodingStrategy\ that implements the persisted \ArrayEncoding\ grammar, while the \encoder/structural\ module provides version-free structural field encoder builders. This change allows the same encoding mechanisms to be composed and accepted across different file versions, with version-specific decisions handled separately.
rust/lance-encoding/src · high confidence
Behavioural changes
Centralized dataset version policies and write validation
The dataset write logic now enforces strict version policies through a new \versions\ module, preventing the mixing of V1 and V2 data files within the same dataset. This change introduces version-specific behaviors for schema comparison, scan stream creation, and blob handling (including legacy vs. V2 blob validation), ensuring that write operations target the correct file format version and that schema evolution rules are applied consistently across different Lance file versions.
rust/lance/src/dataset/versions · high confidence
Input validation and structural writer refactoring in file writing
The file writer now validates that input arrays match the file schema's Arrow types and extensions before encoding, rejecting mismatches with clear error messages to prevent data corruption. Additionally, the underlying writing logic has been refactored to use a new \StructuralFileSink\ for managing page metadata and spilling, which supports the exact current-format writers and improves memory management during large writes.
rust/lance-file/src/writer · high confidence
Introduce Fragment Reuse Index for deferred index remapping
The index subsystem now uses a Fragment Reuse Index (FRI) to track row-address translations during compaction and deletions, allowing index remapping to be deferred until search time. This change introduces a new \IndexSegment\ structure to represent physical index segments with their own provenance (UUID, dataset version, and fragment coverage) and adds a \QueryRowIdRemapper\ that applies FRI history to translate row IDs for both scalar and vector index queries. Legacy index formats that cannot support this batch remapping are automatically excluded from coverage to ensure correct search results.
rust/lance/src/index · high confidence
Introduce internal sharded batch iterator and caching utilities
The internal dataset module now includes a ShardedBatchIterator class that enables reading Lance datasets in parallel across distributed PyTorch workers by sharding data at either the fragment or batch level, along with a CachedDataset helper for streaming data to temporary Arrow files. The ShardedBatchIterator is marked as deprecated in favor of the new Sampler API, and the module exposes these internal APIs for use in PyTorch data loading workflows.
_python/python/lance/\dataset · high confidence
Introduces memory-bounded manifest scanning and typed file schemas
The dataset file handling now uses a shared, memory-bounded manifest walker that lists manifest locations and reads manifests with a 1 GB in-flight memory budget, preventing unbounded memory usage during scans. This infrastructure supports new, strongly-typed Arrow schemas for file metadata: tracked files are now represented with a dictionary-encoded type field (mapping discriminants for manifest, data, deletion, transaction, and index files) to ensure consistent labeling, while all files include size and last-modified timestamp fields. These changes provide a more robust and performant foundation for dataset introspection and file tracking.
rust/lance/src/dataset/files · high confidence
Inverted index rewritten with new format, caching, and search logic
The scalar inverted index implementation has been completely refactored into a new modular structure (cache, doc\_set, flat\_search, format, inverted\_index, partition, posting builders). This introduces a new default FTS format version (V2) with varint-delta posting tails and shared position streams, replacing the legacy Arrow-based storage. The change adds a new slice-aware posting list cache to reduce memory overhead, a new Unicode-aware Levenshtein automaton for fuzzy matching, and a flat search path for unindexed data. Users will see changes in index file formats and potentially different memory usage patterns, with existing V1 indexes being automatically migrated to the new format on update.
rust/lance-index/src/scalar/inverted/index · high confidence
Merge insert now supports conditional updates, row failures, and column-level rewrite modes
The merge insert operation has been significantly enhanced with new behavioral options and execution paths. Users can now specify \WhenMatched::Fail\ to abort the operation if a matching row is found, and \WhenMatched::UpdateIf\ to conditionally update rows based on a provided expression. The system also supports \WhenNotMatchedBySource::DeleteIf\ to conditionally delete target rows that have no corresponding source row. Additionally, the write path now distinguishes between \RewriteRows\ (deleting and re-inserting full rows) and \RewriteColumns\ (attaching new data files for source columns while tombstoning old versions), allowing for more efficient updates when only specific columns change. These changes are implemented via a new action-based logical plan node and sentinel column logic to ensure NULL-safe row detection.
_rust/lance/src/dataset/write/merge\insert · high confidence
New commit retry timeout configuration and refactored I/O module structure
Users can now configure the commit conflict-retry timeout via the LANCE\_COMMIT\_RETRY\_TIMEOUT\_SECS environment variable, allowing long-running maintenance operations to succeed under high contention without hitting the default 30-second limit. Additionally, the I/O module has been restructured: commit handling logic is now centralized in a new commit.rs file with configurable backoff, deletion file reading is extracted into its own module with caching, and execution nodes are reorganized under io/exec.rs to expose updated scan, filter, and KNN execution interfaces.
rust/lance/src/io · high confidence
New conflict resolution logic for tagged fragment reuse indices
The commit handler now includes a comprehensive conflict resolver that correctly handles concurrent operations on tables with tagged fragment reuse indices. This ensures that mutations like deletes, updates, and merge inserts do not break the index's provenance tracking, and that the index coverage is correctly derived from live fragments rather than stale stored provenance.
rust/lance/src/io/commit · high confidence
New internal graph builder and I/O components for HNSW index construction
The \rust/lance-index/src/vector/graph\ module now includes a new \builder.rs\ file defining the \GraphBuilderNode\ struct, which manages neighbor relationships and ranking during the construction of HNSW graphs, and an \io.rs\ file stub for graph input/output operations. These changes provide the internal infrastructure required to build and persist HNSW graph structures, supporting the broader HNSW implementation updates in the vector index.
rust/lance-index/src/vector/graph · medium confidence
Python SDK development environment and documentation overhaul
The Python SDK now mandates \uv\ for all local environment setup, dependency management, and command execution (replacing \pip\, \make test\, etc.), with a new \AGENTS.md\ guide detailing this workflow. A comprehensive \DEVELOPMENT.md\ document has been added to cover building the Rust extension, running tests, and profiling benchmarks, while \CONTRIBUTING.md\ and \README.md\ have been updated to reflect these changes. Additionally, the \Makefile\ has been rewritten to support \uv\-based linting, formatting, and testing, and third-party license files for both Python and Rust dependencies have been added.
python · high confidence
Refactored Lance v1 file I/O into a modular, self-contained implementation
The v1 file format implementation has been reorganized into a dedicated module (\rust/lance-file/src/versions/v1\) with a clear separation of concerns. The new structure includes specific encoding modules for binary, dictionary, and plain data types, a format layer for metadata and page tables, and distinct reader and writer components. This refactoring establishes v1 as the canonical legacy file owner and composes exact current-format readers, ensuring that the v1 wire layout and schema dictionary payloads are handled by these dedicated codecs without leaking into the version-free I/O layer.
rust/lance-file/src/versions/v1 · high confidence
Spill row lineage to data files and validate stable row IDs
Users can now enable the \lance.row\_lineage.spill\ table configuration to move large row lineage sequences (row IDs and version metadata) out of the manifest and into hidden columns of the fragment's data files, preventing manifest bloat during compaction and updates. This change also adds integrity checks in \Dataset::validate\ to ensure stable row IDs remain unique and consistent with physical row counts across fragments.
rust/lance/src/dataset/rowids · high confidence
Stricter validation and error handling for Lance v2.0 encoding
The array encoding module now enforces stricter invariants for the v2.0 format. It rejects lossy null structures in structs and returns explicit errors for unsupported v2.0 field types, ensuring data integrity during encoding. Additionally, the module introduces new logical and physical encoding strategies (including packed struct and basic page schedulers) to support these validation rules and improve the robustness of the encoding pipeline.
_rust/lance-encoding/src/array\encoding · high confidence
Vendored internal tokenizer stack into lance-tokenizer
The \rust/lance-tokenizer\ crate now contains a self-contained implementation of the tokenizer stack, including the core \TextAnalyzer\ API, multiple tokenizer implementations (ICU, Jieba, Lindera, Code, N-gram, Simple, Raw), and a suite of token filters (ASCII folding, lowercasing, stemming, stop-word removal, and more). This change consolidates the tokenization logic previously scattered across the codebase into a single, internal library used by Lance, ensuring consistent text processing for full-text search without relying on external Tantivy components.
rust/lance-tokenizer · high confidence
Fixes
1 commit (1 fix) fixing test\_image\_dataset
A fix in test\_image\_dataset — 1 commit (1 fix), 1 file.
_test\_image\dataset · low confidence · unverified
Test coverage
Add benchmarks for v2 file reader performance and schema reconstruction; Added JNI test helper for FFI validation; Added Java SDK test coverage for async scanning, cleanup, compaction, and commit operations; Added Java SDK tests for dataset operations; Added Java tests for scalar/vector index creation, distributed indexing, and zonemap statistics; Added MemTable vs RocksDB KV point-lookup benchmark; Added benchmarks and tests for cache key performance; Added benchmarks for I/O scheduler performance; Added benchmarks for manifest interning, row ID indexing, and system column streaming; Added historical test fixtures for backward compatibility; Added integration and regression tests for memory management and durability; Added integration tests for GCS, GooseFS, and TOS object stores; Added integration tests for Java MemWAL bindings; Added integration tests for SQL aggregate pushdown optimization; Added integration tests for query execution across data types and index strategies; Added memory usage tests for vector indexing and data writing; Added test coverage for PyTorch integration utilities; Added test fixtures for jieba and lindera tokenization models; Added test resource for Jieba language model; Added test to verify panic location in distance functions; Added test utilities for covering indexes, storage failure injection, and cache serialization; Added tests for InvertedIndexParams serialization and validation; Added tests for Java Full-Text Search API and Scanner integration; Added tests for Java namespace implementations and dynamic context support; Added tests for MemWAL future depth and cross-generation fencing; Added tests for OpenTelemetry metrics bridge; Added tests for binary copy compaction behavior; Added tests for fragment update results and metadata classes; Expanded Python test coverage for query coercion, empty indices, and object store registry; Expanded test coverage for dataset operations and features; New CI benchmark suite for Python API performance; New FineWeb FTS benchmark suite for MemWAL; New LSM vector search benchmarks for MemWAL; New MemTable read performance benchmark; New MemWAL HNSW parity and recall benchmarks; New MemWAL write and replay benchmarks; New Python benchmark suite for data I/O, indexing, and vector operations; New benchmark suite for lance-linalg distance and utility kernels; New benchmarking suite for Lance encoding performance; New benchmarks for concurrent appends, COUNT pushdown, and distributed vector builds; New benchmarks for scalar and vector index performance; New encoding module structure and fuzz testing infrastructure; New library compatibility test suite for cross-version validation.
Dependencies
Add new benchmarking and testing utility packages
Introduces several new packages to support development and testing workflows: a \lance-memtest\ Rust/Python crate for memory allocation testing utilities, and multiple new benchmark projects (including \dbpedia-openai\, \bigann\, \hd-vila\, \sift\, \wiki\, and \full\_report\) with their respective Python dependencies and configurations. Additionally, adds new Rust crates for Arrow scalar support (\lance-arrow-scalar\), Arrow statistics (\lance-arrow-stats\), and various compression algorithms (\lance-bitpacking\, \fsst\), alongside updated example and core library manifests.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 69 → 70 (+1.0)
- Rubric changed (rubric-2026.09.9 → rubric-2026.09.18) — scores are not directly comparable.
Lenses
- Code Health 82 → 82 (-0.0)
- Architecture 97 → 96 (-1.4)
- Maturity 76 → 78 (+2.2)
- Readiness 93 → 75 (-18.5)
- Security 59 → 62 (+3.5)
- Accessibility 72 → 75 (+2.8)
- Performance 82 (new)
Resolved (152)
- BitmapIndexPlugin::streaming_build_and_write (cognitive 16) (rust/lance-index/src/scalar/bitmap.rs)
- Change coupling: feature_flags.rs ↔ transaction.rs (rust/lance-table/src/feature_flags.rs)
- ClassTooLong: BitmapIndexPlugin (rust/lance-index/src/scalar/bitmap.rs)
- ComplexAllNullScheduler::initialize (cognitive 16) (rust/lance-encoding/src/encodings/logical/primitive.rs)
- Dataset::open_vector_index (cognitive 17) (rust/lance/src/index.rs)
- Dataset::open_vector_index (cyclomatic 20) (rust/lance/src/index.rs)
- Duplicated block (10 lines × 2) (rust/lance-index/src/vector/v3/shuffler.rs)
- Duplicated block (10 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
- Duplicated block (10 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/planner.rs)
- Duplicated block (10 lines × 3) (rust/lance-encoding/src/decoder.rs)
- Duplicated block (11 lines × 2) (rust/lance-index/src/scalar/bloomfilter.rs)
- Duplicated block (11 lines × 2) (rust/lance/src/dataset/mem_wal/memtable/scanner/builder.rs)
- Duplicated block (11 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/point_lookup.rs)
- Duplicated block (11 lines × 2) (rust/lance/src/dataset/scanner.rs)
- Duplicated block (11 lines × 3) (rust/lance-linalg/src/distance/cosine.rs)
- Duplicated block (12 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
- Duplicated block (12 lines × 2) (rust/lance/src/dataset/mem_wal/scanner/planner.rs)
- Duplicated block (12 lines × 4) (rust/lance/src/dataset/fragment.rs)
- Duplicated block (12 lines × 4) (rust/lance/src/io/commit/conflict_resolver.rs)
- Duplicated block (12–13 lines × 2) (rust/lance/src/dataset/write/merge_insert.rs)
- …and 132 more
New (330)
- Ambiguous distinction between update_columns and update_columns_with_offsets. Without documentation, it is unclear if the latter is a performance optimization, a different join strategy, or a bug-prone variant. They appear to do the same high-level operation (updating columns via join) but with a signature difference that isn't self-explanatory.
- ClassTooLong: CompoundQueryExec (rust/lance/src/io/exec/fts.rs)
- ClassTooLong: InvertedIndexParams (rust/lance-index/src/scalar/inverted/tokenizer.rs)
- ClassTooLong: LsmFtsSearchPlanner (rust/lance/src/dataset/mem_wal/scanner/fts_search.rs)
- ClassTooLong: Manifest (rust/lance-table/src/format/manifest.rs)
- ClassTooLong: MinHashLshIndex (rust/lance-index/src/scalar/minhash_lsh/index.rs)
- ClassTooLong: U64Segment (rust/lance-table/src/rowids/segment.rs)
- Critical CVE: [GHSA redacted] (java/pom.xml)
- Dataset::open_vector_index_from_metadata (cognitive 20) (rust/lance/src/index.rs)
- Dataset::open_vector_index_from_metadata (cyclomatic 22) (rust/lance/src/index.rs)
- Documentation: no project overview (rust/arrow-stats/README.md)
- Duplicated block (10 lines × 2) (java/src/main/java/org/lance/fragment/RowIdMeta.java)
- Duplicated block (10 lines × 2) (rust/lance-index/src/scalar/inverted.rs)
- Duplicated block (10 lines × 2) (rust/lance-index/src/scalar/label_list.rs)
- Duplicated block (10 lines × 2) (rust/lance-index/src/vector/bq/pairwise.rs)
- Duplicated block (10 lines × 2) (rust/lance-linalg/src/simd/f32.rs)
- Duplicated block (10 lines × 2) (rust/lance-table/src/transaction/manifest_build.rs)
- Duplicated block (10 lines × 2) (rust/lance/src/index/frag_reuse.rs)
- Duplicated block (10 lines × 3) (rust/lance-encoding/src/decoder.rs)
- Duplicated block (10–11 lines × 2) (java/src/main/java/org/lance/OpenDatasetBuilder.java)
- …and 310 more
Changes since last survey
- 204 commits — 131 feature/other, 73 fixes
By area
- rust/lance — 70 commits
- rust/lance-index — 29 commits
- (root) — 22 commits
- python/python — 17 commits
- rust/lance-table — 12 commits
- rust/lance-encoding — 11 commits
- docs/src — 10 commits
- rust/lance-linalg — 9 commits
- rust/lance-io — 5 commits
- .github/workflows — 4 commits
- java/src — 4 commits
- rust/lance-namespace-impls — 3 commits
- rust/lance-core — 2 commits
- rust/lance-datafusion — 2 commits
- benchmarks/auto-ivf-dot — 1 commit
- java/lance-jni — 1 commit
- java/pom.xml — 1 commit
- test_data/v8.0.0 — 1 commit
Notable commits
- fix: fix!: derive merge_columns field ids from the dataset manifest (#9547)
- fix: fix(core): parse fixed-offset timezones in timestamp logical types (#8950)
- fix: fix(datafusion): coerce numeric literals to and from Float16 (#8847)
- fix: fix(datafusion): report the scan range in scan statistics (#9411)
- fix: fix(datafusion): size the memory pool by the effective partition count (#9183)
- fix: fix(dataset): correct BlobFile seek semantics (#9358)
- fix: fix(dataset): fix open branch URI with tag to non-latest version (#9227)
- fix: fix(dataset): handle lance dataset blobs on deep clone (#9185)
- fix: fix(dataset): keep carried index bases through chained shallow clones (#9176)
- fix: fix(dataset): reject a zero add_columns batch size instead of panicking (#9233)
- fix: fix(dataset): reuse the prefetch window in BlobFile::read (#9330)
- fix: fix(dataset): stop the fragment-reuse index blocking row id migration (#9099)
- fix: fix(deps): update rustls for RUSTSEC-2026-0285 (#9212)
- fix: fix(encoding): derive full-zip max_visible_def like the writer (#9254)
- fix: fix(encoding): emit NIL control words when rep/def levels are non-empty but zero-width (#9018)
- fix: fix(encoding): preserve special slots when recording validity (#9268)
- fix: fix(encoding): skip full-width block bitpacking (#9598)
- fix: fix(encoding): split legacy decode batches before i32 offset overflow (#9225)
- fix: fix(exec): preserve exact KNN ordering through late materialization (#9469)
- fix: fix(format): move FLAG_UNKNOWN off the tagged FRI bit (#9420)
- …and 184 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
lance-format/lance was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 5483093251925b8e90703f216e150b263d04f587 — the exact code this score is about.
- Scored under rubric-2026.09.18 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-c4983f2d4e5c.