quickwit-oss/tantivy
76.3
Strong · 28 September 2026
146.3k
lines of production code
Rust
primary language
2
measurements over time
What this system is
Tantivy is a high-performance, full-text search library for Rust, designed as a modular alternative to Lucene. It provides core indexing and search capabilities, including inverted index management, tokenization, and a comprehensive aggregation engine compatible with Elasticsearch-style queries. The system features a modern columnar storage format for fast field access, extensible plugin architectures for custom components, and optimized data structures for efficient memory and disk usage.
How it got here
2016–2018 — Core engine modernization and extensibility
27 changes.
This period focused on a comprehensive architectural overhaul of the Tantivy search engine, introducing a plugin system, a new columnar storage format, and a modular schema API. Significant performance improvements were achieved through block-based inverted index structures, branchless search algorithms, and optimized compression for postings and document stores. The work also expanded query capabilities with phrase and regex support, added new collectors for aggregations, and established a robust foundation for JSON indexing and extensible tokenization.
2019–2023 — Columnar storage and aggregation engine
43 changes.
This period focused on the foundational shift to a new columnar storage format (V2) and the implementation of a comprehensive aggregation engine compatible with Elasticsearch. Key developments included the introduction of SSTable-based term dictionaries, optimized bitpacking with SIMD support, and a modular architecture for tokenizers and low-level primitives. The work also expanded query capabilities with features like MoreLikeThis and PhrasePrefix, while establishing robust testing and benchmarking suites for the new storage and indexing layers.
2024–2026 — Aggregation and query engine optimization
15 changes.
This period focused on significantly enhancing the query engine and aggregation framework through major refactoring and new features. Key developments included rewriting composite and term aggregations for better performance, introducing JIT expression evaluation for calculated fields, and implementing optimized union query strategies. The work also expanded testing coverage with fuzzing and benchmarks while adding modular infrastructure for sorting and document predicates.
Features
Add Histogram and DateHistogram bucket aggregations
Introduces two new bucket aggregation types: \HistogramAggregation\ for numeric data and \DateHistogramAggregation\ for date fields. The \HistogramAggregation\ allows users to chunk data into dynamic buckets based on a specified numeric \interval\, supporting features like \offset\, \min\_doc\_count\, \hard\_bounds\, and \extended\_bounds\. The \DateHistogramAggregation\ provides similar bucketing capabilities for date values but is restricted to fixed time intervals (e.g., \1d\, \1h\) specified via the \fixed\_interval\ parameter, converting date values to milliseconds for calculation. Both aggregations support a \keyed\ parameter to return results as a hash map and are implemented with segment-level collectors for performance.
src/aggregation/bucket/histogram · high confidence
Add built-in stop word lists for multiple languages
The StopWordFilter now supports built-in stop word lists for 13 languages (Danish, Dutch, Finnish, French, German, Hungarian, Italian, Norwegian, Portuguese, Russian, Spanish, Swedish, and English). Users can select a language via the existing Language enum to automatically filter common words, with the lists sourced from the Snowball project and Apache Lucene.
_src/tokenizer/stop\_word\filter · high confidence
Add position tracking for phrase queries
Introduces a new positions module that stores term positions (token ordinals) to enable phrase queries. The implementation uses SIMD bitpacking to compress position deltas in blocks of 128, with a fallback to variable-int encoding for remaining values, and provides a reader optimized for sequential access and efficient skipping.
src/positions · high confidence
Added ArenaHashMap usage example
The stacker/example crate now includes a new Rust source file demonstrating how to use the ArenaHashMap component. The example shows creating a map with a specified capacity and using the mutate\_or\_create method to increment counts for string terms, providing a concrete usage pattern for developers.
stacker/example · high confidence
Added SIMD-accelerated vector filtering for AVX2, NEON, and SVE
The bitpacker module now includes optimized implementations for filtering vectors using SIMD instructions, significantly improving performance on supported hardware. The new \filter\_vec\ module automatically selects the best available instruction set at runtime: AVX2 on x86\_64, NEON on all aarch64, and SVE on non-Apple aarch64 systems, falling back to a scalar implementation if no SIMD is available. This change enhances the speed of range-based filtering operations used internally by the search engine.
_bitpacker/src/filter\vec · high confidence
Added columnar CLI tool for JSON-to-columnar conversion
A new command-line interface has been added to the columnar project, enabling users to convert JSON data into the columnar format. The tool reads JSON lines from a file (defaulting to 'gh\_small.json'), parses the nested structures, and writes the extracted numerical, boolean, and string values into a columnar index. It provides performance metrics for the build and serialization phases, and outputs a summary of the resulting columns and their sizes.
columnar/columnar-cli · high confidence
Added dense and sparse block codecs for optional column indices
The columnar storage layer now includes new \DenseBlockCodec\ and \SparseBlockCodec\ implementations within the optional index set-block module. These codecs provide optimized serialization and lookup mechanisms for optional fields: the dense codec uses bit-vectors for compact storage and fast rank/select operations on dense value sets, while the sparse codec uses sorted arrays with binary search for memory efficiency on sparse data. This change introduces the underlying data structures and algorithms that enable more efficient handling of optional column data in Tantivy's columnar storage.
_columnar/src/column\_index/optional\_index/set\block · high confidence
Added fieldnorm module with compression and per-field control
The \src/fieldnorm\ module has been introduced to manage field norms, which represent the length of a field in a document and are used to adjust BM25 scoring (shorter fields are weighted higher). The implementation uses a 256-entry lookup table to compress field norms into a single byte per document, mirroring Lucene's scheme. This module provides writers to track norms during indexing, serializers to persist them, and readers to retrieve them during search. Users can now control whether field norms are computed for specific text fields via the schema's \TextOptions\ (e.g., \set\_fieldnorms(false)\), allowing for optimization when norm-based scoring is not desired.
src/fieldnorm · high confidence
Added file modification watching for the mmap directory
The mmap directory now monitors the meta file for changes using a background polling thread. This new \FileWatcher\ component computes a checksum of the meta file at regular intervals and triggers registered callbacks when modifications are detected, enabling the directory to react to external updates.
_src/directory/mmap\directory · high confidence
Initial aggregation engine implementation
The \src/aggregation\ module has been introduced to provide Elasticsearch-compatible aggregation capabilities. This includes a full request parsing layer (\agg\_req.rs\) that accepts JSON requests, a segment-level collection architecture (\collector.rs\, \agg\_data.rs\) for executing aggregations, and support for both bucket aggregations (terms, range, histogram, date\_histogram, filter, composite, multi\_terms) and metric aggregations (avg, sum, min, max, count, stats, extended\_stats, percentiles, cardinality, top\_hits). The implementation includes memory and bucket limits (\agg\_limits.rs\) to prevent resource exhaustion, fast-field accessors (\accessor\_helpers.rs\) for efficient data retrieval, and intermediate result handling for merging across segments.
src/aggregation · high confidence
Initial documentation structure for Tantivy search library
The documentation site has been initialized with a new structure in \doc/src\, establishing the foundational content for the Tantivy Rust search library. This includes a foreword clarifying Tantivy's scope as a library rather than a server, core technical chapters on index anatomy (segments, merging, disk access), schema definition, and faceting. It also introduces specific feature documentation for JSON object types (including flattening, encoding, and limitations) and index sorting (covering compression, top-N optimization, and usage).
doc/src · high confidence
Initial project scaffolding and documentation
This change introduces the foundational structure for the Tantivy project, including the \ARCHITECTURE.md\ file that details the library's design inspired by Lucene, a comprehensive \CHANGELOG.md\ documenting versions from 0.22 to 0.27, and standard repository files such as \README.md\, \LICENSE\, \AUTHORS\, \SECURITY.md\, and \CITATION.cff\. It also adds configuration files for Rust tooling (\rustfmt.toml\, \clippy.toml\), release automation (\cliff.toml\, \RELEASE.md\), and build scripts (\Makefile\), along with an updated \.gitignore\ to exclude build artifacts and IDE files.
(repo-wide) · high confidence
Introduce DocPredicateQuery for evaluating document-level conditions
Added a new \DocPredicateQuery\ component that allows queries to evaluate arbitrary boolean conditions on a per-document basis. This includes a generic \FunctionPredicate\ for wrapping custom closures and a \JitExprPredicate\ (behind the \jitexpr\ feature flag) for evaluating JIT-compiled expressions against fast fields. The implementation optimizes performance by using necessary conditions to skip documents that cannot possibly match the predicate and leverages a compilation cache for JIT expressions.
_src/query/doc\_predicate\query · high confidence
Introduce JIT expression framework for calculated fields
This change adds a new JIT expression framework in \jitexpr/src\ that enables the compilation of calculated field expressions into native machine code for faster evaluation. The implementation includes an untyped AST with type inference, a Lisp-like serialization format, and a compilation pipeline using Cranelift. Key features include a thread-safe compilation cache to avoid redundant work, proper memory management for JIT modules, and support for various functions (ADD, IS\_NULL, REGEXP\_EXTRACT, etc.) with safe handling of numeric types and null values.
jitexpr/src · high confidence
Introduce MoreLikeThis query for finding similar documents
Added a new \MoreLikeThisQuery\ and its builder in the \src/query/more\_like\_this\ module, enabling users to find documents similar to a target document or a set of field values. The implementation extracts frequent terms from the target using configurable filters (such as minimum/maximum document frequency, term frequency, word length, and stop words) and constructs a BooleanQuery to retrieve similar results, mirroring the behavior of the Apache Lucene MoreLikeThis query.
_src/query/more\_like\this · high confidence
Introduce PhrasePrefixQuery for phrase searches with a prefix wildcard
Added a new \PhrasePrefixQuery\ component that allows users to search for a specific sequence of words followed by a term where only the prefix is known (e.g., matching "part time" but not "part of"). This feature requires the target field to have positions indexed and supports configuring the maximum number of prefix expansions to limit performance impact.
_src/query/phrase\_prefix\query · high confidence
Introduce columnar storage with multi-value and optional document indexing
The \column\_index\ module is introduced to manage the mapping between document IDs and row IDs in the new columnar storage system. It supports three cardinality modes: \Full\ (one value per document), \Optional\ (zero or one value), and \Multivalued\ (zero or more values). The implementation includes serialization and deserialization logic for persisting these indices, as well as specific index structures like \MultiValueIndexV2\ which uses an optional index to efficiently handle sparse multi-value documents. This change provides the foundational indexing layer required for the columnar storage feature.
_columnar/src/column\index · high confidence
Introduce columnar writer with multivalue and optional cardinality support
The columnar writer now supports optional and multivalued fields by tracking document cardinality and building corresponding value indexes. This allows documents to have zero or multiple values per field, with the writer correctly handling document ordering and value indexing for these cases.
columnar/src/columnar/writer · high confidence
Introduce dictionary-encoded columns for string and byte data
The columnar storage layer now supports dictionary-encoded columns for text and binary data via the new \BytesColumn\ and \StrColumn\ types. This allows string and byte values to be stored as compact term ordinals backed by an SSTable dictionary, enabling more efficient serialization, deserialization, and memory usage compared to raw value storage. The implementation includes dedicated serialization and deserialization logic to handle the dictionary and ordinal column components together.
columnar/src/column · high confidence
Introduce multi-terms aggregation for composite bucketing
Added a new \multi\_terms\ aggregation that creates one bucket per unique combination of values across multiple term fields, mirroring the behavior of Elasticsearch's \multi\_terms\. Users can specify an ordered list of fields to aggregate on, with support for standard parameters like \size\, \segment\_size\, \min\_doc\_count\, and \order\. The response returns composite keys as arrays (e.g., \\["rock", "Product A"\]\) and includes a \key\_as\_string\ representation. Note that byte columns are not supported, and missing values are handled via configurable fallback keys or exclusion.
_src/aggregation/bucket/multi\terms · high confidence
Introduce snippet generation with highlighted search results
Added a new \SnippetGenerator\ component that creates text previews for search results, highlighting matching terms within the context. Users can now generate HTML snippets with customizable prefix/postfix tags (defaulting to \\<b\>\ and \\</b\>\) and limit the output length via \set\_max\_num\_chars\. The \Snippet\ struct provides access to the raw text fragment and the specific character ranges of highlighted terms, enabling rich search result displays.
src/snippet · high confidence
Introduce standalone tokenizer API crate
A new \tokenizer-api\ crate has been added to provide a stable, separate interface for tokenizers, decoupling them from the core tantivy library so that implementers do not need to update their code for every new tantivy version. This change introduces the core \Tokenizer\ and \TokenStream\ traits, along with the \Token\ struct and \TokenFilter\ trait, establishing the foundational API for text tokenization.
tokenizer-api · high confidence
Introduce tantivy-columnar crate with columnar storage and merge capabilities
This change introduces the \tantivy-columnar\ crate, providing a new columnar storage format for Tantivy that enables efficient column-based reads, writing, and merging of multiple segments. The implementation supports various data types including booleans, integers (I64, U64), floats (F64), IP addresses, dates, and strings, with automatic handling of cardinality (full, optional, multivalued). It includes a dictionary encoding mechanism using SSTables, numerical type coercion between I64, U64, and F64, and compatibility tests to ensure format stability across versions. The crate also provides utilities for bit manipulation and byte operations, along with comprehensive test coverage for the new storage format.
columnar/src · high confidence
Introduces a hierarchical skip index for faster document store lookups
The document store now uses a multi-layer skip index structure to accelerate random access to stored documents. This change introduces a new \CheckpointBlock\ for grouping checkpoints and a \SkipIndex\ that organizes these blocks into hierarchical layers. By maintaining pointers to specific byte offsets for ranges of document IDs, the store can quickly seek to the correct location without scanning the entire file, improving performance for large documents and complex queries.
src/store/index · high confidence
Introduces configurable LZ4 and Zstd compression for the document store
The document store now supports block-level compression using LZ4 (block format) and Zstd, replacing the previous default. Users can select the compression algorithm via the \Compressor\ enum (e.g., \Compressor::Lz4\ or \Compressor::Zstd\), with LZ4 as the default when the \lz4-compression\ feature is enabled. Zstd compression allows configuring the compression level (defaulting to 3) to balance speed and size. The store footer now records the active decompressor to ensure correct reading of existing and new indices. This change is gated by feature flags (\lz4-compression\, \zstd-compression\), and if neither is enabled, the store falls back to no compression.
src/store · high confidence
Introduces extensible segment component plugin system
Users can now attach custom data structures to segments through the new \SegmentPlugin\ trait, allowing external code to participate in the segment lifecycle (write, read, and merge) without modifying Tantivy internals. The system manages file extensions, writers, and merging logic, with built-in components like postings and fast fields implemented as plugins. A new \TantivyError::MissingPlugin\ variant is returned if a required segment plugin is not registered before writing, merging, or garbage collecting.
src · high confidence
Introduction of PhraseQuery and RegexPhraseQuery with slop support
Users can now perform phrase searches using \PhraseQuery\ to match specific sequences of terms, and \RegexPhraseQuery\ to match sequences defined by regular expressions. Both query types support a 'slop' parameter, allowing terms to appear within a specified distance of each other rather than requiring strict adjacency. The implementation enforces that the target field must have positions indexed and returns a schema error if this requirement is not met. Additionally, \RegexPhraseQuery\ includes a configurable \max\_expansions\ limit to prevent performance degradation from overly broad regex patterns.
_src/query/phrase\query · high confidence
Introduction of the OwnedBytes crate
A new \ownedbytes\ crate has been added, providing the \OwnedBytes\ struct. This type wraps a \StableDeref\ object to expose its data as a byte slice, supporting operations such as slicing, splitting, advancing, and reading little-endian integers (u8, u32, u64). It also implements standard traits like \Debug\, \PartialEq\, and \Deref\ for seamless integration with existing byte-slice APIs.
ownedbytes · high confidence
Introduction of the TermDictionary abstraction with backend selection
The \src/termdict\ module now provides a unified \TermDictionary\ API that wraps either an FST-based or an SSTable-based storage backend. The specific backend is selected at compile time via the \quickwit\ feature flag: the default build uses the FST implementation, while enabling the \quickwit\ feature switches to the SSTable implementation. The dictionary format includes a footer marker to identify the backend type, and the API exposes standard operations such as term lookup, ordinal mapping, streaming, and range queries, with additional asynchronous methods (e.g., \get\_async\, \warm\_up\_dictionary\) available only when the \quickwit\ feature is enabled.
src/termdict · high confidence
Introduction of the new QueryParser module with comprehensive query syntax support
The \src/query/query\parser\ directory has been restructured to introduce a new \QueryParser\ implementation. This change adds support for a wide range of query syntaxes including unbounded range queries, exists queries (\field:\\), regex queries, phrase prefix queries, and set queries (\IN \[...\]\). It also introduces an \ExistsQuery\ type and handles specific field types like IP addresses and JSON paths. The parser now returns a \LogicalAst\ that represents the parsed query structure, enabling features like boosting, boolean logic, and term sets. This is a foundational change to how user queries are interpreted and executed.
_src/query/query\parser · high confidence
Introduction of the standalone bitpacker crate with blocked encoding support
The \bitpacker\ crate is introduced, providing a new implementation for bit-packing and unpacking numeric data. It includes a \BitPacker\ and \BitUnpacker\ for basic operations, and a \BlockedBitpacker\ that compresses data in 128-element blocks while maintaining an index for efficient random access. The blocked packer uses a base value offset to reduce the number of bits required per block, improving compression ratios for data with non-zero minimums. This change extracts and refactors the bitpacking logic into a dedicated crate, making it reusable and potentially faster through optimized block decoding and unaligned memory reads.
bitpacker/src · high confidence
Introduction of u128 column value serialization with CompactSpace codec
The columnar storage layer now supports serializing and deserializing u128-based column values using a new 'CompactSpace' codec. This change introduces the \U128Header\ structure and associated serialization logic (\serialize\_column\_values\_u128\, \open\_u128\_mapped\) which compresses u128 data by removing holes in the number space. It also provides a specialized accessor (\open\_u128\_as\_compact\_u64\) to retrieve the data as u64 for faster processing, enabling more efficient storage and retrieval of large integer ranges in columnar formats.
_columnar/src/column\_values/u128\based · high confidence
New BitSetDocSet for efficient bitset iteration
Users can now iterate through a BitSet as a DocSet via the new BitSetDocSet implementation. This allows queries to efficiently traverse document sets represented by bitmaps, leveraging bucket-based skipping for performance. The change includes comprehensive tests for sequential iteration, seeking, and empty set handling.
src/query/bitset · high confidence
New ColumnarReader for accessing columnar storage files
Introduces the ColumnarReader component, which enables users to open and inspect columnar data files. This reader parses the file footer to determine the number of documents and format version, then exposes methods to list all available columns or retrieve specific columns by name (including support for JSON subpaths). It provides both synchronous and asynchronous APIs for reading column data, allowing efficient access to the underlying columnar storage structure.
columnar/src/columnar/reader · high confidence
New SSTable merge functionality with configurable value merging
The sstable module now includes a new \merge\ submodule that enables merging multiple SSTables into a single output. This change introduces a heap-based merge algorithm (\heap\_merge.rs\) and defines traits (\ValueMerger\, \SingleValueMerger\) to allow configurable merging strategies. Built-in merge strategies include \KeepFirst\ (which retains the first value encountered for duplicate keys) and \U64Merge\ (which sums u64 values for duplicate keys). The implementation is accompanied by comprehensive tests verifying correct key ordering and value aggregation behavior.
sstable/src/merge · high confidence
New collector module with Count, Facet, Histogram, and Filter collectors
The \src/collector\ module has been introduced, providing a new set of search result collectors. Users can now use \Count\ to retrieve the total number of matching documents, \FacetCollector\ to compute facet breakdowns, and \HistogramCollector\ to generate value distribution histograms. Additionally, \FilterCollector\ allows filtering documents based on fast field predicates before passing them to a wrapped collector, and \MultiCollector\ enables running multiple collectors simultaneously with dynamic type handling.
src/collector · high confidence
New columnar file validation CLI tool
A new command-line utility has been added to inspect and validate columnar files written by Tantivy. Users can now run this tool against a columnar file path to verify its integrity, which includes checking row counts, listing columns, and validating string column metrics such as value counts, term dictionary sizes, and ordinal constraints.
columnar/columnar-cli-inspect · high confidence
New common library module with core data structures and utilities
The \common/src\ directory now contains a new library module providing foundational components for the system. This includes a \DateTime\ type with nanosecond precision and configurable truncation, a \ByteCount\ type for human-readable byte size formatting, and a \FileSlice\ abstraction for efficient, chunked reading of file data. Additionally, the module introduces a \TinySet\ for compact bitset operations, a \JsonPathWriter\ for constructing flattened JSON paths, and utility traits and functions for binary serialization (\BinarySerializable\), variable-length integer encoding (\VInt\), and iterator grouping (\GroupByIteratorExtended\).
common/src · high confidence
New compact space codec for u128 column values
A new \CompactSpace\ codec has been added to the columnar storage engine to compress u128-based column values. This codec maps large, sparse value ranges (such as IP addresses) into a smaller, dense compact space by identifying and excluding 'blank' ranges between actual data points. This allows the underlying bit-packer to use fewer bits per value, improving storage efficiency and potentially query performance for datasets with large value gaps.
_columnar/src/column\_values/u128\_based/compact\space · high confidence
New filter aggregation with programmatic query builder support
Users can now apply a filter aggregation to create a single bucket containing only documents that match a specific query. This new feature supports both standard query strings (parsed via Tantivy's QueryParser) and programmatic query construction through a new \QueryBuilder\ trait, enabling custom query types and distributed aggregation scenarios with full serialization support.
src/aggregation/bucket · high confidence
New metric aggregations: average, count, min, max, sum, stats, extended stats, percentiles, and top\_hits
The \src/aggregation/metric\ module now implements a comprehensive suite of metric aggregations. Users can compute single-value metrics (average, count, min, max, sum) and multi-value statistics (stats, extended\_stats) on numeric fields, with support for a \missing\ parameter to handle null values. The percentiles aggregation uses a DDSketch-based collector for approximate distribution analysis, while the top\_hits aggregation allows retrieving the most relevant documents within buckets based on custom sort criteria and field retrieval. The sum aggregation also introduces a \none\_if\_no\_match\ flag for SQL-style null handling when no values are collected.
src/aggregation/metric · high confidence
New optional index codec for sparse column data
Introduces a new \OptionalIndex\ implementation in \columnar/src/column\_index/optional\_index\ to efficiently store and query sparse (optional) column data. The codec uses a hybrid encoding strategy, switching between dense and sparse block representations based on a configurable threshold (5,120 elements) to optimize storage size and lookup performance. This change provides the underlying data structure for handling columns with missing values, including serialization, deserialization, and efficient rank/select operations via cursors.
_columnar/src/column\_index/optional\index · high confidence
New query types and scoring infrastructure
This change introduces several new query capabilities and scoring components to the query engine. Users can now use \AllQuery\ to match every document in the index, \EmptyQuery\ to match none, and \ExistsQuery\ to find documents where a specific fast field contains a non-null value (with optional JSON subpath support). The \DisjunctionMaxQuery\ allows combining multiple queries where the final score is the maximum of the matching clauses plus a configurable tie-breaker increment. Scoring flexibility is improved with \BoostQuery\ to multiply scores by a factor, \ConstScoreQuery\ to assign a fixed score to a wrapped query's results, and \Bm25Weight\ which exposes the BM25 statistics provider interface for custom scoring logic. Additionally, \AutomatonWeight\ provides a generic weight for automaton-based queries like fuzzy and regex searches, and \Exclude\ enables filtering out documents from a result set based on an exclusion set.
src/query · high confidence
New query-grammar crate with lenient parsing and JSON serialization
The \query-grammar\ crate has been introduced to centralize query parsing logic, exposing \parse\_query\ for strict validation and a new \parse\_query\_lenient\ function that recovers from syntax errors and returns structured \LenientError\ hints. The parsed \UserInputAst\ is now serializable to JSON (using snake\_case and tagged enums), enabling clients to inspect the query structure. The grammar supports new query types including \Exists\, \Regex\, and \Set\ (IN), and handles edge cases like field names with special characters and escape sequences more robustly.
query-grammar/src · high confidence
New space usage reporting API for index components
A new \space\_usage\ module has been added to expose programmatic access to the byte-level storage consumption of Tantivy indexes. This introduces data structures like \SearcherSpaceUsage\ and \SegmentSpaceUsage\ that aggregate usage across segments and individual components (such as term dictionaries, postings, positions, fast fields, field norms, stored documents, and deletions). The API supports serialization for integration with CLI tools and allows users to query total byte counts and per-component breakdowns, though it currently excludes file-system block size overhead.
_src/space\usage · high confidence
New tokenization filters and tokenizers added
The tokenizer module now includes several new components to enhance text processing capabilities. Users can filter tokens to contain only ASCII alphanumeric characters using the new AlphaNumOnlyFilter, or convert non-ASCII Unicode characters to their ASCII equivalents with the AsciiFoldingFilter. Text can be normalized to lowercase using the LowerCaser filter, and long tokens can be discarded via the RemoveLongFilter. Additionally, new tokenizers are available: RegexTokenizer allows splitting text based on custom regular expressions, NgramTokenizer generates character n-grams for fuzzy matching, FacetTokenizer emits hierarchical facet paths, and EmptyTokenizer produces no tokens. The existing Stemmer filter has been updated to use the frostem library, supporting a wider range of languages including Arabic, Danish, Dutch, Finnish, French, German, Greek, Hungarian, Italian, Norwegian, Portuguese, Romanian, Russian, Spanish, Swedish, Tamil, and Turkish.
src/tokenizer · high confidence
New u64-based column codecs for faster range queries and better compression
The columnar storage layer now includes three new codecs for u64-based fast fields: Bitpacked, Linear, and BlockwiseLinear. The Linear and BlockwiseLinear codecs use linear interpolation to approximate values, storing only the deviation from the line, which significantly improves compression for data with trends (like timestamps) and speeds up range queries by allowing the decoder to skip values that fall outside the query range. The system automatically selects the most efficient codec during serialization based on the data distribution.
_columnar/src/column\_values/u64\based · high confidence
Re-enabled and updated example applications
The Tantivy example applications have been re-enabled and updated to reflect the current API. This includes new or refreshed examples for basic search, faceted search, aggregations (including filter and range aggregations), custom collectors, custom tokenizers, date/time fields, document deletion and updates, fuzzy search, and multi-threaded indexing. These examples demonstrate how to define schemas, index documents, and perform various search and aggregation operations using the latest Tantivy interfaces.
examples · high confidence
Restructure index module with extensible segment plugins
The index module has been reorganized into a modular plugin-based architecture. Core index creation is now handled via \IndexBuilder\, and segment file management is driven by \SegmentPlugin\ implementations (such as the built-in \InvertedIndexPlugin\) which own specific file extensions. This change exposes segment file listing on \IndexMeta\ and allows custom plugins to extend segment components, while maintaining backward compatibility for standard index operations.
src/index · high confidence
Support for merging column indexes with shuffled row orders
The columnar module now supports merging column indexes when the resulting row order is shuffled, not just stacked. This change introduces \ShuffleMergeOrder\ handling in the column index merge logic, allowing the system to correctly reconstruct optional and multivalued indexes when rows from different segments are interleaved. It includes logic to detect the resulting cardinality (Full, Optional, or Multivalued) by analyzing alive bitsets and existing cardinalities, ensuring that deleted or empty rows are handled correctly during the merge process.
_columnar/src/column\index/merge · high confidence
Architecture
Isolate stacker components into independent crates
The stacker library has been refactored to isolate its core components—specifically the memory arena, arena-backed hash maps, exponential unrolled linked lists, and fast byte-comparison/copy utilities—into independent crates. This structural change improves modularity and allows these low-level indexing primitives to be reused or tested in isolation without pulling in the entire stacker dependency tree.
stacker/src · high confidence
Isolated FST-based term dictionary implementation
The term dictionary logic has been isolated into a dedicated \fst\_termdict\ module, introducing a new FST-backed storage format (version 1) that separates the sorted term index from term metadata. This change adds a \TermMerger\ to efficiently merge term streams from multiple segments, a \TermStreamer\ for range-based iteration, and a \TermInfoStore\ that uses bitpacking to compress posting and position offsets. Users benefit from a cleaner internal structure for term lookups and merging, with the new format ensuring compatibility via version checks during dictionary opening.
_src/termdict/fst\termdict · high confidence
Behavioural changes
Block-WAND optimization for Boolean Query scoring
Boolean queries now use the Block-WAND algorithm to significantly speed up scoring for top-K document retrieval. This optimization applies to both term intersections (AND) and unions (OR) when all subqueries are TermScorers with frequency data enabled. By using block-max pruning, the engine skips large ranges of documents that cannot exceed the current score threshold, reducing the number of expensive scoring and intersection checks required during search.
_src/query/boolean\query · high confidence
Core engine refactored with new Executor, Searcher, and JSON indexing support
The core module has been significantly restructured to support modern search capabilities. A new \Executor\ abstraction has been introduced to manage task execution, offering both single-threaded and multi-threaded (via Rayon) modes, including asynchronous task spawning for the \quickwit\ feature. The \Searcher\ API has been redesigned to wrap \SegmentReader\s and expose generation tracking via \SearcherGeneration\, enabling features like search warming and consistent snapshotting. Additionally, comprehensive JSON indexing support has been added through \json\_utils\, which handles path-aware position tracking to prevent false positives in phrase queries for nested JSON objects. Legacy modules such as \postings\, \schema\, \directory\, and \writer\ have been removed from this location as their functionality has been migrated or replaced by these new components.
src/core · high confidence
Document model refactored to trait-based architecture with compact storage
The document schema layer has been restructured around new core traits (\Document\, \Value\, \DocumentDeserialize\) to enable zero-allocation indexing of custom document types and improve plugin extensibility. The previous \DocValue\ type has been renamed to \Value\, and the default document implementation is now \CompactDoc\ (aliased as \TantivyDocument\), which stores data in a compact binary format. A new \ErasedDocument\ trait provides object-safe, type-erased access to documents for plugins operating behind \dyn\ boundaries. Additionally, \OwnedValue\ now uses a \Vec\-based representation for arrays and objects, and the binary serialization format has been updated to support these changes, including specific handling for custom field payloads.
src/schema/document · high confidence
Fast fields now use the columnar storage format
The fast field implementation has been migrated from the legacy bitpacked format to the unified columnar storage system. This change replaces the previous \BitpackedFastFieldReader\ and \FastFieldSerializer\ with the \ColumnarReader\ and \ColumnarWriter\, enabling support for a broader range of data types including booleans, dates, IP addresses, and JSON fields. The migration introduces a new \FastFieldsPlugin\ to manage the lifecycle of fast fields, updates the \FastFieldReaders\ API to resolve column names for JSON paths, and standardizes the serialization and merging logic through the columnar codec infrastructure.
src/fastfield · high confidence
Introduces IndexReaderBuilder with configurable reload policies and doc store cache sizing
The \src/reader\ module now provides an \IndexReaderBuilder\ that allows users to configure how index updates are detected and applied. Users can choose between \Manual\ reloads (requiring explicit calls to \reload()\) or \OnCommitWithDelay\ (automatic background reloading after a commit, with a note that it is not synchronized with the commit return). Additionally, the builder exposes \doc\_store\_cache\_num\_blocks\ to control the size of the doc store cache, defaulting to \DOCSTORE\_CACHE\_CAPACITY\. This change replaces previous implicit reader construction with a configurable builder pattern, also introducing a \Warmer\ trait for maintaining segment-level state during reloads.
src/reader · high confidence
Introduces block-based postings and branchless search for inverted index performance
The postings module has been refactored to use a block-based architecture for the inverted index. A new \BlockSegmentPostings\ type iterates over compressed document blocks, and a new \block\_search.rs\ module implements a branchless k-ary search algorithm to locate terms within these blocks more efficiently than traditional binary search. The indexing side now uses an \IndexingContext\ to manage memory arenas for posting lists, and a new \LoadedPostings\ type allows in-memory caching of postings for terms with few documents to reduce memory overhead. Additionally, a \JsonPostingsWriter\ has been added to handle the specific serialization requirements of JSON fields.
src/postings · high confidence
Introduction of Columnar storage format V2
The columnar storage module now supports a new on-disk format version (V2), which is set as the current default. This change introduces a structured \ColumnType\ enum to explicitly define supported data types (including I64, U64, F64, Bytes, Str, Bool, IpAddr, and DateTime) and implements version-aware serialization and deserialization logic with magic bytes for file integrity verification.
columnar/src/columnar · high confidence
Introduction of the new columnar \`ColumnValues\` API and fast-field codecs
The \columnar/src/column\_values\ module has been replaced with a new implementation that provides the \ColumnValues\ trait for accessing dense field columns, replacing the previous interface. This change introduces support for strictly monotonic mappings (via \monotonic\_map\_column\) to enable efficient range queries on types like \DateTime\ and \Ipv6Addr\ by mapping them to \u64\/\u128\ spaces, and adds a \MergedColumnValues\ struct to handle merging data from multiple segments with configurable row ordering. The new API also includes \VecColumn\ for in-memory column representation and \ColumnStats\ for serialization, fundamentally changing how fast-field data is stored, accessed, and merged.
_columnar/src/column\values · high confidence
Major rewrite of the Directory abstraction and file storage format
The \src/directory\ module has been completely rewritten to introduce a new WORM (Write Once Read Many) directory trait, replacing the previous implementation. This change introduces a new file format where every file ends with a JSON-serialized footer containing a CRC32 checksum and the index version, enabling corruption detection and version compatibility checks. A new \ManagedDirectory\ wrapper now tracks index files in a \meta.json\ manifest to support garbage collection of unused files. The \RamDirectory\ implementation has been updated to track active writer memory usage and enforce that files cannot be rewritten once created. Additionally, a new \CompositeFile\ structure allows storing data partitioned by field within a single file, and a \WatchCallback\ system has been added to notify subscribers of file changes.
src/directory · high confidence
New SSTable-based term dictionary implementation
The term dictionary storage has been replaced with a new implementation based on SSTables (Sorted String Tables). This change introduces a new \TermSSTable\ structure that handles the serialization and deserialization of \TermInfo\ objects, including document frequency and posting/position ranges, using a custom binary format. It also adds a \TermMerger\ component to efficiently merge sorted term streams from multiple segments into a single sorted iterator, enabling the system to read and write the term dictionary using this new underlying storage format.
_src/termdict/sstable\termdict · high confidence
New block-based and VInt compression implementations for postings
The postings compression module now introduces \BlockEncoder\ and \BlockDecoder\ structs that utilize the \BitPacker4x\ algorithm to compress sorted and unsorted integer blocks, supporting optional offset-based delta encoding and minus-one encoding for compactness. Additionally, a new VInt encoding layer has been added to handle variable-length integer compression for posting lists, providing specific functions for compressing and decompressing both sorted (with delta) and unsorted sequences. These changes establish the underlying binary format for storing and retrieving inverted index data.
src/postings/compression · high confidence
New columnar merge implementation with row reordering and type coercion
The columnar merge module has been replaced with a new implementation that supports merging multiple columnar tables while allowing rows to be reordered or dropped via \MergeRowOrder\ (stacked or shuffled). It introduces automatic numerical type coercion (e.g., merging \i64\ and \u64\ into a compatible type) and enforces required columns in the output. The merge process now handles dictionary columns by streaming a k-way merge of term dictionaries to build a unified term ordering, and it groups columns by type category to ensure consistent serialization.
columnar/src/columnar/merge · high confidence
New modular sort key infrastructure for collectors
The sorting logic in the collector module has been restructured into a new \src/collector/sort\_key\ directory, introducing a \SortKeyComputer\ trait and specific implementations for sorting by string (\SortByString\), static fast values (\SortByStaticFastValue\), bytes (\SortByBytes\), and similarity score (\SortBySimilarityScore\). This change adds support for sorting by bytes fields and introduces \SortByErasedType\ to allow sorting by fields with types determined at runtime, while also refactoring the score-based top-K collection to use a \BinaryHeap\ for tighter pruning thresholds.
_src/collector/sort\key · high confidence
New optimized union query implementations
The query engine now includes two new union strategies: \BufferedUnionScorer\, which uses a sliding window to pre-buffer document IDs and scores for improved performance on non-scored unions, and \BitSetPostingUnion\, which combines a bitset for fast iteration with individual docsets to support position retrieval (e.g., for regex phrase queries). A \SimpleUnion\ is also provided as a lightweight alternative for queries that perform many seeks. These components replace the previous union logic in \src/query/union\ to address performance regressions and enable more efficient handling of large term unions.
src/query/union · high confidence
Optimized term aggregation with flattened histogram grid
The term aggregation collector now uses a specialized flattened grid structure for the common case of combining terms with a single histogram or date\_histogram sub-aggregation. This optimization significantly improves performance for low-cardinality term fields by reducing memory overhead and avoiding full sorts during segment-level top-k selection, while falling back to the general buffered path when the grid size exceeds cache-friendly limits.
_src/aggregation/bucket/term\agg · high confidence
Range queries now execute via fast fields for significant performance gains
Range queries on fast fields (including U64, I64, F64, Str, Date, IpAddr, Bytes, and Json) now use a new \RangeDocSet\ implementation that scans the columnar storage directly instead of relying on the inverted index. This change introduces a lazy, seek-optimized iteration strategy that can be up to 50x faster for non-overlapping ranges, while also supporting limits on the number of matched terms for inverted index queries and fixing fractional JSON range bounds on integer columns.
_src/query/range\query · high confidence
Refactored indexing architecture with new delete queue and document ID mapping
The indexing subsystem has been restructured to improve concurrency and correctness. A new \DeleteQueue\ implementation replaces the previous delete handling mechanism, utilizing a linked-list structure with cursors to manage delete operations across multiple consumers. Document ID mapping during merges is now handled by dedicated \DocIdMapping\ and \DocToOpstampMapping\ modules, enabling precise tracking of document lifecycles and supporting index sorting. The \IndexWriter\ has been updated to use a builder pattern (\IndexWriterOptions\) for configuring memory budgets and thread counts, and integrates these new components to manage segment lifecycle and delete application more robustly.
src/indexer · high confidence
Rewritten composite aggregation with improved pagination and performance
The composite aggregation implementation has been completely rewritten to improve performance and correctness. The new code introduces a \DynArrayHeapMap\ for efficient memory usage and key eviction, and adds support for calendar-based date histogram intervals (Year, Month, Week) alongside existing fixed intervals. Pagination logic has been refined to correctly handle \after\_key\ scenarios, including explicit missing values and \MissingOrder\ settings, ensuring accurate result ordering. Additionally, the rewrite includes robust handling of numeric type comparisons (i64, u64, f64) to prevent overflow and precision issues during aggregation.
src/aggregation/bucket/composite · high confidence
SSTable index format upgraded to v3 with FST-based lookups and automaton-aware block pruning
The sstable module now uses a new index format (v3) that replaces the previous linear block list with an FST-based index for term lookups, significantly improving lookup performance. This change introduces a new \block\_match\_automaton\ module that allows the index to efficiently prune blocks that cannot match a given automaton during range or pattern queries. The \BlockReader\ and \DeltaReader\ have been updated to handle non-contiguous block slices, preserving term ordinals across pruned gaps to ensure correct iteration. The index version is now 3, with backward compatibility maintained for v2 indexes.
sstable/src · high confidence
Schema module refactored into granular option types and flags
The schema definition logic has been reorganized into distinct modules for field options (e.g., \BytesOptions\, \DateOptions\, \IpAddrOptions\, \JsonObjectOptions\, \NumericOptions\, \TextOptions\, \FacetOptions\) and a new \flags\ module for composable field attributes (\INDEXED\, \STORED\, \FAST\, \COERCE\). This change introduces a \CustomOptions\ type to support plugin-defined field types and updates the \FieldEntry\ and \FieldType\ structures to utilize these new granular configurations, providing a more modular and extensible schema API.
src/schema · high confidence
Term query now falls back to fast field range query for unindexed fields
When a TermQuery targets a field that is not inverted-indexed but is indexed as a fast field, the query now automatically executes a range query on that fast field to find matches. This allows users to search for terms in fields that were previously not searchable via TermQuery, though the fallback is noted to be slower than standard inverted index lookups. This change is specific to cases where scoring is disabled; if scoring is enabled, the query will fail if the field is not indexed.
_src/query/term\query · high confidence
Test coverage
1 commit adding/updating tests in sstable/tests; Added benchmark suite for bitpacking and filter operations; Added benchmark suite for stacker data structures; Added benchmark suite for vint and bitset operations; Added benchmarks for SSTable ord\_to\_term and stream operations; Added columnar access and performance benchmarks; Added compatibility tests for index versions 6 and 7; Added failpoint tests for directory garbage collection and index writer failures; Added fuzz testing for ArenaHashMap; New benchmark suite for aggregations, queries, and tokenization; Removed core intersection test.
Dependencies
Tantivy 0.27 release with Rust 2024 edition and dependency updates
This release updates the Tantivy search engine library to version 0.27.0, migrating several core subcrates (bitpacker, columnar, common, sstable, stacker, query-grammar) to the Rust 2024 edition. It introduces a new JIT expression evaluation crate (jitexpr) and updates key dependencies including base64 to 0.23.0, itertools to 0.14.0, and lru to 0.18.2.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 78 → 76 (-1.4)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.
Lenses
- Code Health 88 → 88 (+0.0)
- Architecture 99 → 94 (-4.9)
- Maturity 67 → 67 (-0.0)
- Readiness 93 → 80 (-13.1)
- Security 82 → 86 (+4.6)
- Performance 100 (new)
Resolved (23)
- Boundary-crossing change coupling: facet_reader.rs ↔ merger.rs (src/fastfield/facet_reader.rs)
- Change coupling: infallible.rs ↔ user_input_ast.rs (query-grammar/src/infallible.rs)
- Documentation: no installation or build instructions (README.md)
- Duplicated block (10 lines × 2) (src/aggregation/bucket/multi_terms/mod.rs)
- Duplicated block (12–15 lines × 3) (src/query/phrase_prefix_query/phrase_prefix_query.rs)
- Duplicated block (5 lines × 4) (jitexpr/src/functions/left.rs)
- Hotspot: jitexpr/src/compile/compile_fn_builder.rs (jitexpr/src/compile/compile_fn_builder.rs)
- Hotspot: jitexpr/src/functions/mod.rs (jitexpr/src/functions/mod.rs)
- Hotspot: src/aggregation/agg_req.rs (src/aggregation/agg_req.rs)
- Hotspot: src/aggregation/intermediate_agg_result.rs (src/aggregation/intermediate_agg_result.rs)
- Hotspot: src/fastfield/writer.rs (src/fastfield/writer.rs)
- Hotspot: src/index/inverted_index_plugin.rs (src/index/inverted_index_plugin.rs)
- Hotspot: src/query/boolean_query/block_wand_intersection.rs (src/query/boolean_query/block_wand_intersection.rs)
- Hotspot: src/query/boolean_query/boolean_weight.rs (src/query/boolean_query/boolean_weight.rs)
- Hotspot: src/query/range_query/range_query_fastfield.rs (src/query/range_query/range_query_fastfield.rs)
- Members sharing a duplicated core (4 members, 50+ identical tokens) (jitexpr/src/functions/left.rs)
- Members sharing a duplicated core (5 members, 50+ identical tokens) (jitexpr/src/functions/left.rs)
- Off-boarding risk: anonymized user #1
- Off-boarding risk: anonymized user #2
- Off-boarding risk: anonymized user #3
- …and 3 more
New (29)
- BitUnpacker::get_range (cognitive 29) (bitpacker/src/bitpacker.rs)
- BitUnpacker::get_range (cyclomatic 18) (bitpacker/src/bitpacker.rs)
- Dependency hygiene PARTLY measured — Cargo dependencies read, no committed lock to grade for currency
- Documentation: no usage examples (src/aggregation/README.md)
- Duplicate functionality with confusing parameter naming. Both methods perform the same logical operation: mapping a batch of DocIds to their corresponding RowIds. The parameter doc_ids_out is semantically identical in both signatures but appears in the middle of the argument list, which is non-standard and confusing compared to typical Rust patterns (usually output buffers are at the end or passed as &mut).
- Duplicated block (10 lines × 2) (query-grammar/src/query_grammar.rs)
- Duplicated block (16–19 lines × 2) (src/query/phrase_prefix_query/phrase_prefix_query.rs)
- Duplicated block (5 lines × 2) (src/aggregation/bucket/multi_terms/mod.rs)
- Duplicated block (5 lines × 4) (jitexpr/src/functions/left.rs)
- Hotspot: bitpacker/src/bitpacker.rs (bitpacker/src/bitpacker.rs)
- Inconsistent mutation patterns. insert returns a new TinySet (immutable/functional style), while insert_mut returns a bool (mutable/in-place style). While distinct, the naming insert vs insert_mut is slightly redundant if insert is clearly immutable. More importantly, BitSet uses insert(el: u32): bool for mutable insertion. This creates a confusing API surface where TinySet has two insertion methods with different semantics and return types, and BitSet has only one.
- Inconsistent naming for optional/batch retrieval. get_vals returns T (implying all values exist), while get_vals_opt returns Option<T>. However, Column::first_vals uses &mut [Option<T>] for batch retrieval, mixing the *_opt suffix pattern with a direct method name. The inconsistency lies between get_vals (no suffix, assumes existence) and first_vals (no suffix, returns Options).
- Inconsistent naming for single-value retrieval. Column::first implies retrieving the first value of a potentially multi-valued field for a document, while ColumnValues::get_val is a generic index-based retrieval. If Column::first is intended for single-value columns, it should be named get or get_val for consistency with ColumnValues. If it is specifically for multi-value columns, the name is clear, but the existence of get_val in the underlying values interface creates a split in terminology for 'getting a value by index/id'.
- Inconsistent return types for similar operations. BitSet::insert returns bool (indicating if the value was already present), while TinySet::insert returns a new TinySet (immutable). This forces users to handle the result differently depending on which set type they are using, despite the operation being conceptually the same.
- JitExprPredicate::doc_predicate (cognitive 17) (src/query/doc_predicate_query/jitexpr_predicate.rs)
- Members sharing a duplicated core (4 members, 50+ identical tokens) (jitexpr/src/functions/left.rs)
- Members sharing a duplicated core (5 members, 50+ identical tokens) (jitexpr/src/functions/left.rs)
- No ADRs found
- Off the main sequence: jitexpr
- Off the main sequence: ownedbytes
- …and 9 more
Changes since last survey
- 38 commits — 36 feature/other, 2 fixes
By area
- src/directory — 11 commits
- src/aggregation — 8 commits
- bitpacker/src — 5 commits
- benches/agg_bench.rs — 3 commits
- src/query — 3 commits
- (repo) — 2 commits
- jitexpr/src — 2 commits
- query-grammar/src — 2 commits
- columnar/src — 1 commit
- jitexpr/Cargo.toml — 1 commit
Notable commits
- fix: Fix cardinality calls for datasketches 0.5
- fix: Fix query parser panics on malformed input
- change: (Calculated fields) Add feature-gated JIT expression document predicate (#3081)
- change: Accelerating jitexpr conditions using necessary conditions (#3129)
- change: Add 26-bit date histogram aggregation benchmark
- change: Added jitexpr compilation cache (#3107)
- change: Addressing -0.0 problem in SafeF64 (#3109)
- change: Apply multi-terms cutoff before final bucket conversion
- change: Batch multi-terms dictionary lookups after pruning
- change: Benchmark multi-terms aggregations across many segments
- change: Buffer RamDirectory writes
- change: Clarify chunk size and packed value type in bit unpacking fast path
- change: Collect decoded multi-term keys before assigning bucket positions
- change: Compile RegexPhraseQuery regexes once per weight (#3135)
- change: Extract bit unpacking load helper with explicit safety contract
- change: Faster columnar data fetching
- change: Finalize mmap writers in directory tests
- change: Format with nightly rustfmt
- change: Make aggregation request-data types crate-internal (#3110)
- change: Merge incoming aggregation buckets into the accumulator
- …and 18 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
quickwit-oss/tantivy was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 047464cf92e5a31d02a696f5158e45f7d34c67eb — the exact code this score is about.
- Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.