zouzias/spark-lucenerdd
61.9
Adequate · 20 September 2026
4.3k
lines of production code
Scala
primary language
1
measurement over time
What this system is
This system is a Scala library that integrates Apache Lucene with Apache Spark to enable distributed full-text search, faceted search, and geospatial analysis. It allows users to create Lucene-backed RDDs for advanced text querying, entity linkage, and spatial operations like nearest-neighbor and radius-based searches. The library provides configurable indexing, similarity scoring, and storage modes, along with utilities to convert search results into Spark DataFrames for further processing.
Features
Add default configuration for LuceneRDD via reference.conf
A new reference.conf file has been added to src/main/resources, establishing default settings for the LuceneRDD library. This configuration defines defaults for the text analyzer (English), similarity scoring (BM25), index storage mode (disk), string field analysis options, query limits (topK), and spatial indexing parameters (quad tree, WKT format). These defaults can be overridden by users via environment variables or application configuration.
src/main/resources · high confidence
Added SparkFacetResultMonoid for aggregating facet results
A new SparkFacetResultMonoid object has been introduced in the aggregate package to handle the aggregation of faceted search results. This component uses Algebird's MapMonoid to combine facet data from executors to the driver, ensuring that results for the same facet name are correctly summed while validating that mismatched facet names are not merged.
src/main/scala/org/zouzias/spark/lucenerdd/aggregate · high confidence
Added example scripts for Spark LuceneRDD usage
New Scala scripts have been added to the scripts directory to demonstrate how to use the Spark LuceneRDD library. These examples include loading text data (Alice in Wonderland, city names, and a word list) into a LuceneRDD, performing basic operations like counting and caching, and executing queries such as term queries and 'more like this' (MLT) searches. Additionally, a script is provided to load H1B Visa CSV data into a LuceneRDD, showcasing integration with Spark SQL and CSV data sources.
scripts · high confidence
Added fuzzy and prefix linkage example scripts
New example scripts have been added to the record linkage module to demonstrate how to perform fuzzy and prefix-based record matching using the LuceneRDD library. The \linkageFuzzyExample.scala\ script shows how to link records using a fuzzy matching function with a defined fuzziness level, while \linkagePrefixExample.scala\ demonstrates linking based on string prefixes. These examples provide users with ready-to-run code for common linkage scenarios.
scripts/recordLinkage · high confidence
Added new geographic dataset files
New Parquet dataset files have been added to the data directory, specifically for capitals, countries bounding boxes, countries polygons, and world cities points. These additions provide users with additional geographic data sources for spatial analysis.
data · high confidence
Added spatial data loading scripts for Solr linkage and Swiss cities
New scripts have been added to the spatial module to facilitate loading and linking geographic data. The \loadSolrSpatialData.scala\ script demonstrates how to load country polygons and city coordinates from Parquet files, create a ShapeLuceneRDD, and link them by radius to find nearby cities. Additionally, \loadSwissCities.scala\ provides an example of loading Swiss city data from a TSV file and indexing it into a ShapeLuceneRDD.
scripts/spatial · high confidence
Configurable Lucene analyzers via configuration file
Users can now specify the text analyzers used for indexing and querying through configuration keys (lucenerdd.index.analyzer.name and lucenerdd.query.analyzer.name). The new AnalyzerConfigurable trait supports a wide range of built-in Lucene analyzers identified by short language codes (e.g., en, de, fr, zh) or by providing the full class name of a custom analyzer present in the classpath. If no analyzer is configured, the system falls back to the StandardAnalyzer.
src/main/scala/org/zouzias/spark/lucenerdd/analyzers · high confidence
Faceted search capability for LuceneRDD
This change introduces the \FacetedLuceneRDD\ class and supporting implicit conversions in the \facets\ package, enabling users to perform faceted search queries on their indexed data. The new \FacetedLuceneRDD\ extends the base \LuceneRDD\ to provide \facetQuery\ and \facetQueries\ methods, which aggregate results using a monoidal structure to return both top documents and facet statistics. The \package.scala\ file adds implicit conversions for primitive types, tuples, maps, and Spark Rows, automatically creating additional indexed fields (suffixed with \\_facet\ or \\_numFacet\) to support these faceted queries. Additionally, a \Versionable\ trait is added to expose project version information to Spark.
src/main/scala/org/zouzias/spark/lucenerdd/facets · high confidence
Initial project scaffolding and documentation
This change introduces the foundational structure for the spark-lucenerdd project. It adds a comprehensive README detailing the library's capabilities (LuceneRDD, FacetedLuceneRDD, ShapeLuceneRDD), supported query types, and compatibility matrix for Spark versions up to 3.5.0. It also includes a Travis CI configuration for Scala 2.12.10 with Oracle and OpenJDK 8, a Scalastyle configuration for code quality, a Docker Compose setup for Zeppelin, and helper scripts (spark-shell.sh, startZeppelin.sh) to run the library locally with Spark 3.2.1. The project version is set to 0.4.1-SNAPSHOT.
(repo-wide) · high confidence
Initial release of LuceneRDD with indexing, querying, and linkage capabilities
This change introduces the core LuceneRDD library, enabling users to create Spark RDDs backed by Lucene indexes for advanced text search and entity linkage. The new \LuceneRDD\ class supports configurable analyzers (including per-field settings), similarity models (TF-IDF/BM25), and provides methods for generic queries, deduplication, and entity linkage via \link\ and \linkDataFrame\. The package also includes implicit conversions for converting Scala primitives, tuples, case classes, and maps into Lucene documents, along with Kryo serialization support and index statistics tracking.
src/main/scala/org/zouzias/spark/lucenerdd · high confidence
Introduce ShapeLuceneRDD for geospatial and full-text search linkage
This change introduces the ShapeLuceneRDD component, enabling users to perform spatial queries (KNN and radius-based) and link geospatial data with other RDDs via DataFrame integration. The new class provides methods like linkByKnn and linkByRadius to connect entities based on proximity, supported by implicit conversions for common shape types (points, rectangles, circles, polygons) and a dedicated Kryo registrator to ensure efficient serialization of spatial objects in Spark clusters.
src/main/scala/org/zouzias/spark/lucenerdd/spatial/shape · high confidence
Introduction of SpatialStrategy trait for Lucene spatial indexing
A new \SpatialStrategy\ trait has been added to the \org.zouzias.spark.lucenerdd.spatial.shape.strategies\ package. This trait extends \PrefixTreeLoader\ and initializes a \RecursivePrefixTreeStrategy\ from Apache Lucene, providing a unified API for indexing and searching spatial shapes with distance calculations. This change establishes the foundational strategy component for spatial data handling within the library.
src/main/scala/org/zouzias/spark/lucenerdd/spatial/shape/strategies · high confidence
New LuceneRDDResponse API for typed results and DataFrame conversion
The library introduces a new \LuceneRDDResponse\ class in the \response\ package to handle search results, replacing the previous approach. This new response type provides a \toDF()\ method that automatically infers the schema from the first result and converts the RDD into a Spark DataFrame, simplifying integration with Spark SQL. It also includes a \take(k)\ method that uses a Top-K monoid to efficiently retrieve the top-k results based on Lucene score ordering, and overrides \collect()\ to handle result aggregation. The implementation relies on new supporting classes \LuceneRDDResponsePartition\ and \FieldType\ to manage partitioned row iterators and data types.
src/main/scala/org/zouzias/spark/lucenerdd/response · high confidence
New TermDocMatrix class for Lucene field term-document analysis
A new TermDocMatrix class has been added to the matrices package, enabling the construction of a term-document matrix from Lucene field data. This component maps terms to matrix rows and documents to columns, allowing users to access the underlying Spark CoordinateMatrix, retrieve row-to-term mappings, and query matrix dimensions and non-zero entry counts.
src/main/scala/org/zouzias/spark/lucenerdd/matrices · high confidence
New model classes for facets, scored documents, and term vectors
Added three new model classes to the library: SparkFacetResult, which wraps Lucene facet results into a Spark-friendly structure with sorted counts; SparkScoreDoc, which extends Lucene score documents with score, document ID, and shard index, and provides conversion to Spark SQL Rows with inferred numeric types; and TermVectorEntry, which represents term vector data with document ID per shard, term text, and count.
src/main/scala/org/zouzias/spark/lucenerdd/models · high confidence
Partition-level Lucene indexing and query API introduced
The partition component now exposes a comprehensive set of search capabilities, including term, prefix, fuzzy, phrase, and multi-term queries, as well as More Like This (MLT) and faceted search. Each partition maintains its own Lucene index and taxonomy, allowing for per-partition analyzer and similarity configuration. Search results are returned via a new LuceneRDDResponsePartition structure, and the partition tracks indexing performance metrics.
src/main/scala/org/zouzias/spark/lucenerdd/partition · high confidence
Behavioural changes
Configurable disk-based Lucene index storage with taxonomy support
The library now supports storing Lucene indexes and taxonomy data on disk rather than exclusively in memory. A new \IndexStorable\ trait manages index and taxonomy directories in the system's temporary folder, using \MMapDirectory\ for disk storage (selected via the \lucenerdd.index.store.mode\ config parameter set to 'disk') and falling back to \RAMDirectory\ for in-memory storage. The \IndexWithTaxonomyWriter\ trait integrates these storage mechanisms with index and taxonomy writers, enabling facet queries to persist data on disk, which improves performance and resource usage for large indexes by leveraging the OS file system cache instead of the Java heap.
src/main/scala/org/zouzias/spark/lucenerdd/store · high confidence
Configurable search similarity and new query helper utilities
Users can now configure the Lucene search similarity algorithm via the \lucenerdd.similarity.name\ setting, choosing between BM25 and the default Classic similarity. Additionally, a new \LuceneQueryHelpers\ object provides utility methods for parsing queries with per-field analyzers, executing faceted text searches using taxonomy readers, and retrieving field names or total document counts from the index.
src/main/scala/org/zouzias/spark/lucenerdd/query · high confidence
Introduce PrefixTreeLoader for spatial grid configuration
A new PrefixTreeLoader trait has been added to the spatial grids module to centralize the configuration of spatial prefix trees. This component initializes the underlying Lucene SpatialPrefixTree based on configurable parameters such as the tree type (e.g., geohash or quad), maximum levels for precision, and maximum distance error, providing a standardized way to load spatial grid structures for shape-based operations.
src/main/scala/org/zouzias/spark/lucenerdd/spatial/shape/grids · high confidence
Introduces configurable traits for LuceneRDD settings
The library now exposes a set of Scala traits (Configurable, LuceneRDDConfigurable, ShapeLuceneRDDConfigurable) that allow users to externalize and customize core indexing and querying parameters via Typesafe Config. This includes defaults for top-K query results, facet counts, string field analysis options (such as term vectors, positions, and index options), and spatial search settings like prefix tree configuration and shape formats. Users can now adjust these behaviors through configuration files rather than hard-coded values.
src/main/scala/org/zouzias/spark/lucenerdd/config · high confidence
Refactored spatial partitioning with explicit search interfaces
The spatial partitioning logic in the shape module has been restructured to introduce an abstract base class (\AbstractShapeLuceneRDDPartition\) and a concrete implementation (\ShapeLuceneRDDPartition\). This change exposes explicit methods for nearest-neighbor search (\knnSearch\), circle search (\circleSearch\), arbitrary spatial search (\spatialSearch\), and bounding-box search (\bboxSearch\), allowing users to perform targeted spatial queries. The implementation also adds logging for indexing start, completion, and duration, providing visibility into indexing performance.
src/main/scala/org/zouzias/spark/lucenerdd/spatial/shape/partition · high confidence
Test coverage
Added LuceneRDDTestUtils for spatial and scoring test assertions; Added comprehensive test suite for LuceneRDD core and linkage features; Added test coverage for ShapeLuceneRDD spatial and linkage operations; Added test data models for Kryo serialization tests; Added test data resources for text and geospatial analysis; Added tests for FacetedLuceneRDD facet queries and case class integration; Added tests for LuceneQueryHelpers; Added tests for ShapeLuceneRDD implicits; Added tests for analyzer configuration; Added unit tests for LuceneRDDResponse.
Dependencies
Initial build configuration for Spark Lucene RDD
The project now includes a \build.sbt\ file that defines the build structure, specifying Scala 2.12.19 and Apache Spark 3.5.9 as provided dependencies. It integrates Apache Lucene 8.11.4 (including facet, analyzer, query parser, expressions, and spatial extras modules) alongside supporting libraries such as JTS Core 1.19.0, Spatial4J 0.8, Algebird 0.13.10, and Joda Time 2.12.7. The configuration also sets up test dependencies using ScalaTest 3.2.20 and spark-testing-base 3.5.1\_1.5.3, enables cross-building for Scala 2.12.19, and configures Maven-style publishing to Sonatype.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 62.
Lenses
- Code Health 97
- Architecture 97
- Maturity 49
- Readiness 62
- Security 73
Changes since last survey
- 300 commits — 268 feature/other, 32 fixes
By area
- (root) — 176 commits
- project/plugins.sbt — 64 commits
- project/build.properties — 24 commits
- src/main — 14 commits
- (repo) — 13 commits
- .github/ISSUE_TEMPLATE — 2 commits
- notebooks/2BXC9TF8J — 2 commits
- project/buildinfo.sbt — 2 commits
- .github/main.workflow — 1 commit
- .github/workflows — 1 commit
- src/test — 1 commit
Notable commits
- fix: (gh actions) fix sbt
- fix: (hotfix) make sbt publish happy
- fix: (hotfix) revert spark-testing ver bump
- fix: (hotfix) revert spark-testing-base version
- fix: HOTFIX: Propagate indexAnalyserPerField and queryAnalyzerPerField to partitions configuration (#367)
- fix: README: fix after release
- fix: README: fix javadoc badge
- fix: README: fix javadoc badge
- fix: Revert "Setting version to 0.2.9-SNAPSHOT"
- fix: Revert "Setting version to 0.3.1-SNAPSHOT"
- fix: Revert "Setting version to 0.3.10-SNAPSHOT"
- fix: Revert "Setting version to 0.3.2-SNAPSHOT"
- fix: Revert "Setting version to 0.3.3-SNAPSHOT"
- fix: Revert "Setting version to 0.3.4-SNAPSHOT"
- fix: Revert "Setting version to 0.3.4-SNAPSHOT"
- fix: Revert "Setting version to 0.3.5-SNAPSHOT"
- fix: Revert "Setting version to 0.3.7-SNAPSHOT"
- fix: Revert "Setting version to 0.3.8-SNAPSHOT"
- fix: Revert "Setting version to 0.3.9-SNAPSHOT"
- fix: Revert "Setting version to 0.4.0-SNAPSHOT"
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
zouzias/spark-lucenerdd was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f101b8a460760a45ec0aa49bde89f03e0f6a343d — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.