Skip to content
CAI
Software that uses CAICheck a score

alexklibisz/elastiknn

63.9

Adequate · 20 September 2026

4.3k

lines of production code

Scala

with Java, Python

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

Elastiknn is an Elasticsearch plugin that extends the search engine with specialized support for vector similarity search, handling both dense float and sparse boolean vectors. It provides exact and approximate nearest-neighbor search capabilities using Locality Sensitive Hashing (LSH) and SIMD-accelerated operations, while also allowing vector similarity to be used as a scoring function for hybrid search. The system includes dedicated Scala and Python client libraries to facilitate indexing and querying, along with comprehensive benchmarking and testing infrastructure to validate performance and correctness.

How it got here

2019–2020 — Initial plugin release and client libraries

19 changes.

This period established the foundational architecture of the Elastiknn Elasticsearch plugin, introducing core vector mappers, similarity models, and optimized search queries. It also delivered the initial releases of dedicated Python and Scala client libraries to enable external interaction with the vector search capabilities.

2022 — Build infrastructure and test coverage expansion

11 changes.

The project established a modern SBT-based build system with Docker support for local development and benchmarking, while introducing Lucene integration components like HashFieldType. Significant effort was dedicated to expanding test coverage across serialization, storage, and query logic to ensure correctness and determinism of vector search operations.

2023–2024 — SIMD acceleration and testing infrastructure

5 changes.

The project introduced SIMD-accelerated vector operations using Project Panama to enhance performance for dense float vector calculations. This work was supported by comprehensive integration tests, JMH benchmarks for performance validation, and new AWS-based development environment tooling.

Features

Add HashFieldType for Lucene integration

Introduces a new HashFieldType class within the elastiknn-lucene module to define how hash-based fields are indexed in Lucene. This type configures fields to be non-tokenized, index document and frequency information, omit norms, and is frozen for performance, enabling the library to properly store and retrieve hash vectors in the search index.

elastiknn-lucene/src/main/java/com/klibisz/elastiknn/lucene · high confidence

Add SIMD-accelerated vector operations via Project Panama

The library now includes a new \PanamaFloatVectorOps\ implementation that uses JDK incubator vector APIs (Project Panama/SIMD) to accelerate dense float vector calculations. This provides high-performance alternatives for dot product, L1/L2 distance, and cosine similarity, utilizing parallel accumulators and vector unrolling to maximize throughput. A \DefaultFloatVectorOps\ fallback and a \FloatVectorOps\ interface are also introduced to support this pluggable architecture, alongside a new \BooleanVectorOps\ class for efficient sorted integer intersection counting.

elastiknn-models/src/main/java/com/klibisz/elastiknn/vectors · high confidence

Add ann-benchmarks integration for performance testing

Adds a new integration directory for the ann-benchmarks tool, enabling users to benchmark Elastiknn's approximate nearest neighbor search performance. This includes a README with setup instructions and Task recipes for running benchmarks (e.g., Fashion MNIST), a submodule reference to the ann-benchmarks repository, a configuration file defining Elastiknn-specific algorithms (exact and L2LSH), a Python script to parse and format benchmark results into tables, and a shell script to scrape Docker container metrics during testing.

ann-benchmarks · high confidence

Added AWS development environment setup

Developers can now provision a remote development instance in AWS using Terraform. The new \development/aws\ directory contains the infrastructure configuration (EC2 instance, security groups, SSH keys) and a setup script that installs required tools (Docker, SBT, Task, Java, Python) and compiles the project, providing an alternative to local development for benchmarking and testing.

development · high confidence

Added SBT build infrastructure and Elasticsearch plugin development support

The project has migrated its build system to SBT (version 1.12.11), introducing a new \project\ directory with \build.properties\ and \plugins.sbt\ (including sbt-jmh, sbt-tpolecat, and sbt-scalafmt). A custom \ElasticsearchPluginPlugin\ was added to streamline development, providing tasks to download the Elasticsearch distribution, bundle the project into an installable plugin zip, and run or debug a local Elasticsearch node with the plugin installed.

project · high confidence

Initial release of the ElastiKnn Python client package

This change introduces the initial packaging structure for the ElastiKnn Python client, enabling users to install and use the library via standard Python tools. The diff establishes the \setup.py\ configuration, \MANIFEST.in\ to include \.pyi\ type hint files and \.proto\ definitions, and a \.gitignore\ to manage build artifacts. This provides the foundational distribution mechanism for the Python client, which implements queries against the ElastiKnn Elasticsearch plugin.

client-python · high confidence

Initial release of the Elastiknn Python client

Introduces the \elastiknn\ Python package, providing a high-level interface for interacting with the Elastiknn Elasticsearch plugin. The package includes an \ElastiknnClient\ for managing index mappings, indexing vector data, and executing nearest-neighbor queries, as well as an \ElastiknnModel\ class that offers a scikit-learn-compatible API (with \fit\ and \kneighbors\ methods) for training and querying vector models using exact and LSH algorithms.

client-python/elastiknn · high confidence

Introduce Docker-based Elasticsearch environment with JDK vector support

This change adds a new Docker setup for running Elasticsearch with the Elastiknn plugin, including a Dockerfile that installs the plugin into Elasticsearch 9.4.1, a health-check script, and Docker Compose files for benchmarking and testing. The benchmarking configuration enables the optional JDK incubator vector module (Project Panama/SIMD) for faster dense float vector operations and exposes JMX ports for monitoring, while the testing configuration sets up a multi-node cluster (master and data nodes) with similar vector support enabled.

docker · high confidence

Introduce Java-based elastiknn-models module with LSH and similarity implementations

The library now includes a new \elastiknn-models\ module containing core vector similarity and hashing logic rewritten in Java. This adds support for Locality Sensitive Hashing (LSH) models including Cosine, Hamming, Jaccard, L2, and Permutation, alongside an \ExactModel\ for precise similarity calculations. The module also introduces new serialization utilities via \ByteBufferSerialization\ and \BitBuffer\ to handle vector data efficiently.

elastiknn-models/src/main/java/com/klibisz/elastiknn/models · high confidence

Introduce elastiknn-api4s subproject with vector search API definitions

This change introduces the new \elastiknn-api4s\ subproject, which defines the core Scala data models and JSON serialization logic for the plugin's vector search capabilities. It adds case classes for vector types (\Vec.DenseFloat\, \Vec.SparseBool\, \Vec.Indexed\), similarity metrics (\Similarity\), and mapping configurations (\Mapping\), along with query structures (\NearestNeighborsQuery\) supporting exact and approximate (LSH) search modes. The subproject also includes Java utility buffers (\FloatArrayBuffer\, \IntArrayBuffer\) for efficient vector handling and an \XContentCodec\ to serialize these types to/from Elasticsearch's XContent format.

elastiknn-api4s/src/main · high confidence

Introduce new Elastic4s-based Scala client for vector operations

The elastiknn-client-elastic4s module now provides a dedicated Scala client built on the elastic4s library, enabling developers to interact with Elasticsearch for vector search tasks using idiomatic Scala patterns. This new client introduces \ElastiknnClient\ and \ElastiknnFutureClient\ abstractions that wrap elastic4s \ElasticClient\ to handle indexing vectors, creating indices with optimized mappings, and executing nearest-neighbor searches. It includes \ElastiknnRequests\ for constructing high-performance requests (such as storing document IDs as doc values for faster retrieval) and \Elastic4sCompatibility\ to seamlessly convert elastiknn query objects into elastic4s \Query\ types. Users can now instantiate the client from an existing Elasticsearch \RestClient\ or directly via host/port, with an optional \strictFailure\ mode to ensure non-2xx responses or bulk errors are raised as exceptions.

elastiknn-client-elastic4s · high confidence

Introduce score functions for nearest-neighbor queries

The plugin now supports using nearest-neighbor similarity as a scoring function within Elasticsearch's function\_score queries. This allows users to re-rank or boost documents based on vector similarity (e.g., cosine, L2, L1) alongside other scoring signals. The implementation adds specific query builders and score function parsers for both exact and approximate (LSH) vector searches, enabling hybrid search scenarios where vector proximity influences the final document ranking.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn/query · high confidence

Introduction of Elastiknn plugin with SIMD vector optimization support

The Elastiknn plugin is now available, providing Elasticsearch with specialized mappers for dense float and sparse boolean vectors, a custom query builder, and a score function for nearest-neighbor search. A key capability is the optional use of JDK incubator vector (Project Panama/SIMD) for accelerated dense float vector operations, controlled via the \elastiknn.jdk-incubator-vector.enabled\ setting. The plugin also introduces a unified \ElastiknnException\ hierarchy for error handling and registers the necessary components (queries, mappers, score functions) within the Elasticsearch plugin lifecycle.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn · high confidence

Introduction of VectorMapper for sparse boolean and dense float vector indexing

This change introduces the VectorMapper class and its specific implementations (SparseBoolVectorMapper and DenseFloatVectorMapper) within the elastiknn plugin. These components handle the parsing, validation, and Lucene field creation for sparse boolean and dense float vector types, supporting exact search and various LSH (Locality Sensitive Hashing) mappings like Jaccard, Hamming, Cosine, L2, and Permutation. The implementation includes a custom FieldType that returns an empty list from the ValueFetcher and utilizes x-content for decoding vectors, ensuring compatibility with Elasticsearch's mapping system while adhering to security constraints by avoiding unsafe runtime dependencies.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn/mapper · high confidence

Introduction of core model abstractions and caching for vector similarity

This change introduces the foundational Scala abstractions for the plugin's similarity models, specifically adding the \ExactSimilarityFunction\ trait and its implementations for Jaccard, Hamming, L1, L2, and Cosine similarities, alongside the \HashingFunction\ trait and a \ModelCache\ for managing LSH model instances. For users, this establishes the internal structure for how exact and approximate vector similarities are computed and cached within the plugin, supporting the various distance metrics available in the API.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn/models · high confidence

Introduces the MatchHashesAndScoreQuery class, which enables approximate nearest neighbor search by matching document hashes and applying a configurable scoring function. This query supports limiting the number of candidate results, provides detailed explanations for similarity scores, and includes performance optimizations such as specialized handling for single-frequency hashes and efficient hit counting.

elastiknn-lucene/src/main/java/org · high confidence

New multimodal search tutorial for Amazon products

Added a new Jupyter notebook tutorial in the \examples/tutorial-notebooks\ directory that demonstrates multimodal search on the Amazon Products dataset. The tutorial uses Elasticsearch 7.9.2 with the Elastiknn plugin to index product metadata and 4096-dimensional image vectors, enabling combined keyword and nearest-neighbor (image similarity) searches. Supporting files include a Dockerfile, docker-compose configuration, and Python utility scripts to facilitate running the demo locally.

examples · high confidence

New vector storage abstraction and reader

The plugin introduces a new internal storage layer for vectors, decoupling the API representation from the physical storage format. This includes a \StoredVec\ trait hierarchy (supporting sparse boolean and dense float types) with typeclass-based codecs for encoding and decoding, and a \StoredVecReader\ class that efficiently retrieves these vectors from Lucene's binary doc values. This change provides a foundation for future optimizations, such as switching serialization methods or supporting streaming reads, without altering the public API.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn/storage · high confidence

Behavioural changes

Added backward-compatible codec wrappers for Lucene 8.4–8.8

The plugin now includes new codec classes (Elastiknn84Codec, Elastiknn86Codec, Elastiknn87Codec, and Elastiknn88Codec) that wrap the corresponding Lucene backward codecs. These are provided for backward compatibility with older index formats, though they are noted as no longer used in Elasticsearch 8.

elastiknn-plugin/src/main/scala/com/klibisz/elastiknn/codec · high confidence

Plugin security policy tightened and utility class added

The plugin's security policy has been restricted by removing the getClassLoader permission, ensuring a more secure execution environment. Additionally, a new utility class, VectorMapperUtil, was introduced to provide a shared empty array constant for FieldMapper parameters, supporting consistent vector mapping configurations.

elastiknn-plugin/src/main/plugin-metadata · high confidence

Refactored hit counting internals with optimized ArrayHitCounter

The search module now uses a new \HitCounter\ interface to abstract document hit counting, replacing previous inline implementations. The primary implementation, \ArrayHitCounter\, introduces a specialized \kthGreatest\ algorithm that computes the k-th highest hit count in linear time, enabling more efficient retrieval of top-k results via a custom \DocIdSetIterator\. An \EmptyHitCounter\ is also provided for cases with zero matches. These changes simplify query internals and optimize performance for scenarios where hash matches are sparse or unique.

elastiknn-lucene/src/main/java/com/klibisz/elastiknn/search · high confidence

Registration of Elastiknn codecs for Lucene 8.4–8.8

The plugin now registers four specific Lucene codecs (Elastiknn84Codec, Elastiknn86Codec, Elastiknn87Codec, and Elastiknn88Codec) via the standard Java SPI mechanism. This enables the plugin to support indexing and searching with these Lucene codec versions, ensuring compatibility with the underlying search engine's storage format.

elastiknn-plugin/src/main/resources · high confidence

Test coverage

Added JMH benchmarks for vector operations, serialization, and hit counting; Added Python client and model tests; Added elastiknn integration tests; Added performance benchmarking tests for HashingQuery; Added tests for ArrayHitCounter and Lucene test utilities; Added tests for MatchHashesAndScoreQuery behavior; Added tests for Panama SIMD vector operations; Added tests for XContentCodec serialization and deserialization; Added tests for exact similarity functions and Permutation LSH model determinism; Added tests for storage serialization components; Added unit tests for elastiknn LSH models.

Dependencies

Upgrade to Elasticsearch 9.4.1 and Scala 3.3.7

The build configuration has been updated to use Elasticsearch 9.4.1, Scala 3.3.7, and Lucene 10.4.0. This upgrade includes corresponding updates to the elastic4s client (9.3.0), Circe (0.14.15), Guava (33.6.0-jre), and Eclipse Collections (13.0.0). The Python client dependencies have also been refreshed to align with the new Elasticsearch version.

(dependencies) · high confidence

Housekeeping

Initial repository structure and configuration

The repository has been initialized with the core project configuration files, including the version file (set to 9.4.1.0), build tool settings (SBT, scalafmt), and CI/CD task definitions (Taskfile.yaml). This includes the addition of .gitignore, .gitattributes, and .tool-versions to manage development environment consistency, as well as the replacement of the placeholder readme.md with a comprehensive README.md detailing the project's purpose as an Elasticsearch plugin for similarity search.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 64.

Lenses

  • Code Health 97
  • Architecture 86
  • Maturity 64
  • Readiness 56
  • Security 69

Changes since last survey

  • 300 commits — 288 feature/other, 12 fixes

By area

  • (root) — 175 commits
  • docs/pages — 22 commits
  • .github/workflows — 21 commits
  • project/build.properties — 21 commits
  • elastiknn-plugin/src — 11 commits
  • project/build.sbt — 10 commits
  • project/plugins.sbt — 9 commits
  • elastiknn-lucene/src — 5 commits
  • elastiknn-models/src — 4 commits
  • elastiknn-api4s/src — 3 commits
  • elastiknn-plugin-integration-tests/src — 3 commits
  • (repo) — 2 commits
  • .claude/skills — 2 commits
  • docs/_config.yml — 2 commits
  • docs/_posts — 2 commits
  • .claude/settings.json — 1 commit
  • .github/scripts — 1 commit
  • ann-benchmarks/README.md — 1 commit
  • client-python/requirements.txt — 1 commit
  • development/aws — 1 commit

Notable commits

  • fix: #785: fix java.security.AccessControlException caused by ClassTag (#786)
  • fix: Bugfix: bump version to release changes from #786 (#787)
  • fix: Build: fix name in release-docs workflow (#668)
  • fix: Build: fix python version for docs release job (#565)
  • fix: Build: fix python version in release workflow (#563)
  • fix: Build: fix python version in release workflow (#564)
  • fix: Build: fix quotes from previous commit
  • fix: Build: fix the Elasticsearch download URL so it works on Linux (#725)
  • fix: Build: fix the target commit when releasing (#475)
  • fix: Fix non-determinism in PermutationLshModelSuite via JIT warm-up (#832)
  • fix: Fix release badges in README.md
  • fix: Plugin: bump version to 8.15.0.1 to release the fix for #715 (#722)
  • change: Add Claude Code setup (#833)
  • change: Add JMH benchmarks for Lucene IntIntHashMap and Eclipse IntShortHashMap (#599)
  • change: Benchmarks: note about thread_pool.search.queue_size
  • change: Benchmarks: point submodule at ann-benchmarks master (#497)
  • change: Benchmarks: re-run benchmarks on EC2 (#587)
  • change: Benchmarks: update ann-benchmarks submodule to use Python 3.10.8 (#521)
  • change: Benchmarks: upgrade ann-benchmarks to use Elastiknn 8.6.2 (#494)
  • change: Build: Reduce scala-steward cron to once per week
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

alexklibisz/elastiknn was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 20 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit bde03df006177794c54e7ff83a6f7e3bbc2fd277 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b51f968c9b10.