Skip to content
CAI
Software that uses CAICheck a score

nixiesearch/nixiesearch

54.0

Adequate · 28 September 2026

13.4k

lines of production code

Scala

primary language

2

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Nixiesearch is a distributed search engine that combines full-text search with semantic vector retrieval and generative AI capabilities. It supports complex querying through boolean logic, aggregations, and reranking, while ingesting data from REST APIs, files, and Kafka streams. The system integrates local and cloud-based AI models for embeddings and text generation, offering flexible deployment modes including standalone, master-slave replication, and AWS Lambda execution.

How it got here

2023 — Architecture overhaul and AI integration

33 changes.

This period focused on a major architectural restructuring of the search engine, introducing a new hierarchical configuration system, a unified query interface, and a refactored core with dedicated indexer and searcher components. Significant features were added to support AI-driven search, including flexible inference providers, semantic embedding caching, and advanced aggregation and filtering capabilities. The work also established production readiness through comprehensive test coverage, Kubernetes deployment manifests, and a structured CLI.

2024 — distributed architecture and AI integration

28 changes.

The project refactored its core architecture to support distributed master-slave index synchronization and pluggable storage backends, while splitting the CLI into distinct operational modes. Significant enhancements included the introduction of native LLaMA.cpp generative model support, expanded embedding infrastructure, and new data ingestion sources like Kafka. These structural changes were accompanied by comprehensive test coverage and UI updates to ensure stability across the new distributed and AI-driven capabilities.

2025 — Search capabilities and infrastructure expansion

27 changes.

This period focused on significantly expanding the search engine's feature set, introducing advanced query types like semantic search and reranking, while adding support for high-dimensional vectors and diverse structured field types. Concurrently, the project enhanced its AI infrastructure by integrating cloud embedding providers, ONNX-based ranking models, and an in-memory caching layer to improve performance and flexibility. The work also included critical infrastructure updates for GraalVM native-image support, AWS Lambda deployment, and comprehensive Prometheus monitoring.

Features

Add AWS Lambda runtime client and periodic evaluation utility

This change introduces support for running the searcher as an AWS Lambda function. It adds a new \LambdaRuntimeClient\ that handles the AWS Lambda Runtime API protocol, including decoding incoming API Gateway V2 requests (with correct base64 decoding for payloads) and posting invocation responses or errors. It also adds a \PeriodicEvalStream\ utility for running periodic background tasks, which supports the standalone mode synchronization mentioned in the commit history.

src/main/scala/ai/nixiesearch/main/subcommands/util · high confidence

Add Kafka and File document sources with configurable offset strategies

Users can now ingest documents from Apache Kafka topics and local files in addition to existing sources. The new Kafka source supports flexible offset configuration, allowing consumers to start from the earliest message, the latest message, a specific timestamp, or a relative duration before the current time. File sources read JSON documents from URLs, supporting recursive directory traversal. Both sources stream documents through a shared parsing pipeline that applies the index mapping schema.

src/main/scala/ai/nixiesearch/source · high confidence

Add ONNX-based ranking model provider with Qwen3-Reranker support

Users can now use ONNX models for text ranking/re-ranking tasks via the new OnnxRankModel provider. This implementation supports both traditional cross-encoder formats and template-based prompts (such as Qwen3-Reranker), automatically configuring padding and logits processing (e.g., Sigmoid) for Qwen models. The feature allows configuring model handles, batch sizes, device placement, and prompt templates through the existing inference configuration system.

src/main/scala/ai/nixiesearch/core/nn/model/ranking/providers · high confidence

Add anonymous usage statistics collection

The application now collects anonymous usage statistics by default, reporting system information (OS, architecture, JVM version), configuration hashes, and uptime metrics to a Google Apps Script webhook. This feature is controlled by the \nixiesearch.core.telemetry.usage\ configuration flag and includes data anonymization for sensitive details like file paths, model names, and index fields.

src/main/scala/ai/nixiesearch/util/analytics · high confidence

Add range and term aggregation support for numeric and text fields

The search engine now supports range and term aggregations, allowing users to bucket results by numeric ranges or distinct term values. Range aggregations are available for integer, long, float, double, date, and datetime fields, supporting inclusive/exclusive boundaries (gte, gt, lte, lt). Term aggregations work on text fields as well as numeric and date fields, with support for an 'all' size parameter to return all available buckets. Results are serialized with proper JSON encoding/decoding for both aggregation types.

src/main/scala/ai/nixiesearch/core/aggregate · high confidence

Add support for OpenAI and Cohere cloud embedding providers

Users can now configure the search engine to use remote embedding models from OpenAI and Cohere in addition to local ONNX models. This change introduces new provider implementations that handle API authentication via the OPENAI\_API\_KEY and COHERE\_API\_KEY environment variables, including validation that keys are present and correctly formatted. Both providers support configurable retry policies with exponential backoff, custom timeouts, and batch sizes, and they integrate with the existing in-memory embedding cache to reduce redundant API calls.

src/main/scala/ai/nixiesearch/core/nn/model/embedding/providers · high confidence

Add support for filtering search results

Users can now filter search results by specifying include and exclude predicates. The new filter API supports term-based filtering for text, boolean, integer, long, date, and datetime fields, as well as boolean logic (AND, OR, NOT) to combine multiple conditions. This allows for more precise retrieval of documents by excluding irrelevant results or restricting searches to specific criteria.

src/main/scala/ai/nixiesearch/api/filter · high confidence

Added in-memory embedding cache for reduced API latency

Users can now enable a heap-based cache for embedding models to avoid redundant API calls for identical text inputs. This change introduces a \MemoryCachedEmbedModel\ that wraps an underlying embedding model, using a Scaffeine-backed cache with a configurable maximum size (defaulting to a batch size of 128). When encoding documents, the system checks the cache for existing embeddings; if found, it returns them immediately, otherwise it fetches them from the underlying model and stores the result for future use.

src/main/scala/ai/nixiesearch/core/nn/model/embedding/cache · high confidence

Expanded field types and nested document support in JSON ingestion

The core document decoder now supports a significantly wider range of field types, including boolean, date, datetime, long, double, float, and geopoint fields, as well as list variants for integers, longs, floats, and doubles. It also introduces support for nested documents by flattening nested objects and arrays into dot-notation fields, and allows text fields to carry pre-computed embeddings. Additionally, the system can now transparently decompress incoming JSON payloads in GZIP, ZSTD, and BZ2 formats, and provides progress logging for indexing operations.

src/main/scala/ai/nixiesearch/core · high confidence

Initial Kubernetes deployment manifests and documentation for Nixiesearch

This release adds official Kubernetes deployment resources for Nixiesearch, providing two distinct modes: a \*\Standalone\\* configuration using a single pod with persistent local storage (PVC), and a \*\Distributed\\* configuration separating searchers and indexers with S3-based index synchronization. The \deploy/kubernetes\ directory includes complete manifests (Deployments, StatefulSets, Services, ConfigMaps, PVCs) and detailed README guides for both modes. Additionally, Docker entrypoint scripts (\nixiesearch.sh\, \nixiesearch-native.sh\) are introduced to configure JVM options, including the addition of the \jdk.incubator.vector\ module flag for vector operations.

deploy · high confidence

Introduce Prometheus metrics for search, inference, indexing, and system health

The application now exposes a comprehensive set of Prometheus metrics for monitoring search operations, inference requests, indexing activity, and system resources. Users can track search and suggest query counts and latencies, embedding and completion request metrics, indexing flush totals and times, and system-level indicators such as CPU load, disk usage, and JVM statistics via the new \Metrics\ registry.

src/main/scala/ai/nixiesearch/core/metrics · high confidence

Introduce distributed index synchronization with master/slave replication

The \src/main/scala/ai/nixiesearch/index/sync\ package now implements a replication architecture for distributed search indices. It introduces \MasterIndex\ for write operations and \SlaveIndex\ for read operations, coordinating state via \StateClient\ interfaces. The system uses manifest-based synchronization to detect sequence number differences and transfer file changes (adds/deletes) between the master and slave nodes, ensuring data consistency across distributed deployments.

src/main/scala/ai/nixiesearch/index/sync · high confidence

Introduce native LLaMA.cpp-based generative model support

Added a new generative model backend that integrates the embedded LLaMA.cpp server to handle text generation and RAG-style prompting. This change introduces the \GenerativeModel\ trait and \LlamacppGenerativeModel\ implementation, which manages the local server lifecycle, tokenization, and streaming chat completions via the \ChatML\ protocol. The \GenerativeModelDict\ now supports loading models from HuggingFace or local directories, automatically detecting GPU availability to enable hardware acceleration, and tracking inference metrics. A new \SSEParser\ handles server-sent event streams for real-time token output.

src/main/scala/ai/nixiesearch/core/nn/model/generative · high confidence

Introduction of ONNX-based ranking model support

The system now supports ranking models via ONNX, introducing a new \RankModel\ trait and an \OnnxRankModel\ provider to handle document scoring. A \RankModelDict\ manages these models, including logic to prevent ONNX usage on GraalVM native images, while a \LogitsProcessor\ allows users to apply sigmoid or noop transformations to model outputs. This change enables the use of ONNX-based rerankers for improved search relevance.

src/main/scala/ai/nixiesearch/core/nn/model/ranking · high confidence

Introduction of a unified Query trait with JSON serialization support

A new \Query\ trait has been added to the \api.query\ package to serve as a common interface for search queries, specifically supporting \RetrieveQuery\ and \RerankQuery\ types. This change introduces JSON encoding and decoding logic for the \Query\ type, allowing the system to parse incoming requests that specify a single query key (either a retrieve type or a rerank type) and deserialize them into the appropriate concrete implementation. This provides a structured way to handle different query modes within the API layer.

src/main/scala/ai/nixiesearch/api/query · high confidence

New CLI argument converters for API mode and log level

The application now supports configuring the API serving mode (HTTP or Lambda) and overriding the log level via command-line arguments. This is enabled by new argument converters that parse these specific settings, allowing users to select the deployment mode and adjust logging verbosity directly from the CLI.

src/main/scala/ai/nixiesearch/main/args · high confidence

New Hugging Face model loading with caching and retry logic

Users can now load models directly from Hugging Face via a new client that implements robust error handling, including automatic retries for HTTP 429 rate-limit responses and proper handling of HTTP 302/307 redirects. The system also introduces a local file cache for downloaded model files, which reduces network usage and speeds up subsequent loads by storing files in a configurable local directory.

src/main/scala/ai/nixiesearch/core/nn/huggingface · high confidence

New REST API endpoints for search, indexing, and inference

The API module now exposes a comprehensive set of HTTP routes for interacting with the search engine. Users can perform full-text search and suggestions via POST /v1/index/{name}/search and /suggest, and manage index data (index, delete, flush, merge) via POST /v1/index/{name}. Administrative operations are available through GET /v1/system/config, /v1/system/telemetry, and /v1/index. The API also supports AI capabilities with POST /inference/embedding and /inference/completion (including streaming), and exposes system health at /health and Prometheus metrics at /metrics. Additionally, a web UI is served at the root path, and CORS is enabled for all origins.

src/main/scala/ai/nixiesearch/api · high confidence

New field codecs for structured data types

The search engine now supports indexing and querying a comprehensive set of structured field types, including Boolean, Date, DateTime, Integer, Long, Float, Double, and their respective list variants, as well as Geopoint and Text fields. These new codecs in the core field package handle the serialization of these types to and from JSON, and manage their storage in Lucene using appropriate field types (e.g., NumericDocValuesField for sorting, StoredField for retrieval, LatLonPoint for geospatial data). This enables users to define schemas with these specific data types and perform operations like filtering, sorting, and faceting on them.

src/main/scala/ai/nixiesearch/core/field · high confidence

New source reader abstraction for local, HTTP, and S3 data ingestion

A new \SourceReader\ trait and its \URLReader\ implementation have been added to handle reading byte streams from local files, HTTP URLs, and S3 buckets. This change introduces support for recursive directory reading on local and S3 sources, while HTTP sources remain non-recursive. The implementation leverages \fs2\ streams and integrates with existing S3 and HTTP clients, providing a unified interface for source data ingestion.

src/main/scala/ai/nixiesearch/util/source · high confidence

New structured query types and reranking capabilities

The search API now supports complex retrieval and reranking queries. Users can construct logical searches using \bool\ (must/must\_not/should), \dis\_max\, and \multi\_match\ (best\_fields/most\_fields) query types. Semantic search is available via the \semantic\ query type, which automatically embeds the query text, and direct vector search is supported via the \knn\ query type, which includes a configurable \num\_candidates\ parameter for performance tuning. Additionally, results can be reranked using Reciprocal Rank Fusion (\rrf\) to merge multiple retrieval results, or via Cross-Encoder (\cross\_encoder\) reranking, which uses a specified model and Jinja template to score and reorder documents based on their content.

src/main/scala/ai/nixiesearch/api/query/retrieve · high confidence

New utility modules for configuration, storage, and inference support

This change introduces a suite of new utility components in the \ai.nixiesearch.util\ package that enhance configuration parsing, external storage access, and AI inference capabilities. Users can now parse boolean environment variables with flexible aliases (e.g., 'y', 'yes') via \BooleanEnv\, and specify data sizes using shorthand units like '1g' or '1m' through the new \Size\ parser. The \Distance\ utility adds support for parsing and encoding geographic distances in various units (km, mi, ly, etc.), while \GPUUtils\ enables runtime detection of CUDA availability and NVIDIA GPU devices. For data persistence, \S3Client\ provides asynchronous, parallel S3 operations including multipart uploads and object listing. Additionally, \DocumentEmbedder\ introduces an in-memory caching mechanism for text embeddings, allowing pre-embedded fields to be handled efficiently and enabling embedding computation outside the main search loop. Supporting utilities like \ExpBackoffRetryPolicy\ for HTTP retries, \JsonUtils\ for strict JSON field validation, and \Version\ for manifest-based version detection are also included.

src/main/scala/ai/nixiesearch/util · high confidence

Support for hierarchical document grouping and merged facet collection

The search core now introduces \DocumentGroup\ to handle hierarchical relationships between documents (parent-child structures), automatically tagging parent and child Lucene documents with role fields and shared IDs for efficient querying. Additionally, \MergedFacetCollector\ has been added to aggregate facet data across multiple field collectors, allowing the system to compute unified facet results when multiple aggregations are requested.

src/main/scala/ai/nixiesearch/core/search · high confidence

Support for higher-dimensional embeddings and new quantization formats

The indexing engine now supports vector fields with up to 8192 dimensions and introduces Int1 (binary) quantization alongside existing Float32, Int8, and Int4 options. This is implemented via new Lucene codec wrappers (Nixiesearch101Codec and Nixiesearch103Codec) and a HighDimVectorFormat that enforces the dimension limit, allowing users to index larger embedding vectors and choose finer-grained compression for storage efficiency.

src/main/scala/ai/nixiesearch/core/codec/compat · high confidence

Support for term and range aggregations with flexible size and boundary parameters

Users can now define term and range aggregations in search queries. Term aggregations accept a size parameter that can be either an integer or the string 'all' to return all terms. Range aggregations support strict and non-strict boundary operators (gt, gte, lt, lte) for defining range facets, allowing for more precise filtering and bucketing of data.

src/main/scala/ai/nixiesearch/api/aggregation · high confidence

Behavioural changes

Default suggestion ranking strategy switched to Reciprocal Rank Fusion (RRF)

The suggestion ranking system in the core suggest module now defaults to Reciprocal Rank Fusion (RRF) for merging and ordering results from multiple suggestion sources (prefix, fuzzy, regex). Previously, the default behavior likely relied on a different ranking method (such as LTR or simple scoring); the new \SuggestionRanker\ implementation explicitly selects \RRFSuggestionRanker\ by default, while still allowing users to opt into LTR via rerank options. This change improves the quality of mixed suggestions by combining scores from different query types more effectively.

src/main/scala/ai/nixiesearch/core/suggest · high confidence

Introduce model handle abstraction for Hugging Face and local paths

The search engine now uses a unified \ModelHandle\ abstraction to identify neural network models, supporting both Hugging Face repository references (e.g., \namespace/name\) and local file directories. This change adds serialization and deserialization logic for these handles, allowing configuration files to specify models via standard Hugging Face identifiers or local file paths, laying the groundwork for flexible model sourcing.

src/main/scala/ai/nixiesearch/core/nn · high confidence

Introduce pluggable state client abstraction for index storage

The index store now uses a new \StateClient\ trait to manage index metadata and file operations, replacing the previous direct implementation. This change introduces three concrete backends: \DirectoryStateClient\ for local Lucene directories, \RemotePathStateClient\ for generic remote filesystem paths, and \S3StateClient\ for Amazon S3 storage (which now supports parallel fetching for files larger than 1MB). Users benefit from a unified interface for index state management that cleanly separates storage logic from the core indexing engine, enabling easier configuration of different storage backends via the \BlockStoreLocation\ configuration.

src/main/scala/ai/nixiesearch/index/store · high confidence

Introduce suggestion deduplication and ranking logic

Added a new ranking module for search suggestions that includes a default deduplication strategy (RRF) and a placeholder for LTR-based ranking. The RRF ranker aggregates scores from prefix, fuzzy, and regex candidate sources, deduplicates suggestions case-insensitively, and re-ranks them by combined score, ensuring users see unique and relevant suggestions by default. The LTR ranker is scaffolded but currently raises a 'not yet' error, indicating it is not yet functional.

src/main/scala/ai/nixiesearch/core/suggest/rank · high confidence

Major configuration schema overhaul with new inference, cache, and storage options

The application configuration structure has been significantly reorganized. The previous flat structure (api, store, index) is replaced by a hierarchical model (core, inference, searcher, schema). Users can now configure an in-memory embedding cache via the new \CacheConfig\ and \EmbedCacheConfig\, allowing for \memory\ or \none\ settings with adjustable sizes. Inference capabilities are exposed through \InferenceConfig\, supporting ONNX, OpenAI, and Cohere embedding providers, as well as Llama.cpp completion models with specific parameters like GPU layers and seed. Storage is now defined by \StoreConfig\, distinguishing between local (disk/memory) and distributed (S3/remote disk) backends. Additionally, the system now supports loading configuration from remote URLs (HTTP, S3, local files) and allows overriding core settings (host, port, log level, telemetry) via environment variables.

src/main/scala/ai/nixiesearch/config · high confidence

Major overhaul of index mapping configuration and field schema support

The index mapping configuration has been significantly restructured to support a wider variety of field types and more granular control over indexing behavior. Field names now support nested paths and wildcards, and the schema now includes dedicated types for Long, Double, and various list fields, alongside support for required fields and text suggestions. Search capabilities have been expanded to allow independent configuration of lexical and semantic (embedding) parameters, including support for cosine and dot distance metrics, multiple quantization levels (int1, int4, int8, float32), and per-field language analyzers. Additionally, users can now configure Lucene merge policies (tiered, byte-size, doc-count, or none) and choose between mmap or nio directory types for storage.

src/main/scala/ai/nixiesearch/config/mapping · high confidence

Native-image support and Lucene codec registration

The application now includes configuration files to support GraalVM native-image builds, specifically adding reachability and reflection metadata for the ai.nixiesearch module. Additionally, the Lucene codec service provider is registered to include both Nixiesearch101Codec and Nixiesearch103Codec, ensuring compatibility with Lucene 10.1 and 10.3. A dummy SLF4J service provider is also registered to replace logback, facilitating the native build process.

src/main/resources/META-INF · medium confidence

New CLI structure with subcommands and runtime detection

The application now uses a structured CLI with distinct subcommands: 'standalone' for the default server, 'index' for indexing data (supporting file, API, and Kafka sources), and 'search' for querying. Users can now override the logging level via the -l flag and specify the search API mode (HTTP or Lambda) for the search subcommand. The startup banner now displays detailed version information, including JDK, architecture, and GPU support status, and the application detects AWS Lambda and GraalVM native image runtimes to adjust behavior accordingly.

src/main/scala/ai/nixiesearch/main · high confidence

New embedding model infrastructure with pooling and caching support

The embedding subsystem has been restructured to support multiple provider types (ONNX, OpenAI, Cohere) through a unified \EmbedModelDict\ registry. This change introduces configurable pooling strategies (mean, CLS, last token) via \EmbedPooling\ and adds an in-memory caching layer (\MemoryCachedEmbedModel\) to optimize repeated embedding requests. Users benefit from more flexible model configuration, potential performance improvements via caching, and support for diverse embedding backends.

src/main/scala/ai/nixiesearch/core/nn/model/embedding · high confidence

New manifest model for tracking index file changes

The index manifest system now uses a new \IndexManifest\ structure that tracks individual index files by name and size, enabling precise detection of added or removed files. This change introduces a \diff\ method that compares the current manifest against a target state to produce a list of file operations (additions or deletions), which supports more accurate synchronization and state management for index data.

src/main/scala/ai/nixiesearch/index/manifest · high confidence

Refactored ONNX model loading with explicit configuration and device support

The ONNX model loading logic has been restructured to support explicit configuration via a new \OnnxConfig\ trait, allowing users to specify the execution device (CPU with thread count or CUDA with GPU ID) and token limits. The \OnnxSession\ now handles model loading from both Hugging Face and local directories, parsing \config.json\ and selecting the appropriate \.onnx\ file, while the \OnnxModelFile\ helper manages file selection logic and serialization. This change introduces a more robust and configurable foundation for ONNX-based inference, replacing the previous implicit loading behavior.

src/main/scala/ai/nixiesearch/core/nn/onnx · high confidence

Refactored document decoding to support wildcard fields and diverse data types

The document visitor logic in the codec layer has been rewritten to replace the previous field-specific writer traits with a unified decoding approach. This change introduces support for wildcard field matching, allowing patterns to match multiple fields during retrieval. It also expands the stored field types beyond text and integers to include long, float, double, and binary data, ensuring these types are correctly decoded from Lucene stored fields. Additionally, the visitor now validates that requested fields are defined and stored in the index mapping before attempting to read them, providing clearer error handling for missing or non-stored fields.

src/main/scala/ai/nixiesearch/core/codec · high confidence

Refactored index core with new Indexer, Searcher, and IndexStats components

The index module has been restructured to replace the legacy Index, IndexBuilder, IndexSnapshot, and IndexStream classes with a new architecture centered on Indexer, Searcher, and IndexStats. The new Indexer handles document ingestion, embedding, and Lucene IndexWriter operations, while the Searcher manages query execution, aggregation, and RAG streaming. Additionally, IndexStats provides detailed metrics on Lucene segments, leaves, and fields, exposing version and codec information for observability.

src/main/scala/ai/nixiesearch/index · high confidence

Replaces Logback with a custom SLF4J print logger

The application no longer relies on Logback for logging; instead, it provides a custom SLF4J implementation (PrintLogger) that outputs log messages directly to standard output with timestamps. This change allows the application to function without external logging configuration libraries, simplifying the runtime environment while maintaining standard SLF4J API compatibility.

src/main/java · high confidence

Separate CLI subcommands for Index, Search, and Standalone modes

The application now supports distinct operational modes via dedicated CLI subcommands. Users can run in 'index' mode to perform offline indexing from file or Kafka sources or expose an indexer-only REST API, in 'search' mode to run a searcher-only instance (including AWS Lambda support), or in 'standalone' mode where indexer and searcher are colocated in a single process. This change replaces the previous monolithic startup logic with specialized entry points for each use case, allowing users to optimize resource usage and deployment topology by choosing the specific mode that fits their needs.

src/main/scala/ai/nixiesearch/main/subcommands · high confidence

Updated web UI styles to Mantine v7

The web UI assets have been updated to use the Mantine v7 design system, introducing new CSS variables for theming, spacing, and typography. This change ensures the interface aligns with the latest Mantine component standards and visual language.

src/main/resources/ui · high confidence

Test coverage

Added API route tests for search, inference, and admin endpoints; Added GPU test resources for CUDA library stubs and device information; Added JSON decoding tests for search query components; Added JSON field decoding tests for all field types; Added JSON filter predicate tests; Added JSON serialization tests for aggregation queries; Added compatibility test fixtures for Lucene 10.1, 10.2, and 10.3; Added compatibility tests for Lucene 10.1, 10.2, and 10.3 indexes; Added comprehensive test coverage for query and search capabilities; Added document decoder benchmark comparing jsoniter and circe; Added end-to-end test for Kafka source integration; Added end-to-end test for ONNX ranking model; Added end-to-end tests for embedding, ranking, RAG, and index synchronization; Added quickstart test script for NixieSearch; Added test coverage for date/datetime field codecs and stored field roundtrips; Added test coverage for field sorting capabilities; Added test coverage for search filter predicates; Added test dataset and conversion script for geo search tests; Added test for ONNX session loading and unloading; Added test for OnStartAnalyticsPayload creation; Added test infrastructure for embedding generation and validation; Added test resources for sentence-transformers model; Added test utilities and unit tests for search infrastructure; Added tests for CLI configuration parsing and argument conversion; Added tests for Directory, Remote, and S3 state clients; Added tests for IndexManifest diff logic; Added tests for LlamaCPP generative model and SSE parser; Added tests for MemoryCachedEmbedModel caching behavior; Added tests for ModelHandle parsing; Added tests for ONNX bi-encoder embedding models; Added tests for SystemMetrics functionality; Added tests for URLReader across local, S3, and HTTP sources; Added tests for document JSON encoding and streaming; Added tests for index mapping configuration and field name parsing; Added tests for index statistics and searcher error handling; Added tests for index, search, and standalone modes; Added tests for local index storage backends; Added tests for master-slave index synchronization; Added tests for range and term aggregations; Added tests for suggestion analysis and candidate generation; Added tests for suggestion case sensitivity, deduplication, and multi-field support; Added tests for the model file cache; Expanded configuration validation tests for search, inference, and storage; Expanded test coverage for document field types and retrieval behaviors.

Dependencies

Major dependency and build system overhaul

The project has significantly updated its core dependencies, upgrading Scala to 3.7.4, cats-effect to 3.6.3, logback-classic to 1.5.20, and Lucene to 10.1, while introducing new libraries for Kafka, S3, Prometheus metrics, and JSON processing. The build configuration now supports GPU-accelerated builds via ONNX Runtime and LlamaCPP, utilizes JDK 25 with Ubuntu 24.04/25.10 base images, and includes a frozen requirements file for the MkDocs documentation site.

(dependencies) · high confidence

Update build tooling and core library dependencies

The build system has been upgraded to sbt 1.12.0, and several core libraries have been updated: http4s to 1.0.0-M46, circe to 0.14.15, circe-yaml to 0.16.1, fs2 to 3.12.2, AWS SDK to 2.40.7, Lucene to 10.3.2, and Prometheus metrics to 1.4.3. New dependencies for ONNX Runtime (1.23.2), DJL (0.35.1), llama.cpp (0.0.4-b5604), and jsoniter-scala (2.38.8) have been added. Additionally, sbt plugins for assembly, updates checking, Docker, build info, and JMH benchmarks have been introduced.

project · high confidence

Upgrade scalafmt to 3.10.3 and switch to Scala 3.6 dialect

The project's code formatter has been updated to version 3.10.3, and the configuration now targets the Scala 3.6 dialect (previously Scala 3). This change ensures that code formatting aligns with the latest scalafmt capabilities and the specific syntax features of Scala 3.6.

(repo-wide) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 63 → 54 (-9.3)
  • Rubric changed (rubric-2026.09.8 → rubric-2026.09.16) — scores are not directly comparable.

Lenses

  • Code Health 84 → 84 (+0.0)
  • Architecture 100 → 68 (-31.9)
  • Maturity 61 → 68 (+7.2)
  • Readiness 58 → 60 (+2.2)
  • Security 63 → 67 (+4.2)
  • Accessibility 42 (new)

Resolved (9)

  • Documentation: contradicts the code (docs/docs/tutorial/upgrade.md)
  • Documentation: hard to navigate (docs/docs/api.md)
  • Hotspot: src/main/scala/ai/nixiesearch/config/mapping/FieldSchema.scala (src/main/scala/ai/nixiesearch/config/mapping/FieldSchema.scala)
  • Hotspot: src/main/scala/ai/nixiesearch/index/sync/SlaveIndex.scala (src/main/scala/ai/nixiesearch/index/sync/SlaveIndex.scala)
  • Hotspot: src/main/scala/ai/nixiesearch/util/DocumentEmbedder.scala (src/main/scala/ai/nixiesearch/util/DocumentEmbedder.scala)
  • Medium IaC: CKV_K8S_21 (deploy/kubernetes/standalone/configmap.yaml)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (src/main/scala/ai/nixiesearch/config/mapping/FieldSchema.scala)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (src/main/scala/ai/nixiesearch/core/field/DoubleListFieldCodec.scala)
  • Off-boarding risk: anonymized user #1

New (12)

  • Members sharing a duplicated core (4 members, 50+ identical tokens) (src/main/scala/ai/nixiesearch/config/mapping/FieldSchema.scala)
  • Members sharing a duplicated core (4 members, 50+ identical tokens) (src/main/scala/ai/nixiesearch/core/field/DoubleListFieldCodec.scala)
  • No ADRs found
  • Off-boarding risk: anonymized user #1
  • Outdated: ch.qos.logback:logback-classic
  • Outdated: com.hubspot.jinjava:jinjava
  • Outdated: commons-codec:commons-codec
  • Outdated: commons-io:commons-io
  • Outdated: org.apache.kafka:kafka-clients
  • Outdated: org.typelevel:cats-effect_3
  • Outdated: org.typelevel:log4cats-slf4j_3
  • Projects may be oversized for their cohesion

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

nixiesearch/nixiesearch was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 31bc0a853a3d5f8421861201af4118834a584a27 — the exact code this score is about.
  • Scored under rubric-2026.09.16 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d46da229e3fd.