chroma-core/chroma
60.4
Adequate · 27 September 2026
330.6k
lines of production code
Rust
with Python, TypeScript
4
measurements over time
What this system is
Chroma is a vector database and embedding platform that manages collections of embeddings with support for dense and sparse vectors, hybrid search, and full-text indexing. It provides a distributed architecture with a Rust-based backend and Go coordinator, enabling multi-tenancy, conditional transactions, and scalable query execution across local and cluster deployments. The system offers client libraries for Python, JavaScript, and Rust, along with extensive tooling for deployment, authentication, and integration with various embedding providers and cloud infrastructure.
How it got here
2022–2023 — Rust backend and distributed architecture
48 changes.
This period focused on establishing the foundational infrastructure for Chroma, including repository scaffolding, CI/CD tooling, and dependency management. It involved a major architectural shift to a Rust-based backend with a new SQLite persistence layer, FastAPI server, and gRPC-based distributed system database. The work also introduced advanced features such as multitenancy, conditional transactions, and a modular segment architecture to support scalable, multi-node deployments.
2024 — distributed architecture and Rust backend
57 changes.
This period focused on establishing the foundational infrastructure for a distributed ChromaDB architecture, introducing a new Go-based coordinator and SysDB alongside a comprehensive Rust backend rewrite. Key developments included implementing distributed storage via Apache Arrow, adding Kubernetes deployment support with Helm charts, and integrating robust observability and authentication systems. The work also expanded the Python client with advanced execution engines and embedding support while laying the groundwork for future performance optimizations.
2025 — Rust service migration and client expansion
58 changes.
This period focused on the extensive migration of core Chroma services—including SysDB, the frontend, logging, and garbage collection—from Python to Rust, establishing a new high-performance backend architecture. Concurrently, the project expanded its client ecosystem by introducing a new JavaScript SDK with advanced query DSLs and embedding support, while also delivering a comprehensive Rust-based CLI for user interaction.
2026 — Foundation API and Rust Agent development
22 changes.
This period focused on building the Rust-native Foundation API and agent infrastructure, introducing core services for workspace initialization, wiki management, and agent-driven research. Significant work included establishing the Spanner schema for system and log services, implementing the Anthropic-based agent core with tool abstractions, and adding comprehensive test coverage for indexing and security components.
Features
Add Anthropic inference model with thinking and caching support
The agent now supports the Anthropic Messages API via a new \AnthropicAgentInferenceModel\ in \rust/agent/src/inference\. This change introduces a provider-agnostic inference interface (\AgentInferenceModel\) and concrete implementation for Anthropic, starting with the \claude-sonnet-4-5-20250929\ model. Key capabilities include support for interleaved thinking blocks (via the \interleaved-thinking-2025-05-14\ beta flag), automatic prompt caching using ephemeral markers to reduce costs, and system prompt injection. The implementation includes robust request diagnostics that redact sensitive content (prompts, reasoning, tool inputs) for safe logging, and parses responses into actions including reasoning, text, and tool calls. It also tracks detailed token usage including cache read/write tokens for billing. Only Anthropic is currently supported.
rust/agent/src/inference · high confidence
Add GCP deployment templates for Chroma
New Terraform files (main.tf and chroma.tfvars) have been added to the GCP deployment directory, enabling users to provision a Chroma server on Google Cloud Platform. The configuration sets up a Debian 11 Compute Engine instance with Docker and Docker Compose, configures firewall rules for SSH and HTTP access, and allows customization of the Chroma version (defaulting to 1.5.9), machine type, and OpenTelemetry settings via variables.
deployments/gcp · high confidence
Add Hugging Face server-side embedding example
A new example in the \examples/server\_side\_embeddings/huggingface\ directory demonstrates how to use Chroma with a Hugging Face Text Embeddings Inference server. The addition includes a \docker-compose.yml\ file that spins up both the Chroma server and the embedding server (defaulting to the \BAAI/bge-small-en-v1.5\ model), along with a \test.ipynb\ notebook showing how to connect the \HuggingFaceEmbeddingServer\ utility to a Chroma HTTP client for embedding generation.
_examples/server\_side\embeddings · high confidence
Add Node.js and Browser example applications for ChromaDB JS client
New example applications have been added to the \clients/js/examples\ directory to demonstrate how to use the ChromaDB JavaScript client in both Node.js and browser environments. The examples illustrate the two available package options: the bundled \chromadb\ package, which includes all embedding libraries as dependencies for ease of use, and the \chromadb-client\ package, which relies on peer dependencies to allow users to manage embedding library versions and keep their dependency tree lean. The provided code shows how to initialize the client, create collections, add and query documents, and list collections using both package variants.
clients/js/examples · high confidence
Add shared Go utility package for logging, process management, and hashing
A new \go/pkg/utils\ package has been introduced to centralize common utilities. This includes a configurable logger using Zerolog with JSON/console output and protobuf support, helpers for running processes and handling shutdown signals, Kubernetes client initialization for both test and production environments, integration test skipping logic based on environment variables, and a rendezvous hashing implementation for key distribution across members.
go/pkg/utils · high confidence
Add simple RBAC authorization provider
Users can now enable Role-Based Access Control (RBAC) via the new SimpleRBACAuthorizationProvider. This component reads a YAML configuration file to map users to roles and roles to specific actions, pre-processing these permissions at startup for efficient runtime checks. It enforces access control by granting or denying requests based on whether the authenticated user's assigned actions match the requested operation, returning a 403 Forbidden error for unauthorized attempts.
_chromadb/auth/simple\_rbac\authz · high confidence
Added OpenTelemetry tracing and metrics initialization with gRPC interceptor
The go/shared/otel package now provides a centralized implementation for distributed tracing and metrics collection using OpenTelemetry. It introduces a gRPC server interceptor that automatically creates spans for incoming RPC requests, enriches context with remote trace/span IDs from metadata, and records error details. Additionally, it includes an initialization function that configures both trace and metric exporters to send data via OTLP/gRPC to a specified endpoint, establishing the global tracer and meter providers for the service.
go/shared/otel · high confidence
Added WAL3 benchmark and bootstrap state-space reasoner examples
Two new example programs are now available in the \rust/wal3/examples\ directory to aid in development and testing of the WAL3 subsystem. The \wal3-bench\ tool provides a configurable benchmark for measuring write throughput and garbage collection performance against an S3-compatible storage backend, allowing users to tune parameters like target throughput and task limits. Additionally, \wal3-bootstrap-reasoner\ offers a static analysis tool that exhaustively explores the state space of the WAL3 bootstrap process (fragment states, manifest initialization, and recovery) to identify potential race conditions, panics, or invalid states, helping developers verify the robustness of the initialization logic.
rust/wal3/examples · high confidence
Added density-based retrieval relevance notebook
A new Jupyter notebook (\density\_relevance.ipynb\) has been added to the experimental directory to demonstrate a data-driven approach for estimating retrieval relevance. This tool helps users understand when a vector search system lacks sufficient information by computing a cumulative density function over the distances between points in a dataset. By evaluating the percentile of a query's result distance against this distribution, users can obtain a uniform, comparable, and automatically adapting measure of relevance, which is particularly useful for preventing performance degradation in retrieval-augmented generation applications.
chromadb/experimental · high confidence
Added official CLI installation scripts for Windows and Unix-like systems
New installation scripts (install.ps1 for Windows and install.sh for Linux/macOS) have been added to the CLI distribution. These scripts automate the download and setup of the Chroma CLI binary (version 1.4.4) from GitHub releases, handling platform-specific assets and providing guidance for adding the binary to the system PATH.
rust/cli · high confidence
Added s3heap benchmark example
A new benchmark example (\s3heap-benchmark.rs\) has been added to the \rust/s3heap\ examples directory. This tool allows users to measure the performance of the S3 heap writer by simulating load with configurable throughput and runtime parameters, providing insights into operation counts and task metrics under stress.
rust/s3heap/examples · high confidence
Added sparse vector benchmark with Block-Max WAND and MaxScore support
A new benchmark example (\sparse\_vector\_benchmark.rs\) has been added to evaluate sparse vector search performance. It compares the Block-Max WAND algorithm against a brute-force baseline using the Wikipedia SPLADE dataset. The benchmark supports various modes, including filtering, profiling via flamegraphs, and testing the new MaxScore reader/writer components. Users can run this example to measure search latency and throughput under different configurations such as document count, query volume, and top-k results.
rust/index/examples · high confidence
Added systemd service examples for Chroma CLI and Docker deployments
New systemd unit files and documentation have been added to the examples directory, providing ready-to-use configurations for running Chroma as a background service. Users can now deploy Chroma using either the native CLI (via \chroma-cli.service\) or Docker Compose (via \chroma-docker.service\), enabling automatic startup on boot and crash recovery without manual intervention.
examples/deployments/systemd-service · high confidence
Automated API client generation and local server startup scripts
The new JavaScript client now includes build scripts to automatically generate the API client code from the server's OpenAPI specification and start a local Chroma server for development or testing. The \gen-api.ts\ script launches the server, fetches the OpenAPI JSON, generates TypeScript types and client code using \@hey-api/openapi-ts\, and applies a specific fix to the \HashMap\ type definition to correctly handle nulls and arrays. Supporting scripts (\start-chroma.ts\, \start-chroma-jest.ts\, and \start-chroma-common.ts\) handle starting the Chroma server either via a pre-built Docker image or by building the Rust binary locally, ensuring the server is ready before client generation or tests begin.
clients/new-js/packages/chromadb/scripts · high confidence
Automated Chroma deployment with Docker and authentication support
A new startup script has been added to the common deployment examples to automate the setup of a Chroma server. The script installs Docker and its dependencies, clones the Chroma repository, and configures the environment based on provided authentication variables. It supports both basic and token-based authentication by generating the necessary .env files and server credentials, then launches the service using Docker Compose.
examples/deployments/common · high confidence
Chroma JS client library v2.1.0 with multi-tenant support and new embedding functions
The @chromadb and @chromadb-client packages have been updated to v2.1.0, regenerated from the Chroma 1.0.0 OpenAPI specification. This release introduces multi-tenant capabilities, including a new AdminClient for managing tenants and databases, and automatic tenant resolution from user identity. It also adds support for Spann index configuration, collection forking, and querying on a subset of IDs. Several new embedding functions are now available, including Cloudflare Workers AI, Mistral, Together AI, and an updated Google Gemini function that supports image inputs. The client now includes a dedicated ChromaQuotaExceededError for billing limits, sends a user-agent header, and enforces mandatory embeddings on upsert operations.
clients/js/packages/chromadb-core · high confidence
ChromaDB JS client now bundles embedding libraries and CLI
The main \chromadb\ package is now a 'thick' client that bundles all embedding libraries (such as \chromadb-default-embed\) and the CLI tool directly into the distribution. This simplifies installation by removing the need for users to manage peer dependencies for embeddings, allowing them to use embedding functions like \OpenAIEmbeddingFunction\ out of the box. The package also includes a bundled CLI entry point, making command-line operations available within the same package.
clients/js/packages/chromadb · high confidence
Embedding function configuration validation via JSON schemas
Chroma now includes a dedicated schemas module that defines JSON Schema Draft-07 specifications for embedding function configurations. This introduces a registry of available schemas and utility functions (such as \validate\_config\_schema\) that allow users and client libraries to strictly validate configuration inputs against these definitions, ensuring cross-language compatibility and preventing configuration errors through strict property validation.
_chromadb/utils/embedding\functions/schemas · high confidence
Foundation API introduces agent query budget debiting
The foundation-api service now tracks and bills agent usage by posting budget debits to the sync-frontend after each agent query completes. This new \budget\ module fetches a cached price card from the sync-frontend, calculates the cost in micro-USD based on input and output token counts for the planner and context-1 models, and posts the debit as a fire-and-forget background task. The system is designed to be fail-open: if the price card is unavailable or the debit fails, the user's response is still streamed and the error is only logged as a warning.
rust/foundation-api/src · high confidence
Initial Azure deployment templates for Chroma
Added Terraform configuration files (main.tf and chroma.tfvars.tf) to the Azure deployment directory, enabling users to provision a Chroma instance on Azure. The setup creates the necessary infrastructure including a Resource Group, Virtual Network, Subnet, NSG, Public IP, and a Debian 11 Linux Virtual Machine with Docker and Docker Compose installed. It configures security rules for SSH and HTTP (port 8000) access and allows customization of the Chroma version, authentication credentials, and OpenTelemetry settings via variables.
deployments/azure · high confidence
Initial CLI entry point and utility functions
This change introduces the initial Python-based CLI structure for Chroma. It adds a main entry point (\chromadb/cli/cli.py\) that handles the \update\ command to check for new versions via GitHub and delegates other commands to the Rust bindings. It also includes utility functions (\chromadb/cli/utils.py\) for managing log file paths and calculating directory sizes, which support CLI operations like vacuuming.
chromadb/cli · high confidence
Initial FastAPI server implementation with multitenancy and quota support
The server backend has been rewritten to use FastAPI, introducing support for multitenancy (tenants and databases) and configurable CORS. This change adds rate limiting and quota enforcement capabilities to the API, while also improving performance through async serialization with orjson and adding OpenTelemetry tracing for better observability.
chromadb/server/fastapi · high confidence
Initial Spanner schema for tenants, databases, and collections
This change introduces the foundational Spanner database schema for the Rust system database (sysdb). It adds 18 migration files that create the core tables: \tenants\, \databases\, \collections\, \collection\_compaction\_cursors\, \collection\_segments\, and \collection\_metadata\. The schema includes necessary unique and list indexes for efficient querying, interleaves child tables for data locality, and sets up default tenant and database records. Additionally, it denormalizes tenant and database IDs into the compaction cursors table and creates a Change Data Capture (CDC) change stream to support multi-region replication and billing tracking.
rust/spanner-migrations/migrations · high confidence
Initial chroma-load daemon stub
A new load-generation daemon named 'chroma-load' has been added to the Rust load module. It includes a default configuration file (chroma\_load\_config.yaml) specifying the service name, OpenTelemetry endpoint, and port 3001. The executable itself is currently a stub that prints a message and sleeps for 24 hours, serving as the foundational structure for the load generation tool.
rust/load · high confidence
Initial release of gRPC service definitions for Chroma components
This change introduces the complete set of Protocol Buffer (proto3) definitions for the Chroma system, establishing the gRPC interfaces for core services including the Coordinator, Log Service, Compactor, Query Executor, Work Queue, and Heap Tender. These IDL files define the data structures and RPC contracts for collection and database management, log operations (push, pull, fork, seal), vector search execution with ranking and grouping, and background task scheduling, serving as the foundational API contract for the Rust-based backend implementation.
idl/chromadb · high confidence
Initial release of the distributed-chroma Helm chart
This change introduces the \distributed-chroma\ Helm chart (version 0.1.93, app version 0.4.24) for deploying Chroma in a distributed Kubernetes environment. The chart includes configuration files (\values.yaml\, \values.dev.yaml\) and templates for core services including the Rust frontend, sysdb, log service, query service, compaction service, work queue, function consumer, garbage collector, and sysdb migration. It also supports optional components like the foundation service and MDAC service, and includes metadata for publishing to Artifact Hub.
k8s/distributed-chroma · high confidence
Initial repository scaffolding and development environment setup
This change establishes the foundational structure for the Chroma project, introducing essential configuration files including \.dockerignore\, \.gitignore\, \.gitattributes\, and \.pre-commit-config.yaml\ to standardize local development, build isolation, and code quality checks. It adds \Dockerfile\ and \docker-compose.yml\ to define the containerized runtime and local deployment, alongside \Tiltfile\ and \docker-bake.hcl\ to orchestrate multi-service builds for the Rust and Go components. Documentation is also initialized with \README.md\, \DEVELOP.md\, \CLAUDE.md\, and \AGENTS.md\ to guide contributors, while \bandit.yaml\ and \RELEASE\_PROCESS.md\ set up security scanning and release procedures.
(repo-wide) · high confidence
Introduce Arrow-backed blockfile storage implementation
This change adds a new production-grade blockfile provider backed by Apache Arrow, replacing the previous storage mechanism. It introduces \ArrowBlockfileProvider\ which manages \ArrowUnorderedBlockfileWriter\ and \ArrowOrderedBlockfileWriter\ instances, utilizing a \BlockManager\ and \RootManager\ for handling data persistence. The implementation includes a new serialization format for the sparse index (migrating from V1 to V1.2) and supports configurable concurrency for block flushes and loads. It also adds specific Arrow key implementations for boolean, float, string, and unsigned integer types to enable efficient columnar storage and retrieval.
rust/blockstore/src/arrow · high confidence
Introduce Arrow-based Block storage with binary search and serialization support
The blockstore now uses a new \Block\ type backed by Apache Arrow RecordBatches to store key-value pairs, replacing previous implementations. This change introduces efficient binary search lookups for sorted data and adds native serialization/deserialization support for \RecordBatchWrapper\ via \serde\_bytes\, enabling blocks to be easily persisted and transferred.
rust/blockstore/src/arrow/block · high confidence
Introduce Foundation API with workspace initialization and wiki management endpoints
This change introduces the \foundation-api\ service, providing a new set of HTTP endpoints for managing the Foundation product. It adds \POST /api/init\ to idempotently bootstrap a tenant's workspace (creating the database, wiki, and source collections) and \POST /api/tenants/{tenant}/foundations\ to create named Foundations. The API also includes wiki management routes: \POST /api/upsert-page\ and \POST /api/apply-patch\ for creating and patching wiki pages, \POST /api/search\ for hybrid dense+sparse search, and \POST /api/read-page\ to reconstruct full pages from chunks. Additionally, it provides \GET /api/tenants/{tenant}/foundations\ and \GET /api/tenants/{tenant}/foundations/{foundation}\ for listing and describing Foundations, along with trajectory and agent-related endpoints.
rust/foundation-api/src/routes · high confidence
Introduce Go build tooling and Dockerfiles for the coordinator and migration services
This change adds the foundational build infrastructure for the Go codebase, including a Dockerfile that installs protoc v31.1 and Go gRPC plugins, a Makefile with targets for generating protobuf code, building the coordinator binary, and running tests sequentially to avoid flakiness, and a separate Dockerfile for running Atlas schema migrations. It also provides a README detailing local Postgres setup and the required Protobuf toolchain versions.
go · high confidence
Introduce Go-based SysDB coordinator model definitions
This change introduces the core data models for the SysDB coordinator in Go, replacing or supplementing previous definitions. It defines the \Collection\ struct with fields for compaction metrics (total records, size bytes, last compaction time), versioning, and lineage, alongside \CollectionToGc\ for garbage collection tracking. It adds a comprehensive \collection\_configuration.go\ that formalizes vector index settings (HNSW and SPANN), embedding functions, and schema structures, including support for quantization and full-text search (FTS) indices. New models for \Database\, \Tenant\, \Segment\, and \Notification\ are also added to support multi-tenancy, database management, segment lifecycle, and internal coordination events.
go/pkg/sysdb/coordinator/model · high confidence
Introduce Kubernetes-backed memberlist manager for cluster coordination
The \go/pkg/memberlist\_manager\ package now provides a new component that maintains a consistent list of active cluster members by watching Kubernetes pods and persisting the state in a custom resource. The manager uses a Kubernetes informer to detect pod lifecycle events, filters for ready pods, and reconciles the resulting member list against a persistent store (backed by a \chroma.cluster/v1\ Custom Resource). This ensures that the system's view of available nodes is kept up-to-date and durable, supporting features like leader election and service discovery within the coordinator.
_go/pkg/memberlist\manager · high confidence
Introduce Rust HTTP client with retry, failover, and metrics
The Rust SDK now includes a new \ChromaHttpClient\ implementation that handles HTTP communication with the Chroma server. This client adds automatic retry logic with exponential backoff and jitter, supports read-only backend failover via multiple endpoints, and provides OpenTelemetry metrics for request latency and retry counts when the \opentelemetry\ feature is enabled. Configuration is managed through \ChromaHttpClientOptions\, allowing users to specify authentication methods (including Cloud API keys), custom endpoints, and retry behavior.
rust/chroma/src/client · high confidence
Introduce Rust SysDB implementation with CLI management tool
This change adds a new Rust-based implementation of the SysDB service, providing an in-memory test backend, a SQLite-based local storage backend, and a gRPC client backend, all unified under a configurable \SysDb\ enum. It also introduces a new \chroma-task-manager\ CLI binary that allows users to attach, detach, and query attached functions on a collection via the SysDB gRPC interface.
rust/sysdb · high confidence
Introduce Rust agent core with provider-agnostic tool and trajectory abstractions
The \rust/agent\ crate now provides the foundational agent driver, replacing the previous Python-based implementation with a Rust-native state machine that manages the infer-act-observe loop. This change introduces a composable behavior system via \AgentBehavior\ hooks, allowing middleware to intercept and modify inference contexts, tool calls, and observations without subclassing. It includes a provider-agnostic tool abstraction (\Tool\/\DynTool\) that automatically generates JSON schemas for model parameters and renders tool definitions into provider-specific formats (currently Anthropic only). The trajectory system records actions and observations, supporting reasoning blocks and tool results, and renders the full history into the Anthropic Messages API format. A dummy \GetWeatherTool\ is included to exercise the full pipeline, and error handling is centralized via \AgentError\ with specific variants for tool failures, HTTP errors, and configuration issues.
rust/agent/src · high confidence
Introduce Rust cache crate with Foyer-backed hybrid caching and partitioned concurrency
The \rust/cache\ crate now provides the core caching infrastructure, introducing a \Cache\ trait and a \PersistentCache\ trait that supports disk-backed storage via the Foyer library. Users can configure memory-only, disk-only, or hybrid caches through \CacheConfig\, with the hybrid option leveraging Foyer's multi-tier storage. The implementation includes an \AsyncPartitionedMutex\ to handle concurrent access efficiently, deterministic hashing by default for stable cache keys across restarts, and utility binaries (\cops-disk-cache-config-writer\, \cops-memory-cache-config-writer\) to generate YAML configurations. This change establishes the foundation for persistent, high-performance caching in the Rust backend.
rust/cache · high confidence
Introduce Rust client library with collection and conditional transaction support
The Rust client library now provides a comprehensive API for interacting with Chroma, centered around the \ChromaCollection\ handle for managing vector embeddings and metadata. Users can perform standard CRUD operations (add, get, query, update, delete) and manage collection metadata. A key addition is support for conditional (optimistic) transactions via \ConditionalCollectionTransaction\, allowing scoped read-check-write operations with automatic retry logic for conflict resolution. The client also exposes functionality for attached functions, including adding inputs to async attached functions, and supports various indexing configurations and search operators.
rust/chroma/src · high confidence
Introduce Rust configuration crate with Spanner connection settings and dependency registry
The new \rust/config\ crate provides a centralized configuration system for the Rust backend, introducing a \Registry\ service locator for dependency injection and a \Configurable\ trait for struct initialization. It adds detailed configuration structs for Google Cloud Spanner connections, including \SpannerSessionPoolConfig\ (tuned for production with 400 max sessions) and \SpannerChannelConfig\ (with gRPC channel and keep-alive settings), as well as a \SpannerEmulatorConfig\ for local development. A helper module adds serde utilities for serializing and deserializing \std::time::Duration\ as seconds.
rust/config/src · high confidence
Introduce Rust frontend with local and distributed execution modes
The Rust frontend now supports two execution backends: a LocalExecutor for single-node deployments using SQLite and in-memory segment managers, and a DistributedExecutor for multi-node clusters that routes requests via gRPC based on assignment policies and tier configurations. This change adds configuration options for executor selection, cache invalidation with TTL, and distributed client management, enabling users to deploy the Rust frontend in either standalone or distributed modes while maintaining feature parity with the Python implementation.
rust/frontend · high confidence
Introduce Rust log service implementation with gRPC and SQLite backends
This change adds a new Rust-based log service implementation located in \rust/log\, providing a unified \Log\ abstraction that can operate via gRPC against a remote log service or locally via SQLite. The implementation includes configuration for connection timeouts, message size limits, and memberlist-based client assignment, along with a local compaction manager that handles backfilling and purging dirty logs using SQLite and HNSW segments. An in-memory log variant is also provided for testing purposes.
rust/log · high confidence
Introduce Rust types crate with conditional transactions and schema validation
The \rust/types\ crate has been established as the central location for core type definitions, build logic, and validation rules. This includes the implementation of conditional transaction state management (buffering writes, tracking read tokens and id sets) and comprehensive schema validation for collections, including support for HNSW and SPANN index configurations, CMEK encryption keys, and FTS indices. The crate also introduces a build-time code generation system that parses Go constants to produce Rust operator constants, ensuring synchronization between the Go and Rust codebases without requiring Go source files in Docker builds.
rust/types · high confidence
Introduce Rust worker service binaries and configuration files
The rust/worker directory now includes the entry-point binaries for the query service, compaction service, function consumer, and work queue service, along with a CLI tool for interacting with the work queue. Configuration files (chroma\_config.yaml, chroma\_mcmr.yaml, chroma\_mcmr2.yaml) are added to define settings for these services, including S3 storage, block cache, and gRPC endpoints. A compaction client binary is also added to manage compaction jobs.
rust/worker · high confidence
Introduce Rust-based Chroma CLI with profile management and interactive commands
The CLI is now implemented in Rust, providing a native binary with version 1.4.4. This release adds a comprehensive set of commands including Browse, Copy, Db, Install, Login, Profile, Run, Update, and Vacuum, along with helpers for Docs and Support. Users can now manage Chroma Cloud profiles via a new config store that persists credentials in TOML and settings in JSON within the \~/.chroma directory. The interface features an interactive terminal with colored output, password masking, and clipboard integration, while also supporting local server management via the Run command.
rust/cli/src · high confidence
Introduce Rust-based Garbage Collector service and CLI tool
This change adds a new Rust-based garbage collector service and a corresponding CLI tool to the \rust/garbage\_collector\ directory. The service exposes a gRPC interface for manual garbage collection and runs a background loop to automatically clean up unused collection versions, attached functions, and logs based on configurable cutoff times and concurrency limits. It supports multiple storage backends (S3 and MCMR regions/topologies) and integrates with the existing SysDb and Log services. The CLI tool provides commands for manual collection, exporting/importing SysDb collections, and downloading collection files for debugging.
_rust/garbage\collector · high confidence
Introduce Rust-based Python bindings for Chroma
This change adds a new Rust implementation of the Python bindings (located in \rust/python\_bindings\), exposing core Chroma components such as the \Bindings\ class, \ConditionalTransaction\, and \ConditionalCommitPayload\ to Python via PyO3. The bindings integrate with the Rust frontend, SQLite storage, and logging systems, providing a native Rust-backed interface for collection operations, conditional transactions, and configuration management, while mapping Rust errors to corresponding Python exceptions.
_rust/python\bindings · high confidence
Introduce Rust-based SQLite persistence layer with schema migrations
This change adds a new Rust implementation for the SQLite storage backend, including the database connection logic, a migration system, and table definitions. It introduces a comprehensive set of SQL migrations for the sysdb (collections, segments, tenants, databases), metadb (embeddings, metadata, full-text search), and embeddings\_queue (log) subsystems. The schema supports scalar and array metadata, full-text search via FTS5, and multi-tenant database isolation. The Rust code handles migration application and validation, configuration via \SqliteDBConfig\, and provides helper functions for metadata upserts and deletions.
rust/sqlite · high confidence
Introduce S3-backed distributed heap for task scheduling
Added the \s3heap\ crate, providing a distributed, persistent heap data structure backed by S3 storage to enable scheduling and processing of tasks at scale. The implementation organizes tasks into one-minute time-based buckets stored as Parquet files, supporting concurrent writes via optimistic concurrency control (ETag-based retries) and automatic deduplication. Key components include \HeapWriter\ for adding tasks, \HeapReader\ for retrieval, and a configurable \HeapScheduler\ trait that allows users to define task completion logic, making it a foundational service for future task scheduling capabilities.
rust/s3heap · high confidence
Introduce S3-based metastore implementation with GCS interoperability support
This change adds a new S3 metastore implementation for SysDB, enabling the storage of collection version and lineage files in S3-compatible object stores. The implementation upgrades the AWS SDK to v2 and includes specific middleware to ensure compatibility with Google Cloud Storage (GCS) by recalculating V4 signatures, addressing header signing issues. It also introduces a new file path structure for version and lineage files organized by tenant, database, and collection, and provides test utilities using a pinned MinIO container for local development and testing.
go/pkg/sysdb/metastore/s3 · high confidence
Introduce basic and token-based server authentication providers
Added new server-side authentication providers for basic (username/password) and token-based (API key) authentication. The basic auth provider reads credentials from a file in htpasswd format and validates them using bcrypt, while the token auth provider supports reading users from a YAML configuration file or a single token, allowing tokens to be passed via the Authorization or X-Chroma-Token headers. Both providers include client-side counterparts to automatically attach the necessary credentials to outgoing requests.
_chromadb/auth/basic\_authn, chromadb/auth/token\authn · high confidence
Introduce conditional transactions and attachable functions for collections
The collection models in \chromadb/api/models\ now support collection-scoped conditional transactions, accessible via \Collection.conditional()\ and \AsyncCollection.conditional()\. These transactions allow users to buffer writes locally and commit them with optimistic conflict detection, including a \run()\ helper that automatically retries on conflicts. Additionally, the new \AttachedFunction\ model enables attaching functions to collections, allowing users to add input collections and manage function parameters, with support for both synchronous and asynchronous collection operations.
chromadb/api/models · high confidence
Introduce configurable blockfile provider abstraction
The blockstore module now exposes a unified \BlockfileProvider\ enum that abstracts over different storage backends, specifically supporting both Arrow-based persistent storage and an in-memory HashMap implementation. This change introduces a new configuration system (\BlockfileProviderConfig\) that allows users to select the provider type (Arrow or Memory) via configuration, with Arrow being the default. The provider handles reading, writing, clearing, and prefetching data through a common interface, delegating to the specific backend implementation, and supports configurable concurrency limits for block flushes and loads when using the Arrow backend.
rust/blockstore/src · high confidence
Introduce dedicated Rust API types module with OCC read tokens and error handling
A new \chroma-api-types\ crate has been added to the Rust codebase, consolidating API data structures into a separate module. This includes support for Optimistic Concurrency Control (OCC) transactional reads via \OccReadToken\ and \OccReadMode\, allowing clients to capture and pin reads to specific log snapshots. The module also defines error types for conditional write conflicts and stale reads, a payload structure for collection forking, a heartbeat response with nanosecond precision, and a user identity response exposing tenant and database access.
rust/api-types · high confidence
Introduce distributed LogService with gRPC client and retry logic
Added a new LogService implementation that acts as a distributed log producer and consumer via gRPC. The service connects to a remote log service host and port, applying OpenTelemetry tracing and a retry interceptor to handle RPC errors. It supports submitting embeddings (with record count telemetry) and pulling logs with a configurable timeout, while subscription methods are currently no-ops and deletion/purge operations are not yet implemented.
chromadb/logservice · high confidence
Introduce gRPC-based distributed system database (SysDB) implementation
Chroma now supports a distributed architecture by adding a new gRPC implementation for the System Database (SysDB), referred to as the 'Coordinator'. This change introduces a client (\GrpcSysDB\) that communicates with a remote coordinator service using generated protobuf stubs, featuring retry logic via \RetryOnRpcErrorClientInterceptor\ and OpenTelemetry tracing. It also provides a mock server implementation (\GrpcMockSysDB\) for testing purposes. This enables the frontend to offload system-level operations (such as managing tenants, databases, collections, and segments) to a dedicated service rather than relying solely on local state.
chromadb/db/impl/grpc · high confidence
Introduce in-memory blockfile storage for testing
The blockstore module now includes a new in-memory implementation (\MemoryBlockfileProvider\, \MemoryBlockfileReader\, \MemoryBlockfileWriter\) backed by \HashMap\ and \BTreeMap\ structures. This allows the system to use RAM-based blockfiles for testing purposes, supporting operations like \set\, \delete\, \get\, \get\_range\, \count\, \contains\, and \rank\ for various data types including strings, byte arrays, and custom records. This change is explicitly noted as not intended for production use and serves to facilitate local development and testing workflows.
rust/blockstore/src/memory · high confidence
Introduce mdac-service for provider token bucket rate limiting
A new internal HTTP service (\mdac-service\) is added to enforce rate limits on foundation-model providers (Baseten, Fireworks, Anthropic, OpenAI) using a token-bucket algorithm. The service exposes a \POST /api/v1/token-bucket/put-back-and-drain\ endpoint that accepts both single objects and arrays of requests, allowing callers to refund excess tokens and request new allowances in a single call. It is configured via YAML and environment variables, supports OpenTelemetry tracing, and is deployed via Tilt and Modal.
rust/mdac-service · high confidence
Introduce new JavaScript/TypeScript client package
Adds the new JS/TS client package for Chroma, providing a JavaScript interface to interact with the Chroma DB backend over REST. This release includes the core source files, build configuration via tsup (supporting both ESM and CJS outputs), Jest testing setup, and updated documentation with examples for Chroma Cloud and local development.
clients/new-js/packages/chromadb · high confidence
Introduce new Python API client architecture with async support and conditional HTTP transactions
The \chromadb/api\ package has been restructured to introduce a modernized client architecture. This change adds a fully asynchronous client (\AsyncClient\) alongside the existing synchronous \Client\, both of which now utilize \httpx\ for HTTP communication and support conditional HTTP transactions for safer concurrent updates. The update also introduces a new \collection\_configuration\ module that allows users to define and validate detailed index settings (such as HNSW and SPANN parameters) and embedding function configurations directly during collection creation or modification.
chromadb/api · high confidence
Introduce new Rust-based SysDb service with multi-region Spanner backend
A new Rust implementation of the SysDb service has been added, providing a gRPC server that handles database, tenant, and collection operations. The service introduces a backend abstraction layer that routes requests to Google Cloud Spanner instances based on topology prefixes found in database names, enabling multi-cloud multi-region (MCMR) support. It includes configuration loading from YAML and environment variables, OpenTelemetry tracing integration, and a dedicated entrypoint binary.
rust/rust-sysdb · high confidence
Introduce new execution engine with local and distributed executors
Chroma now uses a new execution engine for query operations, replacing the previous implementation. This change introduces an abstract \Executor\ interface and two concrete implementations: \LocalExecutor\ for single-node deployments and \DistributedExecutor\ for cluster environments. The \LocalExecutor\ handles count, get, and KNN queries by directly interacting with local segment managers, while the \DistributedExecutor\ routes these same operations to remote nodes via gRPC, featuring round-robin load balancing and automatic retry logic for fault tolerance. Users benefit from a more modular and scalable query execution layer that supports distributed deployments out of the box.
chromadb/execution/executor · high confidence
Introduce procedural macro for defining custom metering capabilities and contexts
The \chroma-metering-macros\ crate now provides a procedural macro, \initialize\_metering\, that allows users to define custom metering capabilities and contexts via attributes. By annotating traits with \\#\[capability\]\ and structs with \\#\[context\]\, the macro generates the necessary trait implementations, marker methods, and runtime support code (such as thread-local context management and async future instrumentation) within the user's project. This enables developers to create type-safe, extensible metering systems that integrate with the existing \chroma\_system\ components.
rust/metering-macros · high confidence
Introduce s3heap-service skeleton and gRPC client
Adds the initial skeleton for the s3heap-service, including a README detailing the heap/sysdb state invariants, a binary entry point for the heap-tender service, and a gRPC client configuration and implementation. The client is configured to connect to the heap service on port 50052 using memberlist-based service discovery and RendezvousHashing assignment to ensure colocation with the log service. The service code is currently disabled (defaulting to \enabled: false\) and marked as non-functional pending nonce-related logic reimplementation, but provides the structural foundation for future heap tender operations.
rust/s3heap-service · high confidence
Introduce shared common package with core types, constants, and error definitions
A new \go/pkg/common\ package has been added to centralize shared system definitions. This includes a \Component\ interface for lifecycle management, standard constants such as \DefaultTenant\ and the \SourceAttachedFunctionIDKey\ schema field, and a comprehensive set of typed error variables for tenants, databases, collections, segments, and attached functions. Additionally, a \UniqueID\ type wrapping UUIDs with parsing and serialization helpers is now available in \go/pkg/types\ to standardize identifier handling across the system.
go/pkg/common · high confidence
Introduce standalone Spanner migration tool with CLI and emulator bootstrap
The \rust/spanner-migrations\ directory now contains a complete, standalone Rust library and CLI for managing Spanner schema migrations. Users can now run migrations via the \spanner\_migration\ binary, which supports applying or validating migrations against specific topologies and filtering by slug (e.g., \spanner\_sysdb\). The tool includes a \generate-sum\ command to create integrity manifests (\migrations.sum\) for migration files. Additionally, it provides automatic bootstrapping for the Spanner emulator, creating the necessary instance and database if they do not exist, and handles legacy hash migrations for backward compatibility.
rust/spanner-migrations · high confidence
Introduce tiered client assignment and Kubernetes-based memberlist provider
The memberlist module now supports tiered client assignment, allowing gRPC clients to be distributed across priority tiers to isolate high-priority workloads. It also introduces a new CustomResourceMemberlistProvider that watches Kubernetes Custom Resources to dynamically maintain the list of active members, replacing previous static or manual discovery methods. This change includes the necessary configuration structures and assignment logic to map nodes to clients based on these new tiers.
rust/memberlist · high confidence
Introduce wal3, a new object-storage-backed write-ahead log library
Adds the wal3 crate, a new write-ahead logging library designed to run on object storage (such as S3) without external coordination. This library provides linearizable logging with separate reader and writer interfaces, supporting high-throughput single-writer scenarios while remaining correct under multiple writers. Key capabilities include fragment-based storage, manifest and snapshot management, three-phase garbage collection, cursors for position pinning, and optional replicated fragment uploads with read repair.
rust/wal3 · high confidence
Introduces a new segment-based architecture with pluggable storage backends
Chroma now uses a modular segment architecture that separates metadata, vector, and record storage into distinct, pluggable components. This change introduces new segment types including SQLite for metadata, HNSW (local memory, local persisted, and distributed) for vectors, and blockfiles for records. It also adds a distributed segment manager with Kubernetes-based memberlist support for multi-node deployments, an LRU cache for managing segment memory and file handles, and a safe unpickler to prevent arbitrary code execution during deserialization.
chromadb/segment · high confidence
Introduces a unified blockfile type system with configurable write ordering
The blockstore now exposes a standardized set of types for managing data persistence, including \BlockfileReader\, \BlockfileWriter\, and \BlockfileFlusher\ enums that abstract over both in-memory and Arrow-backed storage implementations. A key behavioral addition is the \BlockfileWriterOptions\ configuration, which allows users to specify \ordered\_mutations\ or \unordered\_mutations\ when creating writers, enabling performance optimizations for pre-sorted data. The module also defines the \Key\ and \Value\ traits with implementations for common types (such as \f32\, \u32\, \Vec\<f32\>\, \RoaringBitmap\, and \DataRecord\) and introduces specific error types like \BlockfileError\ to handle cases such as missing keys or blocks.
rust/blockstore/src/types · high confidence
Introduces configurable rendezvous hashing for member assignment
The assignment module now supports configurable assignment policies, defaulting to rendezvous hashing using the Murmur3 algorithm. Users can define assignment rules via configuration, which determines how keys are distributed across a set of members (e.g., database nodes or workers). This change provides a standardized, configurable mechanism for data placement that mirrors existing implementations in other language services.
rust/config/src/assignment · high confidence
Introduces core authentication and authorization interfaces
The chromadb/auth module now provides the foundational classes for server-side authentication and client-side header injection. It defines ServerAuthenticationProvider for validating requests and producing UserIdentity objects, ClientAuthProvider for generating auth headers, and a comprehensive set of AuthzAction enums (covering tenant, database, and collection operations) along with AuthzResource to support fine-grained authorization decisions.
chromadb/auth · high confidence
Introduction of SQLite-based embeddings queue storage
ChromaDB now supports an embeddings queue backed by SQLite, enabling durable, transactional processing of embedding operations. This change introduces two new database tables: \embeddings\_queue\, which stores pending operations with sequence IDs, timestamps, topics, IDs, vectors, encoding details, and metadata; and \embeddings\_queue\_config\, which holds configuration data as a JSON string. This provides a reliable mechanism for managing asynchronous embedding tasks.
_chromadb/migrations/embeddings\queue · high confidence
Introduction of SQLite-based persistent storage with connection pooling
ChromaDB now includes a new SQLite implementation (\chromadb.db.impl.sqlite\) that supports persistent storage via configurable connection pools (\LockPool\ for in-memory/shared-cache scenarios and \PerThreadPool\ for file-based persistence). This change introduces a migration system to manage schema evolution, integrates OpenTelemetry tracing for database operations, and adds a warning for users upgrading from versions prior to 0.5.6 to suggest vacuuming their database for optimization.
chromadb/db/impl · high confidence
Introduction of a lightweight HTTP-only Python client
Users can now install a separate, minimal package (chromadb-client) that provides only the HTTP client functionality for connecting to a remote Chroma server, without requiring the full local database or embedding dependencies. This 'thin client' is configured via a new build script and a dedicated flag file, allowing users to interact with Chroma via standard HTTP operations like creating collections and adding documents.
clients/python · high confidence
Introduction of a new telemetry event system with a no-op PostHog implementation
This change introduces a new telemetry infrastructure in the \chromadb/telemetry/product\ module, defining a \ProductTelemetryClient\ base class and specific event models (such as \ClientStartEvent\, \CollectionAddEvent\, and \CollectionQueryEvent\) that capture usage metrics like batch sizes and operation counts. The system includes a \Posthog\ implementation of the telemetry client, but the \capture\ method is currently a no-op (empty), meaning no telemetry data is actually sent to PostHog despite the dependency being present. This establishes the structural foundation for future telemetry features while keeping data collection disabled by default.
chromadb/telemetry/product · high confidence
Introduction of quota enforcement hooks
The system now includes a new quota enforcement component located in chromadb/quota. This introduces a QuotaEnforcer interface that allows the system to intercept and validate API actions (such as creating collections, adding embeddings, or querying) against defined limits. A default SimpleQuotaEnforcer implementation is provided which currently allows all requests, serving as a placeholder for future quota configuration and enforcement logic.
chromadb/quota · high confidence
Introduction of telemetry module with documentation
A new telemetry module has been added to the codebase, including a README that explains the structure: anonymized product telemetry collected by Chroma for usage analysis, and OpenTelemetry configuration for self-hosted metrics and traces that are not sent back to Chroma. The module's \_\init\\_.py file is also introduced, currently empty.
chromadb/telemetry · high confidence
JS client bundles native bindings and CLI entry points
The JavaScript client now includes native platform-specific bindings (for macOS, Linux, and Windows) and a bundled CLI command. Users can now run the \chromadb\ CLI directly from the package, which supports an \update\ subcommand to check for newer versions via npm registry, and the client automatically loads the correct native binary based on the operating system and architecture.
clients/js/packages/chromadb/src · high confidence
JS client now bundles the CLI for standalone distribution
The JavaScript client package now includes a bundled version of the CLI, allowing it to be distributed and executed as a single unit. This change is supported by the addition of a Cargo configuration that enables static CRT linking on Windows (x86\_64-pc-windows-msvc) to ensure the binary is self-contained, along with standard Node.js and macOS/Windows gitignore patterns to clean up the build environment.
_rust/js\bindings · high confidence
Leader election for SysDB
The SysDB service now supports leader election within Kubernetes clusters. A new leader election mechanism has been added to ensure high availability by allowing only one instance to act as the leader at a time. The service name is automatically extracted from the pod name, and the leader lock is configured with a 15-second lease duration, 10-second renewal deadline, and 2-second retry period.
go/pkg/leader · high confidence
New 'Chat with your documents' example using Chroma and OpenAI
Added a new self-contained example in the \examples/chat\_with\_your\_documents\ directory that demonstrates how to build a chat application over local documents. The example uses Chroma for vector storage and retrieval, and OpenAI's API for response generation. It includes sample State of the Union addresses (2022 and 2023) in the \documents\ folder, a \load\_data.py\ script to ingest text files into a Chroma collection, and a \main.py\ interactive chat interface. The chat interface now prompts the user to select a model (defaulting to gpt-4o-mini) and displays source document citations alongside the generated responses.
_examples/chat\_with\_your\documents · high confidence
New /api/agent endpoint for streaming agent-driven research
A new \POST /api/agent\ route has been added to the Foundation API, allowing users to initiate an agent-driven research loop that streams results via Server-Sent Events (SSE). The agent uses the Anthropic API to execute a sequence of inference steps, tool calls (such as search, read page, and sub-agent search), and observations, emitting structured events for reasoning, tool usage, and final answers. The endpoint supports custom system prompts and model selection, handles tenant and database scoping, and integrates with the existing metering system to track usage and debit queries.
rust/foundation-api/src/routes/agent · high confidence
New AWS EC2 deployment example using Terraform
Added a new example deployment for AWS EC2 using Terraform that provisions an Ubuntu 22 instance with Docker Compose, including an attached EBS data volume for persistence. The configuration supports both token and basic authentication, automatically generating credentials and configuring the Chroma server accordingly, while allowing users to specify the Chroma release version, region, and instance type.
examples/deployments/aws-terraform · high confidence
New Arrow serialization implementations for blockstore value types
The blockstore now includes dedicated Arrow serialization and deserialization logic for a wide range of value types, including DataRecord, f32, Vec\<f32\>, String, u32, Vec\<u32\>, RoaringBitmap, SparsePostingBlock, SpannPostingList, and QuantizedCluster. These new modules implement the ArrowWriteableValue and ArrowReadableValue traits, enabling the blockstore to efficiently write and read these types using Apache Arrow's columnar format, which supports optimized storage and retrieval for embeddings, metadata, and sparse structures.
rust/blockstore/src/arrow/block/value · high confidence
New CI tooling for test sharding, runtime reporting, and model caching
The CI infrastructure in bin/ci now includes several new scripts to improve build efficiency and observability. Test execution is optimized via pytest sharding (bin/ci/pytest-shard.py) and a path-based test selection script (bin/ci/determine-tests-to-run.sh) that runs specific language suites only when their code changes. Runtime observability is enhanced with a GitHub Actions runtime report generator (bin/ci/gha-runtime-report.sh) and a converter to Prometheus metrics (bin/ci/gha-runtime-prometheus.sh). Additionally, the default ONNX embedding model is now preloaded and cached in CI (bin/ci/preload\_default\_onnx\_model.py) to speed up embedding-dependent tests, and Python test inventory shards are merged (bin/ci/merge\_python\_test\_inventory.py).
bin/ci · high confidence
New CLI tools for log service cache management and fault injection
Added the \chroma-log-service-purge-cache-entry\ binary to allow administrators to manually evict specific cache entries (cursors, manifests, or fragments) from the Rust log service, and introduced the \chroma-fault\ CLI to inject controlled faults (delays or unavailability) into the log service for testing and debugging purposes.
rust/log-service · high confidence
New DigitalOcean Terraform deployment example
Added a new example deployment blueprint for DigitalOcean using Terraform, which provisions a Droplet with Ubuntu 22.04, mounts a persistent data volume, and configures a firewall for SSH and HTTP access. The configuration supports deploying Chroma via Docker Compose with optional authentication (token or basic auth) and allows users to customize instance size, region, and Chroma release version.
examples/deployments/do-terraform · high confidence
New Foundation MCP endpoint with OAuth and scoped access
A new Model Context Protocol (MCP) server is now available at \/mcp/foundation\ and \/mcp/tenants/{tenant}/foundations/{foundation}\. The bare path uses OAuth bearer tokens for tenant-scoped access, while the scoped path accepts Chroma API keys for direct access to specific Foundations. The endpoint supports search and page-reading tools, with OAuth discovery served at \/.well-known/oauth-protected-resource/mcp/foundation\.
rust/foundation-api/src/routes/mcp · high confidence
New Gemini document chat example
Added a self-contained example in the \examples/gemini\ directory that demonstrates how to build a chat application using Google Gemini and Chroma. The example includes a \load\_data.py\ script to ingest text documents into a Chroma collection and a \main.py\ script to query the collection and generate responses via the Gemini API, along with sample State of the Union addresses and documentation.
examples/gemini · high confidence
New Google Cloud Compute deployment example with authentication and external volumes
Added a new Terraform-based deployment example for Google Cloud Compute that provisions a VM with an attached persistent data volume. The example supports enabling authentication via either token or basic auth, automatically generating credentials and configuring the Chroma server accordingly, and includes instructions for SSH access and public IP retrieval.
examples/deployments/google-cloud-compute · high confidence
New JS embedding packages for BM25, Chroma Cloud Qwen, Chroma Cloud Splade, and Cloudflare Workers AI
The \@chroma-core/ai-embeddings\ workspace introduces four new embedding function packages for the JavaScript client. The \chroma-bm25\ package provides a local, Rust-compatible BM25 sparse embedding implementation with configurable tokenization and stopwords. The \chroma-cloud-qwen\ and \chroma-cloud-splade\ packages enable sparse and dense embeddings via Chroma's hosted cloud service, supporting the Qwen3 and Splade models respectively, with API key hydration from the Chroma client. The \cloudflare-worker-ai\ package adds support for Cloudflare's embedding models. All new packages are re-exported via the new \@chroma-core/all\ umbrella package for convenient installation.
clients/new-js/packages/ai-embeddings · high confidence
New JavaScript client introduces conditional transactions and administrative capabilities
The new JavaScript client adds support for collection-scoped conditional transactions, allowing users to perform read-modify-write operations with automatic conflict retrying via the \collection.conditional().run()\ API or manual commit workflows. It also introduces an \AdminClient\ for managing tenants and databases, and a \CloudClient\ for simplified authentication to hosted Chroma instances. Additionally, the client now supports Customer-Managed Encryption Keys (CMEK) for GCP, configurable read levels (INDEX\_AND\_WAL, INDEX\_ONLY, INDEX\_AND\_BOUNDED\_WAL) for controlling data visibility and latency, and a Next.js webpack helper to handle bundling compatibility.
clients/new-js/packages/chromadb/src · high confidence
New Movies sample app with chat and search capabilities
The sample\_apps/movies directory now contains a complete Next.js application featuring a dual-mode interface for interacting with a movies dataset. Users can switch between a 'Search' tab for direct keyword/hybrid queries and a 'Chat' tab for conversational assistance powered by the AI SDK. The chat interface utilizes a 'searchMovies' tool to perform hybrid retrieval (dense embeddings and BM25) via Chroma, displaying relevant movie results directly within the conversation stream. The app includes dedicated API routes for chat and search, UI components for tabs and input groups, and a library module for handling Chroma client connections and search logic.
_sample\apps/movies · high confidence
New Render.com Terraform deployment example with authentication and SQLite fixes
Added a new Terraform-based deployment example for Render.com that provisions a Chroma web service with persistent disk storage and configurable authentication (token-based by default). The example includes a patch to the Chroma source code to automatically handle older SQLite versions by swapping in pysqlite3-binary, ensuring compatibility on environments like Render where system SQLite may be outdated.
examples/deployments/render-terraform · high confidence
New Rust CLI with comprehensive command set
The CLI has been rewritten in Rust, introducing a full suite of new commands for managing Chroma instances and data. Users can now authenticate via \chroma login\ (supporting both browser-based and API-key flows) and manage multiple user profiles with \chroma profile\. Database administration is handled through \chroma db\ (create, delete, list, and connect with language-specific snippets). Data operations include \chroma copy\ for transferring collections between local and cloud environments, \chroma browse\ for interactive collection inspection, and \chroma vacuum\ for local storage optimization. The CLI also supports installing sample apps (\chroma install\), running a local server (\chroma run\), and self-updating (\chroma update\).
rust/cli/src/commands · high confidence
New Rust client examples for search, embeddings, and conditional transactions
Added three new examples in the Rust client library: \collection\_search.rs\ demonstrates comprehensive search operations including hybrid search with BM25 sparse vectors, filtered queries, and pagination; \embeddings.rs\ shows how to use Chroma Cloud embedding functions (Qwen for dense, Splade for sparse) to create collections and perform searches; and \conditional\_transactions.rs\ illustrates collection-scoped conditional transactions with automatic retry logic and manual commit support.
rust/chroma/examples · high confidence
New Rust index module with HNSW, SPANN, and Bitmap Full-Text search implementations
This change introduces the \rust/index\ crate, providing the core indexing infrastructure for the system. It includes an HNSW vector index provider with configurable caching, parallel loading, and garbage collection policies, as well as a new SPANN index implementation with its own configuration and provider. Additionally, a word-based full-text search index is added, utilizing a tokenizer and a bitmap-based writer/reader for efficient substring and trigram matching. The module also exposes configuration structures for these components, allowing users to tune parameters like cache sizes, parallelism, and GC policies.
rust/index · high confidence
New Rust segment implementation with sharded writers, bloom filters, and local HNSW support
The \rust/segment\ crate introduces a new Rust-based segment layer that replaces the previous Python-backed implementation. This change adds sharded record and metadata writers (\RecordSegmentWriter\, \blockfile\_metadata\) that support concurrent shard operations and integrate with a new \BloomFilter\ abstraction for efficient existence checks. It also introduces dedicated writers and readers for distributed HNSW (\DistributedHNSWSegmentWriter\) and Spann (\SpannSegmentWriterShard\) indices, as well as a \LocalSegmentManager\ that manages an in-memory HNSW index pool and handles legacy \max\_seq\_id\ migration from pickle files for local Chroma deployments.
rust/segment · high confidence
New SQL-based database mixins for embeddings queue and system database
This change introduces two new database mixin implementations: \SqlEmbeddingsQueue\, which uses SQLite to manage the ingestion queue for embeddings (handling submission, purging, and log deletion per collection), and \SqlSysDB\, which provides the SQL-backed implementation for system database operations such as creating, retrieving, listing, and deleting databases and tenants. These mixins allow ChromaDB to use a traditional relational database as the primary ingest queue and system state store, replacing or supplementing previous mechanisms.
chromadb/db/mixins · high confidence
New SysDB coordinator implementation with async function support and heap service integration
The \go/pkg/sysdb/coordinator\ package introduces a new \Coordinator\ component that manages system catalog operations, including database, tenant, and collection lifecycle. This implementation adds support for asynchronous attached functions, featuring a repair flow for completion offsets and a status checking mechanism. It also integrates with a heap service for scheduling, using gRPC and memberlist-based discovery, and exposes metrics for compaction dead-letter queues.
go/pkg/sysdb/coordinator · high confidence
New admission control and rate-limiting primitives in the MDAC library
The \rust/mdac\ crate now exposes core building blocks for managing system load: a \CircuitBreaker\ that rejects requests when saturation is detected, a \TokenBucket\ for precise GCRA-based rate limiting with token refunds, a \Scorecard\ for multi-dimensional traffic tracking and limiting based on labeled rules, and a \Pattern\ utility for glob-style matching used by the scorecard. These components provide the foundational logic for admission control and rate limiting within the MDAC service.
rust/mdac · high confidence
New agent tools for searching and reading wiki pages
The foundation-api now exposes \search\, \read\_page\, and \subagent\_search\ tools for use by AI agents. These tools allow agents to perform hybrid dense+sparse retrieval against the wiki, fetch full page content by slug, and delegate deep research to an external subagent, all while reusing the same underlying retrieval cores as the standard HTTP routes to ensure consistent results.
_rust/foundation-api/src/agent\tools · high confidence
New benchmark infrastructure with multi-dataset support
The benchmarking suite has been restructured to support a modular, extensible dataset system. A new \RecordDataset\ trait and associated utilities handle the asynchronous download, caching, and streaming of data from various sources, including vector datasets (Sift1M, Gist), text corpora (Wikipedia, SciDocs, Microsoft MARCO), and code repositories (The Stack Dedup). The suite now includes specific dataset loaders for these sources, a utility for managing cached dataset files with safe file-renaming, and a \FrozenQuerySubset\ mechanism to pre-compute and cache valid query-corpus pairs for consistent benchmarking.
rust/benchmark · high confidence
New deep-research subagent search endpoint
A new \POST /api/subagent\_search\ route has been added to the Foundation API. It accepts a research query and streams results back via Server-Sent Events (SSE). The endpoint forwards the request to an external deep-research dependency, parsing the raw agent events into typed progress updates (\action\, \observation\) and a final structured \result\ containing ranked documents with justifications. It handles authentication via the caller's Chroma token and tenant scope, and ensures that empty results are treated as valid outcomes rather than errors.
_rust/foundation-api/src/routes/subagent\search · high confidence
New embedding functions and unified registration system
Chroma now includes built-in support for several new embedding providers and sparse vector methods, including Amazon Bedrock, Baseten, Cloudflare Workers AI, Perplexity, Morph, Nomic, and Mistral, alongside new sparse embedding functions for BM25 (both a pure Python implementation and a Fastembed-based one) and Chroma Cloud Splade. To manage this expanded set, the library introduces a centralized registration system in the embedding functions module, exposing a \register\_embedding\_function\ decorator for custom functions and dictionaries (\known\_embedding\_functions\, \sparse\_known\_embedding\_functions\) that map string identifiers to the available implementations, simplifying how users discover and instantiate embedding capabilities.
_chromadb/utils/embedding\functions · high confidence
New example notebooks for authentication, embedding providers, and filtering
The examples/basic\_functionality directory now includes several new Jupyter notebooks that demonstrate key Chroma capabilities. The auth notebook provides a comprehensive guide to setting up and using basic and token-based authentication in client/server deployments. The alternative\_embeddings notebook shows how to integrate third-party embedding providers like OpenAI, Cohere, and Instructor models. The in\_not\_in\_filtering notebook demonstrates the new $in and $not\_in metadata filtering operators, including their interaction with logical operators like $and and $or. Additionally, new notebooks cover local persistence with PersistentClient, basic embedding retrieval using the SciQ dataset, and retrieving collections by ID.
_examples/basic\functionality · high confidence
New examples for conditional transactions, attached functions, and work queue configuration
The examples folder now includes three new demonstration files: \conditional\_transactions.py\ shows how to use the new \collection.conditional()\ API for atomic create-or-update operations and manual transaction commits; \task\_api\_example.py\ demonstrates the new \attach\_function\ and \detach\_function\ methods for automatically processing collections with attached functions like \RECORD\_COUNTER\_FUNCTION\; and \work\_queue\_config.yaml\ provides a sample configuration for the new work queue service, including storage paths, persistence thresholds, and S3 settings.
examples · high confidence
New execution system with CPU/IO affinity and OpenTelemetry metrics
The system crate now includes a new execution framework (under \src/execution/\) that introduces a Dispatcher to manage task distribution to worker threads and an IO runtime. This change adds support for CPU and IO core affinity, allowing worker threads to be pinned to specific CPU cores and IO tasks to be pinned to high-numbered cores, with configuration options \cpu\_affinity\_num\_cores\ and \io\_affinity\_num\_cores\ in \DispatcherConfig\. The new system also integrates OpenTelemetry metrics for tracking dispatcher queue depths, task latencies, worker request counts, and other operational statistics across the Dispatcher, WorkerThread, Orchestrator, and Scheduler components.
rust/system · high confidence
New expression-based search API with hybrid ranking and grouping
The \chromadb/execution/expression\ module introduces a new internal DSL for defining search operations, replacing ad-hoc parameter passing with structured expression objects. Users can now construct complex search queries using a builder pattern or direct instantiation, supporting logical filtering via \Where\ expressions (with MongoDB-style operators like \$and\, \$or\, \$in\), hybrid ranking through \Rank\ expressions (including KNN and Reciprocal Rank Fusion), and result grouping via \GroupBy\. This change provides a more robust and type-safe interface for specifying search parameters such as limits, projections, and filters, laying the groundwork for advanced query capabilities in the Python client.
chromadb/execution/expression · high confidence
New forking example and hardware-optimized image documentation
The examples/advanced directory now includes a new forking.ipynb notebook that demonstrates how to fork a ChromaDB collection to chunk and embed a GitHub repository, apply diffs to a new branch, and use custom code chunking with tree-sitter. Additionally, a hardware-optimized-image.md guide explains how to rebuild the hnswlib from source with AVX support for Intel-based CPUs by setting the REBUILD\_HNSWLIB build argument during Docker image creation.
examples/advanced · high confidence
New gRPC server utility package with mTLS and distributed tracing support
A new \go/pkg/grpcutils\ package has been introduced to centralize gRPC server configuration and lifecycle management. This utility allows services to configure connection limits via \MaxConcurrentStreams\ and \NumStreamWorkers\, and enables mutual TLS (mTLS) authentication when certificate paths are provided. It also integrates optional distributed tracing via the OpenTelemetry interceptor, activated by the \OPTL\_TRACING\_ENDPOINT\ environment variable, and provides standardized helpers for building gRPC error responses.
go/pkg/grpcutils · high confidence
New gRPC service layer for SysDB collection, segment, task, and tenant management
This change introduces the Go gRPC server implementation for the SysDB coordinator, exposing RPC endpoints for managing collections, segments, attached functions (tasks), and tenants/databases. The \collection\_service.go\ and \segment\_service.go\ files handle creation, retrieval, and deletion of vector collections and their internal segments, including metadata conversion and error mapping. The \task\_service.go\ file adds RPCs for attaching functions, managing inputs, and handling async invocation lifecycles. The \tenant\_database\_service.go\ file provides endpoints for CRUD operations on tenants and databases, including soft-delete completion and resource name management. Supporting files include \proto\_model\convert.go\ for bidirectional conversion between protobuf and internal models, and comprehensive test suites (\\\_test.go\) for these services and memberlist configuration.
go/pkg/sysdb/grpc · high confidence
New generative benchmarking sample app for evaluating embedding models
A new sample application located in \sample\_apps/generative\_benchmarking\ provides a toolkit for generating custom benchmarks and comparing embedding model performance. The app includes a configuration file (\config.json\) specifying required environment variables (such as \OPENAI\_API\_KEY\) and startup commands for pip, poetry, and conda. It provides a \compare.ipynb\ notebook that allows users to load and compare metrics (e.g., Recall@3) from different embedding providers like OpenAI, Jina, and Voyage, alongside a \generate\_benchmark.ipynb\ for creating custom benchmarks based on user data. The sample also ships with example data (\chroma\_docs.json\) to immediately test the notebooks.
_sample\apps · high confidence
New ingest stream interfaces for embedding producers and consumers
The \chromadb/ingest\ module now exposes abstract \Producer\ and \Consumer\ components that define the contract for writing and reading embedding records to an ingest stream. The \Producer\ interface includes methods for submitting single or batched \OperationRecord\ embeddings, purging logs, and retrieving the maximum batch size, while the \Consumer\ interface allows subscribing to collection logs with optional start/end sequence ID ranges and callback functions. Utility functions for encoding and decoding vectors using NumPy are also provided to support these interfaces.
chromadb/ingest · high confidence
New integration examples for Cohere, Jina, Ollama, and Roboflow
Added usage examples in the \examples/use\_with\ directory for integrating Chroma with external embedding services. This includes JavaScript and Python notebooks for Cohere (covering basic, multilingual, and multimodal image embedding), a Python notebook for Jina AI demonstrating late chunking, a Markdown guide for local Ollama embedding, and a Python notebook for Roboflow image embeddings using CLIP.
_examples/use\with · high confidence
New interactive TUI collection browser in the CLI
The CLI now includes a new \collection\_browser\ module under \rust/cli/src/tui\ that provides an interactive terminal UI for browsing Chroma collections. This feature introduces a table-based view of records (ID, Document, Metadata) with keyboard navigation, support for dark/light themes and true-color/ansi256 palettes, and a query editor for filtering by IDs, document content, and metadata operators. Users can expand individual cells to inspect full content, paginate through results, and submit searches, with the UI handling async record loading and error states.
rust/cli/src/tui · high confidence
New local development environment with observability and storage services
The \k8s/test\ directory now provides a comprehensive set of Kubernetes manifests to run a local testing and debugging stack on top of production configurations. This includes deployments and services for Grafana (with a pre-configured 'foyer' dashboard), Jaeger (configured with Badger storage to prevent OOMs), Prometheus, MinIO, and a Spanner emulator. It also introduces a custom Postgres deployment that supports multiple databases (\sysdb\, \log\) via a custom entrypoint script, with resource limits tuned to handle connection storms. Additionally, an OpenTelemetry collector is included to bridge traces to Jaeger and metrics to Prometheus, and initial Custom Resource definitions for a distributed 'memberlist' feature are added.
k8s/test · high confidence
New mdac-service with token bucket rate limiter and Modal deployment
Adds the mdac-service, a new Rust-based HTTP service implementing a token bucket rate limiter. The build system includes a dedicated Dockerfile (rust/Dockerfile.mdac) to compile the token\_bucket\_service binary, and a Python deployment script (rust/deploy\_mdac.py) that packages the service for deployment on Modal, exposing it as a web server on port 8000.
rust · high confidence
New metering library with request timing and detailed context tracking
The \rust/metering\ crate has been introduced to centralize usage tracking, providing structured contexts for collection operations (fork, read, write) and search agent usage. This change adds request timing capabilities by capturing start and finish instants to calculate execution time in milliseconds, and enriches metering events with specific metrics such as FTS query length, metadata predicate count, query embedding count, and data transfer sizes (pulled and returned bytes). The library also includes a dedicated receiver mechanism for submitting meter events and exports the necessary types and contexts for integration with the rest of the system.
rust/metering · high confidence
New multimodal retrieval example notebook
Added a new example notebook in the \examples/multimodal\ directory that demonstrates how to create and query a Chroma collection containing both text and images. The notebook installs the \open-clip-torch\ library and shows how to use Chroma's built-in multimodal features to embed and retrieve data across these two modalities.
examples/multimodal · high confidence
New operational tooling for testing, deployment, and CI
This change introduces a suite of new scripts and utilities in the bin directory to support Chroma's testing, deployment, and continuous integration workflows. It adds entrypoint scripts (docker\_entrypoint.sh, docker\_entrypoint.ps1) that standardize server startup and enforce environment variables like CHROMA\_SERVER\_NOFILE. New integration test runners (python-integration-test, rust-integration-test.sh, ts-integration-test.sh, test-remote) allow running tests against local, remote, or containerized instances, including specific support for Rust client backward compatibility and TypeScript client generation checks. Deployment is supported by a CloudFormation generator (generate\_cloudformation.py) for single-instance AWS EC2 setups. CI is enhanced with scripts for Windows SQLite upgrades, log collection (get-logs.sh), and operator constant generation, alongside a version helper and a basic sanity check script.
bin · high confidence
New reasoning trajectory persistence module with Chroma storage
A new \trajectories\ module has been added to the \foundation-api\ to persist user-facing reasoning projections as structured records in Chroma. This module introduces a dedicated data model that prunes raw producer metadata (tools, parameters, observations) down to visible reasoning text, derived page-write facts, and citation attribution. It provides an API for creating open trajectories, appending entries with optimistic concurrency checks, finalizing trajectories, and loading them back. The implementation handles storage details such as chunking large JSON payloads, generating stable base36 identifiers, and managing citation sub-trees, exposing these capabilities through the \foundation-api::trajectories\ public interface.
rust/foundation-api/src/trajectories · high confidence
New search expression DSL for the JavaScript client
The JavaScript client now includes a new expression-based DSL in the execution layer, allowing users to construct complex search queries using a fluent API. This addition introduces typed classes for where clauses (supporting comparison, logical, and regex operators), ranking expressions (enabling arithmetic and mathematical operations on scores), grouping (with min/max aggregation), and result selection. Users can now build search payloads programmatically with strong typing and validation, replacing or supplementing raw JSON objects for query construction.
clients/new-js/packages/chromadb/src/execution · high confidence
New shared OpenTelemetry tracing and configuration library
The \rust/tracing\ crate now provides a centralized library for OpenTelemetry integration, introducing a shared \OpenTelemetryConfig\ struct for managing service names and trace filters. It includes \GrpcClientTraceLayer\ and \GrpcServerTraceLayer\ to automatically propagate trace and span IDs via \chroma-traceid\ and \chroma-spanid\ headers across gRPC boundaries, and a Tower middleware for HTTP services that extracts W3C trace context and records request details. The library also initializes global tracing with a Rustls crypto provider, exports Tokio runtime metrics, and adds a utility to log operations that exceed a specified duration threshold.
rust/tracing · high confidence
New utility modules for async conversion, batch handling, and data conversion
Added a suite of new utility modules to chromadb/utils. This includes async\_to\_sync decorators for bridging async and sync APIs, batch\_utils for splitting large insertions into chunks based on max\_batch\_size, and results utilities to convert QueryResult and GetResult objects into pandas DataFrames. Additional utilities provide rendezvous hashing for replication assignment, LRU caching with eviction callbacks, image loading via PIL, and functions for managing collection statistics (attach, detach, get) with support for filtering by metadata keys.
chromadb/utils · high confidence
New wiki page upsert flow with tree-sitter chunking and SPLADE sparse embeddings
The foundation-api wiki module now supports a new \/upsert-page\ flow that processes wiki content using a tree-sitter-based markdown chunker (splitting pages into chunks capped at 4096 bytes) and generates SPLADE sparse embeddings for each chunk via Chroma Cloud. Page metadata is enriched with author and last-written-by stamps, and source IDs are distributed across chunks to stay within storage limits.
rust/foundation-api/src/wiki · high confidence
New xAI integration example for document-based chat
Added a new example in the \examples/xai\ directory demonstrating how to use Chroma with the xAI SDK. This example implements a Retrieval-Augmented Generation (RAG) workflow where users can add PDF documents to a \docs\ folder, which are then chunked, embedded, and stored in a Chroma collection. The application allows users to chat with their documents by querying the collection for relevant context to answer questions via the xAI API.
examples/xai · high confidence
OpenTelemetry distributed tracing support
Added OpenTelemetry instrumentation to Chroma, enabling distributed tracing for server operations, FastAPI HTTP requests, and gRPC client calls. Users can configure the service name, collection endpoint, and span granularity via settings to export trace data to external backends.
chromadb/telemetry/opentelemetry · high confidence
Rust client adds BM25 sparse embeddings and Ollama integration
The Rust client now supports sparse vector embeddings via a new BM25 implementation (using a configurable tokenizer and MurmurHash3 hasher) alongside existing dense embedding capabilities. Users can generate BM25 sparse vectors locally or use the new Ollama embedding function to connect to a local Ollama instance for dense embeddings. Additionally, the module introduces an \EmbeddingFunction\ trait to standardize text-to-vector conversion and includes Chroma Cloud embedding functions for Qwen (dense) and Splade (sparse) models.
rust/chroma/src/embed · high confidence
Spanner schema for Rust Log Service manifests and fragments
The Spanner database schema for the Rust Log Service has been established with a set of migrations defining the core data structures. This includes the \manifests\ table to track log metadata, the \fragments\ table (interleaved under manifests) to store file segment details, and supporting \manifest\_regions\ and \fragment\_regions\ tables for region tracking. Subsequent migrations refine this structure by dropping the \collected\ column from manifests, adding \initial\_offset\ and \intrinsic\_cursor\ fields to region tables, introducing an \ignore\_dirty\ flag for manifests, adding an \updated\_at\ timestamp, and creating a specific index on fragments to optimize queries by position limit.
_rust/spanner-migrations/log\migrations · high confidence
Unified storage backend with admission-controlled S3 and GCS support
The storage layer now provides a unified interface that supports S3, local filesystem, and Google Cloud Storage (GCS) via a generic object-store backend. A new admission-controlled S3 implementation (\AdmissionControlledS3Storage\) is introduced as the default storage configuration, featuring rate limiting, request coalescing, and priority-based handling to manage load. The system now includes comprehensive OpenTelemetry metrics for all storage operations (get, put, delete, copy, rename, list) and introduces a shared \Stopwatch\ utility for consistent latency tracking across the codebase.
rust/storage · high confidence
Vector distance calculations now use hardware-accelerated SIMD implementations
The \rust/distance\ module now provides optimized distance metrics (cosine, inner product, and Euclidean) using AVX, AVX512, NEON, and SSE instruction sets, falling back to scalar implementations when hardware support is unavailable. This change significantly improves vector search performance on supported CPUs by leveraging SIMD parallelism, with automatic feature detection ensuring compatibility across different processor architectures.
rust/distance · high confidence
Architecture
Extract shared frontend-core library crate
The \frontend-core\ library crate has been extracted to provide shared scaffolding for HTTP frontends, allowing binaries like \chroma-frontend\ and \foundation-api\ to embed common components. This includes a new admission control service (\ac.rs\) that uses a circuit breaker to rate-limit requests, a centralized collection creation planner (\collection\_ops.rs\) that handles segment dispatch and schema reconciliation, and shared logic for attached function operations (\attached\_function\_ops.rs\). The crate also introduces a unified authorization action enum (\auth.rs\) covering system, database, collection, and foundation actions, along with shared HTTP routes for system endpoints (healthcheck, heartbeat, version) and common middleware for JSON error handling.
rust/frontend-core · high confidence
Introduce Atlas-based schema generation and GORM dbmodel layer for SysDB
The SysDB metastore now uses Atlas with the GORM provider to generate database schemas from Go structs, replacing the previous manual or ad-hoc approach. This change introduces a new \dbmodel\ package containing GORM models and database interfaces for core entities including Collections, Databases, Segments, Functions, and Attached Functions, along with their metadata. It also adds autogenerated mock implementations for these interfaces to support testing, and defines constants for built-in function IDs and names to ensure consistency between the Go backend and Rust types.
go/pkg/sysdb/metastore/db/dbmodel · high confidence
Introduction of a new database abstraction layer and system database interface
This change introduces a new internal database architecture for Chroma. It adds a \DB\ interface in \chromadb/db/\_\init\\_.py\ that defines the core vector operations (add, get, delete, update, nearest neighbors) and a \SqlDB\ base class in \chromadb/db/base.py\ that provides a consistent DBAPI 2.0 wrapper and PyPika query builder utilities. It also introduces a \MigratableDB\ class in \chromadb/db/migrations.py\ to handle SQL schema migrations with versioning and hash validation. Additionally, a new \SysDB\ interface in \chromadb/db/system.py\ is added to manage system-level resources, including databases, tenants, collections, and segments, supporting features like multitenancy and segment-based storage.
chromadb/db · high confidence
Behavioural changes
Added utility functions for topic parsing and vector segment migration
A new \utils.py\ module has been added to the ingest implementation, introducing helper functions to parse and construct Pulsar topic names (extracting tenant, namespace, and topic components) and a specific migration trigger for vector segments. The \trigger\_vector\_segments\_max\_seq\_id\_migration\ function allows the system to force the migration of \max\_seq\_id\ data from pickled metadata files to SQLite for unmigrated vector segments, ensuring data consistency during initialization when segments are likely unloaded.
chromadb/ingest/impl · high confidence
ChromaDB 1.5.9 introduces Rust bindings, schema-based collections, and conditional transactions
This release upgrades the Python client to version 1.5.9 and shifts the default backend implementation to Rust bindings (RustBindingsAPI), replacing the previous SegmentAPI for local operations. It introduces a new schema system allowing users to define column types and index configurations (HNSW, SPAN, FTS) at the collection level, alongside support for sparse vectors with optional labels. The update also adds conditional transaction support for optimistic concurrency control, enabling atomic read-modify-write operations with conflict detection. Additionally, it includes a new Search API with Key, Knn, and Rrf operators for hybrid search, rate limiting infrastructure, and stricter error handling for authentication and unique constraints.
chromadb · high confidence
Database schema evolution for multitenancy and configuration storage
The sysdb SQLite schema has been updated to support multitenancy, collection configuration, and boolean metadata. Migration 00004 introduces a tenant/database hierarchy, moving collections from a global scope to be scoped within specific databases under a tenant, and initializes a default tenant and database. Migration 00005 removes the deprecated 'topic' column from both collections and segments. Migration 00006 adds boolean value support to collection and segment metadata tables. Migration 00007 adds a JSON string column to store collection configuration dictionaries. Migration 00008 introduces a maintenance log table to record operations like vacuum. Migration 00009 enforces that segments must always be associated with a collection by making the collection reference non-nullable.
chromadb/migrations/sysdb · high confidence
Introduce configurable coordinator binary with service memberlist and storage flags
The Go coordinator is now a standalone command-line tool that exposes granular configuration flags for its internal services and infrastructure. Users can now explicitly configure the names and pod labels for the query, compaction, garbage collection, log, and function consumer memberlists, allowing for flexible deployment topologies. The binary also supports database connection tuning (including separate read addresses), S3 storage options (such as bucket creation, region, and GCS interop), and heap service settings, providing a unified entry point to manage the coordinator's behavior via CLI arguments.
go/cmd · high confidence
Introduces a unified error handling framework with standardized error codes
The \rust/error\ crate now provides a centralized system for managing errors across the application. It defines 17 standard error codes (aligned with gRPC status codes) plus two custom codes (\VersionMismatch\ and \UnprocessableEntity\), and implements conversions between these codes, HTTP status codes, and gRPC/tonic codes. This ensures consistent error responses for API consumers. The crate also introduces the \ChromaError\ trait and specific wrappers for SQLx, Tonic, and validator errors, allowing internal database and validation failures to be mapped to appropriate, user-facing error codes.
rust/error · high confidence
JS client restructured into a monorepo with separate bundled and peer-dependency packages
The JavaScript client library has been reorganized into a monorepo structure containing three packages: an internal core package (\@internal/chromadb-core\) and two public packages. Users can now choose between the \chromadb\ package, which bundles all embedding library dependencies for a simpler setup, and the \chromadb-client\ package, which uses peer dependencies to keep the dependency tree lean and allow custom embedding library versions. Both packages provide identical functionality and support Node.js and browser environments. The build system now uses pnpm workspaces, and the client code is regenerated from the OpenAPI specification using a new transformation script that handles null types and endpoint response formats.
clients/js · high confidence
New JavaScript client SDK generated from OpenAPI spec
The JavaScript client SDK in \clients/new-js/packages/chromadb/src/api\ has been replaced with a new implementation auto-generated by \@hey-api/openapi-ts\. This introduces a fresh set of TypeScript types and SDK classes (such as \CollectionService\ and \AuthenticationService\) that map directly to the server's OpenAPI specification. For users, this means the client now exposes a comprehensive set of API methods—including collection management, attached functions, and conditional transactions—based on the latest server schema, replacing the previous manual implementation.
clients/new-js/packages/chromadb/src/api · high confidence
New Rust-based CLI client implementation
The CLI now uses a new Rust-based client implementation located in \rust/cli/src/client\. This introduces a \DashboardClient\ that handles authentication flows (including CLI token login and verification) and API interactions for fetching teams and API keys, replacing the previous client logic. The implementation includes utility functions for sending HTTP requests and handling JSON serialization/deserialization.
rust/cli/src/client · high confidence
Refactored block delta storage with optimized ordered writing and precise size tracking
The block delta module has been reorganized to introduce specialized storage backends and improved performance for ordered mutations. A new \OrderedBlockDelta\ type now leverages a \VecBuilderStorage\ for mutations that arrive in sorted order, avoiding the overhead of tree-based structures, while \BTreeBuilderStorage\ remains for unordered data. The system now includes dedicated size trackers (such as \DataRecordSizeTracker\ and \SpannPostingListSizeTracker\) that accurately calculate Arrow buffer sizes—including padding and offsets—during incremental updates, enabling more efficient block splitting and memory management.
rust/blockstore/src/arrow/block/delta · high confidence
SQLite metadata schema migration with performance and type updates
The metadata database schema has been updated through a series of migrations: initial tables for embeddings and metadata are created, a boolean value column is added to metadata, full-text search is reconfigured to use a trigram tokenizer, and performance is improved by adding indices on metadata value columns. Additionally, the sequence ID storage is migrated from a binary blob to a native integer type for more efficient handling.
chromadb/migrations/metadb · high confidence
Support for read replicas and optimized collection queries
The database core layer now supports connecting to a separate read replica via the new \ReadAddress\ configuration field, enabling read operations to be routed to a dedicated instance. Additionally, a feature flag (\EnableOptimizedCollectionQueries\) has been introduced to allow enabling optimized collection queries using Common Table Expressions (CTEs), which is disabled by default.
go/pkg/sysdb/metastore/db/dbcore · high confidence
SysDB schema evolution and function runner support
This update applies a comprehensive series of database migrations to the SysDB metastore, evolving the schema to support a robust function execution system and improved collection management. The schema introduces a new 'functions' entity (replacing the previous 'operators' concept) and 'attached\_functions' to manage function invocations, including columns for tracking failures, readiness, and asynchronous execution modes. Built-in functions such as 'record\_counter', 'statistics', 'http\_generate', 'revision\_history', and 'http\_currents' are seeded into the database. The 'collections' table is expanded with metadata for compaction tracking (size, record counts, timestamps), versioning history, lineage, and schema storage, while 'databases' and 'tenants' gain resource naming and improved pagination indexes. Additionally, legacy tables like 'record\_logs' and 'notifications' are removed, and primary key constraints on 'attached\_functions' are adjusted to support composite keys.
go/pkg/sysdb/metastore/db/migrations · high confidence
Tenant-scoped Foundation management and reasoning trajectory support
The Foundation API now supports creating, listing, and describing Foundations within a specific tenant context using new scoped paths (e.g., \/api/tenants/{tenant}/foundations/{foundation}\), while retaining existing unprefixed paths for default Foundations. It introduces a new reasoning trajectory module that allows saving, appending, and finalizing trajectory records in Chroma, along with example tools to load and verify these trajectories. Additionally, a configuration flag (\foundation.provisioning\_paused\) has been added to temporarily block Foundation creation during catalog migrations.
rust/foundation-api · high confidence
gRPC retry logic and proto conversion utilities added
The proto module now includes a new gRPC client interceptor that automatically retries requests on UNAVAILABLE and UNKNOWN status codes with exponential backoff, improving resilience for distributed communication. Additionally, new conversion utilities were added to handle serialization and deserialization of vectors, metadata, operations, and segments between the internal Python types and the protobuf definitions.
chromadb/proto · high confidence
Test coverage
Add scripts to validate thin client package installation; Added basic integration tests for ChromaClient initialization; Added benchmark for OrderedBlockfileWriter; Added benchmark suite for worker query operators and orchestrators; Added comprehensive API tests for Chroma client functionality; Added comprehensive test suite for the MaxScore sparse index; Added distributed integration tests for conditional transactions, log backpressure, and task APIs; Added full-text search and regex benchmarking infrastructure; Added integration and unit tests for the s3heap scheduling service; Added integration tests for Chroma Cloud embedding functions; Added integration tests for the S3Heap service's HeapTender component; Added k8s integration tests for conditional transactions; Added migration validation test script; Added property-based tests for blockfile writer correctness; Added proptest state-machine tests for the Rust frontend against SQLite; Added regression test cases for garbage collection property tests; Added regression test for HNSW reload failure; Added stress test for creating many collections; Added test coverage for the new JavaScript client; Added test utilities for PostgreSQL container and migration execution; Added test utilities for creating sparse index files; Added tests for SafeUnpickler security and backward compatibility; Added tests for authentication utility logic and RBAC permissions on database deletion; Added tests for client initialization, multitenancy, and concurrency; Added tests for collection configuration and embedding function handling; Added tests for collection operations, delete limits, and fork counts; Added tests for distributed segment directory components; Added tests for the Foundation API agent route; Added tests for worker configuration loading and validation; Expanded property-based test coverage for collection operations and data integrity; Expanded test coverage for embedding functions; Expanded test infrastructure for ChromaDB; Initial test suite for the wal3 replicated log implementation; New test utilities for cross-version persistence, embedding function validation, and result transformation; Refactored SysDB metastore DAO layer with comprehensive test coverage.
Dependencies
Initial commit of Rust workspace and JavaScript client dependency manifests
This change introduces the foundational dependency configuration for the project's Rust backend and JavaScript clients. It adds a \Cargo.lock\ and a comprehensive \Cargo.toml\ workspace definition for the Rust codebase, establishing the initial set of crates and their versions. Simultaneously, it adds \package.json\ and \pnpm-lock.yaml\ files for the JavaScript client ecosystem (including \chromadb\, \chromadb-client\, and \chromadb-core\), defining the Node.js dependencies and peer dependencies for embedding providers.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 39 → 60 (+21.0)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 73 → 82 (+8.2)
- Architecture 77 → 82 (+4.5)
- Maturity 68 → 78 (+10.8)
- Readiness 23 → 63 (+40.2)
- Security 38 → 47 (+9.5)
- Accessibility 82 (new)
Resolved (110)
- Change coupling: collection.ts ↔ utils.ts (clients/new-js/packages/chromadb/src/collection.ts)
- Coverage not measured — test suite did not build
- Critical CVE: [GHSA redacted] (clients/js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (clients/js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (clients/new-js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (clients/js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (clients/new-js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (clients/js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (clients/js/pnpm-lock.yaml)
- Critical CVE: [GHSA redacted] (go/go.mod)
- Critical CVE: [GHSA redacted] (go/go.mod)
- …and 90 more
New (2041)
- AdmissionControlledS3Storage::get_with_e_tag_internal (cognitive 28) (rust/storage/src/admissioncontrolleds3.rs)
- Ambiguous scope of reset. Agent has a reset() method, and AgentBehavior also has a reset() method. It is unclear if calling Agent.reset() internally calls AgentBehavior.reset() or if they manage state independently. This creates a risk of inconsistent state if the user calls one but not the other.
- AppState::search_actions (cognitive 37) (rust/cli/src/tui/collection_browser/app_state.rs)
- AppState::search_actions (cyclomatic 21) (rust/cli/src/tui/collection_browser/app_state.rs)
- ApplyLogsOrchestrator::handle (cognitive 19) (rust/worker/src/execution/orchestration/apply_logs_orchestrator.rs)
- ApplyLogsOrchestrator::initial_tasks (cognitive 18) (rust/worker/src/execution/orchestration/apply_logs_orchestrator.rs)
- ArrowOrderedBlockfileWriter::advance_current_delta_and_get_inner (cognitive 17) (rust/blockstore/src/arrow/ordered_blockfile_writer.rs)
- ArrowUnorderedBlockfileWriter::set (cognitive 20) (rust/blockstore/src/arrow/blockfile.rs)
- AttachedFunctionOrchestrator::handle (cognitive 28) (rust/worker/src/execution/orchestration/attached_function_orchestrator.rs)
- AttachedFunctionOrchestrator::handle (cyclomatic 18) (rust/worker/src/execution/orchestration/attached_function_orchestrator.rs)
- BaseHTTPClient._raise_chroma_error (cognitive 18) (chromadb/api/base_http_client.py)
- Batch.apply (cognitive 22) (chromadb/segment/impl/vector/batch.py)
- BatchManager::take_work (cognitive 16) (rust/wal3/src/interfaces/batch_manager.rs)
- Boundary-crossing change coupling: model_db_convert.go ↔ sysdb.rs (go/pkg/sysdb/coordinator/model_db_convert.go)
- Boundary-crossing change coupling: task_service.go ↔ sysdb.rs (go/pkg/sysdb/grpc/task_service.go)
- Catalog.CreateCollectionAndSegments (cognitive 18) (go/pkg/sysdb/coordinator/table_catalog.go)
- Catalog.DeleteVersionEntriesForCollection (cognitive 26) (go/pkg/sysdb/coordinator/table_catalog.go)
- Catalog.FlushCollectionCompactionForVersionedCollection (cognitive 45) (go/pkg/sysdb/coordinator/table_catalog.go)
- Catalog.FlushCollectionCompactionForVersionedCollection (cyclomatic 23) (go/pkg/sysdb/coordinator/table_catalog.go)
- Catalog.ForkCollection (cognitive 36) (go/pkg/sysdb/coordinator/table_catalog.go)
- …and 2021 more
Changes since last survey
- 59 commits — 38 feature/other, 21 fixes
By area
- rust/worker — 17 commits
- rust/foundation-api — 10 commits
- go/pkg — 7 commits
- rust/wal3 — 4 commits
- rust/agent — 3 commits
- rust/mdac-service — 3 commits
- (root) — 2 commits
- rust/log — 2 commits
- .github/actions — 1 commit
- chromadb/test — 1 commit
- clients/new-js — 1 commit
- docs/mintlify — 1 commit
- docs/scripts — 1 commit
- examples/gemini — 1 commit
- k8s/distributed-chroma — 1 commit
- rust/benchmark — 1 commit
- rust/index — 1 commit
- rust/storage — 1 commit
- rust/tracing — 1 commit
Notable commits
- fix: [BUG] Use UTF-8 for Python reference output (#7504)
- fix: BUG: Drop unpriced Opus from models (#7793)
- fix: BUG: Flush a cached dataset file before renaming it into place (#7717)
- fix: BUG: Don't dead-letter on a missing WorkQueue (#7600)
- fix: BUG: correct GoogleGenerativeAiEmbeddingFunction name in gemini example (#7468)
- fix: BUG: Prefetch records only (#7593)
- fix: BUG: Retry unpublished boundary (#7652)
- fix: BUG: Configurably skip currents on init (#7627)
- fix: BUG: Pin granola collection to 1 dim (#7597)
- fix: BUG: Bound the SPANN reassign recursion depth (#7607)
- fix: BUG: Cap the log client's encode size at the log server's decode limit (#7666)
- fix: BUG: Preserve float metadata precision (#7755)
- fix: BUG: Fix and automate Modal deploy (#7706)
- fix: BUG: remove unused testcontainers dependency (#7467)
- fix: BUG: Field-match racing AttachFunction (#7649)
- fix: BUG: Honor database pagination (#7710)
- fix: BUG: Return segment row read errors (#7785)
- fix: BUG: Bound the backoff test by elapsed time, not a fixed duration (#7718)
- fix: BUG: Fail compactor startup when WorkQueue client init fails (#7619)
- fix: BUG: Pipeline fn consumer dispatches (#7585)
- …and 39 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
chroma-core/chroma was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 27 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit e122af79fd42216d6319ff248914a22142c4ab49 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-7c1cb6328e11.