cocoindex-io/cocoindex
72.1
Strong · 29 September 2026
106.3k
lines of production code
Rust
with Python
2
measurements over time
What this system is
CocoIndex is a data indexing and pipeline framework that provides both Python and Rust SDKs for building incremental, state-managed data workflows. It enables users to ingest data from diverse sources like cloud storage, databases, and message queues, then process it using built-in operations for text chunking, code analysis, and AI-driven embeddings. The system persists results to various targets, including vector databases, graph stores, and traditional databases, while automatically handling change detection and memoization to optimize performance.
How it got here
2025 — Initial SDK and engine scaffolding
23 changes.
This period established the foundational architecture for the CocoIndex project, introducing the v1 API mental model alongside the initial Rust execution engine and Python SDK. It focused on setting up the development environment, dependency management, and core infrastructure such as state management and concurrency utilities. The work also included creating a comprehensive suite of examples and benchmarks to demonstrate capabilities in entity resolution, semantic search, and structured data extraction.
2026 — Rust SDK launch and connector expansion
57 changes.
This period focused on releasing the standalone Rust SDK, mirroring Python capabilities with unified batching, procedural macros, and comprehensive examples. It significantly expanded the connector ecosystem by adding support for vector stores (LanceDB, Qdrant, Valkey), graph databases (Neo4j, FalkorDB, SurrealDB), and cloud storage (S3, Azure, OCI). The work also introduced advanced code-matching infrastructure and robust test coverage for both SDKs.
Features
Add Apache Iggy connector for live data ingestion
Introduces a new connector for Apache Iggy, exposing Iggy streams, topics, and partitions as live data sources within CocoIndex. The implementation includes source and target modules that integrate with the \apache-iggy\ Python SDK, handling partition state tracking, offset management, and asynchronous message consumption to enable real-time data streaming capabilities.
python/cocoindex/connectors/iggy · high confidence
Add Azure Blob Storage source connector
Users can now ingest data from Azure Blob Storage using the new source connector. This change introduces the \AzureBlobFile\, \AzureBlobFilePath\, and \AzureBlobWalker\ classes, along with utility functions like \list\_blobs\ and \read\, enabling the platform to discover and read files directly from Azure Blob containers.
_python/cocoindex/connectors/azure\blob · high confidence
Add FalkorDB graph database connector
Users can now write data to and manage indexes in FalkorDB, a graph database, using the new \falkordb\ connector. This addition introduces support for upserting and deleting graph nodes and relationships, as well as creating and dropping standard and vector-based indexes, enabling integration with FalkorDB's Cypher query engine.
python/cocoindex/connectors/falkordb · high confidence
Add Kafka source and target connectors
Users can now ingest data from and write data to Apache Kafka topics using the new Kafka connectors. The source connector exposes a Kafka topic as a LiveStream of raw messages and as a LiveMapFeed of keyed change events, supporting offset tracking and commit management. The target connector enables writing data to Kafka topics. Both connectors require the confluent\_kafka library.
python/cocoindex/connectors/kafka · high confidence
Add LLM-based entity pair resolver
A new built-in LLM-based pair resolver is available for entity resolution, allowing users to leverage large language models to determine if two entity names refer to the same thing. This component uses the \instructor\ and \litellm\ libraries to generate structured JSON responses, supporting configuration of the model, entity type hints, and extra guidance. It includes a validation-and-retry loop to handle invalid model outputs and memoizes results to persist decisions across runs.
_python/cocoindex/ops/entity\resolution · high confidence
Add LanceDB connector module
Introduces a new LanceDB connector package within the Python CocoIndex library. The module exposes the connector's functionality by re-exporting symbols from the internal \\_target\ implementation, making the LanceDB integration available for use.
_python/cocoindex/connectors/doris, python/cocoindex/connectors/google\drive, python/cocoindex/connectors/lancedb · high confidence
Add PDF-to-Markdown example using docling
The \examples/pdf\_to\_markdown\ directory now contains a complete example that walks a directory of PDFs, converts each one to Markdown using the docling library, and writes the results to an output folder. The example uses CocoIndex's \localfs\ connector and \mount\_each\ to process files incrementally, ensuring that only changed PDFs are reprocessed. It includes a \main.py\ script, a \README.md\ with usage instructions, and configuration files (\.env\, \.gitignore\).
_examples/pdf\_to\markdown · high confidence
Add Qdrant connector
A new Qdrant connector has been added to the Python API, exposing the target implementation via the \cocoindex.connectors.qdrant\ module to allow integration with Qdrant vector databases.
python/cocoindex/connectors/qdrant · high confidence
Add SQLite target connector module
A new SQLite target connector has been added to the Python library, exposing the underlying implementation via the \cocoindex.connectors.sqlite\ package to allow users to write data to SQLite databases.
python/cocoindex/connectors/sqlite · high confidence
Add SurrealDB target connector
A new target connector for SurrealDB has been added to the Python library. This change introduces the \python/cocoindex/connectors/surrealdb\ module, which exposes the underlying target implementation via \\_\init\\_.py\, enabling users to configure and write data to SurrealDB databases.
(repo-wide) · high confidence
Add Valkey vector store target connector
A new connector for the Valkey vector store has been added, exposing the target implementation via the \python/cocoindex/connectors/valkey\ module. This allows users to configure and use Valkey as a target for vector data storage within the CocoIndex framework.
python/cocoindex/connectors/valkey · high confidence
Add v1 examples and benchmarks for entity resolution, file summarization, state store, and cloud storage embeddings
This change introduces a suite of v1-compatible demonstration pipelines and performance benchmarks. It adds benchmark scripts for entity resolution (using synthetic and LLM-based resolvers), file summarization (analyzing codebases and documentation), and state-store latency (measuring cold/warm/drop times for LMDB). It also provides new v1 examples for text embedding from Amazon S3 and Azure Blob Storage, audio-to-text transcription, and code embedding into both PostgreSQL (pgvector) and LanceDB, showcasing the updated connector and target APIs.
cocoindex · high confidence
Added stress test harness for the code matcher
A new stress-test example (\stress.rs\) has been added to the \code\_match\ crate. This tool exercises the matcher and prefilter logic against a real codebase by generating various pattern types (literal, metavar-injected, containment-wrapped, and multi-layer nested) and verifying invariants such as prefilter soundness and bounded match times, helping users detect performance regressions or panics in complex matching scenarios.
_rust/code\match/examples · high confidence
Anonymous usage tracking via Scarf gateway
The core library now collects anonymous usage telemetry by sending events to the Scarf gateway. This feature is active only in release builds and can be disabled by setting the \COCOINDEX\_DISABLE\_USAGE\_TRACKING\ environment variable (unless set to empty or "0"). The tracking is non-blocking and fire-and-forget, ensuring it never interferes with user operations. Additionally, an optional \application\ label can be included in telemetry events by setting the \COCOINDEX\_APPLICATION\_FOR\_TRACKING\ environment variable.
rust/core/src/telemetry · high confidence
Automated documentation builds and Python type-checking via Claude Code hooks
The project now automatically validates documentation and Python code changes during development. When files in the docs/ directory are edited, a hook triggers a documentation build (via yarn) to ensure the docs remain valid. Similarly, when Python (.py or .pyi) files are modified, a hook runs mypy type-checking to catch type errors early. These checks are configured in .claude/settings.json and executed through shell scripts in .claude/hooks/, with local settings and hook artifacts ignored by git.
.claude · high confidence
Configurable code element extraction with per-language hooks
The code\_ast module now supports configurable extraction of code elements (declarations, references, types, namespaces) via a new \LanguageHooks\ trait and \ExtractorConfig\ system. Users can customize how specific AST node types are interpreted per language (C, C++, C\#, Go, Java, JavaScript, Kotlin, Python, Rust, Swift, TypeScript, TSX) through configurable node-kind mappings and exclusion patterns. The system includes language-specific hooks for handling namespace separators, declaration name extraction, body detection, and path resolution, allowing accurate code element identification across diverse programming languages.
_rust/code\ast · high confidence
Expanded language support with precise tokenization for code matching
The code matching engine now supports 29 programming and markup languages, including Bash, C, C++, C\#, CSS, Go, HTML, Java, JavaScript, JSON, Julia, Kotlin, PHP, Python, R, Ruby, Rust, Scala, Solidity, SQL, Swift, and configuration formats like CMake, HCL, TOML, and YAML. Each language is configured with a specific tokenizer profile that accurately handles its unique syntax—such as context-sensitive quote handling in HTML, raw string delimiters in Rust and C++, doubled-quote escaping in SQL and Pascal, and bracket arguments in CMake. This ensures that search patterns match code structures correctly across these diverse languages.
_rust/code\match/src/lang · high confidence
Initial PostgreSQL connector module
A new PostgreSQL connector module has been added to the Python package, exposing both source and target capabilities through the public API. This allows users to read from and write to PostgreSQL databases using the standard import path.
python/cocoindex/connectors/postgres · high confidence
Initial integration of zvec target connector
A new zvec connector has been added to the product, exposing the target implementation via the \python/cocoindex/connectors/zvec\ module. This allows users to utilize the zvec target for data operations.
python/cocoindex/connectors/zvec · high confidence
Initial release of the Cocoindex Python SDK
This change introduces the initial version of the Cocoindex Python SDK, providing the core framework for building and running indexing pipelines. The package includes a version check mechanism to ensure consistency between the Python package and the underlying engine, utilities for serializing objects for the engine, and a robust loader for user applications that supports both file paths and installed module names. It also exposes internal APIs for inspection and defines the public interface for the library.
python/cocoindex · high confidence
Initial release of the Rust core library
This change introduces the \rust/core\ crate, establishing the foundational module structure for the engine. It exposes public modules for the engine, inspection, state management, state storage, and telemetry, while providing a prelude that aggregates common utilities, standard collections, async primitives, and tracing instrumentation for downstream consumers.
rust/core/src · high confidence
Initial release of the Rust execution engine core
This change introduces the foundational Rust implementation of the execution engine, establishing the core architecture for app lifecycle, component processing, and state management. It adds the \App\ struct for managing application instances and updates, the \Component\ system for handling processing units with memoization and activity tracking, and the \Environment\ abstraction for storage and configuration. The engine now supports cooperative deadlines, ID sequencing with exponential batching, and function-level memoization with context dependency tracking. Additionally, it includes the \DeadlineContext\ for timeout management and the \IdSequencerManager\ for efficient, batched ID generation, laying the groundwork for concurrent component execution and state reconciliation.
rust/core/src/engine · high confidence
Initial repository scaffolding and development environment setup
This change establishes the foundational structure for the CocoIndex project, introducing the v1 API mental model and development workflow. It adds configuration files for the \uv\ package manager (\.python-version\, \uv.lock\, \ruff.toml\) and a comprehensive pre-commit hook suite (\.pre-commit-config.yaml\) that enforces formatting, linting, type checking, and security scanning via \gitleaks\. It also provides essential documentation for contributors and AI agents (\AGENTS.md\, \CLAUDE.md\, \context7.json\) and sets up the root \README.md\ with installation instructions and examples.
(repo-wide) · high confidence
Introduce \#\[function\] proc macro with memoization and logic tracking
The Rust SDK now provides the \\#\[function\]\ procedural macro for defining SDK functions, enabling features such as result memoization, explicit versioning for logic-change detection, and configurable logic tracking (full, self-only, or none) to control how transitive dependencies are invalidated. This macro parses function signatures to identify context parameters and argument types, generating the necessary registration and hashing logic to align the Rust SDK's function definition model with the Python SDK.
_rust/sdk/cocoindex\macros · high confidence
Introduce Amazon S3 source connector with read-only API
Users can now read data from Amazon S3 buckets (and S3-compatible services like MinIO) using a new async-only API. This change adds the \S3File\, \S3FilePath\, and \S3Walker\ classes, along with helper functions like \get\_object\ (which accepts both S3 URIs and bucket/key pairs) and \list\_objects\. The implementation ensures correct handling of content fingerprints by encoding S3 ETags to bytes, preventing memo-state round-trip failures during re-runs.
_python/cocoindex/connectors/amazon\s3 · high confidence
Introduce Rust-based Python SDK with core engine bindings
The Python SDK is now implemented in Rust (PyO3), exposing the core engine to Python via new modules. This includes \App\ and \Environment\ management, component processing with memoization and state handling (\use\_state\), and a new batching infrastructure (\BatchQueue\, \Batcher\) for efficient function execution. Code analysis is powered by a shared \CodeSource\ and compiled \CodePattern\ using tree-sitter. Additional capabilities include GPU pool management, deadline/timeout control, and a comprehensive inspection API (\iter\_stable\_paths\, \list\_app\_names\) for debugging and state visibility.
rust/py · high confidence
Introduce internal runtime initialization and batch-splitting retry logic
The Python SDK now initializes its Rust core runtime via a new internal module, registering serialization and stable-path fingerprinting hooks. Additionally, a new batch-splitting mechanism allows batched functions to automatically retry failed batches by halving their size, isolating failures to individual items rather than failing the entire batch.
_python/cocoindex/\internal · high confidence
Introduce local filesystem connector with live watching support
Adds a new local filesystem connector that allows reading files from the local disk, featuring both standard iteration and a live mode that watches for file changes. The implementation includes a \FilePath\ class for stable path resolution using context keys, async file reading to avoid blocking the event loop, and a directory walker that supports recursive traversal and path matching. When live mode is enabled, it provides a \LiveMapView\ for real-time updates on file additions, modifications, and deletions.
python/cocoindex/connectors/localfs · high confidence
Introduce stable path and target state schema for component state management
The state module now defines the core data structures for managing component and target state persistence. This includes the \StableKey\ type, which supports various data types (including Symbols, Uuids, and Bytes) with specific serialization logic for MessagePack, and the \StablePath\ hierarchy used to organize component state. Additionally, the \TargetStatePath\ and \TargetStateProviderGeneration\ structures are introduced to track provider dependencies and handle schema versioning for memoization invalidation. The database schema (\db\_schema.rs\) establishes the key layout for storing user state (distinguishing between regular and live states), memoization info, and child component existence, ensuring proper isolation and prefix scanning capabilities.
rust/core/src/state · high confidence
Introduce structural code-search crate with pattern matching and prefiltering
The new \cocoindex\_code\match\ crate enables matching by-example structural patterns against tree-sitter ASTs across 14 languages. It introduces a pattern lexer that supports metavariables (named, anonymous, and regex-constrained), quantifiers (\\\, \+\, \?\), containment (\\\{{ ... \\}}\), and whole-node boundaries (\\\{ ... \\}\). The matcher uses a memoized dynamic programming algorithm to align flat pattern skeletons against the AST leaf frontier. To optimize performance, a prefilter extracts required literal content (identifiers, string runs, regex literals) into a CNF clause structure using Aho-Corasick, allowing callers to reject non-matching sources before the expensive tree-sitter parse.
_rust/code\match/src · high confidence
New BAML example for extracting patient intake data from PDFs
An example demonstrating how to use BAML to extract structured patient information from PDF intake forms has been added. This includes a schema defining patient details (such as contact info, insurance, and medical history) and a generator configuration that outputs Python Pydantic models, configured to use the Gemini model via the Google AI provider.
_examples/patient\_intake\_extraction\_baml/baml\src · high confidence
New DSPy-based patient intake extraction example
Added a new example in \examples/patient\_intake\_extraction\_dspy\ that demonstrates extracting structured patient records from visual intake forms. The pipeline renders PDF pages to images and uses a DSPy \ChainOfThought\ module with a Gemini vision model to populate a typed Pydantic \Patient\ model, which is then validated and saved as JSON. The example includes the necessary configuration files (\.env\, \.env.example\), a \.gitignore\ for output directories, and sample data to run the async Python pipeline via \cocoindex update\.
_examples/patient\_intake\_extraction\dspy · high confidence
New HackerNews trending topics example
Added a new example demonstrating how to scrape recent HackerNews threads and comments via the Algolia API, use an LLM to extract and normalize topics, and store the results in Postgres for trending analysis. The example includes configuration files for Postgres and Gemini API keys, Pydantic models for topic extraction, and a Python script to run the pipeline and query results.
_examples/hn\_trending\topics · high confidence
New Neo4j connector for graph data integration
Added a new Neo4j connector that enables writing graph data to Neo4j databases. The implementation includes a Cypher generation module that constructs queries for node and relationship upserts, deletions, and index management, specifically targeting Neo4j 5.18+ for vector index support. This allows users to persist and manage graph structures using standard Cypher operations.
python/cocoindex/connectors/neo4j · high confidence
New OCI Object Storage connector with live event watching
Added a new connector for Oracle Cloud Infrastructure (OCI) Object Storage, exposing classes like OCIFile, OCIFilePath, and OCIWalker to list, read, and watch objects. The implementation supports live mode by accepting a LiveStream, which triggers an initial full scan and then watches for changes via OCI Streaming events, allowing users to react to bucket updates in real-time.
_python/cocoindex/connectors/oci\_object\storage · high confidence
New Postgres source example with semantic indexing and composite keys
Added a new example in \examples/postgres\_source\ that demonstrates how to use an existing Postgres table as a source for semantic search. The example reads product rows, derives fields, embeds them using a local sentence-transformer model, and writes the enriched rows and vectors back to Postgres using pgvector. It supports composite primary keys and incremental updates, allowing users to quickly set up a semantic index over structured data with plain async Python.
_examples/postgres\source · high confidence
New Rust SDK examples added
Added a comprehensive suite of Rust examples mirroring the existing Python library, covering data ingestion, transformation, and storage. New examples include embeddings for Amazon S3, Google Drive, OCI Object Storage, and local files; code and image search using pgvector, LanceDB, Qdrant, and ColPali; graph database integrations with SurrealDB, FalkorDB, and Neo4j; message queue interactions with Kafka and Apache Iggy; and utility pipelines for audio transcription, PDF processing, and multi-codebase summarization. Each example includes detailed documentation, configuration instructions, and sample data to demonstrate incremental processing, memoization, and target reconciliation.
examples/rust · high confidence
New Rust SDK examples for CSV-to-message, file transformation, and code summarization
Added Rust implementations of several CocoIndex examples: \csv\_to\_kafka\ and \csv\_to\_iggy\ (publishing CSV rows as messages to Kafka or Apache Iggy with incremental, memoized processing), \files\_transform\ (converting Markdown to HTML via a declarative directory target), \pdf\_to\_markdown\ (extracting text from PDFs and writing Markdown outputs), and \multi\_codebase\_summarization\ (scanning Python projects to extract public classes/functions, generate Mermaid call graphs, and produce project summaries using an LLM). These examples demonstrate the Rust SDK's mount macros, component-memo fast-path, and declarative targets for incremental data pipelines.
(repo-wide) · high confidence
New Rust SDK examples for data ingestion and embedding pipelines
Added Rust ports of Python examples demonstrating how to use the SDK to build data pipelines. These include Amazon S3, Google Drive, OCI Object Storage, and local file embedding workflows using Postgres/pgvector, a code embedding example with tree-sitter chunking, an audio-to-text transcription pipeline using OpenAI Whisper, a HackerNews trending topics scraper with LLM topic extraction, a paper metadata extractor, a PDF embedding example, a SQLite file summarization example, and a Postgres source ingestion example.
(repo-wide) · high confidence
New Rust SDK operations for text, code, and embeddings
The Rust SDK now exposes a first-class \ops\ module mirroring the Python SDK, providing local and remote capabilities for data processing. Users can perform text splitting and language detection via \text.rs\, conduct structural code matching over reusable ASTs via \code.rs\, and generate embeddings locally using \fastembed\ for both text (\sentence\_transformers.rs\) and images (\image.rs\). Additionally, \api.rs\ enables remote embeddings and audio transcription by connecting to any OpenAI-compatible HTTP endpoint. These operations are feature-gated to keep dependencies lightweight.
rust/sdk/cocoindex/src/ops · high confidence
New Rust SDK with unified batching, S3 source, and target connectors
The Rust SDK is now available as a standalone library, providing a unified \Batched\ API for single-value calls with automatic coalescing and memoization, alongside an \Environment\ builder for managing LMDB state and shared resources. New capabilities include an Amazon S3 read-only source connector for listing and reading objects, and target connectors for Apache Doris (table upserts via Stream Load), FalkorDB (graph operations), and Cypher-based graph databases (Neo4j/FalkorDB) with vector and node index management.
rust/sdk/cocoindex/src · high confidence
New Rust benchmark for file summarization workloads
Added a new Rust-based benchmark tool located at benchmarks/file\_summarization/rust that evaluates file summarization performance across different scenarios (codebase vs. docs) and workload profiles (I/O, CPU, mixed). The tool allows users to configure analysis parameters such as shingle span and analysis rounds, and it generates detailed metrics including file counts, section signatures, and token statistics to help assess system performance under varying load conditions.
_benchmarks/file\summarization/rust · high confidence
New Rust examples for Kafka consumption and graph-based meeting note processing
Added three new Rust example applications: a Kafka consumer (\kafka\_consume\) that demonstrates reading from a Kafka topic in both snapshot and live-tailing modes using \topic\_as\_map\; and two graph-database examples (\meeting\_notes\_graph\_falkordb\ and \meeting\_notes\_graph\_neo4j\) that parse markdown meeting notes, extract structured data (meetings, people, tasks), and persist them into FalkorDB and Neo4j graph databases respectively.
_examples/rust/kafka\_consume, examples/rust/meeting\_notes\_graph\_falkordb, examples/rust/meeting\_notes\_graph\neo4j · high confidence
New Rust examples for vector search with LanceDB, Qdrant, and Turbopuffer
Added Rust implementations of the Python vector-search examples, demonstrating how to index and query code, text, and images using the native \cocoindex::lancedb\, \cocoindex::qdrant\, and \cocoindex::turbopuffer\ connectors. The new examples include \code\_embedding\_lancedb\ (code files with tree-sitter chunking), \image\_search\ (CLIP-based image search), \image\_search\_colpali\ (ColPali multi-vector image search), \text\_embedding\_lancedb\, \text\_embedding\_qdrant\, and \text\_embedding\_turbopuffer\ (markdown chunking and embedding). Each example supports incremental indexing and vector search queries, providing reference architectures for integrating these vector databases with the Rust SDK.
(repo-wide) · high confidence
New SEC EDGAR analytics example with hybrid search
Added a new example in \examples/sec\_edgar\_analytics\ that indexes multi-format SEC filings (10-K text and XBRL JSON) into Apache Doris. The pipeline scrubs PII, chunks the data, and creates a single table with both vector (ANN) and full-text (inverted) indexes to enable hybrid search via Reciprocal Rank Fusion (RRF). Includes a Docker Compose setup for Doris and PostgreSQL, a script to generate synthetic sample data, and a search script demonstrating the combined semantic and keyword retrieval.
_examples/sec\_edgar\analytics · high confidence
New agent skills for documentation diagrams, target connectors, and dependency upgrades
Added new agent-skill definitions in dev/agent-skills to guide coding agents on specific tasks: the cocoindex-diagrams skill provides a complete workflow for creating and reviewing inline SVG diagrams in the docs site, including a starter Astro template, layout patterns, and a preview script; the target-connector skill documents the implementation of TargetHandler and TargetActionSink for declarative target state synchronization, including attachment providers and input safety guidelines; and the upgrade-examples skill provides instructions for updating cocoindex package versions across example projects. Additionally, portable build and typecheck scripts were added to dev/agent-checks to support local automation.
dev/agent-skills · high confidence
New audio-to-text transcription example with Postgres and LiteLLM
Added a new example in the \examples/audio\_to\_text\ directory that demonstrates how to transcribe a directory of audio files into a Postgres table using LiteLLM. The example includes a Python script (\main.py\) that walks a local directory, transcribes audio files (e.g., .mp3, .wav) via a LiteLLM speech-to-text model (defaulting to \whisper-1\), and stores the results in a managed Postgres table keyed by filename. It also provides configuration files (\.env\, \.env.example\) and documentation (\README.md\) to guide users through setup, including starting a local Postgres instance and configuring API keys.
_examples/audio\_to\text · high confidence
New benchmarks for entity resolution parallelism, Rust vs Python performance, and state-store latency
Three new benchmarks have been added to the benchmarks directory. The entity resolution benchmark (benchmarks/entity\_resolution) measures the performance of the resolve\_entities operation, specifically validating the parallelization of entity resolution via candidate-graph components against sequential execution using synthetic and OpenAI-backed profiles. The file summarization benchmark (benchmarks/file\_summarization) provides a side-by-side comparison of the Rust and Python SDKs, demonstrating significant speedups for the Rust SDK in cold and warm runs while verifying the incremental update contract. The state-store latency benchmark (benchmarks/state\_store) isolates and measures the overhead of the state store's bookkeeping, memo writes, and drop cascades as the number of components and target states increases.
benchmarks · high confidence
New connectorkits module with async adapters, single-subscriber guards, and fingerprinting utilities
The new \python/cocoindex/connectorkits\ package introduces shared infrastructure for connector implementations. It adds \SingleWatcherGuard\ to enforce a single-active-subscriber contract on live feeds, preventing race conditions when multiple concurrent \watch()\ calls are made. The \async\_adapters\ module provides \sync\_to\_async\_iter\ and \async\_to\_sync\_iter\ utilities that bridge synchronous and asynchronous iterators using background threads and queues, ensuring the event loop is not blocked. Additionally, the \fingerprint\ module exposes deterministic hashing functions (\fingerprint\_bytes\, \fingerprint\_str\, \fingerprint\_object\) for change detection, and \target.py\ defines the \ManagedBy\ StrEnum to specify resource lifecycle management.
python/cocoindex/connectorkits · high confidence
New conversation-to-knowledge example pipeline
Added a new example application in \examples/conversation\_to\_knowledge/conv\_knowledge\ that demonstrates a CocoIndex pipeline for converting YouTube podcast sessions into a structured knowledge graph stored in SurrealDB. The pipeline fetches audio and transcripts (using yt-dlp and AssemblyAI with speaker diarization), uses LLMs to extract metadata and identify speakers, and then extracts thematic statements and entities. It includes an LLM-based entity resolution step to deduplicate persons, technologies, and organizations, finally storing the canonical entities and their relationships (session-to-statement, person-to-session, person-to-statement) in the database.
_examples/conversation\_to\_knowledge/conv\_knowledge, examples/rust/conversation\_to\knowledge · high confidence
New csv\_to\_kafka example for streaming CSV changes to Kafka
Added a new example in examples/csv\_to\_kafka that demonstrates watching a local directory of CSV files and publishing row-level changes to a Kafka topic. The example includes main.py, which uses CocoIndex to detect changes in CSV rows and declare them as target state for a Kafka topic, .env.example for configuration, and sample data files. It supports live mode for continuous monitoring and handles SASL authentication for managed Kafka services.
_examples/csv\_to\kafka · high confidence
New development tooling and environment configuration
The dev directory now includes a suite of scripts and configuration files to streamline local development. A new Python script (\generate\_cli\_docs.py\) automatically generates CLI documentation from Click commands, integrated with the \prek\ hook system. Environment setup is supported by Docker Compose configurations for Neo4j and PostgreSQL. Rust development is improved with scripts to run tests via \uv\, check external Rust workspaces for API breakage, and update their lockfiles. Additionally, a new script detects free-threaded Python (3.13+) to adjust \maturin\ build flags, and a helper script ensures the Rust SDK remains PyO3-free.
dev · high confidence
New files\_transform example for Markdown-to-HTML pipelines
Added a new example in \examples/files\_transform\ that demonstrates a source-to-target pipeline: it watches a directory of Markdown files and renders them to HTML using \markdown-it-py\. The example includes the implementation (\main.py\), sample data, and configuration files (\.env\, \.gitignore\), illustrating how to use \localfs.walk\_dir\ with live watching and \coco.mount\_each\ for concurrent processing.
_examples/files\transform · high confidence
New image search examples using CLIP and ColPali
Added two new examples in the \examples/image\_search\ and \examples/image\_search\_colpali\ directories that demonstrate searching local image folders by meaning using CocoIndex and Qdrant. The \image\_search\ example uses the CLIP model to embed images and text into a shared vector space, while the \image\_search\_colpali\ example uses ColPali for multi-vector embeddings with Qdrant's MaxSim for finer-grained retrieval. Both examples include a FastAPI backend (\api.py\) that runs the index in live mode and a React frontend (\frontend/\) for querying, along with configuration files (\.env\, \.env.example\) and documentation (\README.md\).
_examples/image\search · high confidence
New internal database inspection module for stable paths and target states
Added a new \db\_inspect\ module in \rust/core/src/inspect\ that provides programmatic access to internal database state. This includes functions to stream stable paths with metadata (\iter\_stable\_paths\), list application names (\list\_app\_names\), and resolve target state paths by decoding fingerprinted segments back to original stable keys using an inverted owner index and persisted segment names. The module exposes data structures like \StablePathInfo\, \TargetStateVersion\, and \TargetStateInfoItemSummary\ to support debugging and inspection of the state store.
rust/core/src/inspect · high confidence
New multi-codebase summarization example
Added a new example pipeline that automatically generates self-updating Markdown wikis for Python projects. The pipeline scans subdirectories, uses an LLM to extract public classes, functions, and CocoIndex call graphs, and aggregates them into per-project documentation with Mermaid diagrams. It leverages incremental processing to re-analyze only changed files and concurrent execution to process multiple projects in parallel.
_examples/multi\_codebase\summarization · high confidence
New ops package with LiteLLM, SentenceTransformer, and text processing utilities
The new \cocoindex.ops\ package introduces dedicated modules for common data processing tasks. \liteLLM\ provides \LiteLLMEmbedder\ for vector embeddings and \LiteLLMTranscriber\ for speech-to-text, featuring robust retry logic for transient network errors and terminal handling for timeouts. \sentence\_transformers\ offers \SentenceTransformerEmbedder\ with batched, GPU-accelerated embedding support and automatic retry on out-of-memory errors by splitting batches. \text\ provides \SeparatorSplitter\ and \RecursiveSplitter\ for text chunking, with syntax-aware splitting capabilities via tree-sitter integration for code and custom language configurations.
python/cocoindex/ops · high confidence
New patient intake extraction example using BAML and Gemini
Added a new example demonstrating how to extract structured patient data from PDF intake forms using BAML for type-safe schema validation and Google's Gemini vision model. The example includes a Python pipeline (main.py) that reads PDFs, runs an async extraction function, and writes validated JSON records to disk, along with configuration files (.env, .env.example) and documentation (README.md) to guide users through setup and execution.
_examples/patient\_intake\_extraction\baml · high confidence
New regex-based and syntax-aware text splitting utilities
The \ops\_text\ crate now exposes a new \split\ module providing two distinct text chunking strategies. First, \SeparatorSplitter\ allows splitting text by configurable regex patterns, with options to keep separators on the left or right, trim whitespace, and include empty chunks. Second, \RecursiveChunker\ offers syntax-aware chunking using tree-sitter for code files and a fallback to multi-level regex separators for general text, supporting custom language definitions and configurable chunk sizes and overlaps. Both splitters return \Chunk\ objects containing byte ranges and precise output positions (line/column) for the original text.
_rust/ops\text/src/split · high confidence
New resource modules for embedders, file handling, ID generation, and schema definitions
This change introduces several new modules in the \python/cocoindex/resources\ package that define core protocols and utilities for the SDK. The \embedder\ module adds an \Embedder\ protocol for single-text async embedding, supporting batching patterns. The \file\ module provides a \FileLike\ base class with lazy metadata fetching, content caching, and memoization support via content fingerprints. The \id\ module introduces stable ID and UUID generation utilities, including \IdGenerator\ and \UuidGenerator\ classes that produce distinct IDs on each call even for identical dependencies, alongside idempotent \generate\_id\ and \generate\_uuid\ functions. The \schema\ module defines \VectorSchema\ and \MultiVectorSchema\ structures along with provider protocols to convey vector column metadata. Finally, the \view\ module adds \SourceView\ and \ViewSegment\ dataclasses for representing synthetic text with source-grounded segments, intended for structural chunking and code-match rendering.
python/cocoindex/resources · high confidence
New shared Rust utilities library for batching, concurrency, and error handling
The \rust/utils\ crate has been introduced as a central library providing core infrastructure components used across the platform. It includes a \Batcher\ and \BatchQueue\ for grouping and processing inputs in FIFO order, a \ConcurrencyController\ to limit in-flight rows and bytes, and a \GPUPool\ for fractional GPU capacity management. The library also introduces a structured \Error\ system with support for host-language error tunneling, cancellation detection, and error replication for batched operations, alongside utilities for identifier validation, JSON/MessagePack deserialization with path-aware errors, and HTTP request retry logic with slow-request warnings.
rust/utils · high confidence
New text-embedding example targeting Qdrant
Adds a new example demonstrating semantic search over Markdown files using CocoIndex and Qdrant. The example walks a directory of Markdown files, chunks the text, generates embeddings locally, and upserts the vectors into a managed Qdrant collection. It includes a CLI for indexing (\cocoindex update main\) and querying (\python main.py "query"\), along with environment configuration files for Qdrant connection and PyTorch MPS fallback.
_examples/text\_embedding\_qdrant, examples/text\_embedding\turbopuffer · high confidence
Support for negation patterns in file exclusion
The \ops\_text\ crate now includes a \PatternMatcher\ that supports gitignore-style negation in \excluded\_patterns\. Users can prefix exclusion patterns with \!\ to un-exclude specific paths that would otherwise be filtered out. The implementation handles brace expansion in negation patterns to correctly determine directory inclusion during traversal, ensuring that parent directories of negated files are not prematurely pruned.
_rust/ops\text/src · high confidence
Architecture
Extracted py\_utils into a dedicated Rust package
The \py\_utils\ module has been refactored into a standalone Rust package (\rust/py\_utils\) to better organize Python interoperability utilities. This new package provides structured error handling that preserves full Python tracebacks when exceptions cross the Rust-Python boundary, safe bridging of Python async futures to Rust futures via \call\_soon\_threadsafe\ to avoid event-loop thread violations, and generic serialization helpers for converting between Python and Rust types using \pythonize\/\depythonize\.
_rust/py\utils · high confidence
Behavioural changes
Portable Agent Skills discovery
The .agents directory now includes symbolic links for four agent skills (cocoindex, cocoindex-diagrams, target-connector, and upgrade-examples) that point to their actual implementations in the skills/ and dev/agent-skills/ directories. This change enables the agent system to discover and utilize these skills through a standardized, portable path structure.
.agents · high confidence
Refactored state store with per-app handles and standalone reads
The state store module has been restructured to isolate LMDB-specific logic and introduce an \AppStore\ handle for per-application access. This change replaces the previous API that required callers to manage transactions by providing standalone read methods (e.g., \read\_txn\) that open their own snapshot transactions internally, simplifying usage for memo lookups and inspection. The new architecture also includes a single-writer batcher for write transactions to improve concurrency and supports automatic LMDB map resizing when the database fills up.
_rust/core/src/state\store · high confidence
Test coverage
Added SDK microbenchmarks and state-store read benchmarks; Added comprehensive CLI test suite and fixtures; Added core test suite for logic change detection and memoization; Added integration tests for target-owner cleanup and telemetry shutdown safety; Added integration tests for the structural code-matching engine and prefilter; Added test infrastructure for Docker availability checks and testcontainers stability; Added tests for async iterator cleanup behavior; Added tests for file path matching, ID generation, and LiveMap resources; Added tests for ops module code matching, LiteLLM embedder/transcriber, and entity resolution; Added unit tests for internal memoization and context key APIs; Added unit tests for source and target connectors; Expanded Rust SDK test coverage for sources, targets, and core behaviors; New test utilities for shared test infrastructure.
Dependencies
Initial dependency manifests and lockfiles for Rust workspace, benchmarks, docs, and examples
This change introduces the initial dependency configuration for the project, adding the root \Cargo.toml\ and \Cargo.lock\ for the Rust workspace (using Rust edition 2024 and specifying dependencies like \axum\ 0.8, \pyo3\ 0.29, and \tokio\ 1.48), along with \pyproject.toml\ files for all examples and benchmarks pinning \cocoindex\>=1.0.7\, and \package.json\/\package-lock.json\ files for the Astro-based documentation site and example frontends.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 73 → 72 (-0.9)
- Rubric changed (rubric-2026.09.8 → rubric-2026.09.17) — scores are not directly comparable.
Lenses
- Code Health 86 → 85 (-0.6)
- Architecture 100 → 96 (-4.1)
- Maturity 81 → 81 (+0.1)
- Readiness 68 → 66 (-1.7)
- Security 69 → 76 (+7.5)
- Accessibility 71 (new)
- Performance 100 (new)
Resolved (30)
- Documentation: no installation or build instructions (README.md)
- Documentation: no installation or build instructions (examples/meeting_notes_graph_surrealdb/README.md)
- Documentation: no project overview (README.md)
- Documentation: no usage examples (examples/meeting_notes_graph_surrealdb/README.md)
- Duplicated block (11 lines × 2) (rust/sdk/cocoindex/src/target_state.rs)
- Duplicated block (12 lines × 2) (rust/sdk/cocoindex/src/doris.rs)
- Duplicated block (16 lines × 2) (rust/sdk/cocoindex/src/target_state.rs)
- Duplicated block (6 lines × 2) (rust/sdk/cocoindex/src/target_state.rs)
- Duplicated block (6 lines × 3) (rust/sdk/cocoindex/src/lancedb.rs)
- Duplicated block (7 lines × 2) (rust/sdk/cocoindex/src/target_state.rs)
- Duplicated block (8 lines × 2) (rust/core/src/engine/target_state.rs)
- Further sole-owners (lower concentration)
- Hotspot: docs/src/pages/llms.txt.ts (docs/src/pages/llms.txt.ts)
- Hotspot: rust/code_match/src/lang/mod.rs (rust/code_match/src/lang/mod.rs)
- Hotspot: rust/code_match/src/lang/swift.rs (rust/code_match/src/lang/swift.rs)
- Hotspot: rust/code_match/src/lexer.rs (rust/code_match/src/lexer.rs)
- Hotspot: rust/code_match/src/matcher.rs (rust/code_match/src/matcher.rs)
- Hotspot: rust/core/src/state_store/submit_session.rs (rust/core/src/state_store/submit_session.rs)
- Hotspot: rust/ops_text/src/split/recursive.rs (rust/ops_text/src/split/recursive.rs)
- Hotspot: rust/sdk/cocoindex/src/sqlite.rs (rust/sdk/cocoindex/src/sqlite.rs)
- …and 10 more
New (131)
- Duplicated block (10 lines × 2) (examples/bigquery_target/main.py)
- Duplicated block (10 lines × 2) (python/cocoindex/_internal/function.py)
- Duplicated block (10 lines × 2) (python/cocoindex/connectors/falkordb/_target.py)
- Duplicated block (10 lines × 2) (python/cocoindex/connectors/falkordb/_target.py)
- Duplicated block (10 lines × 2) (python/cocoindex/connectors/postgres/_target.py)
- Duplicated block (10 lines × 3) (examples/amazon_s3_embedding/main.py)
- Duplicated block (10 lines × 3) (python/cocoindex/connectors/falkordb/_target.py)
- Duplicated block (10 lines × 3) (python/cocoindex/connectors/lancedb/_target.py)
- Duplicated block (11 lines × 2) (examples/code_embedding/main.py)
- Duplicated block (11 lines × 2) (python/cocoindex/connectors/bigquery/_target.py)
- Duplicated block (11 lines × 2) (python/cocoindex/connectors/lancedb/_target.py)
- Duplicated block (11 lines × 2) (rust/sdk/cocoindex/src/doris.rs)
- Duplicated block (11 lines × 8) (examples/amazon_s3_embedding/main.py)
- Duplicated block (12 lines × 2) (examples/code_embedding_lancedb/main.py)
- Duplicated block (12 lines × 2) (examples/meeting_notes_graph_falkordb/main.py)
- Duplicated block (12 lines × 2) (python/cocoindex/_internal/function.py)
- Duplicated block (12 lines × 2) (python/cocoindex/connectors/bigquery/_target.py)
- Duplicated block (12 lines × 2) (python/cocoindex/connectors/bigquery/_target.py)
- Duplicated block (12 lines × 2) (python/cocoindex/connectors/falkordb/_target.py)
- Duplicated block (12 lines × 2) (python/cocoindex/connectors/falkordb/_target.py)
- …and 111 more
Changes since last survey
- 21 commits — 11 feature/other, 10 fixes
By area
- python/cocoindex — 6 commits
- .github/workflows — 3 commits
- python/tests — 3 commits
- rust/core — 3 commits
- docs/src — 2 commits
- (root) — 1 commit
- dev/agent-skills — 1 commit
- examples/rust — 1 commit
- rust/sdk — 1 commit
Notable commits
- fix: fix(engine): confine a merged target-action batch failure to the failing component (#2406)
- fix: fix(engine): reuse a concurrent same-key run's memo instead of re-executing (#2412)
- fix: fix(engine): reuse a same-operation memo under full_reprocess too (#2414)
- fix: fix(state_store): keep each LMDB write txn on one OS thread (#2426)
- fix: fix(target_state): key sink identity by type and reject non-weakref-able callbacks (#2404)
- fix: fix(valkey): purge document hashes when an index is deleted (#2419)
- fix: fix: copying a coco function returns the function itself (#2409)
- fix: fix: preserve StableKey variants through MessagePack decoding (#2403)
- fix: fix: preserve typed Python exceptions in component exception handlers (#2383)
- fix: fix: reference-count logic fingerprint registrations (#2405)
- change: Rewrite the GPUPool in Rust (#2243) (#2277)
- change: ci(deps): bump jlumbroso/free-disk-space from 1.3.1 to 2.0.0 (#2428)
- change: ci(deps): bump taiki-e/install-action from 2.87.11 to 2.87.15 (#2427)
- change: ci(deps): bump taiki-e/install-action from 2.87.5 to 2.87.11 (#2415)
- change: docs: document container-deletion semantics (destroy vs abandon) (#2418)
- change: refactor(engine): drop the never-set orphaned flag on target state providers (#2420)
- change: refactor(py): remove unused not_set object from runtime init (#2422)
- change: refactor(target): fulfill child target providers through per-action child slots (#2407)
- change: test(connectors): App-level kafka/iggy topic abandon semantics (#2421)
- change: test(connectors): share kafka/iggy SDK stubs across test modules (#2423)
- …and 1 more
Architecture
- Containers 0 added · 0 removed · contexts 1 added · 0 removed · edges 0 added · 0 removed
Added bounded contexts (1)
- cocoindex
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
cocoindex-io/cocoindex was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 29 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit f3e19d15eb92a1f98349a53829cd90f36ffb8bfb — the exact code this score is about.
- Scored under rubric-2026.09.17 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fbec9b1e08c2.