Cinnamon/kotaemon
59.2
Weak · 18 September 2026
42k
lines of production code
Python
primary language
1
measurement over time
What this system is
Kotaemon is a modular framework and web application for building, managing, and interacting with Large Language Model pipelines. It provides a unified interface for configuring LLMs, embeddings, and vector stores, while supporting advanced reasoning agents like ReAct and ReWOO with tool integration. The system enables users to ingest, index, and retrieve data from various document sources, offering features such as citation generation, reranking, and interactive chat interfaces with thought visualization.
Features
Add web search retrievers via Jina and Tavily APIs
Users can now retrieve information from the web using two new retriever components: one powered by the Jina Reader API and another by the Tavily Search API. These components allow the system to fetch and format external web content as retrieved documents, requiring the respective API keys (JINA\_API\_KEY or TAVILY\_API\_KEY) to be configured. The Jina retriever fetches structured data including titles and descriptions, while the Tavily retriever performs advanced searches and aggregates results.
libs/kotaemon/kotaemon/indices/retrievers · high confidence
Added Alembic configuration for optional database migrations
The ktem library now includes an alembic.ini configuration file, enabling optional database schema migrations via Alembic. This addition allows users to manage database versioning and updates within the library, supporting the project's evolution as an official component.
libs/ktem · high confidence
Added ChromaDB migration script to sync vector metadata with SQLite sources
A new migration utility (\scripts/migrate/migrate\_chroma\_db.py\) has been introduced to synchronize metadata in ChromaDB vector collections with existing source records in SQLite. The script iterates through document sources, identifies their corresponding vector IDs via a 'vector' relation type, and updates the vector metadata to include the correct \file\_id\. This ensures that vector search results are correctly linked back to their original file sources after migration.
scripts/migrate · high confidence
Expanded document ingestion with specialized and AI-enhanced readers
The \kotaemon/loaders\ module now provides a comprehensive suite of document readers, significantly broadening the types of files the system can process. New capabilities include parsing Microsoft Office formats (Word, Excel) via \DocxReader\ and \ExcelReader\, extracting structured content from HTML and MHTML files, and leveraging cloud-based AI services like Adobe PDF Services and Azure AI Document Intelligence for high-accuracy text, table, and figure extraction. The system also integrates advanced OCR solutions, including PaddleOCR and a general \OCRReader\, as well as specialized scientific PDF parsing via Mathpix and layout-aware extraction using Docling. These loaders are unified under a common \BaseReader\ interface and exported via \\_\init\\_.py\, allowing users to easily swap or combine different ingestion strategies for their documents.
libs/kotaemon/kotaemon/loaders · high confidence
Initial release of Kotaemon with multi-provider LLM support and SSO authentication
This entry introduces the Kotaemon application, providing a unified interface for interacting with various Large Language Models (LLMs) and embedding services. The application supports OpenAI, Azure OpenAI, Google, Mistral, VoyageAI, and local models via Ollama, with configuration managed through environment variables and a dedicated settings file. It includes built-in Single Sign-On (SSO) authentication options for Google and Keycloak, allowing secure access in enterprise or demo environments. The release also features a Docker-based deployment strategy with multiple image variants (Lite, Full, PaddleOCR, Ollama-bundled) to accommodate different hardware capabilities and use cases, alongside comprehensive documentation and development tooling.
(repo-wide) · high confidence
Introduce DocumentIngestor for configurable file parsing and indexing
The \libs/kotaemon/kotaemon/indices/ingests\ module now provides a \DocumentIngestor\ component that centralizes the ingestion of local files (PDF, Office documents, HTML, Markdown, images) into indexed text nodes. Users can configure PDF extraction modes (normal, Mathpix, OCR, or multimodal), supply custom document parsers, and adjust text splitting parameters. The component automatically maps file extensions to appropriate readers (including new support for MHTML and integration with Docling, Azure AI Document Intelligence, and PaddleOCR) and handles the full pipeline of reading, splitting, and parsing files for downstream indexing.
libs/kotaemon/kotaemon/indices/ingests · high confidence
Introduce Kotaemon custom theme and PDF.js asset configuration
The application now includes a new 'Kotaemon' theme built on Gradio's Soft theme, defining specific color palettes, typography (Quicksand/IBM Plex Mono), and layout styles for both light and dark modes. Additionally, the assets module now explicitly configures the PDF.js distribution version (defaulting to 4.0.379) and its prebuilt directory path, making these assets available for PDF rendering within the application.
libs/ktem/ktem/assets · high confidence
Introduce Ktem application framework with modular settings and extension support
The Ktem application framework is introduced, providing a structured base for the Kotaemon application. This includes a new settings system (libs/ktem/ktem/settings.py) that organizes configuration into application, index, and reasoning groups, allowing for dynamic UI rendering of settings. The core app logic (libs/ktem/ktem/app.py) now supports extension registration via a pluggy-based plugin system, enabling modular feature addition. Additionally, the framework introduces a ModelPool component (libs/ktem/ktem/components.py) for managing and selecting LLM/embedding models based on criteria like accuracy or cost, and integrates with Gradio for the UI, including support for user management, SSO, and demo modes as indicated by the main app entry point (libs/ktem/ktem/main.py).
libs/ktem/ktem · high confidence
Introduce PromptUI for interactive pipeline configuration and logging
A new PromptUI module has been added to the contribs package, enabling users to automatically generate interactive Gradio-based interfaces for kotaemon pipelines. This feature allows users to configure pipeline parameters, inputs, and outputs via a web UI, run pipelines, and manage chat sessions (for chatbots). It also includes capabilities to save and reload parameter history, view execution logs, and export detailed run results and preferences to Excel spreadsheets for comparison.
libs/kotaemon/kotaemon/contribs · high confidence
Introduce ReAct and ReWOO reasoning agent pipelines with MCP tool support
Users can now build multi-step reasoning agents using two new paradigms: ReAct (sequential thought-action loops) and ReWOO (planning with deferred evidence resolution). The \kotaemon/agents\ module provides \ReactAgent\ and \RewooAgent\ classes, along with a \LangchainAgent\ wrapper for compatibility with LangChain's agent types. Additionally, the system now supports Model Context Protocol (MCP) tools via \MCPTool\, allowing agents to discover and execute tools from external MCP servers. A suite of built-in tools is also included, such as \GoogleSearchTool\, \WikipediaTool\, and \LLMTool\, enabling agents to interact with external data sources and language models directly.
libs/kotaemon/kotaemon/agents · high confidence
Introduce UI-based LLM management with vendor selection and connection testing
Users can now add, edit, and delete language models directly through the application interface. The new LLM Management page allows selecting from supported vendors (such as OpenAI, Anthropic, Gemini, and Ollama) and configuring model specifications via YAML. It includes features to set a default model for the application, rename existing models, and test connectivity to ensure the LLM is reachable before use.
libs/ktem/ktem/llms · high confidence
Introduce base component and schema abstractions for the reasoning pipeline
This change establishes the foundational building blocks for the application's reasoning pipeline by introducing a new \kotaemon.base\ module. It defines a \BaseComponent\ class that serves as the standard interface for pipeline nodes, enabling features like auto-caching, logging, and deployment while enforcing a single output type. Additionally, it provides a set of standardized data schemas—including \Document\, \BaseMessage\ (with \SystemMessage\, \AIMessage\, \HumanMessage\), \RetrievedDocument\, and \LLMInterface\—to ensure consistent data flow and interoperability between components in the reasoning chain.
libs/kotaemon/kotaemon/base · high confidence
Introduce configurable reranking model management
This change adds a new reranking management system within the KTEM library, allowing users to define, store, and select reranking models via a database-backed pool. The implementation includes a database schema for persisting model specifications, a manager class to handle loading and switching between models (including support for vendors like Cohere, TEI, and VoyageAI), and a Gradio-based UI page that enables users to add, edit, rename, delete, and set default reranking models through a web interface.
libs/ktem/ktem/rerankings · high confidence
Introduce dedicated Embedding and MCP management modules
This change introduces two new management modules within the \ktem\ library: \ktem.embeddings\ and \ktem.mcp\. The \ktem.embeddings\ module provides a new database schema (\EmbeddingTable\) and a manager (\EmbeddingManager\) to persist and load embedding model configurations, exposing a Gradio UI (\EmbeddingManagement\) that allows users to add, edit, rename, set defaults, and test connections for various embedding vendors (such as OpenAI, Cohere, Google, Mistral, and TEI). Simultaneously, the \ktem.mcp\ module adds support for the Model Context Protocol by introducing an \MCPTable\ for storing server configurations and an \MCPManager\ for CRUD operations, accompanied by a Gradio UI (\MCPManagement\) that lets users add, edit, and delete MCP server configurations and view their discovered tools.
libs/ktem/ktem/embeddings · high confidence
Introduce dedicated pages for setup, login, help, and settings
The application now provides distinct, dedicated UI pages for first-time setup, user login, help documentation, and settings management. The setup page guides users through initial configuration of LLM and embedding providers (Cohere, Google, OpenAI, Ollama). The login page handles user authentication, supporting both local credentials and SSO integration. The help page displays application version, user guides, and changelogs, fetching remote documentation if local files are missing. The settings page allows users to customize application, index, and reasoning configurations, with user-specific settings support. These pages replace the previous monolithic interface structure, improving organization and user experience.
libs/ktem/ktem/pages · high confidence
Introduce default project template with pre-configured RAG pipelines
Users can now scaffold a new project using the \project-default\ template, which provides a ready-to-use structure including a \pipeline.py\ with pre-configured Question Answering and Indexing pipelines. The template sets up standard development tooling via \.pre-commit-config.yaml\ (Black, Isort, Flake8, MyPy) and \.gitattributes\, and includes a \setup.py\ that installs the \kotaemon\ library from the source repository.
templates · high confidence
Introduce file-based indexing and retrieval infrastructure
This change introduces the core \FileIndex\ component and its associated UI, pipelines, and storage logic, enabling users to upload, index, and retrieve content from local files. The implementation includes a new file storage system, SQL-backed source and index tables, and configurable indexing pipelines (including support for knowledge networks and vector retrieval). The UI provides file upload capabilities, directory indexing, and integration with the chat input for file selection and web search.
libs/ktem/ktem/index/file · high confidence
Introduce modular LLM pipeline components and provider wrappers
The \kotaemon.llms\ package now provides a structured set of components for building language model workflows. It introduces linear and branching pipeline strategies (\SimpleLinearPipeline\, \GatedLinearPipeline\, \SimpleBranchingPipeline\, \GatedBranchingPipeline\) that allow users to chain prompts, LLM calls, and post-processors with optional conditional gating. Chain-of-thought capabilities are exposed via \Thought\ and \ManualSequentialChainOfThought\ classes, enabling multi-step reasoning with manual prompt definition. Additionally, the library adds specific wrapper classes for various LLM providers, including OpenAI, Azure OpenAI, Anthropic, Gemini, Cohere, Ollama, and Llama.cpp, unifying their interfaces under a common \BaseLLM\ abstraction.
libs/kotaemon/kotaemon/llms · high confidence
Introduce modular chat LLM provider implementations
The chat module now provides a structured set of classes for interacting with various Large Language Models, including a base \ChatLLM\ interface, an \EndpointChatLLM\ for generic OpenAI-compatible endpoints, \LlamaCppChat\ for local models via llama-cpp-python (with Hugging Face Hub download support), and LangChain-based wrappers (\LCChatOpenAI\, \LCAnthropicChat\, \LCGeminiChat\, etc.). This change establishes the concrete implementation layer for chat capabilities, allowing users to select and configure specific LLM backends through a unified API.
libs/kotaemon/kotaemon/llms/chats · high confidence
Introduce official chatbot components and CLI interface
This change establishes the core chatbot architecture within the kotaemon library by adding a \BaseChatBot\ abstract class and a \ChatConversation\ component that manages message history and session state. It includes a \SimpleRespondentChatbot\ implementation that wraps a chat LLM, and a new CLI (\kh\) with commands to export pipeline configurations, run a Gradio-based UI with optional authentication and sharing, generate documentation, and scaffold new projects from templates. Additionally, telemetry from PostHog and Haystack is disabled by default via monkey-patching in the package initialization.
libs/kotaemon/kotaemon, libs/kotaemon/kotaemon/chatbot · high confidence
Introduce pluggable document store implementations
The library now provides a \BaseDocumentStore\ interface and four concrete implementations: \InMemoryDocumentStore\ for local testing, \SimpleFileDocumentStore\ for persistent JSON storage, \ElasticsearchDocumentStore\ for scalable BM25 search, and \LanceDBDocumentStore\ for vector/FTS capabilities. Users can now swap storage backends to suit their scale and persistence needs while maintaining a consistent API for adding, querying, and deleting documents.
libs/kotaemon/kotaemon/storages/docstores · high confidence
Introduce structured citation and inline citation support in QA pipelines
The QA index module now includes a \CitationPipeline\ that uses LLM function calling to extract specific evidence quotes for answers, and an \AnswerWithInlineCitation\ component that generates answers with inline citation markers (e.g., 【1】) linked to source text spans. These components are integrated into the \AnswerWithContextPipeline\ via new \enable\_citation\ and \enable\_mindmap\ flags, allowing users to opt into structured citation generation and inline citation rendering in their QA responses.
libs/kotaemon/kotaemon/indices/qa · high confidence
Introduce structured prompt templating and component base classes
The \kotaemon.llms.prompts\ module now provides a structured way to manage LLM prompts. Users can define reusable prompt templates via the new \PromptTemplate\ class, which supports standard string formatting placeholders and validates that all required arguments are provided before rendering. These templates can be used directly or wrapped in \BasePromptComponent\, which handles input validation, type conversion (supporting strings, integers, Documents, and callables), and integrates with the \theflow\ pipeline to return structured \Document\ outputs. This change establishes the foundational API for building consistent, validated prompt workflows within the library.
libs/kotaemon/kotaemon/llms/prompts · high confidence
Introduce unified embedding component library with multi-vendor support
The embeddings module now provides a standardized set of components for generating text embeddings, allowing users to select from various providers and local inference options. The library includes wrappers for OpenAI (including Azure), LangChain-based integrations for OpenAI, Azure, Cohere, Google, Hugging Face, and Mistral, as well as direct support for TEI (Text-Embeddings-Inference) endpoints, Voyage AI, and local fastembed models. This unifies the embedding experience under a common interface, enabling seamless switching between cloud APIs and self-hosted solutions.
libs/kotaemon/kotaemon/embeddings · high confidence
Introduce user-managed index system with UI for creation and configuration
The application now supports creating, updating, and deleting searchable indices (such as file or chat message indexes) through a new Index Management page. Users can define index names, select from available index types, and provide YAML-based configuration specs. The system persists these indices in the database and provides a UI to list, edit, and remove them, with changes requiring a system restart to take effect.
libs/ktem/ktem/index · high confidence
Introduce vector indexing and retrieval components
The \kotaemon.indices\ package now provides \VectorIndexing\ and \VectorRetrieval\ components, enabling users to ingest documents into a vector store and perform similarity-based retrieval. \VectorIndexing\ handles document ingestion, embedding generation, and storage, while \VectorRetrieval\ supports vector, text, and hybrid search modes with configurable top-k results and optional reranking.
libs/kotaemon/kotaemon/indices · high confidence
New GraphRAG indexing modes (LightRAG, Nano-GraphRAG, and standard GraphRAG)
The file indexing system now supports three GraphRAG-based indexing modes: standard GraphRAG, LightRAG, and Nano-GraphRAG. This adds new index classes (GraphRAGIndex, LightRAGIndex, NanoGraphRAGIndex) and their corresponding pipelines, enabling users to index files using graph-based retrieval. LightRAG and Nano-GraphRAG modes allow configuring batch sizes and search types (e.g., local search), while standard GraphRAG requires the GRAPHRAG\_API\_KEY environment variable and supports optional custom settings via settings.yaml. These modes replace or supplement previous vector-based indexing for graph-oriented queries.
libs/ktem/ktem/index/file/graph · high confidence
New LLM-based and Cohere reranking components for document relevance filtering
The \kotaemon.indices.rankings\ module now provides several new components to re-rank or filter documents based on relevance to a query. Users can use \CohereReranking\ to leverage the Cohere API (defaulting to the \rerank-v4.0-fast\ model) for scoring, or choose from LLM-driven approaches: \LLMReranking\ for binary relevance filtering, \LLMScoring\ for adding numerical relevance scores based on log-probabilities, and \LLMTrulensScoring\ which implements a TruLens-inspired 0-10 relevance grading system with context trimming. These components integrate with the existing \BaseComponent\ interface and support concurrent execution for performance.
libs/kotaemon/kotaemon/indices/rankings · high confidence
New PDF and multimodal document processing utilities
Added a suite of utility modules in \libs/kotaemon/kotaemon/loaders/utils\ to support advanced document ingestion. This includes \adobe.py\ for extracting text, tables, and figures from PDFs via the Adobe PDF Services SDK, \gpt4v.py\ for generating text responses from images using Azure OpenAI GPT-4V, and \pdf\ocr.py\ for merging OCR results with structured PDF text using bounding box overlap. The update also introduces \box.py\ for geometric operations on bounding boxes, \table.py\ for parsing and formatting tables into Markdown, and \\\init\\_.py\ (moved from \knowledgehub/contribs\) to expose these new capabilities.
libs/kotaemon/kotaemon/loaders/utils · high confidence
New ReAct and ReWOO reasoning pipelines with thought visualization
The application now includes two advanced reasoning pipelines—ReAct and ReWOO—allowing users to select a more complex, multi-step agent approach for answering questions. These pipelines enable the LLM to plan and execute a sequence of actions (such as searching documents, using Google or Wikipedia, or calling external tools) before generating a final answer, providing greater transparency into the reasoning process. The ReWOO pipeline specifically supports thought visualization, allowing users to see the step-by-step plan and evidence retrieval, while both pipelines integrate with the existing document retrievers and support configurable user settings.
libs/ktem/ktem/reasoning · high confidence
New Resources tab with integrated management for indices, LLMs, embeddings, rerankings, MCP servers, and users
A new 'Index Collections' tab has been added to the application, providing a centralized interface for managing core system components. Users can now configure and manage Index Collections, LLMs, Embeddings, Rerankings, and MCP Servers directly from this location. Additionally, if user management is enabled, a 'Users' tab becomes available to administrators, allowing them to create, edit, and delete user accounts with role-based visibility controls.
libs/ktem/ktem/pages/resources · high confidence
New document extraction and splitting components
The library now includes new modules for document indexing under \indices/extractors\ and \indices/splitters\. Users can utilize \TitleExtractor\ and \SummaryExtractor\ to generate metadata from documents, and \TokenSplitter\ and \SentenceWindowSplitter\ to chunk text, all implemented as wrappers around LlamaIndex core classes.
libs/kotaemon/kotaemon/indices/extractors, libs/kotaemon/kotaemon/indices/splitters · high confidence
New prompt optimization pipelines for question rewriting, decomposition, and mindmaps
The \libs/ktem/ktem/reasoning/prompt\_optimization\ module now includes new pipelines that enhance how user queries are processed. \DecomposeQuestionPipeline\ breaks complex questions into multiple specific sub-questions, while \FewshotRewriteQuestionPipeline\ uses vector-based retrieval of training examples to rewrite questions for better context. Additionally, \CreateMindmapPipeline\ generates structured mindmaps from questions and context, and \SuggestConvNamePipeline\ and \SuggestFollowupQuesPipeline\ provide conversation naming and follow-up question suggestions based on chat history.
_libs/ktem/ktem/reasoning/prompt\optimization · high confidence
New reranking components for Cohere, TEI, and VoyageAI
This change introduces a new modular reranking subsystem within the library, providing three distinct implementations for re-ordering documents by relevance: CohereReranking (defaulting to the rerank-v4.0-fast model with configurable base URL), TeiFastReranking (for self-hosted Text Embeddings Inference services with batch processing and truncation options), and VoyageAIReranking (defaulting to the rerank-2 model). These components share a common BaseReranking interface and are exported from the rerankings package, allowing users to easily swap or integrate different third-party or local reranking services into their pipelines.
libs/kotaemon/kotaemon/rerankings · high confidence
New utility modules for chat, rendering, and visualization
The \ktem/utils\ package introduces several new modules that enhance the user experience in chat interactions, document handling, and data visualization. Users can now see @ mentions (including file tags and the new WebSearch command) highlighted in chat bubbles, with automatic extraction of URLs and clean query preparation for the LLM. Document previews and evidence displays are improved with better markdown rendering, PDF page preview links, and support for image and table types. Additionally, a new citation visualization feature allows users to view retrieved documents and queries projected in a 2D embedding space using UMAP and Plotly, aiding in understanding retrieval relevance. The package also adds support for multiple languages, YAML date handling fixes, and rate limiting for authenticated users.
libs/ktem/ktem/utils · high confidence
New vector store implementations for Milvus and Qdrant
The library now includes dedicated vector store classes for Milvus and Qdrant, allowing users to persist and query embeddings using these specific backends. These new classes, along with existing support for Chroma, LanceDB, and in-memory storage, are unified under a common base interface in the \kotaemon.storages.vectorstores\ module, enabling consistent interaction with different vector databases.
libs/kotaemon/kotaemon/storages/vectorstores · high confidence
One-click installers and local model serving for Linux, Windows, and macOS
The application now includes platform-specific installer scripts (run\_linux.sh, run\_macos.sh, run\_windows.bat) that automate the setup of Miniconda, Python environments, and dependencies, while also handling the download of PDF.js and launching the web UI. A new local model serving capability is introduced via scripts/serve\_local.py and corresponding platform-specific server scripts (server\_llamacpp\_linux.sh, server\_llamacpp\_macos.sh, server\_llamacpp\_windows.bat), which allow users to start a llama-cpp-python inference server for local GGUF models. Additionally, update scripts (update\_linux.sh, update\_macos.sh, update\_windows.bat) are provided to upgrade the application to the latest version.
scripts · high confidence
Optional database migration support via Alembic
The application now includes Alembic migration infrastructure within the \libs/ktem/migrations\ directory, enabling database schema versioning. This feature is opt-in: users must explicitly set the \KH\_ENABLE\_ALEmbic\ configuration flag to \True\ to activate the migration environment. The implementation provides the necessary Alembic configuration (\env.py\), migration script templates (\script.py.mako\), and a placeholder for versioned migration files.
libs/ktem/migrations · high confidence
Removals
Removal of empty base loader classes
The empty base classes DocumentLoader, TextManipulator, and DocumentManipulator have been removed from the knowledgehub loaders module. This cleanup eliminates unused placeholder code from the codebase.
knowledgehub · high confidence
Behavioural changes
Introduce new chat interface with reasoning, suggestions, and feedback
The chat page now features a redesigned UI that supports advanced reasoning with thought visualization, inline citations, and web search capabilities. Users can access chat suggestions, toggle dark mode, and provide feedback on answer correctness directly from the interface. The layout includes a dedicated control panel for managing conversations, a hint section for guidance, and a paper list for browsing popular daily papers in demo mode.
libs/ktem/ktem/pages/chat · high confidence
Introduces customizable database models and optional Alembic migration support
The database layer now allows end-users to define custom table structures for conversations, users, settings, and issue reports via configuration settings (KH\_TABLE\_CONV, KH\_TABLE\_USER, etc.), falling back to built-in base models if not specified. Additionally, database schema creation is now gated by a KH\_ENABLE\_ALEMBIC setting, allowing users to opt into Alembic-based migrations instead of automatic table creation.
libs/ktem/ktem/db · high confidence
New client-side UI initialization and PDF viewer integration
The application now includes a new \main.js\ script that handles initial UI setup, including adding a version badge, setting a favicon, repositioning the information panel expand button and the conversation control sidebar toggle, and converting the 'suggest chat' checkbox into a toggle switch. Additionally, a new \pdf\_viewer.js\ module provides a modal PDF viewer that opens when citation links are clicked, allowing users to view and search within PDF documents directly from the chat interface. The \svg-pan-zoom.min.js\ library has also been added to support interactive zooming and panning for SVG-based visualizations.
libs/ktem/ktem/assets/js · high confidence
New main.css stylesheet for chat UI layout and styling
A new main.css file has been added to the assets directory, introducing comprehensive styling for the chat interface. This includes layout adjustments for the main area height, header bar, and specific tabs (chat, indices, settings, help, resources, login). It also implements custom scrollbar styling, fixes for Firefox Gradio table overflow, and visual tweaks for elements like the chat info panel, conversation settings panel, and message rows.
libs/ktem/ktem/assets/css · high confidence
Test coverage
Added test coverage for conversation mentions, MCP manager, and QA pipeline; Added test infrastructure and fixtures for kotaemon.
Dependencies
Adopt pyproject.toml and uv workspace for dependency management
The project has migrated from legacy requirements files to a structured pyproject.toml configuration, introducing a uv workspace that manages the core kotaemon library, the ktem application layer, and the top-level app package. This change standardizes dependency declarations, enforces specific version constraints for key libraries such as Gradio (\>=4.31.0,\<5), FastAPI (\<=0.112.1), and PyMuPDF (\<=1.24.11), and pins python-multipart to 0.0.12 to prevent installation issues with micropip. It also enables faster installation workflows via the uv package manager and organizes optional dependencies for advanced features like vector stores and OCR into distinct extras.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 59.
Lenses
- Code Health 88
- Architecture 99
- Maturity 70
- Readiness 45
- Security 60
- Domain Modelling 100
Changes since last survey
- 300 commits — 208 feature/other, 92 fixes
By area
- libs/ktem — 104 commits
- (root) — 66 commits
- libs/kotaemon — 59 commits
- .github/workflows — 14 commits
- knowledgehub/pipelines — 9 commits
- docs/images — 5 commits
- knowledgehub/contribs — 5 commits
- knowledgehub/llms — 4 commits
- knowledgehub/loaders — 4 commits
- knowledgehub/storages — 4 commits
- docs/pages — 3 commits
- knowledgehub/agents — 3 commits
- knowledgehub/indices — 3 commits
- scripts/run_windows.bat — 3 commits
- (repo) — 2 commits
- knowledgehub/base — 2 commits
- knowledgehub/embeddings — 2 commits
- .github/ISSUE_TEMPLATE — 1 commit
- docs/index.md — 1 commit
- docs/theme — 1 commit
Notable commits
- fix: (bump:patch) Fix: llama-cpp-python security bug and setup local latest branch in github action (#66)
- fix: Allow users to select reasoning pipeline. Fix small issues with user UI, cohere name (#50)
- fix: Fix UI bugs (#8)
- fix: Fix Yaml datetime format (#79)
- fix: Fix info panel overflow (#33)
- fix: Fix integrating indexing and retrieval pipelines to FileIndex (#155)
- fix: Fix loaders' file_path and other metadata
- fix: Fix subscribing sign-in/out
- fix: bug fix: settings are not persistent
- fix: ci: fix release .env
- fix: ci: revert GH env var
- fix: fix HelloGitHub Badge code (#313)
- fix: fix bug in delete file, remove file delete confirmation (#59)
- fix: fix(docstore): preserve retrieval ranking order in lancedb get() (#745)
- fix: fix: Gradio table display with Firefox
- fix: fix: PDFJS check in run_windows bump:patch
- fix: fix: Remove Collections from all index names (#473)
- fix: fix: UI tab name and reranking process for TeiFastReranking (#576)
- fix: fix: ValueError breaks Ui #629 (#630) #none
- fix: fix: add PDFJS download to Windows setup (#249)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
Cinnamon/kotaemon was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.