Skip to content
CAI
Software that uses CAICheck a score

zylon-ai/private-gpt

64.8

Adequate · 26 September 2026

82.5k

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

PrivateGPT is a self-hosted, modular AI application that provides a private interface for interacting with large language models through document ingestion and retrieval. It exposes an OpenAI-compatible chat API supporting streaming, tool use, and structured outputs, while managing background tasks via Celery and ARQ for reliable document parsing and vector indexing. The system features a flexible configuration model, supports multiple LLM providers, and includes a lightweight static UI for demonstration purposes.

How it got here

2023 — PrivateGPT v1.0 architecture revamp

14 changes.

This period focused on the comprehensive v1.0.0 overhaul of PrivateGPT, introducing a modular architecture with dependency injection, asynchronous worker infrastructure via ARQ, and a new Pydantic-based configuration system. The work established robust streaming chat capabilities, standardized CI/CD and development workflows, and implemented extensive test coverage to validate the new ingestion, chat, and settings subsystems.

2026 — API compatibility and infrastructure modernization

18 changes.

The project modernized its core infrastructure by refactoring the Celery worker subsystem for pluggable backends and introducing a dedicated task for tool execution. It simultaneously revamped the chat interface to support an OpenAI-compatible API with advanced input types and rebuilt the event system for Anthropic-compatible streaming. This period also included the release of a static demo UI and extensive test coverage for the new components.

Features

Introduce PrivateGPT Workbench as a lightweight static demo UI

The \ui\ directory now contains the initial implementation of the PrivateGPT Workbench, a single-file static HTML application (\index.html\) designed to demonstrate the PrivateGPT API capabilities. This change adds the core runtime file along with comprehensive documentation (\PRD.md\, \STYLE\_GUIDE.md\, \SOURCE\_OF\_TRUTH.md\) and agent-specific workflow instructions (\AGENTS.md\, \CLAUDE.md\) to guide future development. The UI provides a local-first interface for users to interact with the PrivateGPT API, featuring a sidebar for navigation, a chat interface, and a debugger, while adhering to a specific visual style guide and API contract defined in the Fern-generated OpenAPI schema.

ui · high confidence

Introduces ARQ-based asynchronous worker infrastructure for resumable chat and tool execution

This change adds a new \private\_gpt/arq\ package that implements an asynchronous task queue using the ARQ library backed by Redis. It introduces a \HeartbeatWorker\ with custom lease management to handle job ownership and stale lock recovery, enabling reliable multi-process execution. The infrastructure supports resumable chat sessions by persisting iteration state and tool results in Redis via \RedisChatCheckpointStore\, allowing the system to recover from interruptions. It also includes components for publishing routes, managing worker health checks, and auto-discovering chat and tool tasks, fundamentally shifting the backend from synchronous or Celery-based processing to a more robust, stateful async model.

_private\gpt/components · high confidence

New PrivateGPT maintenance and deployment scripts

The scripts directory now includes a suite of new tooling: auto\_discover\_models.py automatically detects available LLM and embedding models from remote OpenAI-compatible APIs and writes a settings profile; build\_pip\_index.py generates a static PEP 503 package index for local wheel distribution; extract\_openapi.py exports the application's OpenAPI specification; generate\_homebrew\_formula.py creates Homebrew installation formulas; ingest\_folder.py provides a command-line interface for local folder ingestion with watch and ignore capabilities; set\_version.py synchronizes version numbers across project files; update\_claude\_specs.py keeps Anthropic SDK versions and OpenAPI spec URLs in sync; and worker\_entrypoint configures environment variables for worker processes.

scripts · high confidence

New chat API with streaming, tool use, and structured outputs

The server now exposes a new chat endpoint at POST /v1/messages that supports streaming responses, tool use (including MCP server integration), and structured JSON outputs. Users can configure sampling parameters (temperature, top\_p, seed), enable thinking/reasoning modes, and manage context through system prompts and document mounts. The implementation includes a comprehensive interceptor chain for handling citations, condensation, and tool execution, replacing the previous chat interface.

_private\gpt/server · high confidence

New utility modules for concurrency, batching, and token estimation

The \private\_gpt/utils\ package has been expanded with new modules to support the revamped ingestion and chat pipelines. \concurrency.py\ introduces \bounded\_concurrent\_execute\ and \map\_elements\_in\_parallel\ for controlled async task execution with jitter and worker limits. \async\_utils.py\ and \batches.py\ provide helpers for converting sync iterators to async ones and batching both sync and async iterables with optional stop conditions. \retry.py\ wraps the \retry\_async\ library for robust error handling. \tokens.py\ adds \estimate\_token\_count\ to accurately calculate token usage in chat history, considering tools and reasoning effort, while \token.py\ provides logic for calculating maximum token expansion limits. Additional utilities include \eta.py\ for progress reporting, \pool.py\ for object pooling, and \dataframe.py\ for markdown conversion.

_private\gpt/utils · high confidence

PrivateGPT v1.0.0 revamp: new Docker build, settings schema, and CI/CD artifacts

This release introduces the PrivateGPT v1.0.0 architecture, replacing the previous setup with a multi-stage Dockerfile that uses \uv\ for dependency management and supports optional extras (core, ingest, media, tools, database) to minimize image size. The configuration system is overhauled with a new \settings.yaml\ that exposes granular controls for chat engines, retrieval, document parsing (Docling), and vector stores, alongside dedicated \settings-mock.yaml\ and \settings-test.yaml\ profiles. The project also adds a \Makefile\ for standardized development workflows (test, lint, dev, prod-run), a \SECURITY.md\ for private vulnerability reporting, and a \CITATION.cff\ for academic referencing.

(repo-wide) · high confidence

Behavioural changes

Celery infrastructure refactored with configurable backends and stateful worker support

The Celery worker subsystem has been restructured to support pluggable broker and backend configurations (local filesystem, Redis, or RabbitMQ) via new \broker\_config.py\ and \backend\_config.py\ modules, allowing users to select their preferred infrastructure mode. A new \bootsteps.py\ introduces liveness probes and dedicated warm-up/shutdown logic for stateful workers, ensuring models are loaded safely in preforked child processes. The \base.py\ module now provides a \StatefulBackgroundTask\ base class with controlled retry mechanisms and failure tracking, while \callback.py\ and \notify.py\ implement a unified progress notification system that sends structured updates (including fake progress for long-running tasks) and final results to the API. Additionally, \healthcheck.py\ adds a dedicated health endpoint that monitors worker readiness and heartbeat files, and \task\_helper.py\ provides utilities to inspect and revoke specific ingestion or deletion tasks by collection and artifact.

_private\gpt/celery · high confidence

Dedicated Celery task for tool execution

A new Celery task, \private\_gpt.tools.run\, has been introduced to handle tool execution on a dedicated worker. This change moves tool execution out of the main chat flow, allowing for isolated processing, duplicate execution suppression via checkpoint claiming, and structured error handling. Users benefit from improved reliability and separation of concerns when tools are invoked during chat interactions.

_private\gpt/celery/tasks/tools · high confidence

Ingestion pipeline split into separate parse and vector-store tasks

The background ingestion process has been restructured from a single monolithic operation into two distinct Celery tasks: \parse\_task\ and \store\_vectors\_task\. This change allows the system to handle document parsing and vector indexing as separate steps, improving reliability and enabling features like parse-only mode for chat document conversion. The \parse\_task\ now validates and parses files, then automatically dispatches the \store\_vectors\_task\ to handle embedding generation and persistence, while also managing cleanup of temporary files and handling cancellation scenarios.

_private\gpt/celery/tasks/ingestion · high confidence

Introduces centralized Pydantic-based configuration with profile-based YAML loading

The application now uses a structured settings system where configuration is defined via Pydantic models (e.g., CorsSettings, AuthSettings, ProxySettings) and loaded from YAML files. This change introduces a profile-based loading mechanism that merges a 'default' profile with an optional 'override' profile and any additional profiles specified via the PGPT\_PROFILES environment variable. It also supports environment variable expansion in YAML using the ${VAR} syntax and allows dynamic model discovery via PGPT\MODELS\ environment variables, replacing the previous ad-hoc configuration approach.

_private\gpt/settings · high confidence

New OpenAI-compatible chat API with advanced input and system configuration

The chat module has been revamped to provide a new OpenAI-compatible API interface. This introduces structured input models supporting complex message formats (text, images, audio, tool use), system-level configuration for enabling features like citations, code execution, and skills, and extensions for context filtering and Zylon-format citations. Users can now attach files, control blob visibility modes, and leverage detailed prompt injection controls via the new schema models.

_private\gpt/chat · high confidence

New streaming event models and interceptors for Anthropic-compatible and Zylon-specific responses

The event system has been rebuilt with a new Pydantic-based model layer in \private\_gpt/events/models\ and a new interceptor pipeline in \private\_gpt/events/interceptors\. Users now benefit from a unified streaming event format that supports both standard Anthropic-compatible content blocks (text, images, audio, tool use) and Zylon-specific extensions (source attribution, TLDR summaries, MCP token refresh notifications). The new \FilterZylonEventInterceptor\ automatically filters out Zylon-only events when the response mode is set to Anthropic, ensuring clean wire format compatibility, while the \PingEventInterceptor\ provides reliable SSE keepalive pings during idle periods to prevent connection timeouts.

_private\gpt/events/models · high confidence

Private GPT CLI revamped with new command structure and server management

The Private GPT command-line interface has been restructured to provide a more modular and robust user experience. The new CLI introduces distinct commands for managing the application lifecycle: 'serve' starts the HTTP server with support for PID files (enabling integration with systemd/launchd) and auto-reload for development, while 'run' launches connected code agents (such as Claude Code, OpenClaw, or OpenCode) with automatic model selection and interactive paging. A 'worker' command is now available to handle background tasks, but it is strictly gated behind the presence of the Celery library, preventing execution errors when Celery is not installed. The main entry point also includes improved command validation with typo suggestions for unknown commands.

_private\gpt/cli · high confidence

PrivateGPT v1 revamp: new architecture and streaming support

PrivateGPT has been revamped to version 1, introducing a new modular architecture that replaces the previous global state with a structured dependency injection system (using the \injector\ library) and request-scoped context propagation. This change enables robust, isolated worker profiles (chat, tools, full) with eager component loading to improve startup performance. The update also adds comprehensive Server-Sent Events (SSE) support for real-time chat streaming, including a new event model, formatter, and stream manager. Additionally, the application now uses Python's built-in \logging\ module instead of \loguru\, configures a centralized \PGPT\_HOME\ path for local data and caches, and integrates with LlamaIndex 0.10 for document indexing and retrieval.

_private\gpt · high confidence

Test coverage

Added comprehensive test coverage for chat engines and citation processing; Added comprehensive test coverage for core system components; Added comprehensive test coverage for the chat server subsystem; Added test coverage for context propagation, content retrieval, and dependency injection; Added test fixtures for PrivateGPT integration testing; Added tests for Anthropic model schema parity and drift detection; Added tests for Arq worker infrastructure; Added tests for MCP server OAuth token refresh and tool execution; Added tests for chat interceptor reliability and correctness; Added tests for content block serialization, pruning, and MCP token events; Added tests for model capabilities and primitive chunk retrieval; Added tests for server utility functions; Added tests for settings loading and validation; Added tests for skill validation and wrapper directory handling; Added tests for the new artifact ingestion and conversion system; Added tests for the server files API and Anthropic SDK integration; Added unit tests for SSE streaming components.

Dependencies

Initial release of PrivateGPT v1.0.1 with Poetry dependency management

This change introduces the first official release (v1.0.1) of PrivateGPT, migrating the project to use Poetry for dependency management via a new pyproject.toml file. The configuration defines the core application dependencies, including FastAPI, LlamaIndex, and various LLM providers (OpenAI, Anthropic, Mistral), as well as optional extras for document ingestion, vector stores (Qdrant), and database backends (Postgres, MySQL, MSSQL).

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 46 → 65 (+19.2)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 89 → 85 (-4.0)
  • Architecture 94 → 98 (+4.1)
  • Maturity 51 → 56 (+5.6)
  • Readiness 27 → 71 (+44.5)
  • Security 59 → 69 (+9.7)
  • Domain Modelling 100 (new)
  • Accessibility 67 (new)

Resolved (64)

  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • Duplicated block (10 lines × 2) (private_gpt/chat/input_models.py)
  • Duplicated block (10 lines × 2) (private_gpt/components/engines/chat/async_chat_engine.py)
  • Duplicated block (10 lines × 2) (private_gpt/components/ingest/transformations/markdown_to_tree_transform.py)
  • Duplicated block (10 lines × 2) (tests/components/chat/test_tldr_side.py)
  • Duplicated block (12 lines × 2) (private_gpt/components/engines/chat/chat_engine.py)
  • Duplicated block (12 lines × 2) (private_gpt/components/tabular/pandasai_sandbox.py)
  • Duplicated block (13 lines × 2) (private_gpt/components/vector_store/patched_qdrant_store.py)
  • Duplicated block (14 lines × 2) (private_gpt/components/readers/text/email_reader.py)
  • Duplicated block (14 lines × 2) (private_gpt/components/tools/builders/database_query_builder.py)
  • Duplicated block (15 lines × 2) (private_gpt/components/readers/text/email_reader.py)
  • Duplicated block (16 lines × 2) (private_gpt/server/chat/interceptors/extract_citation_interceptor.py)
  • Duplicated block (17 lines × 2) (tests/components/chat/test_tldr_side.py)
  • Duplicated block (19 lines × 2) (tests/models/anthropic/test_openapi_schema.py)
  • Duplicated block (20 lines × 2) (tests/components/transforms/test_markdown_node_transform.py)
  • Duplicated block (5 lines × 2) (tests/server/events/test_content_block_pruning.py)
  • Duplicated block (6 lines × 2) (private_gpt/components/multimodality/audio_handler.py)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • …and 44 more

New (492)

  • Ambiguous method signature for get_queue_name. In arq.settings, it takes a queue name string (likely for normalization), while in arq.tasks.chat.settings, it takes a Settings object. This inconsistency in parameter types for the same method name across related modules is confusing.
  • AsyncChatEngine._handle_stream_chunk (cognitive 91) (private_gpt/components/engines/chat/async_chat_engine.py)
  • AsyncChatEngine._handle_stream_chunk (cyclomatic 40) (private_gpt/components/engines/chat/async_chat_engine.py)
  • AsyncChatEngine._initialize_run (cognitive 18) (private_gpt/components/engines/chat/async_chat_engine.py)
  • AsyncChatEngine._run_iteration (cognitive 24) (private_gpt/components/engines/chat/async_chat_engine.py)
  • AsyncChatEngine._run_iteration (cyclomatic 20) (private_gpt/components/engines/chat/async_chat_engine.py)
  • Change coupling: embedding_component.py ↔ llm_component.py (private_gpt/components/embedding/embedding_component.py)
  • ChatBody.validate_properties (cognitive 97) (private_gpt/server/chat/chat_models.py)
  • ChatBody.validate_properties (cyclomatic 58) (private_gpt/server/chat/chat_models.py)
  • ChatLoopEngine._handle_stream_chunk (cognitive 91) (private_gpt/components/engines/chat/chat_engine.py)
  • ChatLoopEngine._handle_stream_chunk (cyclomatic 40) (private_gpt/components/engines/chat/chat_engine.py)
  • ChatLoopEngine._run_iteration (cognitive 20) (private_gpt/components/engines/chat/chat_engine.py)
  • ChatLoopEngine._run_iteration (cyclomatic 19) (private_gpt/components/engines/chat/chat_engine.py)
  • CitationRequestInterceptor.intercept (cognitive 20) (private_gpt/server/chat/interceptors/citation_interceptor.py)
  • CondenserContextMemoryStrategy._condense_from_left (cognitive 20) (private_gpt/components/chat/processors/chat_history/memory/strategies/condenser.py)
  • CondenserContextMemoryStrategy._condense_from_left (cyclomatic 16) (private_gpt/components/chat/processors/chat_history/memory/strategies/condenser.py)
  • ContentService._filter_tree_nodes (cognitive 37) (private_gpt/server/content/content_service.py)
  • ContentService._filter_tree_nodes (cyclomatic 17) (private_gpt/server/content/content_service.py)
  • ContentService._retrieve_document_node (cognitive 25) (private_gpt/server/content/content_service.py)
  • ContentService._retrieve_document_node (cyclomatic 16) (private_gpt/server/content/content_service.py)
  • …and 472 more

Changes since last survey

  • 27 commits — 15 feature/other, 12 fixes

By area

  • private_gpt/components — 15 commits
  • (root) — 4 commits
  • private_gpt/server — 4 commits
  • .github/workflows — 2 commits
  • fern/docs — 1 commit
  • private_gpt/arq — 1 commit

Notable commits

  • fix: Update nltk to >=3.10.0 to fix [CVE redacted] (#2336)
  • fix: fix(tabular): force-close sandbox after analysis to avoid leaks (#2357)
  • fix: fix(tldr): never leave a condensation block open on the client (#2387)
  • fix: fix: extended thinking (#2333)
  • fix: fix: keep model download lock PID when acquisition fails (#2362)
  • fix: fix: local resource in sql biulder when code execution is enabled (#2365)
  • fix: fix: openai compatibility (#2340)
  • fix: fix: refresh flag exception (#2341)
  • fix: fix: reject invalid skill metadata (#2374)
  • fix: fix: support Windows model download locking (#2352)
  • fix: fix: tabular sandbox (#2359)
  • fix: fix: worker health (#2358)
  • change: chore(deps): bump astral-sh/setup-uv from 10.0.1 to 10.1.0 (#2371)
  • change: chore(deps): bump astral-sh/setup-uv from 9.0.0 to 10.0.1 (#2337)
  • change: chore(deps): update Claude specs and anthropic SDK (#2347)
  • change: docs: add OrcaRouter OpenAI-compatible gateway configuration example (#2331)
  • change: docs: drop the duplicated word in the chat mapper docstring (#2378)
  • change: docs: fix private_pgt typo in settings.yaml comment (#2361)
  • change: docs: fix star history chart (#2335)
  • change: feat: add compatibility with skill creator (#2342)
  • …and 7 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

zylon-ai/private-gpt was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 01ac43d72bc4c0b7565994197fe23476539059a7 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.