PromtEngineer/localGPT
48.4
Weak · 19 September 2026
18k
lines of production code
Python
with TypeScript
1
measurement over time
What this system is
This system is a local Retrieval-Augmented Generation (RAG) application that enables users to ingest, index, and query multimodal documents using a modular Python backend and a React-based chat interface. It supports hybrid search with reranking, query decomposition, and evidence-gated answer verification, while allowing flexible LLM backends including local Ollama models and IBM Watson X. The platform provides comprehensive evaluation harnesses for measuring retrieval and generation quality, along with robust deployment options via Docker.
Features
Introduce RAG system with multimodal ingestion, Watson X support, and evidence-gated defaults
The \rag\_system\ package is now available, providing the retrieval-augmented generation engine for localGPT. It supports document ingestion via Docling (including OCR for scanned PDFs) and stores indexes in LanceDB. The system integrates with both Ollama and IBM Watson X (using Granite models) as LLM backends. Retrieval uses hybrid search (vector + full-text) with reranking by the Qwen3-Reranker-4B model, which is enabled by default. Query decomposition is active but uses a pooled first-stage strategy by default, and late chunking and answer verification are disabled by default based on recent ablation studies. An ephemeral 'ask folder' mode is also included for one-off queries.
_rag\system · high confidence
Introduce core API client and type definitions for chat and RAG integration
Added new files in src/lib to support the multimodal RAG codebase integration. api.ts provides a client for interacting with the chat and RAG APIs, including health checks, session management, and message sending, along with TypeScript interfaces for chat messages, sessions, and responses. types.ts defines the AttachedFile interface for file handling. utils.ts adds a utility function for merging Tailwind CSS classes.
src/lib · high confidence
Introduce localGPT RAG application shell
The application now includes a root layout and global styling for the 'localGPT' interface, featuring a dark theme with black backgrounds and white text, along with the Geist font family. The main page renders a Demo component, establishing the visual foundation for the local RAG assistant.
src/app · high confidence
Introduce multimodal RAG indexing and session-based chat interface
The application now supports creating and managing document indexes with multimodal file uploads (PDF, DOCX, DOC, HTML, MD, TXT) and configurable chunking strategies (token-based, late-chunk, and Docling high-recall). Users can select embedding and LLM models, adjust retrieval modes (hybrid, vector-only, FTS-only), and enable contextual enrichment. A new landing menu allows users to create new indexes, chat with existing indexes, or start a quick LLM chat. The interface includes an index picker for selecting and managing existing indexes, a session sidebar for chat history, and improved markdown rendering that correctly preserves line breaks in model responses.
src/components · high confidence
Introduce official Docker deployment and Watson X cloud LLM support
Users can now deploy the entire localGPT stack (frontend, backend, RAG API, and optional Ollama) via Docker Compose with a new \docker-compose.yml\, dedicated Dockerfiles, and a \docker.env\ configuration file, eliminating the need for manual local environment setup. Additionally, the system now supports switching the LLM backend from local Ollama to IBM watsonx.ai Granite models by setting \LLM\_BACKEND=watsonx\ and providing Watson X credentials, with full configuration examples and documentation provided in \env.example.watsonx\ and \WATSONX\_README.md\.
(repo-wide) · high confidence
New RAG ingestion pipeline with robust chunking and multi-format support
The ingestion module now includes a new MarkdownRecursiveChunker and DoclingChunker that split documents into token-sized chunks while preserving document boundaries and context, fixing a previous issue where long documents lost approximately half their content. The DocumentConverter has been expanded to support TXT, MD, DOCX, and HTML formats in addition to PDF, utilizing the docling library for conversion and implementing automatic OCR engine selection (supporting OcrMac, EasyOCR, RapidOCR, and Tesseract) with graceful fallbacks when backends are unavailable.
_rag\system/ingestion · high confidence
New agent architecture with optional full-document escalation and local NLI verification
The agent module has been restructured to introduce two new capabilities. First, a full-document escalation feature (roadmap 4.1) is available via the \EscalatingRetrievalPipeline\; when enabled via the \retrieval.document\_escalation.enabled\ flag, it appends the entire source document to the synthesis context if evidence scores remain low after the initial retry, capped by a token budget. Second, the \Verifier\ class now supports a local NLI model backend (roadmap 2.4) via the \VERIFIER\_MODEL\ environment variable, allowing users to switch from the default LLM-prompt verifier to a lightweight local model (e.g., MiniCheck) for groundedness checks, with specific model availability notes provided on failure.
_rag\system/agent · high confidence
New batch processing, logging, and LLM client utilities for RAG
This change introduces three new utility modules in rag\_system/utils. batch\_processor.py provides a generic BatchProcessor with progress tracking, error handling, and optional garbage collection for processing items in batches. logging\_utils.py adds structured logging and a helper to display retrieval results. ollama\_client.py implements an Ollama client with dynamic context-window sizing to prevent silent front-truncation, per-query token usage tracking via context variables, and detection of truncation issues. watsonx\_client.py adds a Watson X AI client with an interface compatible with OllamaClient, including support for temperature settings and image inputs, while reporting zero token counts to maintain consistency with the token tracker.
_rag\system/utils · high confidence
New chat UI components and settings modal
The \src/components/ui\ directory now includes a suite of new interface components: \QuickChat\ and \SessionChat\ for handling chat interactions, \ChatInput\ for message and file attachment, \ConversationPage\ for displaying messages with citations and thinking blocks, \ChatSettingsModal\ for configuring RAG and model options, and supporting UI primitives like \Button\, \Avatar\, \ScrollArea\, \GlassInput\, \GlassToggle\, \InfoTooltip\, \AccordionGroup\, and \ChatBubbleAvatar\. This represents a structural addition of the chat interface layer rather than a modification of existing behavior.
src/components/ui · high confidence
New indexing capabilities: contextual enrichment, cross-references, and late chunking
The indexing pipeline now includes several new features to improve retrieval quality. A new ContextualEnricher uses Ollama to prepend a short summary of surrounding text to each chunk, helping the model understand local context. A CrossRefExtractor identifies and resolves intra-corpus references (like "Exhibit B" or "Section 4.3") at index time, allowing queries to hop to related documents. A LateChunkEncoder generates embeddings by pooling token vectors from the entire document, reducing context loss for long texts. Additionally, an OverviewBuilder creates one-paragraph summaries for each document, stored as a vector sidecar for pre-filtering. The system also now tracks embedding model identity and vector normalization per table in LanceDB to prevent mismatches.
_rag\system/indexing · high confidence
New retrieval subsystem with deterministic filtering, query decomposition, and full-document reassembly
The \rag\_system/retrieval\ package introduces a new retrieval architecture. It adds a deterministic metadata filter DSL (\filters.py\) that compiles to safe, parameter-free SQL to prevent injection and enforce strict pre-filtering on all search legs. A \QueryDecomposer\ (\query\_transformer.py\) now handles multi-turn context resolution by incorporating the last assistant answer to resolve pronouns, while keeping single-turn prompts frozen for benchmark stability. The \MultiVectorRetriever\ (\retrievers.py\) supports hybrid, vector-only, and FTS-only search modes with reciprocal-rank fusion and experimental multi-vector sidecar support. Additionally, \document\_fetch.py\ enables full-document reassembly from indexed chunks in original order, capped by a token budget, to support deep-read escalation.
_rag\system/retrieval · high confidence
Phase 0 evaluation harness and baseline metrics added
The \eval/\ directory now contains a complete evaluation harness for measuring retrieval and generation quality against a gold set. This includes documentation (\BASELINE.md\, \DECISIONS.md\, \README.md\) recording the Phase 0 baseline metrics (using \Qwen/Qwen3-Embedding-0.6B\ and \BAAI/bge-reranker-v2-m3\), the decision log for Phase 1 component adoption, and the gold set generation script (\build\_goldset.py\) which reverse-generates queries from planted facts to ensure answer-bearing text relevance. The harness provides the measurement foundation for all subsequent Phase 1 and Phase 2 changes.
eval · high confidence
Behavioural changes
New modular RAG pipeline architecture with configurable chunking and retrieval
The system introduces a new modular pipeline structure in \rag\_system/pipelines\, separating indexing and retrieval logic into dedicated classes (\IndexingPipeline\ and \RetrievalPipeline\). Users benefit from a new default chunking strategy using Docling (token-based) for more accurate chunk sizing, with a fallback to legacy character-based chunking. The retrieval pipeline now supports hybrid search (dense + FTS) by default, includes a new Qwen3-Reranker-4B for final-stage selection, and adds support for DOCX and HTML file formats via Docling. Configuration is unified across \retrievers\ and \retrieval\ keys, and thread-safety improvements ensure stable concurrent query processing.
_rag\system/pipelines · high confidence
Fixes
Backend gateway introduces deterministic RAG routing and per-request context sizing
The backend server now uses a deterministic, regex-based gate to decide whether to send a message to the RAG API or answer directly via Ollama, replacing the previous enrichment-model router that incorrectly routed queries containing the word 'test' to the direct path. The gate prioritizes RAG for any message with linked indexes unless the input is smalltalk or a question about the assistant itself, and it respects a \force\_rag\ flag. Additionally, the Ollama client now calculates the \num\_ctx\ context window size per request based on conversation length, preventing silent front-truncation of older messages. The backend also includes a new SQLite database layer with automatic path detection for Docker vs. local environments and a test suite verifying the routing logic.
backend · high confidence
Fix doubled whitespace in rendered markdown answers
A new text normalization utility has been added to clean up excessive whitespace in model-generated markdown responses. This ensures that prose sections are rendered with consistent spacing (collapsing multiple blank lines and horizontal space runs) while preserving the exact indentation and formatting inside fenced code blocks, preventing layout issues in both static and streaming answers.
src/utils · high confidence
Dependencies
Initial dependency manifests for multimodal RAG system
This change introduces the dependency configuration files for the new multimodal RAG project. The backend Python environment (requirements.txt) specifies core libraries including PyMuPDF, Pillow, transformers (v4.51.0), torch (v2.4.1), lancedb, and docling, with optional support for IBM Watsonx AI and Anthropic API evaluation. The Docker environment (requirements-docker.txt) mirrors these dependencies but excludes macOS-specific tools like ocrmac in favor of docling's bundled EasyOCR backend for Linux compatibility. The frontend package.json establishes a Next.js 15.3.3 application using React 19, Tailwind CSS v4, and Radix UI components.
(dependencies) · high confidence
Housekeeping
Added placeholder file to public directory
A .gitkeep file was added to the public directory to ensure the folder is tracked by version control even when empty.
public · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 48.
Lenses
- Code Health 65
- Architecture 98
- Maturity 87
- Readiness 32
- Security 76
- Accessibility 50
Changes since last survey
- 300 commits — 259 feature/other, 41 fixes
By area
- (root) — 166 commits
- (repo) — 81 commits
- eval/decisions — 19 commits
- src/components — 5 commits
- Documentation/improvement_plan.md — 4 commits
- eval/corpora — 3 commits
- backend/database.py — 2 commits
- localGPTUI/localGPTUI.py — 2 commits
- localGPTUI/static — 2 commits
- localGPTUI/templates — 2 commits
- rag_system/ingestion — 2 commits
- .github/workflows — 1 commit
- Documentation/api_reference.md — 1 commit
- Documentation/images — 1 commit
- Documentation/research — 1 commit
- Documentation/research_roadmap.md — 1 commit
- SOURCE_DOCUMENTS/constitution.pdf — 1 commit
- backend/ollama_client.py — 1 commit
- gaudi_utils/embeddings.py — 1 commit
- gaudi_utils/pipeline.py — 1 commit
Notable commits
- fix: Fix #789: Update README with instructions for running the quantized Llama3 model
- fix: Fix AutoGPTQ version to 0.2.2 in requirements.txt. Adapt README.md.
- fix: Fix wrong --qa_save reference in README.MD
- fix: Fixed Excel file extension in constants.py
- fix: Merge branch 'main' into fix/database-path-auto-detection
- fix: Merge pull request #141 from Tchekda/fix/default-show-sources
- fix: Merge pull request #153 from PromtEngineer/issue-152-Bug-fix-in-show_sources-flag
- fix: Merge pull request #454 from KonradHoeffner/pr-fix-docker-llama-version
- fix: Merge pull request #649 from andrewyuau/fix-API-external-access
- fix: Merge pull request #847 from PromtEngineer/devin/1752559300-fix-token-based-chunking
- fix: Merge pull request #848 from PromtEngineer/devin/1752564985-fix-markdown-excessive-newlines
- fix: Merge pull request #853 from PromtEngineer/devin/1737063744-docker-setup-fix
- fix: Merge pull request #870 from PromtEngineer/fix/database-path-auto-detection
- fix: Merge pull request #871 from PromtEngineer/fix/lancedb-nan-handling
- fix: Rearchitect for doc/code parity; evidence-gated defaults; E2E fixes
- fix: There was an issue with Persistency, fixed it
- fix: eval: 4.1 escalation re-run post-truncation-fix — reject as default
- fix: eval: completeness-clause A/B (arm J) — reverted, does not fix hr target rows
- fix: eval: fix-set impact measured — authored +4, rfc 21->19 trade, hr_h05 recovered
- fix: eval: strict compose prompt tested and rejected (arm E) — revert, keep arm C
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
PromtEngineer/localGPT was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 19 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 34f3107f56a88203859092312231b945a711c619 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-13a154b7f5d1.