infiniflow/ragflow
44.3
Weak · 6 August 2026
688.7k
lines of production code
Go
with TypeScript, Python
3
measurements over time
What this system is
RAGFlow is a Retrieval-Augmented Generation platform that ingests, structures, and retrieves data from diverse sources including documents, web pages, and cloud storage. The system provides a modular architecture for document parsing, chunking, and embedding, supporting both Python and Go implementations for high-performance processing. It features a comprehensive agent framework with pluggable tools, sandboxed code execution, and workflow orchestration for complex reasoning tasks. Additionally, it offers a unified API and CLI for managing datasets, chat sessions, and administrative configurations.
How it got here
2023–2025 — RAG architecture and SDK expansion
75 changes.
This period focused on restructuring the RAG pipeline into a modular, flow-based execution engine while introducing a comprehensive Python SDK and CLI for programmatic interaction. The work also expanded data source connectors for GitHub, Bitbucket, and Google Drive, and introduced deep document parsing and memory management capabilities.
2026 — Go backend migration and test expansion
138 changes.
This period was dominated by the migration of the RAGFlow backend to Go, covering the CLI, server, and core service layers. Concurrently, the team significantly expanded test coverage across API endpoints, RAG components, and the new Go modules to ensure functional parity and reliability.
Features
Add Bitbucket connector for indexing pull requests
A new Bitbucket connector has been added to the data source layer, enabling the indexing of Bitbucket Cloud pull requests. The implementation includes a connector class that handles authentication, pagination, and checkpointing for resumable indexing, along with utility functions for API interaction and document mapping.
_common/data\source/bitbucket · high confidence
Add Extractor component for document processing
A new 'Extractor' component has been introduced to the RAG flow, enabling the system to process documents by generating a Table of Contents (TOC) or extracting specific fields from text chunks. This addition allows users to configure extraction workflows that can either parse text into structured chunks or generate a hierarchical TOC based on the document's content, with the output format configurable via the 'output\_format' parameter.
rag/flow/extractor · high confidence
Add Go common library with shared utilities and types
The internal/common package now provides a shared Go library containing utility functions and types used across the application. This includes HTTP response helpers (SuccessWithData, ErrorWithCode), a CodedError type for structured error handling, environment variable accessors, formatting utilities (ChunkID, FormatBytes, FormatNumber), JSON unmarshalling for flexible string arrays, and query rewrite prompt builders. The addition of these shared utilities supports the broader Go migration by providing a common layer for logging, error codes, and data formatting.
internal/common · high confidence
Add Go implementation for dataset navigation, query building, and reranking
The NLP service now includes a Go implementation of the dataset navigation system, which manages hierarchical document clusters and supports both in-memory and real search engine backends. A new QueryBuilder handles text preprocessing, stop-word removal, and field weighting for search queries. The Reranker component now supports external reranking models with score normalization to ensure consistent ranking across different providers. Additionally, a new RetrievalService orchestrates hybrid search, reranking, and pagination. These changes are accompanied by comprehensive unit and integration tests for the new components.
internal/service/nlp · high confidence
Add Go parsers for audio, CSV, and DOCX/DOCX-IR
The parser module now includes new Go implementations for handling audio files, CSV spreadsheets, and Microsoft Word documents. Audio files are validated against a whitelist of supported extensions (e.g., .mp3, .wav, .flac) and configured via VLM model IDs. CSV files are parsed into HTML tables, with support for chunking large sheets and cleaning illegal control characters. DOCX parsing is split into a CGO-backed entry point and a pure-Go intermediate representation (IR) that extracts text, headings, images, and tables. The DOCX parser supports both JSON and Markdown output formats, with post-processing to remove headers, footers, and table of contents entries. Tests are added for all new parsers.
internal/parser/parser · high confidence
Add Go-based storage engine implementations for MinIO, S3, OSS, GCS, and memory backends
The internal/storage package now includes Go implementations for multiple storage backends: MinIO, AWS S3, Aliyun OSS, Google Cloud Storage (GCS), and an in-memory store. Each implementation provides the standard storage interface (Put, Get, Remove, ObjExist, ListObjects, GetPresignedURL, BucketExists, RemoveBucket, Copy, Move, and Close). A storage factory (storage\_factory.go) initializes the active backend based on configuration, and a shared types.go defines the Storage interface and StorageType constants. Tests are included for the memory and MinIO backends.
internal/storage · high confidence
Add Google Drive connector for file synchronization
Introduces a new Google Drive connector that enables indexing and synchronization of files from Google Drive, including support for shared drives, personal drives, and folder crawling. The implementation handles file retrieval, permission syncing, and document conversion for various file types (Docs, Sheets, Slides, PDFs, etc.), allowing users to sync their Google Drive content for search and indexing purposes.
_common/data\_source/google\drive · high confidence
Add Infinity Go engine implementation for chunk, document, and metadata operations
The internal engine for Infinity has been implemented in Go, introducing new files for managing chunk storage, document indexing, metadata handling, and SQL query execution. The implementation includes a connection pool with configurable size and timeout settings, a filter translator for metadata queries, and support for skill indexes. Tests are included to verify the filter translation and SQL alias rewriting.
internal/engine/infinity · high confidence
Add Jira connector for syncing issues as Markdown documents
A new Jira connector has been introduced to sync Jira issues as Markdown documents. The implementation includes a main connector class that handles authentication (supporting API tokens, username/password, and scoped tokens), manages checkpointing for incremental updates, and formats issue data (descriptions, comments, attachments) into Markdown. Utility functions handle Atlassian Document Format (ADF) parsing, datetime normalization, and issue filtering based on labels.
_common/data\source/jira · high confidence
Add MCP client and server implementations with SSE and Streamable-HTTP transports
The MCP module now includes new client and server implementations that support both SSE and Streamable-HTTP transports. The client code demonstrates how to connect to the MCP server using either transport, while the server code introduces a \LaunchMode\ enum (\self-host\, \host\) and a \Transport\ enum (\sse\, \streamable-http\) to configure the server's behavior. The server also implements metadata caching for datasets and documents, and handles paginated responses from the backend API, including a fix for handling dict responses in the \list\_chats\ method.
mcp · high confidence
Add Nginx configuration files for HTTP, HTTPS, and backend routing
New Nginx configuration files have been added to the docker/nginx directory to support HTTP and HTTPS proxying. The setup includes a base nginx.conf, a shared proxy.conf, and specific server configurations (ragflow.conf.golang, ragflow.conf.hybrid, ragflow.conf.python, and ragflow.https.conf) that route API requests to the appropriate backend services (ports 9380, 9381, 9383, 9384) and serve static frontend assets.
docker/nginx · high confidence
Add OAuth2 and OIDC authentication support
The auth module now supports OAuth2 and OpenID Connect (OIDC) authentication for third-party identity providers. Users can configure OAuth2 or OIDC providers via a unified client interface, with automatic OIDC configuration discovery and secure JWT token validation. The implementation includes a GitHub OAuth client and a general OAuth2 client, both supporting synchronous and asynchronous user info fetching. The OIDC client enforces a strict allowlist of asymmetric signing algorithms (RS\, ES\, PS\*, EdDSA) to prevent algorithm-confusion attacks, and uses the provider's JWKS URI for key retrieval.
api/apps/auth · high confidence
Add SereneDB document-store engine
The system now supports SereneDB as a new document-store engine, enabling hybrid search (full-text and vector) and metadata management via a PostgreSQL-compatible interface. This adds a new backend for storing and retrieving chunks and metadata, with an integration test suite to verify functionality against a live SereneDB instance.
internal/engine/serenedb · high confidence
Add admin server status tracking
A new \admin\_status.go\ file introduces a global mechanism to track and query the availability of the admin server. The code defines an \AdminStatus\ struct and associated thread-safe functions (\InitAdminStatus\, \GetAdminStatus\, \SetAdminStatus\, \IsAdminAvailable\) that allow the application to check if the admin server is available (status 0) or unavailable (status 1) with a reason. This supports the 'admin server status checking' capability mentioned in the commit messages.
internal/server/local · medium confidence
Add document generation components for DOCX, HTML, Markdown, PDF, and TXT output formats
The agent now supports generating documents in five new formats: DOCX (using a custom, self-contained OOXML implementation), HTML, Markdown, PDF (backed by the signintech/gopdf library), and plain TXT. Each format supports optional headers, footers, watermarks, page numbers, and timestamps. The changes include the Go implementations for each writer, shared helper functions for configuration parsing, and comprehensive unit tests verifying the output structure and content for each format.
internal/agent/component/io · high confidence
Add knowledge-compile message queue and lease management via NATS
The internal/engine/nats package now includes a dedicated knowledge-compile stream, consumer, and KV bucket to handle knowledge compilation events. This adds a new message queue for knowledge compilation tasks, separate from the existing task queue, and introduces a lease mechanism to coordinate concurrent access to knowledge compilation resources.
internal/engine/nats · high confidence
Add skill source adapters for ClawHub, GitHub, and skills.sh
The CLI now supports fetching skills from multiple external registries. A new \source\ package introduces adapters for ClawHub, GitHub, and skills.sh, each implementing the \SkillSource\ interface to handle fetching, inspection, and trust-level evaluation. The \SourceResolver\ in \interface.go\ routes references (e.g., \clawhub://\, \github.com/...\, \skills.sh/...\) to the appropriate adapter. Users can now install skills directly from these hubs via the CLI.
_internal/cli/filesystem/skill\hub · high confidence
Add spaCy-based entity and relationship extraction for GraphRAG
A new spaCy-based extraction pipeline has been added to the GraphRAG module, enabling the system to identify entities and typed relationships from text without requiring LLM calls. This implementation supports seven languages (English, Chinese, German, French, Spanish, Portuguese, and Japanese) and provides both entity recognition and dependency-based relation extraction, including multi-hop inference and confidence scoring.
rag/graphrag/ner · high confidence
Added C++ tokenizer and C API bindings for Go integration
The \internal/binding/cpp\ directory now contains the complete C++ implementation for the RAG tokenizer, including the \CMakeLists.txt\ build configuration, a Makefile for building the C API static library, and the corresponding header and source files (e.g., \analyzer.h\, \darts\_trie.h\, \main.cpp\). This change introduces the C++ components and C API headers required for Go bindings, enabling the Go provider to interface with the C++ tokenizer logic.
internal/binding/cpp · high confidence
Added ClickHouse driver stub for model usage and operation logging
A new file, clickhouse\_ee.go, introduces a global Driver struct with placeholder implementations for CollectModelUsage, SaveOperationLog, and Status. This provides the underlying infrastructure for tracking model usage and saving operation logs, though the current implementation returns nil or empty values.
internal/engine/clickhouse · medium confidence
Added DOCX document parsing support
Users can now parse .docx files. The new parser extracts text, headings, tables, and images from DOCX documents, converting them into a structured format for further processing. This includes support for multi-section documents, nested headings, and table content.
internal/deepdoc/parser/docx · high confidence
Added GitHub connector implementation
The system now supports connecting to GitHub repositories. This change introduces a new connector implementation located in the \common/data\_source/github\ directory, including modules for handling rate limits, utility functions, and data models. This enables the platform to sync data from GitHub sources.
_common/data\source/github · high confidence
Added Go agent DSL parsing and normalization helpers
The internal/agent/dsl package now includes Go implementations for parsing and normalizing the agent canvas DSL. New files (errors.go, extract.go, normalize.go, reset.go) provide functions to extract component input forms, parameters, and names from the DSL map, as well as to normalize the DSL for both the front-end canvas (generating graph nodes/edges from components) and the runtime (folding legacy loop/iteration variants). A reset function clears per-run state (history, retrieval, memory, path) and resets global variables (sys.\* and env.\*) to their defaults, mirroring the Python backend's reset behavior. Test files (extract\_test.go, normalize\_test.go, reset\_test.go) and test fixtures (agent\_msg.json, all.json, dfx\_picture\_parser.json) are added to validate these transformations.
internal/agent/dsl · high confidence
Added Go binding for RAG text analyzer
A new Go binding for the RAG text analyzer has been added to the internal/binding package. The change introduces a full CGo implementation (rag\_analyzer.go) that wraps the native C++ tokenizer, exposing methods for tokenization, fine-grained tokenization, and position tracking. A corresponding stub implementation (rag\_analyzer\_nocgo.go) is also added to handle builds where CGo is disabled, ensuring the API remains consistent across different build configurations.
internal/binding · high confidence
Added Go client for the deepdoc vision service
Introduced a new Go client in internal/deepdoc that implements the DLA (Document Layout Analysis) HTTP contract, including retry logic, exponential backoff, and multipart form handling. The client also provides stubs for OCR and TSR, which return a clear error since those features are local ONNX pipelines with no remote endpoint.
internal/deepdoc · high confidence
Added RAG benchmarking and configuration modules
Introduced new files in the \rag\ package: \\_\init\\_.py\ to initialize the module, \settings.py\ for configuration, and \benchmark.py\ which provides a \Benchmark\ class for evaluating retrieval performance on datasets like MS MARCO, TriviaQA, and MIRACL. This adds the capability to run standardized retrieval benchmarks against a knowledge base.
rag · high confidence
Added Redis engine with strict token-bucket rate limiting
A new Redis engine module has been introduced to the codebase, providing a centralized client for Redis operations including initialization, health checks, and Lua-based scripts for atomic operations. The module implements a strict token-bucket rate limiter that fails closed on transport errors or when the client is uninitialized, ensuring that webhook triggers and other rate-limited components receive explicit errors rather than silent passes.
internal/engine/redis · medium confidence
Added build script for RAGFlow CLI release
A new shell script, build\_cli\_release.sh, has been added to the admin directory. This script automates the build process for the RAGFlow CLI tool, handling the preparation of source code, creation of a release directory structure, and execution of the Python build process to generate the final package.
admin · high confidence
Added chat widget demo pages
Added two new HTML files in the chat\_demo example folder: index.html and widget\_demo.html. These files provide a demonstration of the floating chat widget, including the necessary iframe embedding and JavaScript message handling to control the chat window visibility and scroll behavior.
_example/chat\demo · high confidence
Added empty memory module
An empty \_\init\\_.py file was added to the memory directory, creating a new Python package structure for that location.
memory · high confidence
Added loop and parallel workflow node implementations
The agent workflow engine now supports iterative and parallel execution patterns. A new 'Loop' node allows workflows to repeatedly execute a sub-workflow until a condition is met, with configurable max iterations, stream modes, and checkpointing. A new 'Parallel' node enables concurrent execution of a sub-workflow across a slice of inputs, with configurable concurrency limits, per-item checkpointing, and interrupt/resume support. Both features are implemented in the internal/agent/workflowx package and include comprehensive integration and unit tests.
internal/agent/workflowx · high confidence
Added resume entity parsing modules for companies, degrees, industries, regions, and schools
The resume parser now includes new entity resolution modules for companies, degrees, industries, regions, and schools. These modules load structured data (CSV, JSON) and provide functions to normalize company names, map degree codes to names, look up industry classifications, validate Chinese regions, and match school names against a known list. This enables the resume parser to extract and standardize these specific fields from resumes.
deepdoc/parser/resume/entities · high confidence
Agent component framework and core nodes introduced
The agent execution engine has been refactored into a modular component-based architecture. A new \agent/component\ package provides a dynamic plugin system that automatically discovers and registers all node types. This update introduces the foundational \ComponentBase\ and \ComponentParamBase\ classes, alongside the \Begin\ node for workflow entry points. The \Agent\ node is now implemented as a tool-calling component that manages LLM interactions, tool execution, and structured output formatting. Additionally, new components for data operations, document generation (PDF/DOCX/Markdown), Excel processing, browser control, and categorization are included, enabling complex agent workflows with improved modularity and error handling.
agent/component · high confidence
Agent tools restructured into a pluggable, auto-discovered module
The \agent/tools\ directory has been refactored to use a dynamic import and class-extraction mechanism (\\_\init\\_.py\), allowing new tools to be added simply by creating a new file in the directory. Each tool now inherits from a new \ToolBase\ and \ToolParamBase\ base classes, standardizing how tools are invoked, timed out, and report errors. This change enables the agent to automatically discover and register tools like AkShare, ArXiv, BGPT, Crawler, DeepL, DuckDuckGo, Email, and ExeSQL without manual registration, while also introducing async support and consistent error handling across all agent tools.
agent/tools · high confidence
Canvas debug (dry-run) mode with View result log
Added a canvas debug (dry-run) mode that allows users to preview the results of a pipeline without persisting data. This includes a debug log sink that records component progress and errors to Redis for the front-end 'View result' page, and a debug result DSL builder that formats component outputs for the UI. The debug context is identified by an empty knowledge base ID, which signals the pipeline to skip the persist stage and embedding. Tests verify the debug log sink records trace and end markers, and that the debug page cap is correctly injected into the parser configuration.
internal/ingestion/task · high confidence
Centralized model and configuration management
The \conf/\ directory has been reorganized to support a new, unified model catalog and configuration structure. A new \all\_models.json\ file provides a global registry of supported LLMs, embedding models, and other AI capabilities, each defined with specific attributes like \content\_length\, \max\_output\, and \model\_types\. This is complemented by \llm\_factories.json\, which details the available model providers and their specific model endpoints. Additionally, the directory now includes dedicated mapping configurations for various data stores (Elasticsearch, Infinity, OpenSearch) and system settings, standardizing how the application interacts with different backends and manages internal state.
conf · high confidence
Centralized storage and utility modules for RAG components
The \rag/utils\ package has been reorganized into a set of dedicated modules for storage backends and shared utilities. New implementations have been added for Azure Blob Storage (both SAS and Service Principal Name authentication), Google Cloud Storage (GCS), Elasticsearch, Infinity, and MinIO/S3, each providing a standardized interface for object storage operations. Additionally, the directory now includes utilities for handling encrypted storage, lazy image loading, base64 image processing, and file extraction (PDF, DOCX, OLE/ZIP). These changes consolidate storage logic and provide a unified interface for document and metadata management across the RAG system.
rag/utils · high confidence
Docker deployment configuration and scripts
The docker directory now includes a comprehensive set of configuration files and scripts to support containerized deployment. A new .env file provides environment variables for all supported vector databases (Elasticsearch, OpenSearch, Infinity, SereneDB, OceanBase, SeekDB) and services (MySQL, MinIO, Redis, ClickHouse, Jaeger). The docker-compose files (base, main, macOS, and OpenCloudOS variants) define the service graph, exposing ports for HTTP, HTTPS, MCP, and admin servers. An entrypoint script handles service startup, configuration templating, and argument parsing for disabling components or enabling the MCP/Admin servers. A migration script allows users to backup and restore Docker volumes for all supported databases. Additionally, an example .env file demonstrates how to configure MinIO for single-bucket mode.
docker · high confidence
Expanded data source connector support
The system now supports a significantly larger set of data sources, including Airtable, Asana, Azure Blob Storage, BigQuery, Box, DingTalk AI Table, and others, each with their own connector implementation. This allows users to sync content from these platforms directly into the knowledge base.
_common/data\source · high confidence
Extracted service layer for API endpoints
The \api/apps/services/\ directory now contains a new set of service modules that encapsulate the business logic for various API endpoints. These include \canvas\_replica\_service.py\ for managing canvas runtime state in Redis, \dataset\_api\_service.py\ for dataset creation and deletion, \document\_api\_service.py\ for document updates and status changes, \file\_api\_service.py\ for file and folder management, \memory\_api\_service.py\ for memory configuration and access control, \models\_api\_service.py\ for tenant model resolution and validation, \provider\_api\_service.py\ for managing LLM providers and instances, and \structure\_graph\_common.py\ for graph rendering logic. These services provide the implementation for the corresponding API routes, handling data validation, permission checks, and interactions with the database and external services.
api/apps/services · high confidence
File service implementation in Go
The file service is now implemented in Go, introducing new files for file management including upload, download, delete, and folder operations. The implementation includes a commit system for tracking file changes, content retrieval, and URL-based uploads with SSRF protection. Tests are added for file permissions, content handling, and URL sanitization.
internal/service/document, internal/service/file · high confidence
Go API server: new agent routes and route registration tests
The Go API server now registers the full set of agent canvas endpoints (CRUD, versions, sessions, logs, webhook triggers, and chat completions) via a new \agent\_routes.go\ file, replacing the previous scattered registration. A corresponding test file (\agent\_routes\_test.go\) verifies that all 11 agent routes are correctly wired and that nil-safety is handled. Additional tests confirm that the dataset search, update, searchbot mindmap, chatbot info, and chat mindmap routes are registered in the main router setup.
internal/router · high confidence
Go agent canvas introduces Redis-backed checkpointing, session cancellation, and compile-time parameter overrides
The Go agent canvas now supports Redis-backed checkpointing via a new \RedisCheckPointStore\ that serializes eino workflow states for resume and recovery. A \cancel.go\ module enables cross-process session cancellation by polling a Redis key, allowing workflows to be aborted cleanly. Additionally, the \Compile\ entry point now accepts \CompileOptions\ to inject run-level parameter overrides, enabling pipelines to hot-swap component configurations without mutating shared DSLs. These changes are accompanied by comprehensive unit tests for the checkpoint store, cancellation logic, and compile-time behaviors.
internal/agent/canvas · high confidence
Go chunker implementation with delimiter and title-chunking fixes
The Go-based chunker component is introduced, providing the core text-chunking logic for the ingestion pipeline. This includes the GroupTitleChunker, HierarchyTitleChunker, and TokenChunker variants, each mirroring the corresponding Python implementations. Key fixes ensure that custom delimiters (e.g., backtick-wrapped newlines) are correctly dropped from chunk text across all input formats (text, markdown, HTML, JSON). Additionally, title-based chunking now correctly handles heading hierarchies and section grouping, while image handling supports on-demand PDF cropping and MinIO upload. The implementation includes comprehensive unit tests to verify delimiter case-sensitivity, heading extraction, and metadata preservation.
internal/ingestion/component/chunker · high confidence
Go chunking pipeline gains configurable split strategies and post-processing filters
The Go implementation of the chunking pipeline has been expanded to support multiple split strategies (sentence, character, paragraph, and length-based) and a post-processing stage that can merge chunks to a target size and filter them by minimum length, empty content, or duplicates. A new canonical delimiter parser (mirroring the Python implementation) allows users to define custom multi-character and single-character delimiters, with backtick-wrapped tokens treated as multi-character delimiters. The pipeline also includes a language detection helper (identifying Chinese vs English text) and a preprocessing stage for newline normalization, whitespace stripping, and empty line removal. These changes provide a more flexible and consistent chunking behavior across Go and Python.
internal/parser/chunk · high confidence
Go entity models for conversations, chat, and evaluation
The \internal/entity\ package now includes Go structs that map to the application's database tables, covering conversations (\api\_for\_conversation.go\), chat sessions (\chat.go\), evaluation datasets and runs (\evaluation.go\), and more. These models define the data structures used by the Go backend to interact with the database, ensuring parity with the existing Python implementation.
internal/entity · high confidence
Go implementation of NER and relation extraction for the ingestion pipeline
The ingestion pipeline now includes a Go-based NER and relation extraction system that mirrors the existing Python implementation. This adds a new \extractor\ package under \internal/ingestion/compilation\ containing \ner.go\ (wrapping the C++ ThincNER engine via cgo), \dep\_relation.go\ (dependency-based relation extraction), \ner\_relation.go\ (regex-based relation extraction), and \parser\_go.go\ (C++ parser/tagger wrappers). A \no\_cgo.go\ stub ensures the package compiles without CGO, while \ner\_test.go\ provides unit tests for the new functionality. The \wire.go\ file centralizes component registration for the ingestion pipeline.
internal/ingestion/compilation · high confidence
Go implementation of dataset service with comprehensive test coverage
The dataset service logic in internal/service/dataset has been implemented in Go, providing the backend for dataset creation, listing, and metadata management. This includes validation for parser IDs, pipeline IDs, and embedding models, as well as handling dataset name deduplication and permission checks. The change adds extensive unit tests for the create, CRUD, index, and metadata operations, ensuring parity with the existing Python implementation.
internal/service/dataset · high confidence
Go implementation of the sandbox execution providers
The sandbox package now includes Go implementations for the Aliyun, E2B, and Local sandbox providers, each handling code execution in their respective environments. The Aliyun provider uses the Alibaba Cloud agentrun SDK and a raw REST API for code execution, while the E2B provider utilizes a community Go SDK to manage cloud sandboxes. The Local provider executes code directly on the host via subprocesses. A shared allowlist of artifact file extensions (e.g., .csv, .pdf, .png) is maintained for both the Local and SSH providers to ensure consistent artifact handling. Configuration for these providers is managed through environment variables and admin-panel settings, with a centralized manager coordinating the active provider.
internal/agent/sandbox · high confidence
Go ingestion pipeline schema types for parser, chunker, tokenizer, extractor, and file components
The ingestion pipeline's Go implementation now includes formalized schema types for the wire-level payloads between components. This adds \ParserFromUpstream\ and \ParserOutputs\ for document parsing, \ChunkerFromUpstream\ and \ChunkerOutputs\ for chunking, \TokenizerFromUpstream\ and \TokenizerOutputs\ for tokenization, \ExtractorFromUpstream\ and \ExtractorOutputs\ for metadata/tag extraction, and \FileFromUpstream\ and \FileOutputs\ for file handling. Each type mirrors the corresponding Python Pydantic schemas or runtime contracts, ensuring consistent data shapes across the Go and Python ingestion pipelines.
internal/ingestion/component/schema · high confidence
GraphRAG checkpoint and phase-marker support for community extraction and entity resolution
Added checkpoint/resume support for GraphRAG community extraction and entity resolution, allowing long-running or cancelled tasks to resume from the last saved state rather than recomputing from scratch. The \checkpoints.py\ module introduces Redis-backed checkpoint storage (with configurable page size and TTL) for both community and resolution phases, while \phase\_markers.py\ adds Redis-based phase-completion markers to skip already-completed phases on re-runs. The \entity\resolution.py\ module is updated to accept and use these checkpoints, and the \\\init\\_.py\ file is created to expose the new modules. This change improves reliability and performance for large GraphRAG workloads by preventing redundant computation and enabling safe task cancellation and recovery.
rag/graphrag · high confidence
Improved Docker build performance and developer experience
The project now includes a .dockerignore file to reduce the Docker build context by excluding local artifacts, build outputs, and IDE files. Additionally, a .rooignore file has been added to reduce indexing noise and token waste for AI coding assistants. A new AGENTS.md file provides a comprehensive local operating guide for the codebase, including Go test tier classifications and validation preferences. The main Dockerfile has been updated to use a multi-stage build with a dedicated builder stage, improving cache efficiency and build speed.
(repo-wide) · high confidence
Introduce Agentic RAG orchestration layer
A new \rag/advanced\_rag/harness\ package provides the core components for an agentic research workflow. This includes a \Pipeline\ for unified tool execution and scoping, a \ResearchToolSession\ that routes native or text-based tool calls to the agent loop, and a \Planner\ that decomposes questions into atomic claims. The system supports four thinking modes (low, medium, high, ultra) with distinct strategies like decomposition, agentic loops, and deep research. It also introduces a sufficiency check to determine if collected evidence is enough to answer the user's question, and a routing node to classify query types and suggest knowledge compilation tools.
_rag/advanced\rag/harness · high confidence
Introduce DeepDoc module with table auto-rotation and multi-language documentation
The deepdoc module is introduced, providing documentation and a Python package entry point. The documentation (README.md and translations) details the module's capabilities, including OCR, layout recognition, and table structure recognition (TSR). A key feature is automatic table rotation, which detects and corrects the orientation of rotated tables in scanned PDFs to improve OCR accuracy. This feature is enabled by default and can be controlled via the environment variable TABLE\_AUTO\_ROTATE or the API parameter auto\_rotate\tables. The module also includes a Python \\init\\_.py file that initializes the package with beartype.
deepdoc · high confidence
Introduce Go runtime infrastructure for the agent canvas
Added the \internal/agent/runtime\ package to provide the Go runtime infrastructure for the agent canvas. This includes a component registry (\registry.go\) that supports case-insensitive lookups and category-based filtering, a \CanvasState\ struct to manage per-run shared state (outputs, history, retrieval, etc.) with JSON serialization, and helper functions for timeout enforcement, progress tracking, and safe JSON marshaling. The changes also introduce a \Component\ interface and a factory pattern for component instantiation, along with context-based state passing for message emitters and deferred node completion.
internal/agent/runtime · high confidence
Introduce Go service layer for agent execution, session management, and admin heartbeat reporting
The internal/service package now includes new Go implementations for agent execution (agent.go, agent\_sessions.go, agent\_dbcheck.go), session management, and admin client reporting (admin\_client.go, admin\_client\_ee.go). These additions provide the Go runtime for running agent canvases, managing agent sessions, validating database connections with SSRF guards, and sending heartbeat reports to the admin server. The changes also include comprehensive test coverage for agent cancellation, persistence, and end-to-end execution paths.
internal/service · high confidence
Introduce Go-based Agent Harness framework
The internal/harness directory now contains a new Go framework for building stateful, multi-agent applications. This includes a Pregel-based graph execution engine (graphengine) and a full Agent Development Kit (agentcore) that provides ReAct agents, middleware, workflow orchestration, checkpointing, and streaming support. The framework is organized into three layers: the graph engine, the agent core, and a push-based AgentLoop for chat and streaming applications.
internal/harness · high confidence
Introduce Go-based CLI and server entry points
Added new Go source files (ragflow-cli.go and ragflow\_server.go) that serve as the entry points for the RAGFlow CLI tool and the main server application. The CLI entry point handles argument parsing, logging initialization, and command execution, while the server entry point parses server mode flags (api, admin, ingestor, syncer) and config options, establishing the Go runtime's primary interface for user interaction and service orchestration.
cmd · high confidence
Introduce Go-based CLI with admin mode and benchmarking support
The internal CLI implementation is now available in Go, providing a fully compatible command-line interface that supports both user and administrator modes. Administrators can now manage users, roles, services, and system variables, as well as perform health checks on the server, object store, message queue, and cache. The CLI also includes a benchmarking framework that allows administrators to run performance tests with configurable concurrency and iterations, providing metrics like QPS and success/failure counts. This Go implementation mirrors the existing Python CLI's syntax and functionality, ensuring a consistent experience across both implementations.
internal/cli · high confidence
Introduce Go-based TTS synthesizer with Redis caching
The audio package now includes a new Go implementation for text-to-speech synthesis, replacing the previous placeholder shell-out approach. The \model\_provider\_synthesizer.go\ file wires the \audio.Synthesizer\ interface to the project's model provider service, allowing the Go agent to dispatch TTS requests to HTTP-based providers (Fish, OpenAI, StepFun, Xinference, LiteLLM). The implementation includes a Redis-based caching layer that stores synthesized audio by a SHA-256 hash of the tenant ID, text, voice, and language, with a default 7-day TTL (configurable via \RAGFLOW\_TTS\_CACHE\_TTL\_SECONDS\). The \tts.go\ file defines the \Synthesizer\ interface and error types, while \tts\_dispatch.go\ handles the mapping of audio requests to the model provider's \AudioSpeech\ method. Tests in \model\_provider\_synthesizer\_test.go\, \tts\_dispatch\_test.go\, and \tts\_test.go\ verify the cache key generation, TTL behavior, and dispatcher logic.
internal/agent/audio · high confidence
Introduce Go-based admin server with role and user management
The internal/admin directory now contains a new Go implementation of the admin server, providing a Go-native backend for administrative operations. This includes a complete HTTP handler and router for user management (listing, creating, and deleting users), role management (CRUD for roles and permissions), and system settings (variables, log levels, and environments). The Go server exposes endpoints for managing services, ingestion tasks, and sandbox providers, while also supporting enterprise features like role-based access control, license management, and soft fingerprinting. A test file (service\_variables\_test.go) validates system setting value validation and data type inference.
internal/admin · high confidence
Introduce Go-based chat invoker and navigation service interfaces
Added new Go packages for the chat invoker and navigation service, providing the shared interfaces and singleton registration mechanisms (SetDefaultInvoker, SetNavService) that the agent tools and harness will use. This establishes the Go-side contracts for LLM chat invocation and dataset navigation tree management, replacing the previous Python-based implementations.
internal/agent/chat, internal/service/nav · high confidence
Introduce Go-based dataset-level knowledge compilation and deduplication
The \internal/ingestion/knowledge\_compile\ package now implements the dataset-level post-processing pipeline in Go. This adds a \Consumer\ that claims batches of completed or deleted documents from a MySQL-backed scheduler, performs LLM-assisted deduplication of compiled products, and writes the merged results back to the search engine. The implementation includes a global worker pool for concurrency control, a \Deduper\ interface with an LLM-backed \DecideBatch\ for cross-document merging, and a \Reader\/\Writer\ abstraction for interacting with the DocEngine. Tests verify the consumer's handling of out-of-order events and tombstones.
_internal/ingestion/knowledge\compile · high confidence
Introduce Go-based document and message queue engine architecture
The \internal/engine\ package now provides a Go implementation for document storage and message queuing, replacing or supplementing the previous Python-based approach. Users can configure the system to use Elasticsearch, Infinity, or SereneDB as the document engine, with Elasticsearch being the only fully functional backend at this time. Additionally, the system now supports NATS as the message queue engine, enabling asynchronous task processing and knowledge compilation workflows. The engine is initialized at startup based on configuration, and services like \ChunkService\ can interact with the document engine for search, indexing, and metadata operations.
internal/engine · high confidence
Introduce Go-based knowledge compiler with structured, wiki, tree, and mindmap variants
The knowledge compiler component is now implemented in Go, providing a new runtime component that dispatches to four compilation variants: structure, wiki, tree, and mindmap. The component accepts chunks and template configurations, resolves dependencies (LLM, embedding, tokenization), and returns compiled products merged into the upstream chunk stream. The implementation includes a shared common package with in-memory product storage, tokenization, and batch packing logic, alongside variant-specific runners. Tests and golden fixtures validate the new behavior.
_internal/ingestion/component/knowledge\compiler · high confidence
Introduce Google authentication and API utility modules
Added a new \google\_util\ package containing modules for Google OAuth flow (\oauth\_flow.py\), credential management (\auth.py\), API service builders (\resource.py\), and paginated retrieval utilities (\util.py\). The \auth.py\ module now supports both OAuth2 and Service Account authentication methods, handling token refresh and sanitization. The \oauth\_flow.py\ module implements a local-server OAuth flow with configurable timeouts and scope overrides. The \resource.py\ module provides factory functions to build Google API service clients (Drive, Docs, Admin, Gmail) with automatic token refresh on \RefreshError\. The \util.py\ module adds helpers for paginated API calls and credential loading. These changes enable more robust and flexible Google integration for data sources.
_common/data\_source/google\util · high confidence
Introduce GraphRAG community report extraction and entity embedding capabilities
The \rag/graphrag/general\ directory now contains the core implementation for GraphRAG community report generation and entity embedding. This includes the \CommunityReportsExtractor\ which generates structured community reports from graph data, the \community\_report\_prompt\ defining the LLM instructions for this task, and the \entity\_embedding\ module which uses Node2Vec to generate vector embeddings for graph nodes. These additions enable the system to perform community-level analysis and vector-based entity representation within the GraphRAG pipeline.
rag/graphrag/general · high confidence
Introduce OSS DeepDoc HTTP API service for DLA, OCR, and TSR
Adds a new Python-based HTTP server (deepdoc/server) that serves Document Layout Analysis (DLA), Optical Character Recognition (OCR), and Table Structure Recognition (TSR) models via a unified API using LitServe and OSS ONNX Runtime models. The server exposes endpoints at /predict/dla, /predict/ocr, and /predict/tsr, with a health check at /health and model metadata at /model. It includes adapter classes for each model type, endpoint handlers, and a Dockerfile for containerized deployment on port 9390.
deepdoc/server · high confidence
Introduce Python SDK for RAGFlow API
The Python SDK for RAGFlow is now available, providing a structured client for interacting with the RAGFlow API. The SDK exposes core resources including Datasets, Chats, Sessions, Documents, Chunks, Agents, and Memory, allowing users to programmatically manage these entities through a consistent interface.
_sdk/python/ragflow\sdk · high confidence
Introduce RAGFlow Admin CLI and HTTP client
The RAGFlow Admin Service now includes a dedicated command-line interface (CLI) and Python HTTP client library, enabling users to manage system resources, users, and configurations via the command line. The CLI supports commands for service management (listing, showing, starting, stopping, and restarting services), user management (creating, dropping, and altering users), and model provider management (creating and dropping providers). The implementation adds new files including \http\_client.py\ for HTTP communication, \parser.py\ for command parsing, \ragflow\_cli.py\ as the main CLI entry point, and \user.py\ for authentication logic, all documented in \README.md\ and \command.md\.
admin/client · high confidence
Introduce RESTful API endpoints for chat, agents, and templates
The RESTful API layer is expanded with new endpoints for chat completions, agent bot interactions, and compilation template management. A new \\_generation\_params.py\ module standardizes LLM generation parameters (temperature, top\_p, etc.) and resolves them against user settings. New files \agent\_api.py\, \aimlapi\_api.py\, \bot\_api.py\, \chat\_api.py\, \chat\_channel\_api.py\, \chunk\_api.py\, \compilation\_template\_api.py\, and \compilation\_template\_group\_api.py\ provide the corresponding route handlers. This includes chat channel management, agent completion flows, and CRUD operations for compilation templates and their groups.
_api/apps/restful\apis · high confidence
Introduce agentic RAG with deep-research and reasoning streams
The rag/advanced\_rag module now includes a new agentic RAG pipeline that supports deep-research capabilities, including tree-structured query decomposition, multi-step reasoning, and streaming of internal thought processes to the user. This adds a new retrieval mode that can perform iterative web and knowledge base searches, with the ability to stream intermediate reasoning steps (e.g., 'thinking' tags) to the front end. The implementation includes a LangGraph-based orchestrator, a dedicated logging handler for reasoning logs, and a new 'DeepResearcher' class for deep research workflows.
_rag/advanced\rag · high confidence
Introduce built-in ingestion pipeline templates and component parameter schema
The ingestion pipeline now supports built-in pipeline templates managed via a new registry system. A new \builtin\_registry.go\ file introduces a \Registry\ that loads JSON templates from the \template/\ directory, providing a \List\ and \Get\ API for built-in pipelines. The system includes an alias mechanism (e.g., 'naive' maps to 'general') to maintain backward compatibility with existing dataset rows. Additionally, \component\_params.go\ adds a schema for extracting and validating component parameters from the DSL, excluding internal 'outputs' keys. Tests confirm that the registry correctly loads templates, resolves aliases, and that component parameters are correctly extracted and validated.
internal/ingestion/pipeline · high confidence
Introduce initial Helm chart for RAGFlow deployment
Added a new Helm chart for deploying RAGFlow on Kubernetes, including Chart.yaml, values.yaml, and README. The chart supports configuring the document engine (Infinity, Elasticsearch, or OpenSearch), external services (MySQL, MinIO, Redis), and exposes the admin service. It also allows configuring LLM factories and service settings via the Helm values.
helm · high confidence
Introduce memory services for message and query handling
Added new Python modules in the memory/services directory to manage message storage and retrieval. The new MessageService class provides methods to create, update, delete, and list messages in an Elasticsearch index, including support for filtering by agent ID and keyword search. Additionally, a MsgTextQuery class was introduced to handle text-based search queries, incorporating term weighting, synonym expansion, and special character handling to improve search relevance.
memory/services · high confidence
Introduce modular RAG orchestrator with multiple execution strategies
The RAG harness now features a new orchestrator loop that dispatches to different search strategies based on the selected thinking mode. Low modes use a direct single-pass search, medium modes employ a decompose-and-search approach with parallel search and sufficiency checks, and high/ultra modes utilize an agentic research loop that iteratively verifies claims, performs dynamic claim expansion, and runs sufficiency checks. This modular structure allows the system to adapt its depth and complexity based on the user's selected mode.
_rag/advanced\rag/harness/orchestrator · high confidence
Introduce new RAG flow execution engine
A new RAG flow execution engine has been added, providing a structured way to define and run data processing pipelines. The \rag/flow\ module introduces a \Pipeline\ class that orchestrates the execution of components (such as \File\ processing) within a graph-based workflow. This includes a base \ProcessBase\ class for handling asynchronous execution, timeouts, and logging, along with a dynamic module loader that automatically registers component classes. Users can now define and execute complex RAG workflows with built-in progress tracking and error handling.
rag/flow · high confidence
Introduce new Tokenizer component for document parsing
A new Tokenizer component has been added to the RAG flow, introducing a dedicated module for handling document tokenization and embedding. The implementation includes a Pydantic schema (TokenizerFromUpstream) to validate upstream data formats (JSON, Markdown, Text, HTML, or Chunks) and a Tokenizer class that processes these inputs. The component supports full-text search by generating title and content tokens, and it handles embedding generation for chunks, ensuring that empty chunk lists are processed correctly without errors. This change establishes the foundational structure for the tokenizer within the rag/flow directory.
rag/flow/tokenizer · high confidence
Introduce pluggable sandbox provider architecture
The agent/sandbox module now supports a pluggable architecture for code execution, allowing the system to dynamically select and manage different sandbox providers (e.g., self-managed, Aliyun, E2B) via a central ProviderManager. This change introduces a client interface that loads the active provider from system settings, enabling flexible integration with various execution environments while maintaining a unified API for agent components.
agent/sandbox · high confidence
Introduce plugin system for LLM tools
Added a new plugin architecture in the agent/plugin directory, enabling the loading and execution of LLM tools. This includes the core plugin manager, the LLM tool base class, and a sample 'bad\_calculator' plugin for demonstration. Documentation in English, Chinese, and Turkish explains how to create and integrate custom LLM tools.
agent/plugin · high confidence
Introduce structured output formats for knowledge compilation
The \rag/flow/compiler\ module now supports multiple structured output formats—JSON, Markdown, text, and HTML—via the \CompilerFromUpstream\ schema. This allows the knowledge compilation process to return results in the desired format, enabling downstream consumers to consume compiled knowledge in a standardized, parseable structure rather than relying on a single text-based output.
rag/flow/compiler · high confidence
Introduce unified document store abstraction with new backend connectors
The codebase now includes a new \common/doc\_store\ package that defines a \DocStoreConnection\ base class and concrete implementations for Elasticsearch, Infinity, and OceanBase. This change introduces a standardized interface for document storage, allowing the application to interact with different underlying databases through a consistent API. The new files include connection pool managers and base classes that handle database-specific logic, such as retrying on metadata contention for Infinity or managing full-text search templates for OceanBase.
_common/doc\store · high confidence
Introduced new memory message and tenant model services
Added new Python modules for memory message handling and tenant model management. The \memory\_message\_service.py\ introduces functions to save and extract memory messages, including LLM-based extraction and embedding. The \tenant\_model\_service.py\ provides utilities for managing tenant-specific model configurations, including OCR providers (MinerU, PaddleOCR) and default model resolution. The \user\_account\_service.py\ includes functions for creating new users and deleting user data, including associated tenant and memory data.
_api/db/joint\services · high confidence
Introduces new database service layer for managing user-generated content and system configuration
The codebase now includes a new \api/db/services\ package that provides database access layers for various domain entities. This includes \CommonService\ as a base class with standard CRUD operations and retry logic, alongside specific services for managing user-generated content such as \UserCanvasService\ (for agents/templates), \ChatChannelService\ (for external messaging integrations), \ChunkFeedbackService\ (for adjusting chunk recall weights based on user feedback), and \CompilationTemplateGroupService\ (for organizing compilation templates). These services handle data persistence, retrieval, and updates for these specific features.
api/db/services · high confidence
New CLI commands for managing skills and files
The CLI now supports managing skills and files through a new virtual filesystem interface. Users can install and uninstall skills using \install-skill\ and \uninstall-skill\ commands, which support multiple sources including local paths, GitHub, ClawHub, and skills.sh. The system includes a security scanner that analyzes skill content for threats like data exfiltration, prompt injection, and destructive operations, with a trust-based policy that blocks or requires confirmation for community-sourced skills. Additionally, the \ls\, \cat\, and \search\ commands now operate on datasets, files, and skills, providing a unified way to browse and manage resources.
internal/cli/filesystem · high confidence
New Go agent tools: Agentic Search, AkShare, ArXiv, BGPT, and Code Execution
The Go agent implementation adds five new tools to the canvas: AgenticSearchTool (hybrid/vector/bm25 retrieval), AkShare (East Money stock news), ArXiv academic search, BGPT scientific paper search, and CodeExec (Python/JS code execution). Each tool implements the Eino tool interface, exposing structured inputs and outputs to the agent. Tests cover parsing, URL building, and error handling for each tool.
internal/agent/tool · high confidence
New Go model drivers for 302.AI, Aliyun, Anthropic, Astraflow, and Avian
The Go model layer now includes new provider drivers for 302.AI, Aliyun (Tongyi-Qianwen), Anthropic (Claude), Astraflow (UCloud ModelVerse), and Avian. Each driver implements the ModelDriver interface, supporting chat, streaming, embeddings, and model listing as applicable. Tests verify request formatting, header handling, and response parsing for each provider.
internal/entity/models · high confidence
New Python SDK examples for chat, retrieval, and chunk management
Added four new Python SDK example scripts in the \example/sdk\ directory: \chat\_assistant\_example.py\ demonstrates creating a chat assistant, managing sessions, and performing standard and streaming chat; \chunk\_example.py\ shows how to manage document chunks (add, list, update, delete); \dataset\_example.py\ provides a CRUD example for datasets; and \retrieval\_example.py\ illustrates semantic and keyword-based retrieval workflows. These examples serve as practical references for integrating with the RAGFlow Python SDK.
example/sdk · high confidence
New Python SDK modules for Agents, Chats, Datasets, and Memory
The Python SDK now includes dedicated modules for managing Agents, Chat sessions, Datasets, Chunks, Documents, and Memory. Users can now create and manage Agent sessions, interact with Chat sessions, and perform CRUD operations on Datasets, Documents, and Chunks via the SDK. The SDK also provides methods to manage memory, including listing, forgetting, and updating messages. These changes introduce new capabilities for interacting with the backend APIs for these specific resources.
_sdk/python/ragflow\sdk/modules · high confidence
New and updated document parsers for EPUB, HTML, JSON, Markdown, and DOCX
The deepdoc/parser module now includes dedicated parsers for EPUB, HTML, JSON, Markdown, and DOCX formats, each exposing a class (e.g., RAGFlowEpubParser, RAGFlowHtmlParser, RAGFlowJsonParser, RAGFlowMarkdownParser, RAGFlowDocxParser) that can be imported via the package's \_\init\\_.py. The EPUB parser extracts XHTML content in spine order and delegates chunking to the HTML parser. The HTML parser strips styles/scripts/comments, handles hard line breaks, and splits tables. The JSON parser supports both standard JSON and JSONL formats, converting lists to dictionaries and splitting by size. The Markdown parser extracts and separates tables from fenced code blocks. The DOCX parser extracts text paragraphs and tables, handling image blobs and page breaks. These changes expand supported document types and improve parsing accuracy for these formats.
deepdoc/parser · high confidence
New chat channel runtime and integrations for WhatsApp, DingTalk, Discord, Feishu, LINE, QQ Bot, Telegram, and WeCom
The internal/channels package now provides a new chat-channel runtime that manages the lifecycle of multiple chat integrations. This adds support for WhatsApp, DingTalk, Discord, Feishu, LINE, QQ Bot, Telegram, and WeCom. Each integration is implemented as a channel that handles incoming messages and sends outgoing RAGFlow answers. The runtime reconciles active channels based on configuration, handles start failures with exponential backoff, and manages message queuing and deduplication per platform. Tests cover configuration parsing, message handling, and error conditions for each channel.
internal/channels · high confidence
New chat channels: DingTalk, Discord, Feishu, Line, QQ Bot, Telegram, and WeCom
Users can now connect external messaging bots to RAGFlow via new chat channels. The update introduces a unified channel runtime in api/channels that manages the lifecycle of connected bots, automatically starting, stopping, and restarting them as configurations change. Supported platforms include DingTalk, Discord, Feishu, Line, QQ Bot, Telegram, and WeCom. Each channel implementation handles inbound message routing to the appropriate RAG dialog and sends RAG-generated replies back to the respective platform.
api/channels · high confidence
New deepdoc vision module for document layout and table structure recognition
The deepdoc/vision directory now contains a complete, modularized vision processing pipeline. This includes a new \\_\init\\_.py\ that exports core classes like \OCR\, \Recognizer\, \LayoutRecognizer\, and \TableStructureRecognizer\. The \layout\_recognizer.py\ file implements the \LayoutRecognizer\ and \LayoutRecognizer4YOLOv10\ classes, which handle document layout analysis by downloading models from HuggingFace or using a remote DLA client. The \ocr.py\ file introduces a \TextRecognizer\ class and a \load\_model\ function that supports both CPU and GPU (CUDA) execution with configurable memory limits and thread counts. The \operators.py\ file provides image processing operators such as \DecodeImage\, \StandardizeImage\, and \NormalizeImage\. The \postprocess.py\ file implements post-processing logic for text detection and recognition. The \recognizer.py\ file provides the base \Recognizer\ class with sorting and layout cleanup utilities. Additionally, \seeit.py\ adds visualization capabilities for drawing bounding boxes and labels on images, while \t\_ocr.py\ and \t\_recognizer.py\ serve as test/entry-point scripts for running OCR and layout/table structure recognition tasks. This new structure replaces the previous monolithic implementation with a more organized, reusable, and extensible architecture.
deepdoc/vision · high confidence
New dependency download scripts for the RAGflow build environment
Added \download\_deps.py\ and \download\_go\_deps.py\ scripts, along with a \Dockerfile\ for the \ragflow\_deps\ directory. These scripts automate the download of external artifacts required for building the RAGflow image, including static libraries (pdfium, pdf\_oxide, office\_oxide), the stagehand-server binary, and various system dependencies. The scripts support a \--china-mirrors\ flag to use alternative download sources, ensuring the build process can proceed without network access during CI.
_ragflow\deps · high confidence
New knowledge compilation pipeline for structured and wiki-based document processing
The system now supports a new knowledge compilation workflow that processes documents into structured data and wiki-style pages. This includes a structure compilation pipeline that extracts entities, relations, and claims from documents, as well as a wiki compilation pipeline that generates and maintains wiki pages from document chunks. The implementation introduces several new modules: \structure.py\ for handling document-scoped structure compilation, \wiki.py\ for the MAP/REDUCE/PLAN/REFINE wiki pipeline, \raptor.py\ for RAPTOR-based tree clustering, \mind\_map\_extractor.py\ for mind map extraction, \dataset\_nav.py\ for dataset-level navigation clustering, and \runner.py\ for orchestrating the compilation process. These components work together to enable advanced RAG capabilities with structured knowledge extraction and wiki generation.
_rag/advanced\_rag/knowlege\compile · high confidence
New memory utility modules for search, highlighting, and prompt assembly
Added new utility modules in memory/utils to support memory search and display. This includes aggregation and highlighting helpers (aggregation\_utils.py, highlight\_utils.py) that enable field aggregation and text highlighting for search results, along with connection implementations for Elasticsearch (es\_conn.py), Infinity (infinity\_conn.py), and OceanBase (ob\_conn.py) that map message fields to storage-specific schemas. Additionally, msg\_util.py provides LLM response parsing, and prompt\_util.py assembles system and user prompts for memory extraction.
memory/utils · high confidence
New shell script examples for chat, chunk, retrieval, and dataset management
Added five new shell scripts in the example/http directory that demonstrate how to interact with the RAGFlow API via cURL. The examples cover creating and managing chat assistants and sessions, adding and updating document chunks, performing semantic and keyword-based retrieval, and listing or filtering datasets. These scripts provide ready-to-run references for developers integrating with the platform's core features.
example/http · high confidence
New specialized document parsers and audio transcription support
The RAG application now includes dedicated parsers for specific document types: audio (transcription via Sequence2Txt LLMs), books (handling DOCX, PDF, TXT, HTML, and legacy .doc with Tika), emails (parsing .eml headers and attachments), laws (structured heading and table extraction for legal documents), manuals (headings, tables, and images), and papers (extracting title, authors, abstract, and sections). A generic 'one' parser is also introduced to treat each file as a single chunk, supporting DOCX, PDF, Excel, TXT, Markdown, and HTML. These new modules in the rag/app directory expand the system's ability to ingest and structure diverse document formats.
rag/app · high confidence
New tools for WeChat integration and data migration
Added a ChatGPT-on-WeChat plugin that enables conversational interactions using RAGFlow's retrieval capabilities, and introduced a CLI tool to migrate RAGFlow data from Elasticsearch to OceanBase, including schema conversion, vector data mapping, and resume capability.
tools · high confidence
New utility functions for best-effort execution, captcha rendering, file type detection, and markdown-to-JSON conversion
The \internal/utility\ package introduces several new Go utilities. A \BestEffort\ helper allows non-fatal operations to run without panicking or propagating errors. A PNG-based captcha renderer (\captcha\_png.go\) replaces the previous SVG implementation to prevent answer leakage. File type detection (\file.go\) and conversion helpers (\convert.go\) provide robust type mapping and data transformation. An LRU cache (\embedding\_lru.go\) is added for embeddings. Markdown is converted to JSON (\markdown\_to\_json.go\) for mind-map extraction, and an MCP tool call implementation (\mcp\_call.go\) enables tool invocation via the streamable-HTTP transport. Tests are added for all new components.
internal/utility · high confidence
Port agent canvas, attachment, and debug endpoints to Go
The internal/handler package now includes the Go implementation for the agent canvas, attachment, and debug endpoints. This includes the \AgentHandler\ struct and its methods for listing agents, retrieving component input forms, debugging components, downloading and previewing agent attachments, and executing pipeline debug runs. The implementation mirrors the existing Python API, ensuring parity in response shapes and error codes. Additionally, unit tests have been added to verify the behavior of these endpoints, including edge cases like missing keys, permission errors, and Redis log parsing.
internal/handler · high confidence
Port agentic search logic to Go
The agentic search workflow has been ported from Python to Go, introducing new files in the \internal/agent/harness\ package. This includes the main execution graph (\agentic\_rag.go\), the answer generation logic (\answer.go\), the planner (\planner.go\), the orchestrator (\orchestrator.go\), and routing/navigation components (\route.go\, \datasetnav.go\, \navigation.go\). The Go implementation mirrors the Python logic, handling the full flow from routing and decomposition to search and answer formalization, with corresponding unit tests added for each component.
internal/agent/harness · high confidence
Port chat, session, and template data access to Go
The internal/dao package now includes new Go implementations for accessing conversation, chat, chat session, API token, and template data. This includes the API4ConversationDAO for managing conversation records and statistics, the ChatDAO and ChatSessionDAO for managing chat dialogs and sessions, the APITokenDAO for managing API tokens and beta keys, and the CanvasTemplateDAO and CompilationTemplateDAO for managing and resolving template groups. These changes provide the underlying data access layer for the Go-based agent, chat, and template management features, mirroring the existing Python implementations.
internal/dao · high confidence
Ported Elasticsearch engine implementation to Go
The Elasticsearch engine implementation has been ported from Python to Go, introducing a new Go-based backend for chunk, document, and metadata storage. This includes the core engine client, chunk handling, document indexing, and metadata filtering logic. The Go implementation is accompanied by a comprehensive suite of unit and integration tests to ensure functional parity and correctness.
internal/engine/elasticsearch · high confidence
Ported Go agent canvas components and runtime contract
Added the Go implementation of the RAGFlow agent canvas components, including the core Agent component (multi-turn ReAct agent), Begin, and Browser components, along with the shared Component interface and parameter validation. The Agent component delegates to the eino react agent, supporting tool calls, artifacts, and citation. The Begin component injects request inputs (query, user\_id, webhook\_payload) into the canvas state. The Browser component orchestrates web extraction via Stagehand. Comprehensive unit tests were added for these components, verifying input form generation, artifact collection, and state injection.
internal/agent/component · high confidence
Ported chunk service logic to Go
The chunk service implementation has been ported from Python to Go, introducing a new Go-based service for managing document chunks. This includes the core chunk service logic, vector fetching for search results, and comprehensive unit tests for the new Go implementation.
internal/service/chunk · high confidence
Python SDK: New example scripts and dependency lockfile
Added two new Python scripts to the SDK directory: hello\_ragflow.py, which prints the SDK version, and test.py, which demonstrates a streaming chat interaction with an agent. Additionally, a uv.lock file was introduced to manage Python dependencies, pinning packages such as lxml 6.1.0 and urllib3 \>= 2.7.0.
sdk/python · high confidence
RAG prompts restructured into modular, markdown-based templates
The \rag/prompts\ directory has been refactored to use individual \.md\ template files (e.g., \citation\_prompt.md\, \meta\_filter.md\, \resume\_basic\_info.md\) instead of inline string definitions. A new \generator.py\ module loads these templates using Jinja2, enabling more maintainable and structured prompt engineering for features like metadata filtering, resume parsing, and citation generation.
rag/prompts · high confidence
Unified search request and result types for the Go engine
A new types.go file introduces unified Go structs for search operations, including SearchRequest, SearchResult, SearchMetadataRequest, and related expression types (MatchTextExpr, MatchDenseExpr, FusionExpr). These types define the data structures used by the Go-based document engine to handle search requests, manage pagination, apply filters, and process metadata indices, providing a consistent interface for search and retrieval across different backends.
internal/engine/types · high confidence
Architecture
Admin server restructured with new modular components
The admin server has been restructured into a modular architecture, introducing dedicated files for authentication (auth.py), configuration (config.py), exception handling (exceptions.py), data models (models.py), response utilities (responses.py), role management (roles.py), route definitions (routes.py), and service logic (services.py). This change separates concerns within the admin server, making the codebase more maintainable and extensible for future admin features.
admin/server · high confidence
Centralized common utilities and shared components
The \common\ package has been established as the central repository for shared utilities, constants, and helper modules. This includes new modules for configuration loading (\config\_utils.py\), connection and timeout handling (\connection\_utils.py\), cryptographic operations (\crypto\_utils.py\), HTTP client wrappers (\http\_client.py\), and logging utilities (\log\_utils.py\). Existing functionality such as file utilities, decorators, and exception classes have been migrated into this directory to consolidate cross-cutting concerns and improve code organization.
common · high confidence
Consolidate and harden API utility modules
The \api/utils\ directory has been reorganized into a set of focused, standalone utility modules. This includes new files for handling HTTP request parsing and JSON serialization (\api\_utils.py\, \json\_encode.py\), file and content-type management (\file\_utils.py\, \file\_response.py\), and security hardening for deserialization (\configs.py\ with \RestrictedUnpickler\). Additionally, the codebase introduces dedicated utilities for health checks (\health\_utils.py\), email templates (\email\_templates.py\), and validation (\validation\_utils.py\), providing a more modular and robust foundation for the API layer.
api/utils · high confidence
Database layer refactored into modular components
The database layer in the api/db package has been restructured into distinct modules: constants and enums are now in \_\init\\_.py, the main ORM models and base classes are in db\_models.py, query helpers and bulk insert utilities are in db\_utils.py, and template normalization logic is in template\_utils.py. This separation improves code organization and maintainability of the database interaction layer.
api/db · high confidence
Introduce a new modular parser architecture for the ingestion pipeline
The ingestion pipeline's document parsing logic has been restructured into a new, modular parser framework located in the \rag/flow/parser\ directory. This change introduces a standardized \ParserParam\ configuration that explicitly defines supported output formats (JSON, Markdown, HTML, Text) for each document type (PDF, Docx, Spreadsheet, etc.). The update also adds dedicated modules for handling PDF chunk metadata (extracting and normalizing page positions) and provides utility functions for processing document headers, footers, and tables of contents. This refactoring provides a more consistent and extensible foundation for integrating various document parsers.
rag/flow/parser · high confidence
RAG server refactored with new task execution and synchronization architecture
The RAG server (rag/svr) has been refactored to introduce a new layered task execution architecture. A new task executor (task\_executor.py) manages background tasks, including data synchronization (sync\_data\_source.py) and file caching (cache\_file\_svr.py). The refactoring introduces concurrency limits via a new limiter module (task\_executor\_limiter.py) and adds a Discord bot integration (discord\_svr.py) for chat-based interactions.
rag/svr · high confidence
Refactor server configuration into modular, engine-specific config files
The internal/server/config package has been reorganized into separate files for each engine and service (admin, API server, database, cache, queue, storage, logging, etc.), each handling its own defaults and parsing logic. This modularizes the configuration system, making it easier to add or modify settings for specific components like MySQL, Redis, Elasticsearch, MinIO, and others without touching a monolithic config file.
internal/server/config · high confidence
Refactored task executor into a layered architecture with new chunking and post-processing modules
The task executor has been refactored into a layered architecture to improve modularity and testability. This change introduces several new modules within the \task\_executor\_refactor\ package: \chunk\_builder.py\ for document chunking and parser selection, \chunk\_post\_processor.py\ for keyword extraction, question generation, and metadata tagging, and \chunk\_service.py\ to orchestrate the chunking pipeline. Additionally, \dataflow\_service.py\ handles dataflow pipeline execution, \comparator.py\ provides logic to compare production and dry-run execution contexts, and \dataset\_skill\_generator.py\ and \dataset\_wiki\_generator.py\ handle the generation of skill trees and wiki artifacts respectively. These modules collectively replace the previous monolithic task handler, enabling better separation of concerns and supporting features like LLM-guided semantic rechunking and graph-based knowledge compilation.
_rag/svr/task\_executor\refactor · high confidence
Behavioural changes
Added backward compatibility routes for deprecated API endpoints
The API now includes a backward compatibility layer that maps deprecated endpoints to their new RESTful equivalents, ensuring existing clients continue to function. This includes redirects for chat and agent completions, dataset graph and index operations, file conversions, and system health checks. Each legacy route logs a deprecation warning and forwards the request to the corresponding new API path, such as /api/v1/chats/{chat\_id}/completions now routing to /api/v1/chat/completions.
api/apps · high confidence
Agent module refactored with new DSL migration and canvas execution logic
The agent module has been restructured to support a new internal DSL schema for canvas execution. A migration layer (agent/dsl\_migration.py) automatically rewrites legacy component names (e.g., 'Splitter' to 'TokenChunker') and node types to ensure compatibility with the updated graph structure. The core canvas execution logic (agent/canvas.py) has been refactored to use this normalized DSL, enabling more robust handling of component dependencies, variable references, and execution paths. This change ensures that existing agent configurations are seamlessly upgraded to the new format without manual intervention.
agent · high confidence
Centralized common utilities and exception handling
The api/common directory now contains shared Python modules, including base64 encoding, team permission checks for files and knowledgebases, and centralized admin exception classes. This consolidation ensures consistency across services and reduces code duplication.
api/common · high confidence
Complete Helm chart refactoring and new component support
The Helm chart has been restructured into individual template files for each component (e.g., \elasticsearch.yaml\, \opensearch.yaml\, \infinity.yaml\, \mysql.yaml\, \minio.yaml\, \redis.yaml\), each managing its own StatefulSet, Service, and PVC. A new \infinity\ document engine is now supported alongside \elasticsearch\ and \opensearch\. The \ragflow.yaml\ template now includes an optional admin service and API service, and mounts \service\_conf\ and \llm\_factories\ configuration files. Environment variables are centralized in \env.yaml\ as a Secret, with specific exclusions for sensitive keys to prevent duplicate YAML keys. Redis is now deployed as a StatefulSet with a headless service and PodDisruptionBudget. Helper templates in \\_helpers.tpl\ standardize image repository handling and name generation.
helm/templates · high confidence
Go server introduces centralized configuration and runtime variable management
The Go server now uses a centralized configuration system that reads from service\_conf.yaml and environment variables, supporting database, document engine, storage, cache, queue, and other engine types. It also introduces a runtime variable store for managing dynamic values like secret keys, alongside a new enterprise server structure for tracing and lifecycle management.
internal/server · medium confidence
Improved ingestion task lifecycle and state management
The ingestion service now enforces stricter task lifecycle management: tasks are properly acknowledged on completion or failure to prevent message redelivery, and the heartbeat mechanism ensures in-flight tasks are not redelivered mid-execution. Additionally, document state updates (metadata and counters) are handled with best-effort semantics, preserving existing metadata keys while merging new ones, and the system gracefully handles cancellation and transient failures without crashing the worker process.
internal/ingestion/service · high confidence
Introduce TokenChunker with configurable delimiters and output formats
Users can now use the new TokenChunker for text processing, which supports multiple output formats (JSON, Markdown, text, HTML, and chunks) and allows configuration of delimiters as soft boundaries for chunking. The chunker respects configured delimiters and handles various document types including tables and images, with options for overlapping and context windows.
rag/flow/chunker · high confidence
Introduce structured RAG tool system with phase-based selection
The advanced RAG harness now uses a modular tool system that registers search, navigation, exploration, and inspector capabilities, each with specific schemas and compilation requirements. A new gating mechanism filters and prioritizes tools based on the current search phase (locate, explore, verify, or cross-domain), ensuring the agent selects the most appropriate tool for each stage of the research process.
_rag/advanced\rag/harness/tools · high confidence
New TitleChunker with hierarchy and group chunking strategies
The TitleChunker has been refactored into a new modular implementation that supports two distinct chunking strategies: 'hierarchy' and 'group'. The 'hierarchy' strategy constructs a title-based tree to organize content, while the 'group' strategy groups text by title levels. This change introduces new files including \title\_chunker.py\, \common.py\, \group\_chunker.py\, \hierarchy\_chunker.py\, and \schema.py\ to handle these specific chunking behaviors.
_rag/flow/chunker/title\chunker · high confidence
Porting the Go ingestion pipeline to align with Python
The ingestion pipeline in Go has been refactored to align with the Python implementation. This includes fixing the parser parameters and image VLM system\_prompt handling, clamping the unconfigured embedding limit, and ensuring the parser backends and PDF pipeline match the Python version. Additionally, the extractor component has been ported to Go, including the tagger functionality, and the document storage resolution has been implemented to mirror the Python logic. The DOCX vision dispatch has also been added to enrich parse results with LLM-generated descriptions of embedded images, mirroring the Python's enhance\_media\_sections\_with\_vision function.
internal/ingestion/component · medium confidence
RAG/LLM module refactored to support new model providers and fix integration bugs
The LLM integration layer has been refactored to support a broader range of model providers, including new integrations for Xinference, Ollama, and various cloud-based LLMs. This change introduces a more robust and modular architecture for handling chat, embedding, and vision models, with improved error handling and configuration management. Key fixes include resolving issues with model-specific parameters, improving token counting accuracy, and ensuring compatibility with updated provider APIs.
rag/llm · high confidence
RAGFlow server startup and configuration refactored into the api module
The RAGFlow server startup logic, including the main entry point (ragflow\_server.py), has been moved into the api module. This change centralizes server initialization, configuration loading, and background thread management within the api package, replacing the previous startup mechanism. Additionally, the api module now includes a Python version validation check and automatic NLTK data download during startup.
api · high confidence
Refactored NLP module with unified delimiter parsing and Infinity integration
The \rag/nlp\ module was refactored to consolidate delimiter parsing logic into a canonical \delim.py\ module, which handles custom delimiter fields with consistent normalization, deduplication, and case-sensitive matching. The \query.py\ module was updated to strip special characters from queries to prevent parsing errors in Infinity, and to strip single quotes from synonym terms to avoid tokenization errors. Additionally, \rag\_tokenizer.py\ was modified to bypass tokenization when using the Infinity document engine, and \search.py\ was updated to prune stale chunks from the search results.
rag/nlp · high confidence
Resume parser refactored to normalize and enrich extracted data
The resume parser in \deepdoc/parser/resume\ has been refactored to standardize the structure of extracted resume data. The \\_\init\\_.py\ module now includes a \refactor\ function that cleans up raw data by removing obsolete fields (like \raw\_txt\, \inference\), normalizing lists of education and work experience into consistent dictionary formats, and mapping internal codes to human-readable values (e.g., gender, degree). The \step\_one.py\ module handles the initial extraction and type conversion of resume fields into a flat structure, while \step\_two.py\ performs deeper semantic enrichment, such as identifying school rankings, degrees, and industry names. This change ensures that downstream applications receive a consistent, normalized JSON structure for each parsed resume.
deepdoc/parser/resume · medium confidence
Fixes
Fix tokenizer failing silently when the cl100k BPE table is missing
The Go tokenizer now loads the cl100k BPE table from the local filesystem instead of relying on the \tiktoken-go\ library's default HTTP download. This prevents token counts from silently degrading to zero when the table is missing or inaccessible. The new \bpe\_loader.go\ implementation searches for the table in the working directory, the executable's directory, and configured cache directories, and reports a clear error if the file is not found. A new \usage.go\ module tracks per-run token usage (prompt, completion, and total tokens) via context, enabling accurate token counting for LLM calls. Tests in \bpe\_loader\_test.go\ and \tokenizer\_test.go\ verify the loader's behavior, including cache directory precedence, malformed file rejection, and token count consistency.
internal/tokenizer · high confidence
Fix: Prevent duplicate i18n languageChanged listeners
The web application now prevents the registration of duplicate i18n language change listeners, which resolves potential memory leaks or redundant processing when the user's language preference is updated.
web · high confidence
Fixes Docker build by preserving the bin directory
A .gitkeep file was added to the bin directory to ensure it is included in the Docker image, resolving a build issue where the directory was previously omitted.
bin · high confidence
Test coverage
Add Python SDK test infrastructure for model configuration; Add unit tests for admin service status filtering and configuration loading; Add unit tests for dataset SDK routes; Add unit tests for host address configuration and proxy scheme handling; Add web API test suite and common utilities; Added Playwright end-to-end test infrastructure; Added Playwright end-to-end tests for Next.js apps and model providers; Added Playwright end-to-end tests for authentication flows; Added RAGFlow HTTP benchmark CLI for chat and retrieval latency testing; Added RESTful API test suite for Go proxy compatibility; Added SDK API test suite for dataset, document, and chat assistant management; Added SDK API tests for chunk management within datasets; Added SDK dataset management tests; Added SDK session management tests; Added automated tests for HTTP API session management; Added automated tests for LLM list API endpoints; Added automated tests for chunk API endpoints; Added automated tests for the memory management API; Added comprehensive HTTP API tests for dataset management; Added comprehensive test coverage for the Web API document application; Added comprehensive unit tests for the refactored task executor; Added cross-domain smoke test for the component registry; Added integration tests for incremental wiki (artifacts) build and deletion; Added integration tests for the /api/v1/components endpoint; Added regression tests for RESTful API security and correctness fixes; Added test client for the agent module; Added test infrastructure for Go proxy and authentication; Added test suite for Stop Parse Documents HTTP API; Added test utilities for ingestion; Added tests for SDK message management operations; Added tests for TokenChunker delimiter and output-format behavior; Added tests for admin API user key management; Added tests for message app API endpoints; Added tests for table parser dataset chat functionality; Added tests for tenant model max\_tokens fallback logic; Added unit and contract tests for connector OAuth and Langfuse API endpoints; Added unit and integration tests for system app and API endpoints; Added unit and integration tests for the Search App API; Added unit tests for API utility functions; Added unit tests for Agent CRUD operations; Added unit tests for ChunkFeedbackService; Added unit tests for GraphRAG checkpoint, phase markers, and graph utilities; Added unit tests for MCP server pagination and dataset discovery; Added unit tests for OAuth and OIDC client implementations; Added unit tests for OceanBase database support and template normalization; Added unit tests for OceanBase memory aggregation and highlight utilities; Added unit tests for RAG parsing and NLP components; Added unit tests for RAG prompt generation and security; Added unit tests for RAG utility modules; Added unit tests for advanced RAG components; Added unit tests for agent canvas variable splitting, input reset, and DSL bridge round-trips; Added unit tests for agent components and tools; Added unit tests for agent webhook handling; Added unit tests for canvas app components; Added unit tests for data source connectors; Added unit tests for dataset access permissions, document metadata pagination, and model name parsing; Added unit tests for dataset, model, and document API services; Added unit tests for deepdoc parsers; Added unit tests for deepdoc vision operators and table column matching; Added unit tests for file and file-to-document API routes; Added unit tests for sandbox providers; Added unit tests for table parser column roles and metadata aggregation; Added unit tests for the CAJAL scientific paper agent template; Added unit tests for the Dify retrieval API; Added unit tests for the LLM tools API endpoint; Added unit tests for the MCP server application; Added unit tests for the RSS data source connector; Added unit tests for the check\_files tool; Added unit tests for the file API service; Added unit tests for user and tenant application logic; Expanded unit test coverage for common utilities and connectors; Improved reliability and safety for agent tools; Improved test reliability and coverage for RAGFlow unit tests; New HTTP API tests for chunk management; New test utility modules for assertions, file generation, and data generation; PDF parser: add integration tests and Redis caching for DeepDoc inference; Standardized HTTP API test infrastructure and error handling; Unit tests added for LLM provider integration and configuration.
Dependencies
Add dependency manifests for new Go, Python, and Node.js modules
Added new dependency configuration files to define the build and runtime requirements for several new or isolated components. This includes \pyproject.toml\ for the main Python application and the \ragflow-cli\ admin client, \go.mod\ and \go.sum\ for the Go API server, \requirements.txt\ for the agent sandbox executor, and \package.json\/\package-lock.json\ for the Node.js sandbox base image and the WhatsApp gateway. These files establish the specific library versions and constraints required for these modules to function correctly.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 44.
Lenses
- Code Health 79
- Architecture 56
- Maturity 72
- Readiness 36
- Security 48
- Domain Modelling 60
- Event-Driven 91
- Event Sourcing 61
- Accessibility 46
Changes since last survey
- 300 commits — 114 feature/other, 186 fixes
By area
- web/src — 75 commits
- internal/ingestion — 30 commits
- internal/entity — 29 commits
- api/apps — 17 commits
- internal/service — 15 commits
- rag/advanced_rag — 12 commits
- (root) — 10 commits
- internal/agent — 8 commits
- internal/channels — 8 commits
- internal/handler — 8 commits
- internal/dao — 6 commits
- api/db — 5 commits
- conf/models — 5 commits
- docs/guides — 5 commits
- rag/app — 5 commits
- rag/llm — 5 commits
- agent/sandbox — 4 commits
- agent/templates — 4 commits
- internal/server — 4 commits
- .github/workflows — 3 commits
Notable commits
- fix: fix(agent): make global variable form fields reactive to i18n language changes (#17883)
- fix: fix(document-preview): strip WPS DISPIMG formula from xlsx before preview (#17912)
- fix: Fix: applying model config in one multi-chat card no longer overrides sibling cards (#17866)
- fix: fix(dataset): echo back selected datasets not in the first page (#17616)
- fix: fix(list-filter-bar): correct nested filter to prune non-matching children (#17814)
- fix: CI:fix merge to main,cancel display fail (#17869)
- fix: Fix QA DOCX table parser dropping cells between repeated text (#17497)
- fix: Fix admin user list returning empty on first page (#17483)
- fix: Fix agent stop chat should not cancel the task (#17769)
- fix: Fix attachments not take effect in agentic chat (#17895)
- fix: Fix dataset selector to continue scoll if still have data (#17534)
- fix: Fix deepcopy the chat model (#17560)
- fix: Fix filter dataset by owner_ids not working (#17499)
- fix: Fix go generated token expired in python - 2 (#906) (#17875)
- fix: Fix internal/handler test failures, align response with Python contract, and re-enable handler/storage/agent tests in CI (#17554)
- fix: Fix knowledge compiler Go port regressions (MySQL 1101 + Start deadlock) (#17599)
- fix: Fix multiple model chat, the content overwrite the current chat session (#17566)
- fix: Fix overlapping CJK text lines in docx file preview (#17693)
- fix: Fix parse image in excel as table, the image shows as one column (#17569)
- fix: Fix parsing log display of infinity (#17479)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
infiniflow/ragflow was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 6 August 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 08867c1d73c3d8a81c97e45ba95f23d06ba28293 — the exact code this score is about.
- Scored under rubric-2026.08.19 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer latest.