deepset-ai/haystack
60.6
Adequate · 18 September 2026
47.8k
lines of production code
Python
primary language
1
measurement over time
What this system is
Haystack is an open-source Python framework for building production-ready AI applications, specifically focusing on Retrieval-Augmented Generation (RAG) pipelines and autonomous AI agents. It provides a modular component system that allows developers to orchestrate data ingestion, document storage, embedding, and large language model interactions into flexible, executable workflows. The system supports advanced agent capabilities, including tool use, state management, and human-in-the-loop confirmation strategies, while offering extensive integrations with various vector databases and LLM providers.
How it got here
2019–2023 — Haystack 3.0 architectural overhaul
55 changes.
This period centered on the Haystack 3.0 release, which introduced a major architectural rewrite including a unified Pipeline, native async support, and a modular component structure. Legacy FARM-based code was removed, and core functionality was reorganized into external integration packages with hardened security and serialization. The work also established a comprehensive testing infrastructure and modernized the build system to support these breaking changes.
2024–2025 — Agent framework and documentation overhaul
61 changes.
This period focused on introducing core agent capabilities, including the Agent component, State management, and SuperComponent wrappers, alongside new tools for evaluation, extraction, and query expansion. Concurrently, the project established a comprehensive, versioned documentation website using Docusaurus to support these new features and guide users through migration and advanced usage patterns.
2026 — Agent lifecycle management and documentation expansion
50 changes.
This period focused on establishing a robust lifecycle management system for AI Agents through a new Hooks module, introducing features like Human-in-the-Loop confirmation, context compaction, and token budgeting. Concurrently, the project expanded its documentation ecosystem to cover these advanced agent patterns, multi-agent architectures, and stable releases from version 2.22 through 3.1, while integrating fuzz testing for security.
Features
Add copy button to documentation pages
Documentation pages now include a sticky copy button that allows users to easily copy content. This is implemented by wrapping the original DocItem Content component with a new wrapper that injects a CopyDropdown component, while the Footer component remains unchanged.
docs-website/src/theme/DocItem · high confidence
Add versioned API reference documentation for v2.18
The v2.18 documentation site now includes a dedicated API reference index page. This page organizes technical documentation into three sections: the core Haystack framework API, official integration APIs, and experimental features, providing users with a structured entry point to the v2.18 technical reference.
_docs-website/reference\_versioned\docs/version-2.18 · high confidence
Added backend API routes for search and MCP proxying
The documentation website now includes serverless API endpoints to support interactive features. A new search endpoint proxies queries to the Deepset search API, allowing the site to return results with optional filtering. An MCP (Model Context Protocol) proxy endpoint forwards JSON-RPC requests to the Deepset MCP service, handling session IDs and protocol versions to enable AI tool integration. Additionally, a well-known endpoint returns a JSON 404 for OAuth discovery probes to ensure compatibility with MCP clients that expect JSON responses rather than HTML.
docs-website/api · high confidence
Added documentation templates for components and document stores
New MDX templates have been added to the versioned documentation site to standardize how component and document store pages are structured. The component template provides a structured layout including a key-value table for pipeline positioning and variables, along with sections for overview and usage examples. The document store template defines a consistent format for describing features, initialization steps, supported retrievers, and limitations, ensuring uniform documentation for these specific Haystack elements.
(repo-wide) · high confidence
Added v2.18 documentation for Agents, State, and SuperComponents
The v2.18 documentation site now includes comprehensive guides for building AI agents, including the new \State\ container for sharing data between tools and the \SuperComponent\ feature for wrapping entire pipelines as single components. These pages cover core concepts like component creation, data classes (such as \ChatMessage\ and \Document\), and device management, providing users with the necessary reference material to implement complex agent workflows and reusable pipeline structures in Haystack.
_docs-website/versioned\docs/version-2.18 · high confidence
Docs website structure and configuration initialized
The \docs-website\ directory is now established as the authoritative source for Haystack documentation, powered by Docusaurus 3. This change introduces the core site configuration (\docusaurus.config.js\), navigation structures (\sidebars.js\, \reference-sidebars.js\), and versioning logic that manages both main guides and the auto-generated API reference. It also adds essential infrastructure files such as \.gitignore\, \.gitattributes\, and a \vercel.json\ for deployment redirects, alongside contributor guides (\CONTRIBUTING.md\, \AGENTS.md\) and a landing page redirect to ensure a consistent entry point for users.
docs-website · high confidence
Document Haystack 2.30 experimental and core APIs
The version-2.30 documentation now includes reference pages for the new experimental Agents API (tool-using agents with human-in-the-loop confirmation strategies), experimental memory and chat message stores (InMemoryChatMessageStore, Mem0MemoryStore), and experimental components such as the MarkdownHeaderLevelInferrer, ChatMessageRetriever, ChatMessageWriter, and LLMSummarizer. It also documents the stable Agents API (with Jinja2 user prompts and state schemas) and the Audio API (LocalWhisperTranscriber).
_docs-website/reference\_versioned\docs/version-2.30 · high confidence
Document experimental and stable Agents APIs for Haystack 2.26
The Haystack 2.26 reference documentation now includes dedicated API pages for the Agents component. The stable \agents\_api.md\ page documents the core \Agent\ class, covering its initialization with chat generators and tools, support for Jinja2 prompt templates, and execution logic. Additionally, the new \experiments-api/\ directory introduces reference pages for experimental features, including the \Agent\ extension with human-in-the-loop confirmation strategies, \ChatMessageStore\ and \ChatMessageRetriever\/\Writer\ components for session management, a \Mem0MemoryStore\ for external memory integration, an \OpenAIChatGenerator\ with hallucination risk scoring, a \MarkdownHeaderLevelInferrer\ for document preprocessing, and an \LLMSummarizer\ for text condensation.
_docs-website/reference\_versioned\docs/version-2.26 · high confidence
DocumentWriter now supports async execution and explicit resource cleanup
The DocumentWriter component in haystack.components.writers has been updated to include asynchronous execution capabilities via a new \run\_async\ method, allowing users to write documents to a DocumentStore without blocking the event loop. Additionally, the component now exposes \close\ and \close\_async\ methods to explicitly release resources held by the underlying DocumentStore, ensuring proper cleanup of synchronous and asynchronous connections respectively.
haystack/components/writers · high confidence
Documentation for Haystack 2.20 pipeline components
This update adds comprehensive documentation for the Haystack 2.20 pipeline components, including new pages for the Agent component, audio transcription tools (LocalWhisperTranscriber, RemoteWhisperTranscriber), builders (AnswerBuilder, ChatPromptBuilder, PromptBuilder), caching (CacheChecker), classifiers (DocumentLanguageClassifier, TransformersZeroShotDocumentClassifier), and connectors. The documentation provides usage examples, parameter details, and integration guides for these components.
_docs-website/versioned\_docs/version-2.20/pipeline-components, docs-website/versioned\docs/version-2.21/pipeline-components · high confidence
Documentation for Haystack 2.22 tool integrations and components
Added comprehensive documentation for the Haystack 2.22 tools ecosystem, covering core abstractions like \Tool\, \Toolset\, \ComponentTool\, \PipelineTool\, and \MCPTool\/\MCPToolset\. The new docs also detail the GitHub integration tools (\GitHubFileEditorTool\, \GitHubIssueCommenterTool\, \GitHubIssueViewerTool\, \GitHubPRCreatorTool\, \GitHubRepoViewerTool\), providing usage examples for pipelines, agents, and direct invocation.
_docs-website/versioned\docs/version-2.22/tools · high confidence
Documentation for Haystack 2.27 pipeline components is published
The documentation for version 2.27 has been promoted to the live site, adding comprehensive guides for new and existing pipeline components. This includes detailed pages for the Agent component and its Human-in-the-Loop capabilities, audio processing tools (LocalWhisperTranscriber, RemoteWhisperTranscriber), prompt builders (PromptBuilder, ChatPromptBuilder), and document classifiers. The update also covers caching, builders, and external integrations, providing users with code examples and configuration instructions for these features.
_docs-website/versioned\docs/version-2.27/pipeline-components · high confidence
Haystack 2.20 documentation site launched with core guides and migration resources
The Haystack 2.20 documentation website now includes a complete set of overview pages to help users get started and migrate from previous versions. New content covers installation (pip/conda), building a first RAG application, and a detailed migration guide from Haystack 1.x to 2.x, highlighting key changes like the package name shift from \farm-haystack\ to \haystack-ai\ and the new component-based pipeline architecture. The site also features a breaking change policy, an FAQ with links to archived 1.x and 2.17- documentation, and a comprehensive guide for migrating from LangChain/LangGraph, mapping concepts like nodes to components and graphs to pipelines.
_docs-website/versioned\docs/version-2.20/overview · high confidence
Haystack 2.25 documentation site launched with new overview pages
The documentation website now includes a dedicated set of overview pages for version 2.25, providing users with essential onboarding and reference materials. These new pages cover the installation process via pip, uv, and conda; a comprehensive migration guide for moving from Haystack 1.x to 2.x; a detailed comparison and migration guide for users transitioning from LangChain/LangGraph; a clear definition of the breaking change policy and versioning conventions; an FAQ addressing common setup and support questions; and a transparency page on anonymous telemetry usage and opt-out methods.
_docs-website/versioned\docs/version-2.25/overview · high confidence
Haystack 2.26 documentation site goes live with new overview pages
The Haystack 2.26 documentation website now includes a comprehensive set of overview pages to help users get started and understand the framework. New content includes a detailed breaking change policy, an installation guide supporting pip, uv, and conda, and a migration guide for moving from Haystack 1.x to 2.x. Additionally, a new guide helps users migrating from LangGraph/LangChain by mapping concepts and providing code comparisons, while an FAQ page addresses common questions about GPU usage, telemetry, and support.
_docs-website/versioned\docs/version-2.26/overview · high confidence
Haystack 2.27 documentation site goes live with core guides and migration resources
The Haystack 2.27 documentation website has been published, providing users with essential onboarding and migration resources. New pages include a Getting Started guide with step-by-step RAG and agent examples for OpenAI, Hugging Face, and Anthropic, alongside installation instructions for pip, uv, and conda. A comprehensive migration guide details the transition from Haystack 1.x to 2.x, covering changes to package names, pipeline structures, and component mappings. Additionally, the site now features a breaking change policy, an FAQ, and an extended migration guide for users moving from LangChain/LangGraph, helping developers understand the architectural differences and map existing patterns to Haystack's ecosystem.
_docs-website/versioned\docs/version-2.27/overview · high confidence
Haystack v3.0 documentation published
The documentation for Haystack version 3.0 is now available, introducing new concepts and components such as AI Agents, Multi-Agent Systems, and SuperComponents. The guide covers the updated Component API, including custom component creation and input/output mapping, and details the new data classes like ChatMessage, FileContent, and ImageContent for handling multimodal inputs. It also includes migration tips for upgrading from previous versions and documentation for new integrations like SQLAlchemyTableRetriever, MariaDB Document Store, and various web search tools.
_docs-website/versioned\docs/version-3.0 · high confidence
InMemoryDocumentStore now supports shared process-global storage
The InMemoryDocumentStore now allows multiple instances to share the same underlying document data when initialized with the same index and the new \shared=True\ parameter (which is the default). This enables process-global storage where instances operate on the same documents, while setting \shared=False\ keeps data instance-local and frees it when the instance is garbage collected, preventing unbounded memory growth for frequently created stores.
_haystack/document\_stores/in\memory · high confidence
Initial ClusterFuzzLite integration for Python fuzzing
Added the initial configuration and build scripts for ClusterFuzzLite to enable continuous fuzzing of Haystack's Atheris fuzz targets. This includes a Dockerfile based on the oss-fuzz-base Python builder, a build script that installs the package and compiles fuzz harnesses with necessary dependencies like NumPy, and a project configuration specifying the Python language and address sanitizer.
.clusterfuzzlite · high confidence
Initial documentation for Haystack version 2.19
This release adds the complete set of documentation pages for Haystack version 2.19, covering core concepts such as Agents, State, Components, SuperComponents, Pipelines, and Data Classes (including ChatMessage and StreamingChunk). It also includes guides for Device Management, Custom Components, and Document Stores, providing users with the reference material needed to build and configure pipelines in this version.
_docs-website/versioned\docs/version-2.19 · high confidence
Introduce Agent Hooks for run-loop interception
The \haystack/hooks\ module now provides a new hooking system that allows users to register callbacks at specific points in the Agent's execution lifecycle, including \before\_run\, \before\_llm\, \before\_tool\, \after\_tool\, \on\_exit\, and \after\_run\. Users can define hooks using the \@hook\ decorator or by implementing the \Hook\ protocol, enabling them to inspect or mutate the Agent's live \State\ (such as messages or tool call counts) to influence the run's behavior. The implementation supports both synchronous and asynchronous hooks, handles serialization for persistence, and integrates with the tracing system to log hook invocations.
haystack/hooks · high confidence
Introduce Agent component with stateful execution, hooks, and tracing
This change introduces the new \Agent\ component in \haystack.components.agents\, providing a stateful, multi-step orchestration capability. Users can now build agents that maintain a \State\ object across runs, track metadata like token usage and tool call counts, and execute tools with dedicated tracing spans. The component supports runtime tool selection, lifecycle hooks (before/after run and tool), and integrates with Haystack's tracing and serialization systems.
haystack/components/agents · high confidence
Introduce CacheChecker component for document store caching
A new CacheChecker component is available in haystack.components.caching, allowing users to check for the presence of documents in a Document Store based on a specified metadata field. It returns matching documents as 'hits' and non-matching items as 'misses'. The component supports both synchronous and asynchronous execution via run() and run\_async(), and includes close() and close\_async() methods to properly release underlying Document Store resources.
haystack/components/caching · high confidence
Introduce ChatGenerator protocol for chat components
A new \ChatGenerator\ protocol has been added to define the minimal interface for chat generator components. This protocol specifies that implementations must provide a \run\ method accepting a list of \ChatMessage\ objects and returning a dictionary, allowing for flexible parameter handling in subclasses.
haystack/components/generators/chat/types · high confidence
Introduce EvaluationRunResult for structured evaluation reporting
The \haystack/evaluation\ module now provides the \EvaluationRunResult\ class, which stores inputs and outputs from evaluation pipelines and offers methods to generate aggregated or detailed reports in JSON, CSV, or DataFrame formats. This new component supports comparative analysis between two evaluation runs via \comparative\_detailed\_report\, allowing users to inspect and export metric scores and input data for debugging and performance tracking.
haystack/evaluation · high confidence
Introduce Haystack 2.29 documentation with experimental agent and memory components
The v2.29 reference documentation is now available, introducing a new \agents-api\ section that documents the \Agent\ component for tool-using, provider-agnostic chat workflows. It also adds an \experiments-api\ section covering experimental features, including a tool-using \Agent\ with human-in-the-loop confirmation strategies, an \InMemoryChatMessageStore\ for session isolation, a \Mem0MemoryStore\ for external memory backends, and an \OpenAIChatGenerator\ with hallucination risk scoring. Additional experimental components include a \MarkdownHeaderLevelInferrer\ for normalizing document structure, a \ChatMessageRetriever\ and \ChatMessageWriter\ for managing conversation history, and an \LLMSummarizer\ for recursive text summarization.
_docs-website/reference\_versioned\docs/version-2.29 · high confidence
Introduce JsonSchemaValidator component for LLM output validation
A new \JsonSchemaValidator\ component is now available in \haystack.components.validators\. It validates the JSON content of \ChatMessage\ instances against a specified JSON Schema, routing valid messages to a \validated\ output and invalid ones to a \validation\_error\ output. This enables users to implement recovery loops where LLMs can correct their JSON output based on detailed error messages generated by the validator.
haystack/components/validators · high confidence
Introduce LinkContentFetcher for HTTP-based content fetching
A new \LinkContentFetcher\ component is added to \haystack.components.fetchers\, enabling pipelines to fetch and extract content from URLs. It replaces the previous \requests\ library with \httpx\ to support asynchronous operations and optional HTTP/2. The component handles various content types (text, binary, JSON, images, etc.) by returning \ByteStream\ objects, includes configurable retry logic with exponential backoff, and rotates user agents on failed requests to improve reliability. It also supports custom request headers and timeout settings.
haystack/components/fetchers · high confidence
Introduce OpenAIImageGenerator and internal tracing for ChatGenerators
Users can now generate images using OpenAI's models (defaulting to gpt-image-2) via the new OpenAIImageGenerator component, which supports synchronous and asynchronous execution, configurable quality/size, and custom HTTP client settings. Additionally, internal calls to ChatGenerators are now wrapped in tracing spans, ensuring that LLM token usage metadata is correctly exposed to tracers even when the generator is used internally by other components like rankers or extractors.
haystack/components/generators · high confidence
Introduce QueryExpander component for query expansion
A new QueryExpander component is available in the query module, designed to improve retrieval recall by generating semantically similar alternative queries. It leverages a chat generator (defaulting to OpenAI's gpt-4.1-mini) to produce a JSON response containing expanded queries, which can be configured via parameters like n\_expansions and include\_original\_query. The component handles serialization and deserialization of the chat generator and validates prompt templates for required variables.
haystack/components/query · high confidence
Introduce SkillToolset for agent-driven skill discovery and file reading
A new SkillToolset is available in haystack.tools.skills, enabling Haystack Agents to discover and interact with skills via progressive disclosure. The toolset exposes two fixed tools: load\_skill, which returns a skill's full instructions and a manifest of bundled files after discovering available skills during warm\_up, and read\_skill\_file, which retrieves specific bundled files (including images and PDFs via ImageContent/FileContent) for multimodal consumption. This allows agents to dynamically access specialized instruction sets and their associated resources without system prompt injection.
haystack/tools/skills · high confidence
Introduce State class for managing shared context in Agents
A new \State\ class has been added to \haystack.components.agents.state\ to allow Agents and their tools to share and manage context during execution. The class uses a schema to define typed fields (such as \messages\ for \list\[ChatMessage\]\) and supports custom merge handlers, defaulting to list concatenation for lists and value replacement for other types. It includes serialization support via \to\_dict\ and \from\_dict\ methods, enabling state persistence and snapshotting.
haystack/components/agents/state · high confidence
Introduce SuperComponent to wrap and expose Haystack Pipelines as single components
Users can now wrap an existing Haystack Pipeline inside a SuperComponent, treating the entire pipeline as a single reusable component with its own input and output interfaces. The SuperComponent class and its accompanying decorator allow you to define custom input and output mappings that flatten or reshape the pipeline's internal sockets, enabling seamless integration of complex multi-step workflows into other pipelines or agent tools. It supports both synchronous and asynchronous execution, handles component lifecycle management (warm-up and close), and includes type compatibility checks to ensure mapped inputs and outputs align correctly.
_haystack/core/super\component · high confidence
Introduce TokenBudgetHook to limit Agent token usage
Users can now attach a \TokenBudgetHook\ to an Agent to automatically stop execution when cumulative token usage reaches a specified limit. The hook monitors token counts at the \before\_llm\ stage, supports optional final message insertion, and integrates with the Agent's state to signal a \token\_budget\_exceeded\ stop reason.
haystack/hooks/budget · high confidence
Introduce YAML-based serialization with tuple support
The \haystack/marshal\ module has been introduced to handle component serialization, providing a \YamlMarshaller\ that converts pipeline data to and from YAML strings. This implementation includes custom YAML loaders and dumpers that explicitly support Python tuples, allowing them to be preserved during serialization and deserialization, while enforcing that only basic Python types are used in the serialized data to ensure compatibility.
haystack/marshal · high confidence
Introduce context compaction for long-running Agents
Agents can now automatically manage conversation history to prevent context window overflow. A new \CompactionHook\ monitors the token usage before each LLM call and triggers a compaction strategy when the conversation exceeds a configured threshold. The module provides three compaction strategies: \SlidingWindowCompactor\ (drops oldest messages while preserving structure), \SummarizationCompactor\ (uses an LLM to summarize older turns), and \ToolResultPruningCompactor\ (replaces large tool outputs with placeholders). This allows Agents to maintain longer interactions without hitting model limits.
haystack/hooks/compaction · high confidence
Introduce core component infrastructure with Input/Output Sockets and Variadic types
This change establishes the foundational component model in \haystack/core/component\, introducing the \Component\ protocol, the \@component\ decorator, and the \InputSocket\/\OutputSocket\ data structures that define component interfaces. It adds support for variadic inputs via \Variadic\ and \GreedyVariadic\ type annotations, allowing components to accept multiple connections or run immediately upon receiving the first input. The \Sockets\ class provides a unified view of these inputs and outputs, enabling the pipeline to manage data flow and component registration.
haystack/core/component · high confidence
Introduce modular chat generator components and new LLM/Fallback/Mock generators
The chat generator components are now organized into a dedicated \haystack.components.generators.chat\ package with lazy loading, exposing \OpenAIChatGenerator\, \OpenAIResponsesChatGenerator\, \AzureOpenAIChatGenerator\, \AzureOpenAIResponsesChatGenerator\, \FallbackChatGenerator\, \LLM\, and \MockChatGenerator\. The new \LLM\ component provides a simplified, tool-free text generation interface built on top of \Agent\, dynamically requiring or allowing \messages\ based on prompt configuration. \FallbackChatGenerator\ enables sequential fallback across multiple chat generators with detailed metadata on failures. \MockChatGenerator\ offers a deterministic, API-free replacement for testing, supporting fixed, cycling, and dynamic tool-aware responses. Azure generators now support Azure AD token providers and explicit \max\_retries\/timeout configuration, while all generators expose \SUPPORTED\_MODELS\ and support \tools\_strict\ for strict schema adherence.
haystack/components/generators/chat · high confidence
Introduce new file converters and OutputAdapter component
The \haystack/components/converters\ package now includes a comprehensive set of new components for converting various file formats into Haystack Documents, including CSV, DOCX, HTML, JSON, Markdown, MSG, PPTX, PDF (via PyPDF and PDFMiner), and XLSX. A new \FileToFileContent\ component allows converting files into base64-encoded \FileContent\ objects for chat messages, while \MultiFileConverter\ provides a unified entry point for routing and converting multiple file types. Additionally, the \OutputAdapter\ component enables adapting component outputs using Jinja templates, with support for custom filters and configurable security modes.
haystack/components/converters · high confidence
Introduce tool result offloading to manage agent context window
The \haystack.hooks.tool\_result\_offloading\ module now provides a hook that automatically offloads large tool results from the agent's conversation history to an external store, replacing them with compact pointers to preserve context window space. Users can register the \ToolResultOffloadHook\ on an agent and configure per-tool strategies—such as \AlwaysOffload\, \NeverOffload\, or \OffloadOverChars\—to control which results are offloaded. The module includes a \FileSystemToolResultStore\ for local persistence and supports binary content (images and files) when the store enables it. This feature helps prevent context overflow when agents interact with tools that produce large outputs.
_haystack/hooks/tool\_result\offloading · high confidence
New LLM and Regex text extraction components
Introduces two new components in the \haystack.components.extractors\ package: \LLMMetadataExtractor\, which uses a \ChatGenerator\ to extract structured metadata from documents based on a prompt and expected JSON keys, and \RegexTextExtractor\, which extracts text from strings or chat messages using a regular expression pattern with capture groups.
haystack/components/extractors · high confidence
New LLMDocumentContentExtractor for image-based documents
A new \LLMDocumentContentExtractor\ component is now available in \haystack.components.extractors.image\ to extract textual content and metadata from image-based documents using a vision-enabled LLM. The component converts documents to images and sends them to a \ChatGenerator\, handling JSON responses to populate document content and metadata, while supporting configuration for image detail, size, and secure file path resolution via a \root\_path\ parameter.
haystack/components/extractors/image · high confidence
New and updated document store documentation for Haystack 2.25
The documentation for the \document-stores\ section has been updated for Haystack 2.25, introducing new pages for ArcadeDB, Astra, Azure AI Search, Chroma, Elasticsearch, FAISS, InMemory, MongoDB Atlas, OpenSearch, Pgvector, Pinecone, Qdrant, Valkey, and Weaviate. These pages provide installation instructions, initialization examples, and details on supported retrievers for each integration. Additionally, the FAISS page now includes a troubleshooting section for resolving OpenMP runtime conflicts on macOS.
_docs-website/versioned\docs/version-2.25/document-stores · high confidence
New automation scripts for documentation and release notes
Added two new scripts to the \scripts\ directory: \generate\_platform\_components\_table.py\ automates the creation of the Haystack Enterprise Platform Components MDX table by scanning source code for \@component\-decorated classes and cross-referencing them against the platform schema, and \release\_note\_backticks.py\ acts as a pre-commit hook to enforce double backtick syntax for inline code in reno release notes, preventing RST rendering issues.
scripts · high confidence
New core package with hardened serialization and flexible type connections
The \haystack/core\ package has been introduced, consolidating core functionality into a new structure. This update introduces a security-focused deserialization system that gates component loading through a module allowlist and blocks dangerous built-in functions to prevent remote code execution, with an optional \HAYSTACK\_UNSAFE\_DESERIALIZATION\ environment variable to bypass these checks for trusted environments. Additionally, the pipeline connection logic now supports flexible type conversions, including automatic wrapping and unwrapping of lists, compatibility between \ChatMessage\ and \str\, and support for Python 3.10+ union syntax (\X \| Y\), while stricter validation prevents silent data loss from incompatible type mismatches.
haystack/core · high confidence
New deployment and observability documentation for Haystack 2.22
Added comprehensive guides for deploying Haystack pipelines using Docker, Kubernetes, and OpenShift, including specific instructions for integrating with Hayhooks. The update also introduces new documentation for enabling GPU acceleration, configuring logging (standard, real-time, and structured), and setting up tracing backends such as OpenTelemetry, Datadog, Langfuse, and Weights & Biases.
_docs-website/versioned\docs/version-2.22/development · high confidence
New document store integrations added to Haystack 2.20 documentation
The documentation for Haystack 2.20 now includes dedicated pages for several new document store integrations, providing users with setup guides, initialization examples, and supported retrievers for Astra, Azure AI Search, Chroma, Elasticsearch, In-Memory, MongoDB Atlas, OpenSearch, Pgvector, Pinecone, Qdrant, and Weaviate. These new pages cover installation steps, connection configurations (including environment variables and authentication), and specific retrieval components available for each store, enabling users to easily integrate these vector databases into their Haystack pipelines.
_docs-website/versioned\docs/version-2.20/document-stores · high confidence
New documentation components for image zooming, content copying, and product-specific visibility
The docs website now includes several new React components to enhance the user experience. Users can click on images to view them in a zoomed overlay via the new ClickableImage component. A CopyDropdown component allows users to copy the current page content as Markdown or share it with AI assistants like ChatGPT, Claude, and Perplexity. The ProductOnly component enables conditional rendering of documentation content based on the active product context, ensuring users only see relevant information. Additionally, a YouTubeEmbed component provides responsive video embedding, and the AgentsContent MDX file has been moved to the components directory.
docs-website/src/components · high confidence
New documentation for Agent, Human-in-the-Loop, and Audio components
The documentation site now includes dedicated pages for the \Agent\ component, which enables iterative, tool-using LLM workflows, and the \Human-in-the-loop\ (HITL) system, allowing users to intercept, confirm, reject, or modify agent tool calls before execution. Additionally, new documentation has been added for audio processing, covering \LocalWhisperTranscriber\ and \RemoteWhisperTranscriber\ for local and API-based speech-to-text transcription, along with their respective external integrations.
_docs-website/versioned\docs/version-2.25/pipeline-components · high confidence
New documentation for Agents, Multi-Agent Systems, and Component Templates
The documentation site now includes dedicated guides for building AI agents, including a new Agents concept page, a Multi-Agent Systems guide for coordinator/specialist architectures, and updated State and Human-in-the-loop documentation. Additionally, new component documentation templates (component-template.mdx and document-store-template.mdx) have been added to standardize how component reference pages are structured.
_docs-website/versioned\docs/version-2.28 · high confidence
New documentation for Agents, State, and SuperComponents
The documentation now includes dedicated guides for building AI agents, managing execution state, and wrapping pipelines as components. The new 'Agents' page explains how to create agents using the Agent component, ToolInvoker, and various tool types (Tool, ComponentTool, decorators). A new 'State' guide details how to use the State container to share data and messages between tools in an agent workflow, including schema definition and custom merge handlers. Additionally, 'SuperComponents' documentation explains how to wrap a complete pipeline into a single component using the @super\_component decorator or the SuperComponent class, simplifying complex pipeline interfaces for reuse.
_docs-website/versioned\docs/version-2.23 · high confidence
New documentation for AsyncPipeline, pipeline loops, and breakpoints
The Haystack 2.21 documentation now includes dedicated guides for advanced pipeline features. Users can learn how to use AsyncPipeline for concurrent component execution with synchronous, asynchronous, and generator-based run methods. New pages explain how to implement and control pipeline loops for feedback and self-correction, including safety limits to prevent infinite runs. Additionally, the documentation covers pipeline breakpoints, allowing developers to pause execution at specific components, inspect state via JSON snapshots, and resume workflows, as well as error recovery using automatic snapshots on failure.
_docs-website/versioned\docs/version-2.21/concepts/pipelines · high confidence
New documentation for Haystack 2.22 concepts
The documentation site now includes a new 'concepts' section for Haystack 2.22, introducing pages that explain core architectural elements. This includes an overview of agents, their components (like ToolInvoker and ComponentTool), and how to build them using pipelines or the Agent class. It also introduces the State container for managing shared information and intermediate results during agent execution. Additionally, new pages cover the fundamentals of components, how to create custom components, and the SuperComponent pattern for wrapping entire pipelines as single units. The documentation also details the core data classes (Document, ChatMessage, Answer, etc.) and the framework-agnostic device management system for local model inference.
(repo-wide) · high confidence
New documentation for Haystack 2.25 core concepts
The documentation site now includes a comprehensive set of new concept pages for Haystack 2.25, covering Agents, State, Components, Custom Components, SuperComponents, Data Classes (including ChatMessage), and Device Management. These pages provide users with detailed explanations and code examples for building AI agents, managing shared state between tools, creating and extending custom pipeline components, wrapping pipelines as single units, and handling hardware device allocation for local model inference.
_docs-website/versioned\_docs/version-2.25/concepts, docs-website/versioned\docs/version-2.27/concepts · high confidence
New documentation for Haystack 2.25 tools and integrations
Added comprehensive documentation for Haystack 2.25 tools, including \ComponentTool\, \PipelineTool\, \MCPTool\, \MCPToolset\, and \SearchableToolset\. The release also introduces documentation for new GitHub integration tools (\GitHubFileEditorTool\, \GitHubIssueCommenterTool\, \GitHubIssueViewerTool\, \GitHubPRCreatorTool\, \GitHubRepoViewerTool\) and updates the \MCPTool\ docs to reflect the deprecation of SSE transport in favor of Streamable HTTP.
_docs-website/versioned\docs/version-2.25/tools · high confidence
New documentation for Haystack 2.29 concepts, templates, and data classes
This release adds comprehensive documentation for Haystack version 2.29, introducing new concept pages for Agents (including multi-agent systems), Components (including custom components and SuperComponents), and Data Classes (including detailed guides for ChatMessage and FileContent). It also adds reusable MDX templates for component and document store documentation to standardize future integration guides.
_docs-website/versioned\docs/version-2.29 · high confidence
New documentation for Haystack deployment, observability, and optimization
The documentation site now includes comprehensive guides for deploying Haystack pipelines via Docker, Kubernetes, and OpenShift, along with detailed instructions for using Hayhooks to serve pipelines as REST APIs. It also covers observability features, including logging configuration and tracing integrations with OpenTelemetry, Datadog, Langfuse, and Weights & Biases. Additionally, new sections explain how to enable GPU acceleration and implement advanced RAG techniques like HyDE, providing users with the necessary resources to optimize and monitor their Haystack applications in production.
_docs-website/versioned\docs/version-2.21/development · high confidence
New documentation for advanced RAG techniques and pipeline evaluation
The documentation site now includes dedicated guides for advanced Retrieval-Augmented Generation (RAG) strategies and system evaluation. Users can learn how to implement Hypothetical Document Embeddings (HyDE) to improve retrieval in specialized domains, alongside links to related cookbooks like Query Decomposition and Auto-Merging Retrieval. Additionally, a new evaluation section explains how to assess pipeline performance using both model-based methods (using LLMs or local models for metrics like faithfulness and context relevance) and statistical methods (using metrics like recall and exact match), complete with component references and integration details for frameworks like Ragas and DeepEval.
_docs-website/versioned\_docs/version-2.20/optimization, docs-website/versioned\docs/version-2.27/optimization · high confidence
New documentation for choosing and creating custom document stores
Added two new documentation pages to the Haystack 2.21 concepts guide: 'Choosing a Document Store' and 'Creating Custom Document Stores'. The new 'Choosing' page provides a comparative overview of five vector database categories (vector libraries, pure vector databases, vector-capable SQL/NoSQL, and full-text search databases) along with an in-memory option, helping users select the appropriate storage backend for their AI applications. The 'Creating Custom Document Stores' page offers a comprehensive guide for developers to build and distribute their own Document Store integrations, detailing the required \DocumentStore\ protocol methods, naming conventions, packaging strategies, serialization requirements, and testing procedures.
_docs-website/versioned\docs/version-2.21/concepts/document-store · high confidence
New documentation for deploying Haystack pipelines and configuring observability
Added comprehensive guides for deploying Haystack pipelines using Docker, Kubernetes, and OpenShift, including specific instructions for integrating with the Hayhooks REST API server. The update also introduces new documentation pages for enabling GPU acceleration, configuring standard and structured logging, and setting up tracing backends such as OpenTelemetry, Datadog, Langfuse, MLflow, and Weights & Biases.
_docs-website/versioned\docs/version-2.25/development · high confidence
New documentation for deploying Haystack pipelines and using Hayhooks
Added comprehensive guides for deploying Haystack pipelines via Docker, Kubernetes, and OpenShift, including configuration examples for Hayhooks. Introduced new documentation for Hayhooks, detailing installation, configuration, pipeline deployment, and file upload handling. Also added guides for enabling GPU acceleration, external integrations (Arize, Ray, etc.), and detailed logging and tracing setups with OpenTelemetry, Datadog, and Langfuse.
_docs-website/versioned\_docs/version-2.20/development, docs-website/versioned\_docs/version-2.26/development, docs-website/versioned\docs/version-2.27/development · high confidence
New documentation for multiple document stores and ComponentTool
Added new documentation pages for the ArcadeDB, Astra, Azure AI Search, Chroma, Elasticsearch, FAISS, InMemory, MongoDB Atlas, OpenSearch, Pgvector, Pinecone, Qdrant, Valkey, and Weaviate document stores, providing installation, initialization, and usage guides for each. Additionally, added documentation for the ComponentTool, which allows Haystack components to be used as tools by LLMs.
_docs-website/versioned\_docs/version-2.27/document-stores, docs-website/versioned\docs/version-2.27/tools · high confidence
New documentation pages for Haystack 2.22 overview
The Haystack 2.22 documentation now includes several new overview pages: a Breaking Change Policy detailing versioning and deprecation rules, an FAQ addressing common questions like GPU usage and telemetry, a Get Started guide with code examples for OpenAI, Hugging Face, and Anthropic, an Installation page covering pip/conda setup and optional dependencies, a Migration Guide for moving from Haystack 1.x to 2.x, a LangChain/LangGraph migration guide, and a Telemetry page explaining data collection and opt-out methods.
_docs-website/versioned\_docs/version-2.21/overview, docs-website/versioned\docs/version-2.22/overview · high confidence
New documentation templates and comprehensive agent system guides
The documentation site now includes reusable MDX templates for standardizing component and document store pages, alongside new conceptual guides explaining how to build AI agents, create custom components, and orchestrate multi-agent systems using coordinator and specialist patterns. These additions provide structured reference material and practical examples for integrating agents, tools, and state management into Haystack pipelines.
_docs-website/versioned\docs/version-2.30 · high confidence
New embedder components and lazy-loading module structure
The embedders module now exposes a lazy-importing \\_\init\\_.py\ that makes \AzureOpenAIDocumentEmbedder\, \AzureOpenAITextEmbedder\, \MockDocumentEmbedder\, \MockTextEmbedder\, \OpenAIDocumentEmbedder\, and \OpenAITextEmbedder\ available via \haystack.components.embedders\. This change introduces mock embedders that generate deterministic embeddings without calling external APIs (useful for tests and prototypes), and adds Azure-specific embedders that inherit from the OpenAI base classes to support Azure endpoints, API keys, and Azure AD tokens.
haystack/components/embedders · high confidence
New evaluation components for retrieval and LLM-based metrics
The \haystack.components.evaluators\ package now provides a suite of new components for evaluating RAG pipelines. Retrieval metrics include \DocumentMAPEvaluator\, \DocumentMRREvaluator\, \DocumentNDCGEvaluator\, and \DocumentRecallEvaluator\, all of which support configurable \document\_comparison\_field\ options (content, id, or meta keys). LLM-based evaluation is handled by \LLMEvaluator\ and its specialized subclasses: \FaithfulnessEvaluator\, \ContextRelevanceEvaluator\, \AnswerExactMatchEvaluator\, and \SASEvaluator\. These LLM evaluators use a \chat\_generator\ parameter for the underlying model, support async execution, and expose detailed per-item results and status metadata.
haystack/components/evaluators · high confidence
New experimental and stable API reference documentation for agents, memory, and generators
The Haystack 2.25 documentation now includes detailed API references for the new Agent component, which supports tool-using workflows with provider-agnostic chat models and human-in-the-loop confirmation strategies. It also documents the experimental memory and chat message stores (including Mem0 and in-memory implementations), the ChatMessageRetriever and ChatMessageWriter components for session management, the MarkdownHeaderLevelInferrer for text preprocessing, the LLMSummarizer for text condensation, and the OpenAIChatGenerator with built-in hallucination risk scoring capabilities.
_docs-website/reference\_versioned\docs/version-2.25 · high confidence
New filter merging policy and document store protocol types
The \haystack.document\_stores.types\ module now exposes a \FilterPolicy\ enum with \REPLACE\ and \MERGE\ options, along with an \apply\_filter\_policy\ function that allows runtime filters to be merged with initialization filters (with runtime values overwriting init values). This is accompanied by the \DuplicatePolicy\ enum for controlling document write behavior (NONE, SKIP, OVERWRITE, FAIL) and the \DocumentStore\ protocol definition, which standardizes the interface for document storage, retrieval, and filtering across different backends.
_haystack/document\stores/types · high confidence
New image file and PDF conversion components
This change introduces a new \haystack.components.converters.image\ package containing four components for handling visual content: \ImageFileToDocument\ wraps image file paths into empty Document objects for downstream processing; \ImageFileToImageContent\ converts image files directly into base64-encoded ImageContent objects with optional resizing; \PDFToImageContent\ extracts specific pages from PDF files into ImageContent objects; and \DocumentToImageContent\ processes existing Documents to extract visual content from their referenced image or PDF sources. These components rely on Pillow and pypdfium2 for image processing and PDF rendering.
haystack/components/converters/image · high confidence
New joiner components and package structure
The \haystack.components.joiners\ package has been introduced to centralize joining logic, featuring new components: \AnswerJoiner\ for merging answer lists with optional score-based sorting, \BranchJoiner\ for converging pipeline branches (e.g., loop handling), \ListJoiner\ for flattening multiple lists of any type, and \StringJoiner\ for combining string inputs. The existing \DocumentJoiner\ is also included in this package, supporting modes like concatenate, merge, reciprocal rank fusion, and distribution-based rank fusion.
haystack/components/joiners · high confidence
New preprocessor components for CSV and hierarchical document handling
The \haystack.components.preprocessors\ package now includes \CSVDocumentCleaner\ for removing empty rows and columns from CSV content, \CSVDocumentSplitter\ for splitting CSVs by row-wise or threshold-based empty sections, and \HierarchicalDocumentSplitter\ for creating multi-level block structures. Additionally, \DocumentPreprocessor\ is introduced as a SuperComponent that chains \DocumentSplitter\ and \DocumentCleaner\ into a single pipeline, while \DocumentCleaner\ gains support for unicode normalization, ASCII-only conversion, and minimum content length filtering.
haystack/components/preprocessors · high confidence
New protocol interfaces for text and document embedders
A new \haystack.components.embedders.types\ module has been introduced, defining \TextEmbedder\ and \DocumentEmbedder\ protocols. These interfaces standardize the expected behavior for embedder components: \TextEmbedder\ expects a \run\ method that accepts a string and returns a dictionary containing an embedding, while \DocumentEmbedder\ expects a \run\ method that accepts a list of documents and returns a dictionary containing the embedded documents. This provides a contract for implementing embedders within the Haystack framework.
haystack/components/embedders/types · high confidence
New ranker components and lazy import structure
The rankers module now exposes four components—LLMRanker, LostInTheMiddleRanker, MetaFieldRanker, and MetaFieldGroupingRanker—via a lazy-import structure. LLMRanker reranks documents using a chat generator (defaulting to OpenAI) with a JSON-based prompt and supports async warm-up and lifecycle management. LostInTheMiddleRanker reorders documents so the most relevant appear at the beginning or end of the context, with optional top-k and word-count thresholds, and deduplicates by document ID. MetaFieldRanker sorts documents by a metadata field with configurable weight, ranking mode, sort order, handling of missing metadata, and optional type parsing (float/int/date), also deduplicating before ranking. MetaFieldGroupingRanker groups and optionally subgroups documents by metadata keys and sorts within groups, preserving insertion order when sorting is not possible or when values are non-comparable. All rankers deduplicate documents by ID before processing.
haystack/components/rankers · high confidence
New retriever components and lazy-loading module structure
The \haystack.components.retrievers\ package now exposes several new components: \AutoMergingRetriever\ (merges hierarchical leaf documents into parents based on a threshold), \FilterRetriever\ (retrieves documents using static or dynamic filters), \MultiRetriever\ (runs multiple text retrievers in parallel with optional reciprocal rank fusion), \MultiQueryEmbeddingRetriever\ and \MultiQueryTextRetriever\ (run multiple queries in parallel for embedding- and text-based retrievers respectively), \SentenceWindowRetriever\ (fetches surrounding context documents), and \TextEmbeddingRetriever\ (wraps an embedding retriever with a text embedder). The package \\_\init\\_.py\ has been replaced with a lazy-import structure to optimize import times, making these components available via \from haystack.components.retrievers import ...\.
haystack/components/retrievers · high confidence
New routers package consolidates and introduces document, file, metadata, and LLM-based routing components
The \haystack.components.routers\ package has been introduced to centralize routing logic, providing a unified entry point for \ConditionalRouter\, \DocumentLengthRouter\, \DocumentTypeRouter\, \FileTypeRouter\, \LLMMessagesRouter\, and \MetadataRouter\. This change consolidates previously scattered router implementations into a single, organized module. \ConditionalRouter\ now supports passthrough routing for complex types and custom Jinja filters. \DocumentTypeRouter\ and \FileTypeRouter\ both support regex-based MIME type matching and handle literal types (like \image/svg+xml\) correctly. \MetadataRouter\ now supports routing \ByteStream\ objects in addition to \Document\s and includes a \strict\_datetime\_comparison\ option for precise timezone-aware filtering. \LLMMessagesRouter\ enables LLM-based classification of chat messages with async support. \DocumentLengthRouter\ categorizes documents by content length. All components follow consistent serialization patterns and error handling.
haystack/components/routers · high confidence
New scripts for local docs setup and Python snippet testing
Added \setup-dev.sh\ to automate the local development environment setup for the documentation site by generating and installing dependencies, and introduced \test\_python\_snippets.py\ to validate Python code examples embedded in Markdown/MDX files. The new test script scans documentation directories for Python code blocks, executes them in isolated subprocesses with timeout and safety checks, and supports markers to force-run, skip, or require specific files for individual snippets. Also added \extract\_sidebar.mjs\ to convert Docusaurus sidebar configuration from JavaScript to JSON format.
docs-website/scripts · high confidence
New telemetry module collects anonymous usage data
A new telemetry module has been introduced to Haystack, enabling the collection of anonymous usage statistics to help improve the software. This module automatically gathers system metadata (such as OS, Python version, and containerization status) and tracks pipeline runs and tutorial usage, sending this data to PostHog. Users can opt out of this data sharing by setting the \HAYSTACK\_TELEMETRY\_ENABLED\ environment variable to \false\ or \0\. The implementation includes rate limiting to prevent excessive events and ensures that telemetry failures do not disrupt application execution.
haystack/telemetry · high confidence
New token counting components for estimating API costs
Haystack now includes a \haystack.token\_counters\ module with three implementations to estimate the number of tokens consumed by chat messages and tool schemas: \ApproximateTokenCounter\ (a fast, local character-to-token ratio estimator), \TiktokenCounter\ (a local estimate using OpenAI's tiktoken encoder), and \OpenAITokenCounter\ (which calls OpenAI's input token counting API for an exact count). These components allow users to predict token usage and associated costs before sending requests to language models.
_haystack/token\counters · high confidence
Promote Haystack 2.31 documentation and publish new experimental API references
The version-2.31 documentation is now promoted to stable, making the latest stable API reference available for users. Additionally, a new set of experimental API reference pages has been published under the \experiments-api\ section, documenting new capabilities such as the \Agent\ component with human-in-the-loop confirmation strategies, \InMemoryChatMessageStore\ and \Mem0MemoryStore\ for chat and memory persistence, \ChatMessageRetriever\ and \ChatMessageWriter\ for session management, \LLMSummarizer\ for text summarization, \MarkdownHeaderLevelInferrer\ for document preprocessing, and an \OpenAIChatGenerator\ with hallucination risk scoring.
_docs-website/reference\_versioned\_docs/version-2.27, docs-website/reference\_versioned\docs/version-2.31 · high confidence
Promotes Haystack 2.28 documentation and adds experimental API references
The Haystack 2.28 reference documentation is now promoted to the stable versioned docs, making it the primary reference for this release. Additionally, a new set of experimental API reference pages has been added to the documentation site, covering components such as the Agent (with human-in-the-loop confirmation strategies), ChatMessage Store, Mem0 Memory Store, Generators (including OpenAI hallucination risk scoring), Preprocessors, Retrievers, Summarizers, and Writers.
_docs-website/reference\_versioned\docs/version-2.28 · high confidence
Promotion of Haystack 2.24 documentation with new agent and component concepts
The documentation for Haystack version 2.24 has been promoted to the stable versioned\_docs directory, making the latest features and guides publicly available. This release introduces comprehensive documentation for building AI agents, including the new \State\ container for managing shared information and tool execution context. It also expands the component model with guides on creating custom components, using the \@super\_component\ decorator to wrap pipelines as single units, and details on core data classes like \ChatMessage\ and \StreamingChunk\. Additionally, the docs now cover device management for local model inference and provide updated migration guides and integration examples.
_docs-website/versioned\docs/version-2.24 · high confidence
TopPSampler now supports a minimum document count via min\_top\_k
The TopPSampler component in haystack.components.samplers now accepts a \min\_top\_k\ parameter. This allows users to specify a minimum number of documents to return; if the top-p sampling logic selects fewer documents than this threshold, the sampler automatically includes additional documents with the next-highest scores to meet the minimum. The component also validates that \min\_top\_k\ is a non-negative integer or None, raising a ValueError otherwise.
haystack/components/samplers · high confidence
Removals
Removal of legacy FARM-based Haystack components
The \farm\_haystack\ module has been completely removed, deleting the legacy FARM-based implementation including the \Finder\ orchestration class, the \FARMReader\ and \TfidfRetriever\ components, the Flask-based REST API, the SQL database ORM, and document indexing utilities. This change eliminates the deprecated FARM integration and its associated infrastructure from the codebase.
_farm\haystack · high confidence
Security
Introduce structured \`Secret\` authentication and hardened callable deserialization
The \haystack.utils\ module now exposes a new \Secret\ class for structured authentication, replacing direct string tokens with \TokenSecret\ (which redacts values in logs) and \EnvVarSecret\ (which resolves from environment variables and is serializable). To prevent remote code execution via untrusted pipeline snapshots, callable deserialization is now gated by a strict module allowlist that validates every module and attribute traversed during import, blocking dangerous builtins and internal deserialization machinery. Additionally, the module provides schema-aware serialization utilities for runtime values (like Agent State) and a generic device management abstraction.
haystack/utils · high confidence
Pipeline core restructured into modular package with enhanced security and snapshot controls
The \haystack.core.pipeline\ module has been reorganized into a structured package, separating the orchestration logic into \base.py\ and \pipeline.py\, while introducing dedicated modules for component execution checks, graph visualization, and breakpoint/snapshot management. This restructuring introduces a module allowlist to gate pipeline deserialization, blocking internal interfaces to prevent remote code execution vulnerabilities. Additionally, pipeline snapshot saving is now disabled by default and controlled via the \HAYSTACK\_PIPELINE\_SNAPSHOT\_SAVE\_ENABLED\ environment variable, and the visualization engine now validates Mermaid server responses against expected magic-byte signatures to prevent arbitrary file writes.
haystack/core/pipeline · high confidence
Architecture
Introduce new tools package with lazy imports and core tool abstractions
The \haystack/tools\ package is now a dedicated module that exports core tooling classes (\Tool\, \ComponentTool\, \PipelineTool\, \AgentTool\, \Toolset\, \SearchableToolset\, \SkillToolset\) and utilities (\create\_tool\_from\_function\, \tool\ decorator, \warm\_up\_tools\, serialization helpers). To optimize import performance, the package uses \lazy\_imports.LazyImporter\ for most symbols, while eagerly loading the \tool\ decorator and \Tool\ base class to avoid naming collisions. This change reorganizes the tooling architecture into a separate package, providing a cleaner public API and better startup times for users who do not use tools.
haystack/tools · high confidence
Behavioural changes
Add versioned API reference documentation for Haystack 2.19
A new API reference page has been added for version 2.19, providing a structured overview of the Haystack framework's technical documentation. This entry point organizes the reference material into three distinct sections: the core Haystack API (covering base components, pipelines, document stores, and utilities), the Integrations API (for official packages connecting to external services), and the Experiments API (for features under active development).
_docs-website/reference\_versioned\docs/version-2.19 · high confidence
Code blocks now hide language badges and action buttons based on line count
The documentation website now dynamically adjusts the appearance of code blocks to reduce visual clutter for short snippets. Single-line code blocks will no longer display the language badge, and two-line code blocks will hide the copy and other action buttons, while multi-line blocks retain the full interface.
docs-website/src/theme/CodeBlock · high confidence
Docs now automatically version cross-plugin internal links
A new Remark plugin (\versionedReferenceLinks\) has been added to the documentation build process to automatically inject version prefixes into internal links between the main docs and the reference section. This ensures that when users navigate from a specific documentation version (e.g., v2.19) to the reference, they are directed to the corresponding versioned reference page (e.g., \/reference/v2.19/...\) rather than the latest version. The latest version remains unversioned, while the current/next version uses a \/next/\ prefix, preventing broken links when users browse older documentation versions.
docs-website/src/remark · high confidence
Documentation site assets and configuration updates
The documentation website now includes a .nojekyll file to ensure proper static file serving on GitHub Pages. New SVG icons for the chevron, copy-to-clipboard, and Discord community have been added to the static image directory. Additionally, a new 'Vector Databases' overview diagram has been introduced to help users understand document store architectures.
docs-website/static · high confidence
Haystack 3.0 release with major architectural changes and component migrations
This entry marks the release of Haystack 3.0, introducing significant breaking changes and architectural shifts. The \AsyncPipeline\ has been merged into \Pipeline\, and \ToolInvoker\ has been removed in favor of \Agent\ owning tool calls. Many components (e.g., HuggingFace, Transformers, Whisper, Tika, OpenAPI) have been moved to external integration packages (\haystack-core-integrations\), requiring users to install new packages and update imports. Tracing backends (Datadog, OpenTelemetry) are no longer auto-enabled and must be explicitly configured via new connector components. \GeneratedAnswer\ and \ExtractedAnswer\ serialization formats have changed to flat dictionaries. The release also drops Python 3.9 support, adds native async support to \Pipeline\, and introduces Agent lifecycle hooks and state tracking for step counts and token usage.
(repo-wide) · high confidence
Haystack dataclasses module restructured with new content types and serialization updates
The \haystack.dataclasses\ package has been reorganized into a modular structure with lazy imports, introducing new dataclasses for \ImageContent\ and \FileContent\ to support multimodal chat messages, and adding \ReasoningContent\ to \ChatMessage\ and \StreamingChunk\. \ToolCall\ and \ToolCallDelta\ now include an \extra\ field for provider-specific metadata, while \StreamingChunk\ gains \component\_info\ and \finish\_reason\ fields. \Document\ serialization now uses a deterministic ID generation that ignores metadata key order, and \ByteStream\ supports MIME type guessing and serialization. Pipeline debugging features (\Breakpoint\, \PipelineSnapshot\, \PipelineState\) are now part of the core dataclasses, and legacy fields like \dataframe\ have been removed from \Document\.
haystack/dataclasses · high confidence
Haystack package restructured with new core entry points and logging
The Haystack package has been reorganized, introducing a new \haystack/\_\init\\_.py\ that exposes core components like \Pipeline\, \Document\, \Answer\, and \SuperComponent\ directly at the top level, alongside a new \haystack.errors.FilterError\. A dedicated \haystack.logging\ module now handles logging configuration, including support for structured logging via \structlog\ and environment-variable-based JSON formatting, while \haystack.lazy\_imports\ provides a new context manager for handling optional dependency import errors. The package also includes a \py.typed\ marker file for type checking support and a new \haystack/version.py\ that dynamically reads the version from package metadata.
haystack · high confidence
Human-in-the-Loop confirmation hook moved to the Hooks module
The Human-in-the-Loop (HITL) functionality has been reorganized into the \haystack.hooks.human\_in\_the\_loop\ package, making it accessible as a first-class hook for Agents. This change introduces the \ConfirmationHook\ (restricted to the \before\_tool\ hook point) which allows users to intercept pending tool calls and apply confirmation strategies. The module provides built-in policies (\AlwaysAskPolicy\, \NeverAskPolicy\, \AskOncePolicy\) and strategies (\BlockingConfirmationStrategy\) to control when and how users are prompted. It also includes user interface implementations (\RichConsoleUI\, \SimpleConsoleUI\) for interactive confirmation, rejection, or parameter modification, along with dataclasses (\ToolExecutionDecision\, \ConfirmationUIResult\) to represent the outcomes of these interactions.
_haystack/hooks/human\_in\_the\loop · high confidence
In-memory retrievers now support configurable filter policies
The InMemoryBM25Retriever and InMemoryEmbeddingRetriever components now include a \filter\_policy\ parameter (defaulting to \REPLACE\) that controls how runtime filters interact with initialization filters. Users can set this to \MERGE\ to combine both sets of filters, or keep the default \REPLACE\ behavior to override initialization filters with runtime ones. This change ensures consistent filter application logic across in-memory retrieval operations.
_haystack/components/retrievers/in\memory · high confidence
Introduce pinned, multi-platform base Docker image for Haystack 2.x
The Docker build system has been restructured to provide a new \haystack:base\ image that users can derive their own images from. This base image is built using \docker buildx bake\, supports multi-platform builds (linux/amd64 and linux/arm64), and pins the Python 3.12-slim base image and the \uv\ installation tool by SHA-256 digest for reproducible, supply-chain-safe builds. The image includes a pre-installed virtual environment with Haystack installed from source, and explicitly upgrades \setuptools\ to address [CVE redacted].
docker · high confidence
Introduce pluggable tracing architecture with logging tracer and content tracing controls
The tracing subsystem has been refactored to support pluggable tracer implementations, exposing a core \Tracer\ and \Span\ interface alongside a new \LoggingTracer\ that outputs span details to the application logs. Users can now enable or disable tracing globally via \enable\_tracing\ and \disable\_tracing\, and control the inclusion of sensitive content in traces using the \HAYSTACK\_CONTENT\_TRACING\_ENABLED\ environment variable. The system also includes utilities to coerce complex tag values into serializable formats and prevents large payloads by using string placeholders for big objects.
haystack/tracing · high confidence
New API Reference homepage for Haystack
A new landing page for the API documentation has been added to the reference section. This page serves as the entry point for technical reference material, clearly separating the core framework API (covering base components, pipelines, document stores, and utilities) from the Integrations API (covering official packages that connect Haystack to external services).
docs-website/reference · high confidence
New and updated document store documentation for Haystack 2.26
The documentation for the \docs-website/versioned\_docs/version-2.26/document-stores\ area has been updated to include new guides for ArcadeDB, Astra, Azure AI Search, FAISS, InMemory, MongoDB Atlas, Pgvector, Pinecone, Qdrant, Valkey, and Weaviate document stores. The OpenSearch guide has been updated to reflect version 3.5.0, and the Elasticsearch guide now specifies support for version 8.11.1. These changes provide users with installation instructions, initialization examples, and supported retriever details for these specific integrations.
_docs-website/versioned\docs/version-2.26/document-stores · high confidence
New search modal with product-aware content filtering
The documentation site now features a new search modal (replacing the previous inline search bar) that supports filtering by 'All', 'Documentation', and 'API Reference'. The search experience is product-aware: it detects whether the user is viewing the Professional or Enterprise documentation (via URL path) and automatically hides Table of Contents items and search results that do not belong to the current product context. The modal also includes UI improvements such as branded styling, single-word sidebar link truncation, and a 'Powered by Haystack' attribution.
docs-website/src/theme · high confidence
Promote Haystack 2.21 documentation for Agents, Components, and Data Classes
The documentation for Haystack version 2.21 has been promoted from unstable to stable, making the current concepts officially available for this release. This update introduces new reference pages for the Agent \State\ system, which provides schema-based shared storage for tool execution, and \ChatMessage\, the central abstraction for LLM messages supporting text, images, tool calls, and reasoning content. It also adds guides for creating custom components and wrapping pipelines as \SuperComponent\s, alongside updated content for existing data classes.
_docs-website/versioned\_docs/version-2.21/concepts, docs-website/versioned\_docs/version-2.21/concepts/agents, docs-website/versioned\_docs/version-2.21/concepts/components, docs-website/versioned\docs/version-2.21/concepts/data-classes · high confidence
Promote Haystack 2.26 documentation to stable
The documentation for Haystack version 2.26 has been promoted from unstable to stable, making the current release notes and component guides the default reference for users. This update includes new documentation pages for the \Agent\ component (covering tool usage, state management, and multi-agent patterns), Human-in-the-Loop (HITL) capabilities for intercepting and confirming tool calls, and audio processing components (\LocalWhisperTranscriber\ and \RemoteWhisperTranscriber\). Additionally, the docs now cover the \AnswerBuilder\, \ChatPromptBuilder\, and \PromptBuilder\ components with updated usage examples, alongside guides for classifiers, caching, and various external integrations.
_docs-website/versioned\_docs/version-2.22/pipeline-components, docs-website/versioned\docs/version-2.26/pipeline-components · high confidence
Promoted Haystack 2.31 documentation with new agent and component guides
The documentation for Haystack version 2.31 has been promoted to stable, introducing comprehensive guides for building AI agents and multi-agent systems, including coordinator/specialist architectures. New concept pages explain core building blocks like Components, SuperComponents, and Data Classes (such as ChatMessage and FileContent), while new templates standardize component documentation. The release also updates installation instructions to use the \weave-haystack\ package and removes documentation for components deprecated in version 3.
_docs-website/versioned\docs/version-2.31 · high confidence
Promotes Haystack 3.1 reference documentation
The versioned API reference for Haystack 3.1 is now live on the Docusaurus site, featuring newly added documentation for core components such as Agents (with tool usage, hooks, and prompt templates), Builders (AnswerBuilder and ChatPromptBuilder), Caching, Converters, Data Classes, Document Stores, Document Writers, Embedders, Evaluation, and Evaluators. This update ensures the public documentation aligns with the stable 3.1 release, providing users with accurate usage examples and parameter details for these components.
_docs-website/reference\_versioned\docs/version-3.1 · high confidence
PromptBuilder and ChatPromptBuilder now require all template variables by default
PromptBuilder and ChatPromptBuilder now default \required\variables\ to \"\"\, meaning every variable found in the template is treated as required. Previously, variables were optional by default, which could lead to silent failures or unintended behavior in multi-branch pipelines when variables were missing. Users must now explicitly provide all template variables or set \required\_variables\ to \None\ or a specific list to make variables optional.
haystack/components/builders · high confidence
Updated documentation site styling with Inter font and new color palette
The documentation website now uses the Inter font family for all text and headings, replacing the previous default. A new blue-based color palette has been applied across light and dark modes, affecting primary/secondary colors, admonition backgrounds and borders, and neutral UI components like buttons and breadcrumbs. The active state for navbar items (Docs and API Reference) is now highlighted in the primary blue color with increased font weight, and breadcrumb links follow a similar active-state styling convention.
docs-website/src/css · high confidence
Updated introduction page with rebranded messaging and enterprise promotion
The version 2.21 documentation now includes a new introduction page that rebrands Haystack as an open-source AI framework for building production-ready AI Agents, RAG applications, and multimodal search systems. The content has been updated to reflect the rebranding, correct typos, and promote the Haystack Enterprise Starter and Platform for teams seeking enterprise-grade support and scalable tooling.
_docs-website/versioned\docs/version-2.21 · high confidence
Test coverage
1 commit adding/updating tests in test/test\_files/msg; 5 commits adding/updating tests in test/test\_files/docx; Added HTML test fixtures for document extraction; Added comprehensive test coverage for core dataclasses; Added comprehensive test coverage for retriever components; Added comprehensive test suite for chat generator components; Added comprehensive test suite for the Agent component; Added comprehensive test suites for AnswerBuilder, ChatPromptBuilder, and PromptBuilder; Added comprehensive tests for embedder components; Added end-to-end tests for RAG evaluation and PDF content extraction pipelines; Added fuzz testing for untrusted input entry points; Added integration tests for pipeline breakpoint resumption across various joiner and loop patterns; Added sample CSV test fixtures for CSVToDocument; Added sample components for testing; Added test coverage for Agent Hooks functionality; Added test coverage for LLM and Regex text extractors; Added test coverage for LLMRanker, LostInTheMiddleRanker, MetaFieldGroupingRanker, and MetaFieldRanker; Added test coverage for core serialization, security, and type utilities; Added test coverage for document preprocessing components; Added test coverage for image converter components; Added test coverage for joiner components; Added test coverage for new tooling components and utilities; Added test coverage for router components; Added test fixture files for FileTypeRouter; Added test helper for callable deserialization; Added test suite for context compaction hooks and compactors; Added test suite for evaluators in test/components/evaluators; Added tests for CacheChecker serialization, synchronous, and asynchronous operations; Added tests for DocumentWriter serialization, execution, and lifecycle; Added tests for EvaluationRunResult validation and reporting; Added tests for FileSystemSkillStore and frontmatter parsing; Added tests for Human-in-the-Loop hook components; Added tests for InMemoryDocumentStore and filter policy merging; Added tests for JsonSchemaValidator; Added tests for LLMDocumentContentExtractor; Added tests for LinkContentFetcher component; Added tests for SkillsToolset functionality and multimodal file support; Added tests for SuperComponent and type compatibility utilities; Added tests for TokenBudgetHook behavior; Added tests for TopPSampler validation and min\_top\_k behavior; Added tests for YAML marshalling and type restrictions; Added tests for component and document store factory utilities; Added tests for core component validation and socket behavior; Added tests for core sample components; Added tests for generator components and utilities; Added tests for the tracing subsystem; Added tests for token counter implementations; Added tests for tool result offloading hooks and stores; Added unit tests for the QueryExpander component; Added unit tests for utility functions in test/utils; Centralized behavioral tests for Pipeline execution; Expanded test coverage for core pipeline execution and lifecycle; Expanded test coverage for document converters and utility components; New test infrastructure and guidelines for the test suite; New testing utilities and fixtures for document stores and telemetry; Standardized e2e test environment and telemetry handling.
Dependencies
Migrate build system to pyproject.toml and update documentation site dependencies
The project has replaced the legacy requirements.txt with a modern pyproject.toml configuration, removing outdated dependencies such as FARM, Flask, and sklearn while introducing core dependencies like httpx, jsonschema, and openai\>=2.6.0. The documentation website has been updated to use Docusaurus 3.10, React 19, and @vercel/node 13.0.1, with the sharp image library upgraded to 0.35.0.
(dependencies) · high confidence
Housekeeping
Add Apache 2.0 license header to document stores module; Haystack 3.1 documentation release.
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 61.
Lenses
- Code Health 64
- Architecture 99
- Maturity 76
- Readiness 53
- Security 81
- Accessibility 61
Changes since last survey
- 300 commits — 236 feature/other, 64 fixes
By area
- docs-website/reference_versioned_docs — 77 commits
- haystack/components — 50 commits
- docs-website/docs — 37 commits
- docs-website/reference — 29 commits
- test/components — 23 commits
- (root) — 15 commits
- .github/workflows — 14 commits
- haystack/hooks — 11 commits
- docs-website/versioned_docs — 8 commits
- haystack/core — 7 commits
- docs-website/package.json — 5 commits
- haystack/utils — 4 commits
- test/hooks — 3 commits
- .github/utils — 2 commits
- test/core — 2 commits
- test/tools — 2 commits
- (repo) — 1 commit
- docs-website/docusaurus.config.js — 1 commit
- docs-website/versioned_sidebars — 1 commit
- e2e/conftest.py — 1 commit
Notable commits
- fix: ci: fix docs search sync for 3.x (#12402)
- fix: fix!: require unsafe loading for serialized Jinja filters (#12415)
- fix: fix: AzureOpenAIChatGenerator.to_dict() crashes when response_format is a plain dict (#12407)
- fix: fix: Better HITL tool execution decision and tool call consistency (#12345)
- fix: fix: Fix sliding window compactor to also retain historical user-assistant turns if the budget allows (#12270)
- fix: fix: LLMMetadataExtractor treats an extracted 'error' key as a failure (#12685)
- fix: fix: MarkdownHeaderSplitter drops the leading header of an unsplit document (#12478)
- fix: fix: MarkdownHeaderSplitter page_number drifts upwards when split_overlap is used (#12619)
- fix: fix: accept zero hierarchy meta values in AutoMergingRetriever (#12635)
- fix: fix: add standard split metadata when DocumentSplitter uses split_by="function" (#12505)
- fix: fix: allow FilterRetriever to clear default filters (#12629)
- fix: fix: allow OpenAI generators to clear default tools (#12639)
- fix: fix: allow top-level JSON scalars in JsonSchemaValidator (#12522)
- fix: fix: apply overlap only once in RecursiveDocumentSplitter (#12284)
- fix: fix: avoid duplicating overlap when DocumentSplitter merges a below-threshold segment (#12366)
- fix: fix: avoid mutating filters during FilterPolicy.MERGE (#12501)
- fix: fix: avoid overlap-only trailing chunks in DocumentSplitter token mode (#12661)
- fix: fix: clear stale metadata extraction errors after successful retries (#12533)
- fix: fix: convert images off the event loop in LLMDocumentContentExtractor.run_async (#12359)
- fix: fix: deep-copy document metadata in PythonCodeSplitter (#12424)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
deepset-ai/haystack was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit b717d00654af0dc405e5cec35546d734c832042f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.