microsoft/graphrag
65.0
Weak · 18 September 2026
51.7k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a modular GraphRAG (Retrieval-Augmented Generation) framework that indexes unstructured documents into knowledge graphs and supports multiple search modes, including local, global, and drift search. It provides a unified architecture for document ingestion, text chunking, and LLM interactions, backed by pluggable storage and vector store backends. The system exposes its capabilities through a public Python API and a command-line interface, allowing users to build indexes, tune prompts, and query data via a web application.
How it got here
2024 — monorepo restructuring and legacy removal
66 changes.
The project underwent a major architectural overhaul, migrating from a monolithic package to a multi-package monorepo managed by uv. This period was defined by the extensive removal of legacy indexing, query, LLM, and configuration modules to clear the way for a new system design. Supporting changes included updating the documentation site to MkDocs and modernizing the test infrastructure.
2025–2026 — GraphRAG package modularization and testing
21 changes.
This period focused on structuring the GraphRAG project into distinct, reusable packages for core functionalities such as input ingestion, LLM interaction, storage, and vector management. Concurrently, extensive unit and integration tests were added to validate these new components, including streaming operations, caching factories, and graph algorithms. The work also introduced a unified search application and refined the storage layer with a new Table Provider abstraction for efficient data handling.
Features
Added example notebooks for graphrag-input and graphrag-storage packages
New example notebooks have been added to demonstrate usage of the graphrag-input and graphrag-storage packages. The graphrag-input example shows how to configure input readers for CSV and MarkItDown sources, including handling nested JSON properties, while the graphrag-storage examples demonstrate basic file storage operations and how to register and use custom storage implementations via the factory pattern.
_packages/graphrag-input/example\_notebooks, packages/graphrag-storage/example\notebooks · high confidence
GraphRAG Cache package introduces factory-based caching with custom implementation support
The \graphrag-cache\ package now provides a factory-based caching system that allows users to create and register custom cache implementations. The package includes built-in cache types: \JsonCache\, \MemoryCache\, and \NoopCache\. Users can extend the base \Cache\ class to create custom caches and register them using the \register\_cache\ function. The \create\_cache\ function instantiates caches based on configuration, supporting both built-in and custom types. Example notebooks demonstrate basic usage with JSON cache and custom cache implementation.
packages/graphrag-cache · high confidence
GraphRAG package initialization with public API and CLI
The \packages/graphrag\ directory is introduced, establishing the core GraphRAG product. This includes a public Python API (\graphrag.api\) exposing functions for building indexes (\build\_index\), prompt tuning (\generate\_indexing\_prompts\), and multiple search modes (global, local, drift, basic, each with streaming variants). A command-line interface (CLI) is added via \typer\ with subcommands for project initialization (\init\), indexing (\index\), prompt tuning (\prompt-tune\), and querying (\query\). The package also contains the underlying infrastructure for these features, including configuration loading, workflow callbacks for monitoring pipeline progress, and cache key management.
packages/graphrag · high confidence
Introduce GraphRAG LLM package with completion, embedding, and configuration modules
The new graphrag-llm package provides a unified interface for language model completions and text embeddings, primarily targeting Azure OpenAI via LiteLLM. Users can now create completion and embedding clients using a standardized ModelConfig that supports API key and Azure Managed Identity authentication, along with built-in support for rate limiting, retry strategies, metrics collection, and caching. The package includes example notebooks for basic completion and embedding workflows, and enforces client-side JSON schema validation for structured responses.
packages/graphrag-llm · high confidence
Introduce GraphRAG unified search application
Adds a new Streamlit-based web application that provides a unified interface for querying graph-indexed data. The app supports four distinct search modes—Basic RAG, Local Search, Global Search, and Drift Search—which users can toggle on or off via the sidebar. It handles dataset selection and configuration loading from either local storage or Azure Blob Storage, and features a 'Community Explorer' tab to browse and inspect community reports with interactive graph citations.
unified-search-app/app · high confidence
Introduce dedicated graphrag-storage package with unified storage interface
The storage implementations previously located in graphrag/index/storage have been extracted into a new, standalone graphrag-storage package. This change introduces a unified Storage abstract base class and a StorageFactory that supports pluggable backends, including File, Memory, Azure Blob, and Azure Cosmos DB. The refactoring standardizes the storage API (e.g., find, get, set, has, delete, clear, child, keys) across all providers, simplifies configuration via a centralized StorageConfig model, and enables custom storage implementations to be registered dynamically.
_packages/graphrag-storage/graphrag\storage · high confidence
New Table Provider abstraction with native Cosmos DB, CSV, and Parquet backends
The storage layer now introduces a unified Table Provider interface that enables row-by-row streaming access to tables, replacing the previous bulk-only DataFrame approach. This change adds three native backend implementations: a native Cosmos DB provider that stores data as documents with namespace partitioning (eliminating the previous Parquet serialization round-trip), a CSV provider for file-based storage, and a Parquet provider that simulates streaming. The new architecture supports row transformers (allowing direct mapping to Pydantic models), safe concurrent writes via temporary files for CSV, and lazy container initialization for Cosmos DB, providing a consistent API for both bulk DataFrame operations and memory-efficient row iteration.
_packages/graphrag-storage/graphrag\storage/tables · high confidence
New graphrag-chunking and graphrag-common packages with example notebooks
This change introduces two new packages: graphrag-chunking, which provides text chunking capabilities (SentenceChunker, TokenChunker, and a configuration-driven factory), and graphrag-common, which offers a flexible dependency injection factory and a configuration loading system with Pydantic support. Both packages now include example Jupyter notebooks demonstrating their usage. Additionally, the NLTK bootstrap process has been updated to download the 'punkt\_tab' and 'averaged\_perceptron\_tagger\_eng' models alongside existing resources.
packages/graphrag-chunking, packages/graphrag-common · high confidence
New graphrag-input package for unified document ingestion
The new \graphrag-input\ package introduces a factory-based system for loading input documents, supporting CSV, JSON, JSON Lines, Parquet, plain text, and MarkItDown formats. It provides an \InputReader\ base class with async iteration capabilities and specific readers that map file types to \TextDocument\ objects, allowing users to configure input sources via \InputConfig\ and handle structured data columns or raw text uniformly.
_packages/graphrag-input/graphrag\input · high confidence
New graphrag-vectors package for unified vector store management
The \graphrag-vectors\ package introduces a unified interface for managing vector stores, supporting LanceDB, Azure AI Search, and Azure Cosmos DB. It provides a factory pattern for creating stores, a configuration-driven API via \create\_vector\_store\, and a programmatic \F\ filter builder for metadata filtering. The package also includes example notebooks demonstrating basic usage, custom store registration, and specific backend implementations.
packages/graphrag-vectors · high confidence
Removals
Removal of OpenAI LLM implementation module
The \graphrag/llm/openai\ module has been completely removed, deleting all associated files including client creation, LLM wrappers (chat, completion, embeddings), configuration, and JSON parsing utilities. This change eliminates the built-in OpenAI integration from this location, requiring users to rely on alternative LLM providers or external implementations.
graphrag/llm/openai · high confidence
Removal of custom indexing verb overrides
The custom override implementations for the aggregate, concat, and merge indexing verbs have been removed from the codebase. This change eliminates the local \aggregate\_override\, \concat\_override\, and \merge\_override\ functions, along with their supporting logic (such as the \MergeStrategyType\ enum and JSON merge helper), effectively reverting these specific data processing steps to rely on the standard underlying engine behavior or removing the capability entirely if no replacement is provided in the surrounding context.
graphrag/index/verbs/overrides · high confidence
Removal of graph layout verb and UMAP/zero layout strategies
The \layout\_graph\ verb and its associated layout strategies (UMAP and zero-position) have been removed from the indexing pipeline. This change eliminates the capability to automatically compute 2D node positions for graph visualization based on embeddings or structural properties, meaning users can no longer generate positioned graphs via this specific indexing step.
graphrag/index/verbs/graph/layout · high confidence
Removal of graph merge helper verbs
The graphrag indexing engine has removed the \merge\_graphs\ verb and its associated helper modules (including \defaults.py\ and \typing.py\) from the \graphrag/index/verbs/graph/merge\ location. This change eliminates the capability to merge multiple graphml-formatted graphs into a single graph using configurable node and edge operations (such as replace, skip, concat, sum, max, min, average, and multiply).
graphrag/index/verbs/graph/merge · high confidence
Removal of graphrag.query.input module
The \graphrag/query/input\ package has been removed from the codebase. This change eliminates the module that previously handled GraphRAG orchestration inputs, meaning any code relying on this specific input interface will no longer function and must be updated to use the new input handling mechanisms.
_examples\notebooks, graphrag/query/input · high confidence
Removal of indexing engine example code
The \examples\ directory has been completely removed, deleting all demonstration code for the indexing engine. This includes the main \README.md\ and all subdirectories (\custom\_input\, \custom\_set\_of\_available\_verbs\, \custom\_set\_of\_available\_workflows\, \entity\_extraction\, \interdependent\_workflows\, \multiple\_workflows\, \single\_verb\, \use\_built\_in\_workflows\, and \various\_levels\_of\_configs\). Users will no longer have access to these reference implementations for running pipelines via Python API or configuration files.
examples · high confidence
Removal of internal LLM base implementations
The \graphrag/llm/base\ module has been removed, deleting the \BaseLLM\, \CachingLLM\, and \RateLimitingLLM\ classes along with their cache key generation utilities. This change eliminates the internal base implementations for language model interactions, caching, and rate limiting, indicating a shift to an external or alternative LLM abstraction layer.
graphrag/llm/base · high confidence
Removal of internal LLM limiting and error modules
The \graphrag.llm\ package and its \limiting\ subpackage have been removed from the codebase. This change deletes the internal implementations for LLM rate limiting (including \TpmRpmLLMLimiter\, \CompositeLLMLimiter\, and \NoopLLMLimiter\), the \LLMLimiter\ interface, and the \RetriesExhaustedError\ class. These components are no longer part of the public or internal API surface provided by this location.
graphrag/llm, graphrag/llm/limiting · high confidence
Removal of internal LLM type definitions
The \graphrag/llm/types\ module has been removed, deleting internal type definitions and protocols such as \LLM\, \LLMCache\, \LLMConfig\, \LLMInput\, \LLMOutput\, and related callback and result types. This change aligns with the replacement of the \graphrag.llm\ implementation with \fnllm\, meaning these specific internal typing interfaces are no longer part of the codebase.
graphrag/llm/types · high confidence
Removal of legacy CSV and text input loaders
The \graphrag/index/input\ module has removed the legacy input loading implementations for CSV and text files (\csv.py\, \text.py\) and the central dispatcher (\load\_input.py\). This change eliminates the code responsible for parsing these specific file formats and routing them through the pipeline, indicating a shift away from these input methods in the current version.
graphrag/index/input · high confidence
Removal of legacy DataFrame-based input loaders
The legacy input loader module located at \graphrag/query/input/loaders\ has been removed. This deletion eliminates the previous mechanism for converting pandas DataFrames into GraphRAG model objects (such as Entities, Relationships, and Covariates) and discards the associated utility functions for column validation and type conversion. Users relying on these specific DataFrame-to-object mapping functions will need to adopt the updated input loading strategy provided by the current version of the library.
graphrag/query/input/loaders · high confidence
Removal of legacy GlobalSearch and LocalSearch implementations
The legacy GlobalSearch and LocalSearch modules, along with their associated system prompts, context builders, and callback handlers, have been removed from the codebase. This change eliminates the previous map-reduce orchestration pattern for global search and the single-turn approach for local search, indicating a migration to a new search architecture that replaces these deprecated components.
_graphrag/query/structured\_search/global\search · high confidence
Removal of legacy GraphRAG knowledge model package
The entire \graphrag.model\ package has been removed, deleting all previously defined data models including \Community\, \CommunityReport\, \Covariate\, \Document\, \Entity\, \Relationship\, and \TextUnit\, along with their base protocols (\Identified\, \Named\) and the \TextEmbedder\ type definition. This change eliminates the legacy datamodel classes and their dictionary deserialization logic (\from\_dict\ methods) that were previously used to represent the target data structures for pipelines and analytics tools.
graphrag/model · high confidence
Removal of legacy LLM loading utilities
The \graphrag/index/llm\ module has been removed, deleting the \load\_llm\ and \load\_llm\_embeddings\ functions along with their associated OpenAI and Azure-specific loader implementations. This change eliminates the internal logic for instantiating language models and embeddings directly within the indexing engine, indicating a shift in how LLM providers are integrated or managed elsewhere in the system.
graphrag/index/llm · high confidence
Removal of legacy OpenAI wrapper implementations
The OpenAI-specific LLM and embedding wrappers located in graphrag/query/llm/oai (including base classes, ChatOpenAI, OpenAI, OpenAIEmbedding, and typing definitions) have been removed. This cleanup eliminates the legacy orchestration layer that previously handled synchronous and asynchronous client creation, retry logic, and streaming token processing for OpenAI and Azure OpenAI services.
graphrag/query/llm/oai · high confidence
Removal of legacy Pydantic configuration models
The entire \graphrag/config/models\ directory has been removed, deleting all legacy Pydantic-based configuration classes (including \GraphRagConfig\, \LLMConfig\, \ChunkingConfig\, \EntityExtractionConfig\, and others). This change eliminates the previous configuration parameterization layer, indicating a migration to a new configuration schema or loading mechanism elsewhere in the codebase.
graphrag/config/models · high confidence
Removal of legacy configuration module
The legacy configuration module in graphrag/config has been removed, deleting the entire package including the config input models, Pydantic models, enums, default values, environment reader, and error classes. This cleanup eliminates the old configuration system that was previously used to load and validate GraphRag settings.
graphrag/config · high confidence
Removal of legacy context builder modules
The \graphrag/query/context\_builder\ package has been removed, deleting the legacy context-building infrastructure. This includes the \GlobalContextBuilder\ and \LocalContextBuilder\ abstract base classes, the \build\_community\_context\ function for generating community report context, the \ConversationHistory\ class for managing chat turns, entity extraction and neighbor-finding utilities, and the functions for building entity, relationship, covariate, and text-unit contexts. Users relying on these specific modules for constructing system prompts will need to adopt the new context-building approach.
_graphrag/query/context\builder · high confidence
Removal of legacy covariate extraction verbs
The \extract\_covariates\ verb and its associated implementation files (including the Graph Intelligence strategy, typing definitions, and defaults) have been removed from the indexing engine. This change eliminates the previous mechanism for extracting claims and covariates from text chunks, indicating that this functionality is no longer supported in its prior form within this code path.
graphrag/index/verbs/covariates · high confidence
Removal of legacy default configuration input models
The \graphrag/config/input\_models\ package has been removed, deleting all legacy TypedDict-based configuration classes (such as \GraphRagConfigInput\, \LLMConfigInput\, and specific step configs like \EntityExtractionConfigInput\). This change eliminates the previous default configuration parameterization layer, requiring users to adopt the new configuration schema or input methods that replace these deleted models.
_graphrag/config/input\models · high confidence
Removal of legacy entity extraction and description summarization verbs
The \entity\_extract\ and \summarize\_descriptions\ indexing verbs, along with their underlying implementation strategies (including \graph\_intelligence\ and \nltk\ for extraction, and \graph\_intelligence\ for summarization), have been removed from the \graphrag/index/verbs/entities\ package. This change eliminates the legacy code paths for extracting entities from text and summarizing entity/relationship descriptions, indicating that these capabilities are no longer supported in this location.
graphrag/index/verbs/entities · high confidence
Removal of legacy graph clustering implementation
The legacy graph clustering verb and its associated Leiden strategy implementation have been removed from the indexing engine. This deletion eliminates the \cluster\_graph\ verb, the \GraphCommunityStrategyType\ enum, and the \strategies/leiden.py\ module that previously handled hierarchical graph clustering using the Leiden algorithm. Users relying on this specific clustering workflow will need to adopt the new clustering infrastructure introduced in the same change set.
graphrag/index/verbs/graph/clustering · high confidence
Removal of legacy graph extraction and visualization components
The graph indexing module has removed the legacy implementation files for graph extraction, community report generation, and graph visualization. Specifically, the \graphrag/index/graph/embedding\, \graphrag/index/graph/extractors\ (including claim, community report, graph, and summarize extractors), and \graphrag/index/graph/visualization\ directories have been deleted. This change eliminates the old Node2Vec embedding logic, the previous prompt-based entity and claim extraction pipelines, the legacy community report context preparation and sorting utilities, and the UMAP-based graph visualization tools from the indexing engine.
graphrag/index/graph · high confidence
Removal of legacy graph indexing verbs
The \graphrag/index/verbs/graph\ package has been removed, eliminating legacy graph processing verbs including \create\_graph\, \unpack\_graph\, \compute\_edge\_combined\_degree\, and related clustering, embedding, layout, and community report helpers. Users relying on these specific graph manipulation steps in their indexing pipelines will need to adopt the new graph operations introduced in the current release.
graphrag/index/verbs, graphrag/index/verbs/graph · high confidence
Removal of legacy graph query input retrieval utilities
The \graphrag/query/input/retrieval\ module has been completely removed. This deletes the \\_\init\\_.py\ file and all associated utility modules (\community\_reports.py\, \covariates.py\, \entities.py\, \relationships.py\, and \text\_units.py\) that previously provided functions for retrieving and formatting graph data (such as entities, relationships, and text units) into pandas DataFrames for query inputs.
graphrag/query/input/retrieval · high confidence
Removal of legacy graph report extraction verbs
The \graphrag/index/verbs/graph/report\ module has been removed, eliminating the legacy \create\_community\_reports\ verb and its associated helper verbs (\prepare\_community\_reports\, \prepare\_community\_reports\_claims\, \prepare\_community\_reports\_edges\, \prepare\_community\_reports\_nodes\, and \restore\_community\_hierarchy\). This change also removes the \graph\_intelligence\ strategy implementation and its configuration defaults, effectively deprecating the previous community report generation pipeline in favor of the current indexing engine.
graphrag/index/verbs/graph/report · high confidence
Removal of legacy graphrag index utility modules
The \graphrag/index/utils\ package has been completely removed, deleting all associated utility modules including \\_\init\\_.py\, \dataframes.py\, \ds\_util.py\, \json.py\, \load\_graph.py\, \tokens.py\, and \topological\_sort.py\. This eliminates legacy helper functions for DataFrame manipulation, JSON cleaning, graph loading, token counting, and topological sorting that were previously exposed via the utils namespace.
graphrag/index/utils · high confidence
Removal of legacy graphrag.index package and CLI
The entire \graphrag/index\ package has been removed, including the legacy CLI entry points (\\_\main\\_.py\, \cli.py\), the old pipeline configuration system (\create\_pipeline\_config.py\, \load\_pipeline\_config.py\), and the associated infrastructure for caching, storage, reporting, and table emission. This cleanup eliminates the previous indexing engine implementation in favor of the new architecture.
graphrag/index · high confidence
Removal of legacy pipeline configuration models
The legacy Pydantic-based configuration models for the indexing pipeline have been removed. This includes the deletion of the \graphrag/index/config\ package and all its submodules (\cache\, \input\, \pipeline\, \reporting\, \storage\, \workflow\), which previously defined types such as \PipelineConfig\, \PipelineCacheConfig\, \PipelineInputConfig\, and \PipelineStorageConfig\. Users relying on these specific configuration structures for pipeline setup will need to migrate to the new configuration approach.
graphrag/index/config · high confidence
Removal of legacy prompt tuning templates
The prompt tuning module has removed the standalone Python files containing default prompts for generating personas, entity types, entity relationships, domains, and community reporter roles. This deletion of the \graphrag/prompt\_tune/prompt\ package contents indicates that these specific template definitions are no longer used or have been migrated to a different location or format within the system.
_graphrag/prompt\_tune/prompt, graphrag/prompt\tune/template · high confidence
Removal of legacy query CLI and orchestration module
The legacy command-line interface and orchestration entry points for the query module have been removed. This includes the deletion of the \graphrag.query\ package root, the \\_\main\\_.py\ script that previously handled local and global search execution via argparse, the \cli.py\ module containing the search runner logic, the \factories.py\ file responsible for instantiating search engines and LLM clients, the \indexer\_adapters.py\ for reading indexing outputs, and the \progress.py\ status reporter. Users relying on these specific CLI entry points or direct imports from the \graphrag.query\ package root will need to migrate to the updated query interfaces.
graphrag/query · high confidence
Removal of legacy question generation module
The \graphrag/query/question\_gen\ package, which previously provided the \BaseQuestionGen\, \LocalQuestionGen\, and associated system prompts for generating follow-up questions, has been completely removed. This deletion eliminates the legacy synchronous and asynchronous question generation capabilities that relied on the \LocalContextBuilder\ and specific prompt templates, indicating a shift away from this implementation in the current query flow.
_graphrag/query/question\gen · high confidence
Removal of legacy structured search base classes
The legacy base classes and package initialization for the structured search module have been removed. This eliminates the abstract \BaseSearch\ interface and the \SearchResult\ data structure that previously defined the synchronous and asynchronous search contracts, indicating a shift away from the older search implementation architecture.
_graphrag/query/structured\search · high confidence
Removal of legacy text chunking strategies
The legacy text chunking implementation located in the \graphrag/index/verbs/text/chunk\ directory has been removed. This change eliminates the \tokens\ and \sentence\ chunking strategies, along with their associated configuration defaults (such as \DEFAULT\_CHUNK\_SIZE\ and \DEFAULT\_CHUNK\_OVERLAP\) and the \ChunkStrategy\ type definitions. Users relying on these specific chunking methods within this module will need to adopt the updated chunking approach provided by the new system.
graphrag/index/verbs/text/chunk · high confidence
Removal of legacy text embedding strategies and verb implementation
The \graphrag/index/verbs/text/embed\ module has been removed, eliminating the legacy \text\_embed\ verb and its associated strategy implementations (\openai\ and \mock\). This change removes the ability to generate text embeddings using the previous in-memory or vector-store-based embedding workflows defined in this location.
graphrag/index/verbs/text/embed · high confidence
Removal of legacy text splitting module
The \graphrag/index/text\_splitting\ package, which previously provided text splitting utilities such as \TokenTextSplitter\, \TextListSplitter\, and token-limit checking functions, has been completely removed from the codebase. This deletion eliminates the legacy text chunking logic that relied on \tiktoken\ and custom splitter classes, indicating a shift in how text processing is handled within the indexing engine.
_graphrag/index/text\splitting · high confidence
Removal of legacy v1 workflow definitions and loading infrastructure
The legacy v1 workflow system has been removed from the indexing engine. This change deletes the \default\_workflows.py\ file, which previously registered specific workflow steps (such as \create\_base\_entity\_graph\, \create\_final\_communities\, and \join\_text\_units\_to\_entity\_ids\), along with the \load.py\ module responsible for loading and executing these workflows via topological sorting, and the \typing.py\ module defining the associated data structures. Users relying on the previous v1 workflow registration and loading mechanism will need to migrate to the new workflow generation approach.
graphrag/index/workflows · high confidence
Removal of legacy vector store implementations
The \graphrag/vector\_stores\ package has been removed, eliminating the legacy \AzureAISearch\ and \LanceDBVectorStore\ implementations along with their base classes (\BaseVectorStore\, \VectorStoreDocument\, \VectorStoreSearchResult\) and factory utilities. Users relying on these specific storage backends will need to migrate to the new vector store architecture introduced in this release.
_graphrag/vector\stores · high confidence
Removal of mock LLM implementations
The mock LLM classes (MockChatLLM and MockCompletionLLM) previously available in the graphrag.llm.mock module have been removed. This change eliminates the ability to use these specific mock implementations for testing or development purposes within this package location.
graphrag/llm/mock · high confidence
Removal of prompt tuning configuration and input loaders
The \graphrag/prompt\tune/loader\ module has been removed, deleting the \config.py\, \input.py\, and \\\init\\_.py\ files. This eliminates the previous implementation for reading GraphRAG settings from YAML/JSON files and loading document chunks for prompt tuning, indicating a structural change to how configuration and data ingestion for prompt tuning are handled in the system.
_graphrag/prompt\tune/loader · high confidence
Removal of prompt tuning generator module
The prompt tuning generator module (graphrag/prompt\_tune/generator) has been removed. This eliminates the functionality for automatically generating prompts for community summarization, entity extraction, entity summarization, entity type identification, domain inference, persona creation, and entity relationship examples.
_graphrag/prompt\tune/generator · high confidence
Removal of text\_translate verb and translation strategies
The text\_translate verb, along with its underlying OpenAI and mock translation strategies, has been removed from the indexing engine. Users can no longer use the text\_translate verb to translate text columns into other languages via LLMs or mock implementations.
graphrag/index/verbs/text/translate · high confidence
Removal of v1 workflow definitions
The workflow definition files in the \graphrag/index/workflows/v1\ directory have been removed. This includes the base and final generation steps for documents, text units, entities, relationships, covariates, nodes, and community reports, as well as the helper workflows used to join text units to entity, relationship, and covariate IDs. These changes eliminate the previous v1 indexing pipeline structure from this location.
graphrag/index/workflows/v1 · high confidence
Behavioural changes
Automated dependency management and build script updates
This change introduces new automation scripts to manage workspace dependencies and GitHub operations, while updating existing build tooling. Specifically, it adds scripts to automatically close Dependabot pull requests, open issues for dependency sweeps, and update cross-package dependency versions within the workspace to match the current version. It also adds a script to copy build assets (like the LICENSE file) to package directories for PyPI distribution. Additionally, the semver-check script has been updated to use \uv run\ instead of \poetry run\, reflecting a shift in package management, and the legacy \e2e-test.sh\ script using Poetry has been removed.
scripts · high confidence
Migrate documentation site from Eleventy to MkDocs
The documentation site has switched its static site generator from Eleventy to MkDocs. This change removes the Eleventy configuration files (\.eleventy.js\, \.eleventyignore\, \.gitignore\) and the bundled Yarn binary, replacing them with a MkDocs-based build process. Users will now benefit from MkDocs' native support for static output directories and its integrated search capabilities, which are enabled for both local and global search within the documentation.
docsite · high confidence
Prompt tuning CLI refactored from argparse to Typer
The command-line interface for the prompt tuning feature has been migrated from the standard library \argparse\ to the \Typer\ framework. This change removes the legacy \\_\main\\_.py\ and \cli.py\ entry points, replacing them with a modern, structured CLI implementation that maintains the same functional capabilities for generating indexing prompts while improving command-line usability and code maintainability.
_graphrag/prompt\tune · high confidence
Removal of legacy LLM base classes and text utilities
The \graphrag/query/llm\ module has been cleaned up by removing the \\_\init\\_.py\, \base.py\, and \text\_utils.py\ files. This eliminates the previous abstract base classes for LLMs (\BaseLLM\, \BaseTextEmbedding\) and callbacks, as well as text processing utilities like \num\_tokens\ and \chunk\_text\ that relied on \tiktoken\. Users relying on these specific internal components within this directory will need to migrate to the new LLM provider implementations or utility functions provided elsewhere in the updated package structure.
graphrag/query/llm · high confidence
Test coverage
Added covariates test data for A Christmas Carol; Added integration tests for cache and logging factories; Added integration tests for language model components; Added integration tests for storage backends and factory; Added integration tests for vector store implementations; Added test infrastructure for GraphRAG LLM module; Added unit tests for BlobWorkflowLogger; Added unit tests for CSV and Parquet table providers; Added unit tests for LLM cache middleware event loop handling; Added unit tests for configuration validation and loading; Added unit tests for core GraphRAG components; Added unit tests for graph algorithms; Added unit tests for indexing input loaders and utilities; Added unit tests for indexing operations; Added unit tests for indexing operations and updated config test imports; Added unit tests for prompt tuning document loading; Added unit tests for query context building, entity retrieval, and vector store filtering; Added unit tests for text encoding utilities; Added unit tests for the new cache configuration system; Removal of legacy LLM test helpers; Removal of unit tests for blob and file pipeline storage; Removed indexing config unit tests and fixtures; Removed legacy integration pipeline test suite; Removed unit tests for text split verb; Smoke test infrastructure modernization and stabilization; Updated notebook test infrastructure to use uv and target specific package notebooks; Updated unit tests for community report context sorting; Updated unit tests for graph intelligence entity extraction; Updated unit tests for stable\_lcc to use local graph definitions.
Dependencies
GraphRAG restructured into a multi-package monorepo with uv package management
The GraphRAG project has been reorganized from a single monolithic package into a modular monorepo consisting of eight distinct packages: graphrag, graphrag-cache, graphrag-chunking, graphrag-common, graphrag-input, graphrag-llm, graphrag-storage, and graphrag-vectors, all versioned at 3.1.2. This structural change is accompanied by a migration from Poetry to uv for dependency management, evidenced by the removal of poetry.lock and the addition of \[tool.uv\] configuration in the root pyproject.toml. The unified search application has been updated to depend on graphrag 2.5.0, and the documentation site has switched from an Eleventy-based setup to MkDocs Material, removing the previous docsite dependencies.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 65.
Lenses
- Code Health 91
- Architecture 95
- Maturity 74
- Readiness 54
- Security 66
Changes since last survey
- 300 commits — 244 feature/other, 56 fixes
By area
- .semversioner/next-release — 61 commits
- graphrag/index — 60 commits
- (root) — 45 commits
- graphrag/config — 20 commits
- packages/graphrag — 16 commits
- graphrag/query — 13 commits
- tests/verbs — 8 commits
- .github/workflows — 6 commits
- docs/examples_notebooks — 6 commits
- tests/unit — 6 commits
- graphrag/language_model — 4 commits
- tests/fixtures — 4 commits
- docs/config — 3 commits
- docs/query — 3 commits
- docsite/posts — 3 commits
- graphrag/cli — 3 commits
- packages/graphrag-storage — 3 commits
- unified-search-app/.vsts-ci.yml — 3 commits
- .agents/skills — 2 commits
- docs/blog_posts.md — 2 commits
Notable commits
- fix: Cosmosdb communities bug (#2232)
- fix: Fix API key reference for gh-pages (#1821)
- fix: Fix .strip call (#2431)
- fix: Fix Community ID loading for DRIFT search over existing indexes (#1360)
- fix: Fix DRIFT search on Azure AI Search (#1645)
- fix: Fix JSONL loader handling of blank/invalid lines (#2434)
- fix: Fix StopAsyncIteration catch (#1730)
- fix: Fix broken documentation links. (#2305)
- fix: Fix cannot release un-acquired lock in Blob logger (#2457)
- fix: Fix content embedding container name (#1358)
- fix: Fix cookie consent script missing (#1292)
- fix: Fix deps (#2193)
- fix: Fix documentation for generate_indexing_prompts (#1336)
- fix: Fix drift search edge cases over small input sets (#1310)
- fix: Fix dynamic community selection in global search (#1450)
- fix: Fix encoding issue: Ensure non-ASCII characters are correctly represe… (#1446)
- fix: Fix file path issue in the viz guide (#1372)
- fix: Fix graph creation (#1905)
- fix: Fix id baseline (#2036)
- fix: Fix in load_llm.py (#1508)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
microsoft/graphrag was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 1877d7256b4db4d66f64c491a9b559ef2470d338 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.