FoundationAgents/MetaGPT
54.5
Adequate · 26 September 2026
55.6k
lines of production code
Python
primary language
4
measurements over time
What this system is
This release introduces a comprehensive structured configuration system and a new Context management layer, replacing the previous global state approach. It significantly expands the example suite with new implementations for RAG, Data Interpreter, experience pools, and various simulation environments like Minecraft and Stanford Town. The update also adds support for multiple LLM providers, including ZhiPu AI and AWS Bedrock, while introducing a new strategy module for planning and reasoning.
Features
Add AFlow benchmarking framework for evaluating model performance
Users can now evaluate model performance across multiple datasets (DROP, GSM8K, HotpotQA, HumanEval, MATH, MBPP) using a new benchmarking framework. The change introduces a \BaseBenchmark\ abstract class in \benchmark.py\ that standardizes evaluation logic, including data loading, result logging, and score calculation. Concrete implementations for each dataset provide specific scoring logic (e.g., F1 score for text, exact match for math/code). The \README.md\ provides instructions for creating custom benchmarks by inheriting from \BaseBenchmark\ and implementing abstract methods like \evaluate\_problem\ and \calculate\_score\. This enables consistent evaluation of model outputs against ground truth for various tasks.
metagpt/ext/aflow/benchmark · high confidence
Add AFlow example for agentic workflow generation
Added the AFlow example in the \examples/aflow\ directory, providing a complete setup for automatically generating and optimizing agentic workflows using Monte Carlo tree search. The change introduces a new \optimize.py\ entry point that configures datasets (HumanEval, MBPP, GSM8K, MATH, HotpotQA, DROP) and operators, along with a \config2.example.yaml\ template for LLM configuration. A \README.md\ was also added to document the framework components, datasets, and usage instructions.
examples/aflow · high confidence
Add AFlow optimizer utility modules
Added four new utility classes to the AFlow optimizer pipeline: ConvergenceUtils for detecting optimization convergence via statistical analysis of round scores, DataUtils for managing result data and score sampling, ExperienceUtils for tracking and formatting past optimization successes and failures, and GraphUtils for loading, generating, and writing graph and prompt files. These utilities support the iterative optimization workflow by providing evaluation, data handling, and graph management capabilities.
_metagpt/ext/aflow/scripts/optimizer\utils · high confidence
Add Android Assistant example script
A new \run\_assistant.py\ script has been added to the \examples/android\_assistant\ directory, providing a command-line interface to run the Android Assistant. This entry point allows users to configure and launch the assistant's learning and acting stages, managing the Android environment, device connection, and task execution through the \Team\ and \AndroidEnv\ classes.
_examples/android\assistant · high confidence
Add Chainlit-based UI example for MetaGPT
Introduced a new example in the \examples/ui\_with\_chainlit\ directory that integrates MetaGPT with the Chainlit framework. This adds a web-based user interface where users can input requirements (e.g., 'Create a 2048 game') and receive outputs such as user stories, competitive analysis, and code files directly in the UI. The implementation includes \app.py\ to handle chat profiles and message processing, \init\_setup.py\ to stream logs and role actions to the UI, and a \.gitignore\ file to exclude generated files. This allows users to interact with MetaGPT's software company simulation through a conversational UI rather than the command line.
_examples/ui\_with\chainlit · high confidence
Add Dev Container configuration for local development
Developers can now launch the project in a fully configured, containerized environment using VS Code Dev Containers or GitHub Codespaces. This includes a \devcontainer.json\ that pulls the \metagpt/metagpt:latest\ image, a \docker-compose.yaml\ for multi-container setups, and a \postCreateCommand.sh\ script that automatically installs Mermaid CLI and Python dependencies, ensuring a consistent and reproducible development experience.
.devcontainer · high confidence
Add InfiAgent-DABench example for Data Interpreter
Added a new example demonstrating how to use the Data Interpreter to solve the InfiAgent-DABench benchmark. This includes the core evaluation logic in DABench.py, along with scripts to run single tasks, all tasks serially, or all tasks in parallel. The example also provides a README with setup instructions and notes on configuring the notebook execution environment.
examples/di/InfiAgent-DABench · high confidence
Add Java code review rules for static analysis
Introduces a new code review (CR) module containing 25 Java-specific linting rules. These rules enforce best practices such as avoiding unused variables, preventing empty try/catch/finally blocks, using proper string comparison, and avoiding common pitfalls like using == for String comparison or printStackTrace() for logging. The module includes both English and Chinese JSON files defining the rules and examples.
metagpt/ext/cr · high confidence
Add OmniParse parser for document processing
A new \OmniParse\ parser class has been introduced in \metagpt/rag/parsers/omniparse.py\, enabling the system to parse PDF and other documents via an external OmniParse API. The implementation supports both single-file and concurrent multi-file processing, with configurable parse and result types, and is exported via the \metagpt.rag.parsers\ module.
metagpt/rag/parsers · high confidence
Add Pic2Txt action for converting requirement images to text
A new Pic2Txt action has been added to the requirement analysis module. This feature allows the system to process images depicting user requirements and contextual descriptions, converting them into complete textual user requirements. The implementation includes a new Python module (pic2txt.py) that handles image encoding, calls an LLM to generate text from images, and merges the results with existing textual requirements, evaluation conclusions, and technical specifications.
_metagpt/actions/requirement\analysis/requirement · high confidence
Add SPO (Self-Supervised Prompt Optimization) example with CLI and web interface
Users can now access the SPO (Self-Supervised Prompt Optimization) example, which provides an automated prompt engineering tool for LLMs. The addition includes a comprehensive README with usage instructions, a CLI entry point (optimize.py) for running the optimizer via command line, and a Streamlit-based web interface for a more user-friendly experience. Configuration is managed through a new config2.example.yaml file, allowing users to easily set up API keys and model parameters for optimization, evaluation, and execution.
examples/spo · high confidence
Add SPO utility modules for LLM, evaluation, and data handling
Added new utility modules in metagpt/ext/spo/utils to support the SPO extension. This includes llm\_client.py for managing LLM instances and routing requests, evaluation\_utils.py for executing prompts and evaluating results, data\_utils.py for loading and saving JSON results, prompt\_utils.py for managing prompt files and directories, and load.py for reading YAML configuration files. These utilities provide the core infrastructure for the SPO workflow.
metagpt/ext/spo/utils · high confidence
Add SWE Agent command-line tools for code editing and search
Added a new directory \metagpt/tools/swe\_agent\_commands\ containing shell scripts and Python utilities that provide commands for opening files, editing code with linting checks, searching directories and files, and extracting patches. These tools, adapted from the SWE-agent and OpenDevin projects, enable the agent to navigate, view, and modify source code while enforcing PEP8 standards and syntax validation.
_metagpt/tools/swe\_agent\commands · high confidence
Add Sela experiment scripts and visualization tool
Added three new shell scripts (run\_cls.sh, run\_cls\_mod.sh, run\_reg.sh) that automate the execution of classification, modified classification, and regression experiments using the MCTS mode, alongside a Python script (visualize\_experiment.py) for visualizing the MCTS search tree as a graph.
metagpt/ext/sela/scripts · high confidence
Add Stanford Town environment with action/observation spaces
The Stanford Town environment is introduced with new files defining the action and observation spaces (EnvAction, EnvObsParams, etc.) and the environment implementation (StanfordTownEnv, StanfordTownExtEnv). This provides the structural foundation for interacting with the Stanford Town simulation, including tile-based state representation and event handling.
_metagpt/environment/stanford\town · high confidence
Add ZhiPu AI provider implementation
Introduced a new provider for the ZhiPu AI (ZhiPu) model, including the main module, an async SSE client for streaming responses, and a model API class that supports both synchronous and asynchronous invocation as well as server-sent events (SSE) streaming.
metagpt/provider/zhipuai · high confidence
Add evaluation tools for Werewolf game logs
Added \eval.py\ and \utils.py\ in the \examples/werewolf\_game/evals\ directory to evaluate game performance. The new \Vote\ class parses game logs to calculate the 'good team vote rate' (the percentage of non-werewolf players correctly voting against werewolves) and the 'votewolf\_difficulty' (the ratio of living werewolves to total living players). The \Utils\ class provides helper functions to polish raw logs, extract specific vote data, and format results into DataFrames for analysis.
_examples/werewolf\game/evals · high confidence
Add experience pool example demonstrating caching, initialization, and scoring
The examples/exp\_pool directory now includes a complete set of example scripts showing how to use the experience pool feature. Users can see how to cache function outputs with the @exp\_cache decorator, initialize the pool with sample data, load experiences from logs, manage CRUD operations, and evaluate responses using the SimpleScorer. A README.md provides configuration and usage instructions for these examples.
_examples/exp\pool · high confidence
Add new AI and search tool implementations and a tool registry system
The \metagpt/tools\ directory was expanded with several new tool implementations: Azure TTS, iFLYTEK TTS, text-to-image generation (MetaGPT and OpenAI), text-to-embedding (OpenAI), moderation, and search engines (Bing, DuckDuckGo, Google API, Serper). Additionally, a \tool\_registry.py\ was added to manage and register these tools, along with supporting modules like \tool\_convert.py\ for schema generation and \tool\_recommend.py\ for tool recommendation logic.
metagpt/tools · high confidence
Add object-based ranking capability to RAG pipeline
The RAG module now supports sorting retrieved nodes by a specific object field. A new \ObjectSortPostprocessor\ has been added to the \metagpt.rag.rankers\ package, allowing users to sort results in ascending or descending order based on a JSON field within the node's metadata. This introduces a new behavioral option for post-processing nodes in the retrieval pipeline.
metagpt/rag/rankers · high confidence
Add serializers for the experience pool
A new serializers module has been added to the experience pool, introducing a \BaseSerializer\ interface and concrete implementations (\SimpleSerializer\, \ActionNodeSerializer\, \RoleZeroSerializer\) to handle the serialization and deserialization of experience data. This enables the system to store and retrieve experiences in a standardized format, with specific handling for \ActionNode\ objects and filtered request data for \RoleZero\ scenarios.
_metagpt/exp\pool/serializers · high confidence
Add tree-search algorithms and node management for the Sela experimenter
Added new search algorithms (Greedy, Random, and MCTS) that extend the BaseTreeSearch class to determine the best child node during the search process. The implementation includes a Node class for managing tree state, depth, and ID generation, along with utility functions to initialize the decision tree root and create initial states for tasks. This provides the Sela experimenter with configurable search strategies for exploring the solution space.
metagpt/ext/sela/search · high confidence
Add utility modules for accessibility, HTTP, async execution, cost tracking, and graph repositories
Added several new utility modules in the \metagpt/utils\ directory to support new features and improve code organization. This includes \a11y\_tree.py\ for generating and interacting with browser accessibility trees via Playwright, \ahttp\_client.py\ for async HTTP POST requests, and \async\_helper.py\ for managing separate asyncio event loops. A new \cost\_manager.py\ provides a \CostManager\ class to track and limit API usage costs, while \custom\_decoder.py\ implements a custom JSON decoder to handle non-standard JSON formats. Additionally, \dependency\_file.py\ manages file dependencies, and \di\_graph\_repository.py\ implements a directed graph repository using NetworkX for storing and querying triple-based data.
metagpt/utils · high confidence
Add werewolf game example with human player support
A new example script for the Werewolf game environment is introduced, providing a complete setup for running the game with configurable roles (Villager, Werewolf, Guard, Seer, Witch, Moderator) and human player integration. The script initializes the game environment, hires agents, and runs the game loop, supporting features like reflection, experience, and memory selection.
_examples/werewolf\game · high confidence
Added Android environment and observation/action space definitions
Introduced the Android environment implementation, including the \AndroidEnv\ and \AndroidExtEnv\ classes that handle device interactions via ADB, along with supporting modules for observation and action spaces (\env\_space.py\), constants (\const.py\), and computer vision utilities for text and icon localization (\text\_icon\_localization.py\).
metagpt/environment/android · high confidence
Added Stanford Town simulation example with bootstrap data
The \examples/stanford\_town\ directory now contains a runnable example for the Stanford Town simulation. This includes the \run\_st\game.py\ entry point script, an \\\init\\_.py\ file, and a \storage\ directory containing the initial state for the simulation. The storage data includes environment maps, spatial memory, and associative memory for three personas (Isabella Rodriguez, Klaus Mueller, and Maria Lopez), providing the necessary bootstrap data to run the simulation.
_examples/stanford\town · high confidence
Added Stanford Town simulation integration with multi-agent interaction and reflection actions
The \metagpt/ext/stanford\_town\ directory has been added, introducing a new simulation environment for the Stanford Town project. This includes a suite of actions that enable agents to generate daily and hourly schedules, plan and decompose tasks, and engage in iterative conversations with other agents. The implementation also provides mechanisms for agents to reflect on their own thoughts and the world around them, including summarizing conversations, focusing on key points, and generating event triples and poignancy scores. These components allow MetaGPT to run complex, multi-agent simulations where agents interact, plan, and reflect in a virtual environment.
_metagpt/ext/stanford\town · high confidence
Added baseline runner implementations for SELA
The SELA module now includes a new \runner\ subpackage that provides concrete execution logic for various machine learning baselines. This includes a base \Runner\ class and specific implementations for AIDE, AutoGluon, AutoSklearn, and a custom baseline runner. Additionally, search algorithm runners for MCTS, Random Search, and a base Data Interpreter runner have been added to support automated machine learning workflows.
metagpt/ext/sela/runner · high confidence
Added default prompts for document relevance evaluation
Introduced a new module for default prompts, specifically adding a prompt template for evaluating document relevance to a given question. This includes a system prompt that instructs an assistant to rank documents based on their relevance to a query, outputting a list of document numbers and relevance scores. The implementation uses LlamaIndex's PromptTemplate and PromptType to structure this new capability.
metagpt/rag/prompts · high confidence
Added prompt templates for Aflow workflow optimization and operator execution
New prompt templates have been introduced to support the Aflow workflow's optimization and execution logic. The \optimize\_prompt.py\ file adds prompts for reconstructing and optimizing graph structures and prompts, including instructions for incorporating critical thinking methods, control flow, and machine learning techniques. The \prompt.py\ file introduces a suite of prompts for specific operators, including answer generation, format extraction, ensemble voting (SC and MD), Python code verification, reflection on failed tests, and review/revision of solutions. These templates define the expected input and output formats for the Aflow system's internal processes.
metagpt/ext/aflow/scripts/prompts · high confidence
Added prompt templates for SPO evaluation and optimization
New prompt templates have been introduced for the SPO (Self-Play Optimization) workflow. The system now includes an evaluation prompt that compares two responses against a requirement and a golden answer, outputting an analysis and a choice using XML tags. Additionally, an optimization prompt has been added to reconstruct and improve existing prompts based on requirements, execution results, and expected answers, also utilizing XML tags for structured output.
metagpt/ext/spo/prompts · high confidence
Added requirement analysis framework for generating and evaluating software frameworks
Introduced new components in the requirement analysis module to support RFC243. This includes an \EvaluateAction\ for assessing requirements, a \WriteFramework\ action that generates software framework code from a Technical Requirements Document (TRD), and an \EvaluateFramework\ action that evaluates the quality of the generated framework. Additionally, a \save\_framework\ utility was added to persist the generated framework files to the workspace.
_metagpt/actions/requirement\analysis/framework · high confidence
Added script to download and extract Aflow datasets
A new Python script, download\_data.py, has been added to the aflow data directory. This script provides a convenient way for users to download and extract the Aflow dataset, results, and initial rounds data from Google Drive, automatically handling file extraction and cleanup.
metagpt/ext/aflow/data · high confidence
Expanded Data Interpreter examples for diverse use cases
The Data Interpreter (DI) examples have been significantly expanded to demonstrate a wider range of capabilities. New examples include automated task planning and execution via the TeamLeader, custom tool integration, and specific workflows for data visualization, machine learning modeling, OCR-based receipt processing, email summarization, and GitHub issue fixing. The update also introduces examples for web crawling, webpage imitation, and interactive human-in-the-loop scenarios, providing comprehensive templates for users to build their own DI agents.
examples/di · high confidence
Expanded LLM provider support with new integrations
The provider module now supports a wider range of large language models and cloud services. New integrations include Anthropic (Claude), Amazon Bedrock, Volcengine Ark, Azure OpenAI, Google Gemini, and Alibaba DashScope. The update also introduces a generic HTTP-based API handler for non-standard endpoints and adds support for reasoning content in model responses.
metagpt/provider · high confidence
Expanded example suite with new agent, multi-agent, and tool-use demos
The \examples\ directory was significantly expanded with new demonstration scripts: \agent\_creator.py\ and \build\_customized\_agent.py\ showcase creating and running custom agents; \build\_customized\_multi\_agents.py\ demonstrates a multi-agent team (coder, tester, reviewer); \cr.py\ provides a code review example; \dalle\_gpt4v\_agent.py\ shows image generation and refinement; \debate.py\ and \debate\_simple.py\ illustrate head-to-head roleplay; \hello\_world.py\ and \ping.py\ (renamed from \llm\_hello\_world.py\) cover basic LLM interactions; \invoice\_ocr.py\ and \llm\_vision.py\ demonstrate OCR and vision capabilities; \mgx\_write\_project\_framework.py\ and \write\_design.py\ show project framework and design generation; \research.py\ covers research tasks; \search\_enhanced\_qa.py\ and \search\_with\_specific\_engine.py\ demonstrate search-enhanced QA; \serialize\_model.py\ shows environment serialization; \stream\_output\_via\_api.py\ enables streaming via Flask; \use\_off\_the\_shelf\_agent.py\ uses pre-built roles; \write\_novel.py\ and \write\_tutorial.py\ cover creative writing and tutorials. Additionally, \azure\_hello\_world.py\ and \search\_kb.py\ were removed, and \llm\_hello\_world.py\ was renamed to \ping.py\.
examples · high confidence
Introduce AFLOW inference interface and evaluation scripts
Adds a new AFLOW inference interface (\interface.py\) that loads optimized workflows and executes them for datasets like HumanEval, MBPP, GSM8K, MATH, HotpotQA, and DROP. The change also introduces supporting scripts for evaluation (\evaluator.py\), operator logic (\operator.py\, \operator\_an.py\), graph optimization (\optimizer.py\), and utility functions (\utils.py\), enabling users to run AFLOW-based reasoning and code generation tasks.
metagpt/ext/aflow/scripts · high confidence
Introduce Android Assistant for learning and executing smartphone tasks
The Android Assistant is now available in the \metagpt/ext/android\_assistant\ directory, providing a multi-modal LLM-driven tool that can learn from human demonstrations or auto-explore apps to generate operation documents, and subsequently execute tasks on an Android device via ADB. The implementation includes actions for manual recording (\ManualRecord\), parsing records into documentation (\ParseRecord\), and executing tasks (\ScreenshotParse\, \SelfLearnAndReflect\) with corresponding prompt templates and schema definitions.
_metagpt/ext/android\assistant · high confidence
Introduce Data Interpreter (DI) actions for code execution and planning
Added new actions in the \metagpt/actions/di\ directory to support a Data Interpreter workflow. This includes \ExecuteNbCode\ for running Python code in a notebook environment with real-time output reporting, \WriteAnalysisCode\ for generating and debugging analysis code with optional reflection, \WritePlan\ for generating task plans in JSON format, \AskReview\ for handling human-in-the-loop review prompts, and \WritePlan\ helper functions for updating plans. These components enable the system to plan, execute, and review data analysis tasks interactively.
metagpt/actions/di · high confidence
Introduce MGX environment with public chat and direct messaging
A new MGX environment implementation is added, enabling a public chat mode where all messages are broadcast to all participants. The environment supports direct, private chats with specific roles (bypassing the team leader) and automatically attaches images from user messages. Message routing is handled by the team leader for standard interactions, while direct chats are processed immediately if the target role is idle.
metagpt/environment/mgx · high confidence
Introduce Minecraft environment and Mineflayer integration
Adds a new Minecraft environment module under \metagpt/environment/minecraft\, including Python wrappers (\MinecraftEnv\, \MinecraftExtEnv\) and a bundled \mineflayer\ Node.js bridge for game interaction. The update also includes a modified version of the \mineflayer-collectblock\ plugin to support block collection tasks, along with supporting configuration files and examples for the Minecraft agent workflow.
metagpt/environment/minecraft · high confidence
Introduce RAG pipeline configuration and interfaces
The RAG module now exposes a structured configuration system for retrievers, rankers, and indexes, supporting backends such as FAISS, Chroma, Elasticsearch, and BM25. New protocol interfaces (RAGObject, NoEmbedding) and schema classes (e.g., ElasticsearchStoreConfig, ColbertRerankConfig, ObjectRankerConfig) enable users to configure retrieval and ranking behavior via explicit config objects rather than global state.
metagpt/rag · high confidence
Introduce RoleZero-based roles for data, engineering, and team leadership
Added new role implementations in the \metagpt/roles/di\ directory, all based on the new \RoleZero\ base class. This includes \DataAnalyst\ and \DataInterpreter\ for handling data-related tasks, \Engineer2\ for software development and deployment, \SWEAgent\ for issue resolution, and \TeamLeader\ for team coordination. These roles introduce a unified framework for dynamic thinking, tool recommendation, and long-term memory, enabling agents to plan, reflect, and execute complex multi-step workflows.
metagpt/roles/di · high confidence
Introduce SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
Adds the SELA extension, an automated machine learning system that integrates Monte Carlo Tree Search with LLM-based agents to explore and refine ML pipeline configurations. The change includes the core \Experimenter\ role, MCTS tree search and visualization tools, an insight generator that proposes and refines instructions for model training, and dataset handling for both local and Hugging Face datasets. It also provides a \datasets.yaml\ configuration for various classification and regression tasks, evaluation metrics, and a README with setup and usage instructions.
metagpt/ext/sela · high confidence
Introduce SPO (Self-Supervised Prompt Optimization) app
Added a new Streamlit-based interactive application for the SPO module, enabling users to configure and run self-supervised prompt optimization workflows. The app provides a graphical interface for managing prompt templates, Q&A examples, and LLM settings, and displays optimization results and logs in an expandable, structured format.
metagpt/ext/spo · high confidence
Introduce SPO optimization and evaluation components
Added new Python modules for the SPO (Self-Prompt Optimization) system. The \optimizer.py\ module introduces a \PromptOptimizer\ class that iteratively generates, executes, and evaluates prompt variations across multiple rounds, logging final results. The \evaluator.py\ module provides \QuickExecute\ and \QuickEvaluate\ classes to handle prompt execution and comparative evaluation of answers using LLM-based prompts. These components enable automated prompt refinement and performance assessment within the SPO extension.
metagpt/ext/spo/components · high confidence
Introduce SimpleEngine and FLAREEngine for RAG workflows
Users can now build Retrieval-Augmented Generation (RAG) applications using the new SimpleEngine and FLAREEngine classes. SimpleEngine provides a straightforward, lightweight workflow for document reading, embedding, indexing, and retrieval, supporting both file-based and object-based inputs. FLAREEngine wraps llama-index's FLAREInstructQueryEngine, allowing users to pass other engines as parameters for advanced query instruction-based retrieval. These engines are exported from metagpt.rag.engines.
metagpt/rag/engines · high confidence
Introduce automated code review and modification actions
Added new \code\_review.py\ and \modify\_code.py\ action classes that enable automated code review and code modification based on pull request patches. The \code\_review.py\ file implements a \CodeReview\ action that analyzes patches against a set of review points (standards) and generates structured comments. The \modify\_code.py\ file implements a \ModifyCode\ action that takes the review comments and generates corresponding code patches to fix the identified issues. Supporting utilities in \cleaner.py\ handle patch cleaning and line number addition, while \schema.py\ defines the \Point\ model for review points.
metagpt/ext/cr/actions · high confidence
Introduce base classes for environment and role abstractions
Added new base classes for environment and role abstractions, including \BaseEnvironment\, \BaseRole\, and \BaseSerialization\. These classes provide abstract interfaces for environment interactions (reset, observe, step, publish\_message) and role behaviors (think, act, react, run). The \BaseSerialization\ class enables polymorphic serialization and deserialization using Pydantic, allowing subclasses to be correctly instantiated from serialized data by tracking class names. This establishes a foundational structure for environment and role implementations within the metagpt/base module.
metagpt/base · high confidence
Introduce context builders for experience pool integration
Added new context builder classes (BaseContextBuilder, SimpleContextBuilder, RoleZeroContextBuilder, and ActionNodeContextBuilder) to format and inject past experiences into prompts. These builders handle serializing experiences into specific templates (e.g., numbered lists or structured text) and integrating them into requests for different agent roles, enabling the system to leverage historical data for improved responses.
_metagpt/exp\_pool/context\builders · high confidence
Introduce experience pool for caching and reusing past results
A new experience pool system has been added to the \metagpt.exp\_pool\ package, providing a mechanism to cache, store, and retrieve past function results. The \exp\_cache\ decorator allows functions to check for existing perfect experiences before execution, falling back to running the function and saving the result if none is found. The \ExperienceManager\ handles the storage and retrieval of these experiences, supporting both BM25 and Chroma-based storage backends. Configuration options such as \enabled\, \enable\_read\, and \enable\_write\ control the behavior of the pool. This feature enables the reuse of previous LLM calls or other expensive operations, potentially reducing latency and cost.
_metagpt/exp\pool · high confidence
Introduce experience-based memory for Werewolf agents
The Werewolf extension now equips agents with a memory system that stores and retrieves past game experiences. New files define actions for adding and retrieving these experiences, which are persisted in a local ChromaDB vector store. This allows agents to recall similar past situations to inform their current decisions, moving beyond simple context windows to include historical reflection data.
metagpt/ext/werewolf/actions · high confidence
Introduce new environment abstraction and base classes for external game integration
The \metagpt/environment\ package is introduced, providing a structured way to integrate with external environments (like Android, Werewolf, Stanford Town, and Minecraft games) via a new \ExtEnv\ base class. This class exposes \action\_space\ and \observation\_space\ and registers read/write APIs for observation and step interactions. The \Environment\ class is also added to host roles and manage message routing within the environment. A README is included to guide users on initializing environments and using the API registry for observations and actions.
metagpt/environment · high confidence
Introduce new prompt templates for the DI (Dynamic Intelligence) module
Added new prompt templates for the DI module, including \architect.py\, \data\_analyst.py\, \engineer2.py\, \role\_zero.py\, \swe\_agent.py\, \team\_leader.py\, and \write\_analysis\code.py\. These files define system instructions, examples, and reflection prompts for various agent roles such as the architect, data analyst, engineer, team leader, and interpreter. The changes also include a new \\\init\\_.py\ file to structure the DI prompts package.
metagpt/prompts/di · high confidence
Introduce strategy module with planning, search, and reasoning capabilities
Added a new \metagpt/strategy\ package that provides foundational components for task planning, search algorithms, and reasoning strategies. This includes a \Planner\ for managing task execution and review, a \ThoughtTree\ and \ThoughtNode\ for representing reasoning states, and solvers for algorithms like Tree of Thoughts (ToT) using BFS/DFS. The module also introduces \TaskType\ to categorize and guide different kinds of tasks, and \Command\ definitions for environment interactions and plan updates.
metagpt/strategy · high confidence
Introduce structured ActionNode and ActionGraph components for task execution
Added new files to the \metagpt/actions\ directory that introduce a structured approach to agent actions. This includes \action\_graph.py\ which defines a directed graph to represent dependencies between actions and perform topological sorting. \action\_node.py\ provides the \ActionNode\ class, which wraps LLM outputs into typed, validated structures (using Pydantic) and supports review/revise modes. A registry \action\_outcls\_registry.py\ ensures consistent class generation. Additionally, new action classes like \AnalyzeRequirementsRestrictions\, \ExecuteTask\, \ExtractReadMe\, \FixBug\, \GenerateQuestions\, and \ImportRepo\ are added, each implementing specific logic for requirement analysis, task execution, documentation extraction, bug fixing, question generation, and repository importing.
metagpt/actions · high confidence
Introduce structured Bedrock provider implementations for multiple LLM families
The AWS Bedrock integration is refactored into a modular provider architecture. A new \base\_provider.py\ defines a \BaseBedrockProvider\ abstract class that standardizes request body generation, completion extraction, and streaming handling. Concrete implementations are added for specific model families: \MistralProvider\, \AnthropicProvider\ (supporting Claude and reasoning capabilities), \CohereProvider\ (Command R/R+), \MetaProvider\ (Llama 2/3), \Ai21Provider\ (Jurassic-2 and Jamba), and \AmazonProvider\ (Titan). Additionally, \utils.py\ provides model-specific chat templates for Llama 2 and Llama 3, and a registry of supported model IDs with their maximum token limits.
metagpt/provider/bedrock · high confidence
Introduce structured configuration and context management
Introduces a new structured configuration system via \metagpt/config2.py\, allowing users to define settings such as LLM, embedding, and search parameters through a Pydantic-based \Config\ class that loads from environment variables, YAML files, and CLI arguments. This is accompanied by a new \Context\ class in \metagpt/context.py\ to manage application state, cost tracking, and LLM instances, along with a \ContextMixin\ in \metagpt/context\_mixin.py\ to provide these capabilities to roles and actions. Additionally, \metagpt/\_compat.py\ adds a Windows-specific fix for the asyncio event loop policy to prevent crashes on Windows systems.
metagpt · high confidence
Introduce the Werewolf game environment and its components
Added the core implementation for the Werewolf game environment, including the environment class, external environment integration, and action/observation spaces. This provides the necessary infrastructure for running the Werewolf simulation, with specific support for role types (Villager, Werewolf, Guard, Seer, Witch, Moderator) and game state management.
metagpt/environment/werewolf · high confidence
Introduces a new scoring system for the experiment pool
A new scoring system has been added to the experiment pool, allowing users to evaluate the quality of LLM-generated responses. This includes a base \BaseScorer\ interface and a default \SimpleScorer\ implementation that uses an LLM to assign a 1-10 quality score to responses. The \metagpt.exp\_pool.scorers\ package is now available for importing \BaseScorer\ and \SimpleScorer\.
_metagpt/exp\pool/scorers · high confidence
Introduces structured configuration classes for all major subsystems
The \metagpt/configs\ directory now contains dedicated configuration classes for each subsystem, replacing the previous flat or global configuration approach. This includes \LLMConfig\ for language model providers and parameters, \ModelsConfig\ for managing multiple LLM instances, \EmbeddingConfig\ for vector search, \BrowserConfig\ for headless browsing, \SearchConfig\ for search engines, \CompressMsgConfig\ for message compression strategies, \ExperiencePoolConfig\ for memory/retrieval, \RoleCustomConfig\ and \RoleZeroConfig\ for role-specific settings, and separate configs for Mermaid, Redis, S3, and workspace paths. Users can now configure each component independently via YAML files.
metagpt/configs · high confidence
New Minecraft observation and collection capabilities
Added a structured observation system for the Minecraft environment, introducing base classes and specific observers for status, inventory, chests, chat, errors, and voxels to provide detailed state information. Additionally, implemented a block collection module that enables the bot to automatically mine blocks, manage inventory by storing items in chests, and handle entity pickups with pathfinding and tool management.
metagpt/environment/minecraft/mineflayer, metagpt/environment/minecraft/mineflayer/mineflayer-collectblock, metagpt/environment/software · high confidence
New RAG benchmarking utilities
Added a new \RAGBenchmark\ class in \metagpt/rag/benchmark\ that provides a suite of evaluation metrics for Retrieval-Augmented Generation systems. The implementation includes scoring methods for BLEU, ROUGE-L, semantic similarity, recall, hit rate, and mean reciprocal rank (MRR), along with dataset loading capabilities to facilitate RAG performance testing.
metagpt/rag/benchmark · high confidence
New RAG examples and benchmarking tools
Added four new example scripts in the \examples/rag\ directory to demonstrate Retrieval-Augmented Generation (RAG) capabilities. \rag\_pipeline.py\ provides a comprehensive RAG pipeline example supporting FAISS, ChromaDB, and Elasticsearch backends, including methods for adding documents and objects. \rag\_bm.py\ introduces a RAG benchmarking tool that evaluates retrieval and query performance using metrics like BLEU and ROUGE. \omniparse.py\ demonstrates parsing documents, PDFs, video, and audio using the OmniParse client. \rag\_search.py\ shows how to integrate RAG with an agent (Sales role) for search functionality.
examples/rag · high confidence
New RAG factory module for managing retrievers, rankers, and embeddings
A new \metagpt/rag/factories\ package was introduced to centralize the creation of RAG components. It provides factory classes and helper functions to instantiate embeddings (supporting OpenAI, Azure, Gemini, and Ollama), retrievers (including FAISS, Chroma, Elasticsearch, and BM25), and rankers (such as LLM, Colbert, and Cohere). This refactors the RAG subsystem to use a consistent, configuration-driven approach for building and retrieving these components.
metagpt/rag/factories · high confidence
New RAG retriever architecture with pluggable backends
The RAG retriever module has been restructured to support multiple storage backends. A new base class hierarchy in base.py defines interfaces for modification, persistence, querying, and deletion. Concrete implementations are provided for BM25, Chroma, Elasticsearch, and FAISS, each implementing these interfaces. A SimpleHybridRetriever is also introduced to aggregate results from multiple retrievers.
metagpt/rag/retrievers · high confidence
New TRD generation and evaluation tools for requirement analysis
Added new actions for generating and evaluating Technical Requirements Documents (TRD) as part of RFC243. The \WriteTRD\ action generates or updates TRDs based on user requirements, external interfaces, and interaction events. \DetectInteraction\ identifies interaction events and participants from user requirements. \EvaluateTRD\ assesses the quality of a TRD against user requirements and interaction events. \CompressExternalInterfaces\ extracts and compresses information about external system interfaces. These tools are now available for requirement analysis workflows.
_metagpt/actions/requirement\analysis/trd · high confidence
New assistant roles and task-specific prompts
Added prompt templates for new assistant roles and task types to guide AI behavior. New files include invoice\_ocr.py for processing invoice OCR data, tutorial\_assistant.py for generating technical tutorials, and product\_manager.py for creating PRDs and market research reports. Additionally, task\_type.py introduces specialized prompts for data science workflows (EDA, preprocessing, feature engineering, model training/evaluation) and content generation tasks (image-to-webpage, web scraping). The change also removes legacy Minecraft-related prompts (decompose, structure\_action, structure\_goal, use\_lib\_sop) and updates existing prompts (sales, summarize, generate\_skill) with English translations and formatting improvements.
metagpt/prompts · high confidence
New assistant, researcher, and tutorial roles; refactored architect and engineer
Added new roles: Assistant, InvoiceOCRAssistant, Researcher, Searcher, Teacher, and TutorialAssistant, each providing specialized capabilities such as skill-based assistance, OCR processing, research, search, teaching, and tutorial generation. Refactored the Architect role to inherit from RoleZero and use pydantic Field for attributes, and updated the Engineer role to support incremental development, message filtering, and file name-based data transmission. Updated the \_\init\\_.py to export the new roles and removed the deprecated seacher.py file.
metagpt/roles · high confidence
New environment API registry for managing read and write operations
The metagpt/environment/api module now includes a new environment API registry system. This introduces an EnvAPIRegistry base class that stores and retrieves callable functions by name, with specialized WriteAPIRegistry and ReadAPIRegistry subclasses. The registry supports getting API schemas as strings or objects, and provides a structured way to manage environment interactions through a centralized registry pattern.
metagpt/environment/api · high confidence
New learn module with text-to-image, text-to-speech, and Google search skills
The \metagpt/learn\ directory now includes new modules for text-to-image, text-to-speech, and Google search functionalities. The \text\_to\_image\ module supports image generation via MetaGPT or OpenAI APIs, returning Base64-encoded images or S3 URLs. The \text\_to\_speech\ module provides text-to-speech conversion using Azure AI or iFlyTek services, returning Base64-encoded audio files or S3 URLs. The \google\search\ module wraps the \SearchEngine\ to perform web searches and return results in markdown format. These new capabilities are exported via the \\\init\\_.py\ file, making them accessible as part of the \metagpt.learn\ package.
metagpt/learn · high confidence
New tool library with browser, editor, git, and data preprocessing capabilities
The metagpt/tools/libs directory now contains a comprehensive set of new tools for software development and data processing. Users can now interact with web pages via the Browser tool (clicking, typing, scrolling, navigating), edit and read files using the Editor tool (write, read, search, and modify code), and manage Git repositories by creating pull requests and issues. Additionally, the library introduces data preprocessing and feature engineering tools (e.g., FillMissingValue, MinMaxScale, PolynomialExpansion) for machine learning workflows, a CodeReview tool for analyzing pull requests, and environment variable management utilities.
metagpt/tools/libs · high confidence
New vector store implementations for LanceDB, Qdrant, and Milvus
The document store module now includes new implementations for LanceDB, Qdrant, and Milvus, expanding the system's vector database capabilities. The LanceDB and Qdrant stores are added as new files, while the Milvus store is updated to use the modern \pymilvus\ client API. Additionally, the legacy \document.py\ file is removed, and the \FaissStore\ is refactored to use \llama\_index\ for index management, replacing the previous \langchain\-based approach.
_metagpt/document\store · high confidence
Project infrastructure and development tooling added
The repository now includes configuration files that standardize the development environment and improve code quality. A Dockerfile is provided to build a consistent runtime environment, while .dockerignore and .gitattributes help manage file handling and line endings. Pre-commit hooks are configured to automatically format code using Black, isort, and Ruff. Additionally, a pytest.ini file is added to configure the test runner, and a SECURITY.md file outlines the vulnerability reporting process.
(repo-wide) · high confidence
Refactored memory system with long-term storage and RAG integration
The memory module has been refactored to support long-term memory storage and retrieval. A new \LongTermMemory\ class and \MemoryStorage\ (using FAISS for vector search) have been introduced to handle persistent memory. Additionally, \RoleZeroLongTermMemory\ integrates a RAG engine (Chroma/LLMRanker) for advanced memory retrieval. The base \Memory\ class now uses Pydantic's \BaseModel\ with \SerializeAsAny\ for better serialization, and includes new methods like \find\_news\ and \get\_by\_position\. The \brain\_memory.py\ file has been updated to support Redis caching and simplified functionality.
metagpt/memory · high confidence
Behavioural changes
Added test document for Omniparse example
A new test document (test01.docx) has been added to the Omniparse example data, providing sample content for testing purposes.
examples/data/omniparse · medium confidence
Refactored Werewolf game roles to use a shared BasePlayer
The werewolf game roles (Guard, Seer, Villager, Werewolf, Witch, Moderator, and HumanPlayer) have been refactored to inherit from a new BasePlayer class. This base class centralizes core game logic, including memory management, observation of messages, and action execution, allowing each role to focus only on its specific behaviors and special actions. This change simplifies the codebase and ensures consistent behavior across all player types.
metagpt/ext/werewolf/roles · high confidence
Removed obsolete example scripts
The repository has removed the \scripts/coverage.sh\ script and the \scripts/set\_env\_example.sh\ file. Users will no longer have access to the previous example for running code coverage reports or the example environment variable configuration script.
scripts · high confidence
Simplified SkillManager by removing LLM dependency
The SkillManager no longer instantiates or stores an LLM object, removing the \llm\ field and its import. This simplifies the skill management logic, as skills are now managed purely through the ChromaStore and internal dictionary without direct LLM integration at this layer.
metagpt/management · high confidence
Test coverage
Added UI test data for legal opinion query interface; Added and updated unit tests for document store implementations; Added comprehensive tests for RAG factory components; Added comprehensive unit tests for LLM provider integrations; Added empty test file for FileManager; Added local test runner for MGX environment; Added mock data and documentation for incremental development tests; Added mock implementations for HTTP clients and LLM provider; Added test coverage and run scripts for Data Interpreter roles; Added test coverage for environment API registry; Added test coverage for prompt templates in strategy module; Added test environment setup script; Added test package for requirement analysis; Added test suite and supporting data for the demo 2048 game project; Added tests for Android assistant actions; Added tests for RAG rankers; Added tests for Tree of Thought strategy examples; Added tests for large PDF processing in the RAG module; Added tests for model configuration loading and validation; Added tests for the Pic2Txt requirement analysis action; Added tests for werewolf experience retrieval and management; Added unit tests for AndroidExtEnv; Added unit tests for Data-Intelligent (DI) actions; Added unit tests for MetaGPT actions; Added unit tests for RAG engine and parser components; Added unit tests for RAG retriever components; Added unit tests for Stanford Town simulation components; Added unit tests for StanfordTownExtEnv; Added unit tests for WerewolfExtEnv; Added unit tests for ZhiPu AI provider; Added unit tests for learn module capabilities; Added unit tests for memory components; Added unit tests for serialization and deserialization of core MetaGPT components; Added unit tests for the base environment API registry and execution; Added unit tests for the experience pool subsystem; Added unit tests for the tools/libs module; Added unit tests for utility modules; Added unit tests for various tool modules; Expanded test coverage for MetaGPT core components; Removed obsolete GPT provider tests; Test suite now uses a global response cache and a dedicated mock LLM to eliminate external API calls; Updated skill manager tests to match new constructor and retrieval signatures.
Dependencies
Add example-specific and environment-specific dependency files
New requirements.txt files have been added for the android\_assistant and stanford\_town examples, as well as for the Sela extension and the Minecraft environment, to pin specific Python and JavaScript dependencies for those contexts. The root requirements.txt has also been updated to include a comprehensive set of pinned packages (e.g., openai, faiss-cpu, anthropic, qdrant-client, and various LLM SDKs), ensuring consistent environments across the project.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 36 → 55 (+18.8)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 67 → 78 (+10.5)
- Architecture 80 → 98 (+17.2)
- Maturity 72 → 62 (-10.1)
- Readiness 13 → 45 (+32.1)
- Security 43 → 54 (+11.9)
Resolved (111)
- (anonymous) (cognitive 18) (metagpt/environment/minecraft/mineflayer/index.js)
- Change coupling: index.js ↔ base.js (metagpt/environment/minecraft/mineflayer/index.js)
- Change coupling: status.js ↔ Inventory.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: status.js ↔ Targets.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: status.js ↔ TaskQueue.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: status.js ↔ TemporarySubscriber.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: status.js ↔ Util.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: status.js ↔ index.ts (metagpt/environment/minecraft/mineflayer/lib/observation/status.js)
- Change coupling: voxels.js ↔ BlockVeins.ts (metagpt/environment/minecraft/mineflayer/lib/observation/voxels.js)
- Change coupling: voxels.js ↔ skillLoader.js (metagpt/environment/minecraft/mineflayer/lib/observation/voxels.js)
- Change coupling: voxels.js ↔ utils.js (metagpt/environment/minecraft/mineflayer/lib/observation/voxels.js)
- Coverage not measured — test suite did not build
- Critical CVE: [GHSA redacted] (requirements.txt)
- Critical CVE: [GHSA redacted] (requirements.txt)
- Critical CVE: [GHSA redacted] (requirements.txt)
- Dimension evaluation failed
- Duplicated block (14 lines × 2) (metagpt/environment/stanford_town/stanford_town_ext_env.py)
- Duplicated block (16 lines × 2) (metagpt/ext/stanford_town/actions/gen_action_details.py)
- Duplicated block (5 lines × 2) (metagpt/tools/search_engine_serpapi.py)
- Duplicated block (5 lines × 2) (metagpt/tools/search_engine_serper.py)
- …and 91 more
New (332)
- (anonymous) (cognitive 20) (metagpt/environment/minecraft/mineflayer/index.js)
- (anonymous) (cognitive 33) (metagpt/environment/minecraft/mineflayer/index.js)
- (anonymous) (cyclomatic 28) (metagpt/environment/minecraft/mineflayer/index.js)
- ActionNode.xml_fill (cognitive 26) (metagpt/actions/action_node.py)
- Banned license: crc32c
- Banned license: pylint
- BaseLLM.compress_messages (cognitive 28) (metagpt/provider/base_llm.py)
- CI installs an unverified third-party binary (.github/workflows/fulltest.yaml)
- Change coupling: architect.py ↔ project_manager.py (metagpt/roles/architect.py)
- Change coupling: azure_tts.py ↔ metagpt_oas3_api_svc.py (metagpt/tools/azure_tts.py)
- Change coupling: base_provider.py ↔ bedrock_api.py (metagpt/provider/bedrock/base_provider.py)
- Change coupling: decorator.py ↔ reflection.py (metagpt/exp_pool/decorator.py)
- Change coupling: humaneval.py ↔ workflow.py (metagpt/ext/aflow/benchmark/humaneval.py)
- Change coupling: index.py ↔ retriever.py (metagpt/rag/factories/index.py)
- Change coupling: manager.py ↔ reflection.py (metagpt/exp_pool/manager.py)
- Change coupling: operator_an.py ↔ workflow.py (metagpt/ext/aflow/scripts/operator_an.py)
- Change coupling: product_manager.py ↔ project_manager.py (metagpt/roles/product_manager.py)
- Change coupling: random_search.py ↔ tree_search.py (metagpt/ext/sela/runner/random_search.py)
- Change coupling: run_experiment.py ↔ tree_search.py (metagpt/ext/sela/run_experiment.py)
- Change coupling: search_engine_serpapi.py ↔ search_engine_serper.py (metagpt/tools/search_engine_serpapi.py)
- …and 312 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
FoundationAgents/MetaGPT was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 11cdf466d042aece04fc6cfd13b28e1a70341b1f — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-09659c52afae.