Skip to content
CAI
Software that uses CAICheck a score

FoundationAgents/OpenManus

58.0

Adequate · 11 October 2026

10k

lines of production code

Python

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is an AI agent orchestration platform that executes tasks using specialized agents for web browsing, data analysis, and software engineering. It supports multi-agent workflows and integrates with external Model Context Protocol (MCP) servers to dynamically load tools. The platform ensures safe execution of untrusted code through isolated Docker and Daytona sandboxes with configurable resource limits.

Features

Add A2A protocol server integration for OpenManus

This change introduces a new experimental A2A (Agent-to-Agent) protocol server within the \protocol/a2a\ module, allowing OpenManus to function as an A2A agent. The implementation includes a Starlette-based server (\protocol/a2a/app/main.py\) that exposes an Agent Card and handles non-streaming task requests via JSON-RPC. It integrates the \a2a-sdk\ (version 0.2.5) to manage agent execution, task storage, and push notifications, enabling external A2A clients to send tasks to OpenManus and receive results.

protocol/a2a · high confidence

Add Baidu, Bing, DuckDuckGo, and Google search engine implementations

The search tool now supports four additional web search providers: Baidu, Bing, DuckDuckGo, and Google. These new engines are implemented as distinct classes (BaiduSearchEngine, BingSearchEngine, DuckDuckGoSearchEngine, GoogleSearchEngine) that conform to the existing WebSearchEngine interface, allowing users to perform searches via these specific providers alongside any previously available options.

app/tool/search · high confidence

Add Daytona sandbox integration for agent tasks

Users can now run the agent within a remote Daytona sandbox environment. This change introduces a new \app/daytona\ module containing \sandbox.py\ to manage sandbox lifecycle (creation, start, stop, delete) and a \README.md\ with setup instructions. To use this feature, users must install \daytona==0.21.8\ and \structlog==25.4.0\, configure their Daytona API key and target region in \config.toml\, and run \sandbox\_main.py\. The agent can then execute tasks using sandbox-specific tools (browser, shell, files), with VNC links provided for visual monitoring of browser actions.

app/daytona · high confidence

Add chart visualization tool with PNG/HTML output and insights

The chart visualization tool now supports generating charts in PNG and HTML formats, utilizing Puppeteer to render VChart visualizations. It also includes functionality to save data analysis insights as markdown files alongside the chart outputs.

_app/tool/chart\visualization/src · high confidence

Added Japan travel plan example with multi-format handbook

The examples directory now includes a Japan travel plan use case featuring a detailed itinerary and a comprehensive HTML travel handbook. The handbook is provided in three distinct formats: a detailed desktop version, a mobile-optimized version with collapsible sections and dark mode support, and a print-friendly version with page-break controls. The example also includes a guide document explaining the different versions and their intended use cases, along with a README file demonstrating the prompt used to generate the content.

examples · high confidence

Initial prompt definitions for browser, manus, MCP, planning, visualization, and SWE agents

This change introduces the initial prompt templates for the application's various AI agents. New files are added for the Browser agent (defining interaction rules and JSON response formats), Manus (a general-purpose assistant), MCP (Model Context Protocol agent with error handling), Planning (structured task execution), and Visualization (data analysis). Additionally, the SWE agent's prompt is simplified by removing verbose instructions and templates, retaining only the core constraint to issue single tool calls.

app/prompt · high confidence

Introduce MCP server for tool execution

Added a new Model Context Protocol (MCP) server implementation in app/mcp/server.py that exposes core tools (bash, editor, terminate) via stdio transport, enabling external clients to invoke these capabilities through the MCP standard.

app/mcp · high confidence

Introduce chart visualization tool with Python execution and multi-format output

The \app/tool/chart\_visualization\ module is added, providing three tools: \python\_execute\ for running Python code for data processing and reporting, \visualization\_preparation\ for generating JSON configuration files from data or existing chart insights, and \data\_visualization\ for rendering charts as PNG or HTML using VChart. The visualization tool supports Chinese and English languages and integrates with the LLM via VMind for intelligent chart generation. Documentation is provided in English, Japanese, Korean, and Chinese.

_app/tool/chart\visualization · high confidence

Introduce sandboxed execution, MCP server support, and multi-agent flows

Users can now run the Manus agent in an isolated Docker sandbox via sandbox\_main.py, connect to external Model Context Protocol (MCP) servers interactively or via single prompts using run\_mcp.py, and orchestrate multiple agents (including the new DataAnalysis agent) in a planning flow with a 1-hour timeout via run\_flow.py. The main entry point (main.py) has been updated to use the Manus agent with command-line prompt support and proper resource cleanup, and a Dockerfile is provided for containerized deployment.

(repo-wide) · high confidence

New browser, data analysis, and sandbox agents with MCP integration

Users can now deploy specialized agents beyond the existing SWE and ReAct agents. A new BrowserAgent connects to the Browser Use CLI 3.0 via MCP to automate web tasks, while a DataAnalysis agent leverages chart visualization and Python execution tools for data reporting. A SandboxManus agent provisions isolated Daytona sandboxes with browser, file, shell, and vision tools, and a generic MCPAgent enables connecting to any MCP server (stdio or SSE) to expose remote tools. The general-purpose Manus agent has been updated to support multiple MCP servers and includes human-in-the-loop interaction via an AskHuman tool. Under the hood, the PlanningAgent has been removed, and the core ToolCallAgent now enforces token limits, supports configurable observation length, and handles empty LLM responses gracefully.

app/agent · high confidence

New sandbox execution environment and Amazon Bedrock support

Users can now execute untrusted code in an isolated Docker sandbox with configurable resource limits (CPU, memory, network) and automatic cleanup, and can also use Amazon Bedrock as an LLM provider. The app now enforces Python 3.11–3.13, adds MCP server configuration, and introduces token usage tracking with a configurable max input token limit.

app · high confidence

New sandbox-based tools for browser, file, shell, and vision operations

Added four new tools in the sandbox module: SandboxBrowserTool for web automation (navigation, interaction, scrolling, tab management), SandboxFilesTool for file operations (create, read, update, delete, string replacement), SandboxShellTool for executing shell commands in tmux sessions, and SandboxVisionTool for reading and compressing images (JPG, PNG, GIF, WEBP) into base64. These tools provide secure, sandboxed environments for interacting with web pages, managing files, running CLI commands, and processing visual content.

app/tool/sandbox · high confidence

New tools and sandbox support for file operations

The application introduces several new capabilities: a new 'ask\_human' tool for interactive queries, a 'computer\_use' tool for desktop automation, a 'crawl4ai' tool for web content extraction, and a 'python\_execute' tool for running Python code. File operations are now supported in both local and sandbox environments via a new 'file\_operators' module, allowing the 'str\_replace\_editor' to work within sandboxes. Additionally, the system now supports connecting to multiple MCP servers to dynamically load remote tools, with automatic deduplication to prevent conflicts.

app/tool · high confidence

Behavioural changes

Agent memory limits, sandbox integration, and structured search results

Agents now enforce a configurable message limit in memory to prevent unbounded context growth, and automatically reset their step counter and state to IDLE upon reaching the maximum step count. The agent base class now supports passing base64-encoded images in messages, and a sandbox client is cleaned up after agent execution. The web search tool has been refactored to return structured results with metadata and supports fetching content from result pages, while the flow base class has been updated to use Pydantic for initialization and no longer automatically assigns tools to agents.

openmanus · high confidence

Refactor app utilities: new file handling and logging, removal of legacy extraction and shutdown modules

The app/utils package has been restructured to introduce new capabilities while removing obsolete ones. New files include files\_utils.py, which provides functions to exclude specific files, directories, and extensions from operations and to normalize paths relative to a workspace, and logger.py, which configures structlog with JSON output for production and console output for local environments. Conversely, extract\_html\_content.py and shutdown\_listener.py have been removed, eliminating the previous logic for extracting code from LLM responses and handling application shutdown signals.

app/utils · high confidence

Refactored planning flow to use enums and Pydantic fields

The planning flow now uses the PlanStepStatus enum for step states instead of hardcoded strings, and the FlowType enum is defined locally in the factory. The PlanningFlow class has been refactored to use Pydantic Field definitions for its attributes (llm, planning\_tool, executor\_keys, etc.), simplifying initialization and removing manual tool collection management.

app/flow · high confidence

Test coverage

Added demo scripts for chart visualization testing; Added tests for sandbox execution and MCP tool integration.

Dependencies

Updated Python dependencies and added chart visualization tool

The project's Python dependencies in requirements.txt have been updated, including upgrades to pydantic (2.10.4 to 2.10.6), openai (1.58.1 to 1.66.3), datasets (3.2.0 to 3.4.1), gymnasium (1.0.0 to 1.1.1), and playwright (1.51.0). New packages added include structlog, fastapi, tiktoken, googlesearch-python, baidusearch, duckduckgo\_search, aiofiles, pydantic\_core (pinned to 2.27.2), colorama, docker, pytest, pytest-asyncio, mcp, httpx, tomli, boto3, requests, beautifulsoup4, crawl4ai, huggingface-hub, and setuptools. Additionally, a new chart visualization tool was introduced in app/tool/chart\_visualization, utilizing @visactor/vchart, @visactor/vmind, and puppeteer for generating chart outputs.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 58.

Lenses

  • Code Health 64
  • Architecture 91
  • Maturity 53
  • Readiness 56
  • Security 59
  • Performance 85

Changes since last survey

  • 300 commits — 249 feature/other, 51 fixes

By area

  • (repo) — 87 commits
  • app/tool — 82 commits
  • app/agent — 37 commits
  • (root) — 33 commits
  • app/llm.py — 10 commits
  • .github/workflows — 6 commits
  • app/config.py — 5 commits
  • A2A/Manus — 4 commits
  • app/daytona — 4 commits
  • app/prompt — 4 commits
  • config/config.example.toml — 3 commits
  • protocol/a2a — 3 commits
  • .github/ISSUE_TEMPLATE — 2 commits
  • A2A/common — 2 commits
  • app/utils — 2 commits
  • config/config.example-model-anthropic.toml — 2 commits
  • app/init.py — 1 commit
  • app/main.py — 1 commit
  • app/bedrock.py — 1 commit
  • app/flow — 1 commit

Notable commits

  • fix: Apply pre-commit formatting fixes
  • fix: Apply pre-commit formatting fixes
  • fix: BUGFIX: FIX MCPAgent Bug
  • fix: Fix bug of SWEAgent
  • fix: Fix bugs next_step_prompt accidentally overwritten
  • fix: Fix using isort
  • fix: Fixed the wrong import
  • fix: I have implemented a fix to asyncio resource leak errors when the script finishes. A cleanup method has been added to the agent hierarchy (ToolCallAgent, BrowserAgent, Manus), and main.py now ensures this cleanup is called using a try...finally block before the program exits. This should prevent the ValueError: I/O operation on closed pipe exceptions by properly closing browser resources.
  • fix: Merge branch 'fix-for-search-rate-limits' of https://github.com/a-holm/OpenManus into fix-for-search-rate-limits
  • fix: Merge branch 'main' into fix-for-search-rate-limits
  • fix: Merge branch 'main' of https://github.com/a-holm/OpenManus into fix-for-search-rate-limits
  • fix: Merge branch 'main' of https://github.com/a-holm/OpenManus into fix-for-search-rate-limits
  • fix: Merge branch 'mannaandpoem:main' into fix-error-with-cleaning-up
  • fix: Merge pull request #758 from cyzus/feat-fix-browser-use-click
  • fix: Merge pull request #772 from a-holm/fix-for-search-rate-limits
  • fix: Merge pull request #859 from minbang930/fix/extracted-content-properties
  • fix: Merge pull request #951 from a-holm/fix-error-with-cleaning-up
  • fix: Merge pull request #963 from Jeffrey95/fix/manus-inf-loop
  • fix: Revert "Merge branch 'feat/data_visualization_hackathon' of https://github.com/666haiwen/OpenManus into feat/data_visualization_hackathon"
  • fix: Revert "feat: generate structured analysis reports"
  • …and 280 more

Architecture

  • 0 containers · 1 bounded contexts · 0 dependency edges (baseline)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

FoundationAgents/OpenManus was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 11 October 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 3309bf4e416fb1c74b008f3e86494439a31bad53 — the exact code this score is about.
  • Scored under rubric-2026.10.5 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-fe8540b5da9b.