langchain-ai/deepagents
67.5
Adequate · 18 September 2026
263.8k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a modular framework for building and running autonomous AI agents, centered around the Deep Agents SDK and a terminal-based CLI interface. It provides a pluggable architecture for agent execution, featuring middleware for memory and tool use, a plugin marketplace for extensibility, and support for various sandbox backends like Modal and Vercel. The platform also includes comprehensive evaluation suites for benchmarking agent capabilities and integrations for running agents within text editors and messaging channels.
Features
Add DeepAgents adapter for the continual-learning-bench
Introduces a new evaluation harness that wraps the DeepAgents system as a \ContinualLearningSystem\ for the \continual-learning-bench\ (clbench). The adapter enables the agent to learn across sequential tasks by maintaining persistent memory in an in-state file (\/memory/AGENTS.md\), which is loaded into the prompt each turn and updated by the agent via its own tools. This directory serves as the canonical source, deployed into a clbench checkout via \sync\_to\_clbench.sh\ to run benchmarks like \exploitable\_poker\.
_libs/evals/deepagents\clbench · high confidence
Add context-retrieval-evals dataset with 30 calibrated tasks
A new Harbor dataset named \context-retrieval-evals\ has been added to \libs/evals/datasets\, containing 30 tasks derived from the Context-Bench \cloud\ suite. These tasks are designed to evaluate an agent's ability to extract and reason over information spread across a multi-file corpus (10 files per task). The dataset includes a \calibration.json\ file recording paired performance metrics for the gpt-5.6-terra and gpt-5.6-luna models, and each task is structured with a Dockerfile, instruction, solution, and test case to facilitate automated evaluation.
libs/evals/datasets · high confidence
Add standalone better-harness example for autonomous agent harness optimization
A new \better-harness\ example has been added to the \examples\ directory, providing a research artifact and reference implementation for an autonomous harness-optimization loop. This tool uses an outer Deep Agent to iteratively diagnose failures and propose edits to editable harness surfaces—such as prompts, tools, skills, and middleware—based on explicit train and holdout evaluation cases. The example includes the full Python package implementation (core data models, agent orchestration, patching logic, and pytest/harbor runners), a worked configuration file (\deepagents\_example.toml\) demonstrating how to expose and wire these surfaces, and documentation explaining the optimization workflow.
examples/better-harness · high confidence
Added import validation script for pre-commit checks
A new Python script, check\_imports.py, has been added to the libs/deepagents/scripts directory. This utility allows developers to quickly verify that a list of Python files can be loaded by the interpreter without errors, serving as a fast pre-commit or pre-test check to catch import issues early. It is designed to be integrated into Makefiles and CI pipelines to prevent broken imports from reaching more expensive test suites.
libs/deepagents/scripts · high confidence
Context-Bench evaluation adapter and grading infrastructure
This change introduces the Context-Bench adapter within the Harbor evals system, enabling the generation of 30 calibrated Context-Bench tasks from a vendored filesystem-cloud corpus. It includes a CLI driver to generate tasks, populate environment files, and stamp calibrated difficulty tiers, alongside a faithful reimplementation of the upstream Letta RubricGrader in \judge.py\ to ensure grading scores match the upstream standard.
_libs/evals/harbor\adapters · high confidence
Deep Agents Harbor evals: LangSmith integration, failure classification, and statistical reporting
The \libs/evals/deepagents\_harbor\ package now provides a complete evaluation toolkit for Deep Agents Harbor runs. It integrates with LangSmith to automatically create datasets, experiments, and feedback traces, including deterministic example IDs and safe API key resolution. Eval results are now classified into infrastructure failures (OOM, timeout, sandbox) versus model capability failures using exit codes and text pattern matching. Additionally, score reporting includes Wilson score confidence intervals and minimum detectable effect calculations to provide statistically robust eval metrics.
_libs/evals/deepagents\harbor · high confidence
Enforced branch naming convention via pre-push hook
A new pre-push hook has been added to validate that local branch names follow the repository's naming convention (format: \<github-username\>/\<scope\>/\<short-description\>). The script automatically resolves the developer's GitHub username via git config, the GitHub CLI, or the commit email, and rejects pushes for branches that do not match the expected pattern, except for protected branches (main, master, version tags) and automation/release prefixes (release-please, dependabot, copilot, alpha, beta, rc, dev). This provides immediate local feedback before pushing, complementing the server-side CI check.
.githooks · high confidence
Introduce ACP server integration for Deep Agents
This change adds the \deepagents\_acp\ library, providing an Agent Client Protocol (ACP) server implementation for Deep Agents. It enables Deep Agents to function as an ACP-compliant server, handling session management, tool execution, and streaming responses (including visible reasoning and raw input in tool calls) while ensuring security by blocking dangerous shell patterns in auto-approved commands. The implementation includes version 0.0.11 and supports ACP schema v0.9.0+ with backward compatibility for v0.8.x.
_libs/acp/deepagents\acp · high confidence
Introduce Deep Agents ACP integration for editor-based agent workflows
This change adds the \libs/acp\ package, providing an Agent Client Protocol (ACP) connector that allows Python Deep Agents to run inside ACP-compatible text editors like Zed. It includes a demo coding agent using Anthropic's Claude models with built-in filesystem and shell tools, as well as support for dynamic model switching and persistent session loading via durable LangGraph checkers. The package also ships a prebuilt \dcode\ CLI agent that can expose itself as an ACP server, along with necessary configuration files (\.env.example\, \Makefile\, \uv.lock\) and documentation to guide users through setup and integration.
libs/acp · high confidence
Introduce LangGraph-based agent runtime for Harbor evaluations
The Harbor benchmark now supports running agents via a self-contained LangGraph project, providing a unified registry of agent graphs (dcode, bare, tau3) defined in \langgraph.json\. This change introduces a new LangGraph entrypoint (\langgraph\_agent.py\) that handles model initialization, environment scrubbing for security, and working directory inheritance, while also adding support for the GLM-5.2 harness profile and migrating the Harbor adapter to use \langchain.mcp\ for tool integration.
_libs/evals/deepagents\_harbor/langgraph\project · high confidence
Introduce Modal sandbox backend for Deep Agents
This release adds the \langchain-modal\ partner package (version 0.0.6), providing a \ModalSandbox\ implementation that integrates Modal sandboxes with the Deep Agents backend protocol. Users can now execute shell commands within a Modal sandbox environment, supporting per-command timeout overrides, and perform file uploads and downloads, enabling sandboxed agent execution on Modal infrastructure.
_libs/partners/daytona/langchain\_daytona, libs/partners/modal/langchain\modal · high confidence
Introduce NVIDIA Deep Agent example with GPU-accelerated data processing
Adds a new example demonstrating a multi-model deep agent architecture that uses a frontier model as an orchestrator and NVIDIA Nemotron Super for research. The example includes a data-processor subagent capable of GPU-accelerated data analysis (cuDF), machine learning (cuML), and visualization, executing code within a Modal GPU sandbox. It provides pre-configured skills for analytics, ML, visualization, and document processing, along with a self-improving memory system that allows the agent to update its own skill documentation based on execution experience.
_examples/nvidia\_deep\agent · high confidence
Introduce Runloop sandbox integration for Deep Agents
This change adds the \langchain-runloop\ partner package, providing \RunloopProvider\ and \RunloopSandbox\ classes that allow Deep Agents to create, attach to, and execute commands within Runloop devboxes. The provider supports bootstrapping sandboxes from named blueprints (with automatic Dockerfile building if needed) or attaching to existing devboxes by ID, while the sandbox implementation handles command execution, file uploads, and file downloads via the Runloop API.
_libs/partners/runloop/langchain\runloop · high confidence
Introduce Talon local runtime host with chat-scoped history and background subagents
This change introduces the \deepagents\talon\ package, providing a local runtime host for long-running Deep Agents. It adds a CLI entry point (\\\main\\_.py\) that supports attaching WhatsApp, Telegram, and Discord channel adapters, importing Fleet zip exports, and managing MCP server OAuth logins. The runtime includes a persistent, chat-scoped conversation archive (\archive.py\, \archive\_saver.py\) that wraps LangGraph checkpointers to store and search history, and a background subagent system (\background.py\) that allows the main agent to delegate work to expendable tasks that run independently of the conversation turn.
_libs/talon/deepagents\talon · high confidence
Introduce Vercel Sandbox provider for Deep Agents
Users can now execute shell commands and manage files within isolated Vercel Sandboxes using the new \langchain-vercel-sandbox\ package. This integration provides a \VercelSandbox\ backend that wraps the Vercel SDK, supporting command execution with configurable timeouts, file uploads and downloads, and output truncation to manage large responses. The package is released as version 0.0.2 and requires \deepagents\ 0.7.x.
libs/partners/vercel · high confidence
Introduce \`deepagents\_code\` package with lazy CLI entry point and startup-optimized modules
The \deepagents\_code\ package is now available as a standalone importable library. It provides a lazy \cli\_main\ entry point to avoid loading heavy startup machinery (like \argparse\ and signal handling) when only configuration or widget modules are imported. The package also includes a set of lightweight, dependency-free modules (such as \\_constants\, \\_ask\_user\_types\, and \\_cli\_context\) designed to support the TUI and agent graph without pulling in the full LangChain middleware stack, improving import performance and modularity.
_libs/code/deepagents\code · high confidence
Introduce beta model and harness profile APIs
The \deepagents.profiles\ package now exposes beta APIs that allow users to tailor Deep Agents behavior to specific providers or model specifications. This includes \ProviderProfile\ for controlling model construction (e.g., \init\_chat\_model\ kwargs) and \HarnessProfile\ for controlling runtime agent behavior (e.g., prompt assembly, tool visibility). The system supports additive registration via \register\_provider\_profile\ and \register\_harness\_profile\, allowing users to layer custom configurations over built-in defaults for providers like OpenAI, Anthropic, and NVIDIA, as well as third-party plugins via entry points.
libs/deepagents/deepagents/profiles · high confidence
Introduce built-in skills for thread inspection and skill creation
The Deep Agents Code now ships with built-in skills that are always available at the lowest precedence level, allowing user and project skills to override them. This includes a thread-inspector skill that lets users inspect and explain conversations in the local SQLite session store (reading from $DEEPAGENTS\_HOME/.state/sessions.db) when LangSmith tooling is unavailable, and a skill-creator skill that guides users in creating, structuring, and validating new agent skills with best practices and automated validation scripts.
_libs/code/deepagents\_code/built\_in\skills · high confidence
Introduce experimental Python extension system for agent customization
Adds a new, experimentally gated Python extension system that allows users to register custom tools, LangChain middleware, and backend routes within the agent runtime. The system discovers extensions from user directories, project folders, and installed plugins, enforcing a trust policy (ask, always, or never) for project-authored code. Extensions can dynamically inject tools into the agent's toolset and mount virtual filesystem paths via backend routes, with runtime validation to prevent conflicts with internal storage or sandbox violations. This feature is controlled by the \EXPERIMENTAL\ environment variable and the \\[extensions\]\ configuration section.
_libs/code/deepagents\code/extensions · high confidence
Introduce interactive plugin manager modal
Adds a new \/plugins\ modal screen for managing plugins, featuring tabs for discovering and installing plugins from marketplaces, viewing installed plugins, monitoring errors, and configuring settings. Users can search the plugin list, view detailed component summaries (skills, MCP servers, hooks), enable/disable or uninstall plugins, and manage marketplace sources. The interface displays real-time connection status for MCP-enabled plugins and prompts for a session reload when plugin state changes.
_libs/code/deepagents\_code/tui/modals/plugin\manager · high confidence
Introduce middleware package with filesystem, subagent, and summarization capabilities
The \deepagents.middleware\ package is now available, providing a structured set of middleware components for the Deep Agents SDK. This includes \FilesystemMiddleware\ for managing file operations with permission controls and large-result offloading, \AsyncSubAgentMiddleware\ for launching and monitoring background tasks on remote Agent Protocol servers, and \SummarizationMiddleware\ for context window management. The package also exposes \MemoryMiddleware\ for agent memory, \RubricMiddleware\ for self-evaluation, and \SkillsMiddleware\ for dynamic tool loading, along with internal helpers for message eviction, prompt caching, and tool exclusion.
libs/deepagents/deepagents/middleware · high confidence
Introduce persistent cron scheduling with timezone-aware wall-clock support
Talon now includes a new cron scheduling subsystem that allows agents to create, list, edit, and remove persistent background jobs. Jobs are stored in a structured, versioned format and support both relative intervals (e.g., 'every 15m') and timezone-aware wall-clock schedules (e.g., 'daily at 08:00 America/New\_York') that persist across daylight saving changes. The scheduler ensures reliability by keeping the ticker alive through failed ticks and treats trailing '\[SILENT\]' markers to suppress result delivery to the chat.
_libs/talon/deepagents\talon/cron · high confidence
Introduce pluggable memory backend system
The \deepagents.backends\ package now provides a unified, pluggable system for file storage and execution. It defines a \BackendProtocol\ and \SandboxBackendProtocol\ that standardize file operations (read, write, edit, glob, grep, ls) and command execution. The package ships with several concrete implementations: \FilesystemBackend\ for direct local disk access (with optional virtual path mode), \StateBackend\ for ephemeral storage within LangGraph agent state, \StoreBackend\ for persistent storage with namespace support, \ContextHubBackend\ for storing files in a LangSmith Hub agent repo, \LangSmithSandbox\ for executing commands and managing files within a LangSmith sandbox environment, and \LocalShellBackend\ which combines filesystem access with unrestricted local shell execution. A \CompositeBackend\ is also included to route file operations to different backends based on path prefixes, enabling mixed storage strategies.
libs/deepagents/deepagents/backends · high confidence
Introduce plugin marketplace and management system
Adds a new plugin subsystem that allows users to discover, install, enable, and uninstall plugins from configured marketplaces (local directories, GitHub, Git, or HTTP URLs) via a new \plugin\ CLI command. The implementation includes marketplace source parsing with credential redaction, manifest validation, component discovery (skills, MCP servers, hooks), state storage with file locking, and environment variable substitution for plugin configuration.
_libs/code/deepagents\code/plugins · high confidence
Introduce sandboxed JavaScript REPL middleware for agents
This change adds the \langchain\_quickjs\ package, providing a \CodeInterpreterMiddleware\ that exposes a sandboxed QuickJS REPL to LangChain agents. Users can now execute JavaScript code via an \eval\ tool, with support for programmatic tool calling (PTC) to invoke agent tools directly from JS, and a \task()\ API for dispatching subagents. The middleware includes features like REPL state persistence (thread, turn, or call modes), snapshot-based delta encoding for efficient state storage, HMAC-signed snapshots for integrity, and configurable limits for memory, timeouts, and PTC call budgets.
_libs/partners/quickjs/langchain\quickjs · high confidence
Introduce unified Deep Agents evaluation suite with CLI and structured reporting
The \libs/evals/deepagents\_evals\ package is introduced, providing a centralized evaluation framework for agent capabilities. It includes a unified CLI (\deepagents-evals\) with subcommands for running trials, aggregating results, and generating radar charts. The suite organizes evaluations into specific capability categories (file\_operations, retrieval, tool\_use, memory, conversation, summarization) defined in \categories.json\, and includes a curated \tau3\_subset\ dataset for probing conversation behavior across difficulty tiers. Users can now run structured evals, view per-category correctness matrices in GitHub Actions summaries, and visualize model performance via radar charts with configurable themes.
_libs/evals/deepagents\evals · high confidence
Introduces Harness Profiles for model-specific agent behavior tuning
The \deepagents.profiles.harness\ package now provides a registry of built-in harness profiles that allow users to apply model-specific system prompt suffixes, middleware, and subagent configurations to specific model keys (such as \anthropic:claude-haiku-4-5\, \anthropic:claude-opus-4-7\, \anthropic:claude-sonnet-4-6\, \openai:gpt-5.x-codex\, and \nvidia:nemotron-3-ultra\). These profiles layer guidance like parallel tool usage, grounded investigation, and autonomous action directly onto the agent runtime, ensuring the model behaves according to its specific training characteristics without altering the base system prompt for other models.
libs/deepagents/deepagents/profiles/harness · high confidence
Introduces provider profiles for default model configuration and attribution
The SDK now includes a built-in provider profile system that automatically configures specific model providers. For OpenRouter, the system enforces a minimum \langchain-openrouter\ version (0.2.0) and injects app-attribution headers by default, while also ignoring Azure as an upstream provider to prevent multi-turn reasoning failures (opt-out via \DEEPAGENTS\_OPENROUTER\_ALLOW\AZURE\). For OpenAI, the Responses API is enabled by default for all \openai:\\ models. For NVIDIA NIM, the SDK injects an app-origin attribution header (\X-BILLING-INVOKE-ORIGIN\) into requests. These defaults can be extended or overridden by users via the new \register\_provider\_profile\ API.
libs/deepagents/deepagents/profiles/provider · high confidence
Introduces typed Pydantic models for the Hooks v2 system
This change adds a new \hooks/models\ package that defines the structured data contracts for the Hooks v2 lifecycle. It introduces typed domain models for hook events (such as \SessionStart\, \PreToolUse\, and \PostToolUseFailure\), configuration schemas for command handlers, and wire/transport models for external communication. These models enforce strict validation for hook inputs and outputs, ensuring that hook handlers receive well-defined data structures and that configuration errors are caught early.
_libs/code/deepagents\code/hooks/models · high confidence
New ACP demo agent with local context middleware and model switching
The ACP examples library now includes a new demo coding agent (\demo\_agent.py\) that integrates with the ACP server. This agent features a \LocalContextMiddleware\ (\local\_context.py\) which automatically detects and injects project details (language, package manager, git state, runtimes) into the system prompt. It supports dynamic model switching across OpenAI, Anthropic, and Baseten providers, and allows users to configure session modes (e.g., 'ask before edits', 'accept everything') to control permission prompts for file and shell operations.
libs/acp/examples · high confidence
New CLI command groups for auth, config, tools, and extras
The CLI now exposes dedicated subcommand groups for managing credentials (\dcode auth\), inspecting configuration and their sources (\dcode config\), listing available agent tools (\dcode tools list\), and installing optional extras (\dcode install\). These commands provide non-interactive, scriptable access to capabilities previously only available through the Textual UI, including credential management, configuration resolution, and tool catalog inspection.
_libs/code/deepagents\code/client/commands · high confidence
New LLM Wiki example with persistent, sync-enabled knowledge base
Added a new \examples/llm-wiki\ example that implements a script-first Deep Agents workflow for building and maintaining a persistent, topic-specific wiki. The runner supports four modes: \init\ to create a Context Hub repo with \source=internal\ enforcement, \ingest\ to process raw source files into canonical wiki pages (with an optional human-in-the-loop review step), \query\ to answer questions against the wiki with grounded citations and optional durable filing of answers, and \lint\ to perform health checks, reconcile contradictions, and fix orphan pages. All changes are synced to LangSmith Context Hub, and a structured, append-only \log.md\ tracks every operation for auditability and shell-based parsing.
examples/llm-wiki · high confidence
New Ralph Mode example with non-interactive autonomous looping
Added a new \examples/ralph\_mode\ directory containing a Python script (\ralph\_mode.py\) and documentation that implements the 'Ralph' autonomous looping pattern. This example allows users to run the Deep Agents CLI in a non-interactive loop where each iteration starts with fresh context while persisting filesystem changes, supporting features like iteration limits, custom models, remote sandboxes (AgentCore, Modal, Daytona, Runloop), and shell command allow-lists.
_examples/ralph\mode · high confidence
New RubricMiddleware example with LangSmith tracing
Added a runnable example in \examples/rubric\_middleware\ that demonstrates the \RubricMiddleware\ end-to-end with real models and LangSmith tracing. The script runs an agent that drafts a document, uses a grader model to score it against a rubric, and iteratively revises the output until all criteria are satisfied or the iteration budget is exhausted, providing visibility into grading verdicts and revision prompts via LangSmith.
_examples/rubric\middleware · high confidence
New Text-to-SQL Deep Agent example
Added a new example in \examples/text-to-sql-agent\ that demonstrates a LangChain Deep Agent capable of converting natural language questions into SQL queries against a Chinook SQLite database. The example includes an \agent.py\ entry point that uses Claude Sonnet 4.5, SQL tools, and a filesystem backend with progressive disclosure skills (query-writing, schema-exploration) to plan and execute complex analytical queries, along with supporting documentation, environment configuration, and dependency lock files.
examples/text-to-sql-agent · high confidence
New channel adapters for Discord, Telegram, and WhatsApp
Talon now supports connecting to Discord, Telegram, and WhatsApp, allowing users to interact with the agent through these messaging platforms. The Discord adapter uses the Gateway client and can register chat commands as native slash commands. The Telegram adapter uses the Bot API with long-polling and persists state offsets. The WhatsApp adapter connects via a local Node.js bridge process, handling media, reactions, and message context. All adapters share a common base for message formatting, text chunking, and exposure policies (self, allowlist, or open).
_libs/talon/deepagents\talon/channels · high confidence
New coding agent deployment example with structured skills
Added a new \examples/deploy-coding-agent\ directory that demonstrates deploying an autonomous coding agent using \deepagents deploy\. The example includes an \agent.json\ configuration targeting the \anthropic:claude-sonnet-4-5\ model, environment variable templates for Anthropic and LangSmith API keys, and a set of agent skills (\planning\, \code-review\, \coding-prefs\) that define a structured Plan → Implement → Review → Deliver workflow. It also provides a custom Python lint-check helper script and documentation (\README.md\, \AGENTS.md\) to guide users on deployment, usage, and SDK integration.
examples/deploy-coding-agent · high confidence
New confirmation and management modals for cost, context, and session control
The TUI now presents several new modal dialogs to improve user control over session behavior and costs. A cold-cache warning asks whether to send a turn that may lose prompt-cache savings, offering options to send, suppress the warning for the session, or persistently suppress it in config. Switching models with a large context triggers a confirmation that warns about potential loss of prompt-cache savings and context-limit differences, directing users to /offload to reduce context. Resuming a thread with large context suggests compacting older messages to reduce costs, and handles cases where an old operation was unfinished by offering to cancel it. If a policy blocks resuming a thread at launch, users are prompted to start a new session or exit. Additionally, a new prompt clipboard modal allows searching, previewing, copying, and inserting previously submitted prompts.
_libs/code/deepagents\code/tui/modals · high confidence
New content-writing agent example with file-based configuration
The \examples/content-builder-agent\ directory now contains a complete example of a content-writing agent that generates blog posts, LinkedIn updates, and tweets with cover images. The agent is configured entirely through filesystem primitives: \AGENTS.md\ defines brand voice and style guidelines, \skills/\ directories provide on-demand workflows for specific content types, and \subagents.yaml\ defines a researcher subagent for web search and fact-checking. The \content\_writer.py\ script wires these components together using the \deepagents\ framework, demonstrating how to combine memory, skills, and subagents to produce structured content output in local directories.
examples/content-builder-agent · high confidence
New deep research agent example with Tavily search and strategic reflection
A new \deep\_research\ example has been added to the \examples\ directory, demonstrating a multi-agent research workflow using the \deepagents\ package. The example provides a standalone agent script (\agent.py\) and an interactive Jupyter notebook (\research\_agent.ipynb\) that utilize custom tools: \tavily\_search\ for web discovery and content fetching, and \think\_tool\ for strategic reflection. It includes specific prompt templates (\prompts.py\) that enforce a 5-step research workflow (plan, save, delegate, synthesize, write) with hard limits on tool calls to prevent excessive searching. The example supports deployment via a local LangGraph server (\langgraph.json\) or direct execution, requiring API keys for Anthropic, OpenAI, Tavily, and LangSmith.
_examples/deep\research · high confidence
New developer tooling scripts and hardened installer for deepagents-code
This change introduces several new scripts in \libs/code/scripts\ to improve the development workflow and installation experience. A new \check\_imports.py\ script allows developers to quickly verify that Python files load without errors in an isolated environment, while \check\_process\_cwd.py\ enforces correct working-directory usage across the codebase via an AST-based allowlist. A \debug\_server.sh\ script has been added to help diagnose server startup failures by capturing version, environment, and log details. Additionally, \generate\_commands\_catalog.py\ automates the creation of the \COMMANDS.md\ reference from the slash-command registry. The \install.sh\ script has been significantly enhanced with support for pre-release versions, configurable extras, ripgrep provisioning options, and stricter validation of arguments and downloads.
libs/code/scripts · high confidence
New evals analysis and reporting scripts
Added a suite of new Python scripts in \libs/evals/scripts\ to enhance the evaluation workflow: \analyze.py\ scans trial directories to extract trajectory data and success metrics; \run\_trials.py\ executes the eval suite multiple times to aggregate statistics like mean, median, and standard deviation across trials; \generate\_radar.py\ and \composite\_radar.py\ create visual radar charts from eval results, including the ability to overlay multiple GitHub Actions runs; \generate\_eval\_catalog.py\ and \generate\_model\_groups.py\ auto-generate documentation for the eval catalog and model groups; and \harbor\_langsmith.py\ provides a CLI for integrating Harbor tasks with LangSmith datasets and experiments.
libs/evals/scripts · high confidence
New self-hosted async subagent server example
Added a new example in \examples/async-subagent-server\ that demonstrates how to host a Deep Agents researcher as an async subagent using the Agent Protocol. The example includes a FastAPI server (\server.py\) that implements the required HTTP endpoints (threads, runs, cancellation) with in-memory SQLite persistence, and a supervisor REPL (\supervisor.py\) that connects to the server to delegate research tasks. It also provides an \.env.example\ for configuration and \test\_server.py\ with end-to-end tests for the server's HTTP contract.
examples/async-subagent-server · high confidence
New skills management system with plugin marketplace and trust controls
This change introduces a comprehensive skills management system for Deep Agents Code, enabling users to discover, list, create, and delete skills via CLI commands. It adds support for a plugin marketplace by allowing skills to be sourced from plugin directories with namespaced prefixes. The system enforces strict security through a trust store that manages in-the-moment approvals for symlinked skills, preventing unauthorized file access by verifying resolved paths against trusted roots. Additionally, it provides debug logging for skill name collisions to aid in troubleshooting override behaviors.
_libs/code/deepagents\code/skills · high confidence
Plugin subsystem adapters for hooks, MCP servers, and skills
The plugin subsystem now includes dedicated adapters that translate plugin declarations into the application's internal configuration sources. Users will see plugin-defined hooks, MCP servers, and skills properly discovered and namespaced (e.g., \plugin\\\<id\>\\\<name\>\ for MCP servers, \plugin:id:skill\ for skills) to avoid collisions and ensure compatibility with the core loader. The adapters handle reading hook and MCP configurations from plugin files and manifests, normalizing server settings (such as environment variables and working directories), and discovering skill directories containing \SKILL.md\ files, while gracefully handling discovery or loading failures without breaking other plugins.
_libs/code/deepagents\code/plugins/adapters · high confidence
Architecture
Introduce dedicated module for local LangGraph server lifecycle management
The \deepagents\_code.client.launch\ package has been created to centralize the logic for starting, configuring, and stopping the internal LangGraph development server. This change introduces \server.py\ to handle the server process lifecycle (including ephemeral port selection, health polling, and graceful shutdown with trace flushing) and \server\_manager.py\ to orchestrate the full startup flow, which now includes scaffolding a temporary workspace with a generated \langgraph.json\, \pyproject.toml\, and SQLite checkpointer module. This modularization separates server runtime concerns from the broader client application code.
_libs/code/deepagents\code/client/launch · high confidence
TUI widgets reorganized into a dedicated package
The Textual widget implementations for the TUI have been moved from their previous locations into a new \tui/widgets\ package. This change restructures the codebase by consolidating the UI components—such as the chat input, approval menus, agent selectors, and authentication screens—into a single, organized module, improving maintainability and import clarity for the terminal interface.
_libs/code/deepagents\code/tui/widgets · high confidence
Behavioural changes
Consolidated Textual UI adapter and shared navigation hints
The terminal user interface components have been reorganized into a dedicated \tui\ package, introducing a new \textual\_adapter.py\ that centralizes agent execution logic, Human-in-the-Loop (HITL) decision handling, and session cost tracking, alongside a \key\_hints.py\ module that standardizes modal navigation instructions across screens to ensure consistent keyboard shortcuts and glyph rendering.
_libs/code/deepagents\code/tui · high confidence
Deep Agents SDK v0.7.15 release with new public API surface and internal refactoring
The deepagents package is released as version 0.7.15, establishing a new public API surface in the main \\_\init\\_.py\ that explicitly exports core components such as \DeepAgentState\, \create\_deep\_agent\, various middleware classes (AsyncSubAgent, Filesystem, Memory, Rubric, SubAgent), and profile registration utilities. This release introduces internal structural changes including a dedicated deprecation adapter module to improve warning stack attribution, a custom message delta reducer to optimize checkpoint growth, and refactored model resolution helpers that normalize provider names and support AWS Bedrock detection. The package also includes a \py.typed\ marker file to enable static type checking for consumers.
libs/deepagents/deepagents · high confidence
Introduce Hooks v2 with new event-driven architecture and legacy migration
The hooks subsystem has been replaced by Hooks v2, a new event-driven system that supports client-owned lifecycle events (SessionStart, SessionEnd, UserPromptSubmit) and server-owned tool events (PreToolUse, PostToolUse, PostToolUseFailure, PermissionRequest, Notification). Configuration is now loaded from project, user, and plugin sources with a new JSON schema, and legacy \hooks.json\ documents are automatically migrated to the v2 format at load time. The legacy hook dispatcher is retained for backward compatibility until September 1, 2026, but new integrations should target the v2 config format. The system also includes a capability registry defining event-specific behaviors, a client-side fulfillment ledger for deduplicating server-owned hook invocations, and workspace trust checks for project-scoped hooks.
_libs/code/deepagents\code/hooks · high confidence
Test coverage
Added MemoryAgentBench evaluation tests; Added benchmark tests for CLI startup and local context detection performance; Added benchmarks for deep agent construction and summarization middleware; Added evaluation test data and API stubs for BFCL and Nexus benchmarks; Added integration test suite for deepagents-code; Added integration tests for QuickJS PTC and subagent middleware; Added integration tests for Talon runtime, channels, and history backends; Added integration tests for deepagents SDK components; Added smoke tests for QuickJS system prompt stability; Added smoke tests for system prompt snapshots; Added snapshot test for CLI agent system prompt; Added tau2 airline evaluation suite with tau3 task data; Added test coverage for Talon channel adapters and base utilities; Added test infrastructure and utilities for the Deep Agents SDK; Added test suite for ACP agent server and utilities; Added test suite for Daytona sandbox integration; Added test suite for the Modal partner integration; Added tests for per-user memory isolation in the content writer example; Added tests for the Talon cron job scheduler and job store; Added unit and benchmark tests for QuickJS REPL middleware; Added unit and integration tests for the Runloop partner package; Added unit tests for Hooks v2; Added unit tests for TUI modal interactions; Added unit tests for Talon runtime components; Added unit tests for WhatsApp bridge ID compatibility and message handling; Added unit tests for client commands and shared fixtures; Added unit tests for deepagents backends; Added unit tests for deepagents core components; Added unit tests for evals analysis, assertions, and adapters; Added unit tests for skills CLI, loading, trust, and thread inspector; Added unit tests for the plugin manager modal; Added unit tests for tools module; Initial test suite for Talon agent runtime; New eval test suite for Deep Agents SDK; Unit tests added for deepagents middleware components; Unit tests for plugin marketplace and installation logic.
Dependencies
Dependency minimums raised across Deep Agents packages and examples
The Deep Agents SDK and its ecosystem packages now require newer minimum versions for core dependencies, including LangChain (\>=1.4.1), LangGraph (\>=1.2.8), and LangSmith (\>=0.12.6). The \deepagents-code\ terminal agent has been updated to require Python 3.12 and pins the SDK to version 0.7.15. Several partner integrations (Daytona, Modal, Runloop) now require \deepagents\>=0.7.0\. Additionally, security patches are applied via override dependencies in the evals package, specifically bumping \python-multipart\ to address [CVE redacted] and \protobuf\ to address [CVE redacted].
(dependencies) · high confidence
Housekeeping
New developer documentation for the \`libs/code\` package
Added \AGENTS.md\, \ARCHITECTURE.md\, \COMMANDS.md\, and \DEVELOPMENT.md\ to \libs/code\ to guide contributors. These files define the package's client/server architecture, establish Textual UI and input-surface coding conventions, document the slash command catalog, and provide a quickstart for local development and debugging.
libs/code · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 68.
Lenses
- Code Health 78
- Architecture 96
- Maturity 74
- Readiness 62
- Security 63
- Domain Modelling 100
Changes since last survey
- 300 commits — 175 feature/other, 125 fixes
By area
- libs/code — 121 commits
- libs/talon — 79 commits
- libs/partners — 23 commits
- libs/deepagents — 19 commits
- .github/scripts — 17 commits
- openwiki/.claims — 7 commits
- .github/workflows — 6 commits
- libs/evals — 6 commits
- libs/acp — 5 commits
- (root) — 3 commits
- examples/deep_research — 3 commits
- .github/RELEASING.md — 2 commits
- .github/actions — 2 commits
- .agents/skills — 1 commit
- .github/CODEOWNERS — 1 commit
- .github/PULL_REQUEST_TEMPLATE.md — 1 commit
- .github/SECRETS.md — 1 commit
- .github/topic-labels.json — 1 commit
- examples/nvidia_deep_agent — 1 commit
- examples/talon — 1 commit
Notable commits
- fix: fix(acp): scope cancel() to the requested session (#5107)
- fix: fix(ci): hide superseded release-note warnings (#6055)
- fix: fix(ci): skip repeated runs on curated changelogs (#6054)
- fix: fix(code): accept bare relative paths in marketplace plugin sources (#5959)
- fix: fix(code): align prompt clipboard word deletion (#6345)
- fix: fix(code): attribute dotenv config sources (#6222)
- fix: fix(code): bind request-time working directories (#5968)
- fix: fix(code): bound transcript tail reconciliation (#6057)
- fix: fix(code): cap MCP tool names for provider compatibility (#5953)
- fix: fix(code): clarify unknown effort default (#6337)
- fix: fix(code): complete ASCII UI fallbacks (#5930)
- fix: fix(code): confirm server restart for workspace switch (#6178)
- fix: fix(code): deduplicate footer picker clicks (#6344)
- fix: fix(code): defer the recursion limit to the LangGraph server (#5882)
- fix: fix(code): delay hook status in footer (#6322)
- fix: fix(code): demote no-output hint suppression to debug (#6245)
- fix: fix(code): disable Git terminal prompts in execute (#5878)
- fix: fix(code): drop stale Anthropic thinking blocks (#6300)
- fix: fix(code): expose unknown reasoning effort (#6241)
- fix: fix(code): guard path expansion in the Auto approval gate (#5941)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
langchain-ai/deepagents was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 80ea6ba41aadc0f88b6c843ae0e560039be417c9 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.