Skip to content
CAI
Software that uses CAICheck a score

huggingface/smolagents

76.9

Strong · 18 September 2026

26.1k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a Python library for building and orchestrating intelligent AI agents that execute code or call tools to perform tasks. It provides core agent classes, a local Python executor, and integration with various LLM providers and external services like web search and MCP servers. The project includes extensive examples demonstrating multi-agent orchestration, human-in-the-loop planning, asynchronous web integration, and benchmark evaluation.

Features

Initial release of smolagents v1.27.0.dev0

This change introduces the smolagents library, providing a framework for building and running intelligent agents. It includes core agent classes (CodeAgent, ToolCallingAgent), a local Python executor with security restrictions, a Gradio-based UI for interactive agent demos, and a CLI for quick setup. The library supports various model providers via InferenceClientModel, OpenAIModel, LiteLLMModel, and TransformersModel, and comes with default tools like web search and Python code execution.

src/smolagents · high confidence

Introduce Open Deep Research example with agent-based web research capabilities

Adds a new example in \examples/open\_deep\_research\ that replicates OpenAI's Deep Research functionality. The example provides a \run.py\ script and a \GradioUI\ app (\app.py\) to execute a multi-agent system: a manager \CodeAgent\ coordinates with a \ToolCallingAgent\ specialized in web browsing (using \GoogleSearchTool\ and a text-based browser) to answer complex research questions. It includes setup instructions, a Jupyter notebook for analyzing GAIA benchmark results, and a comparison notebook for visual vs. text-based browsing.

_examples/open\_deep\research · high confidence

Introduce Open Deep Research example with multi-modal agent scripts

Adds a new example in \examples/open\_deep\_research/scripts\ that provides a complete set of tools for an agentic research workflow. This includes a text-based web browser (\text\_web\_browser.py\) with cookie management (\cookies.py\) for navigating and searching the web, a document converter (\mdconvert.py\) for handling various file formats, and specific tools for visual question answering (\visual\_qa.py\) and text inspection (\text\_inspector\_tool.py\). The entry also adds a GAIA dataset scorer (\gaia\_scorer.py\) for evaluation, a reformulator (\reformulator.py\) to synthesize final answers from conversation history, and a runner (\run\_agents.py\) to orchestrate these components.

_examples/open\_deep\research/scripts · high confidence

New Plan Customization example with Human-in-the-Loop support

Added a new example in \examples/plan\_customization\ that demonstrates how to implement Human-in-the-Loop strategies for agent planning. The example shows how to use step callbacks to interrupt execution after a \PlanningStep\, allowing users to review, approve, modify, or cancel the agent's plan in real-time. It also illustrates how to preserve agent memory across runs using \reset=False\ to resume execution with the modified plan.

_examples/plan\customization · high confidence

New Smolagents Chat Server Demo

A new example server has been added to the \examples/server\ directory, providing a web-based chat interface for interacting with an AI code agent. The demo uses the \Qwen/Qwen3-Next-80B-A3B-Thinking\ model and integrates MCP tools via \MCPClient\. It is built with Starlette and AnyIO, featuring an asynchronous request handling loop and a responsive UI for sending messages and viewing agent responses.

examples/server · high confidence

New and updated agent examples with streaming, remote execution, and instrumentation

The examples directory now includes several new demonstration scripts: \agent\_from\_any\_llm.py\ shows how to configure agents with various inference providers (InferenceClient, Transformers, Ollama, LiteLLM, OpenAI); \inspect\_multiagent\_run.py\ demonstrates multi-agent orchestration with OpenTelemetry instrumentation; \multi\_llm\_agent.py\ illustrates load-balancing across LLMs using LiteLLM; \sandboxed\_execution.py\ provides examples for running agent code in remote sandboxes (Blaxel, Docker, E2B, Modal); \structured\_output\_tool.py\ shows integration with MCP servers for structured tool outputs; and \rag.py\ and \rag\_using\_chromadb.py\ offer updated Retrieval-Augmented Generation examples using lexical (BM25) and semantic search respectively. Existing examples like \gradio\_ui.py\, \multiple\_tools.py\, and \text\_to\_sql.py\ have been updated to reflect current APIs, including streaming outputs and new model defaults.

examples · high confidence

New async agent example using Starlette and CodeAgent

Added an example in the \examples/async\_agent\ directory demonstrating how to integrate a synchronous \CodeAgent\ from the \smolagents\ library into an asynchronous Starlette web application. The example shows how to offload the agent's execution to a background thread using \anyio.to\_thread.run\_sync\ to prevent blocking the async event loop, exposing a \/run-agent\ POST endpoint that accepts a task string and returns the result.

_examples/async\agent · high confidence

New smolagents benchmark runner and scoring notebook

Added a new benchmark execution script (\run.py\) and a scoring notebook (\score.ipynb\) for the \smolagents\_benchmark\ example. The runner script allows users to evaluate agents (CodeAgent, ToolCallingAgent, or vanilla LLM) against the \smolagents/benchmark-v1\ dataset using either \InferenceClientModel\ or \LiteLLMModel\, supporting parallel execution and result submission to the Hub. The accompanying notebook provides utilities to score these results against ground truth for tasks like GAIA, MATH, and SimpleQA.

_examples/smolagents\benchmark · high confidence

Test coverage

Initial test suite and test infrastructure

Establishes the foundational test infrastructure for the project, including pytest configuration, shared fixtures for agents and tools, and comprehensive test modules covering agents, CLI, default tools, final answer handling, Gradio UI, documentation code validation, and the local Python executor.

tests · high confidence

Dependencies

Initial release of smolagents with structured dependency management

The smolagents library is introduced as a barebones Python library for building agents that write code to call tools or orchestrate other agents. The package defines a core set of dependencies including huggingface-hub, requests, rich, jinja2, pillow, and python-dotenv, while organizing optional capabilities into extras such as bedrock, blaxel, docker, e2b, gradio, litellm, mcp, mlx-lm, modal, openai, telemetry, toolkit, transformers, vision, and vllm. Example projects for async agents and open deep research are provided with their own specific requirements files to support distinct use cases.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 77.

Lenses

  • Code Health 85
  • Architecture 100
  • Maturity 72
  • Readiness 72
  • Security 90
  • Domain Modelling 100

Changes since last survey

  • 300 commits — 210 feature/other, 90 fixes

By area

  • src/smolagents — 169 commits
  • docs/source — 63 commits
  • (root) — 39 commits
  • .github/workflows — 8 commits
  • .github/ISSUE_TEMPLATE — 3 commits
  • tests/test_remote_executors.py — 3 commits
  • examples/open_deep_research — 2 commits
  • examples/smolagents_benchmark — 2 commits
  • tests/test_agents.py — 2 commits
  • tests/test_models.py — 2 commits
  • .github/dependabot.yml — 1 commit
  • examples/async_agent — 1 commit
  • examples/multi_llm_agent.py — 1 commit
  • examples/plan_customization — 1 commit
  • examples/server — 1 commit
  • tests/test_default_tools.py — 1 commit
  • tests/test_telemetry.py — 1 commit

Notable commits

  • fix: BUG FIX: AmazonBedrockServerModel crashes when thinking mode is enabled (#1632)
  • fix: Bug Fix: Fix continue semantics of LocalPythonExecutor (#1645)
  • fix: Bug Fix: KeyError when agent reaches max steps with image input (#1525)
  • fix: Bugfix: remove duplicate line in prompt of tool calling agent (#1636)
  • fix: CI hotfix: Pin mlx < 0.26.5 (#1586)
  • fix: CI hotfix: Pin openai < 1.100.0 for litellm extra (#1693)
  • fix: CI hotfix: Pin transformers < 4.54.0 (#1620)
  • fix: Documentation: Minor fixes (#1809)
  • fix: Fix @tool decorator for remote Python executor (#1334)
  • fix: Fix Access of content Field in ChatMessage Object (#1533)
  • fix: Fix AmazonBedrockModel with reasoning/thinking content (#1681)
  • fix: Fix AttributeError when trying to log a None (#1786)
  • fix: Fix CI 403 error for Wikipedia page in test_visit_webpage (#1716)
  • fix: Fix CI AttributeError: 'str' object has no attribute 'module' (#2488)
  • fix: Fix CI LiteLLM test_call_different_providers_without_key (#1527)
  • fix: Fix CI PytestUnknownMarkWarning (#1630)
  • fix: Fix CI quality: remove trailing whitespace (#1617)
  • fix: Fix CLI Tool.from_space() call by auto-generating name and description (#1535) (#1859)
  • fix: Fix DockerExecutor connection reset error with server readiness check (#1684)
  • fix: Fix DockerExecutor tests with final_answer by calling send_tools (#1495)
  • …and 280 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

huggingface/smolagents was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 30bb1161095dbae2271e6bc3cc4c219cc3897a57 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.