bojieli/ai-agent-book
35.4
Weak · 18 September 2026
289k
lines of production code
Python
with JavaScript
1
measurement over time
What this system is
This repository is the source code and documentation for the open-source book 'AI Agents in Depth,' providing a comprehensive educational resource on building and evaluating AI agents. It contains runnable experiments across ten chapters that demonstrate core concepts such as context management, tool use, memory systems, retrieval-augmented generation, and multi-agent collaboration. The project also includes a multilingual static website reader, translation orchestration tools, and extensive validation suites to ensure the reproducibility of the experiments.
Features
Add Brazilian Portuguese translation of the AI Agents book
A complete Brazilian Portuguese (pt-BR) translation of the book is now available, including all ten chapters, the afterword, and reference answers. The translation is integrated into the book site and supports PDF generation via a dedicated build script.
book-ptbr · high confidence
Add Chapter 2 attention visualization frontend
A new Next.js-based frontend has been added to visualize attention mechanisms for Chapter 2. It provides a preview heatmap, a full interactive modal with zoom and transform controls (log, sqrt, power, etc.), and a statistics panel showing average attention, maximum attention, and entropy per token. The app loads agent trajectories (including ReAct reasoning steps) from local JSON files and displays prompts, responses, and token-level breakdowns.
_chapter2/attention\visualization/frontend · high confidence
Add Korean edition of the book
The book-ko directory now contains the complete Korean translation of the book, including all chapters, the afterword, and a script to build the PDF.
book-ko · high confidence
Add Russian translation (book-ru) and wire it into i18n
Introduces the Russian edition of the book, including the full translated content (chapters 1–10, afterword), a dedicated PDF build script (build\_pdf.sh) for generating the Russian PDF, and a .gitignore to manage build artifacts.
book-ru · high confidence
Add Traditional Chinese (Taiwan) translation of the book
Introduces a new Traditional Chinese (Taiwan) edition of the book, including all chapters, the afterword, and a PDF build script. This adds a new language variant to the existing multilingual support, ensuring the content is synchronized with the Chinese source material.
book-zhtw · high confidence
Add complete Arabic edition of the book
Introduces the full Arabic translation of the book, including all ten chapters, the afterword, and a build script to generate the PDF. This completes the multilingual support for the publication.
book-ar · high confidence
Add text-to-image workflow experiment with evidence tracking
Introduces a new experiment in \chapter1/image-gen-workflow\ that compares three image generation routes: a workflow using a Kimi LLM to rewrite prompts for DashScope's Wanx model, and two native routes using Google's Gemini 3 Pro Image and OpenAI's GPT-Image 2. The tool runs five specific and broad Chinese language requirements across these routes, capturing detailed API call logs, generated images, and SHA-256 hashes in a structured evidence manifest for reproducibility and validation.
chapter1/image-gen-workflow, chapter5/cad-vs-diffusion · high confidence
Added conversational UI validation run with FastAPI backend and React frontend
This change introduces a new validation run for the conversational UI experiment, providing a complete full-stack chatbot application. The backend is a FastAPI service that supports both a default echo mode and an optional LLM mode (using OpenAI or OpenRouter APIs) to generate real responses. The frontend is a React application that communicates with the backend via a Vite development server, which proxies API requests to the backend. The UI is designed to be customizable via natural language, with UI text and visual styles (colors, fonts) separated into distinct source files to facilitate theme and content modifications.
_chapter5/conversational-ui/validation/runs/20260729T212616Z-5\_11-hmr, chapter5/conversational-ui/validation/runs/20260729T212727Z-5\_11-hmr, chapter5/conversational-ui/validation/runs/20260729T212933Z-5\11-hmr · high confidence
Agent Trajectory Replay Component and Schema
The Agent Lab now includes a self-contained \\<agent-trajectory\>\ Web Component that replays an agent's ReAct loop step-by-step in the browser. This component relies on a new JSON schema (defined in SCHEMA.md) that structures agent runs into thought, action, observation, and answer steps. The entry also includes sample trajectory data files demonstrating successful agent runs and ablation studies (e.g., missing tool results), allowing users to visualize and debug agent behavior directly in documentation.
extras/agent-lab · high confidence
Chapter 1 context experiment: new multi-provider agent with ablation study and grounding checks
The chapter1/context directory now contains a complete, runnable context-aware agent experiment (Experiment 1-1) that supports multiple LLM providers (Doubao, SiliconFlow Qwen, Alibaba Cloud DashScope/Qwen, Kimi K3, DeepSeek, Zhipu GLM) with a universal OpenRouter fallback, and introduces a five-arm ablation study (full, no history, no reasoning, no tool calls, no tool results) to measure how each context component affects agent behavior. The runner persists credential-free API request/response evidence for every turn, and a new grounding module assesses whether the model's numerical answers are supported by the observations it actually saw, reporting outcomes as completed, unsupported numbers, or no terminal response. The experiment includes interactive and batch modes, sample PDF fixtures, and a validation report visualizer to compare behavior across context modes.
chapter1/context · high confidence
Chapter 1 learning-from-experience experiment: RL vs LLM in-context learning on a treasure-hunt game
This location introduces the complete codebase for Experiment 7-2, which compares traditional tabular Q-learning against LLM-based in-context learning on a text-based treasure-hunt game with hidden mechanics. The package includes the game environment, a Q-learning agent, and an LLM agent that supports Kimi K3 (Moonshot), Alibaba Cloud Qwen (Dashscope), and a universal OpenRouter fallback. It provides a CLI experiment runner (\experiment.py\) for full RL/LLM comparison, a quick demo (\quick\_demo.py\) to visualize LLM reasoning, an interactive manual demo (\demo.py\), and a canonical runner (\run\_experiment\_8\_2.py\) with a post-processing validator (\finalize\_experiment\_8\_2.py\) that enforces strict protocol gates and writes structured evidence. The setup uses \uv\ with a \ch1\ extra, loads \.env\ via \python-dotenv\, and guards against division-by-zero errors in empty evaluation windows and zero-episode victory counts.
chapter1/learning-from-experience · high confidence
Chapter 3 documentation and experiment ledger added
This location now includes the Chapter 3 README in multiple languages (Arabic, English, Spanish, Hungarian, Indonesian, Japanese, Korean, Chinese, Russian, Tamil, Turkish, Vietnamese) and a new EXPERIMENT\_LEDGER.md. The ledger defines 12 experiments (3-1 through 3-12) covering user memory, log sanitization, dense/sparse embeddings, retrieval pipelines, structured indexing, agentic RAG, contextual retrieval, and structured knowledge extraction, with acceptance gates and evidence paths for each.
chapter3 · high confidence
Coding agent now supports OpenRouter as a universal fallback provider
The coding agent in chapter5/coding-agent now accepts OpenRouter alongside Anthropic and OpenAI. Users can set PROVIDER=openrouter with an OPENROUTER\_API\_KEY to access multi-model capabilities. Additionally, if a direct provider key (Anthropic or OpenAI) is missing but an OPENROUTER\_API\KEY is present, the agent automatically falls back to OpenRouter, mapping native model IDs (e.g., claude-sonnet-\, gpt-\*) to their OpenRouter equivalents via config.py.
chapter5/coding-agent · high confidence
Contextual Retrieval system with semantic IDs and improved chunking
The chapter3/contextual-retrieval area now provides a complete educational implementation of Anthropic-style contextual retrieval, where LLM-generated context prefixes are prepended to document chunks before indexing to improve BM25 and embedding recall. The system introduces human-readable semantic document IDs (derived from filenames) to replace opaque MD5 hashes, making logs and data easier to debug. It also includes a fix to the sentence-splitting logic in chunking.py to preserve trailing text fragments, and adds a universal OpenRouter fallback in config.py to ensure model availability across providers.
chapter3/contextual-retrieval · high confidence
Experiment 2-6 now supports runtime-agnostic PPT generation via Kimi Code CLI
The Agent Skills PPT demo in chapter2/agent-skills-ppt has been updated to support running the official Anthropic PPTX Skill with Kimi Code CLI, in addition to the existing Claude Code path. This change introduces a new \run\_official\_experiment.py\ script and \experiment\_protocol.json\ that define a runtime-agnostic acceptance policy, allowing users to generate presentations from academic papers using either Anthropic's Claude Code or Moonshot's Kimi Code CLI. The implementation includes a \prepare\_official\_skill.py\ script to fetch and verify the pinned official Skill repository, and updates the \demo.py\ legacy path to use \gpt-5.6-luna\ via OpenRouter when available. Users can now choose their preferred runtime via the \--runtime\ flag (claude or kimi), enabling broader access to the experiment without requiring Anthropic credentials.
chapter2/agent-skills-ppt · high confidence
Hebrew edition of the book is published
The Hebrew translation of the book is now available, including all ten chapters, an afterword, and a script to build a PDF version.
book-he · high confidence
Hungarian edition of the book is published
This change introduces the complete Hungarian translation of the book, including all ten chapters, the afterword, and the PDF build script. The edition localizes text, figure labels, and vector-origin figures into Hungarian, providing a fully translated reading experience for Hungarian-speaking users.
book-hu · high confidence
Indonesian translation of the book and build tooling
Adds the complete Indonesian translation of the book (book-id), including all chapters, the afterword, and the introduction. Includes a build\_pdf.sh script to generate the PDF using Pandoc and XeLaTeX, and a .gitignore to exclude generated artifacts.
book-id · high confidence
Initial Turkish translation of the book
Adds the Turkish (tr) edition of the book, including all chapter source files (chapter1.tr.md through chapter10.tr.md), the afterword, a PDF build script (build\_pdf.sh), and a .gitignore. This provides a complete, localized version of the content for Turkish-speaking users.
book-tr · high confidence
Introduce API-driven smart video editing with Blender and ffmpeg backends
Adds a new Experiment 5-6 module that automates video editing via a multi-agent workflow: a Vision-based sub-agent locates scene boundaries using a two-step sampling approach, a Proposer agent generates a Blender Python API (bpy) script to define cuts, subtitles, and slow-motion effects, and a Reviewer agent validates the result. The system supports both Blender (for 3D/compositing) and ffmpeg (as a lightweight fallback) as execution backends, includes a smoke-test mode for offline verification, and provides bilingual documentation and CLI flags for model and backend selection.
chapter5/video-edit · high confidence
Introduce Agentic RAG experiment with ReAct reasoning and multi-provider support
Adds Experiment 3-8, a new Agentic RAG system that compares single-shot retrieval against a ReAct-style agent capable of iterative tool use (knowledge base search and document retrieval) to improve evidence recall on complex Chinese legal queries. The experiment includes a fully offline BM25 comparison script (\compare\_offline.py\) that requires no API keys, alongside a live campaign runner (\campaign.py\) for reproducible, audited evaluations. The system supports a wide range of LLM providers (Kimi, OpenAI, Alibaba Cloud Qwen, Doubao, OpenRouter, Groq, etc.) with automatic fallback routing through OpenRouter, and includes a legal document indexing script (\index\_local\_laws.py\) with smart paragraph-aware chunking for local knowledge bases.
chapter3/agentic-rag · high confidence
Introduce English edition of the book with PDF build script and afterword
Adds the English translation of the book (book-en), including the full chapter content, a new afterword, and a build\_pdf.sh script to generate the PDF using Pandoc and XeLaTeX. The afterword discusses the co-evolution of models and agents, while the build script handles the compilation of all chapters into a single PDF document.
book-en · high confidence
Introduce Experiment 5-11: Conversational UI Customization
Adds a new lab (Experiment 5-11) that enables users to customize a React/Vite front-end interface through natural language requests. An agent reads the current source files and uses function calling to rewrite specific editable files (such as \src/App.jsx\ and \src/theme.css\), which are then applied instantly via Vite's Hot Module Replacement (HMR). The package includes a FastAPI backend for chat interactions, a baseline snapshot for reproducible runs, and automated verification scripts (\demo.py\, \campaign\_browser.py\) that validate the changes and ensure the application builds successfully.
chapter5/conversational-ui · high confidence
Introduce Experiment 5-1: Code-Assisted Math Solving with AIME 2024 Benchmark
This location adds the complete source and data for Experiment 5-1, which compares pure Chain-of-Thought reasoning against code-assisted solving on a 30-problem AIME 2024 benchmark. The \demo.py\ script orchestrates the comparison using a subprocess sandbox (\sandbox.py\) that executes generated Python code (via sympy, numpy, scipy) to provide exact mathematical results, while \problems.json\ provides a teaching-grade dataset with reference solutions. The entry also includes \build\_aime\_2024.py\ for generating the pinned benchmark from HuggingFace, \test\_campaign.py\ for validating the campaign logic, and \validation/\ files containing the official run evidence (using \doubao-seed-1-6-flash-250615\) which confirms the experiment's execution and statistical results.
chapter5/code-for-math · high confidence
Introduce Experiment 5-8: automated production log diagnosis and regression testing
Adds a new lab in chapter5/log-diagnosis that automates the process of diagnosing production agent trajectories against system architecture and PRD requirements. The pipeline reads trajectory logs, uses an LLM to identify issues (such as missing pre-checks, retry failures, or latency timeouts), generates regression test cases, and executes them against a deterministic system-under-test to verify bug reproduction and fix validation. It also includes a mock GitHub Issue creation step and a live HTTP campaign for real-world validation.
chapter5/log-diagnosis · high confidence
Introduce active tool discovery with hierarchical semantic routing
Adds a new active tool selection experiment in chapter 4 that implements MCP-Zero-style on-demand tool discovery. The system uses a two-stage hierarchical semantic router (server-level then tool-level matching via TF-IDF) to find relevant tools from a knowledge base of 8 servers and 35 tools, allowing agents to iteratively request and load only the tools they need. This replaces the traditional passive approach of injecting all tool schemas into the prompt, significantly reducing context token overhead while maintaining retrieval recall. The location includes the core agent and router implementations, a tool knowledge base, a deterministic offline benchmark for comparing all-tools vs. retrieval strategies, and demo scripts for side-by-side comparison.
chapter4/active-tool-selection · high confidence
Introduce book translation orchestration with multi-agent collaboration and bilingual consistency auditing
This location adds a complete book translation experiment (Experiment 10-2) that demonstrates an orchestration pattern using four specialized agents (Glossary, Translation, Proofreading, and Manager) to translate long documents while controlling context growth and enforcing terminology consistency via a shared glossary. The implementation includes a demo script with a dry-run mode for offline architecture visualization, a bilingual consistency auditor module for checking terminology mapping, code block synchronization, and LaTeX formula preservation, and robust handling of edge cases such as null or non-dict proofread reports and glossary entries.
chapter10/book-translation · high confidence
Introduce context compression strategy comparison experiment and benchmarking module
This location adds the complete codebase for Chapter 2 Experiment 2-10, which compares six context compression strategies (no compression, individual/combined summaries, context-aware summaries, context-aware with citations, and windowed context) for LLM agents. The \agent.py\ and \compression\_strategies.py\ files implement the \ResearchAgent\ and \ContextCompressor\ classes, handling tool execution, trajectory tracking, and compression logic with support for streaming and reasoning-model temperature constraints. The \config.py\ module centralizes configuration, including a new OpenRouter fallback for when the primary Kimi/Moonshot key is missing, and integrates with \agentbook.providers\ for backend resolution. The \experiment.py\ and \run\_all\_strategies.py\ scripts automate the comparison, outputting metrics like token usage, compression ratios, and overflow counts to JSON results. Additionally, \benchmark\_compression.py\ provides a separate harness to evaluate summary, truncation, key-sentence, and observation-filtering strategies on metrics like TTFT and QA retention accuracy.
chapter2/context-compression · high confidence
Introduce legal-domain evaluation framework and dataset builder
The evaluation module now includes a dataset builder that generates a structured evaluation set of Chinese criminal law questions (simple and complex cases) and an evaluation runner that compares agentic versus non-agentic RAG performance using keyword and analysis recall metrics. A new offline QA dataset is provided for reproducible, LLM-free retrieval comparison, and the evaluation logic now safely handles empty expected keywords or analysis lists to prevent division-by-zero errors.
chapter3/agentic-rag/evaluation · high confidence
Introduce proactive tool discovery experiment for Chapter 4
Adds the 'active-tool-discovery' experiment to Chapter 4, which compares three strategies for managing large tool libraries (126 tools): full injection, retrieval prefilter, and active discovery. The active discovery strategy uses semantic embedding to dynamically retrieve and inject only relevant tool schemas on-demand, reducing token usage and improving selection accuracy compared to injecting all schemas at once. The package includes the core agent loop, discovery index, demo runner, and a formal experiment runner that validates the approach against real MCP tools using the Qwen3-4B model.
chapter4/active-tool-discovery · high confidence
Introduces Chapter 5 'Paper-to-PPT' experiment with Proposer-Reviewer dual-agent workflow
Adds a new experiment in \chapter5/paper-to-ppt\ that generates Slidev presentations from academic papers using a dual-agent architecture. A Proposer agent writes the Slidev source code based on paper text and structured text feedback, while a Reviewer agent renders each page to a PNG and uses a Vision LLM to inspect for layout issues like text overflow or overcrowding. The entry includes the core implementation (\demo.py\, \agents.py\, \renderer.py\), a \paper\_source.py\ module for downloading and extracting content from a pinned real PDF, and validation runs comparing this dual-agent approach against a single-agent self-review baseline to demonstrate reduced context window usage.
chapter5/paper-to-ppt · high confidence
Introduces Experiment 5-3: Codified Rules for Small Models
Adds a new experiment in \chapter5/small-model-codified-rules\ that demonstrates how codifying business rules (specifically an airline refund policy) into code-based tool guards allows a small model (e.g., \qwen3:4b\ or \gpt-5.6-luna\) to match the reliability of a large model. The directory includes a controlled simulation environment (\airline\_env.py\) with a 'control' arm (naive tool execution) and a 'codified' arm (server-ground-truth gatekeeper), along with a CLI runner (\demo.py\) and a 60-case factorial test matrix (\tasks.py\) to validate that code-based checks prevent policy violations even when the model's self-reported beliefs are incorrect.
chapter5/small-model-codified-rules · high confidence
Introduces Permission-Embedded Data Objects (PEDO) prototype with BaxBench security evaluation
Adds the Permission-Embedded Data Objects (PEDO) research prototype for Chapter 5, featuring a Python middleware layer over PostgreSQL that enforces authorization and data integrity at the data type level. The package includes core models and a three-tier object store (\pedo/core/\) that handles permission checks, validators, and asynchronous reactions, along with a deterministic demo (\demo.py\) and multi-tenant hiring scenarios. It also introduces a BaxBench-derived security evaluation harness (\pedo/eval/baxbench\_adapter/\) that adapts backend scenarios (SecretStorage, UserCreation, ShoppingCart, ImageTransfer) to test PEDO against raw SQL and insecure implementations, alongside a comprehensive end-to-end hiring pipeline benchmark (\pedo/eval/benchmark\_application.py\).
chapter5/permission-embedded-data-objects · high confidence
Introduces self-healing adaptive log parser with live campaign and validation
Adds a new self-evolving log parsing system that automatically detects unparseable log formats, uses an AI agent to generate and test new parser code, and hot-loads it into the engine for immediate and persistent use. The package includes a live campaign script that generates real-time log streams, validates the parsing results via a browser-based visualization, and records detailed evidence of the self-healing loop in JSON validation files.
chapter5/adaptive-log-parser · high confidence
Introduces voice-enabled multi-agent Werewolf game with strict information isolation and auditability
This change adds a new multi-agent Werewolf simulation system (Experiment 10-6) that supports real-time voice interaction for human players and LLM-driven user simulators. The system enforces strict information isolation by maintaining private, player-specific memory contexts, ensuring that sensitive information (like teammate identities or investigation results) is only delivered to authorized agents. It includes a deterministic Judge orchestrator, an audit log to verify information visibility, and a strategy acceptance audit to validate agent behavior. The implementation supports multiple LLM backends (OpenAI, OpenRouter, ARK, Moonshot) with fallbacks, and provides both offline rule-based strategies and live voice sessions with ASR/TTS capabilities.
chapter10/voice-werewolf/werewolf · high confidence
Japanese edition of the book is added
The Japanese translation of the book (book-ja) is now available, including all chapters (1-10), an afterword, and a script to build the PDF.
book-ja · high confidence
New Astro-based web reader for the AI Agents book
A new static-site reader built with Astro has been added to the project, providing a web interface for the 'AI Agents in Depth' book. This reader supports all 10 chapters across 15 language editions (including English, Chinese, Spanish, and others) and is published on GitHub Pages. It features a persistent reading bar, light/dark themes, adjustable text sizing, and focus mode. Users can highlight text, save notes, and manage highlights via IndexedDB with JSON backup and import capabilities. The reader also includes a language switcher that preserves reading position, figure zooming, and code block copying. The implementation includes specific web-optimized SVG figures for Chapter 2 and a custom Markdown processing pipeline to handle math (KaTeX), footnotes, and code highlighting.
web-astro · high confidence
New Chapter 8 experiment documentation and evaluation tooling
This update introduces comprehensive documentation and evaluation assets for several Chapter 8 experiments. It adds reproduction guides for the AWorld training framework (Experiment 8-15) and the AdaptThink adaptive reasoning method (Experiment 8-10), including a detailed training report with specific W&B run metrics and hardware configurations. It also provides the full README for the Intuitor unsupervised reinforcement learning experiment (Experiment 8-11). Additionally, a new Python evaluation script (\evaluate\_from\_cache.py\) is added to the Intuitor directory to extract and normalize mathematical answers from model outputs, accompanied by a suite of regression tests to ensure correct handling of nested LaTeX, negative currency, and fraction formats.
chapter8 · high confidence
New Collaboration Tools MCP Server with sub-agent, HITL, and notification capabilities
This location introduces the Collaboration Tools MCP Server, a comprehensive implementation providing 19 tools across five categories: browser automation (using browser-use and Playwright), sub-agent management (with sync/async modes and minimal/llm\_generated context strategies), human-in-the-loop (HITL) approval workflows, multi-channel notifications (Email via SMTP/SendGrid, Telegram, Slack, Discord), and timer/scheduling (one-time and recurring with persistent JSON storage). The server is built on FastMCP and includes a unified CLI entry point (main.py) for listing, calling, and demoing tools without starting the MCP server, alongside a Dockerfile for containerized deployment and extensive documentation.
chapter4/collaboration-tools · high confidence
New ERP Agent experiment (5-10) with CLI and multi-database support
Adds a new experiment (5-10) that converts natural language queries into SQL artifacts for an ERP dataset, executed by the database rather than the LLM. The \demo.py\ CLI provides \run\ (online validation), \gold\ (offline SQL verification), \ask\ (interactive query), and \initdb\ (seed data). The agent uses \gpt-5.6-luna\ by default, supports OpenRouter fallback, and includes a \campaign\_postgres.py\ script for running the same logic against PostgreSQL.
chapter5/erp-agent · high confidence
New English video course slides and generation tooling
The slides directory now contains the source files for a 42-lesson English video course on AI Agents, including a course outline, a Slidev-based generator (course.mjs, generate.mjs), and build scripts (build-all.mjs). The generator creates individual lesson decks (e.g., lesson-01.md) from structured metadata, handling slide layout, demo budgets, and PDF/PNG exports. A .gitignore file is added to exclude build artifacts and node\_modules.
slides · high confidence
New Execution Tools MCP server with multi-language support and safety mechanisms
Adds a new MCP server for Chapter 4 that provides execution tools with built-in safety mechanisms, including LLM-based approval for dangerous operations, automatic syntax verification, and long-output truncation with persistence. The server supports multiple programming languages (Python, JavaScript, TypeScript, Go, Java, C++, Rust, PHP, Bash) and includes file system tools, a code interpreter, a virtual terminal, and external integrations for Google Calendar and GitHub. It features a universal OpenRouter fallback for LLM operations when the primary provider's key is missing, and can be run via CLI or as an MCP server with stdio transport.
chapter4/execution-tools · high confidence
New Experiment 5-13: Automated Agent Creation and Validation
Added a new runnable experiment in \chapter5/agent-creator\ that compares two strategies for generating specialized AI Agents: creating one from scratch versus adapting a proven reference template. The tool uses a real LLM to generate the Agent loop, tools, and tests, then runs a suite of structural, compilation, and live-runtime validation gates to determine which strategy is more efficient and reliable. It supports multiple OpenAI-compatible backends (Moonshot/Kimi, Volcengine Ark, OpenAI, OpenRouter) and outputs a detailed comparison report.
chapter5/agent-creator · high confidence
New Memobase Agent with LOCOMO Benchmark and Multi-Provider Support
This location introduces a complete Memobase-inspired memory agent system for Chapter 3, featuring a hand-rolled \MemoryStore\ with episodic, semantic, procedural, and working memory types, persistence via pickle, and compression/consolidation strategies. The agent supports interactive, benchmark, demo, and task modes via \main.py\, and includes a \LOCOMOBenchmark\ suite for evaluating long-context and memory-intensive tasks. Configuration (\config.py\) now supports multiple LLM providers: default Kimi K3, Alibaba Cloud DashScope (Qwen) via \DASHSCOPE\_API\_KEY\, and a universal OpenRouter fallback when primary keys are missing. A separate \profile\_demo.py\ demonstrates the official Memobase SDK (Profile + Event Memory) against a running server or cloud instance. The entry also includes regression tests for empty query handling in the benchmark evaluator.
chapter3/memobase · high confidence
New Vietnamese edition of the book
Adds a complete Vietnamese translation of the book (book-vi), including all chapter source files, an afterword, and a build script to generate the PDF.
book-vi · high confidence
New agentic RAG user-memory experiment with LLM evaluation and hybrid retrieval
Adds a new Chapter 3 experiment (chapter3/agentic-rag-for-user-memory) that implements agentic multi-hop retrieval over conversation memory. The system chunks conversations into \~20-round segments with overlap, indexes them via a local BM25 fallback or an external retrieval pipeline on port 4242 (dense + sparse hybrid), and uses a ReAct agent with tools (search\_memory, get\_conversation\_context, get\_full\_conversation) to answer questions. It integrates automatic LLM-based evaluation (reward 0–1, pass/fail at ≥0.6, with reasoning and required-information checks) via the week2/user-memory-evaluation module, and provides a controlled campaign (campaign.py) over 60 YAML test cases across three layers. The experiment supports multiple LLM providers (Kimi K3, GPT-5.6-luna, Qwen, etc.) with OpenRouter as a universal fallback, and includes offline-demo mode that requires no API keys or external services.
chapter3/contextual-retrieval-for-user-memory · high confidence
New attention visualization tools and experiments for Chapter 2
Adds a complete attention visualization suite for Chapter 2, Experiment 2-2, including a standalone CLI (\attention\_cli.py\) to render self-attention heatmaps (demonstrating the attention sink and causal triangle), an interactive React frontend for viewing agent trajectories, and a ReAct agent (\main.py\) with tool-calling support. The directory also introduces canonical experiment runners (\run\_attention\_experiment.py\, \run\_status\_bar\_experiment.py\) with preregistered protocols (\attention\_experiment\_protocol.json\, \status\_bar\_protocol.json\) to capture and verify real-model attention matrices, along with regression tests for trajectory saving and experiment logic.
_chapter2/attention\visualization · high confidence
New browser-agent example applications and utility modules
This update introduces several new example applications and supporting code. In the frontend utilities, a new \cn\ helper is added to merge Tailwind CSS classes. For the localized timeout task, a Python application package is added with configuration for resolving agent timeouts and a worker module for handling retries. Additionally, new browser-agent demos are provided: an ad generator that creates Instagram and TikTok ads from landing pages, a WhatsApp message scheduler with persistent login and auto-response capabilities, and a news monitor that extracts and analyzes articles with sentiment detection.
(repo-wide) · high confidence
New build and site-generation tooling for the MkDocs documentation site
This change introduces a comprehensive suite of Python and shell scripts to assemble, validate, and optimize the online documentation site. The \build\_site.sh\ script orchestrates the copying of Markdown sources, assets, and translated editions into the \\_web/\ output directory, while \clean\_site\_files.py\ prunes unnecessary files to keep the site small. New hooks handle specific site behaviors: \mkdocs\_pandoc\_strip.py\ removes Pandoc-specific attributes, \git\_revision\_dates.py\ ensures accurate file modification dates, \site\_edit\_urls.py\ fixes broken 'edit this page' links, and \seo\_meta.py\ injects Open Graph and Twitter Card meta tags for rich social sharing. Internationalization is supported by \site\_i18n.py\, which builds browser-side translation catalogs, and \check\_i18n\_consistency.py\, which enforces structural parity across all language editions. Additionally, \split\_search\_index.py\ divides the large search index into per-edition files to improve load times, and \gen\_star\_history.py\ replaces the external star-history.com chart with a self-hosted, GraphQL-powered matplotlib visualization.
scripts · high confidence
New dense embedding service with ANNOY/HNSW comparison and CLI metrics
Adds a new educational HTTP service and CLI tool in chapter3/dense-embedding for Experiment 3-4, enabling dense vector similarity search using the BGE-M3 model with configurable ANNOY or HNSW backends. The service exposes REST endpoints for indexing, searching, deleting, and retrieving document statistics, while the included cli.py allows offline, reproducible evaluation of retrieval quality (recall@k, precision@k, MRR) and direct comparison of ANNOY versus HNSW index performance (build time, query latency, recall) using synthetic vectors or a small cached model.
chapter3/dense-embedding · high confidence
New dynamic form intent clarification experiment (5-9) with offline schema and live browser validation
This location introduces Experiment 5-9, a new capability where the Agent dynamically generates a self-contained HTML form to clarify incomplete user intent in a single submission, rather than asking questions one by one. The form includes cascading logic (e.g., showing return date only for round-trip, updating baggage options based on cabin class) and submits data as JSON. The \demo.py\ script supports both online generation via LLMs (with OpenRouter fallback) and deterministic offline rendering using a built-in flight schema. Additionally, \campaign.py\ provides a live validation path using Playwright/Chromium to execute the form's JavaScript and verify the end-to-end flow, while new tests ensure correct handling of numeric zero baggage counts and prevent trip-type modifiers from contaminating the destination city field.
chapter5/dynamic-form · high confidence
New educational BM25 sparse search engine with CLI, API, and benchmarks
Adds a from-scratch BM25 search engine implementation in chapter3/sparse-embedding, featuring an inverted index, advanced tokenization (handling numbers, codes, technical terms, mixed case, and apostrophes), and a FastAPI HTTP server with an interactive web UI. The change includes a fully offline CLI tool for querying and evaluating recall/precision/MRR on a built-in corpus, a benchmark script that validates scoring against hand calculations and measures exact-keyword vs. synonym retrieval behavior, and a quickstart/demo script for interactive exploration. The engine supports BM25 and SPLADE (learned sparse) retrieval methods, with configurable k1/b parameters and detailed educational logging.
chapter3/sparse-embedding · high confidence
New experiment harness for logic reasoning with Knights & Knaves CSP solver
This location introduces the complete code and data infrastructure for Experiment 5-2, which evaluates whether code-assisted constraint solving improves logical reasoning on Knights & Knaves puzzles. The package includes a deterministic offline CSP solver (\csp\_solver.py\) using \python-constraint\, a sandboxed code interpreter (\sandbox.py\) for executing model-generated Python, and a main demo script (\demo.py\) that compares pure natural-language reasoning against code-assisted approaches. It ships with a curated set of 12 puzzles (\puzzles.json\) and a stratified 84-puzzle test set (\hf\_test\_stratified\_84.json\) sourced from the \K-and-K/perturbed-knights-and-knaves\ Hugging Face dataset, along with scripts to build and validate these datasets. The experiment also includes comprehensive test coverage for puzzle generation, answer parsing (including boolean JSON values), and silent resident handling.
chapter5/code-for-logic · high confidence
New experiments for cross-provider trajectory handoff and streaming interruption recovery
Added a new experimental module in chapter5/provider-failover that implements two scenarios for handling AI agent failures: Experiment 5-1 tests cross-provider trajectory handoff (switching from one model provider to another mid-execution) using three strategies (direct pass-through, stripping reasoning, and a neutral format), while Experiment 5-2 tests recovery from streaming output interruptions (cutting off at reasoning, text, or tool argument boundaries) using three recovery methods (full resend, prefix continuation, and meta-instruction). The module includes a neutral trace format to standardize reasoning and tool calls across Moonshot, Anthropic, and Gemini, along with minimal HTTP clients, renderers, streaming cut-off logic, and offline tests to validate the rendering and interruption logic without hitting external APIs.
chapter5/provider-failover · high confidence
New hybrid retrieval pipeline with offline evaluation and campaign support
Introduces a complete hybrid retrieval pipeline for Chapter 3 that combines dense (BGE-M3), sparse (BM25), fusion (RRF and weighted), and neural reranking (BGE-Reranker-v2). The location provides a FastAPI service (main.py) for real-time search, an offline evaluation script (evaluate.py) for benchmarking recall/MRR/nDCG without external services, a canonical campaign runner (campaign.py) for reproducible evidence, and educational demos (demo.py). It also includes the core pipeline logic, document store, and fusion strategies.
chapter3/retrieval-pipeline · high confidence
New judicial case analysis pipeline with offline reproducibility
Adds a complete four-stage experiment for extracting latent knowledge from structured legal data: bottom-up factor discovery, structured extraction, archetype clustering, and a conversational sentencing-advice Agent. The directory now includes a canonical \campaign.py\ that deterministically samples from the official CAIL2018 dataset for reproducible validation, alongside a legacy synthetic demo (\demo.py\) for quick local walkthroughs. Robustness improvements include explicit error handling for empty inputs and zero-archetype scenarios, plus a \.gitignore\ to keep generated artifacts and large upstream archives out of version control.
chapter3/structured-knowledge-extraction · high confidence
New log sanitization experiment with offline regex and local LLM engines
Adds a new Chapter 3 experiment that detects and redacts secrets and PII from agent logs using two complementary engines: a deterministic offline regex engine (covering API keys, JWTs, credentials, and PII like credit cards and IDs) and a local LLM engine (defaulting to Ollama with qwen3:0.6b, with an OpenRouter fallback) for semantic Level 3 PII detection. The experiment includes a benchmarking campaign to compare regex, LLM, and hybrid approaches, along with performance metrics and validation evidence.
chapter3/log-sanitization · high confidence
New multi-role transfer experiment comparing system-prompt handoffs versus skill loading
Adds a new Chapter 10 experiment (chapter10/multi-role-transfer) that runs a controlled A/B comparison of two multi-role agent architectures over the same shared conversation history. Path 1 uses a system-prompt transfer mechanism where the orchestrator swaps the system prompt and tool set on each autonomous handoff via a \transfer\_to\_agent\ tool. Path 2 uses a skill-loading mechanism where a fixed system prompt and tool catalog remain in place, and role capabilities are appended as \SKILL.md\ tool results via a \load\_skill\ tool. The package includes a demo entry point, a paired comparison runner, a blind position-swapped quality judge, deterministic evaluation rubrics, and an experiment protocol to measure cost, latency, and instruction-following across both arms.
chapter10/multi-role-transfer · high confidence
New multimodal agent experiment for Chapter 4
Adds a new \chapter4/multimodal-agent\ directory containing Experiment 4-3, which compares three extraction paradigms: native multimodal processing, extract-to-text, and tool-based analysis. The implementation includes a \MultimodalAgent\ core (\agent.py\) supporting PDF, image, and audio inputs via Gemini, OpenAI, and Doubao providers, with a configurable OpenRouter fallback. A \campaign.py\ script provides a live, checkpointed comparison of these modes against specific visual data questions, while \create\_sample.py\ generates offline test assets (charts and reports) to measure extraction fidelity without API keys. The package also includes interactive and demo entry points (\main.py\, \demo.py\) and regression tests for tool execution robustness and interactive toggles.
chapter4/multimodal-agent · high confidence
New paper-to-video experiment with guided narration and visual review
The chapter5/paper-to-video location introduces Experiment 5-5, a pipeline that generates narrated lecture videos from paper slides. It features a formal campaign runner (campaign.py) that uses Kimi K3 for narration, Qwen-VL-Max for independent visual review, and Fish Audio S1 for TTS, producing H.264/AAC videos with strict duration and drift gates. A legacy demo (demo.py) provides a self-contained path using OpenAI APIs or offline placeholders. The location also includes a protocol definition (experiment\_protocol.json), validation artifacts, and tests for robustness (e.g., null bullet handling, ffprobe N/A errors).
chapter5/paper-to-video · high confidence
New parallel real-browser research experiment with multi-provider LLM fallback
Added Experiment 10-4, a new multi-agent research demo that launches independent Playwright Chromium browser contexts to scrape live university faculty pages in parallel, using an async message bus for coordination and a cascade-terminate mechanism to stop workers once the first target is found. The experiment includes a CLI demo, a provenance-complete official runner that records raw browser observations and LLM receipts, and a multi-provider LLM extraction layer that falls back through Volcengine ARK, Moonshot (kimi-k3), OpenAI, and OpenRouter if earlier endpoints fail. Validation data and tests confirm measured parallel speedup over serial execution, single-locked settlement, and proper resource cleanup.
chapter10/parallel-web-research · high confidence
New prompt engineering ablation study framework for Chapter 2
This location introduces a complete experimental framework for the Chapter 2 prompt engineering ablation study. It adds an \AblationAgent\ that supports tone modifications (Trump-style, casual, or default), wiki rule randomization, and tool description removal to evaluate their impact on agent task success. The suite includes a runner (\run\_ablation.py\) and analyzer (\analyze\_results.py\) to execute these experiments against the tau-bench airline environment, with a frozen protocol (\experiment\_protocol.json\) defining the specific arms and metrics. It also provides configuration guidance for GPT-5 via OpenRouter to minimize token usage through \reasoning\_effort\ settings, along with necessary project metadata (LICENSE, MANIFEST.in, .gitignore) and bilingual documentation.
chapter2/prompt-engineering · high confidence
New prompt injection attack and defense experiment for Chapter 2
This location introduces a complete, runnable experiment (Experiment 2-5) that demonstrates how layered defenses reduce prompt injection success rates. It provides an Agent with sensitive capabilities (internal secret, high-risk file/email tools) and three attack vectors: direct system-prompt leakage, indirect injection via webpage content, and memory injection via poisoned persistent notes. Four progressive defense configurations are implemented: D1 (baseline rules), D2 (prompt hardening against external content), D3 (source tagging with XML wrappers to separate data from instructions), and D4 (runtime execution guards requiring explicit user confirmation for high-risk tools). The suite includes a CLI-driven demo script, a deterministic robustness evaluator for offline testing, and a campaign runner that generates reproducible evidence (JSON traces, workspace inventories, provider receipts) to quantify attack success rates across all 12 attack-defense combinations.
chapter2/prompt-injection · high confidence
New site extras for language switching, auto-translation, and theme styling
This change introduces a suite of client-side scripts and styles in the extras directory to enhance the reading experience. It adds a custom language switcher (lang-switcher.js/css) that populates a header dropdown, handles URL rewriting for translated editions, and localizes sidebar labels. It also introduces an opt-in machine-translation tier (auto-translate.js/css) that uses a third-party library to translate pages for languages without built editions, complete with a notice indicating translation status. Additionally, it provides a flat, 'fenix-inspired' theme override (book-theme.css) that removes Material's card-like decorations, implements a full-width layout with a collapsible sidebar (nav-collapse.js), and fixes math rendering (mathjax.js) and diagram loading (mermaid-init.js). Finally, it includes a search index router (search-index-router.js) to split the large search index per edition for better performance.
extras · high confidence
New staged system-prompt experiment with phase-based role switching
Adds a new Chapter 10 experiment (10-1) that runs a single Coding Agent through three distinct phases—Requirements Clarification, Code Implementation, and Code Review—by swapping system prompts and tool sets at each stage while preserving conversation history. The agent uses specific signal tools (complete\_requirements\_analysis, submit\_for\_review, request\_revision, approve\_code) to transition between phases, with a fallback mechanism that returns review issues to the implementation phase. The demo supports running the full flow or starting from specific stages, includes a simulated user for unattended operation, and provides offline validation tests for tool argument coercion and stage transitions. The implementation defaults to gpt-5.6-luna with OpenRouter fallback and includes credential-free evidence recording for reproducibility.
chapter10/staged-system-prompt · high confidence
New structured indexing experiment with RAPTOR and GraphRAG
Chapter 3 now includes a new experiment (3-7) that compares two advanced indexing approaches for large technical documents: RAPTOR (a hierarchical tree with recursive summarization) and GraphRAG (a knowledge graph with entity extraction and multi-hop traversal). This addition provides a new HTTP API service for building and querying these indexes, along with a campaign script for running comparative retrieval experiments on the Intel SDM manual.
chapter3/structured-index · high confidence
New user-memory-evaluation framework with three-layer test suite and LLM-as-judge
Adds a complete evaluation framework for testing AI agent memory capabilities across three progressive complexity layers: basic recall, contextual disambiguation, and cross-session synthesis. The framework includes 60 test cases with realistic business conversations, an LLM-as-judge evaluator supporting Kimi and OpenAI models (with OpenRouter fallback), offline keyword-recall comparison mode, and programmatic/interactive/batch execution modes for scoring agent responses.
chapter3/user-memory-evaluation · high confidence
Perception Tools MCP server introduced with 18+ tools and zero-API-key defaults
A new perception-tools MCP server for Chapter 4 is added, providing a modular set of tools across five categories: search (web, knowledge base, file download), multimodal understanding (webpage, document, image, video), filesystem operations (read, grep, summarize), public data sources (weather, stocks, crypto, currency, location, POI, Wikipedia, ArXiv, Wayback), and private data integrations (Google Calendar, Notion). The server uses MCP SDK v2 (protocol 2026-07-28), ships with a unified CLI (cli.py) for listing, inspecting, and calling tools, and defaults to free, open APIs (DuckDuckGo, Open-Meteo, CoinGecko, Nominatim, Overpass) so most features work out-of-the-box without API keys. Optional paid/third-party integrations include Xquik for X post search, and the repo includes a Dockerfile, architecture docs, and experiment validation scripts.
chapter4/perception-tools · high confidence
Repository restructured for v2.0 with multi-language support and EPUB builds
The repository has been reorganized to support version 2.0 of the book, which restructures chapters (e.g., merging previous chapters 4 and 9 into a new Chapter 6 on Interaction) and expands to 15 languages. This change introduces a unified EPUB build pipeline (\build\_epub.sh\) and documentation (\EPUB.md\), adds a centralized \.env.example\ for configuring multiple AI provider keys (OpenRouter, Ollama, Moonshot, etc.), and adopts the Apache 2.0 license. It also integrates CodeRabbit for automated PR reviews via \.coderabbit.yaml\ and updates the READMEs to reflect the new content structure and language availability.
(repo-wide) · high confidence
Reproduction of the Stanford Generative Agents society experiment with modern LLMs
This location introduces a complete, self-contained reproduction of the Stanford Generative Agents (Smallville) society simulation, pinned to the upstream commit fe05a71d. The code replaces the obsolete GPT-3 API surface with current OpenAI-compatible chat and embedding endpoints (defaulting to DashScope's qwen3.7-flash and text-embedding-v4) via a runtime adapter that reads credentials exclusively from environment variables. It executes three experimental arms—baseline, custom\_goal, and no\_reflection—each running 17,280 ten-second steps from a shared 25-persona history seed. The implementation includes a compatibility layer for legacy action-arena response parsing, robust retry logic for transient provider errors, and a full analysis and evidence-packaging pipeline that retains credential-free provider receipts, movement logs, memory states, and blind plausibility judgments.
chapter10/generative-agents · high confidence
Spanish edition of the book is now available
The book is now fully translated into Spanish, including all chapters, the afterword, glossary, and reference answers. A build script is provided to generate a PDF of the complete Spanish edition.
book-es · high confidence
Tamil edition of the book is introduced
A new Tamil translation of the book is added, including all chapters (1–10), the afterword, and a dedicated script to build the PDF. This brings the total number of supported language editions to thirteen.
book-ta · high confidence
User Memory System now supports multiple LLM providers with OpenRouter fallback
The User Memory System in chapter3/user-memory has been expanded to support multiple LLM providers, including Kimi/Moonshot (default), SiliconFlow, Doubao, Alibaba Cloud DashScope (Qwen), and OpenRouter. Users can now select their preferred provider via the --provider flag or PROVIDER environment variable, with specific API keys required for each (e.g., MOONSHOT\_API\_KEY, SILICONFLOW\_API\_KEY, DASHSCOPE\_API\_KEY, DOUBAO\_API\_KEY, OPENROUTER\_API\_KEY). A universal OpenRouter fallback is implemented: if the primary provider's API key is missing but OPENROUTER\_API\_KEY is set, the system automatically routes requests through OpenRouter using mapped model IDs (e.g., kimi-k3 maps to moonshotai/kimi-k2.6, GPT models map to openai/gpt-5.6-luna). The system also includes provider-specific default models and base URLs, and supports both command-line and Python API usage across all providers.
chapter3/user-memory · high confidence
Voice Werewolf experiment introduces real-LLM user simulator and live voice integration
The chapter10/voice-werewolf directory now provides a complete, end-to-end playable Werewolf game that supports two distinct user paths: a live human player interacting via real-time microphone capture, OpenAI TTS/ASR, and barge-in detection, and an unattended automated user simulator that drives the game using a real LLM. The simulator strictly enforces an audio-action boundary by converting LLM tool outputs into synthesized audio, transcribing that audio back via ASR, and failing closed if the ASR transcript does not match the intended action. The game logic includes a deterministic Judge that manages state, enforces information isolation (verified by automated post-game audits), and correctly handles edge cases such as simultaneous deaths (reporting UNDECIDED) and preventing a killed Witch from using poison in the same night. The project includes comprehensive validation evidence, independent audio evaluation tools, and tests covering strategy acceptance, privacy isolation, and simulator trace integrity.
chapter10/voice-werewolf · high confidence
Removals
Removal of Kimi-based Web Search Agent
The Kimi Web Search Agent implementation in the \week1/web-search-agent\ directory has been removed. This deletes the core agent logic (\agent.py\), the main entry point (\main.py\), configuration files (\config.py\, \env.example\), and documentation (\README.md\) that previously enabled users to perform web searches using the Kimi (Moonshot AI) API.
week1 · high confidence
Behavioural changes
Added missing experiment evidence and protocols for Chapter 2 runs
This change adds the missing experiment protocol and evidence files for Chapter 2 experiments 2-2 and 2-7, ensuring that the run artifacts for the Qwen3-0.6B model are complete and reproducible. Specifically, it includes the full run data for experiment 2-2 (testing attention on simple and reasoning prompts) across three versions, and the comparison data for experiment 2-7 (testing the impact of a status bar on tool-use refusal).
_chapter2/attention\visualization/runs · high confidence
Autonomous phone registration now uses local WebRTC instead of PSTN
The autonomous phone registration experiment (Experiment 10-3) has replaced its default PSTN (Twilio) transport with a local WebRTC call. This change allows the system to orchestrate a real LLM-driven Computer Agent and Phone Agent using local browser audio tracks and RTP, eliminating the need for a phone number, PSTN provider, or public webhook. The system now negotiates a local WebRTC offer/answer pair to carry agent and participant audio, while retaining optional legacy transports (local microphone and Twilio) for backward compatibility.
chapter10/autonomous-phone-registration · high confidence
Book restructured into 10 chapters with new afterword and PDF build automation
The book content has been reorganized into a 10-chapter structure (chapters 1–10) with a new afterword, replacing the previous versioning scheme. A new build script (build\_pdf.sh) automates PDF generation using Pandoc and the ElegantBook class, outputting a versioned PDF file. The repository now serves as the primary open-source source for the book, with all chapters, figures, and experiment code consolidated in this location.
book · high confidence
Centralized provider registry with unified resolution and new provider support
The \agentbook/providers\ package now provides a single source of truth for LLM provider configuration, replacing scattered, per-experiment setup code with a shared registry. Users can now select from a broader set of supported providers (including Krill AI and Atlas Cloud) via a unified \resolve\_backend\ interface, which automatically handles credential resolution, base URL overrides, and model ID mapping. The system intelligently routes \gpt-5.x\ requests through OpenRouter to bypass direct API restrictions, while also supporting zero-cost local execution via Ollama or free OpenRouter models. A legacy shim ensures backward compatibility for existing chapter experiments during migration.
agentbook/providers · high confidence
Chapter 1 search-codegen companion upgraded to GPT-5.6 Sol with multi-provider Responses API support
The search-codegen experiment in Chapter 1 has been updated to use the GPT-5.6 Sol model and now supports the OpenAI Responses API alongside Alibaba Model Studio (DashScope) as an eligible acceptance backend. The agent implementation has been refactored to use the native \/v1/responses\ protocol, ensuring that hosted \web\_search\ and \code\_interpreter\ tool calls are properly tracked and cited. Configuration now includes specific settings for DashScope, including streaming requirements and model-specific tool shapes, while retaining OpenRouter as a diagnostic-only path. The companion also includes updated validation logic to verify multi-round search and code execution workflows, such as the ASEAN capitals and Bitcoin analysis tasks.
chapter1/search-codegen · high confidence
Chapter 6 documentation restructured with experiment ledger and multilingual READMEs
Chapter 6 has been reorganized to introduce a formal experiment ledger (EXPERIMENT\_LEDGER.md) that tracks the status, canonical run paths, and SHA-256 manifests for experiments 6-1 through 6-3, while also mapping current experiment numbers to their archived identifiers. The chapter now features a comprehensive bilingual README (English and Chinese) that details 14 companion projects—ranging from event-driven agents and asynchronous frameworks to voice interaction, computer use, and robotic manipulation—complete with specific reproduction instructions, external repository anchors, and hardware requirements. This documentation is synchronized across 11 additional languages (Arabic, Spanish, Hungarian, Indonesian, Japanese, Korean, Russian, Tamil, Turkish, Vietnamese, and others) to ensure consistent access to experiment status and setup guides for all users.
chapter6 · high confidence
Chapter 7 introduces a formal experiment ledger and standardized README structure
Chapter 7 now includes an \EXPERIMENT\_LEDGER.md\ that serves as a centralized acceptance record, explicitly mapping each experiment (7-1 through 7-14) to its specific evidence files, audit status, and completion boundaries. This ledger is referenced by the newly structured \README.md\ files across all supported languages (English, Chinese, Arabic, Spanish, Hungarian, Indonesian, Japanese, and Korean), which now feature a standardized 'Companion Projects' table and a 'Project Types' legend to clarify which experiments are standalone, require external cloning, or are still in design.
chapter7 · high confidence
Chapter 9 documentation and experiment catalog refreshed
The Chapter 9 README has been rewritten to provide a structured guide for navigating the self-evolution experiments, introducing a three-tier reading path (Starter, Builder, Maintainer) and a standardized project-type legend (Standalone, Reproduction Guide, Design Doc). The companion project table has been updated to reflect the current status of experiments 9-1 through 9-9, including specific validation evidence for the trajectory verifier, tau2-escalation-experience, and self-modifying-agent. Additionally, the \ai-style-skill\ project has been added to document the open-ended extraction of writing rules from user feedback, and supplementary cases like \self-evolving-tools\ and \ai-style-skill\ have been integrated into the chapter overview.
chapter9 · high confidence
Custom search index routing via template override
A new \overrides/main.html\ template has been added to inject a custom JavaScript file (\extras/search-index-router.js\) into the page configuration block. This script runs before the main bundle to enable dynamic routing of search index requests, supporting the per-book edition search index split.
overrides · high confidence
Introduce agentbook package for shared dependency management
The agentbook package has been added to centralize dependency declarations via the root pyproject.toml, allowing users to install specific chapter requirements (e.g., chapter 1 or chapter 7) using optional dependency groups instead of managing separate requirements.txt files. This package handles plumbing tasks such as provider resolution, environment loading, and trace printing, while keeping teaching code isolated within individual chapter directories.
agentbook · medium confidence
Introduce parallel tool execution and structured streaming for local LLM tool calls
The local LLM serving experiment now executes independent tool calls in parallel using a thread pool, significantly reducing latency when multiple tools are invoked in a single turn. Additionally, streaming responses now expose structured chunks—including distinct types for thinking, tool calls, tool results, and content—allowing users to observe the agent's internal reasoning and tool usage in real time. The demo also includes a system compatibility checker to guide users toward the correct backend (vLLM for Linux/GPU, Ollama for macOS/Windows) and a benchmark script to measure throughput, TTFT, and KV-cache effects.
_chapter2/local\_llm\serving · high confidence
KV Cache experiment now supports Kimi K3 and GPT-5 reasoning models with automatic temperature handling
The KV Cache demonstration in chapter2/kv-cache has been updated to use the current Kimi K3 family (kimi-k2.5, kimi-k2.6, kimi-k2.7, kimi-k3) and GPT-5 as the default models, replacing the deprecated Kimi K2. These models are identified as reasoning models that require temperature=1 and report cached\_tokens, which the experiment uses to measure cache hit rates. The agent code now automatically enforces temperature=1 for these models and increases the max\_tokens limit to 4096 to accommodate hidden reasoning tokens, ensuring tool calls are not truncated. Additionally, the experiment now supports OpenRouter as a fallback provider if the primary Moonshot API key is not set.
chapter2/kv-cache · high confidence
Mem0 v3 agent with Kimi K3 and OpenRouter fallback
The chapter3/mem0 experiment now uses the Mem0 v3 memory framework with Kimi K3 as the primary language model, supporting append-only fact extraction and hybrid retrieval for persistent, cross-session memory. It includes a universal OpenRouter fallback that routes the chat LLM through OpenRouter when the primary Kimi API key is missing, and adds contract tests to verify the v3 memory filter and retrieval behavior.
chapter3/mem0 · high confidence
Replace native language selector with custom dropdown and expose config to JavaScript
The site now uses a custom language switcher dropdown in the header instead of the default Material for MkDocs selector. This change injects a new UI component between the search bar and repository link, and exposes site language configuration, the site root URL, and optional machine-translation settings to the browser via JavaScript variables (window.LANG\_CONFIG, window.SITE\_ROOT, and window.AUTO\_TRANSLATE\_CONFIG) to support the new client-side switching logic.
overrides/partials · high confidence
Retained TalkAct reproduction campaign with Anthropic caller validation
The chapter 10 TalkAct reproduction experiment (10-3) has been completed and its validation artifacts retained. Because the original Gemini caller credentials were invalid, the campaign used an Anthropic caller (\claude-sonnet-4-5-20250929\) instead. The results show that the concurrent duplex configuration reduced median voice latency by approximately 5.40× compared to the single-model strawman control, though it did not improve task success rates. A new \validate\_campaign.py\ script and acceptance report verify the integrity of the 16-episode run, including source pins, token usage, and concurrency events.
chapter10/talkact-reproduction · high confidence
System-Hint agent upgrades to Kimi K3 and adds offline status-bar preview
The default model for the System-Hint agent experiment has been switched from the retired Kimi K2 to Kimi K3, with provider aliases (kimi/moonshot) now resolving to kimi-k3 and reasoning-model constraints (temperature=1, max\_tokens=8192) enforced via \_reasoning\_safe\_temperature(). A new --mode preview CLI option renders the five status-bar techniques (timestamps, tool counter, TODO list, detailed errors, system state) as before/after comparisons without API keys or LLM calls. Trajectory logging now captures last\_llm\_messages (the full context including dynamic system hints) alongside conversation\_history, and all timestamps use real system time rather than simulated delays.
chapter2/system-hint · high confidence
Validation suite for voice-based Werewolf multi-agent campaigns
The validation directory now contains a comprehensive set of acceptance reports, game logs, and audit traces for the Chapter 10 voice-based Werewolf agent system. These files document the results of experiments (primarily experiment 10-6) that test the system's core behaviors, including information isolation between roles (e.g., wolves, seers, villagers), real-time speech-to-text (ASR) and text-to-speech (TTS) integration, and LLM-driven strategy. The reports cover various execution modes, from offline all-AI diagnostics to simulated user runs using providers like OpenRouter and Volcengine ARK, capturing detailed metrics on game cycles, role consistency, and safety gates to verify the system's functionality and compliance.
chapter10/voice-werewolf/validation · high confidence
Web search agent upgraded to Kimi K3 with OpenRouter fallback and increased timeout
The web search agent now uses the Kimi K3 model instead of the discontinued Kimi K2, and includes a universal OpenRouter fallback (defaulting to GPT-5.6) when the primary Moonshot API key is missing. The default search timeout has been increased from 30 to 180 seconds to accommodate Kimi K3's reasoning latency, and the agent now distinguishes between rate-limit errors and timeouts in its error messages.
chapter1/web-search-agent · high confidence
Fixes
Introduce coding-agent tools module with robust file, shell, and notebook operations
The \chapter5/coding-agent/tools\ package now provides a complete set of tool implementations for the coding agent, including Bash, Read, Write, Edit, MultiEdit, Grep, Glob, LS, NotebookEdit, and BashOutput. This release fixes several behavioral issues: the Bash tool now supports sub-second timeouts and treats timeout values of 0 or null as omitting the timeout; the BashOutput tool correctly returns only new output since the last check; the NotebookEdit tool preserves newlines when writing cell source; the Grep tool clamps negative context counts to zero, treats negative head limits as unlimited, and honors head\_limit=0 as zero results; the Read tool treats negative limits as reading to EOF and no longer reports non-empty files as empty when the limit is 0; the Edit and MultiEdit tools reject empty old\_string on existing files and ensure MultiEdit creates are atomic; and the Bash tool now supports shell commands on Windows via PowerShell or cmd.
chapter5/coding-agent/tools · high confidence
Test coverage
Added comprehensive test suite for the web search agent; Added formal validation experiment and stricter judging criteria for async steering; Added regression and basic tests for learning experiment components; Added tests for the tier-3 machine translation feature; Comprehensive test suite for Coding Agent tools; Expanded test coverage for site asset cleanup and chapter-specific modules.
Dependencies
Explicit dependency manifests added for all chapter experiments
The repository now includes dedicated \requirements.txt\ files for every experiment across chapters 1 through 10, as well as \package.json\ and \package-lock.json\ files for the Node.js-based frontend components (such as the attention visualization and conversational UI). This change ensures that each experiment has a pinned, explicit list of its Python and JavaScript dependencies, replacing any previous reliance on transitive or root-level dependencies and guaranteeing reproducible environments for the book's code samples.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 35.
Lenses
- Code Health 53
- Architecture 97
- Maturity 70
- Readiness 17
- Security 51
- Domain Modelling 100
- Accessibility 41
Changes since last survey
- 300 commits — 231 feature/other, 69 fixes
By area
- assets/star-history-dark.png — 40 commits
- (root) — 22 commits
- book-ar/chapter1.ar.md — 14 commits
- book-ar/chapter2.ar.md — 12 commits
- book-ar/chapter10.ar.md — 8 commits
- book-ar/chapter3.ar.md — 8 commits
- book-ar/chapter6.ar.md — 8 commits
- book/images — 8 commits
- book-ar/chapter7.ar.md — 7 commits
- book-ar/chapter9.ar.md — 7 commits
- book-ar/chapter4.ar.md — 6 commits
- book-ar/images — 6 commits
- chapter1/context — 6 commits
- book-ar/chapter8.ar.md — 5 commits
- chapter2/local_llm_serving — 5 commits
- book-ar/chapter5.ar.md — 4 commits
- book/chapter1.md — 4 commits
- book/chapter2.md — 4 commits
- book/chapter3.md — 4 commits
- book/chapter6.md — 4 commits
Notable commits
- fix: Fix stale highlights backup exports across browser tabs (#1083)
- fix: [verified] fix(epub): unblock PT-BR artifact publication (#1053)
- fix: fix(book): add the missing deletion marker in the fig5-4 diff example (#885)
- fix: fix(ch1): guard against empty evaluation windows in RL and LLM learning agents (#900)
- fix: fix(ch1): handle null response in search codegen agent (#786)
- fix: fix(ch1): 修正实验 1-1 的消融设计与度量,并按原理重写正文结论 (#971) (#984)
- fix: fix(ch1): 实验 1-2 默认超时从 30 秒提到 180 秒,并区分速率限制与超时 (#1058)
- fix: fix(ch10): guard against empty source and formula blocks in bilingual consistency auditor (#897)
- fix: fix(ch10): handle non-JSON payloads in Envelope.short log representation (#778)
- fix: fix(ch10): harden voice werewolf acceptance gates (#764)
- fix: fix(ch2): fix math label rendering (#905)
- fix: fix(ch2): 修复 Ollama 工具调用消息历史顺序 (#1035)
- fix: fix(ch2): 修正图 2-10 KV Cache 前缀复用机制图的语义与排版问题 (#1029)
- fix: fix(ch2): 补交实验 2-2/2-8 缺失的协议 JSON 与运行证据 (#1023)
- fix: fix(ch2): 让 gpt-5.x 改走 OpenRouter,而不是关掉推理 (#980)
- fix: fix(ch2): 说明 V4 回传 reasoning_content 的强制条件是携带 tools (#982)
- fix: fix(ch2): 适配新版 ollama 的 client.list() 返回类型 (#1041)
- fix: fix(ch2,ch3): 统一适配新版 ollama 的 client.list() 返回类型 (#1047)
- fix: fix(ch3): guard BM25 calculate_term_score against zero avgdl and calculate_raw_idf (#780)
- fix: fix(ch3): handle empty conversation chunks in Agentic RAG user memory tools (#899)
- …and 280 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
bojieli/ai-agent-book was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 4da3246bec9e9531df6712354a7c79f278925f33 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.