wshobson/agents
64.6
Adequate · 26 September 2026
11.2k
lines of production code
Python
with C#
5
measurements over time
What this system is
This system is a framework for authoring, evaluating, and distributing portable AI coding plugins and skills across multiple harnesses like Cursor, Codex, and Copilot. It provides specialized capabilities for backend development, Kubernetes operations, LLM application design, and file conversion, alongside robust governance tools for policy enforcement and human-gated review workflows. The platform also includes a quality scoring framework to certify plugin standards and ensures cross-harness compatibility through unified adapter generation.
How it got here
2025 — expansion of specialized development skills
4 changes.
This period focused on significantly expanding the project's plugin ecosystem by introducing new specialized skills for backend development, Kubernetes operations, LLM application design, and .NET contributions. These additions provided developers with structured guidance, templates, and best practices for modern architectural patterns, infrastructure management, and AI-driven application development.
2026 — Plugin ecosystem expansion and governance
9 changes.
This period focused on significantly expanding the plugin ecosystem with diverse tools for file conversion, presentation generation, and specialized hardware operations. Concurrently, it established a robust governance framework by introducing evaluation metrics, human-gated approval workflows, and cryptographic audit trails to ensure plugin quality and security. The work also standardized plugin development through multi-harness adapters and unified dependency management.
Features
Add .NET backend development skill with patterns and templates
A new skill named 'dotnet-backend-patterns' has been added to the .NET contribution plugin, providing guidance and code templates for building production-grade C\#/.NET backends. The skill covers Clean Architecture project structures, dependency injection (including .NET 8 keyed services), async/await best practices, and configuration via IOptions. It includes reference documentation and implementation templates for both Dapper (high-performance SQL) and Entity Framework Core, demonstrating patterns for repository and service layers, caching, validation, and error handling.
plugins/dotnet-contribution/skills · high confidence
Introduce PluginEval: a quality scoring framework for plugins and skills
The \plugins/plugin-eval\ directory now contains a complete evaluation framework that scores Claude Code plugins and skills across up to three layers: a deterministic static lint, an experimental LLM judge, and an experimental Monte Carlo simulation. Users can run the Typer CLI (\plugin-eval score\, \certify\, \compare\, \init\) to generate composite scores, assign quality badges (Platinum, Gold, Silver, Bronze), and identify anti-patterns. The framework includes a batch evaluation script (\eval\_all.py\) for CI reporting, a corpus management system with Elo ranking, and a set of Claude Code agents (\eval-judge\, \eval-orchestrator\) to assist with the evaluation process.
plugins/plugin-eval · high confidence
Introduce file-conversion plugin with ChangeThisFile integration
Adds a new file-conversion plugin that enables users to convert files between 999 supported formats (such as PDF to Word, HEIC to JPG, and MP4 to MP3) using the free ChangeThisFile API. The plugin provides MCP-aware tools for direct integration and includes a bundled shell script fallback for environments without MCP support. It enforces a 25 MB input limit, applies strict validation to prevent path traversal and argument injection in the conversion script, and handles file downloads securely.
plugins/file-conversion · high confidence
Introduce multi-harness plugin adapters for Codex, Cursor, Copilot, OpenCode, Antigravity, and Pi
The tools/adapters module now provides a unified framework to generate native plugin artifacts for six AI coding assistants from a single set of canonical Markdown sources. This includes new adapter implementations for Google Antigravity CLI (emitting agy plugins with TOML commands and subagents), GitHub Copilot (emitting agent profiles and runnable skills), and Pi (emitting prompt templates and skills), alongside updated adapters for Codex, Cursor, and OpenCode that handle harness-specific constraints like tool-name casing, permission blocks, and body size caps. A central capabilities matrix defines the supported features for each harness, ensuring portable content authoring across the ecosystem.
tools · high confidence
Introduce protect-mcp plugin for policy-gated tool execution and cryptographic audit trails
The new protect-mcp plugin enforces Cedar authorization policies on every Claude Code tool call and produces Ed25519-signed receipts for offline verification. A PreToolUse hook evaluates each call against a user-defined policy file (defaulting to allow if missing) and blocks execution with exit code 2 on deny; a PostToolUse hook appends a signed receipt to ./receipts/receipts.jsonl. The plugin ships with slash commands (/verify-receipt, /audit-chain) and helper agents (policy-enforcer, receipt-verifier) to author policies and verify integrity, along with test fixtures that validate the full evaluate-sign-verify loop.
plugins/protect-mcp · high confidence
Introduce review-agent-governance plugin for human-gated AI review actions
Adds the \review-agent-governance\ plugin (v0.1.4) to enforce human approval before AI agents can post PR reviews, comments, merges, or edit CI configuration. The plugin uses \protect-mcp\ (pinned to 0.7.4) and Cedar policies to block review-surface actions by default, requiring a \.review-approved\ flag file or the \/approve-review\ slash command to open a temporary approval window. It includes \PreToolUse\ and \PostToolUse\ hooks to evaluate policies and sign Ed25519 receipts for every allowed tool call, along with a \/list-pending\ command to view blocked actions. The plugin is available for both Claude Code and Codex environments.
plugins/review-agent-governance · high confidence
New Cursor-specific agent authoring and project conventions
Added three new rule files to the Cursor configuration that define project conventions, Python tooling standards (uv, ruff, ty), and guidelines for authoring portable agent skills and commands. These rules ensure that plugin content is compatible across multiple AI harnesses, including Cursor, by specifying frontmatter structures, content length limits, and tool-agnostic language requirements.
.cursor · high confidence
New DGX Spark Ops plugin for GB10 environment setup and preflight
Adds the \dgx-spark-ops\ plugin, providing specialized skills and an agent for NVIDIA DGX Spark (GB10 Grace Blackwell) systems. This includes an environment setup skill for managing the aarch64/CUDA-13 stack and container workflows, a preflight command that runs automated checks (G1–G10) to verify hardware identity, memory headroom, and thermal status, and detailed reference materials for unified memory accounting and known failure modes.
plugins/dgx-spark-ops · high confidence
New Kubernetes operations skills for GitOps, Helm, and manifest generation
Added three new documentation-based skills to the Kubernetes operations plugin: 'gitops-workflow' provides guidance for implementing GitOps with ArgoCD and Flux CD, including installation, sync policies, and App-of-Apps patterns; 'helm-chart-scaffolding' offers templates and references for creating production-ready Helm charts with security best practices; and 'k8s-manifest-generator' supplies templates and detailed references for generating secure Kubernetes Deployments, Services, and ConfigMaps. These skills provide structured, step-by-step guidance for common Kubernetes deployment and management tasks.
plugins/kubernetes-operations/skills · high confidence
New LLM application development skills for embedding, hybrid search, and LangGraph architecture
Added new skill documentation and reference templates for building modern LLM applications. The \embedding-strategies\ skill provides guidance on selecting and optimizing embedding models (including Voyage AI and OpenAI) and chunking strategies for RAG. The \hybrid-search-implementation\ skill introduces patterns for combining vector and keyword search using methods like Reciprocal Rank Fusion (RRF) and linear combination, with specific examples for PostgreSQL. The \langchain-architecture\ skill details the use of LangChain 1.x and LangGraph for building agents, managing state, and implementing memory systems. Additionally, the \llm-evaluation\ skill offers strategies for automated metrics and LLM-as-judge evaluation, while \prompt-engineering-patterns\ includes templates for few-shot learning, chain-of-thought, and structured outputs.
plugins/llm-application-dev/skills · high confidence
New PPTX Deck Creation plugin for generating editable PowerPoint decks
The \plugins/pptx-deck-creation\ plugin is now available, enabling the creation of production-ready, editable PowerPoint decks through a spec-first workflow. It provides skills for preparing business narratives, authoring coordinate-explicit slide specifications, analyzing reference decks in read-only mode, and applying quality gates for geometry and accessibility. The plugin uses a task-local \python-pptx\ builder to generate native editable objects (text, shapes, tables) based on inch-based JSON contracts, ensuring source decks are never mutated and supporting safe OOXML package inspection.
plugins/pptx-deck-creation · high confidence
New backend development skills for API design, architecture, and event-driven patterns
The backend-development plugin now includes five new specialized skills: \api-design-principles\ provides guidance and a production-ready FastAPI template for building REST and GraphQL APIs; \architecture-patterns\ covers Clean, Hexagonal, and Domain-Driven Design with implementation examples; \cqrs-implementation\ details Command Query Responsibility Segregation patterns; \event-store-design\ guides the setup of event stores for event-sourced systems; and \saga-orchestration\ (referenced in architecture patterns) supports distributed transaction management. These skills equip developers with structured approaches to designing scalable, maintainable backend systems.
plugins/backend-development/skills · high confidence
Dependencies
Migrate plugin-eval to pyproject.toml and update yt-design-extractor dependencies
The plugin-eval tool has been scaffolded with a pyproject.toml manifest, replacing previous dependency management methods and specifying Python 3.12+ requirements along with core dependencies like pydantic, typer, rich, and pyyaml, as well as dev tools including pytest, ruff, and ty. The yt-design-extractor tool has also been migrated to pyproject.toml, updating its dependency list to include yt-dlp (version 2026.7.4 or higher), youtube-transcript-api, Pillow, pytesseract, and colorthief, while maintaining an optional easyocr extra. Additionally, a new requirements.txt file was added for the pptx-deck-creation plugin's reference deck analysis skill, pinning defusedxml to versions \>=0.7 and \<1.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 51 → 65 (+13.2)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 92 → 90 (-1.7)
- Architecture 92 → 84 (-7.5)
- Maturity 59 → 81 (+21.5)
- Readiness 30 → 45 (+14.9)
- Security 72 → 80 (+8.3)
Resolved (20)
- Coverage not measured — test suite did not build
- Dimension evaluation failed
- Duplicated block (6 lines × 2) (tools/tests/test_adapters.py)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- LLM evaluation failed
- Low IaC: KSV-0004 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/deployment-template.yaml)
- Low IaC: KSV-0020 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/deployment-template.yaml)
- Low IaC: KSV-0021 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/deployment-template.yaml)
- Low IaC: KSV-0021 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/deployment-template.yaml)
- Low IaC: KSV-0106 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/deployment-template.yaml)
- Low: security finding (details withheld)
- Medium IaC: KSV-01010 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml)
- Medium IaC: KSV-01010 (plugins/kubernetes-operations/skills/k8s-manifest-generator/assets/configmap-template.yaml)
- No exposed public API
- No tests found
- Off-boarding risk: anonymized user
- Test reliability not included
New (121)
- Dependency hygiene PARTLY measured — Python dependencies read, no exact pin to grade for currency
- Documentation: contradicts the code (evals/README.md)
- Duplicated block (10 lines × 2) (tools/adapters/copilot.py)
- Duplicated block (10–12 lines × 3) (tools/validate_generated.py)
- Duplicated block (10–14 lines × 3) (tools/adapters/codex.py)
- Duplicated block (11 lines × 2) (tools/doc_gardener.py)
- Duplicated block (11 lines × 2) (tools/validate_generated.py)
- Duplicated block (11 lines × 3) (tools/install_antigravity.py)
- Duplicated block (11 lines × 3) (tools/install_antigravity.py)
- Duplicated block (12 lines × 2) (tools/install_copilot.py)
- Duplicated block (12–15 lines × 2) (tools/validate_generated.py)
- Duplicated block (13 lines × 3) (tools/validate_generated.py)
- Duplicated block (14 lines × 2) (tools/validate_generated.py)
- Duplicated block (18 lines × 4) (tools/install_antigravity.py)
- Duplicated block (28 lines × 2) (tools/yt-design-extractor/yt-design-extractor.py)
- Duplicated block (5 lines × 2) (tools/adapters/copilot.py)
- Duplicated block (5 lines × 2) (tools/doc_gardener.py)
- Duplicated block (5 lines × 2) (tools/yt-design-extractor/yt-design-extractor.py)
- Duplicated block (5 lines × 3) (plugins/pptx-deck-creation/skills/pptx-reference-deck-analysis/scripts/inspect.py)
- Duplicated block (5 lines × 3) (tools/adapters/antigravity.py)
- …and 101 more
Changes since last survey
- 61 commits — 41 feature/other, 20 fixes
By area
- plugins/plugin-eval — 15 commits
- (root) — 13 commits
- .github/workflows — 7 commits
- tools/yt-design-extractor — 4 commits
- docs/agents.md — 2 commits
- plugins/hermes-tweet — 2 commits
- plugins/protect-mcp — 2 commits
- docs/usage.md — 1 commit
- plugins/avoid-ai-writing — 1 commit
- plugins/code-refactoring — 1 commit
- plugins/context-management — 1 commit
- plugins/documentation-standards — 1 commit
- plugins/llm-application-dev — 1 commit
- plugins/meigen-ai-design — 1 commit
- plugins/pptx-deck-creation — 1 commit
- plugins/python-development — 1 commit
- plugins/review-agent-governance — 1 commit
- plugins/runapi-mcp — 1 commit
- plugins/superself — 1 commit
- tools/adapters — 1 commit
Notable commits
- fix: fix(adapters): quote YAML scalars in OpenCode and Copilot frontmatter (#700)
- fix: fix(audit-trails): make the protect-mcp audit trail work with Claude Code hooks and protect-mcp 0.7.4 (#732)
- fix: fix(ci): accept ref/sha pins on git-subdir marketplace entries (#679)
- fix: fix(commands): give every command an imperative description (#726)
- fix: fix(copilot): do not clear caches when install fails (#663) (#666)
- fix: fix(garden): ignore model tier and named variants in agent divergence (#724)
- fix: fix(hermes-tweet): sync catalog with 0.1.12 (#672)
- fix: fix(hermes-tweet): sync catalog with 0.1.13 (#686)
- fix: fix(hooks): read the Claude Code hook payload from stdin in protect-mcp and review-agent-governance (#706)
- fix: fix(plugin-eval): count each fenced code block once, not once per fence (#710)
- fix: fix(plugin-eval): populate model_usage from judge and Monte Carlo layers (#660) (#668)
- fix: fix(plugin-eval): resolve nested skill references (#662)
- fix: fix(plugin-eval): stop counting errored runs as activations (#652)
- fix: fix(pptx-deck-creation): drop the agents entry that stops the plugin from loading (#736)
- fix: fix(review-agent-governance): make the default Cedar policy deny under protect-mcp 0.7 (#728)
- fix: fix(test): keep real-CLI smoke tests out of the ordinary test target (#664) (#667)
- fix: fix: describe six Claude plugin commands (#719)
- fix: fix: issue triage — grounded-vault skill, $ARGUMENTS framing, agent copy reconciliation (#694)
- fix: fix: point python-design-patterns related-skills link at the existing python-project-structure skill (#636)
- fix: fix: stop optimize() early when no variation improves the prompt (#637)
- …and 41 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
wshobson/agents was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 26 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 8edbd648a62a04899307ae181367fb2e6dd0cd47 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-d0929f7ac71f.