Skip to content
CAI
Software that uses CAICheck a score

SebastienDegodez/copilot-instructions

53.8

Adequate · 21 September 2026

2.2k

lines of production code

TypeScript

with Python, C#

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is an automated evaluation and documentation infrastructure for AI-assisted software development. It provides specialized plugins, skills, and instructions for .NET, C\#, and DDD, alongside tools for behavior-first development and minimal context usage. The system continuously assesses the quality of these assets through an automated evaluation pipeline that generates detailed performance dashboards and reports.

Features

Add automated evaluation benchmarks dashboard and detail views

The website now includes a new automated evaluation pipeline with a benchmarks dashboard. A \BenchmarkPanel\ component displays score, pass rate, and trend for skills and instructions. The \/benchmarks\ page lists all evaluated assets (skills, plugins, and instructions) with their latest scores, pass rates, and trends. Each asset's detail page (skills, instructions) now includes the \BenchmarkPanel\ to show specific evaluation results. The \summary.json\ data file and content configuration for skills, instructions, agents, and plugins are also added to support this feature.

website/src · high confidence

Added AI agent and instruction templates for architecture, C\#, and DDD

Introduced new configuration files to guide AI agents and code generation. In the \agents\ directory, added \architect.agent.md\ and \slides-wizard.agent.md\ to define specialized AI roles for software architecture and presentation design. In the \instructions\ directory, added multiple \.instructions.md\ files providing templates and rules for Clean Architecture, C\# coding style, CQRS, Domain-Driven Design, modular monoliths, and Marp presentations. These files establish structural and behavioral guidelines for developers and AI assistants.

agents, instructions, plugins · high confidence

Initial website configuration and structure

The website project is initialized with a standard Astro minimal template, including a strict TypeScript configuration and a Tailwind CSS setup that defines a custom primary color palette and typography plugins. The build configuration is set to use a specific base path for production deployments, ensuring correct asset routing.

website · high confidence

Introduce Superpowers Whetstone plugin for behavior-first development

The new Superpowers Whetstone plugin provides a suite of skills for behavior-first development, including outside-in TDD, mutation testing, and red-synthesize-green cycles. It introduces four new skills: \gherkin-gate\ for capturing observable behavior in business language, \red-synthesize-green\ for enforcing a two-step AI TDD cycle, \outside-in-tdd\ for driving implementation from approved scenarios, and \mutation-testing\ to verify test suite effectiveness. The plugin also includes reference materials for CQRS patterns, test examples, and a testing strategy for DDD.

plugins/superpowers-whetstone · high confidence

Introduce automated evaluation pipeline for Copilot skills, plugins, and instructions

The \tools/evaluator\ tool now provides a complete automated evaluation pipeline. It includes a CLI with \discover\, \evaluate\, and \report\ commands to scan for changed assets, run LLM-based evaluations against defined scenarios, and generate markdown or JSON reports. The implementation adds adapters for file reading, Git operations, and LLM interactions (OpenAI), along with core logic for scenario execution, judging, and benchmark summary management. This enables continuous, automated assessment of AI-generated code responses.

tools/evaluator · high confidence

Introduce minimal-context-tools plugin with pre-tool-use hook

Adds the minimal-context-tools plugin, which provides a curated collection of CLI tools (such as ast-grep, bat, fd, fzf, jq, and tokei) optimized for minimal context usage and one-shot patterns with LLMs. The plugin includes a \preToolUse\ hook (\snip.json\) and a shell script (\snip-rewrite.sh\) that transparently rewrites supported commands through the \snip\ tool to save context. It also introduces multiple new skills for code analysis, file finding, and data querying.

plugins/minimal-context-tools · high confidence

New C\# Clean Architecture development plugin with DDD, CQRS, and testing guidance

Users can now install the \csharp-clean-architecture-development\ plugin to receive comprehensive guidance on implementing Clean Architecture, Domain-Driven Design, and CQRS in .NET projects. The plugin provides a \clean-architecture-dotnet\ skill that covers layer responsibilities, interface-based dependency injection, and architectural rules. It also clarifies that testing guidance previously in \application-layer-testing\ is now part of the \outside-in-tdd\ skill in the \superpowers-whetstone\ plugin, which users should also install for testing best practices.

plugins/csharp-clean-architecture-development · high confidence

New skills and evaluation dashboards for clean-architecture-dotnet, creating-dotnet-mcp-servers, and mutation-testing

Added baseline skill documentation and evaluation infrastructure for three new areas: the clean-architecture-dotnet skill (including a changelog and results dashboard), the creating-dotnet-mcp-servers skill (including a dashboard), and the mutation-testing skill (including a changelog). These files provide the reference documentation for each skill and the automated evaluation results that track their quality and improvement over time.

.autoresearch · high confidence

New skills for .NET MCP servers, Microcks mocking, and prompt migration

Added new skill documentation for building .NET Model Context Protocol (MCP) servers, generating Microcks OpenAPI mock samples, and migrating prompts to the SKILL.md format. The .NET MCP skill provides guidance on AOT compatibility, transport configuration, and error handling. The Microcks skill offers examples for JSON\_BODY, JavaScript, and Groovy dispatchers. The migration skill outlines rules for frontmatter and scoping.

skills · high confidence

Test coverage

Added unit tests for the evaluator tool

Added unit tests for the evaluator tool, covering the benchmark summary merge logic, the discovery command and discoverer logic, the judge scenario evaluation, the report command, the evaluation runner, and the Git client. These tests validate the core evaluation pipeline components, including how changed assets are discovered, how results are merged into a benchmark summary, and how scenarios are judged and reported.

tools · high confidence

Dependencies

Added automated evaluation pipeline and website dependencies

Introduced a new \@copilot-instructions/evaluator\ tool for automated evaluation, adding dependencies on commander, openai, pino, pino-pretty, yaml, and zod, along with development tools like vitest and eslint. Additionally, the website project was configured with Astro, Tailwind CSS, and related dependencies.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 52 → 54 (+1.4)
  • Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 92 → 84 (-7.6)
  • Architecture 100 → 93 (-7.1)
  • Maturity 74 → 72 (-2.8)
  • Readiness 52 → 50 (-2.2)
  • Security 55 → 72 (+16.8)
  • Accessibility 41 → 43 (+2.0)

Resolved (23)

  • Coverage not measured — analyzer environment
  • Critical CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • Dependency hygiene not measured — dependency manifest found but not parsed for hygiene
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Low vulnerability: [GHSA redacted] (tools/evaluator/package-lock.json)
  • Medium vulnerability: [GHSA redacted] (tools/evaluator/package-lock.json)
  • No exposed public API
  • Test reliability not included
  • Tests co-located / outside the solution
  • …and 3 more

New (40)

  • Critical CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • Documentation: no installation or build instructions (README.md)
  • Documentation: no usage examples (README.md)
  • FunctionTooLong: cli.buildCLI (tools/evaluator/src/cli.ts)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Job token omits the contents scope its checkout needs
  • Low vulnerability: [GHSA redacted] (tools/evaluator/package-lock.json)
  • Medium CVE: [GHSA redacted] (tools/evaluator/package-lock.json)
  • …and 20 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

SebastienDegodez/copilot-instructions was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 21 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 0f0dccf10029993f877672336c23ae73caa5548e — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-b84573e22831.