ccheney/robust-skills
48.0
Weak · 21 September 2026
2.5k
lines of production code
Python
with JavaScript
4
measurements over time
What this system is
This system is a framework for defining, validating, and evaluating AI assistant capabilities, referred to as 'skills.' It provides a structured environment for creating specific skills, such as one for Microsoft Teams Adaptive Cards, and includes a comprehensive local evaluation suite to test model behavior against deterministic tasks and routing requests. The system enforces data integrity and test coverage through validation scripts and manages its own runtime dependencies for these evaluations.
Features
Expanded local skill evaluation suite with 112 tasks and 226 routing requests
The \evals\ directory now contains a comprehensive, locally runnable evaluation suite covering all 11 skills, including 112 automated artifact tasks and 226 routing requests. This expansion introduces deterministic artifact graders and bounded Gemini skill evaluations (defaulting to Gemini 3.8 Flash) to detect missing skill guidance, routing mistakes, and regressions. The suite includes detailed coverage documentation, validation evidence, and specific assertion files for skills like modern-javascript, clean-ddd-hexagonal, and slack-mrkdwn, allowing developers to verify model behavior against positive controls and semantic counterexamples without relying on CI or external model providers.
evals · high confidence
New Microsoft Teams Adaptive Cards skill for building and validating card payloads
A new skill has been added to help build, validate, and repair Adaptive Card layouts, actions, forms, and delivery wrappers specifically for Microsoft Teams. This skill provides structured guidance for selecting the correct transport (Bot Framework, Workflows webhook, Microsoft Graph, or Message Extensions) and matching card elements to Teams client capabilities, including a version policy that defaults to 1.2 for broad compatibility and warns against unsupported versions like 1.6. It includes detailed references on action compatibility across surfaces, responsive design principles for narrow clients, and specific rules for Graph attachment formatting. To ensure payload correctness, the skill ships with a local Node.js validation script (\check-teams-card.mjs\) that checks JSON structure and transport-specific constraints for bots, webhooks, and Graph messages.
skills/teams-adaptive-cards · high confidence
New skill validation script for metadata and evaluation coverage
A new \scripts/validate\_skills.py\ script has been added to enforce consistency across skills. It validates that each skill's metadata (name, description) is well-formed, ensures all linked resources are reachable and contained within the skill's directory, and checks that evaluation coverage files (\evals/routing.json\ and \evals/workflows.json\) include required test cases for every skill. This helps maintain data integrity and ensures skills are properly tested before deployment.
scripts · high confidence
Dependencies
Add evaluation runtime dependencies
The \evals/runtime\ directory now includes a Node.js environment with dependencies for database access (drizzle-orm 0.45.2), diagram rendering (mermaid 11.16.1), browser automation (playwright 1.56.1), and TypeScript (5.9.3), alongside a Python requirement for PyYAML 6.0.2 in the \evals\ directory.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Score
- CAI 43 → 48 (+5.3)
- Rubric changed (rubric-2026.08.19 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 40 (new)
- Architecture 97 (new)
- Maturity 61 → 70 (+9.1)
- Readiness 15 → 33 (+17.4)
- Security 100 → 100 (+0.0)
Resolved (8)
- Dependency hygiene not measured — no supported dependency manifest was read
- LLM evaluation failed
- No automated tests
- No exposed public API
- No tests found
- Test reliability not included
- no production source files with tracked history to analyse
- no production source files with tracked history to analyse
New (45)
- (anonymous) (cognitive 20) (evals/runtime/run.mjs)
- (anonymous) (cognitive 35) (evals/runtime/run.mjs)
- (anonymous) (cyclomatic 22) (evals/runtime/run.mjs)
- (anonymous) (cyclomatic 34) (evals/runtime/run.mjs)
- Catalog.invoke (cognitive 17) (evals/catalog.py)
- Gemini.chat (cognitive 61) (evals/providers.py)
- Gemini.chat (cyclomatic 23) (evals/providers.py)
- Hotspot: evals/graders.py (evals/graders.py)
- Hotspot: evals/providers.py (evals/providers.py)
- Hotspot: evals/runner.py (evals/runner.py)
- Hotspot: evals/validation.py (evals/validation.py)
- Hotspot: scripts/validate_skills.py (scripts/validate_skills.py)
- No ADRs found
- Outdated: pyyaml
- check-teams-card.collectCards (cognitive 18) (skills/teams-adaptive-cards/scripts/check-teams-card.mjs)
- check-teams-card.validateCard (cognitive 22) (skills/teams-adaptive-cards/scripts/check-teams-card.mjs)
- check-teams-card.validateCard (cyclomatic 16) (skills/teams-adaptive-cards/scripts/check-teams-card.mjs)
- check-teams-card.validateElement (cognitive 66) (skills/teams-adaptive-cards/scripts/check-teams-card.mjs)
- check-teams-card.validateElement (cyclomatic 54) (skills/teams-adaptive-cards/scripts/check-teams-card.mjs)
- checks.evaluate (cognitive 52) (evals/checks.py)
- …and 25 more
Changes since last survey
- 25 commits — 23 feature/other, 2 fixes
By area
- (root) — 7 commits
- (repo) — 4 commits
- skills/slack-block-kit — 4 commits
- evals/README.md — 2 commits
- skills/modern-javascript — 2 commits
- skills/slack-mrkdwn — 2 commits
- evals/fixtures — 1 commit
- evals/reference_outputs — 1 commit
- skills/bazel — 1 commit
- skills/clean-ddd-hexagonal — 1 commit
Notable commits
- fix: fix: complete local eval trials and select free Gemini models per run
- fix: fix: run skill evals locally with verified Google credentials and bounded retries
- change: Clarify Slack header rollout and unfurl safety
- change: Document current Slack surfaces and Work Objects
- change: Merge pull request #6 from ccheney/update/slack-mrkdwn-docs
- change: Merge pull request #7 from ccheney/update/slack-block-kit-docs
- change: Merge pull request #8 from ccheney/update/modern-javascript-2026
- change: Merge pull request #9 from ccheney/feature/skill-evals
- change: Refresh Block Kit component references
- change: Resolve Slack documentation precision gaps
- change: Update Slack Block Kit skill guidance
- change: Update Slack mrkdwn guidance
- change: chore: default local skill evals to Gemini 3.8 Flash
- change: ci: validate skill evals and add opt-in free-tier benchmarks
- change: docs: add skills.sh install badge
- change: docs: clarify the maturity of skill evaluations
- change: docs: record the skill authoring audit and v4 migration notes
- change: docs: refresh modern JavaScript through ES2026
- change: docs: tighten async and Temporal guidance
- change: feat: add Bazel monorepo skill
- …and 5 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
ccheney/robust-skills was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 21 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 0eea7b060d12e3520556a73b15fbfa79480987d1 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-fa71c66cabd8.