addyosmani/agent-skills
54.9
Adequate · 25 September 2026
1.6k
lines of production code
JavaScript
primary language
4
measurements over time
What this system is
This system is an evaluation and validation framework for an AI agent's skill-based capabilities, designed to ensure consistency and correctness across its operational environment. It provides a suite of validation scripts and libraries to enforce structural integrity in skills, commands, and artifacts, while maintaining a comprehensive set of behavioral fixtures for testing agent performance. Additionally, it includes specific agent skills, such as structured ideation, and infrastructure hooks to manage session context and caching.
Features
Expanded evaluation fixtures for behavioral and skill-based testing
The \evals/fixtures\ directory has been significantly expanded with new scenario files and code samples to support a broader range of evaluation capabilities. This update introduces fixtures for browser testing with devtools (including a signup reproduction app), CI/CD automation (slug generation), code review (user search diff), code simplification (config parser), context engineering (session audit), debugging (pagination and time-pressure scenarios), deprecation/migration (API inventory), documentation (ADR context), doubt-driven development (migration plan), frontend UI engineering (React button component), git workflow (working tree patch), incremental implementation (report filtering and export drafts), observability (payment retry), performance optimization (product rendering), planning (notifications spec), security (webhook preview), shipping/launch (checkout status), source-driven development (Express session task), spec-driven development (billing and portal briefs), test-driven development (split payment and ledger utilities), and agent skill usage (login regression). These additions provide concrete, executable, or descriptive contexts for evaluating agent behaviors across diverse software engineering tasks.
evals/fixtures · high confidence
Introduce Idea Refine skill for structured ideation
Added the \idea-refine\ skill, which guides users through a three-phase process (divergent expansion, evaluation/convergence, and sharpening) to turn vague concepts into actionable, stress-tested plans. The skill produces a structured markdown one-pager (Problem Statement, Recommended Direction, Key Assumptions, MVP Scope, and Not Doing list) and includes supporting reference materials for ideation frameworks and evaluation criteria. It is triggered by phrases like 'refine this idea' or 'stress-test my plan' and can be initialized via a provided shell script.
skills/idea-refine · high confidence
New CI validation scripts for skills, commands, and artifact paths
This change introduces a suite of new Node.js scripts in the \scripts/\ directory to enforce consistency and correctness across the project's skills, slash-commands, and evaluation artifacts. \validate-skills.js\ and \validate-reference-links.js\ ensure skills follow defined anatomy and that links to shared checklists resolve correctly, preventing broken references. \validate-commands.js\ guarantees that command definitions remain in sync across Claude, Gemini, and Antigravity tool directories. \validate-artifact-paths.js\ protects the spec-to-build pipeline by ensuring all commands and skills reference a canonical set of artifact paths. Additionally, \run-evals.js\ and its test harness \run-evals-test.js\ provide a framework for running and validating skill evaluations, including tiered checks for trigger ranking and behavioral correctness.
scripts · high confidence
New session-start hook and opt-in citation cache for source-driven development
The hooks directory now includes a session-start hook that emits the standard SessionStart envelope with agent-skills context, replacing the previous auto-injection of the meta-skill. Additionally, an opt-in citation cache (sdd-cache) is introduced for the source-driven-development skill: it intercepts WebFetch calls to cache and revalidate documentation content via HTTP 304 responses, reducing redundant fetches while preserving freshness guarantees. A new simplify-ignore hook is also added to protect specific code blocks from being simplified by the model, using placeholders that are expanded back to original content on session stop.
hooks · high confidence
New skill validation library and test suite
Introduced \scripts/lib/skill-lint.js\ as the shared source of truth for SKILL.md validation rules, replacing the previous monolithic CLI script. This library enforces required sections, validates YAML frontmatter (including rejecting invalid YAML that breaks host parsers), strips fenced code blocks to prevent false positives in section checks, and ensures exemption logic relies on an explicit allowlist rather than Object.prototype keys. A comprehensive test suite (\skill-lint-test.js\) was added to verify these behaviors, including edge cases for fence parsing and trigger clause validation.
scripts/lib · high confidence
Behavioural changes
Skills directory now uses a symlink for proper resolution
The .opencode/skills path has been converted from a regular directory to a symbolic link pointing to ../skills/. This change ensures that the opencode skills are located and loaded correctly by resolving the symlink to the actual skills directory.
.opencode · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.
Score
- CAI 41 → 55 (+14.4)
- Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.
Lenses
- Code Health 35 → 48 (+13.5)
- Architecture 69 → 91 (+22.4)
- Maturity 74 → 80 (+5.6)
- Readiness 26 → 41 (+14.8)
- Security 78 → 100 (+21.5)
Resolved (17)
- Dimension evaluation failed
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- Hotspot: evals/fixtures/code-simplification/config-parser.js (evals/fixtures/code-simplification/config-parser.js)
- Hotspot: scripts/lib/skill-lint.js (scripts/lib/skill-lint.js)
- Hotspot: scripts/run-evals.js (scripts/run-evals.js)
- Hotspot: scripts/validate-commands.js (scripts/validate-commands.js)
- LLM evaluation failed
- No automated tests
- No exposed public API
- No tests found
- Scanner failed to run — not a clean result
- Test reliability not included
New (28)
- Documentation: no installation or build instructions (README.md)
- Documentation: no usage examples (README.md)
- FunctionTooLong: run-evals.runDeterministic (scripts/run-evals.js)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- High: security finding (details withheld)
- No dependency advisory monitoring
- Scanner failed to run — not a clean result
- Scanner failed to run — not a clean result
- Workflow token permissions not restricted
- run-evals.parseGrading (cognitive 26) (scripts/run-evals.js)
- run-evals.parseGrading (cyclomatic 26) (scripts/run-evals.js)
- run-evals.runBehavioral (cognitive 29) (scripts/run-evals.js)
- run-evals.runBehavioral (cyclomatic 20) (scripts/run-evals.js)
- …and 8 more
Changes since last survey
- 161 commits — 115 feature/other, 46 fixes
By area
- (repo) — 67 commits
- (root) — 14 commits
- scripts/run-evals-test.js — 8 commits
- .github/workflows — 7 commits
- scripts/lib — 5 commits
- skills/security-and-hardening — 5 commits
- .claude/commands — 4 commits
- docs/advanced-per-agent-configuration.md — 4 commits
- evals/cases — 4 commits
- scripts/validate-reference-links-test.js — 4 commits
- docs/getting-started.md — 3 commits
- skills/context-engineering — 3 commits
- skills/interview-me — 3 commits
- skills/shipping-and-launch — 3 commits
- docs/adoption-guide.md — 2 commits
- hooks/SIMPLIFY-IGNORE.md — 2 commits
- references/performance-checklist.md — 2 commits
- skills/frontend-ui-engineering — 2 commits
- skills/observability-and-instrumentation — 2 commits
- .gemini/commands — 1 commit
Notable commits
- fix: Merge #469: fix references/ links and add CI validator (#468)
- fix: Merge #501: add hook regression coverage and a bash 5.2 simplify-ignore fix
- fix: Merge #510: fix and CI-wire the SessionStart envelope test
- fix: Merge #570: fix stale skill count in README heading
- fix: Merge pull request #434 from ayobamiseun/fix/rank1-description-vocab
- fix: Merge pull request #505 from abhisheksharma2411/fix/skill-lint-frontmatter-and-exempt-lookup
- fix: Merge pull request #531 from ayobamiseun/fix/518-plan-clobber-guard
- fix: Potential fix for pull request finding
- fix: fix(agents): align code-reviewer severity labels with the code-review skill
- fix: fix(commands): mirror incomplete-plan guard across tools
- fix: fix(context): address nucliweb review — merge duplicate red flag, wire Level 5 to budget section
- fix: fix(context): correct lost-in-the-middle mechanism per federicobartoli review
- fix: fix(evals): add missing owner to negative trigger tests
- fix: fix(evals): bind grader results to declared expectations by id
- fix: fix(evals): canonicalize expectation text and recompute pass_rate
- fix: fix(evals): clear result slot up front instead of on rejection
- fix: fix(evals): record executor model and timestamp in grading.json
- fix: fix(evals): reject incomplete grader results
- fix: fix(evals): reject null grader expectations without crashing
- fix: fix(evals): remove stale grading.json when grading is rejected
- …and 141 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
addyosmani/agent-skills was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 25 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit bcab6a1b8503100e8618c3b4e32cc78de43de769 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-dd72cc24c749.