Skip to content
CAI
Software that uses CAICheck a score

DietrichGebert/ponytail

54.0

Adequate · 25 September 2026

4k

lines of production code

JavaScript

with Python

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

Ponytail is a configuration system that injects a 'lazy senior dev' coding ruleset into various AI coding agents and IDEs. It provides a unified interface to manage injection modes and skills across platforms like Cursor, VS Code, and command-line agents via plugins, hooks, and an MCP server. The system includes a comprehensive benchmarking suite to evaluate agent behavior, code minimalism, and safety against these injected rules.

Features

Introduce OpenCode plugin for Ponytail integration

Adds a new OpenCode plugin that injects Ponytail rulesets into the system prompt, persists mode switches via a local state file, and registers slash commands and skills directories. The plugin module now exports only the plugin function, with the command-file frontmatter parser moved to a separate module to prevent OpenCode's legacy loader from misinterpreting it as a plugin.

.opencode/plugins · high confidence

Introduce Ponytail extension with mode management and status bar integration

The pi-extension now includes the Ponytail extension, which adds a /ponytail command for managing runtime modes (lite, full, ultra) and setting a default mode. It registers specific skill commands (ponytail-review, ponytail-audit, ponytail-gain, ponytail-debt, ponytail-help) that delegate to their respective skills. The extension integrates with the UI status bar to display the current mode and activity state, with guards to handle missing themes or UI contexts. It also supports deactivating Ponytail via specific commands and offers an opt-out for the startup notification via the quietStartup configuration.

pi-extension · high confidence

Introduce ponytail-mcp server for on-demand ruleset injection

Added a new MCP server (ponytail-mcp) that exposes Ponytail's lazy-senior-dev ruleset to MCP-compatible hosts via a user-invoked prompt and a read-only tool. The server supports three intensity modes (lite, full, ultra) and resolves the mode using the same configuration sources as the existing hooks and Pi extension, ensuring consistent instructions across hosts. It is designed for hosts that can only inject context through the prompt menu or tools, rather than replacing the always-on adapters.

ponytail-mcp · high confidence

Introduction of Ponytail rule for efficient, minimal code generation

A new 'Ponytail' rule has been added to the Cursor configuration, establishing a 'lazy senior dev mode' that prioritizes the simplest working solution. This rule instructs the AI to avoid unnecessary abstractions, new dependencies, and boilerplate by reusing existing code, standard libraries, or platform features. It enforces a specific decision ladder (YAGNI, reuse, stdlib, etc.) before writing code, mandates fixing root causes rather than symptoms, and requires marking deliberate simplifications with 'ponytail:' comments. The rule also emphasizes that 'lazy' does not mean careless, requiring thorough problem understanding and leaving a small runnable check for non-trivial logic.

.cursor · high confidence

Native hooks for Cursor, Copilot, Codex, and Qoder

The hooks directory now includes dedicated configuration files and scripts to support Cursor, VS Code Copilot, Codex, and Qoder alongside Claude Code. New JSON manifests (cursor-hooks.json, copilot-hooks.json, claude-codex-hooks.json, qoder-hooks.json) define the specific hook events and command structures required by each agent. The underlying JavaScript modules (ponytail-activate.js, ponytail-mode-tracker.js, ponytail-runtime.js) have been updated to detect the running environment via environment variables (e.g., CURSOR\_VERSION, COPILOT\_PLUGIN\_ROOT) and emit output in the correct format for each host, ensuring the ponytail mode and ruleset are correctly injected regardless of the IDE.

hooks · high confidence

New agentic benchmark for measuring code minimalism, safety, and completeness

A new agentic benchmark has been added to evaluate coding agents by running them as headless Claude Code sessions against a seeded codebase, rather than relying on single-shot prompt completions. This benchmark introduces two evaluation tiers: an LOC tier that measures code volume via git diffs, and a safety tier that executes produced code against adversarial inputs to detect vulnerabilities like path traversal or SQL injection. To prevent agents from 'winning' by writing less code without delivering features, a new completeness judge (complete.py) uses an LLM to verify that submissions actually implement the requested functionality. Additionally, an over-engineering judge (judge.py) assesses whether solutions are appropriately minimal. The benchmark includes specific tasks such as safe-path, rate-limit, and critic-email, and provides a portable plugin directory resolution mechanism to ensure reproducibility across different environments.

benchmarks/agentic · high confidence

New benchmark arms for caveman and ponytail communication styles

Added new benchmark arms in the \benchmarks/arms\ directory to evaluate specific communication styles. The 'caveman' arm (\caveman.js\ and \caveman-SKILL.md\) implements an ultra-compressed communication mode that reduces token usage by dropping articles and filler words while maintaining technical accuracy, supporting various intensity levels. The 'ponytail' arm (\ponytail.js\) uses the repository's own \SKILL.md\ as the system prompt, serving as a baseline for the standard assistant behavior. A 'baseline' arm (\baseline.js\) was also added, which sends only the user task without any system prompt.

benchmarks/arms · high confidence

New benchmarking suite with behavioral gates and local model support

The benchmarks directory now includes a comprehensive suite for evaluating the ponytail ruleset against baseline and caveman arms. This adds behavioral gates (behavior.js/behavior.yaml) to verify specific output qualities like hardware calibration, explanation depth, and runnable checks. It introduces a local benchmark script (benchmark-local.py) for running tests against Ollama models without API keys. The suite also features robustness audits (robustness-audit.js) and cross-model email validation checks (claude-email.js, model-email.js). Several metrics have been refined: loc.js now correctly handles CRLF line endings and strips block comments, while correctness.js now supports unfenced code blocks and uses python3 for better portability.

benchmarks · high confidence

New scripts for version consistency, rule drift checks, and host-specific integration

The scripts directory now includes several new tools to improve reliability and integration. check-versions.js ensures all seven plugin manifests and package files share a single pinned version and validates it against release tags. check-rule-copies.js verifies that rule copies across hosts (Cursor, Windsurf, Cline, etc.) match the canonical AGENTS.md and that key rule invariants are present in SKILL.md. cursor-hooks.js manages installing and uninstalling ponytail hooks in Cursor's hooks.json, preserving other user hooks. build-openclaw-skills.js and publish-openclaw-skills.js generate and publish OpenClaw/ClawHub skill packages with short descriptions. uninstall.js cleans up ponytail's state outside plugin files, including mode flags, config, statusLine entries, and Cursor hooks, handling malformed JSON gracefully.

scripts · high confidence

Ponytail v4.10.0 release with Hermes plugin and expanded agent support

This release bumps the version to 4.10.0 and introduces a new Hermes plugin (via \\_\init\\_.py\ and \plugin.yaml\) that provides always-on context injection, slash commands (e.g., \/ponytail lite\|full\|ultra\|off\), and bundled skills. The product now officially supports Gemini CLI (via \gemini-extension.json\) and OpenCode (via \opencode.json\), alongside existing support for Claude Code, Codex, and GitHub Copilot CLI. The main README has been updated to reflect support for 20 agents and includes a new 'Already built with Ponytail' section featuring the Retriever app.

(repo-wide) · high confidence

Test coverage

Added comprehensive test suite for Ponytail adapters, hooks, and benchmarks; Added tests for the Ponytail VS Code extension.

Dependencies

Initial release of Ponytail v4.10.0 with MCP server and Pi extension support

This release introduces Ponytail as an installable npm package (@dietrichgebert/ponytail) compatible with OpenCode and Pi, alongside a new private ponytail-mcp server for AI agents and a dedicated pi-extension. The package includes native hooks for Cursor and Qoder, and the MCP server dependency @modelcontextprotocol/sdk is updated to ^1.26.0 to address [CVE redacted].

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

This is the PUBLIC form of this artifact. Findings are listed in full, but the details of SECURITY findings — which rule fired, in which file, on which line, and how to fix it — are deliberately withheld, and any secret-scanner results are excluded entirely. Where detail is absent here it was REMOVED FOR PUBLICATION; it is not missing from the analysis. The complete artifact is available from the repository owner.

Score

  • CAI 42 → 54 (+12.4)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 48 → 63 (+15.2)
  • Architecture 89 → 90 (+1.8)
  • Maturity 72 → 62 (-10.5)
  • Readiness 24 → 42 (+17.8)
  • Security 58 → 67 (+8.7)

Resolved (17)

  • Boundary-crossing change coupling: ponytail-config.js ↔ index.js (hooks/ponytail-config.js)
  • Change coupling: ponytail-activate.js ↔ ponytail-config.js (hooks/ponytail-activate.js)
  • Dimension evaluation failed
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • LLM evaluation failed
  • Low: security finding (details withheld)
  • No automated tests
  • No exposed public API
  • No tests found
  • Scanner failed to run — not a clean result
  • Test reliability not included
  • finish (cognitive 57) (hooks/ponytail-mode-tracker.js)
  • finish (cyclomatic 37) (hooks/ponytail-mode-tracker.js)

New (28)

  • (anonymous) (cognitive 18) (.opencode/plugins/ponytail.mjs)
  • (anonymous) (cyclomatic 18) (.opencode/plugins/ponytail.mjs)
  • Dependency hygiene PARTLY measured — npm pinning read, dependency currency not (no committed lockfile, so no resolved version to grade)
  • Further sole-owners (lower concentration)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • High: security finding (details withheld)
  • Hotspot: benchmarks/agentic/judge.py (benchmarks/agentic/judge.py)
  • Hotspot: benchmarks/agentic/run.py (benchmarks/agentic/run.py)
  • Hotspot: benchmarks/agentic/tasks.py (benchmarks/agentic/tasks.py)
  • No dependency advisory monitoring
  • Off-boarding risk: anonymized user #1
  • PR-triggered workflow without a permissions block
  • Repeated repair: .opencode/plugins/ponytail.mjs (.opencode/plugins/ponytail.mjs)
  • Repeated repair: hooks/ponytail-config.js (hooks/ponytail-config.js)
  • Repeated repair: hooks/ponytail-instructions.js (hooks/ponytail-instructions.js)
  • Repeated repair: pi-extension/index.js (pi-extension/index.js)
  • …and 8 more

Changes since last survey

  • 18 commits — 16 feature/other, 2 fixes

By area

  • (root) — 15 commits
  • assets/retriever-icon.png — 1 commit
  • hooks/claude-codex-hooks.json — 1 commit
  • hooks/ponytail-runtime.js — 1 commit

Notable commits

  • fix: fix: detect VS Code Copilot via CLAUDE_PLUGIN_ROOT fallback (#528) (#579)
  • fix: fix: drop commandWindows from hooks.json for Claude.ai marketplace validation (#593) (#601)
  • change: chore: release v4.10.0 (#870)
  • change: chore: release v4.9.0 (#703)
  • change: docs: Already built with Ponytail as a real section, left aligned
  • change: docs: Built with Ponytail, Retriever (#830)
  • change: docs: Retriever as one logo image, icon and name together, dark and light
  • change: docs: Retriever icon and name on one line
  • change: docs: Retriever logo with the name below the icon
  • change: docs: add Trendshift badge to READMEs (#801)
  • change: docs: add Trendshift monthly ranking badge (#802)
  • change: docs: crop the transparent margin off the Retriever icon
  • change: docs: drop the tagline under the Retriever icon
  • change: docs: icon and name in one block, rule after the section
  • change: docs: put daily, weekly and monthly Trendshift badges in one row (#803)
  • change: docs: rename the Retriever logo files so GitHub serves the new version
  • change: feat: add Grok Build native skills adapter (revive #561) (#661)
  • change: feat: add native Cursor hooks via hooks.json (#817) (#869)

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

DietrichGebert/ponytail was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 25 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit e3ba2aa6f1e6f0bc4d69eb09c9f0d0a93af56156 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-dd72cc24c749.