Skip to content
CAI
Software that uses CAICheck a score

garrytan/gstack

52.4

Adequate · 28 September 2026

83.9k

lines of production code

TypeScript

primary language

4

measurements over time

CAI band scale
CAI trend line
CAI lens gauges

What this system is

This system is a modular skill framework for AI coding agents that orchestrates the entire software development lifecycle through specialized, composable workflows. It provides capabilities for automated code review, design system generation, visual QA, and performance benchmarking, while enforcing safety through destructive command guardrails and directory-scoped edit restrictions. The platform also supports deployment automation, iOS device testing, and structured documentation management, integrating with multiple external AI providers to enhance code quality and developer experience.

Features

Added Claude Code skill for external diff reviews and consultations

Introduced a new \SKILL.md.tmpl\ file that enables the system to invoke Claude Code for code reviews, adversarial challenges, and repository consultations from outside the native Claude Code environment. This skill defines the execution boundary, allowing users to trigger reviews via \claude review\ or \claude challenge\ commands, which route to the \gstack-claude-code\ runner to analyze diffs against a base branch with strict security constraints (no tools, read-only access for consultations) and output formatting.

claude-code · high confidence

Added Hacker News frontpage skill with browser automation client

The browser-skills/hackernews-frontpage location now includes a new browse-client SDK (browse-client.ts) that allows the skill to drive a browser via a local gstack daemon. This client handles authentication through environment variables or a state file, manages timeouts, and provides methods for navigation and interaction, enabling the Hacker News frontpage skill to automate browser tasks.

browser-skills/hackernews-frontpage · high confidence

Added Windows Node.js server build script and extension ID utility

This change introduces two new scripts in the browse/scripts directory to support Windows environments and extension security. The build-node-server.sh script generates a Node.js-compatible server bundle for Windows, addressing Playwright compatibility issues on that platform by transpiling the server code and injecting Bun API polyfills. Additionally, the extension-id.ts script provides a utility to derive the Chrome extension ID from the public key in the manifest, enabling the browse server to verify the Origin header against a known extension identity for secure token handling.

browse/scripts · high confidence

Added contributor documentation for creating new host configurations

A new template file (SKILL.md.tmpl) has been added to the contrib/add-host directory to guide contributors in adding support for new AI coding agents. This document outlines the steps to create a host configuration file, including gathering host-specific information, creating the TypeScript config, registering it in the index, updating .gitignore, and verifying the setup through documentation generation and tests.

contrib · high confidence

Added design prototype script for UI mockup generation

A new prototype script (design/prototype.ts) has been added to validate the generation of UI mockups using the OpenAI GPT-4o image generation capabilities. This tool sends design briefs for a dashboard, landing page, and mobile app to the OpenAI Responses API, saving the resulting high-quality images to a temporary directory for review.

design · high confidence

Added shared-code extraction audit skill

The deslop-shared-libs area now includes a new SKILL.md and its template (SKILL.md.tmpl) that define a read-only audit process for identifying opportunities to extract shared code. This skill guides the review of recent commits and open pull requests to find candidates for code reuse, enforcing strict safety boundaries (no file writes, no execution of project code) and requiring verification of callers and source revisions before recommending any shared-library extractions.

deslop-shared-libs · high confidence

Added shell script helpers for binary discovery and remote identification

The browse/bin directory now includes two new executable scripts: find-browse, which acts as a shim to delegate to a compiled binary or fall back to discovering the browse tool via standard agent skill directories, and remote-slug, which extracts the owner-repo identifier from the git remote URL to support project-specific path resolution.

browse/bin · high confidence

CEO review skill adds Selective Expansion mode and cross-session context memory

The plan-ceo-review skill now includes a fourth review mode, Selective Expansion, allowing users to cherry-pick specific scope expansions while holding the rest of the plan. The skill also gains cross-session memory via gbrain context queries, automatically retrieving prior CEO plans, recent design docs, and past review activity to provide continuous, informed reviews. Additionally, it now supports WebSearch for landscape checks and includes a handoff mechanism to resume paused reviews from /office-hours sessions.

plan-ceo-review · high confidence

Codex skill introduces Challenge and Consult modes with robust error handling

The Codex skill now supports two new interaction modes alongside the existing Review mode: Challenge mode (adversarial testing for edge cases and security holes) and Consult mode (free-form codebase questions with session continuity). To ensure reliability, all Codex execution paths now include explicit hang detection (logging and surfacing messages for timeouts), stderr capture to surface authentication errors, and a three-way completeness check that distinguishes between stated failures and network disconnects. The skill also enforces a fail-closed gate in Review mode, requiring explicit severity tags to verify results, and mandates a synthesis recommendation line to guide user action.

codex/sections · high confidence

Declarative multi-host platform for external AI agents

The system now supports a declarative configuration model for integrating external AI coding agents. A new \defineHost\ factory standardizes host definitions, allowing each agent to specify its own path rewrites, tool-name mappings, frontmatter rules, and resolver behaviors. This change introduces built-in support for OpenCode, Slate, Cursor, OpenClaw, Hermes, GBrain, Codex, and Factory Droid, ensuring that generated skills and instructions are correctly adapted to each platform's specific runtime and CLI requirements.

hosts · high confidence

Initial repository scaffolding and configuration

The repository is initialized with essential configuration files to support development and security workflows. A \.env.example\ file is added to document the \ANTHROPIC\_API\_KEY\ required for LLM-as-judge evaluations. A \.gitattributes\ file enforces LF line endings across all text, script, and template files to prevent cross-platform test failures and execution errors on Windows. GitLab CI configuration (\.gitlab-ci.yml\) is introduced to mirror GitHub protections, enforcing version-gate checks and PR title synchronization. Finally, an \.osv-scanner.toml\ file is added to explicitly suppress specific security advisories that are deemed unreachable or unfixable without disproportionate risk, ensuring the security scanner does not fail on known, accepted vulnerabilities.

(repo-wide) · high confidence

Introduce 'careful' skill for destructive command guardrails

A new 'careful' skill has been added to provide safety guardrails for destructive commands. When activated (via triggers like 'be careful' or 'safety mode'), it intercepts Bash commands to warn users before executing high-risk operations such as \rm -rf\, \DROP TABLE\, force-pushes, or \kubectl delete\. The skill distinguishes between MEDIUM warnings (which allow user override) and HIGH tier denials (hard stops for catastrophic actions like deleting the home directory or force-pushing to the default branch). It also supports additive, project-specific custom patterns via configuration files.

careful · high confidence

Introduce 'learn' skill for managing project knowledge

A new 'learn' skill has been added to the system, allowing users to review, search, prune, and export project learnings captured across sessions. The skill provides commands to show recent insights, search by query, prune stale or contradictory entries, export knowledge to markdown, and view statistics. It integrates with the gstack infrastructure via a preamble script and supports both interactive and spawned session modes, including specific handling for AskUserQuestion failures and plan mode operations.

learn · high confidence

Introduce /autoplan auto-review pipeline skill

Adds the \autoplan\ skill, an automated review pipeline that sequentially executes CEO, design, DX, and engineering review phases. The skill uses six decision principles to auto-approve mechanical choices and surfaces only taste-based disagreements or user challenges at a final approval gate, allowing users to run a full review gauntlet without answering intermediate questions. It includes robust handling for interactive, headless, and spawned sessions, with specific fallbacks for AskUserQuestion failures and strict phase execution ordering.

autoplan · high confidence

Introduce /qa skill for systematic web application testing and bug fixing

The new /qa skill enables the agent to systematically test web applications and iteratively fix bugs found during testing. It supports three testing tiers (Quick, Standard, Exhaustive) to balance coverage and speed, and automatically enters diff-aware mode when run on a feature branch to verify recent changes. The skill uses the Aside browser for interactive testing, captures screenshots as evidence, and produces structured reports with before/after health scores and a ship-readiness summary. It includes robust handling for AskUserQuestion decisions, including auto-decision rules for spawned sessions and prose fallbacks for interactive sessions, ensuring reliable operation across different execution contexts.

qa · high confidence

Introduce /setup-gbrain skill for coding-agent onboarding

Adds the \setup-gbrain\ skill, providing a guided, step-by-step workflow for users to install the gbrain CLI, initialize a local PGLite or Supabase database, or connect to a remote MCP brain, and register the tool with the coding agent. The skill includes robust state detection, broken-engine remediation, and detailed documentation on memory ingest, transcript handling, and per-repo trust policies to ensure secure and correct configuration.

setup-gbrain · high confidence

Introduce Codex skill with multi-mode review, challenge, and consult capabilities

Adds the Codex skill, enabling users to invoke the OpenAI Codex CLI for an independent second opinion via three modes: review (diff analysis with pass/fail gate), challenge (adversarial testing), and consult (general questions with session continuity). The skill includes a robust preamble that handles environment setup, degraded-mode fallbacks, and strict AskUserQuestion protocols (including prose fallbacks for Conductor sessions and auto-decision rules for spawned sessions) to ensure reliable user interaction across different execution contexts.

codex · high confidence

Introduce GStack Browser skill for AI-controlled Chromium automation

Users can now launch an AI-controlled Chromium browser instance via the new 'open-gstack-browser' skill, triggered by commands like 'open gstack browser' or 'launch chromium'. This skill provisions a visible, headed browser window with a built-in sidebar extension for real-time activity feeds and chat, featuring anti-bot stealth capabilities. The workflow handles setup, pre-flight cleanup of stale processes, connection verification on port 34567, and guides the user through pinning the extension, allowing for direct observation and control of browser actions by the AI agent.

open-gstack-browser · high confidence

Introduce dedicated design binary with persistent board daemon and vision-based quality gates

Adds a new stateless design CLI (\design/src/cli.ts\) for AI-powered UI mockup generation, featuring a persistent background daemon (\design/src/daemon.ts\) that hosts comparison boards for user review and feedback. The toolchain includes vision-based quality checks (\design/src/check.ts\) using GPT-4o to verify text readability and layout completeness, visual diffing (\design/src/diff.ts\) to compare mockups against live implementations, and a design-to-code prompt generator (\design/src/design-to-code.ts\). Authentication is handled via a new \design/src/auth.ts\ module that resolves API keys from \\~/.gstack/openai.json\ or environment variables, with specific safeguards to warn users if their key matches a local \.env\ file to prevent silent billing of other projects. The daemon client (\design/src/daemon-client.ts\) manages the lifecycle of the persistent board server, including identity verification and version mismatch handling to prevent accidental process termination or data loss.

design/src · high confidence

Introduce gstack browse Chrome extension with terminal, inspector, and ref overlays

This change adds the initial files for the 'gstack browse' Chrome extension (Manifest V3), providing a side panel that hosts an interactive Claude Code terminal (via xterm.js) and debug tools including a CSS inspector and element ref overlays. The extension connects to a local browse server on port 34567, handling authentication via a pinned extension identity and managing connection state through health polling. It includes a background service worker for command proxying and token management, content scripts for rendering ref badges and status pills on web pages, and a popup for port configuration.

extension · high confidence

Introduce make-pdf CLI for markdown-to-PDF conversion

This change adds the core implementation for the make-pdf tool, enabling users to convert Markdown documents into publication-quality PDFs. The new codebase includes a CLI entry point with argument parsing, a command registry, and an orchestrator that manages the rendering pipeline. Key capabilities include rendering Mermaid and Excalidraw diagrams, handling image inlining and size policies (including automatic landscape promotion for wide images), and supporting features like tables of contents, cover pages, watermarks, and page numbering. The tool uses the Aside browser (or a headless fallback) for rendering and includes a pdftotext wrapper for CI validation.

make-pdf/src · high confidence

Introduce make-pdf skill for publication-quality PDF generation

Adds the make-pdf skill, enabling users to convert markdown files into high-quality PDFs with features like proper margins, page numbers, clickable tables of contents, and draft watermarks. The skill supports multiple output formats including HTML and DOCX, handles embedded diagrams (Mermaid, Excalidraw) and images, and intelligently manages font rendering for emoji and special characters across Linux, macOS, and Windows. It integrates with the gstack browser infrastructure, using the Aside browser when available or falling back to a headless browser, and includes robust error handling and CI-mode strictness for asset validation.

make-pdf · high confidence

Introduce plan-tune skill for question sensitivity and developer profile management

Adds the plan-tune skill, enabling users to tune how often gstack asks questions, set per-question preferences (never-ask, always-ask, or one-way), and view a dual-track developer profile that compares declared preferences against inferred behavior. The skill includes a consent gate, a 5-question setup wizard for initial profile declaration, and a 'dream cycle' feature to distill free-text feedback into actionable profile updates. It also defines strict formatting and fallback rules for AskUserQuestion prompts, including ELI10 explanations, completeness scores, and prose fallbacks for interactive sessions where tool calls fail.

plan-tune · high confidence

Introduces fail-closed freeze boundary enforcement hooks

Adds \check-freeze.sh\ and \freeze-state.sh\ to the \freeze/bin\ directory to enforce file-edit boundaries with strict security guarantees. The new \check-freeze.sh\ hook validates that tool inputs target files within the allowed directory, resolving symlinks to prevent boundary escapes and failing closed (denying access) if parsing fails, helpers are missing, or the state is ambiguous. The \freeze-state.sh\ script manages the boundary state file with atomic writes and mutex locking to prevent race conditions, ensuring that the freeze configuration is preserved or released safely.

freeze/bin · high confidence

New 'health' skill for code quality dashboards

Added a new \health\ skill that acts as a code quality dashboard. It automatically detects or uses a configured 'Health Stack' of tools (type checkers, linters, test runners, dead code detectors, shell linters, and GBrain) to run checks, compute a weighted 0-10 composite score, and track trends. The skill handles complex decision flows via \AskUserQuestion\, including auto-decision logic for spawned/conductor sessions and prose fallbacks for interactive sessions, ensuring robust operation across different execution environments.

health · high confidence

New 'investigate' skill for systematic root-cause debugging

Adds the 'investigate' skill (v1.0.0), a structured workflow for debugging that enforces a four-phase process: investigate, analyze, hypothesize, and implement, with the 'Iron Law' that no fixes are applied without first identifying the root cause. The skill integrates with the gstack infrastructure to load prior investigation history and project learnings, and includes a 'Scope Lock' mechanism (using check-freeze.sh) to prevent scope creep during edits. It features robust decision-making via AskUserQuestion, supporting both tool-based and prose fallbacks, with specific handling for spawned/Conductor sessions and destructive actions.

investigate · high confidence

New 'skillify' skill codifies scrape flows into permanent browser skills

A new 'skillify' skill (SKILL.md and SKILL.md.tmpl) allows users to convert the most recent successful /scrape flow into a permanent, deterministic browser skill on disk. This enables future /scrape calls with the same intent to run in \~200ms by executing the codified script instead of re-driving the page. The skill synthesizes a TypeScript script, a test file, and a fixture, runs the test in a temporary directory, and only commits the skill to disk after a passing test and explicit user approval, ensuring no broken skills are left on the system.

skillify · high confidence

New /benchmark skill for performance regression detection

A new \benchmark\ skill has been added to the gstack suite, enabling automated performance regression detection. It establishes baselines for page load times, Core Web Vitals (TTFB, FCP, LCP), and resource sizes, then compares current metrics against those baselines to identify regressions. The skill uses the Aside browser to collect real browser metrics and supports modes for capturing baselines (\--baseline\), quick checks (\--quick\), and diff-based benchmarking (\--diff\).

benchmark · high confidence

New /design-consultation skill for end-to-end design system creation

Introduces the /design-consultation skill, which guides users through creating a complete design system (aesthetic, typography, color, layout, spacing, motion) and generating a DESIGN.md file. The skill features a multi-phase workflow: gathering product context, optional competitive research via the Aside browser, and generating a full proposal with independent 'outside voices' from Codex and Claude. It includes robust anti-slop guardrails against common AI design patterns, font verification procedures, and a preview phase that generates either AI mockups or a self-contained HTML preview page for user approval.

design-consultation · high confidence

New /diagram skill for generating editable diagram triplets

A new /diagram skill has been added that converts English descriptions or Mermaid source code into a complete diagram triplet: a Mermaid source file (.mmd), an editable Excalidraw scene file (.excalidraw), and rendered SVG and PNG images. The skill operates fully offline using a local render bundle and supports round-trip editing via Excalidraw, allowing users to modify the visual layout and re-render the output without losing fidelity.

diagram · high confidence

New /document-release skill for post-ship documentation updates

A new \/document-release\ skill has been added to handle post-ship documentation updates. It automatically analyzes code diffs to identify changes in features, behavior, and infrastructure, then cross-references these against existing documentation (README, ARCHITECTURE, CONTRIBUTING, CLAUDE.md) to detect drift and coverage gaps using a Diataxis-based audit lens. The skill polishes CHANGELOG voice, cleans up TODOs, and surfaces documentation debt in the PR body, while ensuring version bumps and risky narrative changes require explicit user confirmation via AskUserQuestion.

document-release · high confidence

New /freeze skill to restrict file edits to a specific directory

A new 'freeze' skill has been added that allows users to lock file editing to a specific directory for the current session. When invoked, the user specifies a target path, and the system installs PreToolUse hooks on the Edit and Write tools to check every operation against this boundary. Any attempt to edit or write a file outside the allowed directory is blocked (denied), while operations within the directory proceed normally. The restriction persists until explicitly removed via the /unfreeze command and is designed to prevent accidental changes to unrelated code during debugging or focused work.

freeze · high confidence

New /guard skill combines destructive command warnings and directory-scoped edit restrictions

A new 'guard' skill has been added that activates both destructive command warnings (via the /careful hook) and directory-scoped edit restrictions (via the /freeze hook) in a single mode. Users can invoke this 'full safety mode' to protect against accidental deletions or modifications outside a specified directory boundary, with the system prompting for the target directory path upon activation.

guard · high confidence

New /setup-deploy skill for automated deployment configuration

A new 'setup-deploy' skill has been added to configure deployment settings for the /land-and-deploy workflow. This skill detects the user's deploy platform (Fly.io, Render, Vercel, Netlify, Heroku, GitHub Actions, or custom), identifies the production URL, health check endpoints, and deploy status commands. It then writes this configuration to CLAUDE.md, enabling automatic future deploys. The skill includes robust handling for plan mode, specific AskUserQuestion formatting for decision briefs, and fallback logic for degraded skill-start states.

setup-deploy · high confidence

New /spec skill: five-phase spec authoring with quality gates and redaction

The \spec\ skill is now available to turn vague requests into precise, executable specifications. It guides users through five phases: understanding the 'why', scoping boundaries, technical interrogation (reading code first), draft review, and a quality gate before filing. The quality gate includes a semantic content review and a fail-closed redaction scan for secrets/PII/legal patterns, ensuring no sensitive data is persisted or dispatched. The skill can file the spec as a GitHub issue, archive it locally, and optionally spawn an agent to implement it. It supports plan mode, deduplication checks, and configurable execution paths.

spec · high confidence

New /unfreeze skill to clear edit restrictions

A new 'unfreeze' skill has been added to the system, allowing users to remove the directory edit restrictions previously set by the /freeze command without ending their session. This skill triggers on commands like 'unfreeze', 'unlock edits', or 'remove edit restrictions', and executes a script to clear the freeze state while logging the action to analytics. It ensures that existing /freeze hooks remain registered but effectively allow all edits until a new freeze is applied.

unfreeze · high confidence

New Claude Code hooks for AskUserQuestion reliability, user preferences, and telemetry repair

This change introduces a suite of new hooks in the Claude Code host to improve AskUserQuestion (AUQ) reliability, enforce user preferences, and fix telemetry leaks. The \auq-error-fallback-hook\ detects when an AUQ call fails or returns a missing result, injecting context to guide the model toward a safe fallback (auto-choosing in spawned/headless sessions, or rendering a prose brief in interactive ones). The \question-preference-hook\ enforces user-defined never-ask/always-ask rules for specific questions, allowing users to bypass prompts for routine decisions. The \question-log-hook\ ensures that AUQ interactions are reliably logged for analytics, regardless of agent behavior. Additionally, the \timeline-stop-hook\ repairs dangling timeline entries by marking 'started' events as 'completed' when a session ends unexpectedly, preventing telemetry leaks. These hooks are supported by \spawn-bin.ts\ for robust, cross-platform subprocess execution and \spawned-directive.ts\ for consistent behavior in orchestrated environments.

hosts/claude · high confidence

New Hacker News front-page scraping skill

A new browser skill named \hackernews-frontpage\ has been added to the \browser-skills\ directory. When executed, it navigates to the Hacker News front page, parses the top 30 stories, and outputs a JSON document containing each story's rank, title, URL, point count, and comment count. The implementation includes a dedicated parser script, a bundled HTML fixture for testing, and a test suite that validates the parsing logic against the fixture without requiring a live network connection.

browser-skills · high confidence

New QA skill sections for test framework bootstrap and core QA patterns

The QA skill now includes a registry manifest and documentation for two new sections: 'Test Framework Bootstrap' and 'Core QA patterns'. The Test Framework Bootstrap section provides a structured workflow for detecting a project's runtime and test framework, offering to install and configure a testing suite (including CI pipeline generation) if none is found, while respecting existing test configurations. The Core QA patterns section defines the QA baseline workflow (Phases 1-6), detailing modes such as Diff-aware, Full, Quick, and Regression, along with specific instructions for browser-based testing, issue documentation, and health scoring.

qa/sections · high confidence

New browser driver and skill composition resolvers

Added new resolver scripts in scripts/resolvers/ to manage the Aside browser integration and skill composition. The aside.ts and browse.ts resolvers define the Aside browser driver contract (detection, rules, and cookbook) and the gstack headless browser fallback, including a unified untrusted-content warning to prevent prompt injection. The composition.ts resolver introduces the {{INVOKE\_SKILL}} directive for composable skills and host-specific autoplan hooks. Additional resolvers (confidence.ts, constants.ts, design-checklist.ts, design-doc-discovery.ts) provide confidence calibration for reviews, shared Codex and Claude Code constants, a generated design review checklist, and design document discovery logic.

scripts/resolvers · high confidence

New browser-skill infrastructure with secure, scoped execution and auditability

This change introduces the core backend for the browser-skill system in browse/src, enabling agents and users to install, run, and manage Playwright-based automation scripts. It adds a scoped authentication model via browse-client.ts, which resolves daemon tokens from environment variables or a state file to ensure skills only access the browser with limited, per-spawn permissions. Security is enforced through browser-skill-commands.ts, which spawns scripts with scrubbed environments and locked working directories for untrusted skills, and browser-skill-write.ts, which uses atomic staging and commit logic to prevent partial or malicious skill installations. The system also includes robust operational features: activity.ts provides a real-time, privacy-filtered (redacting passwords/tokens) streaming feed of browser commands for the Chrome extension, while audit.ts maintains a persistent, append-only JSONL log of all commands for forensic review. Finally, cdp-allowlist.ts and cdp-bridge.ts implement a strict default-deny policy for Chrome DevTools Protocol access, ensuring only audited, safe methods are exposed to skills.

browse/src · high confidence

New cross-model benchmark skill for comparing AI providers

Added a new \benchmark-models\ skill that allows users to run side-by-side comparisons of AI models (Claude, GPT, Gemini) on specific prompts or existing skills. The skill guides users through selecting a prompt, verifying provider authentication via a dry-run, optionally enabling a quality judge, and executing the benchmark to compare latency, cost, and output quality. It includes logic to handle plan mode, telemetry logging, and saving results for future trend analysis.

benchmark-models · high confidence

New design-html skill for Pretext-native HTML generation

The design-html skill is now available to generate production-quality HTML/CSS using the Pretext engine for dynamic text layout. It accepts input from approved mockups, CEO plans, or direct user descriptions, and includes a one-time offer to install the 'impeccable' design detector engine for automated CSS validation. The skill provides specific JavaScript patterns for height computation, tight-fit containers, and editorial layouts, ensuring text reflows and heights adjust correctly to content.

design-html · high confidence

New design-review skill for visual QA and automated fixes

A new \design-review\ skill (v2.0.0) has been added to the gstack suite. It performs visual design audits—checking for inconsistencies, spacing issues, hierarchy problems, and AI-generated patterns—and iteratively fixes them in source code with atomic commits and before/after verification. The skill supports plan-mode reviews, integrates with the Aside browser for live site inspection, and uses a structured AskUserQuestion format for decision briefs, including prose fallbacks for interactive sessions and auto-decision rules for spawned or conductor sessions.

design-review · high confidence

New design-shotgun skill for visual design exploration and iteration

Introduces the \design-shotgun\ skill, enabling users to generate multiple AI design variants, view them side-by-side in a comparison board, and collect structured feedback to iterate on visual directions. The skill includes a preamble for session detection and degraded-mode handling, a robust \AskUserQuestion\ protocol with prose fallbacks for decision briefs, and a 'taste memory' system that learns from approved/rejected variants to bias future generations. It also incorporates a UX principles doctrine to guide design decisions based on observed user behavior.

design-shotgun · high confidence

New development and usage tooling for gstack

This change introduces several new command-line utilities to the \bin/\ directory to improve the developer experience and provide local usage insights. \bin/dev-setup\ and \bin/dev-teardown\ allow developers to easily switch between local repository skills and the global installation by managing symlinks in \.claude/skills/\ and \.agents/skills/\, while also handling environment variable propagation for worktrees. \bin/gstack-analytics\ provides a local, privacy-preserving dashboard that aggregates usage data from local JSONL logs, displaying skill run counts, success rates, and average durations. Additionally, \bin/gstack-artifacts-init\ and \bin/gstack-artifacts-url\ facilitate the setup of a private Git repository for syncing gbrain artifacts (like CEO plans and designs) to remote hosts like GitHub or GitLab, including helper functions for canonical URL management.

bin · high confidence

New devex-review skill for live developer experience audits

A new \devex-review\ skill has been added to the gstack suite, enabling live, automated testing of the developer experience. This skill navigates documentation and getting-started flows in the Aside browser, times to-first-helpful-work (TTHW), captures screenshots of error messages, and evaluates CLI help text to produce a scored DX scorecard with evidence. It includes a 'boomerang' feature that compares current live results against prior \/plan-devex-review\ predictions, and provides detailed instructions for handling AskUserQuestion interactions, including prose fallbacks for constrained environments like Conductor sessions.

devex-review · high confidence

New document-generate skill for creating structured documentation

Added the /document-generate skill, which uses the Diataxis framework to produce tutorials, how-tos, references, and explanations for features, modules, or entire projects. The skill supports standalone invocation or integration with /document-release to fill coverage gaps, and includes robust handling for plan mode, spawned sessions, and AskUserQuestion failures via prose fallbacks.

document-generate · high confidence

New documentation release section for per-file audits and CHANGELOG polishing

Added a new \document-release/sections\ manifest and \release-body\ skill that guides the automated documentation release process. This section defines a structured workflow for auditing individual documentation files (README, ARCHITECTURE, CONTRIBUTING) against code diffs, applying factual auto-updates, and polishing CHANGELOG entries without overwriting existing content. It also includes steps for cross-document consistency checks, TODO cleanup, and version bump decisions, ensuring documentation stays aligned with code changes while preserving author intent.

document-release/sections · high confidence

New gstack-upgrade skill for safe, automated version updates

A new skill named 'gstack-upgrade' (v1.1.0) has been added to handle upgrading the gstack installation. It supports both git-based and vendored install types, automatically detecting the current setup to apply the correct upgrade logic. The skill features an inline upgrade flow that respects user preferences, including an auto-upgrade mode, snooze functionality with escalating backoff (24h, 48h, 1 week), and the ability to disable update checks entirely. For git installs, it prioritizes fast-forward merges and safely handles local changes by stashing or requiring explicit user confirmation before destructive resets. For vendored installs, it uses a backup-and-replace strategy with automatic restoration on failure, ensuring users are not left with a broken installation.

gstack-upgrade · high confidence

New iOS Debug Bridge templates for automated QA

Added a suite of Swift and Objective-C templates to the ios-qa toolchain that generate a local debug server and UI overlay for iOS app testing. The generated code includes a loopback StateServer for reading and writing app state, a touch-synthesis engine for reliable SwiftUI and UIKit interaction, and a visual DebugOverlay to indicate active testing sessions. These templates are strictly gated to DEBUG builds to prevent private API exposure in release binaries.

ios-qa/templates · high confidence

New iOS QA and Debug Bridge skills for live-device testing

Added \ios-qa\ and \ios-sync\ skills to enable live-device quality assurance and debug bridge management for SwiftUI apps. The \ios-qa\ skill connects to a real iPhone via USB (using CoreDevice) to run a vision-driven agent loop that captures screenshots, analyzes the accessibility tree, and performs automated tap/swipe/type interactions against an embedded StateServer, with optional remote access via Tailscale. The \ios-sync\ skill allows users to regenerate the typed \@Observable\ state accessors and bridge artifacts when adding new ViewModels or upgrading the gstack framework. Both skills include robust fallback logic for \AskUserQuestion\ failures, supporting spawned, headless, and interactive session modes, and handle plan-mode execution constraints.

ios-qa, ios-sync · high confidence

New iOS design review and autonomous bug-fix skills

Two new gstack skills are introduced to close the loop on iOS visual quality and defects. The ios-design-review skill connects to a real iPhone via the gstack StateServer to screenshot every screen and evaluate it against Apple HIG, DESIGN.md, and best practices, scoring ten dimensions (typography, spacing, color, touch targets, loading/empty/error states, accessibility, animation, iOS idioms, information density, and AI-slop) on a 0–10 scale with actionable feedback. The ios-fix skill takes a bug reported by /ios-qa, captures a reproducing state snapshot as a regression fixture, locates the root cause in Swift source, applies a minimal fix, rebuilds and redeploys to the device, verifies the fix against the snapshot, and adds a regression test in test/fixtures/ios-fix. Both skills use the same preamble protocol (gstack-skill-start) for session/onboarding/telemetry, support plan mode with AskUserQuestion decision briefs (including prose fallbacks and spawned/Conductor handling), and are triggered by voice/text commands such as 'review the iOS design' and 'fix this iOS bug'.

ios-design-review, ios-fix · high confidence

New interactive developer experience (DX) plan review skill

Adds the \/plan-devex-review\ skill, an interactive tool for auditing developer-facing plans (APIs, CLIs, SDKs, libraries, platforms, and docs). The skill guides users through a structured review process—starting with developer persona identification and competitive benchmarking, then moving through eight specific review passes (Getting Started, API/CLI Design, Error Messages, Documentation, Migration Paths, Tooling, Community, and Measurement). It enforces rigorous decision-cadence rules using \AskUserQuestion\ with structured decision briefs (including ELI10 explanations, completeness scores, and pros/cons) to ensure explicit user approval for scope and design choices. The skill includes a reference \dx-hall-of-fame.md\ with industry gold standards and anti-patterns to ground the review in evidence rather than opinion.

plan-devex-review · high confidence

New interactive plan-design-review skill for pre-implementation design critique

A new /plan-design-review skill has been added, allowing users to review design plans before implementation. The skill acts as a senior product designer, rating design dimensions on a 0-10 scale, identifying missing decisions, and suggesting fixes. It features a strict scope gate to confirm the review target (branch diff, plan doc, or specific path), integrates with the gstack designer to generate visual mockups for UI plans, and enforces a structured decision-brief format for user interactions. The skill is triggered by phrases like 'design plan review' or 'review ux plan' and is designed to work in plan mode, ensuring design quality is addressed early in the development cycle.

plan-design-review · high confidence

New ios-clean skill to remove DebugBridge instrumentation

Added a new \ios-clean\ skill (SKILL.md and template) that provides a guided workflow to remove the DebugBridge SPM package, \#if DEBUG wiring, StateServer/DebugOverlay artifacts, and generated StateAccessor files from an iOS app. The skill handles inventory, removal, and verification (including a Release build check) while deferring to the structural Package.swift guard for safety, and integrates with the gstack preamble for session context and AskUserQuestion decision flows.

ios-clean · high confidence

New land-and-deploy skill for automated PR merging and production verification

Introduces the land-and-deploy skill, which automates the workflow of merging a Pull Request, waiting for CI and deployment, and verifying production health via canary checks. The skill triggers on commands like 'merge', 'land', or 'deploy', taking over after the /ship command creates the PR. It includes a robust preamble that initializes the gstack environment, handles degraded modes if the start script is unavailable, and manages user interactions through a structured AskUserQuestion format with prose fallbacks for interactive sessions. The workflow enforces pre-flight checks (authentication, PR state, branch matching), pre-merge validation (CI status, conflicts), and a first-run dry-run validation to confirm deployment configuration before proceeding. It specifically notes that GitLab support is not yet implemented, requiring manual merging in that case.

land-and-deploy · high confidence

New land-and-deploy skill sections for first-run validation, pre-merge readiness, and merge/deploy execution

The land-and-deploy skill now includes a structured set of sections to guide the user through the entire landing and deployment workflow. A new first-run validation step (Step 1.5) detects the deployment platform, validates read-only access, and confirms the setup before any changes are made. A pre-merge readiness gate (Step 3.5) enforces safety by checking review staleness, test results, and PR accuracy, offering an inline review if needed. Finally, the merge and deploy sections (Steps 4-5) handle the actual PR merge with robust state readbacks, manage merge queues, and detect CI auto-deployments to provide clear status updates to the user.

land-and-deploy/sections · high confidence

New landing-report skill for version queue visibility

Added a new \landing-report\ skill that provides a read-only dashboard of the version queue. It shows which version slots are claimed by open PRs, identifies sibling Conductor workspaces with work likely to ship soon, and indicates which version slot \/ship\ would claim next. The skill uses the \gstack-next-version\ utility to query micro, patch, minor, and major bump levels, rendering a table of claimed versions, active siblings, and potential collisions, with offline fallbacks when the queue is unavailable.

landing-report · high confidence

New library modules for rendering, project identity, and Claude Code integration

This change introduces several new core library modules: \lib/aside-render.ts\ provides a unified HTML rendering engine that uses the Aside browser as the primary path and falls back to the gstack browse daemon; \lib/bin-context.ts\ implements native project root and remote resolution logic to derive consistent project slugs; \lib/claude-bin.ts\ handles cross-platform resolution of the Claude Code binary with environment overrides; \lib/claude-code.ts\ and \lib/claude-code-windows-job.ts\ provide supervised, restricted invocation of Claude Code for outside reviews, including Windows process containment; \lib/claude-code-migration.ts\ manages the renaming and migration of gstack-claude skills to gstack-claude-code; and \lib/claude-public-transcript.ts\ and \lib/autoplan-phase-publication.ts\ add parsing for Claude's native transcript data and phase completion detection.

lib · high confidence

New offline diagram rendering engine supporting Mermaid and Excalidraw

The system now includes a self-contained, offline-capable diagram rendering engine located in lib/diagram-render. This engine bundles Mermaid (v11.16.1) and Excalidraw (v0.18.1) into a single HTML page, enabling the conversion of Mermaid diagrams to editable Excalidraw scenes and subsequent rasterization to PNG. It is integrated into the product via the Aside browser and the gstack headless browser, allowing users to generate diagrams for PDFs and the /diagram skill without external network dependencies.

lib/diagram-render · high confidence

New pair-agent skill for remote browser sharing

Added a new 'pair-agent' skill that allows a local AI agent to share its browser session with a remote AI agent (such as OpenClaw, Hermes, Codex, or Cursor). The skill generates a one-time setup key and provides instructions for the remote agent to connect via an ngrok tunnel (for remote machines) or direct config write (for local machines), enabling the remote agent to browse the web using the local browser's tabs and cookies.

pair-agent · high confidence

New plan design review section with 7-pass evaluation and AI-slop hard rules

The plan-design-review skill now includes a dedicated review-sections module that enforces a structured 7-pass design review process. This module introduces strict anti-skip and anti-shortcut rules, requiring individual user decisions for every unresolved finding rather than batching approvals. It adds a new Pass 4 focused on 'AI Slop Risk,' which classifies UI modes (Persuade, Operate, Read, Experience) and applies specific hard-rejection criteria and litmus checks to prevent generic, template-like designs. The review also mandates explicit handling of interaction states, user journey emotional arcs, design system token alignment, and responsive/accessibility specs, with all outcomes recorded in a standardized completion summary.

plan-design-review/sections · high confidence

New plan review section manifest and 11-section review guide

Added a new \plan-ceo-review/sections\ area containing a \manifest.json\ that registers the '11-section deep review' skill and a \review-sections.md\ file (generated from a template) that defines the review workflow. This guide instructs the reviewer to evaluate plans across 11 specific sections—including Architecture, Error & Rescue, Security, and Data Flow—using a structured 'Analyze → Resolve → Apply' decision gate process, while enforcing strict anti-skip rules and mode-specific scope handling.

plan-ceo-review/sections · high confidence

New plan-eng-review sections manifest and review workflow documentation

Added a new \plan-eng-review/sections\ area containing a \manifest.json\ that registers the 'plan-eng-review' skill and a \review-sections.md\ file (generated from a template) that defines the review preparation, record/write policy, prior learnings, retrospective learning, confidence calibration, and decision procedure for the review process.

plan-eng-review/sections · high confidence

New report-only QA skill for bug detection without code changes

A new 'qa-only' skill has been added to the system, allowing users to request a structured QA report (including health scores, screenshots, and repro steps) without triggering any code fixes. This skill is triggered by phrases like 'just report bugs' or 'qa report only' and is designed to run in diff-aware mode by default when on a feature branch, producing reports in \.gstack/qa-reports/\. It operates as a read-only counterpart to the full \/qa\ skill, ensuring that testing activities do not inadvertently modify the codebase.

qa-only · high confidence

New review sections: plan completion audit, specialist dispatch, and adversarial review

The review skill now includes three new sections that enhance the review process. The plan completion audit (plan-completion.md) performs a deep pass to discover plan files, extract actionable items, and cross-reference them against the diff to identify scope drift and missing requirements. The specialist dispatch (review-army.md) introduces a 'Review Army' that automatically detects the project stack and test framework, then dispatches parallel subagents for testing, maintainability, security, performance, data migration, API contract, design, and simplification reviews, with adaptive gating based on past hit rates. The adversarial review (adversarial.md) adds an always-on adversarial review step that runs both a Claude subagent and optional Codex passes (with robust preflight checks for Codex availability, authentication, and model usability), providing an additional layer of security-focused analysis before persisting the review result.

review/sections · high confidence

New scrape skill for read-only web data extraction via Aside browser

A new \scrape\ skill has been added to the gstack suite, enabling the system to pull structured data from web pages using the user's existing Aside browser sessions. The skill is strictly read-only; it refuses any intent that implies mutating actions (such as form submissions or logins) and instead directs users to the \/qa\ flow for such tasks. It supports both structured extraction (via DOM selectors) and fuzzy queries (via Aside's built-in agent), returning results as a single JSON document on stdout. The skill includes robust safety measures, such as treating all page content as untrusted input, handling sign-in walls by prompting the user to log in manually, and ensuring browser tabs are closed automatically to leave the browser state unchanged.

scrape · high confidence

New scripts for analytics, browser packaging, and build automation

This change introduces a suite of new scripts to support product capabilities and infrastructure. It adds an analytics CLI (\scripts/analytics.ts\) to view skill usage statistics and safety hook events from local JSONL data. It includes a browser launcher (\scripts/app/gstack-browser\) and a macOS build script (\scripts/build-app.sh\) to package the GStack Browser application with bundled Chromium and extensions. Additionally, it provides build automation for native components (\scripts/build-cso.sh\, \scripts/build-cso-windows.ps1\), a general build script (\scripts/build.sh\), and CI/testing helpers like \scripts/compare-pr-version.ts\ and \scripts/capture-baseline.ts\. The \scripts/archetypes.ts\ file defines psychographic archetypes for future planning features, and \scripts/brain-cache-spec.ts\ establishes the cache layer specification for brain-aware planning skills.

scripts · high confidence

New ship workflow sections for adversarial review, Apple release, and changelog generation

Added new documentation and templates for the ship workflow: adversarial review (Claude and Codex), Apple App Store release, and changelog generation. The adversarial review section introduces a preflight check for Codex availability (handling cases like disabled, not installed, under Codex, not authed, broken install, model unusable, or ready) and dispatches a Claude adversarial subagent for every diff, with optional Codex adversarial and structured reviews for large diffs. The Apple release section provides a comprehensive guide for releasing to the App Store/TestFlight using fastlane, including authorization, preflight, store assets, archive/upload, and storefront completion, with specific handling for credentials and idempotency. The changelog section defines a workflow for auto-generating CHANGELOG entries by enumerating commits, reading diffs, grouping by theme, and writing concise user-facing bullet points.

ship/sections · high confidence

New skill to import browser cookies for authenticated headless browsing

Added a new \setup-browser-cookies\ skill that allows users to import cookies from their real Chromium browser into the headless browse session. This enables QA testing of authenticated pages by providing an interactive picker to select cookie domains and accounts. The skill handles browser setup verification, safe cookie extraction with explicit user consent for options like clearing storage, and reports import status honestly without exposing sensitive data.

setup-browser-cookies · high confidence

New structured manifest and 8-pass review sections for DevEx reviews

Added a new \manifest.json\ registry entry for the \plan-devex-review\ skill and introduced \review-sections.md\ (generated from \review-sections.md.tmpl\) to define the eight required DevEx review passes. This change provides the structural definition and detailed evaluation criteria for passes covering Getting Started, API/CLI/SDK Design, Error Messages, Documentation, Upgrade Paths, Developer Environment, Community, and the final Review Report, ensuring these checks are systematically applied after Step 0 investigation.

plan-devex-review/sections · high confidence

New sync-gbrain skill to keep the agent's code search index current

A new \sync-gbrain\ skill has been added to allow users to manually refresh the gbrain code index and update agent search guidance. Triggered by commands like \/sync-gbrain\, \sync gbrain\, or \reindex repo\, this skill wraps the \gstack-gbrain-sync\ orchestrator to probe state, register native code surfaces, and reindex the repository. It supports various modes including incremental syncs, full reindexes, and call-graph generation, while handling local engine pre-flight checks and remote brain configurations to ensure the coding agent has up-to-date context for code search.

sync-gbrain · high confidence

Office-hours skill introduces dual-mode design workflow with security-scanned repo handoff

The office-hours skill now operates in two distinct modes—Startup (for founders) and Builder (for side projects/hackathons)—using a structured diagnostic process to generate a design document. A key behavioral change is the introduction of a dual-write mechanism for the final design doc: it saves a private copy in \~/.gstack for cross-session memory and, when run inside a git repository, writes a team-shareable copy to docs/designs/. This public-facing copy is strictly gated by a pre-write security scan (gstack-redact) to prevent accidental leakage of secrets, with a non-blocking fallback if the scan fails or the repo is read-only. The skill also implements a tiered relationship closing that adapts its tone based on the user's session history.

office-hours · high confidence

Post-deploy canary monitoring skill

A new 'canary' skill has been added to the gstack suite to provide post-deploy visual monitoring. It watches the live application for console errors, performance regressions, and page failures by taking periodic screenshots and comparing them against pre-deploy baselines. The skill supports user-invocable commands like \/canary \<url\>\ with options for custom durations, baseline capture, and specific page monitoring. It automatically discovers pages, captures baselines, and alerts on anomalies such as new console errors, load time regressions, or broken links, ensuring reliability between deployment and verification.

canary · high confidence

Review skill restructured with Fix-First workflow and expanded checklist

The review skill has been significantly expanded and restructured. The main SKILL.md now includes a comprehensive AskUserQuestion format with ELI10 explanations, completeness scoring, and support for split questions. The workflow now implements a 'Fix-First Review' approach where mechanical fixes are applied automatically and ambiguous issues are batched into single user questions. The checklist has been updated with new critical categories including Race Conditions & Concurrency, Shell Injection (Python-specific), and Enum & Value Completeness. Additional features include Greptile comment triage integration, workspace-aware queue status checking, and slop scanning for AI code quality issues. The skill now supports parallel specialist reviewers for categories like Test Gaps and Dead Code.

review · high confidence

iOS QA daemon introduces secure, capability-tiered remote access via Tailscale

The iOS QA daemon now supports secure remote access over Tailscale, gated by a capability-tiered allowlist and session tokens. An owner-granted CLI (\gstack-ios-qa-mint\) manages the allowlist, while the daemon's \/auth/mint\ endpoint issues short-lived session tokens to remote agents whose identities are verified via Tailscale's WhoIs. All remote interactions are proxied through a CoreDevice tunnel to the iOS device, with strict rate limiting, audit logging, and attempt tracking to ensure security and observability.

ios-qa/daemon · high confidence

Behavioural changes

Autoplan review phases now use dual-voice reviews with Codex outside checks

The autoplan review phases (CEO, Design, DX, and Engineering) have been updated to always run a dual-voice review process. In addition to the native Claude subagent review, each phase now invokes an outside review via Codex (using the \gstack-codex-probe\ script and \outside-review-result.ts\). The phases generate consensus tables comparing Claude and Codex findings, with specific degradation and error-handling policies (e.g., falling back to single-voice reviews if one provider is unavailable). This change applies to the \ceo-phase\, \design-phase\, \dx-phase\, and \eng-phase\ section definitions and their corresponding templates.

autoplan/sections · high confidence

CSO helper introduces strict resource limits and hardened lease management for reproduction groups

The CSO helper now enforces aggregate CPU, memory, PID, and writable storage limits per reproduction group and role (e.g., anchor, app, verifier) to prevent resource exhaustion. It also adds a new admission system that manages reproduction lease slots with atomic, race-resistant claim acquisition and release, ensuring that only one helper can own a slot at a time and that stale or crashed leases are safely reclaimed. Additionally, the helper validates input files and paths with strict bounds (size, ownership, permissions, no symlinks) and uses bounded, stable reads to prevent snapshot races during file operations.

lib/cso · high confidence

CSO security audit skill upgraded to v3 with evidence-based reporting and comprehensive repair workflows

The CSO skill has been updated to version 3.0.0, introducing a stricter evidence-based audit workflow that separates supported findings from labeled hypotheses. This update adds comprehensive mode capabilities, allowing the tool to generate isolated reproduction steps and up to three repair candidates when a qualified runtime catalog profile is available. The skill now enforces a private, isolated execution environment using a trusted launcher, ensuring that source inspection and scanner execution do not bypass redaction or affect the host. Reporting artifacts have been standardized to include stable finding IDs, explicit assurance labels (such as \self\_reported\ vs. \runtime\_tested\), and detailed coverage records, while legacy v2 report imports are supported in read-only mode.

cso · high confidence

Deterministic generation of app-owned bridge accessors

The iOS QA scripts now include a new tool (with a TypeScript fallback) that scans Swift source files for @Observable classes marked with @Snapshotable and automatically generates StateAccessor.swift files. This process is fully deterministic: it uses a composite cache key that includes the Swift version, tool revision, platform, and source content hash to ensure consistent output across different environments and builds, while strictly validating declarations to fail fast on invalid or duplicate keys.

ios-qa/scripts · high confidence

Improved interactive element detection in browse snapshots

The browse snapshot command now automatically includes cursor-interactive scanning when the interactive flag is used, ensuring that dropdowns, popovers, and other non-ARIA interactive elements are captured. This change also adds specific detection for floating containers (like React portals) and removes the previous behavior that skipped elements with ARIA roles if they were missed by the accessibility tree, resulting in more complete and accurate snapshots for modern web applications.

browse · high confidence

Plan-eng-review skill receives major overhaul with scope gate and structured decision briefs

The plan-eng-review skill has been significantly updated to enforce a strict 'Scope gate' that auto-selects or prompts for the review target (branch diff, plan doc, or specific path) before any other work begins. The skill now includes a comprehensive 'AskUserQuestion Format' for structured decision briefs, complete with D-numbering, ELI10 explanations, completeness scores, and pros/cons, along with specific fallback behaviors for different session types (spawned, headless, interactive). Additionally, the skill's preamble has been enhanced to handle degraded modes, context recovery, and telemetry more robustly, and it now includes a 'Design Doc Check' and 'Scope Challenge' step to ensure thorough review preparation.

plan-eng-review · high confidence

Preamble instructions refactored into modular, context-aware generators

The system's preamble instructions have been restructured from monolithic inline scripts into a set of dedicated TypeScript generators (e.g., \generate-ask-user-format.ts\, \generate-completeness-section.ts\, \generate-preamble-bash.ts\). This change introduces context-aware rendering: the generated instructions now adapt to the specific skill (e.g., \plan-eng-review\), host environment (\gbrain\, \hermes\), and configuration (\explainLevel\). Key behavioral updates include a refined \AskUserQuestion\ protocol with specific handling for spawned vs. interactive sessions, a new \Completeness\ scoring principle for decision briefs, and a shift of heavy runtime bootstrapping (session bookkeeping, artifacts sync) to the \gstack-skill-start\ binary, leaving the preamble with only the interpretation rules the model needs. This results in a more modular, token-efficient, and robust instruction set that dynamically adjusts based on the current execution context.

scripts/resolvers/preamble · high confidence

Renamed /checkpoint resume to /context-restore

The context-restore skill has been renamed from /checkpoint resume to /context-restore. This change was made because the /checkpoint command is now treated as a native rewind alias in current environments, which caused conflicts. The new skill loads the most recent saved state (preferring the current branch, falling back across branches) so users can pick up where they left off, even across Conductor workspace handoffs. It is designed to be used in conjunction with /context-save.

context-restore · high confidence

Renamed /checkpoint to /context-save to resolve command conflicts

The context-save skill has been renamed from /checkpoint to /context-save. This change resolves a conflict where the native /checkpoint command in Claude Code was shadowing the skill's functionality. The skill now triggers on 'save progress', 'save state', 'save my work', and 'context save', and it pairs with /context-restore to resume work in future sessions.

context-save · high confidence

Retro skill v2.0.0: new report structure, gbrain context, and robust date handling

The /retro skill has been upgraded to version 2.0.0, introducing a significantly more detailed retrospective report format that includes a tweetable summary, shipping velocity analysis, code quality signals, test health tracking, plan completion metrics, and a personal deep-dive for the user alongside a team breakdown. The skill now leverages gbrain to load prior retrospectives, recent timeline events, and learnings for trend analysis, while fixing date window calculations to use local midnight alignment instead of wall-clock time to prevent stale data issues. It also adds support for cross-project retrospectives, compare modes, and integrates with external context files like Greptile history and TODO lists to provide a more comprehensive view of engineering performance.

retro · high confidence

Ship workflow now detects base branch dynamically and supports Apple distribution

The ship skill has been refactored to replace the hardcoded 'main' branch assumption with dynamic base branch detection (Step 1), allowing it to operate correctly on any repository default or feature branch. Additionally, a new Apple target detection step (Step 0.9) has been added to handle App Store and TestFlight distribution workflows for Swift/Xcode projects, ensuring proper release pipelines are followed for iOS/macOS artifacts.

ship · high confidence

Telemetry schema expansion and security lockdown

The system now collects detailed usage and security telemetry, including new columns for tracking prompt injection attempts (attack domains, payload hashes, and verdicts) alongside standard skill run data. To protect this data, read access for anonymous users has been revoked via tightened Row Level Security policies, while write access is maintained through scoped, column-limited policies to support both legacy PostgREST clients and new edge functions. A cache table for community pulse aggregation has also been added to prevent performance degradation from repeated queries.

supabase/migrations · high confidence

Fixes

Fixes command parsing and decision envelope in the /careful safety hook

The /careful bash hook now uses a shared JSON extractor (hook-extract.sh) to correctly parse command arguments containing escaped quotes, preventing destructive commands from being silently truncated and allowed. It also ensures safety decisions are nested under hookSpecificOutput so Claude Code does not ignore them, and fails closed when the input cannot be parsed or helpers are missing.

careful/bin · high confidence

Fixes console window flashing on Windows when launching browsers

This patch for playwright-core (v1.62.1) prevents unwanted console windows from appearing on Windows when launching browser processes. It achieves this by setting the \windowsHide\ option to \true\ in both the initial process spawn and the subsequent \taskkill\ command used for cleanup, ensuring a cleaner user experience on Windows systems.

patches · high confidence

Introduce autoplan publication hook for reliable phase validation

Added a new executable shell script and TypeScript implementation in the autoplan/bin directory to serve as a native read barrier for Autoplan phases. This hook ensures that broken installations return a deny decision rather than failing silently, and it validates that phase artifacts (methodologies, snapshots, and implementations) match their expected identities and content hashes before allowing consumption, thereby making checks more reliable and everyday validation faster.

autoplan/bin · high confidence

Migration scripts for gstack upgrades v0.15.2.0 through v1.86.0.0

This update adds a suite of 16 migration scripts to the \gstack-upgrade/migrations\ directory, covering versions from v0.15.2.0 to v1.86.0.0. These scripts automate the transition for existing users by fixing skill directory structures for unprefixed discovery (v0.15.2.0), merging per-project resource logs into the builder profile (v0.16.2.0), introducing a one-time prompt for the new V1 writing style (v1.0.0.0), and removing the stale \/checkpoint\ skill to prevent shadowing Claude Code's native \/rewind\ alias (v1.1.3.0). Further migrations wire existing brain-sync repositories as gbrain federated sources (v1.17.0.0), rename \gstack-brain-\\ to \gstack-artifacts-\\ including repo and config updates (v1.27.0.0), and notify users about split-engine gbrain with local PGLite code search (v1.37.0.0). Additional scripts patch allowlists and privacy maps for new artifact patterns (v1.38.1.0, v1.40.0.0), register Conductor AskUserQuestion hooks (v1.58.0.0), repair poisoned Chrome-for-Testing bundles on macOS (v1.65.0.0), restore tracked files dirtied by legacy in-place gbrain renders (v1.67.0.0), migrate feature-discovery acknowledgements to \GSTACK\_HOME\ (v1.78.0.0), and update \/claude\ to \/claude-code\ on non-Claude hosts (v1.86.0.0).

gstack-upgrade/migrations · high confidence

New telemetry and update-check edge functions with security and reliability improvements

Added three new Supabase Edge Functions: \telemetry-ingest\ for validating and storing telemetry events, \community-pulse\ for aggregating community statistics, and \update-check\ for logging install pings. The \telemetry-ingest\ function improves security by using the Supabase anon key with Row Level Security (RLS) instead of the over-privileged service role key, and enforces payload size and batch limits. The \community-pulse\ function fixes a reliability issue where database errors previously returned fake zero counts; it now explicitly checks for query errors and serves stale cache or a 503 error instead. It also introduces k-anonymity filtering for security event data to prevent re-identification of individual users' attack logs. The \update-check\ function ensures the current version is always returned to the client, even if the background logging to the database fails.

supabase/functions · high confidence

Supabase RLS verification and configuration scripts

Added \config.sh\ to store public Supabase connection details and \verify-rls.sh\ to validate Row Level Security policies. The verification script ensures that read and update operations are denied via the anonymous key while insert operations remain allowed for backward compatibility, confirming the security posture of the deployed RLS rules.

supabase · high confidence

Test coverage

Added comprehensive test coverage for the design subsystem; Added comprehensive test coverage for the make-pdf module; Added end-to-end gate tests for make-pdf output quality and format fidelity; Added provider adapter tests for Claude, Gemini, and GPT; Added test coverage for browse module security, activity, and browser management; Added test fixtures for autoplan, repository carving, and payment processing; Expanded test coverage for core infrastructure and evaluation systems; Expanded test fixtures for browser automation and security validation; New test fixtures for ARM benchmark and plan-ceo-review auto-decision flows; New test harness infrastructure for E2E and evaluation reliability.

Dependencies

Major dependency overhaul and new build/test tooling

The project has upgraded core dependencies, notably bumping Playwright to v1.62.1 and diff to v9.0.0, while adding new libraries for PDF generation (html-to-docx), terminal emulation (xterm), and AI integration (@huggingface/transformers, @anthropic-ai/sdk). Security is addressed by overriding vulnerable packages like sharp, adm-zip, and ip-address. The build system has been significantly expanded with new scripts for diagram rendering, PDF generation, and a comprehensive, sharded test suite (gate, periodic, evals) that supports multiple platforms and CI profiles. Additionally, new Swift and Node.js fixtures have been added to support iOS QA and ARM benchmark testing.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Score

  • CAI 41 → 52 (+11.4)
  • Rubric changed (rubric-2026.08.15 → rubric-2026.09.15) — scores are not directly comparable.

Lenses

  • Code Health 70 → 61 (-8.8)
  • Architecture 88 (new)
  • Maturity 73 → 84 (+11.2)
  • Readiness 24 → 70 (+45.5)
  • Security 46 → 77 (+31.3)
  • Accessibility 34 (new)

Resolved (121)

  • (anonymous) (cognitive 21) (extension/sidepanel-terminal.js)
  • (anonymous) (cognitive 34) (extension/background.js)
  • (anonymous) (cyclomatic 28) (extension/background.js)
  • Coverage not measured — test suite did not build
  • Dimension evaluation failed
  • FileTooLong: extension/background.js (extension/background.js)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (lib/diagram-render/bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • High CVE: [GHSA redacted] (bun.lock)
  • …and 101 more

New (1190)

  • (anonymous) (cognitive 20) (extension/sidepanel.js)
  • (anonymous) (cognitive 44) (extension/background.js)
  • (anonymous) (cognitive 51) (lib/dom-dump.js)
  • (anonymous) (cyclomatic 35) (extension/background.js)
  • (anonymous) (cyclomatic 47) (lib/dom-dump.js)
  • (anonymous)::buildSelector (cognitive 21) (extension/inspector.js)
  • (anonymous)::captureBasicData (cognitive 47) (extension/inspector.js)
  • (anonymous)::captureBasicData (cyclomatic 30) (extension/inspector.js)
  • (anonymous)::forceRestart (cognitive 29) (extension/sidepanel-terminal.js)
  • (anonymous)::forceRestart (cyclomatic 25) (extension/sidepanel-terminal.js)
  • (anonymous)::onClick (cognitive 16) (extension/inspector.js)
  • BrowseClient.command (cognitive 16) (browse/src/browse-client.ts)
  • BrowseClient.command (cognitive 16) (browser-skills/hackernews-frontpage/_lib/browse-client.ts)
  • BrowserManager.close (cognitive 30) (browse/src/browser-manager.ts)
  • BrowserManager.close (cyclomatic 21) (browse/src/browser-manager.ts)
  • BrowserManager.launchHeaded (cognitive 46) (browse/src/browser-manager.ts)
  • BrowserManager.launchHeaded (cyclomatic 31) (browse/src/browser-manager.ts)
  • BrowserManager.promoteToHeaded (cognitive 35) (browse/src/browser-manager.ts)
  • BrowserManager.promoteToHeaded (cyclomatic 28) (browse/src/browser-manager.ts)
  • BrowserManager.restoreState (cognitive 37) (browse/src/browser-manager.ts)
  • …and 1170 more

Changes since last survey

  • 46 commits — 18 feature/other, 28 fixes

By area

  • (root) — 21 commits
  • browse/test — 11 commits
  • test/fixtures — 5 commits
  • browse/src — 2 commits
  • scripts/resolvers — 2 commits
  • ship/sections — 2 commits
  • lib/cso — 1 commit
  • make-pdf/test — 1 commit
  • test/helpers — 1 commit

Notable commits

  • fix: v1.60.2.0 fix: free-suite drift on dev machines (eval-list cwd, gemini regex, observability floor) (#2470)
  • fix: v1.61.0.0 fix wave: guards failing open / silent failures (9 fixes, 4 community PRs absorbed) (#2472)
  • fix: v1.63.0.0 feat: GStack 2 fork port wave — egress receipts, context-bill, sharded gate, /health fix (#2541)
  • fix: v1.64.0.0 fix wave: full tracker audit — 90 fixes, 52 issues closed, ~50 community PRs absorbed (#2571)
  • fix: v1.64.1.0 v1.64.1.0: the code-smell fix wave — every pipeline guard now provably fires (net −24,943 lines) (#2572)
  • fix: v1.65.0.0 feat: fork port wave 2 — feature fixes, session persistence, Apple releases, supply-chain CI (#2577)
  • fix: v1.67.0.0 fix: the tracker wave — XProtect self-heal, complete installs, brain-sync integrity, 31 community PRs credited (#2604)
  • fix: v1.67.1.0 fix: external-contributor security sweep — 6 findings hardened, regression-pinned (#2605)
  • fix: v1.68.0.0 fix: next tracker wave — 16 verified fixes in, 90 stale PRs and 21 issues closed with receipts (#2632)
  • fix: v1.68.1.0 fix: phantom AskUserQuestion hooks — canonical-only registration + self-healing settings.json (#2631)
  • fix: v1.68.2.0 fix: tunnel revoke exists and revokes everything — setup keys included, verified live (#2646)
  • fix: v1.68.3.0 fix(pairing): re-pair to narrow revokes the old grant on the spot (#2665)
  • fix: v1.69.0.0 fix: the silent-failure wave — 6 fixes, 5 community PRs absorbed, tracker closed with receipts (#2666)
  • fix: v1.70.1.0 fix: ship names the /document-release subagent at every decision point (tripwire + gate E2E) (#2700)
  • fix: v1.76.0.0 fix: ship doc-sync survives Conductor — spawned subagent sessions reachable (#2733) (#2741)
  • fix: v1.78.0.0 fix: the two-red-lanes wave — AUQ collapse rooted, OSV green from 105, 18 community PRs absorbed, upgrade path can't eat installs (#2752)
  • fix: v1.79.0.0 fix: ship subagent dispatches can no longer strand the run (#497/#2440 class) (#2772)
  • fix: v1.80.0.0 fix: setup survives a failed Chromium install, hooks share one state root, gstack never clobbers a skill it did not create (#2802)
  • fix: v1.84.1.0 fix: default Codex and Claude to frontier models (#2835)
  • fix: v1.87.1.0 fix: update vulnerable sharp and adm-zip overrides (#2873)
  • …and 26 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

garrytan/gstack was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 28 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 01593aa67c94780528e8f5121e47362502410ced — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-d00c643c3f66.