Skip to content
CAI
Software that uses CAICheck a score

virgiliojr94/book-to-skill

73.0

Strong · 18 September 2026

3.7k

lines of production code

Python

primary language

1

measurement over time

CAI band scale
CAI lens gauges

What this system is

This system is a toolkit for extracting content from books and documents into structured 'skills' for AI agents, supporting multiple formats like PDF, EPUB, and DOCX with robust security and encoding handling. It provides utilities to measure context efficiency, validate agent compatibility, and perform deterministic offline evaluations of generated skills. The project includes comprehensive testing and modular tooling to ensure reliable extraction, scanning, and scoring workflows.

Features

Initial release of book\_to\_skill extraction tool

This change introduces the book\_to\_skill package, providing a CLI tool to extract text and metadata from books in various formats (PDF, EPUB, DOCX, HTML, etc.). Key features include a new PDF inspection layer using pdf-inspector for trusted native text extraction, comprehensive sanitization to strip invisible Unicode code points and bidirectional formatting controls for security, and robust chapter detection supporting multiple languages and scripts (including CJK, Thai, Hindi, Bengali, and others). The tool also handles concurrent execution safely via per-process work directories and manages dependencies with automatic installation prompts.

_book\_to\skill · high confidence

New deterministic evaluation tools for manifest generation, paper processing, and trajectory scoring

The \tools/evals\ directory now includes four new Python scripts that enable offline evaluation workflows. \manifest.py\ generates deterministic, secret-free run manifests by validating configuration and hashing source files. \paper\_flat.py\ builds experiment-only skill packs by splitting source text into chunks based on chapter headings. \replay.py\ allows deterministic replay of synthetic trajectory fixtures. \score.py\ provides pure scoring and accounting for these trajectories, ensuring that usage counts are never estimated from missing data and that scoring does not crash on incomplete recorded data.

tools/evals · high confidence

New tools for measuring context efficiency, scanning generated skills, and validating agent compatibility

This change introduces three new command-line utilities in the \tools/\ directory. \discovery\_tax.py\ quantifies the token cost of answering questions using three strategies (context-dump, discovery-loop, and book-to-skill) by reusing the extractor's chapter and table-of-contents detection logic. \scan\_generated\_skill.py\ performs an advisory security scan on generated skills to detect prompt injection patterns, unsafe authority overrides, and exfiltration attempts, while also ensuring the scan scope is bounded to prevent symlink traversal. \validate\_skill.py\ audits \SKILL.md\ files against specific host rules (lenses) for Claude Code, GitHub Copilot CLI, Amp, Hermes Agent, and OpenClaw, checking for valid frontmatter, reserved names, and allowed tool grants, with support for UTF-8 BOM files and configurable encoding.

tools · high confidence

Behavioural changes

Improved search engine visibility and social sharing previews

The documentation site now generates dynamic, SEO-friendly page titles that include descriptive keywords rather than just short navigation labels, ensuring better search engine results. Additionally, all pages now share a single, consistent social media preview card (for platforms like Twitter and Facebook) that displays the correct page title and description, replacing the previous behavior which often showed generic or misleading previews.

overrides · high confidence

Script refactored into modular extractor package with UTF-8 output fix

The extract.py script has been significantly refactored from a monolithic file into a structured extractor package, simplifying the entry point to a backward-compatible wrapper. This change also ensures that extracted text and CLI output are forced to UTF-8 encoding, preventing UnicodeEncodeError issues on Windows systems with legacy code pages.

scripts · high confidence

Fixes

Introduce dedicated parsers for DOCX, EPUB, HTML, PDF, RTF, and text with security and accuracy fixes

The \book\_to\_skill/parsers\ package now provides dedicated extraction modules for DOCX, EPUB, HTML, PDF, RTF, and plain text. DOCX parsing includes self-defending XML safety checks to prevent XXE attacks and reconstructs tables as tab-joined rows. EPUB extraction now respects the spine reading order and reports unextracted images. HTML parsing uses trafilatura for robust boilerplate removal, with a stdlib fallback that correctly emits block boundaries at end tags and decodes entities once. PDF extraction cleans pdftotext output by dehyphenating words and stripping running headers/footers only at page edges, while also aborting early on scanned PDFs. RTF parsing decodes unicode escapes and skips metadata destination groups. Text file reading now correctly decodes UTF-16 and UTF-32 files based on their Byte Order Mark (BOM).

_book\_to\skill/parsers · high confidence

Test coverage

Added comprehensive test coverage for extraction resilience, format handling, and host integration; Added tests for evaluation manifest stability, paper-flat packing, trajectory replay, and scoring robustness.

Dependencies

Introduce pyproject.toml with dependency and tooling configuration

The project now uses a pyproject.toml file to define its build system (hatchling), project metadata, and optional dependencies for various formats (HTML, EPUB, PDF, DOCX, RTF, technical). It also configures ruff for linting and pytest for testing, ensuring local development matches CI requirements.

(dependencies) · high confidence

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

How this codebase got here

Baseline

  • First survey — no prior run to compare against. CAI 73.

Lenses

  • Code Health 83
  • Architecture 99
  • Maturity 63
  • Readiness 78
  • Security 84

Changes since last survey

  • 189 commits — 124 feature/other, 65 fixes

By area

  • (root) — 98 commits
  • book_to_skill/utils.py — 30 commits
  • .github/workflows — 12 commits
  • book_to_skill/parsers — 11 commits
  • scripts/extractor — 9 commits
  • (repo) — 5 commits
  • book_to_skill/pdf_inspector_integration.py — 3 commits
  • book_to_skill/sanitize.py — 3 commits
  • .github/dependabot.yml — 2 commits
  • book_to_skill/dependencies.py — 2 commits
  • scripts/extract.py — 2 commits
  • tests/evals — 2 commits
  • tests/test_output_dir_security.py — 2 commits
  • .github/FUNDING.yml — 1 commit
  • .github/PULL_REQUEST_TEMPLATE.md — 1 commit
  • docs/assets — 1 commit
  • docs/index.md — 1 commit
  • evals/fixtures — 1 commit
  • tests/test_scan_coverage.py — 1 commit
  • tests/test_scan_generated_skill.py — 1 commit

Notable commits

  • fix: Potential fix for pull request finding
  • fix: fix(changelog): stop duplicating the PR number in generated entries (#131)
  • fix: fix(cli): support documented help flags (#97)
  • fix: fix(config): give each run its own workdir so concurrent extractions cannot clobber each other (#184)
  • fix: fix(deps): diagnose a module installed in an isolated environment (#151)
  • fix: fix(deps): honor alternative parser availability (#208)
  • fix: fix(discovery_tax): reuse the extractor's multilingual ToC detection (#138)
  • fix: fix(docs): reserve the banner's space on the site build
  • fix: fix(docx): self-defend the DOCX parsers against XXE regardless of caller (#99)
  • fix: fix(epub): read content in spine order in the stdlib extractor (#64)
  • fix: fix(epub): report unextracted source images (#134)
  • fix: fix(evals): stop scoring crashing on, and inventing counts from, recorded data (#225)
  • fix: fix(extract): count supplementary-plane CJK in the token estimate (#136)
  • fix: fix(extract): detect a table of contents in any source, not just the first (#114)
  • fix: fix(extract): report which method produced the chapter count (#150)
  • fix: fix(extract): write metadata.json as UTF-8 (#105)
  • fix: fix(extractor): count numbered headings as chapters when they carry a chapter's weight (#149)
  • fix: fix(extractor): defer annotation evaluation so import works on Python 3.9 (#34)
  • fix: fix(extractor): detect Markdown-prefixed chapter headings (#91) (#92)
  • fix: fix(extractor): detect full-width Arabic digits in CJK chapter headings (#46)
  • …and 169 more

Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.

Survey your own repository

virgiliojr94/book-to-skill was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.

About this page

  • The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
  • Measured at commit 526f362552562d88c1a8bbf8012d2cee93f831d5 — the exact code this score is about.
  • Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
  • Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.