virgiliojr94/book-to-skill
73.0
Strong · 18 September 2026
3.7k
lines of production code
Python
primary language
1
measurement over time
What this system is
This system is a toolkit for extracting content from books and documents into structured 'skills' for AI agents, supporting multiple formats like PDF, EPUB, and DOCX with robust security and encoding handling. It provides utilities to measure context efficiency, validate agent compatibility, and perform deterministic offline evaluations of generated skills. The project includes comprehensive testing and modular tooling to ensure reliable extraction, scanning, and scoring workflows.
Features
Initial release of book\_to\_skill extraction tool
This change introduces the book\_to\_skill package, providing a CLI tool to extract text and metadata from books in various formats (PDF, EPUB, DOCX, HTML, etc.). Key features include a new PDF inspection layer using pdf-inspector for trusted native text extraction, comprehensive sanitization to strip invisible Unicode code points and bidirectional formatting controls for security, and robust chapter detection supporting multiple languages and scripts (including CJK, Thai, Hindi, Bengali, and others). The tool also handles concurrent execution safely via per-process work directories and manages dependencies with automatic installation prompts.
_book\_to\skill · high confidence
New deterministic evaluation tools for manifest generation, paper processing, and trajectory scoring
The \tools/evals\ directory now includes four new Python scripts that enable offline evaluation workflows. \manifest.py\ generates deterministic, secret-free run manifests by validating configuration and hashing source files. \paper\_flat.py\ builds experiment-only skill packs by splitting source text into chunks based on chapter headings. \replay.py\ allows deterministic replay of synthetic trajectory fixtures. \score.py\ provides pure scoring and accounting for these trajectories, ensuring that usage counts are never estimated from missing data and that scoring does not crash on incomplete recorded data.
tools/evals · high confidence
New tools for measuring context efficiency, scanning generated skills, and validating agent compatibility
This change introduces three new command-line utilities in the \tools/\ directory. \discovery\_tax.py\ quantifies the token cost of answering questions using three strategies (context-dump, discovery-loop, and book-to-skill) by reusing the extractor's chapter and table-of-contents detection logic. \scan\_generated\_skill.py\ performs an advisory security scan on generated skills to detect prompt injection patterns, unsafe authority overrides, and exfiltration attempts, while also ensuring the scan scope is bounded to prevent symlink traversal. \validate\_skill.py\ audits \SKILL.md\ files against specific host rules (lenses) for Claude Code, GitHub Copilot CLI, Amp, Hermes Agent, and OpenClaw, checking for valid frontmatter, reserved names, and allowed tool grants, with support for UTF-8 BOM files and configurable encoding.
tools · high confidence
Behavioural changes
Improved search engine visibility and social sharing previews
The documentation site now generates dynamic, SEO-friendly page titles that include descriptive keywords rather than just short navigation labels, ensuring better search engine results. Additionally, all pages now share a single, consistent social media preview card (for platforms like Twitter and Facebook) that displays the correct page title and description, replacing the previous behavior which often showed generic or misleading previews.
overrides · high confidence
Script refactored into modular extractor package with UTF-8 output fix
The extract.py script has been significantly refactored from a monolithic file into a structured extractor package, simplifying the entry point to a backward-compatible wrapper. This change also ensures that extracted text and CLI output are forced to UTF-8 encoding, preventing UnicodeEncodeError issues on Windows systems with legacy code pages.
scripts · high confidence
Fixes
Introduce dedicated parsers for DOCX, EPUB, HTML, PDF, RTF, and text with security and accuracy fixes
The \book\_to\_skill/parsers\ package now provides dedicated extraction modules for DOCX, EPUB, HTML, PDF, RTF, and plain text. DOCX parsing includes self-defending XML safety checks to prevent XXE attacks and reconstructs tables as tab-joined rows. EPUB extraction now respects the spine reading order and reports unextracted images. HTML parsing uses trafilatura for robust boilerplate removal, with a stdlib fallback that correctly emits block boundaries at end tags and decodes entities once. PDF extraction cleans pdftotext output by dehyphenating words and stripping running headers/footers only at page edges, while also aborting early on scanned PDFs. RTF parsing decodes unicode escapes and skips metadata destination groups. Text file reading now correctly decodes UTF-16 and UTF-32 files based on their Byte Order Mark (BOM).
_book\_to\skill/parsers · high confidence
Test coverage
Added comprehensive test coverage for extraction resilience, format handling, and host integration; Added tests for evaluation manifest stability, paper-flat packing, trajectory replay, and scoring robustness.
Dependencies
Introduce pyproject.toml with dependency and tooling configuration
The project now uses a pyproject.toml file to define its build system (hatchling), project metadata, and optional dependencies for various formats (HTML, EPUB, PDF, DOCX, RTF, technical). It also configures ruff for linting and pytest for testing, ensuring local development matches CI requirements.
(dependencies) · high confidence
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
How this codebase got here
Baseline
- First survey — no prior run to compare against. CAI 73.
Lenses
- Code Health 83
- Architecture 99
- Maturity 63
- Readiness 78
- Security 84
Changes since last survey
- 189 commits — 124 feature/other, 65 fixes
By area
- (root) — 98 commits
- book_to_skill/utils.py — 30 commits
- .github/workflows — 12 commits
- book_to_skill/parsers — 11 commits
- scripts/extractor — 9 commits
- (repo) — 5 commits
- book_to_skill/pdf_inspector_integration.py — 3 commits
- book_to_skill/sanitize.py — 3 commits
- .github/dependabot.yml — 2 commits
- book_to_skill/dependencies.py — 2 commits
- scripts/extract.py — 2 commits
- tests/evals — 2 commits
- tests/test_output_dir_security.py — 2 commits
- .github/FUNDING.yml — 1 commit
- .github/PULL_REQUEST_TEMPLATE.md — 1 commit
- docs/assets — 1 commit
- docs/index.md — 1 commit
- evals/fixtures — 1 commit
- tests/test_scan_coverage.py — 1 commit
- tests/test_scan_generated_skill.py — 1 commit
Notable commits
- fix: Potential fix for pull request finding
- fix: fix(changelog): stop duplicating the PR number in generated entries (#131)
- fix: fix(cli): support documented help flags (#97)
- fix: fix(config): give each run its own workdir so concurrent extractions cannot clobber each other (#184)
- fix: fix(deps): diagnose a module installed in an isolated environment (#151)
- fix: fix(deps): honor alternative parser availability (#208)
- fix: fix(discovery_tax): reuse the extractor's multilingual ToC detection (#138)
- fix: fix(docs): reserve the banner's space on the site build
- fix: fix(docx): self-defend the DOCX parsers against XXE regardless of caller (#99)
- fix: fix(epub): read content in spine order in the stdlib extractor (#64)
- fix: fix(epub): report unextracted source images (#134)
- fix: fix(evals): stop scoring crashing on, and inventing counts from, recorded data (#225)
- fix: fix(extract): count supplementary-plane CJK in the token estimate (#136)
- fix: fix(extract): detect a table of contents in any source, not just the first (#114)
- fix: fix(extract): report which method produced the chapter count (#150)
- fix: fix(extract): write metadata.json as UTF-8 (#105)
- fix: fix(extractor): count numbered headings as chapters when they carry a chapter's weight (#149)
- fix: fix(extractor): defer annotation evaluation so import works on Python 3.9 (#34)
- fix: fix(extractor): detect Markdown-prefixed chapter headings (#91) (#92)
- fix: fix(extractor): detect full-width Arabic digits in CJK chapter headings (#46)
- …and 169 more
Written by watchdog.canine.dev from the codebase's own history, inside the signed delivery this page is composed from.
Survey your own repository
virgiliojr94/book-to-skill was measured the same way every project in this corpus was: the same rubric, at a pinned commit, with the result published in full. Point a surveyor at a repository you know and see whether you agree with it.
About this page
- The score is its most recent published measurement, taken on 18 September 2026 at a pinned commit. It is not a live figure and does not change until the project is measured again.
- Measured at commit 526f362552562d88c1a8bbf8012d2cee93f831d5 — the exact code this score is about.
- Scored under rubric-2026.09.15 — the same rubric and the same method as every other entry in this index.
- Measured by watchdog.canine.dev using codehealth-analyzer preprod-5d04157a340d.